当前位置: 开发笔记 > 编程语言 > 正文

文字检测与识别资源

作者：love_xiao奇 | 来源：互联网 | 2023-08-16 19:54

综述[2015-PAMI-Overview]TextDetectionandRecognitioninImagery:ASurvey[paper][2014-Front.Comput.

综述

[2015-PAMI-Overview]Text Detection and Recognition in Imagery: A Survey[paper]

[2014-Front.Comput.Sci-Overview]Scene Text Detection and Recognition: Recent Advances and Future Trends[paper]

自然场景文字检测
[2017-arXiv]R2CNN: Rotational Region CNN for Orientation Robust Scene Text Detection[paper]

[2017-CVPR]EAST: An Efficient and Accurate Scene Text Detector [paper]

[2017-arXiv]Cascaded Segmentation-Detection Networks for Word-Level Text Spotting[paper]

[2017-arXiv]Deep Direct Regression for Multi-Oriented Scene Text Detection[paper]

[2017-CVPR]Detecting oriented text in natural images by linking segments [paper]

[2017-CVPR]Deep Matching Prior Network: Toward Tighter Multi-oriented Text Detection[paper]

[2017-arXiv]Arbitrary-Oriented Scene Text Detection via Rotation Proposals [paper]

[2017-AAAI]TextBoxes: A Fast Text Detector with a Single Deep Neural Network[paper][code]

[2016-arXiv]Accurate Text Localization in Natural Image with Cascaded Convolutional TextNetwork [paper]

[2016-arXiv]DeepText : A Unified Framework for Text Proposal Generation and Text Detectionin Natural Images [paper] [data]

[2016-arXiv]TextProposals: a Text-specific Selective Search Algorithm for Word Spotting in the Wild [paper] [code]

[2016-arXiv] SceneText Detection via Holistic, Multi-Channel Prediction [paper]

[2016-CVPR] CannyText Detector: Fast and Robust Scene Text Localization Algorithm [paper]

[2016-CVPR]Synthetic Data for Text Localisation in Natural Images [paper] [data][code]

[2016-ECCV]Detecting Text in Natural Image with Connectionist Text Proposal Network[paper][demo][code]

[2016-TIP]Text-Attentional Convolutional Neural Networks for Scene Text Detection [paper]

[2016-IJDAR]TextCatcher: a method to detect curved and challenging text in natural scenes[paper]

[2016-CVPR]Multi-oriented text detection with fully convolutional networks [paper]

[2015-TPRMI]Real-time Lexicon-free Scene Text Localization and Recognition[paper]

[2015-CVPR]Symmetry-Based Text Line Detection in Natural Scenes[paper][code]

[2015-ICCV]FASText: Efficient unconstrained scene text detector[paper][code]

[2015-D.PhilThesis] Deep Learning for Text Spotting [paper]

[2015 ICDAR]Object Proposals for Text Extraction in the Wild [paper] [code]

[2014-ECCV] Deep Features for Text Spotting [paper] [code] [model] [GitXiv]

[2014-TPAMI] Word Spotting and Recognition with Embedded Attributes [paper] [homepage] [code]

[2014-TPRMI]Robust Text Detection in Natural Scene Images[paper]

[2014-ECCV] Robust Scene Text Detection with Convolution Neural Network Induced MSER Trees [paper]

[2013-ICCV] Photo OCR: Reading Text in Uncontrolled Conditions[paper]

[2012-CVPR]Real-time scene text localization and recognition[paper][code]

[2010-CVPR]Detecting Text in Natural Scenes with Stroke Width Transform [paper] [code]

自然场景文字识别
[2017-AAAI-网络图片]Detection and Recognition of Text Embedded in Online Images via Neural Context Models[paper][project]

[2017-arvix 文档识别] Full-Page TextRecognition : Learning Where to Start and When to Stop[paper]

[2016-AAAI]Reading Scene Text in Deep Convolutional Sequences [paper]

[2016-IJCV]Reading Text in the Wild with Convolutional Neural Networks [paper] [demo] [homepage]

[2016-CVPR]Recursive Recurrent Nets with Attention Modeling for OCR in the Wild [paper]

[2016-CVPR] Robust Scene Text Recognition with Automatic Rectification [paper]

[2016-NIPs] Generative Shape Models: Joint Text Recognition and Segmentation with Very Little Training Data[paper]

[2015-CoRR] AnEnd-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition [paper] [code]

[2015-ICDAR]Automatic Script Identification in the Wild[paper]

[2015-ICLR] Deep structured output learning for unconstrained text recognition [paper]

[2014-NIPS]Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition [paperhomepage] [model]

[2014-TIP] A Unified Framework for Multi-Oriented Text Detection and Recognition [paper]

[2012-ICPR]End-to-End Text Recognition with Convolutional Neural Networks [paper] [code] [SVHN Dataset]

数据集

COCO-Text (ComputerVision Group, Cornell) 2016
63,686images, 173,589 text instances, 3 fine-grained text attributes.
Task:text location and recognition
COCO-Text API
Synthetic Data for Text Localisation in Natural Image (VGG)2016
         800k thousand images
         8 million synthetic word instances
         download
Synthetic Word Dataset (Oxford, VGG) 2014
9million images covering 90k English words
Task:text recognition, segmentation
download
IIIT 5K-Words 2012
5000images from Scene Texts and born-digital (2k training and 3k testing images)
Eachimage is a cropped word image of scene text with case-insensitive labels
Task:text recognition
download
StanfordSynth(Stanford, AI Group) 2012
Smallsingle-character images of 62 characters (0-9, a-z, A-Z)
Task:text recognition
download
MSRA Text Detection 500 Database(MSRA-TD500) 2012
500 natural images(resolutions of the images vary from 1296x864 to 1920x1280)
Chinese,English or mixture of both
Task:text detection
Street View Text (SVT) 2010
350 high resolution images (average size 1260 × 860) (100 images for training and 250 images for testing)
Onlyword level bounding boxes are provided with case-insensitive labels
Task:text location
KAIST Scene_Text Database 2010
3000images of indoor and outdoor scenes containing text
Korean,English (Number), and Mixed (Korean + English + Number)
Task:text location, segmentation and recognition
Chars74k 2009
Over74K images from natural images, as well as a set of synthetically generatedcharacters
Smallsingle-character images of 62 characters (0-9, a-z, A-Z)
Task:text recognition
ICDARBenchmark Datasets

Dataset

Discription

Competition Paper

ICDAR 2015

1000 training images and 500 testing images

paper

ICDAR 2013

229 training images and 233 testing images

paper

ICDAR 2011

229 training images and 255 testing images

paper

ICDAR 2005

1001 training images and 489 testing images

paper

ICDAR 2003

181 training images and 251 testing images(word level and character level)

paper

开源库

Tesseract: c++ based tools for documents analysis and OCR,support 60+ languages [code]

Ocropy: Python-based tools for document analysis and OCR [code]

CLSTM : A small C++ implementation of LSTM networks,focused on OCR [code]

Convolutional Recurrent Neural Network,Torch7 based [code]

Attention-OCR: Visual Attention based OCR [code]

Umaru: An OCR-system based on torch using the technique of LSTM/GRU-RNN, CTC and referred to the works of rnnlib and clstm [code]

其他

DeepFont:Identify Your Font from An Image[paper]

Writer-independent Feature Learning for Offline Signature Verification using Deep Convolutional Neural Networks[paper]

End-to-End Interpretation of the French Street Name Signs Dataset [paper] [code]

Extracting text from an image using Ocropus [blog]

手写字识别

[2016-arXiv]Drawingand Recognizing Chinese Characters with Recurrent Neural Network [paper]

Learning Spatial-Semantic Context with Fully Convolutional Recurrent Network for Online Handwritten Chinese Text Recognition [paper]

Stroke Sequence-Dependent Deep Convolutional Neural Network for Online Handwritten Chinese Character Recognition [paper]

High Performance Offline Handwritten Chinese Character Recognition Using GoogLeNet and Directional Feature Maps [paper] [github]

DeepHCCR:Offline Handwritten Chinese Character Recognition based on GoogLeNet and AlexNet (With CaffeModel) [code]

如何用卷积神经网络CNN识别手写数字集？[blog][blog1][blog2] [blog4] [blog5] [code6]

Scan,Attend and Read: End-to-End Handwritten Paragraph Recognition with MDLSTMAttention [paper]

MLPaint:the Real-Time Handwritten Digit Recognizer [blog][code][demo]

caffe-ocr: OCR with caffe deep learning framework [code] (单字分类器)

牌照等识别

ReadingCar License Plates Using Deep Convolutional Neural Networks and LSTMs [paper]

Numberplate recognition with Tensorflow [blog] [code]

end-to-end-for-plate-recognition[code]

ApplyingOCR Technology for Receipt Recognition[blog][mirror]

破解验证码

[2017-Arvix]Using Synthetic Data to Train NeuralNetworks is Model-Based Reasoning[paper]

Using deep learning to break a Captcha system [blog] [code]

Breakingreddit captcha with 96% accuracy [blog] [code]

I'mnot a human: Breaking the Google reCAPTCHA [paper]

NeuralNet CAPTCHA Cracker [slides] [code] [demo]

Recurrentneural networks for decoding CAPTCHAS [blog] [code] [demo]

Readingirctc captchas with 95% accuracy using deep learning [code]

端到端的OCR：基于CNN的实现 [blog]

IAm Robot: (Deep) Learning to Break Semantic Image CAPTCHAs [paper]

参考
[1]http://handong1587.github.io/deep_learning/2015/10/09/ocr.html
[2]https://github.com/chongyangtao/Awesome-Scene-Text-Recognition

推荐阅读

java
您的数据库配置是否安全？DBSAT工具助您一臂之力！

本文探讨了Oracle提供的免费工具DBSAT，该工具能够有效协助用户检测和优化数据库配置的安全性。通过全面的分析和报告，DBSAT帮助用户识别潜在的安全漏洞，并提供针对性的改进建议，确保数据库系统的稳定性和安全性。 ... [详细]

蜡笔小新 2024-11-11 14:44:47
case
在范围[0..n-1]中产生m个不同的随机数 - Generating m distinct random numbers in the range [0..n-1]

Ihavetwomethodsofgeneratingmdistinctrandomnumbersintherange[0..n-1]我有两种方法在范围[0.n-1]中生 ... [详细]

蜡笔小新 2024-11-13 09:49:14
const
FFMpeg学习进阶：音频处理基础理论与重采样技术详解

在Android平台中，播放音频的采样率通常固定为44.1kHz，而录音的采样率则固定为8kHz。为了确保音频设备的正常工作，底层驱动必须预先设定这些固定的采样率。当上层应用提供的采样率与这些预设值不匹配时，需要通过重采样（resample）技术来调整采样率，以保证音频数据的正确处理和传输。本文将详细探讨FFMpeg在音频处理中的基础理论及重采样技术的应用。 ... [详细]

蜡笔小新 2024-11-09 13:46:55
const
三角测量计算三维坐标的代码_双目三维重建——层次化重建思考

双目三维重建——层次化重建思考FesianXu2020.7.22atANTFINANCIALintern前言本文是笔者阅读[1]第10章内容的笔记，本文从宏观的角度阐 ... [详细]

蜡笔小新 2024-11-13 19:31:37
export
更新vuex的数据为什么用mutation?

更新vuex的数据为什么用mutation?,Go语言社区,Golang程序员人脉社 ... [详细]

蜡笔小新 2024-11-13 18:30:04
case
单片微机原理P3：80C51外部拓展系统

　　外部拓展其实是个相对来说很好玩的章节，可以真正开始用单片机写程序了，比较重要的是外部存储器拓展，81C55拓展，矩阵键盘，动态显示，DAC和ADC。0.IO接口电路概念与存 ... [详细]

蜡笔小新 2024-11-12 19:51:29
java
深入解析 Lifecycle 的实现原理

本文将详细介绍 Android Jetpack 中 Lifecycle 组件的实现原理，帮助开发者更好地理解和使用 Lifecycle，避免常见的内存泄漏问题。 ... [详细]

蜡笔小新 2024-11-12 14:05:19
io
在AX2012中使用自定义查询在数据网格视图中显示数据

本文介绍了如何在AX2012中通过自定义查询在数据网格视图中显示所有记录的方法。 ... [详细]

蜡笔小新 2024-11-12 12:02:50
const
poj 3352 Road Construction

poj 3352 Road Construction ... [详细]

蜡笔小新 2024-11-12 11:24:39
io
掌握MySQL数据库的基础语法与核心操作

本文详细介绍了MySQL数据库的基础语法与核心操作，涵盖从基础概念到具体应用的多个方面。首先，文章从基础知识入手，逐步深入到创建和修改数据表的操作。接着，详细讲解了如何进行数据的插入、更新与删除。在查询部分，不仅介绍了DISTINCT和LIMIT的使用方法，还探讨了排序、过滤和通配符的应用。此外，文章还涵盖了计算字段以及多种函数的使用，包括文本处理、日期和时间处理及数值处理等。通过这些内容，读者可以全面掌握MySQL数据库的核心操作技巧。 ... [详细]

蜡笔小新 2024-11-11 23:39:51
const
使用 Matplotlib 保存 Python 动态图像为视频文件的方法与技巧

本文介绍了如何利用 `matplotlib` 库中的 `FuncAnimation` 类将 Python 中的动态图像保存为视频文件。通过详细解释 `FuncAnimation` 类的参数和方法，文章提供了多种实用技巧，帮助用户高效地生成高质量的动态图像视频。此外，还探讨了不同视频编码器的选择及其对输出文件质量的影响，为读者提供了全面的技术指导。 ... [详细]

蜡笔小新 2024-11-11 22:11:30
io
MySQL Decimal 类型的最大值解析及其在数据处理中的应用艺术

在关系型数据库中，表的设计与SQL语句的编写对性能的影响至关重要，甚至可占到90%以上。本文将重点探讨MySQL中Decimal类型的最大值及其在数据处理中的应用技巧，通过实例分析和优化建议，帮助读者深入理解并掌握这一重要知识点。 ... [详细]

蜡笔小新 2024-11-11 19:36:19
const
Codeforces竞赛解析：Educational Round 84（Div. 2评级），题目A：奇数和问题

Codeforces竞赛解析：Educational Round 84（Div. 2评级），题目A：奇数和问题 ... [详细]

蜡笔小新 2024-11-11 14:02:18
io
MSP430F5438 ADC12模块应用与学习心得

在最近的实践中，我深入研究了MSP430F5438的ADC12模块。尽管该模块的功能相对简单，但通过实际操作，我对MSP430F5438A和MSP430F5438之间的差异有了更深刻的理解。本文将分享这些学习心得，并探讨如何更好地利用ADC12模块进行数据采集和处理。 ... [详细]

蜡笔小新 2024-11-10 11:48:42
java
C#编程指南：利用ASP.NET和JavaScript实现带有Fingerprint功能的Web应用登录系统

本指南介绍了如何在ASP.NET Web应用程序中利用C#和JavaScript实现基于指纹识别的登录系统。通过集成指纹识别技术，用户无需输入传统的登录ID即可完成身份验证，从而提升用户体验和安全性。我们将详细探讨如何配置和部署这一功能，确保系统的稳定性和可靠性。 ... [详细]

蜡笔小新 2024-11-09 18:14:37

love_xiao奇

这个家伙很懒，什么也没留下！

Tags | 热门标签

RankList | 热门文章

Dataset	Discription	Competition Paper
ICDAR 2015	1000 training images and 500 testing images	paper
ICDAR 2013	229 training images and 233 testing images	paper
ICDAR 2011	229 training images and 255 testing images	paper
ICDAR 2005	1001 training images and 489 testing images	paper
ICDAR 2003	181 training images and 251 testing images(word level and character level)	paper

文字检测与识别资源

手写字识别

牌照等识别

破解验证码

参考[1]http://handong1587.github.io/deep_learning/2015/10/09/ocr.html[2]https://github.com/chongyangtao/Awesome-Scene-Text-Recognition var cpro_id = "u6885494";

参考
[1]http://handong1587.github.io/deep_learning/2015/10/09/ocr.html
[2]https://github.com/chongyangtao/Awesome-Scene-Text-Recognition