Handwritten Chinese Recognition Method and System Based on Text Detection and Text Recognition

By adopting a two-stage method based on text detection and text recognition in handwritten Chinese recognition technology, combining the global fine-grained feature extraction module and the multi-head self-attention mechanism, the missed and misidentified recognition problems of long text recognition in the existing technology are solved, and the recognition accuracy and model stability are improved.

CN119625707BActive Publication Date: 2025-06-10SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411705868.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-06-10
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Existing handwritten Chinese recognition technology is prone to misrecognition and misrecognition when processing long texts. The CTC-based method has the problem of insufficient RNN performance, and the attention mechanism method will experience attention drift when detecting long texts.

Method used

Using a two-stage method based on text detection and text recognition, text boxes are generated through the trained text detection model, and text recognition is used to use the global fine-grained feature extraction module to combine pixel reorganization strategy and multi-head self-attention mechanism to improve the detection and recognition performance of the model.

Benefits of technology

Effectively replace manual segmentation, improves the accuracy of text lines recognition, reduces the occurrence of missed and misidentified, enhances the model's ability to recognize long texts, and solves the problem of attention drift.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625707B_ABST
    Figure CN119625707B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for handwritten Chinese character recognition based on text detection and text recognition. The method includes: inputting a document image to be recognized into a trained text detection model to obtain a text detection result; the trained text detection model first performs a preprocessing operation on the document image to be recognized, extracts features from the preprocessing operation result, performs feature fusion on the extracted features, performs a line regression operation and a probability map prediction on the fused features, and generates a text box based on the predicted probability map; inputting the text detection result into a trained text recognition model to obtain a text recognition result; performing a preprocessing operation on the text region image within the text box, performing an image augmentation process on the preprocessed text region image, performing a block embedding operation on the augmented image, extracting global fine-grained features from the image after the block embedding operation, and performing recognition on the extracted global fine-grained features to obtain a text recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Chinese character recognition, and in particular to a handwritten Chinese character recognition method and system based on text detection and text recognition. Background Art

[0002] Computer text recognition, commonly known as optical character recognition (OCR), is a subfield of computer vision and pattern recognition. It is a technology that uses a computer to automatically recognize character instances contained in an image and process them into information that can be recognized by the computer and understood by people. With the development and popularization of artificial intelligence, especially computer vision technology, many practical applications related to text recognition have greatly facilitated our lives, such as license plate recognition, bank card recognition, invoice recognition, picture-to-text recognition, and so on. The text recognition task is usually divided into two stages: text detection and text recognition. Text detection is a subtask of object detection, which is used to obtain the position information of multiple text instances contained in an image as the front end of the text recognition task; text recognition is used to recognize and transcribe the detected text instances into structured data that can be recognized by the computer and understood by people.

[0003] With the popularization of computers and the Internet among the public, intelligent marking has gradually shown its brilliance in the field of education. Handwritten intelligent recognition, as an important foundation of intelligent marking technology, has also received significant attention. This technology aims to transcribe the scanned image of the examinee's answer content into text as the input for intelligent marking to achieve the final score, and can optimize irregular handwritten text to a certain extent. The handwritten Chinese character recognition method is more complex than the ordinary handwritten character recognition method, which is specifically reflected in the diversity of Chinese characters and their structures. The similarity between the two is that due to the great differences in the writing styles of different writers, a large amount of data samples are needed to cover various writing styles and variants.

[0004] The handwritten Chinese character recognition method mainly includes traditional machine learning methods and deep learning-based methods, etc. Early traditional machine learning methods relied on manually designed feature extractors, including methods such as edge detection, contour feature extraction, and structural feature extraction. Then, the extracted features were encoded and sent to a classifier for recognition. The classifiers used in this stage mainly included support vector machines (SVM), random forests, and naive Bayes classifiers, etc. This traditional method requires manual design and extraction of features. When dealing with handwritten Chinese text, due to the complex structure and great diversity of Chinese characters, the generalization ability of the model will be limited, and the performance of the model will be relatively poor when dealing with unknown or other writing style data.

[0005] With the development of deep learning, many deep learning methods including those based on CTC, Attention, and Transformer have achieved good performance in scene text recognition, which has also promoted the vigorous development of handwritten Chinese character recognition. These deep learning-based models can automatically learn the optimal feature representations from data, have strong generalization ability, and can handle large datasets. This undoubtedly simplifies the steps of handwritten Chinese character recognition and makes it possible to achieve end-to-end handwritten Chinese character recognition. However, current handwritten Chinese character recognition mainly focuses on text line recognition, that is, it can only predict the results of a single line of text images. And due to the performance issues of the RNN used in current CTC-based methods, and the inevitable attention drift that occurs when the attention mechanism-based methods detect long texts, it will lead to problems such as missed recognition and misrecognition in long text line scenarios. Summary of the Invention

[0006] To solve the deficiencies of the prior art, the present invention provides a method and system for handwritten Chinese character recognition based on text detection and text recognition;

[0007] On the one hand, a method for handwritten Chinese character recognition based on text detection and text recognition is provided, including:

[0008] Obtain a document image to be recognized;

[0009] Input the document image to be recognized into a trained text detection model to obtain a text detection result; wherein, the trained text detection model first performs a preprocessing operation on the document image to be recognized, then extracts features from the preprocessing operation result, then performs feature fusion on the extracted features, then performs line regression operation and probability map prediction on the fused features, and finally generates a text box based on the predicted probability map;

[0010] Input the text detection result into a trained text recognition model to obtain a text recognition result; perform a preprocessing operation on the text region image within the text box, perform image augmentation processing on the preprocessed text region image, perform block embedding operation on the augmented image, perform global fine-grained feature extraction on the image after the block embedding operation, and perform recognition on the extracted global fine-grained features to obtain a text recognition result.

[0011] On the other hand, a system for handwritten Chinese character recognition based on text detection and text recognition is provided, including:

[0012] An acquisition module, which is configured to: obtain a document image to be recognized;

[0013] A text detection module, which is configured to: input a document image to be recognized into a trained text detection model to obtain a text detection result; wherein, the trained text detection model first performs a preprocessing operation on the document image to be recognized, then extracts features from the result of the preprocessing operation, then performs feature fusion on the extracted features, then performs line regression operation and probability map prediction on the fused features, and finally generates a text box based on the predicted probability map;

[0014] A text recognition module, which is configured to: input the text detection result into a trained text recognition model to obtain a text recognition result; perform a preprocessing operation on the text region image within the text box, perform image augmentation processing on the preprocessed text region image, perform block embedding operation on the augmented image, perform global fine-grained feature extraction on the image after the block embedding operation, and perform recognition on the extracted global fine-grained features to obtain a text recognition result.

[0015] On the other hand, an electronic device is also provided, including:

[0016] A memory for non-temporarily storing computer-readable instructions; and

[0017] A processor for running the computer-readable instructions,

[0018] wherein, when the computer-readable instructions are run by the processor, the method described in the first aspect above is executed.

[0019] On the other hand, a storage medium is also provided, which non-temporarily stores computer-readable instructions, wherein when the non-temporary computer-readable instructions are executed by a computer, the method described in the first aspect is executed.

[0020] On the other hand, a computer program product is also provided, including a computer program, and the computer program is used to implement the method described in the first aspect above when running on one or more processors.

[0021] The above technical solutions have the following advantages or beneficial effects:

[0022] (1) For page-level handwritten recognition, this method proposes a two-stage method of detection and recognition. The excellent detection module can not only effectively replace manual segmentation, but also the novel recognition module can more effectively recognize the detected text lines.

[0023] (2) For the task of detecting text lines in page handwritten recognition, the present invention designs a novel line regression module, which can directly utilize the rich information of different feature dimensions contained in the fused feature map to obtain the position where the text line appears, helping the model better locate the text region.

[0024] (3) The present invention adopts a pixel recombination strategy to replace the commonly used interpolation upsampling strategy for the final predicted probability map, which helps to retain more information and improve the detection performance.

[0025] (4) In the preprocessing of the detection stage, the present invention shrinks the label area to avoid the occurrence of text box adhesion, and combined with post-processing, it can effectively help the detection model divide different text line areas.

[0026] (5) In the recognition stage, the present invention proposes an effective global fine-grained feature extraction module, which can fully extract character fine-grained information and image global information. This module stacks multiple sub-modules, in each sub-module, grouped convolution and multi-head self-attention mechanism are alternately connected, and the sub-modules are connected by a width downsampling module, so that the model can not only utilize the character fine-grained structural features and the context information provided by the image global, but also fully obtain the rich information contained at the text line level, and then better complete the recognition task.

[0027] (6) The present invention proposes a novel sequence loss function to suppress the negative samples contained in the final recognition result, assist the CTC decoder to jointly optimize the objective function, and help to improve the model recognition performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention, and the schematic embodiments and descriptions thereof are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0029] Figure 1 is the flowchart of the method in Embodiment 1;

[0030] Figure 2 is the structure diagram of the handwritten Chinese detection model in Embodiment 1;

[0031] Figure 3 is the internal structure diagram of the fusion pyramid module in Embodiment 1;

[0032] Figure 4 is the internal structure diagram of the line regression module in Embodiment 1;

[0033] Figure 5 is the internal structure diagram of the width pooling module in Embodiment 1;

[0034] Figure 6 is the internal structure diagram of the probability map prediction module in Embodiment 1;

[0035] Figure 7 is the main flowchart of the text box generation module in Embodiment 1;

[0036] Figure 8 is the structure diagram of the handwritten Chinese recognition model in Embodiment 1;

[0037] Figure 9 Internal structure diagram of the block embedding module in the first embodiment;

[0038] Figure 10 Internal structure diagram of the global fine-grained feature extraction module in the first embodiment;

[0039] Figure 11 Internal structure diagram of the feature extraction module in the first embodiment;

[0040] Figure 12 Internal structure diagram of the downsampling layer in the first embodiment;

[0041] Figure 13 Internal structure diagram of the height pooling layer in the first embodiment;

[0042] Figure 14 Internal structure diagram of the output head module in the first embodiment;

[0043] Figure 15 Schematic diagram of the label file in the first embodiment;

[0044] Figure 16 Schematic diagram of shrinking each line of text area corresponding to the label inward by a certain distance in the first embodiment. Detailed implementation manners

[0045] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0046] First embodiment

[0047] This embodiment provides a handwritten Chinese character recognition method based on text detection and text recognition;

[0048] As Figure 1 shown, the handwritten Chinese character recognition method based on text detection and text recognition includes:

[0049] S101: Obtain a document image to be recognized;

[0050] S102: Input the document image to be recognized into the trained text detection model to obtain a text detection result; wherein, the trained text detection model first performs a preprocessing operation on the document image to be recognized, then performs feature extraction on the result of the preprocessing operation, then performs feature fusion on the extracted features, then performs line regression operation and probability map prediction on the fused features, and finally generates a text box based on the predicted probability map;

[0051] S103: Input the text detection result into the trained text recognition model to obtain the text recognition result; perform preprocessing operations on the text region image within the text box, perform image augmentation processing on the preprocessed text region image, perform block embedding operations on the augmented image, perform global fine-grained feature extraction on the image after the block embedding operation, and perform recognition on the extracted global fine-grained features to obtain the text recognition result.

[0052] Further, S101: Obtain the document image to be recognized, including: obtaining the document image by means of scanning or photographing. The document image includes an exam paper, and the exam paper includes a handwritten answer area.

[0053] The basic architecture of the handwritten Chinese detection model used in the present invention is a deep learning model based on a fusion pyramid module. In practical application scenarios such as exam paper images, there are usually multiple subjective question answer areas, and the answer areas usually contain several text region lines, and at the same time contain irrelevant information such as question numbers and answer boxes. This requires using a text detection model to detect the text regions. Compared with the traditional method of directly regressing the text box based on regression or the text detection method based on segmentation that directly returns a probability binary map, this detection model can better distinguish text instances and greatly improve the performance of the detection model.

[0054] The handwritten Chinese detection model performs text detection on the input scanned document image. The detection model includes a preprocessing module, a feature extraction module, a fusion pyramid module, a line regression module, a probability map prediction module, and a post-processing module. The basic structure diagram is as Figure 2 shown.

[0055] Further, the trained text detection model includes:

[0056] A first preprocessing module, a feature extraction module, a fusion pyramid module, a probability map prediction module, and a post-processing module connected in sequence; the post-processing module includes: a binarization sub-module and a text box generation sub-module; wherein, the output end of the fusion pyramid module is also connected to the input end of the line regression module.

[0057] Further, the first preprocessing module includes: performing grayscale processing, binarization processing, and size adjustment processing on the document image to be recognized.

[0058] Further, the first preprocessing module is used to preprocess the scanned document image and the label file: for the scanned document image, perform grayscale and binarization (i.e., 0, 1) processing on it, and uniformly adjust the height and width.

[0059] For the label file, since the lines written by some candidates are relatively dense in practice, when directly training with the original text line label file, there will be a sticking phenomenon where multiple lines of text are detected as one line. The label file is a txt file, and each line of it contains the file name and the coordinate value annotation of the text area of the corresponding file. For example (file name 1x1.y1.x2.y2, x3.y3.x4.y4,...), the corresponding example diagram has been given, and the position is at Figure 15 。

[0060] Such as Figure 16 shown, in order to avoid this phenomenon, the present invention innovatively adopts the method of shrinking each line of text area corresponding to the label inward by a certain distance. Here, the area D area 、perimeter D perimeter and the scaling factor α (α is a hyperparameter) are used to calculate the shrinking distance distance, and all the boundary points of the original text area are shrunk inward by the distance distance:

[0061]

[0062] After that, a blank image of the same size as the original image is generated, and the scaled text area is marked in this blank image, and then it is also binarized. The scaled text area, such as Figure 16 shown.

[0063] Further, the feature extraction module includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a fifth convolutional layer connected in sequence.

[0064] Further, the feature extraction module includes:

[0065]

[0066] Among them, conv1 represents the first convolutional layer, conv2 represents the second convolutional layer, conv3 represents the third convolutional layer, conv4 represents the fourth convolutional layer, and conv5 represents the fifth convolutional layer;

[0067] The image input to the feature extraction module is the grayscale image preprocessed from the document image to be recognized where H and W are the height and width of the image respectively, and F 1 、F 2 、F 3 、F 4 、F 5 are the output results of the five layers of convolution of ResNet-50 respectively, and among them, F 2 、F 3 、F 4 、F 5Feed into the fusion pyramid module.

[0068] It should be understood that the feature extraction module is used to extract low-dimensional and high-dimensional features of the preprocessed document image to be recognized, and then input these features into the fusion pyramid module to fuse the feature information of each dimension. This module uses the Residual Network ResNet-50 to extract low-dimensional and high-dimensional features of the image.

[0069] ResNet (Residual Network) is a deep neural network architecture. Its key innovation is the introduction of residual connections. In this way, the network can more easily learn the identity mapping, making the network very deep.

[0070] Furthermore, the fusion pyramid module includes:

[0071] The features F 2 , F 3 , F 4 and F 5 extracted from the last four layers of the feature extraction module have different numbers of channels. First, their numbers of channels are unified to C (C is a hyperparameter) through 1×1 convolutions respectively, obtaining feature maps F' 2 , F' 3 , F' 4 and F' 5 ;

[0072] Then, F' 5 is upsampled by a factor of 2. The upsampling strategy uses nearest neighbor interpolation, and then it is fused with F' 4 to obtain a sub-fusion feature map F'' 4 ;

[0073] After that, F'' 4 is upsampled by a factor of 2 and fused with F' 3 to obtain a sub-fusion feature map F'' 3 ;

[0074] Then, F'' 3 is upsampled by a factor of 2 and fused with F' 2 to obtain a sub-fusion feature map F'' 2 ;

[0075]

[0076] It should be understood that directly upsampling the original features may lead to information loss and feature confusion, affecting the quality of the final feature representation. Therefore, a step-by-step upsampling strategy is used, which allows the network to merge information at different scales. This ensures the alignment of feature maps at different resolutions, enabling better fusion of features at different levels. It helps to retain the semantic information contained in high-dimensional features and the detailed information contained in low-dimensional features, thereby better detecting text regions at different scales.

[0077] After that, F″ 2 、F″ 3 、F″ 4 、F″ 5 go through a 3×3 convolution layer. The purpose is to reduce their number of channels to C / 4. Using feature compression here helps reduce the model complexity, improve computational efficiency, and also prevent overfitting caused by an overly complex model.

[0078] Then, Conv2d(F″ 3 ), Conv2d(F″ 4 ), and Conv2d(F″ 5 ) are upsampled to the same spatial size as Conv2d(F″ 2 ). Finally, F″′ 2 and the upsampling results of F″′ 3 、F″′ 4 、F″′ 5 are all fed into the connection layer and concatenated in the channel dimension to generate the final fused feature map f.

[0079]

[0080] It should be understood that the fusion pyramid module is designed based on the feature pyramid module and is used to fuse feature information from low-dimensional to high-dimensional, enabling the model to better distinguish text regions and non-text regions and detect short texts. Using the last four outputs F 2 、F 3 、F 4 、F 5 of the feature extraction module as inputs, after convolution and upsampling, the final fused feature map is obtained, and its specific structure is as Figure 3 shown.

[0081] The main idea of the feature pyramid is to construct multi-scale feature maps in the neural network, enabling the network to detect targets at different scales and retain high-level semantic information. Such a design allows the network to detect targets of different sizes and effectively identify them even if there are significant differences in the size of the targets in the image. In actual application scenarios, the lengths of text lines vary, and using the feature pyramid strategy can more effectively help detect text regions of different shapes.

[0082] Furthermore, the line regression module includes:

[0083] Taking the output f of the fusion pyramid module as the input of the line regression module, first obtain f through a 3×3 convolutional layer with a stride of 1 in the height dimension and a stride of 2 in the width dimension, a batch normalization layer, and a ReLu activation function layer 1 , whose width dimension is compressed to W / 8, where W is the width of the original image.

[0084]

[0085] Next, f 1 Then obtain f through a 1×1 convolutional layer, a batch normalization layer, and a ReLu activation function layer at the second level 2 , whose channel number dimension is compressed to C / 4, where C is the number of channels of f.

[0086]

[0087] Then, f 2 Then obtain f through a 3×3 convolutional layer with a stride of 1 in the height dimension and a stride of 2 in the width dimension, a batch normalization layer, and a ReLu activation function layer at the third level 3 , whose width dimension is compressed to W / 16.

[0088]

[0089] Finally, f 3 Obtain the height information of the text region through the width pooling module.

[0090] The width pooling module is used to compress the feature map, enabling the model to concentrate on extracting vertical features. Its specific structure is as Figure 5 shown.

[0091] The width pooling module first performs global pooling on f 3The width is compressed to 1, and then the compressed image is fed into a 1×1 convolution to change the number of channels to 1. After passing through a ReLu activation function layer, the squeeze function is used to remove the width dimension and the channel dimension. Here, squeeze is a function used to remove dimensions with a size of 1 in a tensor. Finally, the compressed result is fed into a linear layer to obtain the upper-left corner height information and the lower-right corner height information of D text regions, and the final information is stored in Height Info (where D is a hyperparameter).

[0092]

[0093] f squeeze = Squeeze(ReLu((Conv2d(f 4 ))) (20)

[0094]

[0095] It should be understood that the present invention innovatively proposes a line regression module. This module uses the features of the fusion graph to predict the vertical coordinates of the text lines for the input text line data, thereby improving the performance of the detection model. Its basic structure is as Figure 4 shown.

[0096] It should be understood that Squeeze is a function used to remove dimensions with a size of 1 in a tensor. For example, for a tensor with a dimension of (1, 3, 1, 5), after using the Squeeze function to compress the first and third dimensions, the dimension of this tensor becomes (3, 5). The parameter Height Info is an array of size D×2. D represents the number. In each 1×2 array, the first number represents the upper-left corner height information, and the second number represents the lower-right corner height information.

[0097] Furthermore, the probability map prediction module includes:

[0098] Taking the output f of the fusion pyramid module as the input of the probability map prediction module, first, a feature map f′ with the number of channels compressed to 1 / 4 of the original number of channels is obtained through a 3×3 convolutional layer, a batch normalization layer, and a ReLu activation function layer:

[0099]

[0100] Then, f′ is passed through a 3×3 convolutional layer and a batch normalization layer to obtain f″, where the number of channels of f″ is 16:

[0101]

[0102] Finally, f″ is sent into the variable-resolution layer to obtain f″′, where the number of channels of f″′ is stretched to 1, and its spatial resolution is the same as that of the document image to be recognized. The variable-resolution layer uses a pixel reorganization strategy, which increases the height and width of f″ proportionally by rearranging the channel dimension information to other dimensions, thus improving the image resolution. Compared with directly upsampling the image, the pixel reorganization strategy can retain more information.

[0103]

[0104] Finally, the image f″′ output by the variable-resolution layer is sent into the Sigmoid activation function layer, and the probability map P is obtained, where the value of each pixel is in the range of 0-1:

[0105] P = Sigmoid(f″′)(25).

[0106] It should be understood that the present invention uses the fused feature map output by the fusion pyramid module to predict the probability map. Here, the pixel reorganization strategy is innovatively used to replace the direct upsampling or interpolation strategy used in the previous methods for improving the image resolution. Its specific structure is as Figure 6 shown.

[0107] It should be understood that the present invention greatly improves the performance of the detection model in distinguishing different text instances through the innovative generation strategy designed in the post-processing module. Specifically, this module includes a binarization sub-module and a text box generation sub-module.

[0108] Further, the binarization sub-module includes:

[0109] The probability map P output by the probability map prediction module is binarized to obtain the binarized map B. The specific method is to give a threshold λ (λ is a hyperparameter), traverse each pixel of the probability map, and when the value of the pixel is greater than or equal to λ, its value is changed to 1, otherwise it is changed to 0:

[0110]

[0111] Further, the text box generation sub-module includes:

[0112] (1) Find the contour set contours corresponding to the binarized map: First, multiply the pixel values of the binarized image by 255 to obtain an operable binarized image, and then obtain the contour set contours corresponding to all text regions of the operable binarized image. Contours initially include several contours, and these contours are all irregular shapes, and each contour contains several points.

[0113] (2) Find the minimum bounding rectangle of each contour: For each contour contour in contours, calculate the minimum bounding rectangle rectangle of the corresponding point set of the contour. If the length or width of the minimum bounding rectangle rectangle is less than the threshold l (l is a hyperparameter), then filter out the minimum bounding rectangle rectangle.

[0114] (3) Create a dynamic contour that can perform offset operations: Create a rectangular box box with the same size as the minimum bounding rectangle rectangle but capable of performing offset operations.

[0115] In the preprocessing module, to avoid the occurrence of adhesion phenomena, the original label file was scaled. Therefore, the obtained rectangle rectangle that meets the threshold l is scaled back to the size corresponding to the original text box.

[0116] (4) Operate on the rectangular box box to finally obtain the text box: For the obtained rectangular box box, calculate the area box area and perimeter box perimeter , and then the moving distance distance can be calculated jointly with the scaling factor α:

[0117]

[0118] Finally, move all the boundary points of the rectangular box box outward by a distance of distance to obtain the corrected rectangular box; then calculate the minimum bounding rectangle of the corrected rectangular box. The minimum bounding rectangle of the corrected rectangular box is the rectangular text box corresponding to the contour. Finally, in the image, the area inside the rectangular text box is regarded as the text area, and other parts are regarded as non-text areas.

[0119] Adopt the method of finding the text box corresponding to each text area according to the binary image B. The main process is as Figure 7 shown: First, draw the contours of each text object in the binary image and store the results in contours; then find the minimum bounding rectangle rectangle of each contour contour in contours. After that, construct a dynamic contour box that can perform offset operations according to the four corner points of the rectangle. Finally, calculate the moving distance distance according to the rectangle area, perimeter, and the scaling factor α (α is a hyperparameter), adjust the box according to distance, and finally horizontally correct the adjusted box to obtain the rectangular text box corresponding to the contour.

[0120] Furthermore, S102: Input the document image to be recognized into the trained text detection model to obtain the text detection result; for the trained text detection model, the training process includes:

[0121] Construct a first training set, where the first training set is a document image with known text detection regions;

[0122] Input the first training set into the text detection model to train the text detection model. When the loss function value of the text detection model no longer decreases, or the iteration reaches the set number of times, stop training to obtain the trained text detection model.

[0123] Furthermore, the loss function of the text detection model includes:

[0124] L = L p + λL r (28)

[0125] where L p is the cross-entropy loss between the result of the post-processing module and the original unscaled label, and L r is the SmoothL 1 loss between the height information of the text region predicted by the line regression module and the original unscaled label. λ is a weight parameter, and the overall loss function L drives the training of the text detection model.

[0126] where L p The specific formula is as follows:

[0127]

[0128] where y i represents the label value of sample i, and p i is the probability that sample i is predicted as the positive class.

[0129] L r The specific formula is as follows:

[0130]

[0131] where x represents the difference between the predicted value Height Info and the top-left height information and bottom-right height information of the true text region.

[0132] Furthermore, in S103: Input the text detection result into the trained text recognition model to obtain the text recognition result, where the trained text recognition model includes:

[0133] A second preprocessing module, an image augmentation module, a block embedding module, a global fine-grained feature extraction module, and an output head module connected in sequence.

[0134] The basic architecture of the handwritten Chinese recognition model used in the present invention is a deep learning model based on a global fine-grained feature extraction module, which effectively solves the problems of missed recognition and misrecognition of long texts by other recognition models. Moreover, compared with previous methods, the present invention enables the model to fully utilize the relationships between the text regions and non-text regions in the picture, allowing the model to fully perceive global features and effectively solving the problem of attention drift encountered by other recognition models.

[0135] The handwritten Chinese recognition model for document images recognizes the independent text region images obtained by the detection model. The recognition model includes a preprocessing module, an image augmentation module, a block embedding module, a global fine-grained feature extraction module, and an output head module. The basic structure is as Figure 8 shown.

[0136] Furthermore, the second preprocessing module includes:

[0137] First, the text line image is grayscaled, and then its width and height are scaled proportionally until the height is the same as the set height. At this time, if the width of the image is greater than the set width, it is directly compressed in the width dimension; otherwise, blanks are filled on the right side of the image until the set width is reached.

[0138] It should be understood that the preprocessing module is used to preprocess the text line images, dictionaries, and labels output by the detection module: Since the lengths of the input text line regions are different, for shorter text line images, if they are directly scaled to the size required by the model, misrecognition or non-recognition may occur. Therefore, when processing text line images, first grayscale them, and then scale their width and height proportionally until the height is the same as the height of the input image required by the model. At this time, if the width of the image is greater than the width of the input image required by the model, it is directly compressed in the width dimension; otherwise, blanks are filled on the right side of the image until the width of the image required by the model is reached. For the label file, the correspondence between the image and the label is constructed, and the label corresponding to each image is mapped to the index of the character in the corresponding dictionary file, and then the blank filler <pad>Unify the lengths of all tags.

[0139] Furthermore, the image augmentation module includes: an image random cropping operation, an image random rotation operation, and an image Gaussian blur operation;

[0140] The image random cropping operation means that since there may be slight errors in the detection module, the text lines segmented may have incomplete character display or excessive blank space in the height dimension. Therefore, the original text line image is randomly cropped to simulate the situation of slight errors in image segmentation in the usage scenario, where Δh is the height cropping amount:

[0141] y t = y + Δh (31)

[0142] The image random rotation operation means that since the images in the actual application scenario may have character tilt or overall tilt of the text line, rotating the original image around the image center by a certain angle can simulate this tilt phenomenon, where θ is the rotation angle, x o , y o is the image center point:

[0143]

[0144] The image Gaussian blur operation means that in order to reduce image noise and lower its level of detail, the original image is subjected to Gaussian blur to weaken the influence of noise on feature extraction.

[0145] The so-called Gaussian blur is a process of weighted averaging the entire image. The value of each pixel point is obtained by weighted averaging its own value and the values of other pixels in the neighborhood.

[0146] Its specific operation is: scanning each pixel in the image with a mask, and replacing the value of the mask center pixel with the weighted average gray value of the pixels in the neighborhood determined by the mask. Its mathematical representation is as follows, where x is the distance to the pixel center and σ is the standard deviation:

[0147]

[0148] It should be understood that since the images in the actual application scenario may have tilts, incomplete character display, etc., in order to improve the generalization ability of the model, at the same time, it can also increase sample diversity and weaken the occurrence of model overfitting. Using the image augmentation module can effectively improve the performance of the recognition model. Considering the actual situation of the handwritten Chinese text line images in the document images, the image augmentation module of the present invention uses a series of image augmentation operations that retain the overall structure and local feature information of the characters.

[0149] Furthermore, the block embedding module includes:

[0150] First, input the original image I into the patch embedding module, where H and W are the height and width of the original image respectively; the original image I passes through a 3×3 convolutional layer with a stride of 2, a batch normalization layer, and a ReLu activation function layer to obtain P′ emb , where C 0 is the number of channels of the patch to be input into the global fine-grained feature extraction module (C 0 is a hyperparameter).

[0151] P′ emb = ReLu(BatchNorm2d(Conv2d(I))) (34)

[0152] Then P′ emb passes through another 3×3 convolutional layer with a stride of 2, a batch normalization layer, and a ReLu activation function layer again to obtain the patch embedding P″ enb , where

[0153] P″ emb = ReLu(BatchNorm2d(Conv2d(P′ emb ))) (35)

[0154] Finally, combine P″ emb with the position embedding to obtain the final patch embedding P emb , where the position embedding is a learnable embedding. Specifically, a trainable vector is created for each position. Using the position embedding can explicitly provide the model with the position information of the elements in the sequence.

[0155] P emb = P″ emb + Position Embedding (36)

[0156] It should be understood that the patch embedding module divides the original image data into several non-overlapping patches, changing the smallest unit of the picture from pixels to patches. The patch embedding is obtained by performing an overlapping patch embedding operation using two 3×3 two-dimensional convolutions with a stride of 2. Its specific structure is as Figure 9 shown.

[0157] Furthermore, the global fine-grained feature extraction module includes: a first stage, a second stage, a third stage, and a height pooling layer connected in sequence;

[0158] The first stage includes: N 1 successively connected feature extraction sub-modules;

[0159] The second stage includes: a first downsampling layer and N connected in series 2 feature extraction sub-modules;

[0160] The third stage includes: a second downsampling layer and N connected in series 3 feature extraction sub-modules.

[0161] Among them, the feature extraction sub-module adopts stacked grouped convolutions and self-attention mechanism, which enables the model to capture features at different granularities. The specific process is as follows. First, the output image of the block embedding module or the downsampling layer is sent into the feature extraction module. The feature extraction sub-module contains sub-modules with stacked grouped convolutions and self-attention connected alternately. After the image passes through the grouped convolution, it is reshaped into a vector form, and then after passing through multiple groups of self-attention layers, it is reshaped into the original image format. If it is in the first or second stage, it is sent to the downsampling layer of the next stage for downsampling the height; otherwise, it is sent to the height pooling layer to compress the height to 1. The structure of the feature extraction sub-module is as Figure 11 shown.

[0162] For the image P i-1 input into the feature extraction sub-module, first it is sent into the grouped convolution layer with layer normalization. The specific steps are to evenly divide the feature a i obtained by layer normalization into k parts in the spatial dimension, perform 3×3 convolution operations on each part separately, and finally splice the results after convolution of each group and introduce a skip connection to obtain a' i :

[0163] a i = LayerNorm(P i-1 )(37)

[0164]

[0165] After that, a' i is sent into a multi-layer perceptron (MLP) with layer normalization, and a skip connection is introduced again to obtain the feature c i extracted by the grouped convolution:

[0166] c i = reshape(MLP(LayerNorm(a' i )) + a' i )(40)

[0167] Next, a reshape operation is performed on c i to transform it into a vector form b i , where reshape is a function used to change the shape of the tensor. The main operation here is to reshape c i The last two dimensions are merged into one dimension to obtain b i After that, b i is sent to the layer normalization module:

[0168] b i = reshape(c i )(41)

[0169] b' i = LayerNorm(b i )(42)

[0170] After that, b' i is sent to the self-attention layer to obtain dependencies, and then a skip connection is introduced to generate the intermediate variable h i . The multi-head self-attention mechanism is used in the self-attention layer, and its process is as follows:

[0171] Q, K, V = b' i (43)

[0172]

[0173] MultiHead(Q, K, V) = Concat(head 1 , …, head h )W O (45)

[0174] head i = Attention(QW i Q , KW i K , VW i V (46)

[0175] h i = MultiHead(Q, K, V) + b i (47)

[0176] Among them, d k is the dimension of K, h is the number of self-attention heads (h is a hyperparameter), Among them, d model is the number of channels of P i-1 , d v is the dimension of V, and there is d k = d v = d model / h.

[0177] Finally, h i It is fed into a multi-layer perceptron (MLP) with normalization, and skip connections are introduced again. Finally, a reshape operation is performed to obtain P'. i :

[0178] P' i = reshape(MLP(LayerNorm(h i )) + h i )(48)

[0179] Next, if P' i is not in the last stage, it is fed into the downsampling layer of the next stage to downsample the height. Its specific structure is as Figure 12 shown. There are mainly two reasons for only downsampling the height and not the width here. One is that since the recognition module mainly recognizes text line images, the text information contained in the width dimension of the image is very rich, and downsampling the width is likely to cause information loss; the other is that the width is related to the character length, and the larger the width, the sparser the final prediction result of the recognition module, which helps to make the CTC decoding in the subsequent output head module more effective.

[0180] Specifically, P' i first passes through a 3×3 convolution with a stride of 2 in the height dimension and a stride of 1 in the width dimension, which will halve the height dimension; then it is fed into layer normalization to obtain P i , and it is fed into the next stage:

[0181]

[0182] Among them, H i-1 , W i-1 are the height and width of P i-1 respectively, and C i is the number of channels input to the next stage (C i is a hyperparameter). If P' i is in the last stage, it is fed into the height pooling layer. Its specific structure is as Figure 13 shown. The height pooling layer first compresses the image height to 1 through a global pooling operation in one step, and then feeds the compressed image into a 3×3 convolution to change the number of channels to C out (C out is a hyperparameter), and finally passes through an activation function to obtain the output result P out of the global fine-grained feature extraction module:

[0183]

[0184] Among them, W i-1 is the width of P u-1 , and then P out Feeding it into the output head module can obtain the prediction result of the recognition model.

[0185] It should be understood that Transformer generally refers to a model architecture based on the self-attention mechanism (Self-Attention Mechanism), which first showed excellent performance in the field of natural language processing and has recently shone in the field of image processing. The Transformer model is different from traditional convolutional neural networks (CNNs) or recurrent neural networks (RNNs). It mainly consists of self-attention layers (Self-Attention Layers). The self-attention mechanism can find the dependencies between positions in the input sequence, allowing the model to simultaneously focus on different positions in the input sequence. Since the text line images contain texts of different lengths, this inevitably leads to problems such as missed recognition and misrecognition during model recognition. And due to the existence of longer texts, this makes the model inevitably have the problem of attention drift, affecting the recognition effect.

[0186] To solve these problems, the present invention innovatively designs a global fine-grained feature extraction module, which is used to accurately perceive the text and its fine-grained features in the image, and can capture the feature differences between the text region and the non-text region, so as to accurately recognize long texts. Its basic structure is as Figure 10 shown.

[0187] Furthermore, the output head module includes:

[0188] Using the output P out of the global fine-grained feature extraction as the input to predict the final output result sequence. Its specific structure is as Figure 14 shown. Specifically, the output head module first uses the squeeze operation to compress the height dimension, and then obtains pred′ through a linear layer, and its number of channels is C char (C char is the dictionary size), and then uses CTC decoding, which is a decoding method for sequence alignment. By calculating the CTC loss (L ctc ), it guides the alignment of the model output result and the label. Finally, the text sequence pred is obtained:

[0189]

[0190] pred = CTC(pred′)(53)

[0191] Furthermore, in S103: Input the text detection result into the trained text recognition model to obtain the text recognition result. The training process of the trained text recognition model includes:

[0192] Construct a second training set, where the second training set is an image of a text region with known text recognition results;

[0193] Input the second training set into the text recognition model to train the text recognition model. When the loss function value of the text recognition model no longer decreases or the iteration reaches the set number of times, stop the training to obtain the trained text detection model.

[0194] The loss function L of the text recognition model is as follows:

[0195] L = L ctc + λL seq (54)

[0196] where L ctc is the CTC decoding loss, λ is the weight parameter, and L seq is the sequence loss function. The overall loss function L drives the training of the recognition model. Among them, L ctc The specific formula is as follows:

[0197]

[0198] L ctc = -∑logp(l|x)(57)

[0199] Here, represents the probability of having character π t at time step t, π represents the entire prediction path. Multiplying the prediction results at each time step can obtain the probability p(π|x) of the entire prediction path π. π ∈ B -1 (l) represents the prediction path π that can obtain the final sequence l after the B transformation. Since the prediction path π contains the separator "-", only one character is considered inside two separators. For example, (hh-e-ll-l-oo) and (h-e-lll-ll-o) have the same length and are both decoded as (hello), which means that there are multiple prediction paths π that can obtain the final sequence l after the B transformation. Therefore, the probability p(l|x) of the prediction result being l is represented as the sum of the probabilities p(π|x) of all prediction paths π that can pass through the B transformation to obtain l. Finally, the objective function L ctc is the negative log-likelihood of the conditional probability p(l|x) of the label sequence l.

[0200] The specific formula of the sequence loss function L seq is as follows:

[0201]

[0202] k = (k 1 , k 2 , …, k |D| )(59)

[0203]

[0204] Here, l is the label, D is the dictionary sequence file, and y is the prediction result of the recognition model. ∑ |l| y represents the summation. The meaning of k is that if the guidance exists in the label, the guidance position value is 0; otherwise, it is 1. This can make the results in y that exist in the label not participate in the calculation, that is, the loss function L seq only targets the prediction results that do not appear in the label. Therefore, using the loss function L seq can suppress the results in y that do not exist in the label and can guide the recognition model to predict the correct results.

[0205] Embodiment 2

[0206] This embodiment provides a handwritten Chinese character recognition system based on text detection and text recognition, including:

[0207] An acquisition module, which is configured to: acquire a document image to be recognized;

[0208] A text detection module, which is configured to: input the document image to be recognized into a trained text detection model to obtain a text detection result; wherein, the trained text detection model first performs a preprocessing operation on the document image to be recognized, then performs feature extraction on the result of the preprocessing operation, then performs feature fusion on the extracted features, then performs a line regression operation and a probability map prediction on the fused features, and finally generates a text box based on the predicted probability map;

[0209] A text recognition module, which is configured to: input the text detection result into a trained text recognition model to obtain a text recognition result; perform a preprocessing operation on the text region image within the text box, perform an image augmentation operation on the preprocessed text region image, perform a block embedding operation on the augmented image, perform global fine-grained feature extraction on the image after the block embedding operation, and perform recognition on the extracted global fine-grained features to obtain a text recognition result.

[0210] It should be noted here that the above acquisition module, text detection module, and text recognition module correspond to steps S101 to S103 in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0211] Embodiment 3 This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in Embodiment 1 above.

[0212] Embodiment 4 This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in Embodiment 1 is completed.

[0213] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.< / pad>

Claims

1. A handwritten Chinese recognition method based on text detection and text recognition, characterized in that: include: Obtaining a document image to be recognized; The document image to be identified is input into the trained text detection model to obtain the text detection result; wherein the trained text detection model first performs a preprocessing operation on the document image to be identified, then extracts features from the preprocessing operation result, then fuses the extracted features, then performs a row regression operation and a probability map prediction on the fused features, and finally generates a text box based on the predicted probability map; The text detection result is input into the trained text recognition model to obtain the text recognition result; the text area image in the text box is preprocessed, the preprocessed text area image is image augmented, the augmented image is block embedded, the image after the block embedding is extracted from the global fine-grained feature, the extracted global fine-grained feature is recognized to obtain the text recognition result; The block embedding operation includes a block embedding module, and the block embedding module includes: First, the original image Input to the block embedding module, where , and are the height and width of the original image respectively; the original image After a step of 2 Convolutional layer, batch normalization layer and ReLu activation function layer acquisition ,in , is the number of channels of the block to be input into the global fine-grained feature extraction module; Then After another step of 2 Convolutional layer, batch normalization layer, and ReLu activation function layer to obtain block embedding ,in ; Finally Combined with the position embedding, we get the final block embedding , 。 2. The handwritten Chinese recognition method based on text detection and text recognition as claimed in claim 1, characterized in that: The trained text detection model includes: A first preprocessing module, a feature extraction module, a fusion pyramid module, a probability map prediction module and a post-processing module connected in sequence; the post-processing module includes: a binarization submodule and a text box generation submodule; wherein the output end of the fusion pyramid module is also connected to the input end of the row regression module; The first preprocessing module is used to preprocess the scanned document image and the label file: for the scanned document image, grayscale and binarization are performed on it, and the height and width are uniformly adjusted; Shrink each line of text area corresponding to the label inward by a set distance, using the original text line area Area ,perimeter and the scaling factor Calculate the shrinkage distance , shrink the boundary points of the original text area inward distance: Then, a blank image of the same size as the original image is generated, the scaled text area is marked in the blank image, and then the blank image is also binarized; The feature extraction module comprises: in, represents the first convolutional layer, represents the second convolutional layer, represents the third convolutional layer, represents the fourth convolutional layer, Represents the fifth convolutional layer; the image input to the feature extraction module is the grayscale image of the document image to be recognized after preprocessing ,in and are the height and width of the image, respectively. , , , , They are the output results of the five-layer convolution of ResNet-50, where , , , Feed into the fusion pyramid module.

3. The handwritten Chinese recognition method based on text detection and text recognition as claimed in claim 1, characterized in that: Fusion pyramid module, including: Features extracted from the last four layers of the feature extraction module , , and The number of channels is different. First, Convolution unifies their number of channels to , get the uniform channel number feature map , , and ; Then Perform 2 times upsampling, the upsampling strategy uses the nearest neighbor interpolation, and then Fusion is performed to obtain a sub-fusion feature map ; Then Upsample by 2, and Fusion obtains sub-fusion feature map ; Again Upsample by 2, and Fusion obtains sub-fusion feature map ; Afterwards, , , , Through a layer Convolution, then, , , Upsample to The same space size, finally and , , The upsampling results are sent to the connection layer and spliced ​​in the channel dimension to generate the final fusion feature map ; 。 4. The handwritten Chinese recognition method based on text detection and text recognition as claimed in claim 1, characterized in that: The row regression module comprises: The output of the fusion pyramid module As the input of the row regression module, it first passes through the first level with a height dimension step of 1 and a width dimension step of 2 Convolutional layer, batch normalization layer and ReLu activation function layer acquisition , whose width dimension is compressed to ,in is the width of the initial image; then, Pass the second level Convolutional layer, batch normalization layer and ReLu activation function layer acquisition , whose channel number dimension is compressed to ,in for The number of channels; Then, Then, the third level has a height dimension step of 1 and a width dimension step of 2. Convolutional layer, batch normalization layer and ReLu activation function layer acquisition , whose width dimension is compressed to ; at last, Get the text area height information through the width pooling module; The width pooling module first uses a global pooling operation to The width is compressed to 1, and then the compressed image is sent to Convolution, change the number of channels to , and then through a layer After the activation function, the width dimension and channel dimension are removed using the squeeze function, where squeeze is a function used to remove dimensions with a size of 1 in the tensor; Finally, the compressed result is sent to the linear layer to obtain The upper left corner height information and the lower right corner height information of the text area are stored in middle; 。 5. The handwritten Chinese recognition method based on text detection and text recognition as claimed in claim 1, characterized in that: The probability map prediction module comprises: The output of the fusion pyramid module As the input of the probability map prediction module, first pass The number of channels obtained by the convolution layer, batch normalization layer, and ReLu activation function layer is compressed to the original number of channels Feature map : Then, use pass The convolutional layer and batch normalization layer get ,in The number of channels is : Finally Send it to the variable resolution layer to get ,in The number of channels is stretched to 1, and its spatial resolution is the same as that of the document image to be recognized; Finally, the image output by the variable resolution layer Send it to the Sigmoid activation function layer to get the probability map , where each pixel value is in the range of 0-1: ; The document image to be recognized is input into the trained text detection model to obtain the text detection result; the training process of the trained text detection model includes: Constructing a first training set, wherein the first training set is a document image of a known text detection area; Inputting the first training set into the text detection model, training the text detection model, and stopping the training when the loss function value of the text detection model no longer decreases or when the iteration reaches a set number of times, thereby obtaining a trained text detection model; The loss function of the text detection model includes: in, is the cross entropy loss between the post-processing module result and the original unscaled label, The height information of the text area predicted by the line regression module is the difference between the original unscaled label loss, is the weight parameter, the overall loss function Drive text detection model training; in, The specific formula is as follows: in, Representation sample The label value of For sample The probability of predicting the positive class; The specific formula is as follows: in, Represents the predicted value The difference between the upper left corner height information and the lower right corner height information of the actual text area.

6. The handwritten Chinese recognition method based on text detection and text recognition as claimed in claim 1, characterized in that: The text detection result is input into the trained text recognition model to obtain the text recognition result, wherein the trained text recognition model includes: A second preprocessing module, an image augmentation module, a block embedding module, a global fine-grained feature extraction module, and an output head module connected in sequence; The second preprocessing module includes: firstly graying the text line image, then scaling its width and height in equal proportion until the height is the same as the set height, and if the image width is greater than the set width, directly compressing the width dimension; otherwise, filling blanks on the right side of the image until the set width is reached; The image augmentation module includes: an image random cropping operation, an image random rotation operation and an image Gaussian blur operation.

7. The handwritten Chinese recognition method based on text detection and text recognition as claimed in claim 1, characterized in that: The global fine-grained feature extraction module includes: a first stage, a second stage, a third stage and a high pooling layer connected in sequence; The first stage includes: Feature extraction submodules connected in series; The second stage includes: a first downsampling layer and a feature extraction submodule; The third stage includes: a second downsampling layer and a feature extraction submodule; Among them, the feature extraction submodule adopts stacked group convolution and self-attention mechanism, which enables the model to capture features of different granularities; the specific process is to first send the block embedding module or the downsampling layer output image to the feature extraction module, and the feature extraction submodule contains submodules that are alternately connected with stacked group convolution and self-attention; the image is reshaped into a vector form after group convolution, and then reshaped into the original image format after multiple sets of self-attention layers. If it is in the first or second stage, it is sent to the downsampling layer of the next stage to downsample the height; otherwise, it is sent to the height pooling layer to compress the height to 1; For the image input into the feature extraction submodule , first send it to the grouped convolutional layer with layer normalization. The specific steps are to normalize the features obtained by layer In the spatial dimension, Each copy is processed separately Convolution operation, finally concatenate the results of each group of convolutions, and introduce skip connections to obtain : Afterwards, It is sent to a multi-layer perceptron with layer normalization, and the skip connection is introduced again to obtain the features extracted by group convolution. : Next, yes Perform a reshape operation to convert it into a vector form , where reshape is a function used to change the shape of a tensor. The main operation here is to The last two dimensions of ; Afterwards Feed into the layer normalization module: Afterwards Send it to the self-attention layer to obtain dependencies, and then introduce skip connections to generate intermediate variables ; A multi-head self-attention mechanism is used in the self-attention layer, and the process is as follows: in, for The dimension of is the number of self-attention heads, , , , ,in for The number of channels, for Dimensions, and there are ; Finally, It is sent to a normalized multi-layer perceptron, and the skip connection is introduced again, and finally the reshape operation is performed to obtain : Next, if If it is not in the last stage, it is sent to the downsampling layer of the next stage to downsample the height; First, it passes through a step size of 2 in the height dimension and 1 in the width dimension. Convolution, which will halve the height dimension; then it is fed into the layer normalization to get , and will be sent to the next stage: in, , They are The height and width of is the number of channels input to the next stage; if At the last stage, it is sent to the height pooling layer. The height pooling layer first compresses the image height to 1 through a global pooling operation, and then sends the compressed image to Convolution, change the number of channels to Finally, after a layer of activation function, the output result of the global fine-grained feature extraction module is obtained. : in, for width, then Send it to the output head module to get the prediction result of the recognition model; The output head module includes: using global fine-grained feature extraction output As input, we first use the squeeze operation to compress the height dimension, and then pass it through a linear layer to get , the number of channels is , then use CTC decoding to align the model output results with the label by calculating the CTC loss; finally, the text sequence is obtained : 。 8. The handwritten Chinese recognition method based on text detection and text recognition as claimed in claim 1, characterized in that: The text detection result is input into the trained text recognition model to obtain the text recognition result. The training process of the trained text recognition model includes: Constructing a second training set, wherein the second training set is text region images of known text recognition results; Input the second training set into the text recognition model to train the text recognition model. When the loss function value of the text recognition model no longer decreases or the iteration reaches a set number of times, stop the training to obtain a trained text detection model. Loss function of text recognition model as follows: in is the CTC decoding loss, is the weight parameter, is the sequence loss function, the overall loss function Drive recognition model training; The specific formula is as follows: in, Indicates that at time step With characters The probability of Represents the entire prediction path, and the prediction results of each time step Multiply them together to get the entire predicted path Probability ; It means after After transformation, the final sequence is obtained The predicted path The prediction result is Probability is represented by all the Transformation The predicted path Probability The sum of; finally, the objective function It is the tag sequence The conditional probability The negative log-likelihood of ; Sequence loss function The specific formula is as follows: in, For labels, is a dictionary sequence file, To identify the model prediction results; It represents peace; The meaning is that if the guide exists in the label, the guide position value is 0, otherwise it is 1.

9. A handwritten Chinese recognition system based on text detection and text recognition, characterized in that: include: An acquisition module is configured to: acquire a document image to be recognized; The text detection module is configured to: input the document image to be identified into the trained text detection model to obtain the text detection result; wherein the trained text detection model first performs a preprocessing operation on the document image to be identified, then performs feature extraction on the preprocessing operation result, then performs feature fusion on the extracted features, then performs row regression operation and probability map prediction on the fused features, and finally generates a text box based on the predicted probability map; The text recognition module is configured to: input the text detection result into the trained text recognition model to obtain the text recognition result; perform a preprocessing operation on the text area image in the text box, perform image augmentation processing on the preprocessed text area image, perform a block embedding operation on the augmented image, perform global fine-grained feature extraction on the image after the block embedding operation, and recognize the extracted global fine-grained features to obtain the text recognition result; The block embedding operation includes a block embedding module, and the block embedding module includes: First, the original image Input to the block embedding module, where , and are the height and width of the original image respectively; the original image After a step of 2 Convolutional layer, batch normalization layer and ReLu activation function layer acquisition ,in , is the number of channels of the block to be input into the global fine-grained feature extraction module; Then After another step of 2 Convolutional layer, batch normalization layer, and ReLu activation function layer to obtain block embedding ,in ; Finally Combined with the position embedding, we get the final block embedding , 。

Citation Information

Patent Citations

  • Video character end-to-end detection and recognition method based on deep learning

    CN113361432A

  • Traffic text detection and recognition method and device, equipment and storage medium

    CN117593733A