Model Training Method, Text Recognition Method, Apparatus, Medium and Device

By combining text recognition models connecting the timing classification network and attention mechanism network, the problem of low text recognition accuracy in OCR in educational scenarios is solved, and a more efficient text recognition effect is achieved.

CN114973267BActive Publication Date: 2025-07-11BEIJING ZHIYUAN HANGCHENG SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210615726.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-07-11
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

The prior art text recognition method of OCR in educational scenarios cannot effectively solve the problems of affine transformation, large scale changes, text bends, background interference, font variations, insufficient lighting and blurred shooting in the image of the photographed document, resulting in a low recognition accuracy.

Method used

A text recognition model combining a connection timing classification network and an attention mechanism network is adopted to extract feature map samples through a high-resolution network, and perform sequence conversion and weighted calculations. The connection timing classification network is used to force monotonous alignment of input and output, avoiding the attention mechanism network ignoring the overall information and improving the recognition accuracy.

Benefits of technology

Taking into account the alignment of overall information and input and output, the accuracy of text recognition is improved, the calculation overhead is reduced, and the model training time is shortened.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973267B_ABST
    Figure CN114973267B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a model training method, a text recognition method, an apparatus, a medium, and a device to solve the problems existing in the related art. The model training method includes: obtaining a sample image, where the sample image includes a sample text, and the sample text is labeled with a sample text label; inputting the sample image into a text recognition model to obtain a sample text recognition result output by the text recognition model, where the text recognition model includes a connectionist temporal classification network and an attention mechanism network, and the sample text recognition result includes a first sample recognition result output by the connectionist temporal classification network and a second sample recognition result output by the attention mechanism network; calculating a loss value according to the first sample recognition result, the second sample recognition result, and the sample text label; and adjusting the parameters of the text recognition model according to the loss value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular, to a model training method, a text recognition method, an apparatus, a medium, and a device. Background Art

[0002] OCR (Optical Character Recognition) is one of the important directions in computer vision. OCR can perform character recognition on objects such as scanned documents, and can also perform character recognition on natural scenes, that is, scene text recognition (STR). OCR in the education scenario usually performs character recognition on images such as test papers and scanned documents taken by mobile devices. Among them, the photographed document images have problems such as affine transformation, large scale change, text bending, background interference, variable fonts, insufficient illumination, blurred shooting, and multiple languages, which pose great technical challenges to text recognition. The methods for text recognition in the education scenario in the related art cannot well solve these problems, so the accuracy of text recognition is relatively low. Summary of the Invention

[0003] The purpose of the present disclosure is to provide a model training method, a text recognition method, an apparatus, a medium, and a device to solve the problems existing in the related art.

[0004] To achieve the above purpose, according to the first aspect of the embodiments of the present disclosure, a model training method is provided, and the method includes:

[0005] Obtain a sample image, where the sample image includes a sample text, and the sample text is labeled with a sample text label;

[0006] Input the sample image into a text recognition model to obtain a sample text recognition result output by the text recognition model, where the text recognition model includes a connectionist temporal classification network and an attention mechanism network, and the sample text recognition result includes a first sample recognition result output by the connectionist temporal classification network and a second sample recognition result output by the attention mechanism network;

[0007] Calculate a loss value according to the first sample recognition result, the second sample recognition result, and the sample text label;

[0008] Adjust the parameters of the text recognition model according to the loss value.

[0009] Optionally, the text recognition model further includes a high-resolution network. Inputting the sample image into the text recognition model to obtain the first sample text recognition result and the second sample recognition result output by the text recognition model includes:

[0010] Input the sample image into the high-resolution network to obtain an initial feature map sample;

[0011] Determine a text contour feature map sample, a text direction feature map sample, and a text character category feature map sample according to the initial feature map sample;

[0012] Perform a sequence conversion operation on the text contour feature map sample, the text direction feature map sample, and the text character category feature map sample to obtain a feature vector sample sequence;

[0013] Input the feature vector sample sequence into the attention mechanism network and the connectionist temporal classification network to obtain a first character category distribution probability sample of each feature vector sample in the feature vector sample sequence output by the connectionist temporal classification network, and obtain a second character category distribution probability sample of each feature vector sample in the feature vector sample sequence output by the attention mechanism network;

[0014] Obtain the first sample recognition result according to the first character category distribution probability sample, and obtain the second sample recognition result according to the second character category distribution probability sample.

[0015] Optionally, the attention mechanism network includes an encoder and a decoder, and the inputting the feature vector sample sequence into the attention mechanism network and the connectionist temporal classification network includes:

[0016] Input the feature vector sample sequence into the encoder of the attention mechanism network to obtain an encoded feature vector sample sequence;

[0017] Input the encoded feature vector sample sequence into the decoder of the attention mechanism network and the connectionist temporal classification network.

[0018] Optionally, the method further includes:

[0019] Determine a text boundary offset feature map sample according to the initial feature map sample, and the text boundary offset feature map sample represents the position of the sample text recognition result in the sample image.

[0020] Optionally, the training parameter includes a loss weight value, and the loss value is calculated by the following formula:

[0021] L MTL = λL CTC +(1 - λ)L Attention ;

[0022] where, L MTL is the loss value, L CTCis the loss function of the connection timing classification network, L Attention is the loss function of the attention mechanism network, and λ is the loss weight value.

[0023] Optionally, the text recognition model is a fully convolutional point aggregation network model.

[0024] According to the second aspect of the embodiments of the present disclosure, there is provided a text recognition method, the method comprising:

[0025] Obtain an image to be recognized, where the image to be recognized includes text to be recognized;

[0026] Input the image to be recognized into a text recognition model to obtain a text recognition result output by the text recognition model, where the text recognition model is trained by the model training method described in any one of the first aspect.

[0027] Optionally, the text recognition result includes a first recognition result and a second recognition result, and the method further comprises:

[0028] Perform weighted calculation according to the first recognition result and the second recognition result to obtain a target text recognition result.

[0029] According to the third aspect of the embodiments of the present disclosure, there is provided a model training device, the device comprising:

[0030] A first acquisition module, configured to acquire a sample image, where the sample image includes sample text, and the sample text is labeled with a sample text label;

[0031] A first input module, configured to input the sample image into a text recognition model to obtain a sample text recognition result output by the text recognition model, where the text recognition model includes a connection timing classification network and an attention mechanism network, and the sample text recognition result includes a first sample recognition result output by the connection timing classification network and a second sample recognition result output by the attention mechanism network;

[0032] A calculation module, configured to calculate a loss value according to the first sample recognition result, the second sample recognition result, and the sample text label;

[0033] An adjustment module, configured to adjust parameters of the text recognition model according to the loss value.

[0034] According to the fourth aspect of the embodiments of the present disclosure, there is provided a text recognition device, the device comprising:

[0035] A second acquisition module, configured to acquire an image to be recognized, where the image to be recognized includes text to be recognized;

[0036] A second input module, configured to input the image to be recognized into a text recognition model, and obtain a text recognition result output by the text recognition model, where the text recognition model is trained by the model training method described in any item of the first aspect.

[0037] According to a fifth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any item of the first aspect or any item of the second aspect are implemented.

[0038] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0039] A memory, on which a computer program is stored;

[0040] A processor, configured to execute the computer program in the memory to implement the steps of the method described in any item of the first aspect or any item of the second aspect.

[0041] Through the above technical solution, a sample image is input into a text recognition model to obtain a sample text recognition result. The text recognition model includes a connectionist temporal classification network and an attention mechanism network, and the sample text recognition result includes a first sample recognition result output by the connectionist temporal classification network and a second sample recognition result output by the attention mechanism network. Since the loss value is calculated according to the first sample recognition result output by the connectionist temporal classification network and the second sample recognition result output by the attention mechanism network, the connectionist temporal classification network can be used to force monotonic alignment between the input and the output, and the attention mechanism network can be used to avoid ignoring the overall information when the connectionist temporal classification network predicts local information. Therefore, text recognition is performed while considering both the overall information and the alignment of the input and output, improving the accuracy of text recognition.

[0042] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification. They are used to explain the present disclosure together with the following specific implementation, but do not constitute a limitation to the present disclosure. In the drawings:

[0044] Figure 1 is a flowchart of a model training method shown in an exemplary embodiment of the present disclosure.

[0045] Figure 2 is a flowchart of a text recognition method shown in an exemplary embodiment of the present disclosure.

[0046] Figure 3It is a schematic diagram of text recognition results shown in an exemplary embodiment of the present disclosure.

[0047] Figure 4 It is a block diagram of a model training device shown in an exemplary embodiment of the present disclosure.

[0048] Figure 5 It is a block diagram of a text recognition device shown in an exemplary embodiment of the present disclosure.

[0049] Figure 6 It is a block diagram of an electronic device shown in an exemplary embodiment of the present disclosure. Detailed implementation manners

[0050] The following will describe in detail the specific implementation manners of the present disclosure with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the present disclosure, and are not used to limit the present disclosure.

[0051] It should be noted that all actions of obtaining signals, information or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located, and with the authorization given by the owner of the corresponding device.

[0052] The related technology for OCR in the education scenario involves traditional methods and deep learning methods, where deep learning includes a single-stage method (i.e., end-to-end text detection and recognition) and a two-stage method (i.e., text detection and text recognition).

[0053] Traditional OCR methods rely on image processing technology and statistical machine learning methods. Usually, manual feature extraction methods (such as SWT, MSER, etc.) are used for text detection, and then template matching or model training methods are used to recognize the detected text. Traditional OCR methods can be divided into three stages: image preprocessing, text recognition, and post-processing. Among them, image preprocessing can be used to complete text region localization, text correction, character segmentation, etc. Text recognition can be used to recognize the segmented text. Usually, manually designed features (such as HOG features, etc.) or CNN can be used to extract features, and then a machine learning classifier (such as SVM) is used for recognition. Post-processing can use preset rules, language models, etc. to correct the recognition results for layout restoration, recognition correction, etc. However, traditional methods have poor adaptability and anti-interference ability to text shape changes (such as text blur, stroke adhesion, broken strokes, uneven black and white, and ink back penetration, etc.), and cannot adapt to text recognition in different scenarios (when traditional methods face different scenarios, they usually need to independently design the parameters of each module, and it is difficult to design a model with good generalization performance when facing complex scenarios).

[0054] OCR using deep learning methods usually employs convolutional neural networks to replace manual feature extraction methods for text detection, and then uses neural networks to recognize the detected text. The basic idea of the single-stage approach lies in designing a model that has both a detection unit and a recognition unit. The detection unit and the recognition unit share CNN features (i.e., features extracted by convolutional neural networks) and are jointly trained. In the model application stage, the end-to-end recognition model can predict the position and content information of the text in the image in a single forward pass. However, the training of the single-stage approach often requires character-level annotations (such as Mask textSpotter and CharNet), which is costly. And the single-stage approach usually predefines the text reading direction (such as TextDragon and Mask textSpotter strongly assume that the reading direction of the text region is: from left to right or from top to bottom), so the recognition effect for texts with non-traditional reading directions is not good.

[0055] The two-stage approach splits text detection and text recognition into two parts. It locates text lines through text detection and then recognizes the content of the located text lines through text recognition. For text detection, methods based on regression or segmentation can be used. Among them, the regression-based method has poor detection effects for irregularly shaped texts or sparsely distributed texts (for example, CTPN has poor detection effects for tilted and curved texts, and SegLink has poor detection effects for sparsely distributed texts). The post-processing of the segmentation-based algorithm is complex, has operation performance problems, and has poor detection effects for overlapping texts. Text recognition can include regular text recognition (such as recognizing texts on the horizontal line like printed fonts and scanned texts) and irregular text recognition (i.e., recognizing texts not on the horizontal line). For regular text recognition, network models such as CTC or Sequence2Sequence can be used for text recognition. For irregular text recognition, network models such as STAR-Net, RARE, and Transformer can be used for text recognition. However, the two-stage approach tunes multiple stages during training and involves time-consuming operations such as non-maximum suppression (NMS) and ROI (region of interest), which brings non-negligible computational overhead and affects the training effect of the model.

[0056] In view of this, the present disclosure provides a model training method, a text recognition method, an apparatus, a medium, and a device. The text recognition model is trained in a single-stage manner to avoid the computational overhead caused by the two-stage manner, and a connectionist temporal classification network and an attention mechanism network are used to decode the feature map. Thus, the connectionist temporal classification network can be used to enforce monotonic alignment between the input and the output, and the attention mechanism network can be used to avoid ignoring the overall information when the connectionist temporal classification network predicts local information. In this way, text recognition can be performed while considering both the overall information and the alignment between the input and the output, improving the accuracy of text recognition.

[0057] The technical solutions of the present disclosure will be described in detail with reference to the following embodiments.

[0058] Figure 1 FIG. is a flowchart of a model training method shown in an exemplary embodiment of the present disclosure. As Figure 1 shown, the method includes:

[0059] S101, obtaining a sample image, where the sample image includes a sample text, and the sample text is labeled with a sample text label.

[0060] Among them, the sample image may be an image such as a test paper or a scanned document photographed by a mobile device in an educational scenario, and these images may include a sample text. The sample text label may be the ground truth of the text in the sample image (i.e., the sample text in the sample image), which may be manually labeled before model training or labeled by a machine learning method in related technologies.

[0061] S102, inputting the sample image into the text recognition model to obtain a sample text recognition result output by the text recognition model, where the text recognition model includes a connectionist temporal classification network and an attention mechanism network, and the sample text recognition result includes a first sample recognition result output by the connectionist temporal classification network and a second sample recognition result output by the attention mechanism network.

[0062] Among them, the connectionist temporal classification network can be a CTC (Connectionist Temporal Classification) network, and the attention mechanism network can be an Attention network. The CTC network can include a forward-backward algorithm, which can enforce monotonic alignment of the input and output, so that a stable alignment effect can be ensured even when there is too much noise in the input data. However, the CTC network assumes that the outputs at different time steps are independent, and each output is the probability of an independent single character, which leads to the neglect of the overall information. The Attention network does not introduce any alignment constraints. When aligning, it selects the target input corresponding to the output from all inputs. Decoding in this case may cause misalignment between the input and the output.

[0063] The model training method provided by the embodiments of the present disclosure combines a connectionist temporal classification network and an attention mechanism network, so that the connectionist temporal classification network can be used to enforce monotonic alignment of the input and output, and the attention mechanism network can be used to avoid the neglect of the overall information when the connectionist temporal classification network predicts local information. Thus, text recognition is performed while considering both the overall information and the alignment of the input and output, improving the accuracy of text recognition. Moreover, since the alignment time consumption of the attention mechanism network is reduced, the model convergence can be accelerated and the model accuracy can be improved.

[0064] S103, calculate a loss value according to the first sample recognition result, the second sample recognition result, and the sample text label.

[0065] S104, adjust the parameters of the text recognition model according to the loss value.

[0066] It is not difficult to understand that by calculating the loss value through the first sample recognition result output by the connectionist temporal classification network and the second sample recognition result output by the attention mechanism network, and adjusting the parameters of the text recognition model according to the loss value, the text recognition model can perform text recognition while considering both the overall information and the alignment of the input and output, improving the accuracy of text recognition.

[0067] Through the above technical solution, the sample image is input into the text recognition model to obtain the sample text recognition result. The text recognition model includes a connectionist temporal classification network and an attention mechanism network. The sample text recognition result includes a first sample recognition result output by the connectionist temporal classification network and a second sample recognition result output by the attention mechanism network. Since the loss value is calculated based on the first sample recognition result output by the connectionist temporal classification network and the second sample recognition result output by the attention mechanism network, the connectionist temporal classification network can be used to force the monotonic alignment of the input and output, and the attention mechanism network can be used to avoid ignoring the overall information when the connectionist temporal classification network predicts local information. Therefore, text recognition can be performed while considering both the overall information and the alignment of the input and output, improving the accuracy of text recognition.

[0068] Optionally, the text recognition model is a fully convolutional point aggregation network model.

[0069] The fully convolutional point aggregation network model may refer to PGNet (Point Gathering Network), which is an end-to-end character recognition model.

[0070] It should be noted that PGNet in the related art is mainly applied to English OCR in natural scenes for detecting and recognizing English and numbers. When PGNet is used for English OCR, a total of 37 text character categories are set (26 English letters, 10 numbers, and 1 background). In this case, PGNet uses ResNet for image feature extraction. In the OCR recognition of Chinese-related educational scenarios, since there are more than 6,000 commonly used Chinese characters, it is more difficult for the model to extract image features. On this basis, the effect of using ResNet to extract image features is not good. This is because there are regional and pixel-level problems in visual feature extraction. Classification networks such as ResNet and VGGNet learn low-resolution representations, so the discrimination of the restored high-resolution representation space is not strong enough. This makes it difficult for ResNet to obtain accurate prediction results in tasks sensitive to spatial accuracy. Therefore, in addition to improving the decoding network of PGNet in the related art (that is, combining the connectionist temporal classification network and the attention mechanism network for decoding), the present disclosure also improves the network for image feature extraction of PGNet.

[0071] Optionally, the text recognition model further includes a high-resolution network. The steps of inputting the sample image into the text recognition model to obtain the first sample text recognition result and the second sample text recognition result output by the text recognition model may include:

[0072] Input the sample image into the high-resolution network to obtain an initial feature map sample;

[0073] Determine a text contour feature map sample, a text direction feature map sample, and a text character category feature map sample according to an initial feature map sample;

[0074] Perform a sequence conversion operation on the text contour feature map sample, the text direction feature map sample, and the text character category feature map sample to obtain a feature vector sample sequence;

[0075] Input the feature vector sample sequence into an attention mechanism network and a connectionist temporal classification network to obtain a first character category distribution probability sample for each feature vector sample in the feature vector sample sequence output by the connectionist temporal classification network, and obtain a second character category distribution probability sample for each feature vector sample in the feature vector sample sequence output by the attention mechanism network;

[0076] Obtain a first sample recognition result according to the first character category distribution probability sample, and obtain a second sample recognition result according to the second character category distribution probability sample.

[0077] Among them, the high-resolution network can be HRNet (High-Resolution Network). HRNet makes a fundamental change to the network structure, improving the traditional serial connection of high and low-resolution convolutions to a parallel connection of high and low-resolution convolutions, so that high resolution can be maintained during convolution and rich high-resolution representations and accurate spatial information can be learned through multiple information exchanges of high and low-resolution representations, obtaining a feature map with stronger representation ability and higher resolution.

[0078] It is not difficult to understand that the embodiments of the present disclosure use HRNet for image feature extraction. Compared with ResNet, HRNet can obtain a feature map with stronger representation ability and higher resolution.

[0079] It should be noted that after extracting features from the sample image through a high-resolution network, an initial feature map sample can be obtained. Based on this initial feature map sample, a text contour feature map sample, a text direction feature map sample, a text character category feature map sample, and a text boundary offset feature map sample can be obtained through parallel multi-task learning. Among them, the text contour feature map sample can be a 1-channel feature map, used to represent the text centerline (TCL, text centerline). The text direction feature map sample can be a 2-channel feature map, used to represent the text direction offset (TDO, text direction offset), where the text direction offset can refer to the offset of each pixel of the TCL in the text contour feature map to the next text reading position. The text character category feature map sample can be an n-channel feature map (n is the number of character categories, for example, n is 37 in English OCR), used to represent the character classification (TCC, text character classification). The text boundary offset feature map sample can be a 4-channel map, used to represent the position of the sample text recognition result in the sample image, where the text boundary offset (TBO, text border offset) can refer to the offset of each pixel of the TCL in the text contour feature map from the upper and lower boundary points of the text area.

[0080] After obtaining the text contour feature map sample, the text direction feature map sample, and the text character category feature map sample, the added feature map (channel numbers added) of the text contour feature map sample, the text direction feature map sample, and the text character category feature map sample can be calculated, and a sequence conversion operation (Map-to-Seq, an operation to convert the feature map into a feature vector sequence) can be performed on this added feature map to obtain a feature vector sample sequence. On this basis, the feature vector sample sequence can be input into the attention mechanism network and the connectionist temporal classification network to obtain the first character category distribution probability sample of each feature vector sample in the feature vector sample sequence output by the connectionist temporal classification network, and the second character category distribution probability sample of each feature vector sample in the feature vector sample sequence output by the attention mechanism network. Among them, the character category distribution probability sample is used to represent the probability sample of the feature vector sample corresponding to each character category.

[0081] It can be understood that the character category with the highest probability in the character category distribution probability sample of the feature vector can be the predicted character corresponding to the feature vector (i.e., the sample recognition result). On this basis, the first sample recognition result can be obtained according to the first character category distribution probability sample, and the second sample recognition result can be obtained according to the second character category distribution probability sample. Moreover, the first sample recognition result and the second sample recognition result can be weighted and calculated to obtain the target sample recognition result. The target sample recognition result can be used to represent the sample text of the sample image. The weights for the weighted calculation of the first sample recognition result and the second sample recognition result can be the training parameters of the model, and the training parameters can be adjusted according to the results of the iterative training of the model.

[0082] Optionally, the technical solution provided by the embodiments of the present disclosure may further include:

[0083] Determine a text boundary offset feature map sample according to the initial feature map sample, where the text boundary offset feature map sample represents the position of the sample text recognition result in the sample image.

[0084] It should be noted that the sample recognition result may include the sample text and the sample text box, and the sample text box can represent the position of the sample text in the sample image. It is not difficult to understand that the sample text box can be obtained according to the text boundary offset feature map sample.

[0085] In addition, it should also be noted that the text contour feature map sample, the text direction feature map sample, and the text boundary offset feature map sample can be supervised and learned by the label feature map of the same scale, so that the text recognition model can predict more accurate text contour feature map samples, text direction feature map samples, and text boundary offset feature map samples. Moreover, since the training text recognition model learns the text direction feature map sample, the recognition of non-traditional reading directions by the model can be extended, and the accuracy of the model for text recognition can be improved.

[0086] Optionally, the attention mechanism network includes an encoder and a decoder. The steps of inputting the feature vector sample sequence into the attention mechanism network and the connectionist temporal classification network may include:

[0087] Input the feature vector sample sequence into the encoder of the attention mechanism network to obtain an encoded feature vector sample sequence;

[0088] Input the encoded feature vector sample sequence into the decoder of the attention mechanism network and the connectionist temporal classification network.

[0089] It should be noted that in the embodiments of the present disclosure, by enabling the connection timing classification network to share the encoder of the attention mechanism network, and inputting the encoded feature vector sample sequence encoded by the encoder of the attention mechanism network into the decoder of the attention mechanism network and the connection timing classification network, the connection timing classification network and the attention mechanism network can be combined. On this basis, the first sample recognition result output by the connection timing classification network and the second sample recognition result output by the attention mechanism network can be combined in a beam search algorithm to eliminate irregular alignment. In this way, the problem that the connection timing classification network ignores the overall information when predicting the local information in the first sample recognition result (i.e., the prediction of the probability of a single character), and the problem that the decoding of the attention mechanism network is not restricted by alignment can be effectively solved. Moreover, since the alignment time consumption of the attention mechanism network is reduced, the model convergence can be accelerated and the model accuracy can be improved.

[0090] Optionally, the training parameters include loss weight values, and the loss value can be calculated by the following formula:

[0091] L MTL = λL CTC +(1 - λ)L Attention ;

[0092] Where L MTL is the loss value, L CTC is the loss function of the connection timing classification network, L Attention is the loss function of the attention mechanism network, and λ is the loss weight value.

[0093] Among them, the loss weight value is a training parameter of the text recognition model and can be adjusted according to the results of model iterative training. The loss function of the connection timing classification network can be the loss function of the CTC network, and the loss function of the attention mechanism network can be the loss function of the Attention network. The loss functions of the CTC network and the Attention network are prior arts and will not be elaborated here.

[0094] It should be noted that the pixel-level text character category feature map samples can be obtained through training with L MTL , so that the text recognition model can get rid of character-level annotation and NMS and ROI operations.

[0095] By calculating the loss value using the above formula to adjust the training parameters of the model, it is possible to use the connection timing classification network to enforce the monotonic alignment of input and output, and at the same time, avoid the attention mechanism network from ignoring the overall information when the connection timing classification network predicts local information, so as to perform text recognition while considering both the overall information and the alignment of input and output, and improve the accuracy of text recognition.

[0096] Through the above technical solution, the sample image is input into the text recognition model to obtain the sample text recognition result. The text recognition model includes a connectionist temporal classification network and an attention mechanism network. The sample text recognition result includes a first sample recognition result output by the connectionist temporal classification network and a second sample recognition result output by the attention mechanism network. Since the loss value is calculated based on the first sample recognition result output by the connectionist temporal classification network and the second sample recognition result output by the attention mechanism network, the connectionist temporal classification network can be used to force the monotonic alignment of the input and output, and the attention mechanism network can be used to avoid the neglect of the overall information when the connectionist temporal classification network predicts local information. Therefore, text recognition can be performed while considering both the overall information and the alignment of the input and output, improving the accuracy of text recognition.

[0097] Figure 2 is a flowchart of a text recognition method shown in an exemplary embodiment of the present disclosure. As Figure 2 shown, the method includes:

[0098] S201, obtain the image to be recognized, where the image to be recognized includes the text to be recognized.

[0099] S202, input the image to be recognized into the text recognition model to obtain the text recognition result output by the text recognition model, where the text recognition model is trained by the above model training method.

[0100] It is not difficult to understand that the text recognition model trained by the above model training method can avoid the computational overhead brought by the two-stage method, reduce the time consumption, and improve the accuracy of text recognition. On this basis, by inputting the image to be recognized into the text recognition model, a text recognition result with higher accuracy can be obtained.

[0101] Optionally, the text recognition result may include a first recognition result and a second recognition result. On this basis, the technical solution provided by the embodiments of the present disclosure may further include:

[0102] Perform weighted calculation according to the first recognition result and the second recognition result to obtain the target text recognition result.

[0103] It should be noted that after inputting the image to be recognized into the text recognition model, the first recognition result can be obtained according to the connectionist temporal classification network of the text recognition model, and the second recognition result can be obtained according to the attention mechanism network of the text recognition model. On this basis, weighted calculation can be performed according to the first recognition result and the second recognition result to obtain the target text recognition result, which can be used to represent the text to be recognized in the image to be recognized.

[0104] It should also be noted that the target text recognition result may further include a recognized text box, which may represent the position of the text to be recognized in the image to be recognized.

[0105] Figure 3 It is a schematic diagram of a text recognition result shown in an exemplary embodiment of the present disclosure. As Figure 3 shown, the text recognition result may include a recognized text box and recognized text (such as Figure 3 the Chinese text and English text in the recognized text box shown).

[0106] The text recognition model is trained by the above model training method. Since the loss value is calculated according to the first sample recognition result output by the connectionist temporal classification network and the second sample recognition result output by the attention mechanism network during training, the connectionist temporal classification network can be used to enforce the monotonic alignment of the input and output, and the attention mechanism network can be used to avoid the neglect of the overall information when the connectionist temporal classification network predicts local information. Therefore, text recognition can be performed while considering both the overall information and the alignment of the input and output, improving the accuracy of text recognition.

[0107] Based on the same inventive concept, the present disclosure also provides a model training device. Refer to Figure 4 , Figure 4 It is a block diagram of a model training device shown in an exemplary embodiment of the present disclosure. As Figure 4 shown, the model training device 100 includes:

[0108] A first acquisition module 101, configured to acquire a sample image, where the sample image includes a sample text, and the sample text is labeled with a sample text label;

[0109] A first input module 102, configured to input the sample image into the text recognition model to obtain a sample text recognition result output by the text recognition model, where the text recognition model includes a connectionist temporal classification network and an attention mechanism network, and the sample text recognition result includes a first sample recognition result output by the connectionist temporal classification network and a second sample recognition result output by the attention mechanism network;

[0110] A calculation module 103, configured to calculate a loss value according to the first sample recognition result, the second sample recognition result, and the sample text label;

[0111] An adjustment module 104, configured to adjust parameters of the text recognition model according to the loss value.

[0112] Through the above technical solution, the sample image is input into the text recognition model to obtain the sample text recognition result. The text recognition model includes a connectionist temporal classification network and an attention mechanism network. The sample text recognition result includes a first sample recognition result output by the connectionist temporal classification network and a second sample recognition result output by the attention mechanism network. Since the loss value is calculated based on the first sample recognition result output by the connectionist temporal classification network and the second sample recognition result output by the attention mechanism network, the connectionist temporal classification network can be used to force the monotonic alignment of the input and output, and the attention mechanism network can be used to avoid ignoring the overall information when the connectionist temporal classification network predicts local information. Therefore, text recognition can be performed while considering both the overall information and the alignment of the input and output, improving the accuracy of text recognition.

[0113] Optionally, the text recognition model further includes a high-resolution network. The first input module 102 is configured to:

[0114] Input the sample image into the high-resolution network to obtain an initial feature map sample;

[0115] Determine a text contour feature map sample, a text direction feature map sample, and a text character category feature map sample according to the initial feature map sample;

[0116] Perform a sequence conversion operation on the text contour feature map sample, the text direction feature map sample, and the text character category feature map sample to obtain a feature vector sample sequence;

[0117] Input the feature vector sample sequence into the attention mechanism network and the connectionist temporal classification network to obtain a first character category distribution probability sample of each feature vector sample in the feature vector sample sequence output by the connectionist temporal classification network, and obtain a second character category distribution probability sample of each feature vector sample in the feature vector sample sequence output by the attention mechanism network;

[0118] Obtain the first sample recognition result according to the first character category distribution probability sample, and obtain the second sample recognition result according to the second character category distribution probability sample.

[0119] Optionally, the attention mechanism network includes an encoder and a decoder. The first input module 102 is configured to:

[0120] Input the feature vector sample sequence into the encoder of the attention mechanism network to obtain an encoded feature vector sample sequence;

[0121] Input the encoded feature vector sample sequence into the decoder of the attention mechanism network and the connectionist temporal classification network.

[0122] Optionally, the apparatus 100 further includes:

[0123] A determination module, configured to determine a text boundary offset feature map sample according to the initial feature map sample, where the text boundary offset feature map sample characterizes the position of the sample text recognition result in the sample image.

[0124] Optionally, the training parameter includes a loss weight value, and the loss value is calculated by the following formula:

[0125] L MTL = λL CTC + (1 - λ)L Attention ;

[0126] where L MTL is the loss value, L CTC is the loss function of the connection time series classification network, L Attention is the loss function of the attention mechanism network, and λ is the loss weight value.

[0127] Optionally, the text recognition model is a fully convolutional point aggregation network model.

[0128] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0129] Based on the same inventive concept, the present disclosure further provides a model training apparatus. Refer to Figure 5 , Figure 5 , which is a block diagram of a text recognition apparatus shown in an exemplary embodiment of the present disclosure. As Figure 5 shown, the text recognition apparatus 200 includes:

[0130] A second acquisition module 201, configured to acquire an image to be recognized, where the image to be recognized includes text to be recognized;

[0131] A second input module 202, configured to input the image to be recognized into a text recognition model to obtain a text recognition result output by the text recognition model, where the text recognition model is trained by the above model training method.

[0132] The text recognition model is trained by the above model training method. Since the loss value is calculated based on the first sample recognition result output by the connectionist temporal classification network and the second sample recognition result output by the attention mechanism network during training, the connectionist temporal classification network can be used to enforce monotonic alignment between the input and output, and the attention mechanism network can be used to avoid ignoring the overall information when the connectionist temporal classification network predicts local information. Therefore, text recognition can be performed while considering both the overall information and the alignment of the input and output, improving the accuracy of text recognition.

[0133] Optionally, the text recognition result includes a first recognition result and a second recognition result, and the apparatus 200 further includes:

[0134] A weighted calculation module, configured to perform weighted calculation based on the first recognition result and the second recognition result to obtain a target text recognition result.

[0135] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0136] Based on the same inventive concept, an embodiment of the present disclosure further provides an electronic device, which includes:

[0137] A memory, on which a computer program is stored;

[0138] A processor, configured to execute the computer program in the memory to implement the steps of the above model training method or text recognition method.

[0139] Figure 6 is a block diagram of an electronic device shown in an exemplary embodiment of the present disclosure. As Figure 6 shown, the electronic device 300 may include: a processor 301, a memory 302. The electronic device 300 may further include one or more of a multimedia component 303, an input / output (I / O) interface 304, and a communication component 305.

[0140] Among them, the processor 301 is used to control the overall operation of the electronic device 300 to complete all or part of the steps in the above-mentioned model training method or text recognition method. The memory 302 is used to store various types of data to support the operation of the electronic device 300. These data may include, for example, instructions for any application or method operating on the electronic device 300, as well as application-related data, such as contact data, sent and received messages, pictures, audio, video, and so on. The memory 302 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The multimedia component 303 may include a screen and an audio component. Among them, the screen may be a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, and the microphone is used to receive external audio signals. The received audio signals may be further stored in the memory 302 or sent through the communication component 305. The audio component further includes at least one speaker for outputting audio signals. The I / O interface 304 provides an interface between the processor 301 and other interface modules. The above-mentioned other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 305 is used for wired or wireless communication between the electronic device 300 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, 4G or 5G, NB-IoT (Narrow Band Internet of Things), or a combination of one or more of them. Therefore, the corresponding communication component 305 may include: a Wi-Fi module, a Bluetooth module, an NFC module, and so on.

[0141] In an exemplary embodiment, the electronic device 300 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, and is used to execute the above-mentioned model training method or text recognition method.

[0142] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above-mentioned model training method or text recognition method are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 302 including program instructions, and the above-mentioned program instructions can be executed by the processor 301 of the electronic device 300 to complete the above-mentioned model training method or text recognition method.

[0143] Regarding the computer-readable storage medium in the above-mentioned embodiments, the steps of implementing the model training method or text recognition method when the computer program stored thereon is executed have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0144] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above-mentioned model training method or text recognition method when executed by the programmable device.

[0145] The preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.

[0146] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present disclosure will not separately describe various possible combination manners.

[0147] In addition, any combination can be made among various different embodiments of the present disclosure, as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.

Claims

1. A model training method, characterized in that, The method includes: Obtain a sample image, where the sample image includes sample text, and the sample text is labeled with a sample text label; Input the sample image into a text recognition model to obtain a sample text recognition result output by the text recognition model, where the text recognition model includes a connectionist temporal classification network and an attention mechanism network, and the sample text recognition result includes a first sample recognition result output by the connectionist temporal classification network and a second sample recognition result output by the attention mechanism network; Calculate a loss value according to the first sample recognition result, the second sample recognition result, and the sample text label; Adjust the parameters of the text recognition model according to the loss value; Wherein, the text recognition model is a fully convolutional point aggregation network model; The text recognition model further includes a high-resolution network. Inputting the sample image into the text recognition model to obtain the first sample recognition result and the second sample recognition result output by the text recognition model includes: Input the sample image into the high-resolution network to obtain an initial feature map sample; Determine a text contour feature map sample, a text direction feature map sample, and a text character category feature map sample according to the initial feature map sample; Perform a sequence conversion operation on the text contour feature map sample, the text direction feature map sample, and the text character category feature map sample to obtain a feature vector sample sequence; Input the feature vector sample sequence into the attention mechanism network and the connectionist temporal classification network to obtain a first character category distribution probability sample of each feature vector sample in the feature vector sample sequence output by the connectionist temporal classification network, and obtain a second character category distribution probability sample of each feature vector sample in the feature vector sample sequence output by the attention mechanism network; Obtain the first sample recognition result according to the first character category distribution probability sample, and obtain the second sample recognition result according to the second character category distribution probability sample.

2. The method according to claim 1, wherein The attention mechanism network includes an encoder and a decoder. Inputting the feature vector sample sequence into the attention mechanism network and the connectionist temporal classification network includes: Input the feature vector sample sequence into the encoder of the attention mechanism network to obtain an encoded feature vector sample sequence; Input the encoded feature vector sample sequence into the decoder of the attention mechanism network and the connectionist temporal classification network.

3. The method according to claim 1, wherein The method further includes: Determine a text boundary offset feature map sample according to the initial feature map sample, and the text boundary offset feature map sample represents the position of the sample text recognition result in the sample image.

4. The method according to claim 1, wherein The training parameter includes a loss weight value, and the loss value is calculated by the following formula: ; wherein, is the loss value, is the loss function of the connection timing classification network, is the loss function of the attention mechanism network, and λ is the loss weight value.

5. A text recognition method, characterized in that, The method includes: Obtain an image to be recognized, where the image to be recognized includes text to be recognized; Input the image to be recognized into a text recognition model to obtain a text recognition result output by the text recognition model, where the text recognition model is trained by the model training method described in any one of claims 1-4.

6. The method according to claim 5, wherein The text recognition result includes a first recognition result and a second recognition result, and the method further includes: Performing weighted calculation according to the first recognition result and the second recognition result to obtain a target text recognition result.

7. A model training device, characterized in that, The device includes: A first acquisition module, configured to acquire a sample image, where the sample image includes a sample text, and the sample text is labeled with a sample text label; A first input module, configured to input the sample image into a text recognition model to obtain a sample text recognition result output by the text recognition model, where the text recognition model includes a connectionist temporal classification network and an attention mechanism network, and the sample text recognition result includes a first sample recognition result output by the connectionist temporal classification network and a second sample recognition result output by the attention mechanism network; A calculation module, configured to calculate a loss value according to the first sample recognition result, the second sample recognition result, and the sample text label; An adjustment module, configured to adjust parameters of the text recognition model according to the loss value; wherein, the text recognition model is a fully convolutional point aggregation network model; The text recognition model further includes a high-resolution network, and the first input module is configured to: Input the sample image into the high-resolution network to obtain an initial feature map sample; Determine a text contour feature map sample, a text direction feature map sample, and a text character category feature map sample according to the initial feature map sample; Perform a sequence conversion operation on the text contour feature map sample, the text direction feature map sample, and the text character category feature map sample to obtain a feature vector sample sequence; Input the feature vector sample sequence into the attention mechanism network and the connectionist temporal classification network to obtain a first character category distribution probability sample of each feature vector sample in the feature vector sample sequence output by the connectionist temporal classification network, and obtain a second character category distribution probability sample of each feature vector sample in the feature vector sample sequence output by the attention mechanism network; Obtain the first sample recognition result according to the first character category distribution probability sample, and obtain the second sample recognition result according to the second character category distribution probability sample.

8. A text recognition device, characterized in that, The device includes: A second acquisition module, configured to acquire an image to be recognized, where the image to be recognized includes a text to be recognized; A second input module, configured to input the image to be recognized into a text recognition model to obtain a text recognition result output by the text recognition model, where the text recognition model is trained by the model training method according to any one of claims 1-4.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1-4 or any one of claims 5-6.

10. An electronic device, characterized in that, Including: A memory, on which a computer program is stored; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1-4 or any one of claims 5-6.

Citation Information

Patent Citations

  • Character recognition method based on an attention mechanism and linkage time classification loss

    CN109492679A

  • Handwriting model training method, text recognition method and apparatus, device, and medium

    WO2019232869A1