Method, apparatus, electronic device, and storage medium for generating text recognition model

By acquiring the training data set and correcting the feature extraction network and text recognition subnet, the recognition problem of optical character recognition technology when the image is not clear or the text direction changes is achieved, and higher text recognition accuracy and reliability are achieved.

CN114998882BActive Publication Date: 2025-07-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210597022.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-07-04
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

When existing optical character recognition technologies recognize text in images, especially when the image is not captured clearly and the text direction is rotated or offset, it is difficult to accurately recognize text.

Method used

By acquiring the training data set, including the annotation result of the first image and the classification label of the second image, the feature extraction network and the text recognition subnet are used for feature extraction and text recognition, and the network is corrected based on the difference between the annotation result and the prediction result, as well as the difference between the prediction label and the classification label, to improve the performance of the text recognition model.

Benefits of technology

Improves the accuracy and reliability of the text recognition model, ensuring that text in the image can be more accurately recognized under complex conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114998882B_ABST
    Figure CN114998882B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, electronic device, and storage medium for generating a text recognition model, which relates to the field of computer technologies, and particularly to artificial intelligence technologies such as natural language processing and deep learning. The method includes: obtaining a training data set, where the training data set includes a first image and a corresponding annotation result, and a second image and a corresponding classification label; inputting the first image into a preset feature extraction network to obtain a prediction result output by a text recognition sub-network; inputting the second image into the feature extraction network to obtain a prediction label output by a classification sub-network; and correcting the feature extraction network and the text recognition sub-network according to a first difference between the annotation result and the prediction result and a second difference between the prediction label and the classification label. Thus, the training of the text recognition model is assisted by a classification task, thereby improving the performance of the text recognition model, and further improving the accuracy and reliability of obtaining a text recognition result based on the text recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and particularly to artificial intelligence technologies such as natural language processing and deep learning. It can be applied to scenarios such as optical character recognition, and specifically relates to a method, apparatus, electronic device, and storage medium for generating a text recognition model. Background Art

[0002] Optical Character Recognition (OCR) is a technology that can convert picture information into text information that is easier to edit and store. Currently, OCR is widely used in various scenarios, such as bill recognition, bank card information recognition, formula recognition, etc. In addition, OCR also helps many downstream tasks, such as subtitle translation, security monitoring, etc.; at the same time, OCR also contributes to other visual tasks, such as video search.

[0003] Therefore, how to accurately recognize text from images has become a key research direction. Summary of the Invention

[0004] The present disclosure provides a method, apparatus, electronic device, and storage medium for generating a text recognition model.

[0005] According to a first aspect of the present disclosure, there is provided a method for generating a text recognition model, including:

[0006] Obtain a training data set, where the training data set includes a first image, an annotation result corresponding to the first image, a second image, and a classification label corresponding to the second image;

[0007] Input the first image into a preset feature extraction network to obtain a prediction result output by a text recognition sub-network connected to the preset feature extraction network;

[0008] Input the second image into the preset feature extraction network to obtain a predicted label output by a classification sub-network connected to the preset feature extraction network;

[0009] According to a first difference between the annotation result and the prediction result, and a second difference between the predicted label and the classification label, respectively correct the feature extraction network and the text recognition sub-network to obtain a trained feature extraction network and text recognition sub-network.

[0010] According to a second aspect of the present disclosure, there is provided a device for generating a text recognition model, including:

[0011] A first acquisition module, configured to acquire a training data set, where the training data set includes a first image, an annotation result corresponding to the first image, a second image, and a classification label corresponding to the second image;

[0012] A second acquisition module, configured to input the first image into a preset feature extraction network to obtain a prediction result output by a text recognition sub-network connected to the preset feature extraction network;

[0013] A third acquisition module, configured to input the second image into the preset feature extraction network to obtain a predicted label output by a classification sub-network connected to the preset feature extraction network;

[0014] A fourth acquisition module, configured to correct the feature extraction network and the text recognition sub-network respectively according to a first difference between the annotation result and the prediction result, and a second difference between the predicted label and the classification label, so as to obtain a trained feature extraction network and a text recognition sub-network.

[0015] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for generating a text recognition model as described in the first aspect.

[0019] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the method for generating a text recognition model as described in the first aspect.

[0020] According to a fifth aspect of the present disclosure, there is provided a computer program product, including computer instructions, and the computer instructions, when executed by a processor, implement the steps of the method for generating a text recognition model as described in the first aspect.

[0021] The method, device, electronic device, and storage medium for generating a text recognition model provided by the present disclosure have the following beneficial effects:

[0022] In the embodiments of the present disclosure, first, a first image, an annotation result corresponding to the first image, a second image, and a classification label corresponding to the second image are obtained. Then, the first image is input into a preset feature extraction network to obtain a prediction result output by a text recognition sub-network connected to the preset feature extraction network. The second image is input into the preset feature extraction network to obtain a predicted label output by a classification sub-network connected to the preset feature extraction network. Finally, according to a first difference between the annotation result and the prediction result, and a second difference between the predicted label and the classification label, the feature extraction network and the text recognition sub-network are respectively corrected to obtain a trained feature extraction network and a text recognition sub-network. Thus, the training of the text recognition model is assisted according to the classification task, thereby improving the performance of the text recognition model, and further improving the accuracy and reliability of obtaining the text recognition result based on the text recognition model.

[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0025] Figure 1 is a flowchart of a method for generating a text recognition model according to an embodiment of the present disclosure;

[0026] Figure 2 is a flowchart of a method for generating a text recognition model according to another embodiment of the present disclosure;

[0027] Figure 3 is a flowchart of a method for generating a text recognition model according to another embodiment of the present disclosure;

[0028] Figure 4 is a schematic structural diagram of a device for generating a text recognition model according to an embodiment of the present disclosure;

[0029] Figure 5 is a block diagram of an electronic device for implementing the method for generating a text recognition model of the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0031] The embodiments of the present disclosure relate to the fields of artificial intelligence technologies such as computer vision and deep learning.

[0032] Artificial Intelligence, abbreviated as AI in English, is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.

[0033] Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. The ultimate goal of deep learning is to enable machines to have the ability to analyze and learn like humans, and be able to recognize data such as text, images, and sounds.

[0034] Natural language processing is to use a computer to process, understand, and apply human languages (such as Chinese, English, etc.). It is an interdisciplinary field of computer science and linguistics, and is often referred to as computational linguistics. Since natural language is the fundamental sign that differentiates humans from other animals, and without language, human thinking would be impossible to talk about, so natural language processing reflects the highest task and realm of artificial intelligence. That is to say, only when a computer has the ability to process natural language can the machine be considered to have achieved true intelligence.

[0035] The following describes a method, apparatus, electronic device, and storage medium for generating a text recognition model according to an embodiment of the present disclosure with reference to the accompanying drawings.

[0036] It should be noted that the execution subject of the method for generating the text recognition model in this embodiment is the apparatus for generating the text recognition model. This apparatus can be implemented in software and / or hardware, and this apparatus can be configured in an electronic device, and the electronic device can include but is not limited to a terminal, a server, etc.

[0037] Figure 1 It is a schematic flowchart of a method for generating a text recognition model according to an embodiment of the present disclosure.

[0038] As Figure 1 shown, the method for generating the text recognition model includes:

[0039] S101: Obtain a training data set, where the training data set includes a first image, an annotation result corresponding to the first image, a second image, and a classification label corresponding to the second image.

[0040] Among them, the first image may contain text, and the annotation result corresponding to the first image is the correct text contained in the first image.

[0041] Among them, the second image also contains text, and the second image may be an image obtained by rotating the original image. For example, the second image may be an image obtained by randomly rotating the original image clockwise or counterclockwise by 0 degrees, 90 degrees, 180 degrees, or 270 degrees, etc. The present disclosure does not limit this.

[0042] Among them, the classification label corresponding to the second image is used to annotate the rotation angle of the second image. For example, when the rotation angle of the second image is 0 degrees, the corresponding classification label is 0; when the rotation angle of the second image is 90 degrees, the corresponding classification label is 1; when the rotation angle of the second image is 180 degrees, the corresponding classification label is 2; when the rotation angle of the second image is 270 degrees, the corresponding classification label is 3. The present disclosure does not limit this.

[0043] S102: Input the first image into a preset feature extraction network to obtain a prediction result output by a text recognition sub-network connected to the preset feature extraction network.

[0044] Among them, the preset feature extraction network can perform feature extraction on the first image to obtain image features corresponding to the first image.

[0045] Among them, the text recognition sub-network can analyze the image features corresponding to the first image to predict the text contained in the first image and obtain a prediction result.

[0046] Optionally, the preset feature extraction network may include multiple Convolution layers and multiple MaxPooling layers. In the embodiments of the present disclosure, the number of Convolution layers and the number of MaxPooling layers are not limited. For example, the feature extraction network may include 7 Convolution layers and 4 MaxPooling layers.

[0047] Optionally, the text recognition sub-network may be a Long Short Term Memory (LSTM) network or a Gate Recurrent Unit (GRU). The present disclosure does not limit this.

[0048] S103: Input the second image into a preset feature extraction network to obtain a prediction label output by a classification sub-network connected to the preset feature extraction network.

[0049] It can be understood that in the embodiments of the present disclosure, the preset feature extraction network can not only extract features from the first image, but also extract features from the second image to obtain the image features corresponding to the second image.

[0050] Among them, the classification sub-network connected to the preset feature extraction network can analyze the image features corresponding to the second image to predict the category of the second image and obtain a prediction label.

[0051] Optionally, the classification sub-network may include a fully connected layer, a softmax layer, etc. The present disclosure does not limit this.

[0052] S104: According to the first difference between the annotation result and the prediction result, and the second difference between the prediction label and the classification label, respectively correct the feature extraction network and the text recognition sub-network to obtain the trained feature extraction network and the text recognition sub-network.

[0053] Among them, the trained feature extraction network and the text recognition sub-network constitute a text recognition model. The text recognition model can be used to recognize the text in the image to be recognized to obtain an accurate text recognition result.

[0054] In the embodiments of the present disclosure, the preset feature extraction network can extract features from both the first image and the second image. Therefore, according to the first difference between the annotation result and the prediction result, and the second difference between the prediction label and the classification label, the feature extraction network can be corrected to obtain the trained feature extraction network, and the text recognition sub-network can be corrected according to the first difference to obtain the trained text recognition sub-network. Further, the classification sub-network can also be corrected according to the second difference to improve the performance of the classification sub-network.

[0055] It can be understood that the classification task is used to assist the text recognition task to train the feature extraction network, thus considering the situation that in the text recognition process, due to unclear image shooting, the text direction may rotate, shift, etc., which may make it difficult to accurately recognize the text. Therefore, by using the classification task to assist the training of the feature extraction network, the feature extraction network can accurately extract the features in the image, thereby improving the accuracy of the text recognition model.

[0056] Optionally, any loss function that can determine the first difference between the annotation result and the prediction result can be used to obtain the first difference, and the present disclosure does not limit this. For example, the CTC (Connectionist Temporal Classification) loss function can be used to determine the first difference.

[0057] Optionally, any loss function that can determine the second difference between the predicted label and the classification label can be used to obtain the second difference, and the present disclosure does not limit this. For example, the cross-entropy loss function can be used to determine the second difference.

[0058] Optionally, the first weight value corresponding to the text recognition sub-network and the second weight value corresponding to the classification sub-network can be obtained first, and then the feature extraction network can be corrected according to the sum of the product of the first difference and the first weight value and the product of the second difference and the second weight value. That is, the feature extraction network is corrected according to the weighted sum of the first difference and the second difference.

[0059] Among them, the first weight value may be equal to the second weight value or may not be equal to the second weight value, and the present disclosure does not limit this.

[0060] Optionally, a fifth weight value can also be set for the classification sub-network, and the fifth weight can decrease as the number of iterations increases. Then, the feature extraction network is corrected according to the first difference, the first weight value, the second difference, the second weight value, and the third weight value. Specifically, the feature extraction network is corrected according to the product of the first difference and the first weight, and the product of the second difference and the second weight and the third weight.

[0061] In the embodiments of the present disclosure, correcting the feature extraction network according to the first difference between the annotation result and the prediction result and the second difference between the predicted label and the classification label can improve the encoding accuracy of the feature extraction network to obtain more accurate features. Thereby improving the performance of the text recognition model and providing conditions for accurately recognizing the text contained in the image.

[0062] In an embodiment of the present disclosure, first, a first image, an annotation result corresponding to the first image, a second image, and a classification label corresponding to the second image are obtained. Then, the first image is input into a preset feature extraction network to obtain a prediction result output by a text recognition sub-network connected to the preset feature extraction network. The second image is input into the preset feature extraction network to obtain a predicted label output by a classification sub-network connected to the preset feature extraction network. Finally, according to a first difference between the annotation result and the prediction result, and a second difference between the predicted label and the classification label, the feature extraction network and the text recognition sub-network are respectively corrected to obtain a trained feature extraction network and a text recognition sub-network. Thus, the training of the text recognition model is assisted according to the classification task, thereby improving the performance of the text recognition model, and further improving the accuracy and reliability of obtaining the text recognition result based on the text recognition model.

[0063] Figure 2 is a schematic flowchart of a method for generating a text recognition model according to another embodiment of the present disclosure;

[0064] As Figure 2 shown, the method for generating the text recognition model includes:

[0065] S201: Obtain a first image, an unannotated third image, and an annotation result corresponding to the first image.

[0066] Among them, the unannotated third image may be an image that contains text but has no corresponding annotation result for the text.

[0067] S202: Perform a first augmentation process on the third image to obtain a second image.

[0068] Among them, the first augmentation process may include a first random rotation, normalization, etc., and the present disclosure does not limit this.

[0069] Specifically, to perform the first augmentation process on the third image, first perform a first random rotation on the third image, that is, randomly rotate clockwise by 0 degrees, 90 degrees, 180 degrees, or 270 degrees, or randomly rotate counterclockwise by 0 degrees, 90 degrees, 180 degrees, or 270 degrees. Then, normalize the image obtained after the first random rotation to obtain the second image. Thus, not only can the second image be increased to make the training data more abundant, but the data after normalization can also make the model converge more easily.

[0070] S203: Determine a classification label corresponding to the second image according to the type of the second image.

[0071] Optionally, the type of the second image may be determined according to the rotation angle of the second image. That is, the type of the second image may include rotation by 0 degrees, rotation by 90 degrees, rotation by 180 degrees, or rotation by 270 degrees, etc.

[0072] It should be noted that for different types of the second image, the corresponding classification labels are also different. For example, if the type of the second image is rotated by 0 degrees, the corresponding classification label can be 0 or θ0; if the type of the second image is rotated by 90 degrees, the corresponding classification label can be 1 or θ1; if the type of the second image is rotated by 180 degrees, the corresponding classification label can be 2 or θ2; if the type of the second image is rotated by 270 degrees, the corresponding classification label can be 3 or θ3. The present disclosure does not limit this.

[0073] S204: Perform a second augmentation process on the third image to obtain a blurred image corresponding to the third image.

[0074] Among them, the blurred image can be an image with some parts being blurred and unclear obtained by blurring some parts of the third image.

[0075] Optionally, the second augmentation process may include second random rotation, Gaussian blur, random masking, etc. The present disclosure does not limit this.

[0076] Among them, the angle of the second random rotation cannot exceed a preset threshold range. For example, the preset threshold range can be [-10°, 10°], or [-15°, 15°], etc. If the preset threshold range is [-10°, 10°], then the angle of clockwise rotation of the third image cannot exceed 10°, and the angle of counterclockwise rotation cannot exceed 10°. The present disclosure does not limit this.

[0077] Optionally, the ratio of the area of the random mask to the area of the third image can be greater than a first threshold and less than a second threshold. Among them, the first threshold can be 0.5, the second threshold can be 0.7, etc. The present disclosure does not limit this.

[0078] For example, the third image can be first subjected to a second random rotation, the image obtained after the second random rotation can be subjected to Gaussian blur, and then the image after Gaussian blur can be subjected to random masking to obtain a blurred image.

[0079] It should be noted that the above examples are only simple illustrations and cannot be used as specific limitations on the second augmentation process of the third image in the embodiments of the present disclosure.

[0080] S205: Input the blurred image into a preset feature extraction network to obtain a predicted image output by an image prediction sub-network connected to the preset feature extraction network.

[0081] In the embodiments of the present disclosure, a preset feature extraction network can extract features from a blurred image to obtain image features corresponding to the blurred image. Subsequently, an image prediction sub-network connected to the preset feature extraction network can analyze the image features corresponding to the blurred image to predict unclear content in the blurred image and obtain a predicted image.

[0082] Optionally, the image prediction sub-network may include: a supervised attention module, a full-resolution network, etc. The present disclosure does not limit this.

[0083] S206: Input the first image into the preset feature extraction network to obtain the prediction result output by the text recognition sub-network connected to the preset feature extraction network.

[0084] S207: Input the second image into the preset feature extraction network to obtain the prediction label output by the classification sub-network connected to the preset feature extraction network.

[0085] Among them, for the specific implementation forms of steps S206 and S207, reference may be made to the detailed descriptions in other embodiments of the present disclosure, and details will not be elaborated here.

[0086] S208: According to the first difference between the annotation result and the prediction result, the second difference between the prediction label and the classification label, and the third difference between the predicted image and the third image, respectively correct the feature extraction network and the text recognition sub-network to obtain the trained feature extraction network and text recognition sub-network.

[0087] In the embodiments of the present disclosure, the preset feature extraction network can extract features from the first image, can also extract features from the second image, and can also extract features from the blurred image. Therefore, when correcting the feature extraction network, it is necessary to simultaneously consider the first difference between the annotation result and the prediction result, the second difference between the prediction label and the classification label, and the third difference between the predicted image and the third image, so that the feature extraction network can better extract features from the first image, the second image, and the blurred image at the same time. That is, correct the feature extraction network according to the first difference, the second difference, and the third difference.

[0088] It can be understood that the training of the feature extraction network is assisted by the classification task and the image restoration task, thereby considering the situation that in the text recognition process, due to unclear image shooting, the text direction may rotate, shift, or the text may be blurred or missing, which may make it difficult to accurately recognize the text. Thus, by assisting the training of the feature extraction network through the classification task and the image prediction task, the feature extraction network can accurately extract features in the image, further improving the accuracy of the text recognition model.

[0089] Optionally, any loss function that can determine the third difference between the predicted image and the third image can be used to obtain the third difference, and the present disclosure does not limit this. For example, the mean-squared loss (MSE) can be used to determine the first difference.

[0090] Optionally, corresponding weights can be set for the text recognition sub-network, the classification sub-network, and the image prediction sub-network respectively. Then, the feature extraction network is corrected according to the weighted sum of the first difference, the second difference, and the third difference.

[0091] For example, the weight corresponding to the text recognition sub-network can be 0.5, the weight corresponding to the classification sub-network can be 0.25, and the weight corresponding to the image prediction sub-network can be 0.25. Then the weighted sum is 0.5 * the first difference + 0.25 * the second difference + 0.25 * the third difference. The present disclosure does not limit this.

[0092] In the embodiments of the present disclosure, the text recognition sub-network can also be corrected according to the first difference between the annotation result and the prediction result; the classification sub-network can be corrected according to the second difference between the predicted label and the classification label; and the image prediction sub-network can be corrected according to the third difference between the predicted image and the third image. Thereby, the performance of the text recognition sub-network, the classification sub-network, and the image prediction sub-network can be improved.

[0093] In the embodiments of the present disclosure, first, a first image, the annotation result corresponding to the first image, an unannotated third image, a second image, the classification label corresponding to the second image, and a blurred image are obtained. Then, the blurred image is input into a preset feature extraction network to obtain a predicted image output by an image prediction sub-network connected to the preset feature extraction network. The first image is input into the preset feature extraction network to obtain a prediction result output by a text recognition sub-network connected to the preset feature extraction network. The second image is input into the preset feature extraction network to obtain a predicted label output by a classification sub-network connected to the preset feature extraction network. Finally, according to the first difference between the annotation result and the prediction result, the second difference between the predicted label and the classification label, and the third difference between the predicted image and the third image, the feature extraction network and the text recognition sub-network are corrected respectively to obtain the trained feature extraction network and text recognition sub-network. Thus, the training of the text recognition model is assisted by the classification task and the image restoration task, thereby further improving the performance of the text recognition model, and further improving the accuracy and reliability of the text recognition result obtained based on the text recognition model.

[0094] Figure 3 It is a schematic flowchart of a method for generating a text recognition model according to another embodiment of the present disclosure;

[0095] AsFigure 3 As shown in Figure 3 , the method for generating the text recognition model includes:

[0096] S301: Obtain a first image, an unlabeled third image, and the annotation result corresponding to the first image.

[0097] S302: Perform a first augmentation process on the third image to obtain a second image.

[0098] S303: Determine the classification label corresponding to the second image according to the type of the second image.

[0099] S304: Perform a second augmentation process on the third image to obtain a blurred image corresponding to the third image.

[0100] S305: Input the blurred image into a preset feature extraction network to obtain a predicted image output by an image prediction sub-network connected to the preset feature extraction network.

[0101] S306: Input the first image into a preset feature extraction network to obtain a predicted result output by a text recognition sub-network connected to the preset feature extraction network.

[0102] S307: Input the second image into a preset feature extraction network to obtain a predicted label output by a classification sub-network connected to the preset feature extraction network.

[0103] Among them, for the specific implementation forms of step S301 and step S307, reference may be made to the detailed descriptions in other embodiments of the present disclosure, and details are not described herein again.

[0104] S308: Correct the text recognition sub-network according to the first difference between the annotation result and the predicted result to obtain a trained text recognition sub-network.

[0105] In the embodiments of the present disclosure, after obtaining the predicted result corresponding to the first image output by the text recognition sub-network, the text recognition sub-network can be corrected according to the first difference between the annotation result and the predicted result to improve the robustness and convergence of the text recognition sub-network, and further improve the performance of the text recognition sub-network.

[0106] S309: Determine the correction gradient of the feature extraction network according to the first difference, the second difference between the predicted label and the classification label, the third difference between the predicted image and the third image, and the weight value corresponding to each sub-network.

[0107] Among them, the correction gradient can be used to indicate the direction for correcting the parameters in the feature extraction network.

[0108] Optionally, the first correction gradient can be determined based on the first difference and the first weight value corresponding to the text recognition sub-network. Then, based on the second difference, the third difference, the second weight value corresponding to the classification sub-network, and the third weight value corresponding to the image prediction sub-network, the second correction gradient can be determined. When the training stop condition is that the accuracy of the text recognition sub-network is greater than the threshold, the current iteration number of the feature extraction network can be determined. Finally, based on the first correction gradient, the second correction gradient, and the current iteration number, the correction gradient of the feature extraction network can be determined.

[0109] Among them, the first weight value, the second weight value, and the third weight value can be pre-set weight values. The numerical values of the first weight value, the second weight value, and the third weight value can be the same or different, and the present disclosure does not limit this. For example, the first weight value can be 1, the second weight value can be 0.5, and the third weight value can be 0.5, etc.

[0110] Among them, the first correction gradient can be the product of the first difference and the first weight value.

[0111] Among them, the second correction gradient can be the second difference * the second weight value + the third difference * the third weight value.

[0112] It can be understood that since the training of the text recognition model focuses on the training of the text recognition task, therefore, as the number of iterations increases, the influence of the classification task and the image prediction task on the training of the feature extraction network can be appropriately reduced. That is, as the number of iterations increases, the influence of the second correction gradient on the correction gradient of the feature extraction network can be reduced.

[0113] Optionally, based on the mapping relationship between the current iteration number and the fourth weight value corresponding to the second correction gradient, the fourth weight value currently corresponding to the second correction gradient can be determined. Then, the correction gradient of the feature extraction network is determined as: the first correction gradient + the second correction gradient * the fourth weight value.

[0114] Optionally, the first correction gradient can also be determined first based on the first difference and the first weight value corresponding to the text recognition sub-network. Then, based on the second difference, the third difference, the second weight value corresponding to the classification sub-network, and the third weight value corresponding to the image prediction sub-network, the second correction gradient can be determined. When the training stop condition is to reach the target iteration number, the current iteration number of the feature extraction network can be determined, and based on the current iteration number and the target iteration number, the fourth weight value currently corresponding to the second correction gradient can be determined. Finally, based on the first correction gradient, the second correction gradient, and the fourth weight value, the correction gradient of the feature extraction network can be determined.

[0115] In an embodiment of the present disclosure, when the training stop condition is that the number of training times of the feature extraction network reaches the target number of iterations, the current fourth weight value of the second corrected gradient can be determined according to the current number of iterations and the target number of iterations. The fourth weight value is: 1 - current number of iterations / target number of iterations.

[0116] Among them, the corrected gradient of the feature extraction network can be:

[0117] Loss = α * Loss1 + (β * Loss2 + γ * Loss3) * (1 – Curr_it / Total_it)

[0118] Among them, Loss is the corrected gradient of the feature extraction network, Loss1 is the first difference, Loss2 is the second difference, Loss3 is the third difference, α is the first weight, β is the second weight, γ is the third weight, (1 – Curr_it / Total_it) is the fourth weight, Curr_it is the current number of iterations, and Total_it is the target number of iterations.

[0119] S310: Based on the corrected gradient, correct the feature extraction network to obtain a corrected feature extraction network.

[0120] In an embodiment of the present disclosure, after determining the corrected gradient of the feature extraction network, the feature extraction network can be corrected according to the corrected gradient to improve the convergence and robustness of the feature extraction network.

[0121] In an embodiment of the present disclosure, first, a first image, an annotation result corresponding to the first image, an unannotated third image, a second image, a classification label corresponding to the second image, and a blurred image are obtained. Then, the blurred image is input into a preset feature extraction network to obtain a predicted image output by an image prediction sub-network connected to the preset feature extraction network. The first image is input into the preset feature extraction network to obtain a predicted result output by a text recognition sub-network connected to the preset feature extraction network. The second image is input into the preset feature extraction network to obtain a predicted label output by a classification sub-network connected to the preset feature extraction network. Finally, according to a first difference between the annotation result and the predicted result, the text recognition sub-network is corrected. According to the first difference, a second difference between the predicted label and the classification label, a third difference between the predicted image and the third image, and weight values respectively corresponding to each sub-network, a correction gradient of the feature extraction network is determined, and then the feature extraction network is corrected. Thus, in the process of using the classification task and the image restoration task to assist in the training of the text recognition model, as the number of iterations increases, the influence of the classification task and the image restoration task on the feature extraction network can be reduced, so that the feature extraction network focuses more on the text recognition task, thereby further improving the performance of the text recognition model, and further improving the accuracy and reliability of the text recognition result obtained based on the text recognition model.

[0122] Figure 4 is a schematic structural diagram of a generating device for a text recognition model according to an embodiment of the present disclosure;

[0123] As Figure 4 shown, the generating device 400 for the text recognition model includes:

[0124] A first obtaining module 410, configured to obtain a training data set, where the training data set includes a first image, an annotation result corresponding to the first image, a second image, and a classification label corresponding to the second image;

[0125] A second obtaining module 420, configured to input the first image into a preset feature extraction network to obtain a predicted result output by a text recognition sub-network connected to the preset feature extraction network;

[0126] A third obtaining module 430, configured to input the second image into a preset feature extraction network to obtain a predicted label output by a classification sub-network connected to the preset feature extraction network;

[0127] A fourth obtaining module 440, configured to correct the feature extraction network and the text recognition sub-network respectively according to a first difference between the annotation result and the predicted result, and a second difference between the predicted label and the classification label, so as to obtain a trained feature extraction network and a trained text recognition sub-network.

[0128] In some embodiments of the present disclosure, the first acquisition module 410 is specifically configured to:

[0129] Acquire a first image, an unlabeled third image, and an annotation result corresponding to the first image;

[0130] Perform a first augmentation process on the third image to obtain a second image;

[0131] Determine a classification label corresponding to the second image according to the type of the second image.

[0132] In some embodiments of the present disclosure, it further includes:

[0133] A fifth acquisition module, configured to perform a second augmentation process on the third image to obtain a blurred image corresponding to the third image;

[0134] A sixth acquisition module, configured to input the blurred image into a preset feature extraction network to obtain a predicted image output by an image prediction sub-network connected to the preset feature extraction network;

[0135] The fourth acquisition module 440 is specifically configured to:

[0136] According to the first difference, the second difference, and the third difference between the predicted image and the third image, respectively correct the feature extraction network and the text recognition sub-network to obtain a trained feature extraction network and a trained text recognition sub-network.

[0137] In some embodiments of the present disclosure, the fourth acquisition module 440 includes:

[0138] A first acquisition unit, configured to correct the text recognition sub-network according to the first difference to obtain a trained text recognition sub-network;

[0139] A first determination unit, configured to determine a correction gradient of the feature extraction network according to the first difference, the second difference, the third difference, and the weight value corresponding to each sub-network;

[0140] A second acquisition unit, configured to correct the feature extraction network based on the correction gradient to obtain a corrected feature extraction network.

[0141] In some embodiments of the present disclosure, the first determination unit is specifically configured to:

[0142] Determine a first correction gradient according to the first difference and the first weight value corresponding to the text recognition sub-network;

[0143] Determine a second correction gradient according to the second difference, the third difference, the second weight value corresponding to the classification sub-network, and the third weight value corresponding to the image prediction sub-network;

[0144] In response to the training stop condition that the accuracy of the text recognition sub-network is greater than the threshold, determine the current iteration number of the feature extraction network;

[0145] Determine the corrected gradient of the feature extraction network according to the first corrected gradient, the second corrected gradient, and the current iteration number.

[0146] In some embodiments of the present disclosure, the first determination unit is specifically configured to:

[0147] Determine the first corrected gradient according to the first difference and the first weight value corresponding to the text recognition sub-network;

[0148] Determine the second corrected gradient according to the second difference, the third difference, the second weight value corresponding to the classification sub-network, and the third weight value corresponding to the image prediction sub-network;

[0149] In response to the training stop condition being that the target iteration number is reached, determine the current iteration number of the feature extraction network;

[0150] Determine the current fourth weight value of the second corrected gradient according to the current iteration number and the target iteration number;

[0151] Determine the corrected gradient of the feature extraction network according to the first corrected gradient, the second corrected gradient, and the fourth weight value.

[0152] It should be noted that the foregoing explanation of the method for generating the text recognition model also applies to the apparatus for generating the text recognition model in this embodiment, and will not be elaborated here.

[0153] In the embodiments of the present disclosure, first obtain a first image, an annotation result corresponding to the first image, a second image, and a classification label corresponding to the second image. Then input the first image into a preset feature extraction network to obtain a prediction result output by a text recognition sub-network connected to the preset feature extraction network, and input the second image into the preset feature extraction network to obtain a predicted label output by a classification sub-network connected to the preset feature extraction network. Finally, correct the feature extraction network and the text recognition sub-network respectively according to the first difference between the annotation result and the prediction result, and the second difference between the predicted label and the classification label, so as to obtain the trained feature extraction network and text recognition sub-network. Thus, the training of the text recognition model is assisted according to the classification task, thereby improving the performance of the text recognition model, and further improving the accuracy and reliability of obtaining the text recognition result based on the text recognition model.

[0154] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0155] Figure 5FIG. 0 shows a schematic block diagram of an exemplary electronic device 500 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smartphone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0156] As Figure 5 shown, the device 500 includes a computing unit 501 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0157] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0158] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the method for generating a text recognition model. For example, in some embodiments, the method for generating a text recognition model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method for generating a text recognition model described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the method for generating a text recognition model in any other suitable manner (e.g., by means of firmware).

[0159] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0160] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0161] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0162] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0163] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0164] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server may also be a server of a distributed system or a server combined with a blockchain.

[0165] In this embodiment, first, a first image, an annotation result corresponding to the first image, a second image, and a classification label corresponding to the second image are obtained. Then, the first image is input into a preset feature extraction network to obtain a prediction result output by a text recognition sub-network connected to the preset feature extraction network. The second image is input into the preset feature extraction network to obtain a predicted label output by a classification sub-network connected to the preset feature extraction network. Finally, according to the first difference between the annotation result and the prediction result, and the second difference between the predicted label and the classification label, the feature extraction network and the text recognition sub-network are respectively corrected to obtain the trained feature extraction network and text recognition sub-network. Thus, the training of the text recognition model is assisted according to the classification task, thereby improving the performance of the text recognition model, and further improving the accuracy and reliability of obtaining the text recognition result based on the text recognition model.

[0166] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0167] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this disclosure, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. In the description of this disclosure, the words "if" and "when" can be interpreted as "when...", "when...", "in response to determining", or "in the case of...".

[0168] The above specific embodiments do not limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.

Claims

1. A method for generating a text recognition model, comprising: Obtaining a training data set, wherein the training data set includes a first image, an annotation result corresponding to the first image, a second image, and a classification label corresponding to the second image, wherein the annotation result corresponding to the first image is the correct text included in the first image, and the classification label corresponding to the second image is used to annotate the rotation angle of the second image; Inputting the first image into a preset feature extraction network to obtain a prediction result output by a text recognition sub-network connected to the preset feature extraction network, wherein the preset feature extraction network extracts features from the first image to obtain image features corresponding to the first image, and the text recognition sub-network analyzes the image features corresponding to the first image to predict the text included in the first image and obtain a prediction result; Inputting the second image into the preset feature extraction network to obtain a prediction label output by a classification sub-network connected to the preset feature extraction network; According to a first difference between the annotation result and the prediction result, and a second difference between the prediction label and the classification label, respectively correct the feature extraction network and the text recognition sub-network to obtain a trained feature extraction network and text recognition sub-network.

2. The method according to claim 1, wherein, The obtaining of the training data set includes: Obtaining the first image, an unannotated third image, and the annotation result corresponding to the first image; Performing a first augmentation process on the third image to obtain the second image; Determining a classification label corresponding to the second image according to the type of the second image.

3. The method according to claim 2, wherein After obtaining the training data set, further comprising: Performing a second augmentation process on the third image to obtain a blurred image corresponding to the third image; Inputting the blurred image into the preset feature extraction network to obtain a predicted image output by an image prediction sub-network connected to the preset feature extraction network; The obtaining of the trained feature extraction network and text recognition sub-network includes: According to the first difference, the second difference, and a third difference between the predicted image and the third image, respectively correct the feature extraction network and the text recognition sub-network to obtain a trained feature extraction network and text recognition sub-network.

4. The method according to claim 3, wherein The obtaining of the trained feature extraction network and text recognition sub-network includes: Correcting the text recognition sub-network according to the first difference to obtain a trained text recognition sub-network; Determining a correction gradient of the feature extraction network according to the first difference, the second difference, the third difference, and weight values corresponding to each sub-network; Based on the correction gradient, correcting the feature extraction network to obtain a corrected feature extraction network.

5. The method according to claim 4, wherein, The determining of the correction gradient of the feature extraction network according to the first difference, the second difference, the third difference, and weight values corresponding to each sub-network includes: Determining a first correction gradient according to the first difference and a first weight value corresponding to the text recognition sub-network; Determine a second correction gradient according to the second difference, the third difference, the second weight value corresponding to the classification sub-network, and the third weight value corresponding to the image prediction sub-network; In response to the training stop condition being that the accuracy of the text recognition sub-network is greater than a threshold, determine the current iteration number of the feature extraction network; Determine the correction gradient of the feature extraction network according to the first correction gradient, the second correction gradient, and the current iteration number.

6. The method according to claim 4, wherein The determining the correction gradient of the feature extraction network according to the first difference, the second difference, the third difference, and the weight values corresponding to each sub-network respectively includes: Determine a first correction gradient according to the first difference and the first weight value corresponding to the text recognition sub-network; Determine a second correction gradient according to the second difference, the third difference, the second weight value corresponding to the classification sub-network, and the third weight value corresponding to the image prediction sub-network; In response to the training stop condition being that a target iteration number is reached, determine the current iteration number of the feature extraction network; Determine the current fourth weight value of the second correction gradient according to the current iteration number and the target iteration number; Determine the correction gradient of the feature extraction network according to the first correction gradient, the second correction gradient, and the fourth weight value.

7. An apparatus for generating a text recognition model, comprising: A first acquisition module, configured to acquire a training data set, where the training data set includes a first image, an annotation result corresponding to the first image, a second image, and a classification label corresponding to the second image, where the annotation result corresponding to the first image is the correct text included in the first image, and the classification label corresponding to the second image is used to annotate the rotation angle of the second image; A second acquisition module, configured to input the first image into a preset feature extraction network to obtain a prediction result output by a text recognition sub-network connected to the preset feature extraction network, where the preset feature extraction network performs feature extraction on the first image to obtain an image feature corresponding to the first image, and the text recognition sub-network analyzes the image feature corresponding to the first image to predict the text included in the first image and obtain a prediction result; A third acquisition module, configured to input the second image into the preset feature extraction network to obtain a prediction label output by a classification sub-network connected to the preset feature extraction network; A fourth acquisition module, configured to correct the feature extraction network and the text recognition sub-network respectively according to a first difference between the annotation result and the prediction result, and a second difference between the prediction label and the classification label, so as to obtain a trained feature extraction network and a text recognition sub-network.

8. The device according to claim 7, wherein, The first acquisition module is specifically configured to: Acquire the first image, an unannotated third image, and the annotation result corresponding to the first image; Perform a first augmentation process on the third image to obtain the second image; Determine the classification label corresponding to the second image according to the type of the second image.

9. The device according to claim 8, wherein Further includes: A fifth acquisition module, configured to perform a second augmentation process on the third image to obtain a blurred image corresponding to the third image; A sixth acquisition module, configured to input the blurred image into the preset feature extraction network to obtain a predicted image output by an image prediction sub-network connected to the preset feature extraction network; The fourth acquisition module is specifically configured to: According to the first difference, the second difference, and a third difference between the predicted image and the third image, respectively correct the feature extraction network and the text recognition sub-network to obtain a trained feature extraction network and a text recognition sub-network.

10. The device according to claim 9, wherein, The fourth acquisition module includes: A first acquisition unit, configured to correct the text recognition sub-network according to the first difference to obtain a trained text recognition sub-network; A first determination unit, configured to determine a correction gradient of the feature extraction network according to the first difference, the second difference, the third difference, and weight values respectively corresponding to each sub-network; A second acquisition unit, configured to correct the feature extraction network based on the correction gradient to obtain a corrected feature extraction network.

11. The apparatus according to claim 10, wherein, The first determination unit is specifically configured to: Determine a first correction gradient according to the first difference and a first weight value corresponding to the text recognition sub-network; Determine a second correction gradient according to the second difference, the third difference, a second weight value corresponding to the classification sub-network, and a third weight value corresponding to the image prediction sub-network; In response to that the training stop condition is that the accuracy rate of the text recognition sub-network is greater than a threshold, determine the current iteration number of the feature extraction network; Determine the correction gradient of the feature extraction network according to the first correction gradient, the second correction gradient, and the current iteration number.

12. The apparatus according to claim 10, wherein, The first determination unit is specifically configured to: Determine a first correction gradient according to the first difference and a first weight value corresponding to the text recognition sub-network; Determine a second correction gradient according to the second difference, the third difference, a second weight value corresponding to the classification sub-network, and a third weight value corresponding to the image prediction sub-network; In response to that the training stop condition is that a target iteration number is reached, determine the current iteration number of the feature extraction network; Determine a current fourth weight value of the second correction gradient according to the current iteration number and the target iteration number; Determine the correction gradient of the feature extraction network according to the first correction gradient, the second correction gradient, and the fourth weight value.

13. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

15. A computer program product comprising computer instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Recognition model training and application method and device, computing equipment and storage medium

    CN110569359A

  • End-to-end text detection and recognition method

    CN112733822A