Text recognition method, device, computer equipment and storage medium
By detecting and eliminating noise training images during the training process of text recognition model, the problem of poor performance of text recognition model in the prior art is solved, and the effect of improving model performance and generalization ability is achieved.
Patent Information
- Application Number
- CN202011487467.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-12-16
AI Technical Summary
In the prior art, the text recognition model has noise training samples during training, resulting in poor model performance.
The text content on the picture is determined by obtaining the picture to be identified and inputting it into the preset text recognition model. The text recognition model is obtained based on training picture sets. The training picture set includes multiple training pictures and labeled text content corresponding to each training picture. The multiple training pictures include noise training pictures and non-noise training pictures. During the training process, the performance of the model is improved by detecting the image quality categories of the training image and removing noise training images.
By increasing the diversity of training samples and removing noise training pictures during the training process, the performance of the text recognition model is improved, the model is overfitted, and the model generalization ability is improved.
Smart Images

Figure CN112580495B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and particularly to a text recognition method, apparatus, computer device, and storage medium. Background Art
[0002] In daily work and study, in order to recognize the text on pictures and the like, a text recognition model is usually used for recognition, and the text recognition model is usually trained with a large number of training samples. Since there are many training samples, and generally most of them are data obtained through manual annotation or from the Internet, there are often noisy training samples obtained during the process of obtaining training samples.
[0003] In related technologies, when training a text recognition model, most often, the noisy training samples are removed and other methods are used to train the text recognition model with noise-free training samples to obtain a trained text recognition model.
[0004] However, the above technologies have the problem that the performance of the trained text recognition model is not good. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a text recognition method, apparatus, computer device, and storage medium that can improve the performance of the trained text recognition model.
[0006] A text recognition method, the method includes:
[0007] Obtain a picture to be recognized;
[0008] Input the picture to be recognized into a preset text recognition model for text recognition, and determine the text content on the picture to be recognized;
[0009] Wherein, the preset text recognition model is trained based on a training picture set, the training picture set includes a plurality of training pictures and the corresponding labeled text content of each training picture, and the plurality of training pictures include noisy training pictures and non-noisy training pictures.
[0010] In one embodiment, the training method of the text recognition model includes:
[0011] Use the training picture set as the input of the initial text recognition model, use the corresponding labeled text content of each training picture as the reference output of the initial text recognition model, and perform a preset number of trainings on the initial text recognition model to obtain an intermediate text recognition model;
[0012] According to the picture quality category of each training picture, detect whether each training picture is a noisy training picture;
[0013] If the above training images are noisy training images, then use other non-noisy training images among the above to continue training the above intermediate text recognition model to obtain the above text recognition model.
[0014] In one embodiment, the above-mentioned steps of using the above training image set as the input of the initial text recognition model, using the labeled text content corresponding to each of the above training images as the reference output of the initial text recognition model, and training the initial text recognition model for a preset number of times to obtain an intermediate text recognition model include:
[0015] Input each training image in the above training image set into the above initial text recognition model to obtain the predicted text content corresponding to each of the above training images;
[0016] According to the predicted text content of the above training image and the corresponding labeled text content, modify the labeled text content of the above training image to obtain the new labeled text content corresponding to the above training image;
[0017] Train the above initial text recognition model for a preset number of times according to the new labeled text content corresponding to the above training image to obtain an intermediate text recognition model.
[0018] In one embodiment, the above-mentioned steps of modifying the labeled text content of the above training image according to the predicted text content of the above training image and the corresponding labeled text content to obtain the new labeled text content corresponding to the above training image include:
[0019] Calculate the first loss between the predicted text content of the above training image and the corresponding labeled text content;
[0020] Modify the labeled text content of the above training image according to the above first loss to obtain the new labeled text content corresponding to the above training image.
[0021] In one embodiment, the above-mentioned steps of modifying the labeled text content of the above training image according to the above first loss to obtain the new labeled text content corresponding to the above training image include:
[0022] Compare the above first loss with a preset first loss threshold;
[0023] If the above first loss is less than the above first loss threshold, then modify the labeled text content of the above training image to obtain the new labeled text content corresponding to the above training image.
[0024] In one embodiment, the above-mentioned steps of training the above initial text recognition model for a preset number of times according to the new labeled text content corresponding to the above training image to obtain an intermediate text recognition model include:
[0025] Calculate a second loss between the new annotation text content corresponding to the above training images and the predicted text content corresponding to the above training images;
[0026] Train the above initial text recognition model a preset number of times according to the above first loss and the above second loss to obtain the above intermediate text recognition model.
[0027] In one embodiment, the detecting whether each of the above training images is a noisy training image according to the image quality category of each of the above training images includes:
[0028] Input each of the above training images into a preset classifier for classification to obtain the image quality category corresponding to the above training image; wherein, the preset classifier is trained based on a first training image set, and the first training image set includes first training images and the labeled image quality categories corresponding to the first training images;
[0029] If the image quality category corresponding to the above training image is the first category and the first loss corresponding to the above training image is not less than the above first loss threshold, determine that the above training image is a noisy training image.
[0030] A text recognition device, the device includes:
[0031] An acquisition module, configured to acquire an image to be recognized;
[0032] A recognition module, configured to input the above image to be recognized into a preset text recognition model for text recognition to determine the text content on the above image to be recognized; wherein, the preset text recognition model is trained based on a training image set, the training image set includes a plurality of training images and the labeled text content corresponding to each training image, and the plurality of training images include noisy training images and non-noisy training images.
[0033] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0034] Acquire an image to be recognized;
[0035] Input the above image to be recognized into a preset text recognition model for text recognition to determine the text content on the above image to be recognized;
[0036] Wherein, the preset text recognition model is trained based on a training image set, the training image set includes a plurality of training images and the labeled text content corresponding to each training image, and the plurality of training images include noisy training images and non-noisy training images.
[0037] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the following steps are implemented:
[0038] Obtain a picture to be recognized;
[0039] Input the picture to be recognized into a preset text recognition model for text recognition, and determine the text content on the picture to be recognized;
[0040] Among them, the preset text recognition model is trained based on a training picture set, the training picture set includes multiple training pictures and the corresponding labeled text content for each training picture, and the multiple training pictures include noisy training pictures and non-noisy training pictures.
[0041] The above text recognition method, device, computer device and storage medium obtain a picture to be recognized, input the picture to be recognized into a preset text recognition model for text recognition, and determine the text content on the picture to be recognized; among them, the text recognition model is trained based on a training picture set, the training picture set includes multiple training pictures and the corresponding labeled text content for each training picture, and the multiple training pictures include noisy training pictures and non-noisy training pictures. In this method, since the text recognition model is trained based on noisy training pictures and non-noisy training pictures during training, the diversity of training samples can be increased, thereby improving the performance of the trained text recognition model. Description of the Drawings
[0042] Figure 1 It is the internal structure diagram of a computer device in an embodiment;
[0043] Figure 2 It is the flowchart of the text recognition method in an embodiment;
[0044] Figure 3 It is the flowchart of the text recognition method in another embodiment;
[0045] Figure 4 It is the flowchart of the text recognition method in another embodiment;
[0046] Figure 4a It is the flowchart of the process of modifying the labeled text content in another embodiment;
[0047] Figure 5 It is the flowchart of the text recognition method in another embodiment;
[0048] Figure 5a It is the flowchart of the process of network optimization using a classifier in another embodiment;
[0049] Figure 5bIt is a schematic diagram of the specific network structure of the text recognition method in another embodiment;
[0050] Figure 6 It is a block diagram of the structure of the text recognition device in one embodiment. Specific implementation manners
[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0052] The text recognition method provided by the embodiments of the present application can be applied to a computer device, which can be a terminal or a server. Taking the terminal as an example, its internal structure diagram can be as Figure 1 shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication) or other technologies. When the computer program is executed by the processor, it realizes a text recognition method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0053] Those skilled in the art can understand that Figure 1 the structure shown in
[0054] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0055] It should be noted that the execution subject of the embodiments of the present application can be a computer device or a text recognition model. The following embodiments will describe the technical solutions of the present application with the computer device as the execution subject.
[0055] In one embodiment, a text recognition method is provided. This embodiment relates to the specific process of how to recognize a picture to be recognized. As Figure 2 shown, the method may include the following steps:
[0056] S202, Obtain the image to be recognized.
[0057] Among them, the image to be recognized can be a PDF-formatted image, a JPG-formatted image, a PNG-formatted image, etc., and can also be a single frame image obtained from video images, etc.
[0058] In addition, the image to be recognized can be an image that only contains text content; it can also be an image that includes not only text content but also portraits of animals, plants, people, etc.; it can also be an image that includes other content.
[0059] Specifically, image acquisition devices such as cameras can be used to collect images of the content to be recognized to obtain the image to be recognized. It can also be that the image to be recognized is pre-stored in the cloud or a local database, and when needed, the image to be recognized can be directly retrieved from the cloud or the local database.
[0060] S204, Input the above-mentioned image to be recognized into a preset text recognition model for text recognition to determine the text content on the above-mentioned image to be recognized; among them, the above-mentioned preset text recognition model is trained based on a training image set, the above-mentioned training image set includes multiple training images and the corresponding labeled text content for each training image, and the above-mentioned multiple training images include noise training images and non-noise training images.
[0061] In this step, the text recognition model can be a neural network model, which can include a convolutional neural network model, a recurrent neural network model, etc. Optionally, the text recognition model here can be a neural network model composed of VGG (Visual Geometry Group) and Bi LSTM (Bi-directional Long Short-Term Memory) networks. Of course, it can also be other types of networks.
[0062] Before recognizing the image to be recognized, the text recognition model can be pre-trained. Here, the training process is simply given. The training process of this text recognition model is as follows: First, obtain a training image set, which includes multiple training images and the corresponding labeled text content for each training image. Among the multiple training images, there are at least one noise image and at least one non-noise image; then, use each training image in the training image set as the input of the initial text recognition model, and use the corresponding labeled text content of each training image as the reference input of the initial text recognition model to train the initial text recognition model to obtain a trained text recognition model.
[0063] When training a text recognition model here, it is possible to directly train the text recognition model using each training image; or it is possible to train a preliminary model of the text recognition model using each training image, and then process each training image (for example, removing noisy training images, performing preprocessing such as translation on each training image, etc.), and then continue to train the preliminary model using the processed training images to finally obtain a trained text recognition model; of course, there can also be other training methods, which are not specifically limited in this embodiment.
[0064] In addition, the above-mentioned noisy training images generally refer to training images whose corresponding labeled text content is incorrect, and non-noisy training images generally refer to training images whose corresponding labeled text content is correct.
[0065] Furthermore, in the process of training the text recognition model above, since noisy training images are used, this can increase the diversity of training images, thereby making the performance of the trained text recognition model better.
[0066] Specifically, after training the text recognition model as above, the image to be recognized can be input into the trained text recognition model for text recognition to obtain the text content on the image to be recognized.
[0067] In the above text recognition method, by obtaining the image to be recognized, inputting the image to be recognized into a preset text recognition model for text recognition, and determining the text content on the image to be recognized; wherein, the text recognition model is trained based on a training image set, and the training image set includes multiple training images and the corresponding labeled text content for each training image, and the multiple training images include noisy training images and non-noisy training images. In this method, since the text recognition model is trained based on noisy training images and non-noisy training images during training, this can increase the diversity of training samples, thereby improving the performance of the trained text recognition model.
[0068] The above embodiments briefly introduce the training process of the text recognition model. The following will elaborate on the specific training process of the text recognition model.
[0069] In another embodiment, another text recognition method is provided. This embodiment relates to the specific process of how to train a text recognition model. On the basis of the above embodiments, as Figure 3 shown, the training process of the text recognition model can include the following steps:
[0070] S302. Use the above training image set as the input of the initial text recognition model, and use the annotation text content corresponding to each of the above training images as the reference output of the initial text recognition model. Train the initial text recognition model for a preset number of times to obtain an intermediate text recognition model.
[0071] In this step, when training the text recognition model, it is divided into two stages for training, the early training stage and the late training stage. This step belongs to the early training stage.
[0072] In the early training stage, since overfitting generally does not occur in the network, all training images in the training image set can be made to participate in the training process. In the early training stage, all training images can be input into the initial text recognition model, and the initial text recognition model can be trained with reference to the annotation text content corresponding to each training image. Here, when inputting each training image, it can be input in batches or individually.
[0073] In addition, when using each training image to train the initial text recognition model in the early stage, in order to make the finally trained model more accurate, a preset number of trainings can be carried out in the early stage. The preset number here is less than the number of trainings when the text recognition model is finally trained well. Usually, it is a value taken before the text recognition model is not yet trained well. For example, when the number of trainings of the text recognition model reaches 10,000 times, it can be considered trained well. Then the preset number here can be taken as 5,000 times, 6,000 times, etc. Similarly, when the training reaches the preset number of times, the text recognition model is not yet trained well. At this time, the obtained model is a model in the intermediate process, denoted as the intermediate text recognition model.
[0074] S304. According to the image quality categories of each of the above training images, detect whether each of the above training images is a noisy training image.
[0075] In this step, the image quality categories of each training image can be set according to the actual situation. For example, it can include two categories (respectively simple samples and difficult samples. Simple samples refer to training images with less background noise and clear pictures, and difficult samples refer to training images with more background noise and blurred pictures); for another example, it can also include three categories (respectively simple samples, intermediate samples, and difficult samples, where intermediate samples refer to training images with relatively less background noise and relatively clear pictures, etc.); of course, it can also include other categories. Here is only an example. Taking the above two categories as an example, the image quality categories of the training images here can be represented by two numbers. For example, 1 represents a simple sample and 0 represents a difficult sample.
[0076] In addition, the quality categories of the above-mentioned training images can be obtained in advance through a trained classifier, or can be obtained in advance through manual annotation, and of course can also be obtained through other means. In short, the image quality categories of the training images can be obtained.
[0077] After obtaining the image quality categories of the training images, it is possible to directly determine whether each training image is a noisy training image based on the image quality categories of the training images. For example, if the image quality category of a training image is difficult sample 0, it can be determined that the training image is a noisy training image. Or it is also possible to comprehensively determine whether each training image is a noisy training image by combining the image quality categories of the training images and the training results of the training images in the training process of the intermediate text recognition model in S302 above. For example, if the image quality category of a training image is simple sample 1, but the training result in the training process of the intermediate text recognition model is poor / more noisy, then it can be considered that the training image is a noisy training image. Of course, it can also be detected by other means. In short, it is possible to detect whether each training image is a noisy training image.
[0078] Of course, the above-mentioned detection of whether each training image is a noisy training image is generally carried out after the intermediate text recognition model is trained.
[0079] S306, if the above-mentioned training image is a noisy training image, then use the other non-noisy training images among the above to continue training the above-mentioned intermediate text recognition model to obtain the above-mentioned text recognition model.
[0080] In this step, it is possible to detect whether each training image is a noisy training image through several methods in S304 above. Then, after obtaining the intermediate text recognition model above, it is necessary to continue training the intermediate text recognition model. When continuing to train, the results of the detected noisy training images can be combined for training. This continued training belongs to the later training stage mentioned above.
[0081] Taking a training image as a noisy training image as an example for illustration, when it is detected that the training image is a noisy training image above, then in the later training stage, parameters such as the loss of the training image do not participate in backpropagation, but the losses of other non-noisy training images in the training images are used to participate in backpropagation. Finally, through the parameter backpropagation of the training images and continuously adjusting the parameters of the intermediate text recognition model, the trained text recognition model can be obtained.
[0082] The text recognition method in this embodiment can train the initial text recognition model with all training images in the early stage to obtain an intermediate text recognition model. According to the image quality category of each training image, it is detected whether each training image is a noisy training image. When the training image is a noisy training image, its parameters do not participate in backpropagation, and the parameters of other non-noisy training images are used for backpropagation to continue training the intermediate text recognition model to obtain a trained text recognition model. In this embodiment, since the model can be trained with noisy training samples and non-noisy samples in the early stage of model training, the diversity of training samples can be improved, and the performance of the trained text recognition model can be guaranteed. Further, in the later stage of model training, noisy training images can be screened out in combination with the image quality category of the training images to prevent noisy training images from participating in the later model training too much, which can prevent the text recognition model from overfitting during training and improve the generalization ability of the trained text recognition model.
[0083] The above embodiment introduces the training process of combining the early and late training of the text recognition model. The following will detail the early training process of the text recognition model.
[0084] In another embodiment, another text recognition method is provided. This embodiment relates to the specific process of how to train the initial text recognition model in the early stage. On the basis of the above embodiment, as Figure 4 shown, the above S302 may include the following steps:
[0085] S402, input each training image in the above training image set into the above initial text recognition model to obtain the predicted text content corresponding to each above training image.
[0086] In this step, after obtaining each training image, each training image is input into the initial text recognition model in sequence or in batches for text recognition processing, and the predicted text content corresponding to each training image can be obtained.
[0087] S404, modify the labeled text content of the above training image according to the predicted text content and the corresponding labeled text content of the above training image to obtain the new labeled text content corresponding to the above training image.
[0088] In this step, after obtaining the predicted text content of each training image, correspondingly, the labeled text content corresponding to each training image can also be obtained. Then, optionally, the following steps A1 and A2 can be used to modify the labeled text content of the training image:
[0089] Step A1, calculate the first loss between the predicted text content of the above training image and the corresponding labeled text content.
[0090] Step A2: Modify the annotation text content of the above training images according to the above first loss to obtain the new annotation text content corresponding to the above training images.
[0091] In steps A1 - A2, reference can be made to Figure 4a As shown, when inputting the training images into the initial text recognition model to obtain the predicted text content, a loss function can be used to calculate the loss between the predicted text content of each training image and the corresponding annotation text content. Here, the loss obtained for each training image is recorded as the first loss, that is, Figure 4a the loss 1 in , and the loss function is not specifically limited here and can be a mean square error function, a covariance function, etc.
[0092] After obtaining the first loss corresponding to each training image, taking one training image as an example, optionally, the above first loss can be compared with a preset first loss threshold; if the above first loss is less than the above first loss threshold, then modify the annotation text content of the above training image to obtain the new annotation text content corresponding to the above training image.
[0093] That is to say, after obtaining the first loss corresponding to the training image, the first loss of the training image can be compared with the first loss threshold. Here, the comparison can be in ways such as taking the difference, taking the quotient, or directly comparing the magnitudes. In short, after the comparison, a comparison result can be obtained. In a possible implementation manner, continue to refer to Figure 4a As shown, it can be determined whether the first loss is less than the threshold (the first loss threshold). If the first loss of the training image is less than the first loss threshold, it indicates that the training image should be a non - noise training image, and the training image can participate in parameter backpropagation. Then, the annotation text content corresponding to the training image can be modified to obtain the new annotation text content corresponding to the training image, denoted as the new annotation text content. In another possible implementation manner, if the first loss of the training image is not less than the first loss threshold, that is, the first loss of the training image is greater than or equal to the first loss threshold, then it indicates that the training image may be a noise training image, and its annotation text content is not modified temporarily. It should be noted that all training images can be operated in this way.
[0094] Exemplarily, assume that the annotation text content corresponding to a training image is ABCD. Then, one or more of these four letters can be randomly changed to obtain the new annotation text content. For example, change B to E to obtain the new annotation text content AECD.
[0095] In addition, the first loss threshold here can be determined according to the actual situation, and the first loss threshold here is different from the loss threshold corresponding to the finally trained text recognition model. Generally, the first loss threshold here is greater than the loss threshold corresponding to the finally trained text recognition model. For example, if the loss threshold corresponding to the finally trained text recognition model is 0.001, the first loss threshold here should be greater than 0.001, for example, it can be 0.05.
[0096] S406. Train the initial text recognition model a preset number of times according to the new labeled text content corresponding to the above training pictures to obtain an intermediate text recognition model.
[0097] In this step, after obtaining the new labeled text content corresponding to the training pictures, optionally, the second loss between the new labeled text content corresponding to the above training pictures and the predicted text content corresponding to the above training pictures can be calculated; the initial text recognition model is trained a preset number of times according to the first loss and the second loss to obtain the above intermediate text recognition model.
[0098] That is to say, taking a training picture as an example for illustration, continue to refer to Figure 4a As shown, the loss function can be used to calculate the loss between the new labeled text content of the training picture and the corresponding predicted text content. The loss obtained for each training picture is recorded as the second loss, that is Figure 4a Loss 2 in the figure. The loss function is not specifically limited here and can be the same as the loss function for calculating the first loss above. For example, it can be the mean square error function, covariance function, etc.
[0099] It should be noted that when calculating the second loss here, if the new labeled text content of the training picture is the same as the corresponding predicted text content, the second loss is 0; if not, the loss function can be used to continue the calculation.
[0100] After calculating the first loss and the second loss of the training pictures above, the first loss and the second loss can be summed (including simple summation, weighted summation, etc.), and the parameters of the initial text recognition model are adjusted using the obtained loss sum value, that is, the initial text recognition model is trained. When the training reaches the preset number of times, an intermediate text recognition model is obtained. Here, two losses are used to train the initial text recognition model, and more backpropagation parameters are relied on, so the finally trained text recognition model will be more accurate and the performance of the model will be better.
[0101] In the text recognition method of this embodiment, the training images can be input into the initial text recognition model to obtain the predicted text content. Then, the labeled text content of the training images is modified according to the predicted text content and the labeled text content, and the initial text recognition model is trained according to the obtained new labeled text content to obtain the intermediate text recognition model. In this embodiment, since the labeled text content of the training images can be modified, the diversity of the training samples can be further increased, thereby further improving the performance of the obtained text recognition model. Further, the first loss can be used to determine whether to modify the label, which is relatively accurate and effective, thus avoiding the non-modification of the label and improving the accuracy of label modification. Furthermore, the first loss and a threshold can be compared to determine whether to modify the label, which can further refine the modification process, thereby further ensuring the accuracy of label modification.
[0102] The above embodiments introduce the specific process of training the initial text recognition model in the early stage. The following will detail the detection of whether the training images are noise images before the later training stage.
[0103] In another embodiment, another text recognition method is provided. This embodiment involves the specific process of detecting whether the training images are noise images by the image quality category of the training images. On the basis of the above embodiments, as Figure 5 shown, the above S304 may include the following steps:
[0104] S502, input each of the above training images into a preset classifier for classification to obtain the image quality category corresponding to the above training images; wherein, the above preset classifier is trained based on a first training image set, and the first training image set includes first training images and the labeled image quality categories corresponding to the first training images.
[0105] Among them, the classifier here can be pre-trained. Taking the above two categories as an example, the training process of the classifier can include: obtaining a first training image set, which can be the same as the training image set in the above S204. The first training image set includes multiple first training images, and each first training image is labeled with an image quality category. For example, some are labeled 0 and some are labeled 1. Then, each first training image can be input into the initial classifier for classification to obtain the predicted image quality category corresponding to each first training image, calculate the loss between the predicted image quality category and the corresponding labeled image quality category of each first training image, and use the calculated loss for parameter backpropagation to train the initial classifier. When the number of training times reaches the classification training times threshold, or the loss no longer changes or reaches the classification loss threshold, the parameters of the classifier are fixed, and the trained classifier can be obtained.
[0106] After the classifier is trained, each training image can be input into the trained classifier for classification, and the image quality category corresponding to each training image can be obtained.
[0107] It should be noted that, as shown in Figure 5a here, the classifier can also be a part of the above text recognition model and can be used as a classification branch in the text recognition model (i.e., the classification branch in Figure 5a ), and the above text recognition process can be used as a recognition branch in the text recognition model. When backpropagating the parameters using the first loss and the second loss, the network parameters of both the classification branch and the recognition branch can also be adjusted.
[0108] S504, if the image quality category corresponding to the above training image is the first category and the first loss corresponding to the above training image is not less than the above first loss threshold, then determine that the above training image is a noise training image.
[0109] In this step, taking the classifier with two categories as an example, the first category here refers to the category of simple sample 1, that is, the background noise of the training image is less, the image is clear, and the image quality is relatively high.
[0110] In addition, during the training process of the above text recognition model, taking a training image as an example, the comparison result between the first loss corresponding to the training image and the first loss threshold can also be obtained. If the image quality category of the training image here is simple sample 1 and its first loss is greater than or equal to the first loss threshold, then it can be determined that the training image is a noise training image. During the later stage of training the intermediate text recognition model, the loss and other parameters of this training image are not used for backpropagation, reducing the impact of noise samples on the model performance and avoiding overfitting of the model.
[0111] Correspondingly, as shown in Figure 5a here, if the image quality category of the training image here is difficult sample 0 and its first loss is less than the first loss threshold, then the label can be modified in the above Figure 4a way, continue to calculate the second loss, and train the intermediate text recognition model through the first loss and the second loss.
[0112] In addition, combined with the classification branch here, a specific network structure example is given. As shown in Figure 5b obtain the image to be recognized (such as the input images in Figure 5b ), input it into the text recognition model (such as vgg and bilstm in Figure 5b ) for text recognition, and at the same time combine the classification result of the classification branch (such as in Figure 5bthe fc|softmax(class) in and the recognition result of the recognition branch (such as Figure 5b calculate the loss in fc|ctc(recog) in, and train the text recognition model to obtain the text recognition model.
[0113] In the text recognition method of this embodiment, each training picture can be classified by a classifier to obtain the picture quality category corresponding to each training picture. If the picture quality category of the training picture is the first category and the first loss is not less than the first loss threshold, it is determined that the training picture is a noise training picture. In this embodiment, the picture quality category of the training picture is determined by the classifier, and the category result is relatively accurate. Then the noise training pictures determined subsequently are also accurate, so that the noise training pictures can be more accurately removed to prevent the text recognition model from overfitting in the later training stage.
[0114] It should be understood that although Figures 2 - 5 the steps in the flowchart of Figures 2 - 5 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0115] In one embodiment, as Figure 6 shown, a text recognition device is provided, including: an acquisition module 10 and a recognition module 11, where:
[0116] The acquisition module 10 is used to acquire the picture to be recognized;
[0117] The recognition module 11 is used to input the above-mentioned picture to be recognized into a preset text recognition model for text recognition, and determine the text content on the above-mentioned picture to be recognized; wherein, the above-mentioned preset text recognition model is trained based on a training picture set, the above-mentioned training picture set includes a plurality of training pictures and the labeled text content corresponding to each training picture, and the above-mentioned plurality of training pictures include noise training pictures and non-noise training pictures.
[0118] For the specific limitations of the text recognition device, reference can be made to the limitations on the text recognition method in the above text, which will not be elaborated here.
[0119] In another embodiment, another text recognition device is provided. Based on the above embodiment, the device may further include a model training module, which includes an initial training unit, a detection unit, and a subsequent training unit, where:
[0120] The initial training unit is configured to use the above training picture set as the input of the initial text recognition model, use the annotation text content corresponding to each of the above training pictures as the reference output of the initial text recognition model, and perform training on the initial text recognition model for a preset number of times to obtain an intermediate text recognition model;
[0121] The detection unit is configured to detect whether each of the above training pictures is a noisy training picture according to the picture quality category of each of the above training pictures;
[0122] The subsequent training unit is configured to, when the above training picture is a noisy training picture, continue to train the intermediate text recognition model using other above non-noisy training pictures to obtain the above text recognition model.
[0123] In another embodiment, another text recognition device is provided. Based on the above embodiment, the initial training unit may include a recognition subunit, a modification subunit, and an initial training subunit, where:
[0124] The recognition subunit is configured to input each training picture in the above training picture set into the above initial text recognition model to obtain the predicted text content corresponding to each of the above training pictures;
[0125] The modification subunit is configured to modify the annotation text content of the above training picture according to the predicted text content and the corresponding annotation text content of the above training picture to obtain the new annotation text content corresponding to the above training picture;
[0126] The initial training subunit is configured to perform training on the above initial text recognition model for a preset number of times according to the new annotation text content corresponding to the above training picture to obtain an intermediate text recognition model.
[0127] Optionally, the above modification subunit is specifically configured to calculate a first loss between the predicted text content and the corresponding annotation text content of the above training picture; modify the annotation text content of the above training picture according to the first loss to obtain the new annotation text content corresponding to the above training picture.
[0128] Optionally, the above modification subunit is specifically configured to compare the first loss with a preset first loss threshold; when the first loss is less than the first loss threshold, modify the annotation text content of the above training picture to obtain the new annotation text content corresponding to the above training picture.
[0129] Optionally, the above initial training subunit is specifically configured to calculate a second loss between the new annotation text content corresponding to the above training picture and the predicted text content corresponding to the above training picture; and train the above initial text recognition model a preset number of times according to the above first loss and the above second loss to obtain the above intermediate text recognition model.
[0130] In another embodiment, another text recognition device is provided. On the basis of the above embodiment, the above detection unit may include a classification subunit and a determination subunit, where:
[0131] The classification subunit is configured to input each of the above training pictures into a preset classifier for classification to obtain the picture quality category corresponding to the above training picture; where the above preset classifier is trained based on a first training picture set, and the first training picture set includes first training pictures and the labeled picture quality categories corresponding to the above first training pictures;
[0132] The determination subunit is configured to determine the above training picture as a noise training picture when the picture quality category corresponding to the above training picture is the first category and the first loss corresponding to the above training picture is not less than the above first loss threshold.
[0133] For the specific limitations of the text recognition device, reference may be made to the limitations on the text recognition method above, which will not be elaborated here.
[0134] Each module in the above text recognition device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to the above respective modules.
[0135] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0136] Obtain a picture to be recognized; input the above picture to be recognized into a preset text recognition model for text recognition to determine the text content on the above picture to be recognized; where the above preset text recognition model is trained based on a training picture set, and the training picture set includes multiple training pictures and the labeled text content corresponding to each training picture, and the multiple training pictures include noise training pictures and non-noise training pictures.
[0137] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0138] Use the above training image set as the input of the initial text recognition model, use the annotation text content corresponding to each of the above training images as the reference output of the above initial text recognition model, and train the above initial text recognition model for a preset number of times to obtain an intermediate text recognition model; according to the image quality category of each of the above training images, detect whether each of the above training images is a noisy training image; if the above training image is a noisy training image, then use other above non-noisy training images to continue training the above intermediate text recognition model to obtain the above text recognition model.
[0139] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0140] Input each training image in the above training image set into the above initial text recognition model to obtain the predicted text content corresponding to each of the above training images; according to the predicted text content and the corresponding annotation text content of the above training image, modify the annotation text content of the above training image to obtain the new annotation text content corresponding to the above training image; according to the new annotation text content corresponding to the above training image, train the above initial text recognition model for a preset number of times to obtain an intermediate text recognition model.
[0141] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0142] Calculate the first loss between the predicted text content and the corresponding annotation text content of the above training image; according to the above first loss, modify the annotation text content of the above training image to obtain the new annotation text content corresponding to the above training image.
[0143] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0144] Compare the above first loss with a preset first loss threshold; if the above first loss is less than the above first loss threshold, then modify the annotation text content of the above training image to obtain the new annotation text content corresponding to the above training image.
[0145] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0146] Calculate the second loss between the new annotation text content corresponding to the above training image and the predicted text content corresponding to the above training image; according to the above first loss and the above second loss, train the above initial text recognition model for a preset number of times to obtain the above intermediate text recognition model.
[0147] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0148] Input each of the above training pictures into a preset classifier for classification to obtain the picture quality category corresponding to the above training picture; wherein, the above preset classifier is trained based on a first training picture set, and the first training picture set includes first training pictures and the labeled picture quality categories corresponding to the first training pictures; if the picture quality category corresponding to the above training picture is the first category and the first loss corresponding to the above training picture is not less than the above first loss threshold, then determine that the above training picture is a noise training picture.
[0149] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0150] Obtain a picture to be recognized; input the above picture to be recognized into a preset text recognition model for text recognition to determine the text content on the above picture to be recognized; wherein, the above preset text recognition model is trained based on a training picture set, the training picture set includes a plurality of training pictures and the labeled text content corresponding to each training picture, and the plurality of training pictures include noise training pictures and non-noise training pictures.
[0151] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0152] Use the above training picture set as the input of an initial text recognition model, use the labeled text content corresponding to each of the above training pictures as the reference output of the initial text recognition model, train the initial text recognition model for a preset number of times to obtain an intermediate text recognition model; according to the picture quality category of each of the above training pictures, detect whether each of the above training pictures is a noise training picture; if the above training picture is a noise training picture, then continue to train the intermediate text recognition model with other above non-noise training pictures to obtain the above text recognition model.
[0153] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0154] Input each training picture in the above training picture set into the above initial text recognition model to obtain the predicted text content corresponding to each of the above training pictures; according to the predicted text content and the corresponding labeled text content of the above training picture, modify the labeled text content of the above training picture to obtain the new labeled text content corresponding to the above training picture; train the initial text recognition model for a preset number of times according to the new labeled text content corresponding to the above training picture to obtain an intermediate text recognition model.
[0155] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0156] Calculate a first loss between the predicted text content of the above training images and the corresponding labeled text content; modify the labeled text content of the above training images according to the above first loss to obtain the new labeled text content corresponding to the above training images.
[0157] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0158] Compare the above first loss with a preset first loss threshold; if the above first loss is less than the above first loss threshold, modify the labeled text content of the above training images to obtain the new labeled text content corresponding to the above training images.
[0159] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0160] Calculate a second loss between the new labeled text content corresponding to the above training images and the predicted text content corresponding to the above training images; train the above initial text recognition model a preset number of times according to the above first loss and the above second loss to obtain the above intermediate text recognition model.
[0161] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0162] Input each of the above training images into a preset classifier for classification to obtain the image quality category corresponding to the above training images; wherein, the above preset classifier is trained based on a first training image set, and the first training image set includes first training images and the labeled image quality categories corresponding to the first training images; if the image quality category corresponding to the above training images is the first category and the first loss corresponding to the above training images is not less than the above first loss threshold, determine the above training images as noise training images.
[0163] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0164] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0165] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A text recognition method, characterized in that, The method includes: Obtaining the picture to be recognized; Inputting the picture to be recognized into a preset text recognition model for text recognition to determine the text content on the picture to be recognized; Wherein, the preset text recognition model is trained based on a training picture set, the training picture set includes multiple training pictures and the corresponding labeled text content of each training picture, and the multiple training pictures include noise training pictures and non-noise training pictures; The training method of the text recognition model includes: Inputting each training picture in the training picture set into an initial text recognition model to obtain the predicted text content corresponding to each training picture; Calculating a first loss between the predicted text content of the training picture and the corresponding labeled text content; comparing the first loss with a preset first loss threshold; if the first loss of the training picture is less than the first loss threshold, it indicates that the training picture is a non-noise training picture; If the first loss is less than the first loss threshold, randomly modify one character in the labeled text content of the training picture to obtain the new labeled text content corresponding to the training picture; Performing preset times of training on the initial text recognition model according to the new labeled text content corresponding to the training picture to obtain an intermediate text recognition model; Detecting whether each training picture is a noise training picture through the picture quality category of each training picture and the training result of each training picture during the training process of the intermediate text recognition model; the picture quality category represents the clarity of the picture; If the training picture is a noise training picture, continue to train the intermediate text recognition model with other non-noise training pictures to obtain the text recognition model.
2. The method according to claim 1, characterized in that The detection of whether each training picture is a noise training picture is performed after the intermediate text recognition model is trained.
3. The method according to claim 1, wherein The performing preset times of training on the initial text recognition model according to the new labeled text content corresponding to the training picture to obtain an intermediate text recognition model includes: Calculating a second loss between the new labeled text content corresponding to the training picture and the predicted text content corresponding to the training picture; Performing preset times of training on the initial text recognition model according to the first loss and the second loss to obtain the intermediate text recognition model.
4. The method according to claim 1, characterized in that, The detecting whether each training picture is a noise training picture according to the picture quality category of each training picture includes: Inputting each training picture into a preset classifier for classification to obtain the picture quality category corresponding to the training picture; wherein, the preset classifier is trained based on a first training picture set, and the first training picture set includes first training pictures and the corresponding labeled picture quality categories of the first training pictures; If the picture quality category corresponding to the training picture is the first category and the first loss corresponding to the training picture is not less than the first loss threshold, determine that the training picture is a noise training picture.
5. The method according to claim 4, characterized in that, The classifier is a classification branch in the text recognition model.
6. A text recognition device, characterized in that, The device includes an acquisition module and an identification module, where: The acquisition module is used to acquire the picture to be identified; The identification module is used to input the picture to be identified into a preset text recognition model for text recognition, and determine the text content on the picture to be identified; where, the preset text recognition model is trained based on a training picture set, and the training picture set includes multiple training pictures and the corresponding labeled text content for each training picture, and the multiple training pictures include noisy training pictures and non-noisy training pictures; The device further includes a model training module, and the model training module includes an initial training unit, a detection unit and a subsequent training unit, where the initial training unit further includes an identification subunit, a modification subunit and an initial training subunit; The identification subunit is used to input each of the training pictures in the training picture set into an initial text recognition model to obtain the predicted text content corresponding to each of the training pictures; The modification subunit is specifically used to calculate the first loss between the predicted text content of the training picture and the corresponding labeled text content; compare the first loss with a preset first loss threshold; if the first loss of the training picture is less than the first loss threshold, it indicates that the training picture is the non-noisy training picture; if the first loss is less than the first loss threshold, randomly modify one character in the labeled text content of the training picture to obtain the new labeled text content corresponding to the training picture; The initial training subunit is used to perform a preset number of trainings on the initial text recognition model according to the new labeled text content corresponding to the training picture to obtain an intermediate text recognition model; The detection unit is used to detect whether each of the training pictures is a noisy training picture through the picture quality category of each training picture and the training result of each training picture during the training process of the intermediate text recognition model; the picture quality category represents the clarity of the picture; The subsequent training unit is used to, when the training picture is a noisy training picture, continue to train the intermediate text recognition model with other non-noisy training pictures to obtain the text recognition model.
7. The device according to claim 6, characterized in that, The initial training subunit is specifically used for: Calculating the second loss between the new labeled text content corresponding to the above-mentioned training picture and the predicted text content corresponding to the above-mentioned training picture; performing a preset number of trainings on the initial text recognition model according to the above-mentioned first loss and the above-mentioned second loss to obtain the above-mentioned intermediate text recognition model.
8. The device according to claim 7, characterized in that, The device further includes a classification subunit and a determination subunit, where: The classification subunit is used to input each of the above-mentioned training pictures into a preset classifier for classification to obtain the picture quality category corresponding to the above-mentioned training picture; where, the preset classifier is trained based on a first training picture set, and the first training picture set includes first training pictures and the corresponding labeled picture quality categories for the first training pictures; The determining subunit is configured to determine the training picture as a noise training picture when the picture quality category corresponding to the training picture is the first category and the first loss corresponding to the training picture is not less than the first loss threshold.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Label correction method and device for sample picture, equipment and storage medium
CN111382798A
Noise sample recognition method, device and apparatus for pedestrian re-recognition, and storage medium
CN111414952A
Text detection model training method and apparatus, text region determination method and apparatus, and text content determination method and apparatus
WO2020221298A1