A method, device, electronic device, and storage medium for rejecting recognition of text line noise
By introducing a text line noise refusal model in OCR technology, identifying and processing text-like noise introduced during text line detection, the problem of noisy disordered text output is solved, and the recognition rate and user experience are improved.
Patent Information
- Application Number
- CN202210646919.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-06-09
AI Technical Summary
In OCR technology, text-like noise is easily introduced during text line detection, resulting in the text line recognition model outputting noise disordered text, affecting the user experience.
A text line noise refusal method is adopted. By obtaining the text line image to be detected, the text line detection model is input to obtain the text line image area to be identified, and then inputting it into the well-trained text line noise refusal model for processing. The model is trained based on image area samples with text noise and without text noise and corresponding text annotations, including target convolutional recurrent networks and confidence scoring networks, and is trained using a joint learning mutual supervision strategy.
By identifying the image area with text-like noise and outputting blank recognition results, the existing recognition model is avoided to output noise disordered text messages, improve the recognition rate of text images containing text-like noise, and improve the user experience.
Smart Images

Figure CN114973270B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method for rejecting text line noise recognition, a device for rejecting text line noise recognition, an electronic device, and a computer-readable storage medium. Background Art
[0002] In OCR (Optical Character Recognition) technology, the recognition of text images mainly includes two steps. First, text line detection is performed, and then text line recognition is carried out. However, when performing text line detection on text images, it is easy to introduce text-like noises such as musical note curves and low-resolution texts that are unrecognizable by the human eye, resulting in the text line recognition model outputting noisy disordered texts after recognizing the text-like noises, which affects the user experience. Summary of the Invention
[0003] In view of the above problems, embodiments of the present invention are proposed to provide a method for rejecting text line noise recognition, a device for rejecting text line noise recognition, an electronic device, and a computer-readable storage medium that overcome the above problems or at least partially solve the above problems.
[0004] To solve the above problems, embodiments of the present invention disclose a method for rejecting text line noise recognition, and the method includes:
[0005] Obtain a text line image to be detected;
[0006] Input the text line image to be detected into a text line detection model to obtain a text line image region to be recognized; the text line image region to be recognized includes an image region with text-like noise;
[0007] Input the text line image region to be recognized into a text line noise rejection model for processing to obtain a recognition result for the text line image region to be recognized; wherein, the text line noise rejection model is trained based on image region samples with text-like noise, image region samples without text-like noise, and text annotations corresponding to the samples; the text line noise rejection model includes a target convolutional recurrent network and a confidence scoring network connected to the target convolutional recurrent network, and is trained using a co-learning mutual supervision strategy;
[0008] If the text line image region to be recognized is an image region without text-like noise, output a text line recognition result and the confidence of the text line recognition result;
[0009] If the text line image region to be recognized is an image region with text-like noise, output a blank recognition result.
[0010] Optionally, the target convolutional recurrent network includes a convolutional network and a recurrent network; the process of inputting the text line image region to be recognized into the text line noise rejection recognition model to obtain the recognition result for the text line image region to be recognized includes:
[0011] Extracting the convolutional features of the text line image region to be recognized through the convolutional network; extracting the text line sequence features of the text line image region to be recognized through the recurrent network based on the convolutional features; performing decoding processing on the text line sequence features through the connectionist temporal classification algorithm CTC to obtain the recognition result for the text line image region to be recognized.
[0012] Optionally, if the text line image region to be recognized is an image region without text-like noise, outputting the text line recognition result and the confidence of the text line recognition result includes:
[0013] If the text line image region to be recognized is an image region without text-like noise, inputting the text line recognition result for the text line image region to be recognized into the confidence scoring network to obtain the confidence of the text line recognition result;
[0014] Outputting the text line recognition result and the confidence of the text line recognition result.
[0015] Optionally, the text line noise rejection recognition model is trained in the following manner:
[0016] Obtaining a text line image region sample and the text annotation corresponding to the text line image region sample; the text line image region sample includes an image region sample with text-like noise and an image region sample without text-like noise;
[0017] Taking the text line image region sample as the input of the text line noise rejection recognition model; the text line noise rejection recognition model includes a plurality of preset convolutional recurrent networks and a confidence scoring network connected to each preset convolutional recurrent network; each preset convolutional recurrent network includes a convolutional network and a recurrent network; extracting the convolutional features of the text line image region sample through the convolutional network; extracting the text line sequence features of the text line image region sample through the recurrent network based on the convolutional features; performing decoding processing on the text line sequence features through the connectionist temporal classification algorithm CTC, and outputting the CTC greedy decoding result;
[0018] Training the text line noise rejection recognition model based on the output of each preset convolutional recurrent network by adopting a co-learning mutual supervision strategy.
[0019] Optionally, the text annotation includes the true label corresponding to the text line image region sample; training the text line noise rejection model by adopting a co-learning and mutual supervision strategy based on the outputs of the respective preset convolutional recurrent networks includes:
[0020] Performing mutual learning based on the convolutional features extracted by the convolutional network and the text line sequence features extracted by the recurrent network in the respective preset convolutional recurrent networks, so that the CTC greedy decoding results output by the respective preset convolutional recurrent networks approach the true label;
[0021] Calculating the edit distance between the CTC greedy decoding results output by the respective preset convolutional recurrent networks and the true label, and using it as the supervision label of the confidence scoring network connected to the respective preset convolutional recurrent networks;
[0022] Fitting a confidence interval based on the supervision label through the confidence scoring network;
[0023] Calculating the loss by using CTC, and adjusting the parameters of the text line noise rejection model to train the text line noise rejection model.
[0024] Optionally, the method further includes:
[0025] Regarding multiple preset convolutional recurrent networks as the target convolutional recurrent network;
[0026] Or, obtaining a text line image region test sample;
[0027] Inputting the text line image region test sample into a pre-trained text line noise rejection model for testing, and obtaining test results corresponding to the respective preset convolutional recurrent networks and the confidence scoring network connected to the respective preset convolutional recurrent networks;
[0028] Regarding the preset convolutional recurrent network with the highest test result score as the target convolutional recurrent network.
[0029] An embodiment of the present invention also discloses a text line noise rejection device, and the device includes:
[0030] An acquisition module, configured to acquire a text line image to be detected;
[0031] A detection module, configured to input the text line image to be detected into a text line detection model to obtain a text line image region to be recognized; the text line image region to be recognized includes an image region with class text noise;
[0032] An identification module, configured to input the text line image region to be identified into a text line noise rejection recognition model for processing, and obtain an identification result for the text line image region to be identified; wherein, the text line noise rejection recognition model is trained based on image region samples with text-like noise, image region samples without text-like noise, and the text annotations corresponding to the samples; the text line noise rejection recognition model includes a target convolutional recurrent network and a confidence scoring network connected to the target convolutional recurrent network, and is trained using a co-learning mutual supervision strategy;
[0033] A text line output module, configured to output a text line recognition result and the confidence of the text line recognition result if the text line image region to be identified is an image region without text-like noise;
[0034] A blank output module, configured to output a blank recognition result if the text line image region to be identified is an image region with text-like noise.
[0035] Optionally, the target convolutional recurrent network includes a convolutional network and a recurrent network; the identification module includes:
[0036] A feature extraction sub-module, configured to extract convolutional features of the text line image region to be identified through the convolutional network; extract text line sequence features of the text line image region to be identified based on the convolutional features through the recurrent network; and perform decoding processing on the text line sequence features through a connectionist temporal classification algorithm CTC to obtain an identification result for the text line image region to be identified.
[0037] Optionally, the text line output module includes:
[0038] A scoring sub-module, configured to input the text line recognition result for the text line image region to be identified into the confidence scoring network if the text line image region to be identified is an image region without text-like noise, and obtain the confidence of the text line recognition result;
[0039] A result output sub-module, configured to output the text line recognition result and the confidence of the text line recognition result.
[0040] Optionally, the text line noise rejection recognition model is trained through the following modules:
[0041] A sample acquisition module, configured to acquire text line image region samples and the text annotations corresponding to the text line image region samples; the text line image region samples include image region samples with text-like noise and image region samples without text-like noise;
[0042] A decoding result output module, configured to use the text line image region sample as the input of a text line noise rejection recognition model; the text line noise rejection recognition model includes a plurality of preset convolutional recurrent networks and a confidence score network connected to each preset convolutional recurrent network; each of the preset convolutional recurrent networks includes a convolutional network and a recurrent network; the convolutional network is used to extract the convolutional features of the text line image region sample; the recurrent network is used to extract the text line sequence features of the text line image region sample based on the convolutional features; through the connectionist temporal classification algorithm CTC, decode the text line sequence features to output the CTC greedy decoding result;
[0043] A training module, configured to train the text line noise rejection recognition model based on the outputs of the respective preset convolutional recurrent networks by using a co-learning and mutual supervision strategy.
[0044] Optionally, the text annotation includes the true label corresponding to the text line image region sample; the training module includes:
[0045] A mutual learning sub-module, configured to perform mutual learning based on the convolutional features extracted by the convolutional network and the text line sequence features extracted by the recurrent network in each of the preset convolutional recurrent networks, so that the CTC greedy decoding results output by each of the preset convolutional recurrent networks approach the true label;
[0046] A supervised label determination sub-module, configured to calculate the edit distance between the CTC greedy decoding result output by each of the preset convolutional recurrent networks and the true label as the supervised label of the confidence score network connected to each of the preset convolutional recurrent networks;
[0047] A fitting sub-module, configured to fit a confidence interval through the confidence score network based on the supervised label;
[0048] A parameter adjustment module, configured to calculate the loss by using CTC and adjust the parameters of the text line noise rejection recognition model to train the text line noise rejection recognition model.
[0049] Optionally, the apparatus further includes:
[0050] A target network determination module, configured to use the plurality of preset convolutional recurrent networks as the target convolutional recurrent networks;
[0051] Or, a test sample acquisition module, configured to acquire a text line image region test sample;
[0052] A test module, configured to input the text line image region test sample into a pre-trained text line noise rejection model for testing, and obtain test results corresponding to each of the preset convolutional recurrent networks and the confidence scoring networks connected to the preset convolutional recurrent networks respectively;
[0053] A network determination module, configured to use the preset convolutional recurrent network with the highest test result score as the target convolutional recurrent network.
[0054] An embodiment of the present invention also discloses an electronic device, including: a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the steps of the text line noise rejection method described above are implemented.
[0055] An embodiment of the present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the text line noise rejection method described above are implemented.
[0056] The embodiments of the present invention have the following advantages:
[0057] In the embodiment of the present invention, by inputting a text line image region to be recognized, which includes an image region with text-like noise, into a text line noise rejection model trained based on image region samples with text-like noise, image region samples without text-like noise, and text annotations corresponding to the samples for processing, an identification result for the text line image region to be recognized is obtained; if the image region is an image region without text-like noise, a text line recognition result and the corresponding confidence are output; if the image region is an image region with text-like noise, a blank recognition result is output. Thus, the image region with text-like noise can be recognized by the text line noise rejection model and a blank recognition result can be output, avoiding the output of noisy disordered text caused by using the existing recognition model, thereby improving the recognition rate on text images containing text-like noise and further enhancing the user experience. Description of the Drawings
[0058] Figure 1 is a flowchart of the steps of a text line noise rejection method provided by an embodiment of the present invention;
[0059] Figure 2 is a schematic structural diagram of a text line noise rejection model provided by an embodiment of the present invention;
[0060] Figure 3 is a flowchart of the steps of a training method for a text line noise rejection model in an embodiment of the present invention;
[0061] Figure 4It is a structural block diagram of a text line noise rejection device provided by an embodiment of the present invention. Specific embodiments
[0062] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0063] In OCR technology, when performing text line detection on a text image, it is easy to introduce text-like noises such as musical note curves and low-resolution texts that are unrecognizable to the human eye, resulting in the text line recognition model outputting noisy and disordered texts after recognizing the text-like noises. Currently, the text line recognition model is usually trained based on the noises existing in the original text image itself, such as image bending deformation, a small amount of background noise, and image blurring. Therefore, it can only recognize the noises existing in the original text image itself, and usually outputs disordered and garbled texts for the text-like noises introduced by text line detection.
[0064] The core concept of the embodiment of the present invention is to input the text line image region to be recognized, including the image region with text-like noise, into a text line noise rejection model trained based on the image region samples with text-like noise, the image region samples without text-like noise, and the text annotations corresponding to the samples for processing, to obtain the recognition result for the text line image region to be recognized; if the image region is an image region without text-like noise, the text line recognition result and the corresponding confidence level are output; if the image region is an image region with text-like noise, a blank recognition result is output. Thus, the image region with text-like noise can be recognized by the text line noise rejection model and a blank recognition result can be output, avoiding the output of noisy and disordered texts caused by using the existing recognition model, thereby improving the recognition rate on the text image containing text-like noise and further enhancing the user experience.
[0065] Refer to Figure 1 , which shows a step flowchart of a text line noise rejection method provided by an embodiment of the present invention. The method may specifically include the following steps:
[0066] Step 101, obtain the text line image to be detected.
[0067] OCR recognition technology obtains the text image information on paper through optical input methods such as scanning and photography, and uses various pattern recognition algorithms to analyze the morphological characteristics of the text. It can convert bills, newspapers, books, manuscripts, and other printed materials into text image information, and then use image recognition technology to convert the text image information into computer input that can be used.
[0068] The text line image can be an image with at least one line of text. In an embodiment of the present invention, the text line noise rejection method can be applied to a scenario where a server obtains a text line image to be detected for detection, performs text line noise rejection on the obtained text line image region to be recognized, and thus outputs a blank recognition result for an image region with pseudo-character noise.
[0069] Step 102: Input the text line image to be detected into a text line detection model to obtain a text line image region to be recognized; the text line image region to be recognized includes an image region with pseudo-character noise.
[0070] In an embodiment of the present invention, when the obtained text line image to be detected is input into the text line detection model, the text line detection model can locate the region where the text is located in the text line image to be detected, and perform cutting according to the located region to obtain the text line image region to be recognized.
[0071] The text line image region to be recognized can include an image region with pseudo-character noise and an image region without pseudo-character noise. Pseudo-character noise can be noise introduced during the process of the text line detection model processing the input text line image to be detected. For example, pseudo-character noise can be musical note curves, low-resolution text that cannot be recognized by the human eye, etc.
[0072] Step 103: Input the text line image region to be recognized into a text line noise rejection model for processing to obtain a recognition result for the text line image region to be recognized; wherein, the text line noise rejection model is trained based on an image region sample with pseudo-character noise, an image region sample without pseudo-character noise, and the text annotation corresponding to the sample; the text line noise rejection model includes a target convolutional recurrent network and a confidence scoring network connected to the target convolutional recurrent network, and is trained using a co-learning mutual supervision strategy.
[0073] In an embodiment of the present invention, the text line image region to be recognized can be input into a pre-trained text line noise rejection model, and the trained text line noise rejection model can recognize the text line image region to be recognized, so as to determine whether the text line image region to be recognized is an image region with pseudo-character noise according to the recognition result.
[0074] A text line noise rejection model can be trained based on image region samples with text-like noise, image region samples without text-like noise, and text annotations corresponding to the image region samples. Among them, the image region samples without text-like noise can be training samples where the image has other noises but no text-like noise. The other existing noises can include image bending deformation, a small amount of background noise, image blurring, etc. The image region samples with text-like noise can be training samples collected or simulated by introducing text-like noise to the text line detection model. For example, the text-like noise can be musical note curves, low-resolution text that cannot be recognized by the human eye, etc.
[0075] The text line noise rejection model can include a target convolutional recurrent network and a confidence scoring network connected to the target convolutional recurrent network, and is trained using a co-learning and mutual supervision strategy. Among them, the target convolutional recurrent network can be one or more convolutional recurrent networks determined from multiple trained convolutional recurrent networks. Specifically, after the model training is completed, the trained model can be tested, and the optimal convolutional recurrent network can be determined as the target convolutional recurrent network according to performance requirements; in another example, after the model training is completed, all convolutional recurrent networks can be used as the target convolutional recurrent network, and the output results of all convolutional recurrent networks can be weighted and averaged for the final output.
[0076] Step 104, if the image region of the text line to be recognized is an image region without text-like noise, output the text line recognition result and the confidence of the text line recognition result.
[0077] If the image region of the text line to be recognized is an image region without text-like noise, then this region is a normal image region. At this time, output the text line recognition result and the confidence of this text line recognition result.
[0078] Step 105, if the image region of the text line to be recognized is an image region with text-like noise, output a blank recognition result.
[0079] If the image region of the text line to be recognized is an image region with text-like noise, then text-like noise is introduced when performing text line detection on the text line image to be detected. At this time, output a blank recognition result.
[0080] In an optional embodiment, step 103 may include: extracting convolutional features of the image region of the text line to be recognized through the convolutional network; extracting text line sequence features of the image region of the text line to be recognized based on the convolutional features through the recurrent network; and decoding the text line sequence features through the connectionist temporal classification algorithm CTC to obtain a recognition result for the image region of the text line to be recognized.
[0081] Reference Figure 2 As shown, it is a schematic structural diagram of a text line noise rejection recognition model provided by an embodiment of the present invention. The text line noise rejection recognition model may include a target convolutional recurrent network and a confidence scoring network connected to the target convolutional recurrent network. The target convolutional recurrent network may include a convolutional network and a recurrent network. Inputting the text line image region to be recognized into the text line noise rejection recognition model, convolutional features of the text line image region to be recognized can be extracted through the convolutional network; based on the convolutional features of the text line image region to be recognized, the text line sequence features of the image region can be extracted through the recurrent network; through the connectionist temporal classification algorithm CTC (Connectionist Temporal Classification), decoding processing is performed on the text line sequence features to obtain the recognition result for the text line image region to be recognized.
[0082] In an alternative embodiment, step 104 may include the following sub-steps S11 - S12:
[0083] Sub-step S11, if the text line image region to be recognized is an image region without text-like noise, then input the text line recognition result for the text line image region to be recognized into the confidence scoring network to obtain the confidence of the text line recognition result.
[0084] Sub-step S12, output the text line recognition result and the confidence of the text line recognition result.
[0085] In the embodiment of the present invention, when the text image region to be recognized is an image region without text-like noise, the text line recognition result for the text line image region to be recognized can be input into the confidence scoring network. The confidence scoring network performs confidence scoring on the text line recognition result to obtain the confidence of the text line recognition result, and outputs the text line recognition result and the confidence of the text line recognition result. Thus, when it is determined that the text image region to be recognized does not have text-like noise, the text line recognition result and its confidence are output to improve the reliability of the confidence of the text line recognition result.
[0086] In an embodiment of the present invention, by inputting a text line image region including an image region with text-like noise into a text line noise rejection recognition model trained based on an image region sample with text-like noise, an image region sample without text-like noise, and the text annotation corresponding to the sample, a recognition result for the text line image region to be recognized is obtained; if the image region is an image region without text-like noise, a text line recognition result and the corresponding confidence level are output; if the image region is an image region with text-like noise, a blank recognition result is output. Thus, an image region with text-like noise can be recognized by the text line noise rejection recognition model and a blank recognition result can be output, avoiding the output of noisy disordered text caused by using an existing recognition model, thereby improving the recognition rate on a text image containing text-like noise and further improving the user experience.
[0087] Refer to Figure 3 , which shows a flowchart of a training method for a text line noise rejection recognition model in an embodiment of the present invention. The training method for the text line noise rejection recognition model includes:
[0088] Step 301, obtain a text line image region sample and the text annotation corresponding to the text line image region sample; the text line image region sample includes an image region sample with text-like noise and an image region sample without text-like noise.
[0089] The image region sample without text-like noise can be a training sample in which the image has other noises but no text-like noise. Among them, the other noises existing in the image can include image bending and deformation, the presence of a small amount of background noise, image blurring, etc.
[0090] The image region sample with text-like noise can be a training sample obtained by collecting or simulating the text-like noise introduced by a text line detection model. For example, the text-like noise can be musical note curves, low-resolution text that cannot be recognized by the human eye, etc.
[0091] In one example, before obtaining the text line image region sample, 5 million lines of image region samples without text-like noise and 500,000 lines of image region samples with text-like noise can be collected or simulated, and 95% of the samples are randomly selected as text line image region training samples for model training, and 5% of the samples are used as test samples for model performance testing.
[0092] The sampling ratio of the image region training sample without text-like noise to the image region training sample with text-like noise in the training batch can be set. Exemplarily, if the sampling ratio is set to 10:1 and the size of each training batch is 66, 60 lines of image region training samples without text-like noise and 6 lines of image region training samples with text-like noise can be obtained for training.
[0093] Those skilled in the art should understand that the sampling ratio of the training samples set in the above training batch is only an example of the present invention. Those skilled in the art can set different sampling ratios of the training samples according to actual needs, and the present application does not limit this here.
[0094] Step 302: Use the text line image region training sample as the input of the text line noise rejection model; the text line noise rejection model includes a plurality of preset convolutional recurrent networks and a confidence score network connected to each preset convolutional recurrent network; each preset convolutional recurrent network includes a convolutional network and a recurrent network; extract the convolutional features of the text line image region sample through the convolutional network; extract the text line sequence features of the text line image region training sample based on the convolutional features through the recurrent network; perform decoding processing on the text line sequence features through the connectionist temporal classification algorithm CTC, and output the CTC greedy decoding result.
[0095] Before inputting the text line image region sample into the text line noise rejection model for training, each preset convolutional recurrent network can be initialized respectively. Exemplarily, a CNN (Convolutional Neural Network) composed of mobilev3 and an RNN (Recurrent Neural Network) composed of a bidirectional LSTM (Long Short-Term Memory) can be used as the preset convolutional recurrent network; a 2-layer fully connected layer can be used as the confidence score network.
[0096] In the embodiment of the present invention, the text line noise rejection model can include a plurality of preset convolutional recurrent networks and a confidence score network connected to each preset convolutional recurrent network; the preset convolutional recurrent network can include a convolutional network and a recurrent network. Input the text line image region sample into the convolutional network, and the convolutional network can extract the convolutional features of the text line image region training sample; then the recurrent network extracts the text line sequence features of the image region training sample based on the convolutional features of the text line image region training sample; finally, perform decoding processing on the text line sequence features through the connectionist temporal classification algorithm CTC, and output the CTC greedy decoding result.
[0097] Step 303: Based on the outputs of the respective preset convolutional recurrent networks, train the text line noise rejection model using a co-learning mutual supervision strategy.
[0098] Inputting the text line image region sample into the text line noise rejection recognition model, multiple outputs of each preset convolutional recurrent network in the model for the text line image region sample can be obtained. Based on the multiple outputs, a co-learning and mutual supervision strategy can be adopted to train the text line noise rejection recognition model.
[0099] In the embodiments of the present invention, by adopting multiple preset convolutional recurrent networks for mutual learning, it is beneficial to increase the convergence performance of the network and improve the recognition rate of text images containing pseudo-character noise.
[0100] In an optional embodiment, the text annotation includes the true label corresponding to the text line image region sample, and step 303 may include the following sub-steps S21 - S24:
[0101] Sub-step S21, perform mutual learning based on the convolutional features extracted by the convolutional network and the text line sequence features extracted by the recurrent network in each of the preset convolutional recurrent networks, so that the CTC greedy decoding results output by each of the preset convolutional recurrent networks approach the true label.
[0102] Specifically, the Kullback-Leibler divergence (KL divergence, also known as relative entropy) can be used to perform mutual learning based on the convolutional features extracted by the convolutional networks in each of the preset convolutional recurrent networks and the text line sequence features extracted by the recurrent networks in each of the preset convolutional recurrent networks, so that the CTC greedy decoding results output by each of the preset convolutional recurrent networks approach the true label.
[0103] Sub-step S22, calculate the edit distance between the CTC greedy decoding result output by each of the preset convolutional recurrent networks and the true label, and use it as the supervision label of the confidence scoring network corresponding to each of the preset convolutional recurrent networks.
[0104] The edit distance is a quantitative measurement of the difference degree between two strings. The measurement method is to see how many times of processing are required at least to change one string into another string. In the embodiments of the present invention, the edit distance between the CTC greedy decoding result output by each of the preset convolutional recurrent networks and the true label corresponding to the text line image sample can be calculated respectively, and the edit distance is used as the supervision label of the confidence scoring network corresponding to each of the preset convolutional recurrent networks.
[0105] Sub-step S23, through the confidence scoring network, fit the confidence interval based on the supervision label.
[0106] After determining the supervision label for each of the preset convolutional recurrent networks, the confidence interval can be fitted based on the supervision label through the confidence scoring network connected to each of the preset convolutional recurrent networks.
[0107] Sub-step S24: Calculate the loss using CTC, and adjust the parameters of the text line noise rejection model to train the text line noise rejection model.
[0108] Specifically, the parameters of the recurrent network in the text noise rejection model can be adjusted by calculating the gradient according to the CTC criterion to train the text line noise rejection model.
[0109] In an alternative embodiment, the method may further include: using multiple preset convolutional recurrent networks as the target convolutional recurrent network; or, obtaining text line image region test samples; inputting the text line image region test samples into a pre-trained text line noise rejection model for testing to obtain test results corresponding to each of the preset convolutional recurrent networks and the confidence scoring networks connected to the preset convolutional recurrent networks; and using the preset convolutional recurrent network with the highest test result score as the target convolutional recurrent network.
[0110] In one example, after the text line noise rejection model is trained, multiple preset convolutional recurrent networks can be used as the target convolutional recurrent network. After inputting the text line image region to be recognized into the target convolutional recurrent network, i.e., multiple preset convolutional recurrent networks, multiple output results can be obtained. The result obtained by weighted averaging the multiple output results can be used as the final output result.
[0111] In another example, text line image region test samples can be obtained, and the text line image region test samples can be input into the trained text line noise rejection model for testing to obtain test results corresponding to each preset convolutional recurrent network. The preset convolutional recurrent network with the highest test result score, i.e., the preset convolutional recurrent network with the best test performance among the multiple preset convolutional recurrent networks, can be used as the target convolutional recurrent network.
[0112] In the embodiments of the present invention, by adding image region samples with pseudo-character noise to train the text line noise rejection model during the training process, the recognition rate of text images containing pseudo-character noise can be effectively improved in practical applications, blank recognition results can be output for image regions with pseudo-character noise, a series of garbled characters can be reduced, and the user experience can be improved.
[0113] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0114] Refer to Figure 4 , which shows the structural block diagram of a text line noise rejection device provided by an embodiment of the present invention. Specifically, it may include the following modules:
[0115] An acquisition module 401, configured to acquire an image of a text line to be detected;
[0116] A detection module 402, configured to input the image of the text line to be detected into a text line detection model to obtain an image region of the text line to be recognized; the image region of the text line to be recognized includes an image region with text-like noise;
[0117] An identification module 403, configured to input the image region of the text line to be recognized into a text line noise rejection model for processing to obtain an identification result for the image region of the text line to be recognized; wherein, the text line noise rejection model is trained based on image region samples with text-like noise, image region samples without text-like noise, and text annotations corresponding to the samples; the text line noise rejection model includes a target convolutional recurrent network and a confidence scoring network connected to the target convolutional recurrent network, and is trained using a co-learning mutual supervision strategy;
[0118] A text line output module 404, configured to output a text line recognition result and the confidence of the text line recognition result if the image region of the text line to be recognized is an image region without text-like noise;
[0119] A blank output module 405, configured to output a blank recognition result if the image region of the text line to be recognized is an image region with text-like noise.
[0120] In an embodiment of the present invention, the target convolutional recurrent network includes a convolutional network and a recurrent network; the identification module includes:
[0121] A feature extraction sub-module, configured to extract convolutional features of the image region of the text line to be recognized through the convolutional network; extract text line sequence features of the image region of the text line to be recognized based on the convolutional features through the recurrent network; and perform decoding processing on the text line sequence features through a connectionist temporal classification algorithm CTC to obtain an identification result for the image region of the text line to be recognized.
[0122] In an embodiment of the present invention, the text line output module includes:
[0123] A scoring sub-module, configured to input the text line recognition result for the image region of the text line to be recognized into the confidence scoring network to obtain the confidence of the text line recognition result if the image region of the text line to be recognized is an image region without text-like noise;
[0124] A result output sub-module, configured to output the text line recognition result and the confidence of the text line recognition result.
[0125] In an embodiment of the present invention, the text line noise rejection recognition model is trained through the following modules:
[0126] A sample acquisition module, configured to acquire text line image region samples and text annotations corresponding to the text line image region samples; the text line image region samples include image region samples with text-like noise and image region samples without text-like noise;
[0127] A decoding result output module, configured to use the text line image region samples as the input of the text line noise rejection recognition model; the text line noise rejection recognition model includes a plurality of preset convolutional recurrent networks and a confidence scoring network connected to each preset convolutional recurrent network; each of the preset convolutional recurrent networks includes a convolutional network and a recurrent network; the convolutional features of the text line image region samples are extracted through the convolutional network; the text line sequence features of the text line image region samples are extracted through the recurrent network based on the convolutional features; through the connectionist temporal classification (CTC) algorithm, the text line sequence features are decoded to output a CTC greedy decoding result;
[0128] A training module, configured to train the text line noise rejection recognition model based on the outputs of the respective preset convolutional recurrent networks by adopting a co-learning mutual supervision strategy.
[0129] In an embodiment of the present invention, the text annotation includes the true label corresponding to the text line image region sample; the training module includes:
[0130] A mutual learning sub-module, configured to perform mutual learning based on the convolutional features extracted by the convolutional networks in the respective preset convolutional recurrent networks and the text line sequence features extracted by the recurrent networks, so that the CTC greedy decoding results output by the respective preset convolutional recurrent networks approach the true label;
[0131] A supervised label determination sub-module, configured to calculate the edit distance between the CTC greedy decoding result output by each preset convolutional recurrent network and the true label, and use it as the supervised label of the confidence scoring network connected to each preset convolutional recurrent network;
[0132] A fitting sub-module, configured to fit a confidence interval through the confidence scoring network based on the supervised label;
[0133] A parameter adjustment module, configured to calculate the loss by using CTC and adjust the parameters of the text line noise rejection recognition model to train the text line noise rejection recognition model.
[0134] In an embodiment of the present invention, the device further includes:
[0135] A target network determination module, configured to use multiple preset convolutional recurrent networks as target convolutional recurrent networks;
[0136] Or, a test sample acquisition module, configured to acquire a test sample of a text line image region;
[0137] A test module, configured to input the test sample of the text line image region into a pre-trained text line noise rejection recognition model for testing, and obtain test results corresponding to each of the preset convolutional recurrent networks and a confidence scoring network connected to each of the preset convolutional recurrent networks;
[0138] A network determination module, configured to use the preset convolutional recurrent network with the highest test result score as the target convolutional recurrent network.
[0139] In an embodiment of the present invention, by inputting a text line image region to be recognized, which includes an image region with text-like noise, into a text line noise rejection recognition model trained based on an image region sample with text-like noise, an image region sample without text-like noise, and text annotations corresponding to the samples for processing, an identification result for the text line image region to be recognized is obtained; if the image region is an image region without text-like noise, a text line recognition result and a corresponding confidence level are output; if the image region is an image region with text-like noise, a blank recognition result is output. Thus, an image region with text-like noise can be recognized by the text line noise rejection recognition model and a blank recognition result can be output, avoiding the output of disordered noise text by using an existing recognition model, thereby improving the recognition rate on text images containing text-like noise and further improving the user experience.
[0140] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment.
[0141] An embodiment of the present invention further provides an electronic device, including:
[0142] A processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, it implements each process of the above-mentioned text line noise rejection recognition method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0143] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-described embodiment of the text line noise rejection method and can achieve the same technical effect. To avoid repetition, it will not be described in detail here.
[0144] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0145] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0146] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0147] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0148] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide for implementing the functions specified in one Figure 1One or more processes and / or boxes Figure 1 Steps of functions specified in one or more boxes.
[0149] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the embodiments of the present invention.
[0150] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.
[0151] The above has introduced in detail a method, device, electronic device and storage medium for rejecting text line noise provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for rejecting recognition of text line noise, characterized in that, the method includes: obtaining an image of a text line to be detected; inputting the image of the text line to be detected into a text line detection model to obtain a text line image region to be recognized; the text line image region to be recognized includes an image region with text-like noise; inputting the text line image region to be recognized into a text line noise rejection model for processing to obtain a recognition result for the text line image region to be recognized; wherein, the text line noise rejection model is trained based on image region samples with text-like noise, image region samples without text-like noise, and text annotations corresponding to the samples; the text line noise rejection model includes a target convolutional recurrent network and a confidence scoring network connected to the target convolutional recurrent network, and is trained using a co-learning mutual supervision strategy; if the text line image region to be recognized is an image region without text-like noise, output the text line recognition result and the confidence of the text line recognition result; if the text line image region to be recognized is an image region with text-like noise, output a blank recognition result; the target convolutional recurrent network includes a convolutional network and a recurrent network; the step of inputting the text line image region to be recognized into the text line noise rejection model for processing to obtain a recognition result for the text line image region to be recognized includes: extracting convolutional features of the text line image region to be recognized through the convolutional network; extracting text line sequence features of the text line image region to be recognized based on the convolutional features through the recurrent network; decoding the text line sequence features through a connectionist temporal classification algorithm CTC to obtain a recognition result for the text line image region to be recognized.
2. The method according to claim 1, characterized in that, the step of if the text line image region to be recognized is an image region without text-like noise, output the text line recognition result and the confidence of the text line recognition result includes: if the text line image region to be recognized is an image region without text-like noise, input the text line recognition result for the text line image region to be recognized into the confidence scoring network to obtain the confidence of the text line recognition result; output the text line recognition result and the confidence of the text line recognition result.
3. The method according to claim 1, characterized in that, the text line noise rejection model is trained in the following manner: obtaining text line image region samples and text annotations corresponding to the text line image region samples; the text line image region samples include image region samples with text-like noise and image region samples without text-like noise; Use the text line image region sample as the input of the text line noise rejection recognition model; the text line noise rejection recognition model includes multiple preset convolutional recurrent networks and a confidence score network connected to each preset convolutional recurrent network; each of the preset convolutional recurrent networks includes a convolutional network and a recurrent network; extract the convolutional features of the text line image region sample through the convolutional network; extract the text line sequence features of the text line image region sample based on the convolutional features through the recurrent network; perform decoding processing on the text line sequence features through the connectionist temporal classification (CTC) algorithm, and output the CTC greedy decoding result; Based on the outputs of the respective preset convolutional recurrent networks, train the text line noise rejection recognition model using a co-learning mutual supervision strategy.
4. The method according to claim 3, wherein, the text annotation includes the true label corresponding to the text line image region sample; the training of the text line noise rejection recognition model using a co-learning mutual supervision strategy based on the outputs of the respective preset convolutional recurrent networks includes: Perform mutual learning based on the convolutional features extracted by the convolutional network and the text line sequence features extracted by the recurrent network in each of the preset convolutional recurrent networks, so that the CTC greedy decoding results output by each of the preset convolutional recurrent networks approach the true label; Calculate the edit distance between the CTC greedy decoding result output by each of the preset convolutional recurrent networks and the true label, and use it as the supervision label of the confidence score network connected to each of the preset convolutional recurrent networks; Fit the confidence interval based on the supervision label through the confidence score network; Use CTC to calculate the loss, adjust the parameters of the text line noise rejection recognition model, and train the text line noise rejection recognition model.
5. The method according to claim 4, wherein, the method further includes: Use multiple preset convolutional recurrent networks as the target convolutional recurrent network; or, obtain a text line image region test sample; Input the text line image region test sample into the pre-trained text line noise rejection recognition model for testing, and obtain test results corresponding to each of the preset convolutional recurrent networks and the confidence score network connected to each of the preset convolutional recurrent networks; Use the preset convolutional recurrent network with the highest test result score as the target convolutional recurrent network.
6. A text line noise rejection recognition device, wherein, the device includes: An acquisition module, configured to acquire a text line image to be detected; A detection module, configured to input the text line image to be detected into a text line detection model to obtain a text line image region to be recognized; the text line image region to be recognized includes an image region with text-like noise; An identification module, configured to input the text line image region to be identified into a text line noise rejection recognition model for processing, and obtain an identification result for the text line image region to be identified; wherein, the text line noise rejection recognition model is trained based on image region samples with text-like noise, image region samples without text-like noise, and text annotations corresponding to the samples; the text line noise rejection recognition model includes a target convolutional recurrent network and a confidence scoring network connected to the target convolutional recurrent network, and is trained using a co-learning mutual supervision strategy; A text line output module, configured to output a text line recognition result and the confidence of the text line recognition result if the text line image region to be identified is an image region without text-like noise; A blank output module, configured to output a blank recognition result if the text line image region to be identified is an image region with text-like noise; The target convolutional recurrent network includes a convolutional network and a recurrent network; the identification module includes: a feature extraction sub-module, configured to extract convolutional features of the text line image region to be identified through the convolutional network; extract text line sequence features of the text line image region to be identified based on the convolutional features through the recurrent network; and perform decoding processing on the text line sequence features through a connectionist temporal classification algorithm CTC to obtain an identification result for the text line image region to be identified.
7. An electronic device, characterized in that, it includes: a processor, a memory, and a computer program stored on the memory and capable of running on the processor, where when the computer program is executed by the processor, the steps of the text line noise rejection recognition method according to any one of claims 1-5 are implemented.
8. A computer-readable storage medium, characterized in that, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the text line noise rejection recognition method according to any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Complex character recognition method based on deep learning
CN104966097A
Text recognition method and device
CN112749695A