Method, apparatus, storage medium and computer device for determining text recognition result

By updating the probability distribution of easily mixed characters in the recognition results, the filtering problem of recognition results caused by easily mixed characters in the prior art is solved, and the accuracy of the recognition results is improved.

CN115147847BActive Publication Date: 2025-07-01JIANGSU SEUIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210885785.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-26
Publication Date
2025-07-01
Estimated Expiration
2042-07-26

AI Technical Summary

Technical Problem

When the prior art scores the confidence level of the recognition results, the ambiguity of easily mixed characters leads to too low overall scores, resulting in the correct recognition results being filtered, reducing the accuracy of the recognition results.

Method used

By OCR recognition of the text image to be recognized, the characters corresponding to the maximum probability value in the probability distribution of each time step are obtained, and whether the character is a character in the pre-configured error-prone character grouping list. If so, the maximum probability value is updated based on the probability value of the error-prone characters paired with the maximum probability value in the error-prone character grouping list, thereby adjusting the maximum probability value of the output character.

Benefits of technology

It effectively avoids the overall score due to the ambiguity of easily mixed characters, reduces the situation where the correct recognition results are filtered, thereby improving the accuracy of the recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147847B_ABST
    Figure CN115147847B_ABST
Patent Text Reader

Abstract

The method, device, storage medium, and computer equipment for determining the text recognition result provided by this application can, when performing OCR recognition on the text image to be recognized, first determine the character corresponding to the extremely high probability value in the probability distribution of each time step, and then, for each time step, determine whether the character corresponding to the extremely high probability value in the probability distribution of this time step is a character in the pre-configured error-prone character grouping list. If not, directly use the character corresponding to the extremely high probability value in the probability distribution of this time step as the output character of this time step; if so, update the extremely high probability value based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution of this time step, and use the character after updating the extremely high probability value as the output character of this time step. In this way, the situation where the overall score is too low due to the ambiguity of easily confused characters can be avoided, thereby effectively improving the accuracy of the recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of text recognition, and in particular, to a method, apparatus, storage medium, and computer device for determining text recognition results. Background Art

[0002] OCR (Optical Character Recognition) refers to a technology that can recognize text to be recognized in an image to be recognized. OCR is widely used in fields such as logistics, medical care, finance, and insurance, and can also be carried on a PDA to achieve more applications. Among them, PDAs include consumer PDAs and industrial PDAs. Consumer PDAs include smartphones, tablets, handheld game consoles, etc.; industrial PDAs are mainly used in fields such as factory manufacturing, logistics warehousing, and outdoor material inspection. Common ones include barcode scanners (also known as gun scanners), RFID readers, POS machines, etc., all of which can be called PDAs. Industrial PDAs can be used in many places with relatively harsh environments, and at the same time, many optimizations have been made for industrial use characteristics, supporting RFID reading and barcode scanning functions, and having an industrial grade of IP54 or above, which are not possessed by consumer handheld terminals.

[0003] Currently, when applying OCR recognition technology to industrial PDAs, generally, a confidence score is given to the recognition result of OCR, and the scoring result is used to detect whether the current recognition result is reliable. The prediction of the usual deep learning CTC algorithm scores by taking the best value for all time series, but some characters are easily confused characters in the dictionary, such as 1 and I, 0 and O, and the ambiguity of easily confused characters will cause the overall score to be too low, which easily filters out some correct recognition results, thereby reducing the accuracy of the recognition result. Summary of the Invention

[0004] The purpose of this application aims to at least solve one of the above technical defects, especially the technical defect that in the prior art, when giving a confidence score to the recognition result, the ambiguity of easily confused characters will cause the overall score to be too low, which easily filters out some correct recognition results, thereby reducing the accuracy of the recognition result.

[0005] This application provides a method for determining text recognition results, and the method includes:

[0006] Perform OCR recognition on the text image to be recognized to obtain the probability distribution of all character categories output at each time step of the text image;

[0007] Traverse the probability distribution of all character categories output at each time step, and obtain the character corresponding to the extremely high probability value in the probability distribution of each time step;

[0008] For each time step, determine whether the character corresponding to the extremely high probability value in the probability distribution of this time step is a character in a pre-configured error-prone character grouping list;

[0009] If so, based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution of this time step, update the extremely high probability value, and use the character after updating the extremely high probability value as the output character of this time step; wherein, the updated extremely high probability value is greater than the extremely high probability value before the update;

[0010] If not, directly use the character corresponding to the extremely high probability value in the probability distribution of this time step as the output character of this time step;

[0011] Concatenate the output characters of all time steps to obtain a character sequence, and decode the character sequence to obtain the final recognition result.

[0012] Optionally, the method further includes:

[0013] According to the recognition scenario corresponding to the text image, correct the recognition result to obtain a corrected result.

[0014] Optionally, the correcting the recognition result according to the recognition scenario corresponding to the text image to obtain a corrected result includes:

[0015] Determine the recognition scenario corresponding to the text image and the list of characters prone to appear corresponding to the recognition scenario;

[0016] Compare each output character in the recognition result with the characters prone to appear in the list of characters prone to appear to determine whether there is an output character in the recognition result that is not included in the list of characters prone to appear;

[0017] If there is, replace the output character not included in the list of characters prone to appear with the error-prone character paired with this output character in the error-prone character grouping list;

[0018] Use the recognition result after replacing the output characters as the corrected result.

[0019] Optionally, the performing OCR recognition on the text image to be recognized to obtain the probability distribution of all character categories output at each time step includes:

[0020] Input the text image to be recognized into a pre-configured CRNN network for OCR recognition, and obtain multiple feature maps by extracting features of the text image through the convolutional layer of the CRNN network;

[0021] After converting each feature map into a feature vector, each feature vector is sequentially input into the recurrent layer of the CRNN network, and the recurrent layer predicts the characters corresponding to each feature vector to obtain the probability distribution of all character categories output by the recurrent layer at each time step.

[0022] Optionally, updating the extremely high probability value based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list at this time step includes:

[0023] Determine the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list, and the probability value of the error-prone character in the probability distribution at this time step;

[0024] Add the probability value of the error-prone character in the probability distribution at this time step to the extremely high probability value, and update the extremely high probability value according to the addition result.

[0025] Optionally, decoding the character sequence to obtain the final recognition result includes:

[0026] Traverse the character sequence to determine whether the character sequence contains a placeholder;

[0027] If not, merge the repeated output characters in consecutive time steps, and use the merged character sequence as the final recognition result;

[0028] If so, merge the repeated output characters in consecutive time steps in the character sequence according to the placeholder, and remove the placeholder in the merged character sequence to obtain the final recognition result.

[0029] Optionally, the method further includes:

[0030] Multiply the extremely high probability values corresponding to each output character in the recognition result to obtain a product result;

[0031] Use the product result as the confidence score of the recognition result.

[0032] This application also provides a device for determining a text recognition result, including:

[0033] A text recognition module, configured to perform OCR recognition on a text image to be recognized, and obtain the probability distribution of all character categories output by the text image at each time step;

[0034] A character acquisition module, configured to traverse the probability distribution of all character categories output at each time step, and acquire the character corresponding to the extremely high probability value in the probability distribution at each time step;

[0035] An error-prone character determination module, configured to determine, for each time step, whether the character corresponding to the extremely high probability value in the probability distribution of this time step is a character in a pre-configured error-prone character grouping list;

[0036] A first character determination module, configured to, if so, update the extremely high probability value based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution of this time step, and use the character after updating the extremely high probability value as the output character of this time step; wherein, the updated extremely high probability value is greater than the extremely high probability value before the update;

[0037] A second character determination module, configured to, if not, directly use the character corresponding to the extremely high probability value in the probability distribution of this time step as the output character of this time step;

[0038] An identification result determination module, configured to splice the output characters of all time steps to obtain a character sequence, and decode the character sequence to obtain the final identification result.

[0039] The present application further provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the method for determining a text recognition result as described in any one of the above embodiments.

[0040] The present application further provides a computer device, including: one or more processors, and a memory;

[0041] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the method for determining a text recognition result as described in any one of the above embodiments are executed.

[0042] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0043] The method, apparatus, storage medium, and computer device for determining a text recognition result provided by this application can, when performing OCR recognition on a text image to be recognized, first obtain the probability distribution of all character categories output at each time step of the text image. Then, this application can determine the character corresponding to the extremely high probability value in the probability distribution of each time step. Next, for each time step, it is determined whether the character corresponding to the extremely high probability value in the probability distribution of this time step is a character in a pre-configured error-prone character grouping list. If not, the character corresponding to the extremely high probability value in the probability distribution of this time step is directly used as the output character of this time step, which can improve the accuracy of the recognition result to a certain extent. If so, the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution of this time step is used to update the extremely high probability value, and the character after updating the extremely high probability value is used as the output character of this time step. Since the extremely high probability value of the output character of this time step is adjusted compared to the original extremely high probability value, when the output characters of all adjusted time steps are concatenated and then decoded, it is possible to avoid the situation where the overall score is too low due to the ambiguity of easily confused characters and the correct recognition result is filtered, thereby effectively improving the accuracy of the recognition result. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0045] Figure 1 It is a schematic flowchart of a method for determining a text recognition result provided by an embodiment of this application;

[0046] Figure 2 It is a schematic diagram of a text image provided by an embodiment of this application;

[0047] Figure 3 It is a display diagram of the recognition result of the text image provided by an embodiment of this application;

[0048] Figure 4 It is a display diagram of the corrected result of the text image provided by an embodiment of this application;

[0049] Figure 5 It is a schematic structural diagram of an apparatus for determining a text recognition result provided by an embodiment of this application;

[0050] Figure 6 It is a schematic internal structure diagram of a computer device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0052] Currently, when applying the OCR recognition technology to industrial PDAs, generally, a confidence score is given to the recognition result of OCR, and the scoring result is used to detect whether the current recognition result is reliable. The prediction of the usual deep learning CTC algorithm scores by taking the best value for all time series. However, some characters are easily confused characters in the dictionary, such as 1 and I, 0 and O. The ambiguity of easily confused characters will cause the overall score to be too low, which easily filters out some correct recognition results, thereby reducing the accuracy of the recognition result.

[0053] Based on this, the present application proposes the following technical solutions. For details, please refer to the following text:

[0054] In one embodiment, as Figure 1 shown, Figure 1 is a schematic flowchart of a method for determining a text recognition result provided by an embodiment of the present application. The present application provides a method for determining a text recognition result, and the method may include:

[0055] S110: Perform OCR recognition on the text image to be recognized to obtain the probability distribution of all character categories output by the text image at each time step.

[0056] In this step, when the text image to be recognized is obtained, the present application can perform OCR recognition on the text image and obtain the probability distribution of all character categories output by the text image at each time step.

[0057] Among them, the text image to be recognized in the present application includes, but is not limited to, text images corresponding to bank card numbers, identity documents, weight information, steel seals, waybills, etc. When the present application performs OCR recognition on the text image, the text block to be recognized in the text image can be obtained. For example, after cropping the text block in the specified area of the text image, the text block to be recognized is obtained, and then OCR recognition is performed on the text block to be recognized. In this way, the influence of the complex background in the text image on the recognition result can be filtered out, and the recognition efficiency of OCR can be improved to a certain extent, saving the waiting time of the user and improving the user experience.

[0058] It can be understood that the present application provides an OCR recognition component, which uses a sight alignment as a reference point and captures several text blocks that appear in the field of view (FOV) in the sample image. Therefore, the present application can select the text block closest to the center point of the text image and the text blocks in the same line as this text block as the text blocks to be recognized.

[0059] Furthermore, when performing OCR recognition on a text image, the present application can select different recognition models for text recognition. However, since the text blocks in the text image have temporal characteristics, therefore, regardless of which recognition model is selected for text recognition, it is necessary to pay attention to the text sequence with temporal characteristics in the text image and predict the text at different positions in the text sequence in turn, so as to obtain the probability distribution of all character categories output at each time step of the text image, and predict the text content of the text image through the probability distribution of all character categories output at each time step of the text image.

[0060] In a specific implementation manner, the present application can select a CRNN (Convolutional Recurrent Neural Network) network to perform OCR recognition on the text image. This CRNN network is mainly used to recognize an indefinite-length text sequence end-to-end, without the need to cut individual characters first, but instead transforms text recognition into a sequence learning problem with temporal dependence, so that accurate prediction can be achieved by combining the context information in the text image.

[0061] S120: Traverse the probability distribution of all character categories output at each time step, and obtain the character corresponding to the extremely high probability value in the probability distribution at each time step.

[0062] In this step, after performing OCR recognition on the text image to be recognized through S110 and obtaining the probability distribution of all character categories output at each time step of the text image, the present application can also traverse the probability distribution of all character categories output at each time step and obtain the character corresponding to the extremely high probability value in the probability distribution at each time step.

[0063] It can be understood that after the present application obtains the probability distribution of all character categories output at each time step, in order to further improve the accuracy of OCR recognition, the present application can traverse the probability distribution of all character categories output at each time step and screen out the character corresponding to the extremely high probability value in the probability distribution at each time step. Determining the final recognition result based on the character corresponding to the extremely high probability value can effectively improve the accuracy of OCR recognition.

[0064] S130: For each time step, determine whether the character corresponding to the extremely high probability value in the probability distribution of this time step is a character in a pre-configured error-prone character grouping list; if so, execute S140; if not, execute S150.

[0065] In this step, after obtaining the character corresponding to the extremely high probability value in the probability distribution of each time step through S120, this application can determine, for each time step, whether the character corresponding to the extremely high probability value in the probability distribution of this time step is a character in a pre-configured error-prone character grouping list, so that the extremely high probability value in one or more time steps can be adjusted according to the determination result, avoiding the situation where the overall score is too low due to the ambiguity of easily confused characters and the correct recognition result is filtered out.

[0066] Specifically, when determining whether the character corresponding to the extremely high probability value in the probability distribution of this time step is a character in the error-prone character grouping list, the character corresponding to the extremely high probability value in the probability distribution of this time step can be compared with the characters in the error-prone character grouping list respectively. If the character corresponding to the extremely high probability value is the same as the character in the error-prone character grouping list, it indicates that the character corresponding to the extremely high probability value is a character in the error-prone character grouping list; otherwise, it is not.

[0067] Among them, the error-prone character grouping list of this application refers to the grouping list corresponding to the characters that are prone to errors when recognizing text images in the scenario of numbers + letters. This grouping list can include (1, I, i, 1), (0, o, O), (Z, z, 2), (C, c), (W, w), (V, v), (U, u). When the character corresponding to the extremely high probability value appears in this grouping list, it means that the character corresponding to the extremely high probability value is a character in the error-prone character grouping list.

[0068] S140: Based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the probability distribution of this time step in the error-prone character grouping list, update the extremely high probability value, and use the character after updating the extremely high probability value as the output character of this time step.

[0069] In this step, after determining through S130 whether the character corresponding to the extremely high probability value in the probability distribution of each time step is a character in a pre-configured error-prone character grouping list, if the character corresponding to the extremely high probability value in the probability distribution of one or more time steps is a character in the error-prone character grouping list, the extremely high probability value can be updated according to the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the probability distribution of this time step in the error-prone character grouping list, and the character after updating the extremely high probability value is used as the output character of this time step.

[0070] It can be understood that when the present application performs OCR recognition on a text image, the obtained is the probability distribution of all character categories output by the text image at each time step. Therefore, this probability distribution not only contains the character corresponding to the extremely high probability value, but also contains the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list and the probability value corresponding to this error-prone character.

[0071] When the character corresponding to the extremely high probability value in the probability distribution of one or more time steps is a character in the error-prone character grouping list, the extremely high probability value can be updated according to the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution of this time step. For example, add the extremely high probability value and the probability value of the error-prone character, or perform weighted summation. If both the extremely high probability value and the probability value of the error-prone character are greater than 0, the extremely high probability value can also be updated through a power function. Of course, the extremely high probability value can also be updated in other ways, as long as the updated extremely high probability value is greater than the extremely high probability value before the update, and no restrictions are imposed here.

[0072] S150: Directly use the character corresponding to the extremely high probability value in the probability distribution of this time step as the output character of this time step.

[0073] In this step, after determining whether the character corresponding to the extremely high probability value in the probability distribution of each time step is a character in the pre-configured error-prone character grouping list through S130, if there is no time step in which the character corresponding to the extremely high probability value in the probability distribution is a character in the error-prone character grouping list, the character corresponding to the extremely high probability value in the probability distribution of each time step can be directly used as the output character of this time step.

[0074] S160: Concatenate the output characters of all time steps to obtain a character sequence, and decode the character sequence to obtain the final recognition result.

[0075] In this step, after obtaining the output character of each time step through S140 and S150, the present application can concatenate the output characters of all time steps and obtain a character list. Then, the character list can be decoded to obtain the final recognition result.

[0076] It can be understood that in order to improve the OCR recognition accuracy, the present application can concatenate the output characters of all time steps to obtain a character sequence. In this way, this character sequence is the sequence formed by the combination of the characters with the extremely high probability value at each time step. By decoding this character sequence, a recognition result with relatively high accuracy can be obtained.

[0077] Among them, the decoding process refers to the process of translating a character sequence into a final recognition result, which may include removing redundant information, etc. Generally, when performing temporal classification, a lot of redundant information will inevitably appear. For example, a letter is recognized twice consecutively. At this time, a redundancy removal mechanism needs to be used to remove the redundant information in the character sequence to obtain the final recognition result.

[0078] In the above embodiment, when performing OCR recognition on the text image to be recognized, all probability distributions of character categories output at each time step of the text image can be obtained first. Then, the present application can determine the character corresponding to the extremely high probability value in the probability distribution of each time step. Then, for each time step, it is determined whether the character corresponding to the extremely high probability value in the probability distribution of this time step is a character in a pre-configured error-prone character grouping list. If not, the character corresponding to the extremely high probability value in the probability distribution of this time step is directly used as the output character of this time step, which can improve the accuracy of the recognition result to a certain extent; if so, the extremely high probability value is updated based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution of this time step, and the character after updating the extremely high probability value is used as the output character of this time step. Since the extremely high probability value of the output character of this time step is adjusted compared to the original extremely high probability value, when the output characters of all adjusted time steps are concatenated and then decoded, the situation where the overall score is too low due to the ambiguity of easily confused characters and the correct recognition result is filtered can be avoided, thereby effectively improving the accuracy of the recognition result.

[0079] In one embodiment, the method may further include:

[0080] S170: According to the recognition scenario corresponding to the text image, correct the recognition result to obtain a corrected result.

[0081] In this embodiment, after obtaining the recognition result, the present application can also correct the recognition result according to the recognition scenario corresponding to the text image. In this way, the obtained corrected result is more accurate than the original recognition result, thereby further improving the accuracy of the OCR recognition result.

[0082] In one embodiment, in S170, according to the recognition scenario corresponding to the text image, correcting the recognition result to obtain a corrected result may include:

[0083] S171: Determine the recognition scenario corresponding to the text image and the error-prone character list corresponding to the recognition scenario.

[0084] S172: Compare each output character in the recognition result with the frequently occurring characters in the frequently occurring character list respectively to determine whether there is an output character not included in the frequently occurring character list in the recognition result; if there is, execute S173; if not, directly use the recognition result as the corrected result.

[0085] S173: Replace the output characters not included in the frequently occurring character list with the error-prone characters paired with these output characters in the error-prone character group list.

[0086] S174: Use the recognition result after replacing the output characters as the corrected result.

[0087] In this embodiment, when correcting the recognition result, the recognition scene corresponding to the text image and the frequently occurring character list corresponding to this recognition scene can be determined first. Then, each output character in this recognition result is compared with the frequently occurring characters in the frequently occurring character list respectively to determine whether there is an output character not included in the frequently occurring character list in the recognition result. If there is, it indicates that the output characters in this recognition result are very likely to be misrecognized characters. At this time, these output characters can be replaced with the error-prone characters paired with them in the error-prone character group list, and the recognition result after replacing the output characters is used as the corrected result; if not, it indicates that the output characters in this recognition result basically have no recognition errors. At this time, this recognition result can be directly used as the corrected result.

[0088] For example, when the text image is the text image obtained in the license plate scene, since the character O or o will not appear in the license plate scene, if O or o appears in the recognition result at this time, it can be replaced with 0. Further, as Figure 2 、 3 、shown in 4, Figure 2 is the schematic diagram of the text image provided by the embodiment of the present application, Figure 3 is the display diagram of the recognition result of the text image provided by the embodiment of the present application, Figure 4 is the display diagram of the corrected result of the text image provided by the embodiment of the present application; as can be seen from Figure 2 、 Figure 2 , the format of the second-line text of the text image in Figure 3 is number (letter letter) number number number number number number number. The recognition result obtained after the first recognition is as shown in Figure 4 . According to the correction method in the present application, the recognition result is corrected. After replacing the characters in the fourth to tenth positions with 0, the corrected result is as shown in

[0089] . After such processing, the correct rate of the recognition result is greatly improved, and the recognition result will not be filtered out due to the low score caused by easily confused characters.It should be noted that the present application can pre-construct a list of characters prone to appear based on the characters that are relatively common in text images in different scenarios. In this way, after obtaining the recognition result, the current recognition result can be judged according to the list of characters prone to appear, and adjusted in case of inaccuracy, which is beneficial to further improve the accuracy of text recognition.

[0090] In one embodiment, in S110, OCR recognition is performed on the text image to be recognized, and the probability distribution of all character categories output at each time step can include:

[0091] S111: Input the text image to be recognized into a pre-configured CRNN network for OCR recognition, and perform feature extraction on the text image through the convolutional layer of the CRNN network to obtain multiple feature maps.

[0092] S112: After converting each feature map into a feature vector, input each feature vector into the recurrent layer of the CRNN network in sequence, and predict the character corresponding to each feature vector through the recurrent layer to obtain the probability distribution of all character categories output by the recurrent layer at each time step.

[0093] In this embodiment, when performing OCR recognition on a text image, a CRNN network can be used. The CRNN network can include a convolutional layer, a recurrent layer, etc. Through the convolutional layer of the CRNN network, feature extraction can be performed on the text image to obtain multiple feature maps, and the multiple feature maps can be converted into feature vectors. The feature vectors can be used to perform character prediction through the recurrent layer of the CRNN network, so as to obtain the probability distribution of all character categories output by the text image at each time step.

[0094] For example, the height of the text image input to the CRNN network is 32. After convolution by the convolutional layer, the height of the text image becomes 1, and the width can be 160. Therefore, the size of the text image input to the convolutional layer can be (channel, height, width) = (1, 32, 160), and the output size can be (512, 1, 40). That is, 512 feature maps can be obtained after feature extraction by the convolutional layer, and the height of each feature map is 1 and the width is 40. Since the input to the recurrent layer is a feature sequence, after obtaining the feature maps, a sequence of feature vectors can be extracted from the feature maps. Each feature vector is generated column by column from left to right on the feature map, and each column contains 512-dimensional features. This means that the i-th feature vector is the concatenation of the pixels in the i-th column of all the feature maps, and these feature vectors form a sequence. Each column of the feature map (i.e., a feature vector) corresponds to a rectangular region (referred to as the receptive field) of the text image, and these rectangular regions have the same order as the corresponding columns from left to right on the feature map. Therefore, each vector in the feature vector sequence is associated with a receptive field.

[0095] Next, each feature vector in the feature vector sequence can be used as the input of the recurrent layer at a time step. It can be understood that through the above steps, 40 feature vectors can be obtained in this application, and the length of each feature vector is 512. In the recurrent layer, only one feature vector is input at a time step for classification. Therefore, the recurrent layer can have a total of 40 time steps. At each time step, the recurrent layer can predict which character the currently input feature vector is and output the probability distribution of all character classes as a vector with a length equal to the number of character classes. After 40 vectors with a length equal to the number of character classes are output, a posterior probability matrix can be formed by the 40 vectors with a length equal to the number of character classes, and this posterior probability matrix can be transcribed into the recognition result of the text image through the method for determining the text recognition result of this application.

[0096] In one embodiment, in S140, based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution at this time step, updating the extremely high probability value may include:

[0097] S141: Determine the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list, and the probability value of the error-prone character in the probability distribution at this time step.

[0098] S142: Add the probability value of the error-prone character in the probability distribution at this time step to the extremely high probability value, and update the extremely high probability value according to the addition result.

[0099] In this embodiment, when the character corresponding to the extremely high probability value in the probability distribution of one or more time steps is a character in the error-prone character grouping list, the extremely high probability value can be updated according to the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution of this time step.

[0100] Specifically, when the present application updates the extremely high probability value, it can first determine the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list, and the probability value of this error-prone character in the probability distribution of this time step. Then, the present application can add the probability value of the error-prone character in the probability distribution of this time step to the extremely high probability value, and use the added probability value as the new extremely high probability value to update the original extremely high probability value. Compared with the original extremely high probability value, the updated extremely high probability value has a relatively large value, because it can avoid the situation that the overall score is too low due to the ambiguity of similar characters and the correct recognition result is filtered out.

[0101] In one embodiment, after decoding the character sequence in S160 to obtain the final recognition result, it may include:

[0102] S161: Traverse the character sequence to determine whether the character sequence contains a placeholder; if not, execute S162; if so, execute S163.

[0103] S162: Merge the repeated output characters in consecutive time steps, and use the merged character sequence as the final recognition result.

[0104] S163: Merge the repeated output characters in consecutive time steps in the character sequence according to the placeholder, and remove the placeholder in the merged character sequence to obtain the final recognition result.

[0105] In this embodiment, after obtaining the character sequence formed by splicing the extremely high probability values in each time step, the present application can first traverse the character sequence to determine whether the character sequence contains a placeholder. If so, merge the repeated output characters in consecutive time steps in the character sequence according to the placeholder, and remove the placeholder in the merged character sequence to obtain the final recognition result; when the character sequence does not contain a placeholder, the repeated output characters in consecutive time steps can be directly merged, and the merged character sequence is used as the final recognition result.

[0106] It is understandable that during the process of decoding the character sequence and obtaining the final recognition result, since the same text may be repeatedly recognized when the character sequence is subjected to temporal classification, redundant information may thus appear in the character sequence. For example, when the text "ab" needs to be recognized, if the text is divided into 5 time steps for recognition, it may be that at times t0, t1, and t2, it is mapped to "a", and at times t3 and t4, it is mapped to "b". After connecting these character sequences, "aaabb" is obtained. In this character sequence, the letter "a" is repeated three times and the letter "b" is repeated twice. Therefore, when decoding subsequently, the repeated characters need to be merged in order to obtain the final recognition result "ab".

[0107] Furthermore, since there are texts with repeated characters in some text images themselves, such as words like "book" and "hello", if consecutive characters are merged in the above manner, "bok" and "helo" will be obtained, which will lead to inaccurate recognition results. Therefore, the "-" symbol can be used to represent a placeholder. When outputting the character sequence, a "-" can be inserted between the repeated characters in the text label. For example, if the output sequence is "bbooo-ookk", it will finally be mapped to "book", that is, if there are placeholders separating them, consecutive identical output characters do not need to be merged. Only the consecutive identical output characters without placeholders separating them need to be merged, and then the placeholders in the merged character sequence are deleted to obtain the final recognition result.

[0108] In one embodiment, the method may further include:

[0109] S181: Multiply the most probable values corresponding to each output character in the recognition result to obtain a product result.

[0110] S182: Use the product result as the confidence score of the recognition result.

[0111] In this embodiment, after obtaining the final recognition result, in order to verify the reliability of the recognition result, the present application may also perform a confidence score on the recognition result, and the user can judge the strength of the reliability of the current recognition result through the confidence score.

[0112] Specifically, when scoring the recognition result, since each output character in the recognition result is the character corresponding to the most probable value, therefore, the present application can directly multiply the most probable values corresponding to each output character and use the product result as the confidence score of the recognition result.

[0113] The computational complexity of the above scoring method of the present application is the same as that of the greedy algorithm, but it takes into account the correlation between different times, and thus is more robust to error-prone characters.

[0114] The following describes the apparatus for determining the text recognition result provided by the embodiments of the present application. The apparatus for determining the text recognition result described below can be correspondingly referred to the method for determining the text recognition result described above.

[0115] In one embodiment, as Figure 5 shown, Figure 5 is a schematic structural diagram of an apparatus for determining a text recognition result provided by an embodiment of the present application; The present application also provides an apparatus for determining a text recognition result, which may include a text recognition module 210, a character acquisition module 220, an error-prone character judgment module 230, a first character determination module 240, a second character determination module 250, and a recognition result determination module 260, specifically including the following:

[0116] The text recognition module 210 is configured to perform OCR recognition on the text image to be recognized, and obtain the probability distribution of all character categories output by the text image at each time step.

[0117] The character acquisition module 220 is configured to traverse the probability distribution of all character categories output at each time step, and obtain the character corresponding to the extremely high probability value in the probability distribution of each time step.

[0118] The error-prone character judgment module 230 is configured to determine, for each time step, whether the character corresponding to the extremely high probability value in the probability distribution of this time step is a character in a pre-configured error-prone character grouping list.

[0119] The first character determination module 240 is configured to, if so, update the extremely high probability value based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution of this time step, and use the character after updating the extremely high probability value as the output character of this time step; wherein, the updated extremely high probability value is greater than the extremely high probability value before updating.

[0120] The second character determination module 250 is configured to, if not, directly use the character corresponding to the extremely high probability value in the probability distribution of this time step as the output character of this time step.

[0121] The recognition result determination module 260 is configured to splice the output characters of all time steps to obtain a character sequence, and decode the character sequence to obtain the final recognition result.

[0122] In the above embodiments, when performing OCR recognition on the text image to be recognized, all probability distributions of character categories output at each time step of the text image can be obtained first. Then, the present application can determine the character corresponding to the extremely high probability value in the probability distribution of each time step. Then, for each time step, it is determined whether the character corresponding to the extremely high probability value in the probability distribution of this time step is a character in the pre-configured error-prone character grouping list. If not, the character corresponding to the extremely high probability value in the probability distribution of this time step is directly used as the output character of this time step, which can improve the accuracy of the recognition result to a certain extent; if so, the extremely high probability value is updated based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution of this time step, and the character after updating the extremely high probability value is used as the output character of this time step. Since the extremely high probability value of the output character of this time step is adjusted compared with the original extremely high probability value, when the output characters of all adjusted time steps are concatenated and then decoded, the situation where the overall score is too low due to the ambiguity of easily confused characters and the correct recognition result is filtered can be avoided, thereby effectively improving the accuracy of the recognition result.

[0123] In one embodiment, the device may further include:

[0124] A result correction module, configured to correct the recognition result according to the recognition scenario corresponding to the text image to obtain a corrected result.

[0125] In one embodiment, the result correction module may include:

[0126] A scenario and list determination module, configured to determine the recognition scenario corresponding to the text image and the list of characters prone to appear corresponding to the recognition scenario.

[0127] A character comparison module, configured to compare each output character in the recognition result with the characters prone to appear in the list of characters prone to appear, and determine whether there is an output character in the recognition result that is not included in the list of characters prone to appear.

[0128] A character replacement module, configured to, if any, replace the output character not included in the list of characters prone to appear with the error-prone character paired with the output character in the error-prone character grouping list.

[0129] A corrected result determination module, configured to use the recognition result after replacing the output characters as the corrected result.

[0130] In one embodiment, the text recognition module 210 may include:

[0131] A feature extraction module for inputting a text image to be recognized into a pre-configured CRNN network for OCR recognition, and obtaining a plurality of feature maps after feature extraction of the text image through a convolutional layer of the CRNN network.

[0132] A character prediction module for converting each feature map into a feature vector and then sequentially inputting each feature vector into a recurrent layer of the CRNN network, and predicting a character corresponding to each feature vector through the recurrent layer to obtain a probability distribution of all character classes output by the recurrent layer at each time step.

[0133] In one embodiment, the first character determination module 240 may include:

[0134] A character and probability value determination module for determining an error-prone character paired with a character in the error-prone character grouping list corresponding to the extremely high probability value, and a probability value of the error-prone character in the probability distribution at this time step.

[0135] A probability addition module for adding the probability value of the error-prone character in the probability distribution at this time step to the extremely high probability value, and updating the extremely high probability value according to the addition result.

[0136] In one embodiment, the recognition result determination module 260 may include:

[0137] A traversal module for traversing the character sequence to determine whether the character sequence contains a placeholder.

[0138] A first merging module for, if not, merging repeated output characters in consecutive time steps and using the merged character sequence as the final recognition result.

[0139] A second merging module for, if so, merging repeated output characters in consecutive time steps in the character sequence according to the placeholder, and removing the placeholder in the merged character sequence to obtain the final recognition result.

[0140] In one embodiment, the apparatus may further include:

[0141] A product module for multiplying the extremely high probability values corresponding to each output character in the recognition result to obtain a product result.

[0142] A confidence score module for using the product result as the confidence score of the recognition result.

[0143] In one embodiment, the present application further provides a storage medium storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the method for determining the text recognition result as described in any one of the above embodiments.

[0144] In one embodiment, the present application further provides a computer device, including: one or more processors, and a memory.

[0145] The memory stores computer-readable instructions that, when executed by the one or more processors, perform the steps of the method for determining the text recognition result as described in any one of the above embodiments.

[0146] Schematically, as Figure 6 shown, Figure 6 is a schematic internal structure diagram of a computer device provided by an embodiment of the present application. The computer device 300 can be provided as a server. Referring to Figure 6 , the computer device 300 includes a processing component 302, which further includes one or more processors, and memory resources represented by a memory 301 for storing instructions executable by the processing component 302, such as application programs. The application programs stored in the memory 301 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 302 is configured to execute instructions to perform the method for determining the text recognition result of any of the above embodiments.

[0147] The computer device 300 may further include a power component 303 configured to perform power management of the computer device 300, a wired or wireless network interface 304 configured to connect the computer device 300 to a network, and an input / output (I / O) interface 305. The computer device 300 may operate based on an operating system stored in the memory 301, such as Windows Server TM, Mac OS XTM, Unix TM, Linux TM, Free BSDTM, or the like.

[0148] Those skilled in the art can understand that Figure 6 the structure shown in

[0149] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0150] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0151] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for determining a text recognition result, characterized in that, The method includes: Performing OCR recognition on the text image to be recognized to obtain the probability distribution of all character categories output by the text image at each time step; Traversing the probability distribution of all character categories output at each time step, and obtaining the character corresponding to the extremely high probability value in the probability distribution of each time step; For each time step, determining whether the character corresponding to the extremely high probability value in the probability distribution of this time step is a character in a pre-configured error-prone character grouping list; wherein, the error-prone character grouping list refers to a grouping list corresponding to characters that are prone to errors when recognizing text images in the scenario of numbers + letters, and each group of characters includes combinations of easily confused letters, or combinations of easily confused letters and numbers; If so, based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution of this time step, updating the extremely high probability value, and using the character after updating the extremely high probability value as the output character of this time step; wherein, the updated extremely high probability value is greater than the extremely high probability value before updating, and the pairing refers to the letters or numbers that are easily confused with the character corresponding to the extremely high probability value in the same group of characters; If not, directly using the character corresponding to the extremely high probability value in the probability distribution of this time step as the output character of this time step; Concatenating the output characters of all time steps to obtain a character sequence, and decoding the character sequence to obtain the final recognition result.

2. The method for determining the text recognition result according to claim 1, wherein The method further includes: According to the recognition scenario corresponding to the text image, correcting the recognition result to obtain a corrected result.

3. The method for determining the text recognition result according to claim 2, wherein The correcting the recognition result according to the recognition scenario corresponding to the text image to obtain a corrected result includes: Determining the recognition scenario corresponding to the text image and the list of characters that are likely to appear corresponding to the recognition scenario; wherein, the characters that are likely to appear are pre-constructed according to the common characters in text images in different scenarios; Comparing each output character in the recognition result with the characters that are likely to appear in the list of characters that are likely to appear, and determining whether there are output characters in the recognition result that are not included in the list of characters that are likely to appear; If there are, replacing the output characters that are not included in the list of characters that are likely to appear with the error-prone characters paired with these output characters in the error-prone character grouping list; Using the recognition result after replacing the output characters as the corrected result.

4. The method for determining the text recognition result according to claim 1, characterized in that, The performing OCR recognition on the text image to be recognized to obtain the probability distribution of all character categories output by the text image at each time step includes: Inputting the text image to be recognized into a pre-configured CRNN network for OCR recognition, and obtaining multiple feature maps by extracting features of the text image through the convolutional layer of the CRNN network; After converting each feature map into a feature vector, sequentially inputting each feature vector into the recurrent layer of the CRNN network, and predicting the character corresponding to each feature vector through the recurrent layer to obtain the probability distribution of all character categories output by the recurrent layer at each time step.

5. The method for determining the text recognition result according to claim 1, wherein Updating the extremely high probability value based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list at this time step includes: Determining the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list, and the probability value of the error-prone character in the probability distribution at this time step; Adding the probability value of the error-prone character in the probability distribution at this time step to the extremely high probability value, and updating the extremely high probability value according to the addition result.

6. The method for determining the text recognition result according to claim 1, characterized in that The obtaining of the final recognition result after decoding the character sequence includes: Traversing the character sequence to determine whether the character sequence contains a placeholder; If not, merging the repeated output characters in consecutive time steps, and using the merged character sequence as the final recognition result; If so, merging the repeated output characters in consecutive time steps in the character sequence according to the placeholder, and removing the placeholder in the merged character sequence to obtain the final recognition result.

7. The method for determining the text recognition result according to any one of claims 1-6, characterized in that, The method further includes: Multiplying the extremely high probability values corresponding to each output character in the recognition result to obtain a product result; Using the product result as the confidence score of the recognition result.

8. An apparatus for determining a text recognition result, characterized in that including: A text recognition module for performing OCR recognition on the text image to be recognized to obtain the probability distribution of all character categories output by the text image at each time step; A character acquisition module for traversing the probability distribution of all character categories output at each time step to obtain the character corresponding to the extremely high probability value in the probability distribution at each time step; An error-prone character judgment module for determining, for each time step, whether the character corresponding to the extremely high probability value in the probability distribution at this time step is a character in a pre-configured error-prone character grouping list; wherein, the error-prone character grouping list refers to a grouping list corresponding to characters that are prone to errors when recognizing text images in the digital + letter scenario, and each group of characters includes a combination of easily confused letters, or a combination of easily confused letters and numbers; A first character determination module for, if so, updating the extremely high probability value based on the probability value of the error-prone character paired with the character corresponding to the extremely high probability value in the error-prone character grouping list in the probability distribution at this time step, and using the character after updating the extremely high probability value as the output character at this time step; wherein, the updated extremely high probability value is greater than the extremely high probability value before updating, and the pairing refers to the letters or numbers that are easily confused with the character corresponding to the extremely high probability value in the same group of characters; A second character determination module for, if not, directly using the character corresponding to the extremely high probability value in the probability distribution at this time step as the output character at this time step; A recognition result determination module for splicing the output characters of all time steps to obtain a character sequence, and decoding the character sequence to obtain the final recognition result.

9. A storage medium, characterized in that: The storage medium stores computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to execute the steps of the method for determining the text recognition result according to any one of claims 1 to 7.

10. A computer device, characterized in that, Comprising: One or more processors, and a memory; The memory stores computer-readable instructions, which, when executed by the one or more processors, execute the steps of the method for determining the text recognition result according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • OCR recognition result confidence determination method and device, and electronic equipment

    CN110765870A

  • Text recognition method and device, storage medium and computer equipment

    CN112508102A