OCR Correction Method, Device, Computer Equipment and Storage Medium Based on CTC Loss

Through the OCR recognition model based on CTC loss, the CTC character submatrix with character confidence threshold is constructed and filtered, which improves the OCR recognition accuracy in complex environments, and solves the problem of the decrease in OCR recognition accuracy in complex environments.

CN114529902BActive Publication Date: 2025-07-04HUNAN XINGHAN DIGITAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210150555.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-18
Publication Date
2025-07-04
Estimated Expiration
2042-02-18

AI Technical Summary

Technical Problem

The accuracy of the existing OCR technology in complex environments (such as document, document photography, etc.), has decreased, especially due to factors such as blurred font printing, local highlighting of photography, and physical wear and wrinkles.

Method used

Using an OCR recognition model based on CTC loss, a CTC decoded character sequence matrix is ​​constructed by extracting the text to be identified, a CTC decoded character sequence matrix is ​​filtered, a CTC character submatrix that satisfies the character confidence threshold, a string sequence is constructed based on the confidence threshold relationship, and finally the optimal string sequence is output as the OCR correction result.

Benefits of technology

The recognition accuracy of OCR in complex scenarios is improved, and the confidence of the possible character values ​​is combined with the threshold value is obtained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529902B_ABST
    Figure CN114529902B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of artificial intelligence image and text recognition, and provides an OCR correction method, device, computer device and storage medium based on CTC loss. The method includes: extracting the text to be recognized, using an OCR recognition model based on CTC loss to perform character recognition on the text to be recognized, and obtaining a CTC decoded character sequence matrix; screening characters that meet the character confidence threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix, and obtaining a set of CTC character sub-matrices; traversing the set of CTC character sub-matrices, screening the possible values of the characters based on the magnitude relationship between the confidence of each possible value of each character in each CTC character sub-matrix and the character confidence threshold, constructing a string sequence, and obtaining a set of string sequences; screening the optimal string sequence from the set of string sequences, and outputting the optimal string sequence as the OCR correction result. Using this method can improve the accuracy of OCR recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence image and text recognition, and particularly relates to an OCR correction method, device, computer device, and storage medium based on CTC loss. Background Art

[0002] OCR (Optical Character Recognition) is one of the branches in the field of computer vision research, and its essence is image recognition. Specifically, it is a process of using electronic devices such as scanners or digital cameras to check the characters printed on paper, determining their shapes by detecting dark and bright patterns, and then translating the shapes into computer text using character recognition methods, and converting the text in the image into a text format through recognition software for word processing software to edit and process.

[0003] At present, the OCR recognition accuracy in simple environments such as PDF (Portable Document Format) images and web screenshots is already relatively high. However, in actual social life, there are more and more application requirements for OCR recognition in complex environments. However, in complex environments such as taking pictures of certificates and bills, due to factors such as blurred font printing, local highlights in the photo, and wear and wrinkles of the entity, the accuracy of OCR recognition still decreases. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide an OCR correction method, device, computer device, and storage medium based on CTC loss that can improve the OCR recognition accuracy.

[0005] The present invention provides an OCR correction method based on CTC loss, including:

[0006] Extracting the text to be recognized, and using an OCR recognition model based on CTC loss to perform character recognition on the text to be recognized to obtain a CTC decoded character sequence matrix;

[0007] Screening the characters that meet the character confidence threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix, and obtaining a set of CTC character sub-matrices;

[0008] Traversing the set of CTC character sub-matrices, screening the possible values of the characters based on the magnitude relationship between the confidence of each possible value of each character in each CTC character sub-matrix and the character confidence threshold, and constructing a string sequence to obtain a set of string sequences;

[0009] Screening the optimal string sequence from the set of string sequences, and outputting the optimal string sequence as the OCR correction result; specifically including:

[0010] Traverse each string sequence in the set of string sequences, and determine the string sequences that meet the adaptive requirements as candidate string sequences;

[0011] Take the weighted average confidence of the candidate string sequences as the first score, and take the edit distance between the candidate string sequences and the elements in the first row of the CTC decoded character sequence matrix as the second score;

[0012] Weight the first score and the second score to obtain the final score of the candidate string sequences;

[0013] Determine the optimal string sequence from the candidate string sequences according to the final score;

[0014] Output the optimal string sequence in the character format of the text to be recognized to obtain the OCR correction result.

[0015] In one embodiment, the extracting the text to be recognized and using an OCR recognition model based on the CTC loss to perform character recognition on the text to be recognized to obtain a CTC decoded character sequence matrix includes:

[0016] Detect and intercept the text box in the image to be detected to obtain the text to be recognized;

[0017] Use an OCR character recognition model based on the CTC loss to perform character recognition on the text to be recognized, and obtain the CTC decoded character sequence and the possible values of each character in the CTC decoded character sequence;

[0018] Construct a CTC decoded character sequence matrix according to the length of the CTC decoded character sequence and the possible values of each character.

[0019] In one embodiment, the screening characters that meet the character credibility threshold from the CTC decoded character sequence matrix to construct a CTC character submatrix to obtain a set of CTC character submatrices includes:

[0020] Compare the confidence of the first possible value of each character in the CTC decoded character sequence matrix with the character credibility threshold respectively to determine whether the character is correct;

[0021] When the character is correct, represent the character with a credible mode character, and when it is incorrect, represent the character with an incredible mode character, and construct a mode character sequence corresponding to the CTC decoded character sequence matrix;

[0022] According to the pattern character sequence, enumerate all pattern character sequences containing a trusted subsequence to obtain a set of pattern character sequences; the trusted subsequence is composed of trusted pattern characters;

[0023] Traverse the set of pattern character sequences, and construct a CTC character sub-matrix by screening characters from the CTC decoded character sequence matrix according to the pattern character sequence to obtain a set of CTC character sub-matrices.

[0024] In one embodiment, the step of enumerating all pattern character sequences containing a trusted subsequence according to the pattern character sequence to obtain a set of pattern character sequences includes:

[0025] Take the pattern character sequence as a binary character sequence, use the decimal value obtained by converting the binary as the start node, and determine the value according to the length of the character sequence as the end node;

[0026] Traverse from the start node to the end node, perform a bitwise OR operation on each integer with the decimal value and convert it into a binary sequence to obtain a pattern character sequence containing a trusted subsequence;

[0027] Combine all the pattern character sequences containing a trusted subsequence to obtain a set of pattern character sequences.

[0028] In one embodiment, the step of enumerating all pattern character sequences containing a trusted subsequence according to the pattern character sequence to obtain a set of pattern character sequences includes:

[0029] Construct an empty set of pattern character sequences, and traverse each pattern character in the pattern character sequence;

[0030] If the currently traversed pattern character is the first one and is a trusted pattern character, add a pattern character sequence with the first character being a trusted pattern character to the set of pattern character sequences;

[0031] If the currently traversed pattern character is the first one and is an untrusted pattern character, add a pattern character sequence with the first character being a trusted pattern character and a pattern character sequence with the first character being an untrusted pattern character to the set of pattern character sequences; if the currently traversed pattern character is not the first one and is a trusted pattern character, append a trusted pattern character to the end of each existing pattern character sequence in the set of pattern characters;

[0032] If the currently traversed pattern character is not the first character and is an untrusted pattern character, then append a trusted pattern character to the end of each existing pattern character sequence in the set of pattern characters, and copy the existing pattern character sequence. Append an untrusted pattern character to the end of the copied pattern character sequence, doubling the number of pattern character sequences in the set of pattern characters.

[0033] In one embodiment, traversing the set of pattern character sequences, and screening characters from the CTC decoded character sequence matrix according to the pattern character sequences to construct CTC character sub-matrices, obtaining a set of CTC character sub-matrices, including:

[0034] Construct an empty CTC character sub-matrix;

[0035] Traverse the pattern characters in the pattern character sequence. If the k-th pattern character currently traversed is a trusted pattern character, then add the k-th column in the CTC decoded character sequence matrix to the empty CTC character sub-matrix to obtain a CTC character sub-matrix corresponding to the CTC decoded character sequence matrix;

[0036] Combine the CTC character sub-matrices of each pattern character sequence into a set of CTC character sub-matrices.

[0037] In one embodiment, traversing the set of CTC character sub-matrices, screening character possible values based on the magnitude relationship between the confidence levels of the possible values of each character in each CTC character sub-matrix and the character credibility threshold, and constructing a string sequence to obtain a set of string sequences, including:

[0038] For each CTC character sub-matrix in the set of CTC character sub-matrices, respectively construct an empty string sequence with a sequence length equal to the number of characters in the CTC character sub-matrix;

[0039] Compare the confidence levels of the first possible values of each character in the CTC character sub-matrix with the character credibility threshold respectively;

[0040] If the confidence level corresponding to the first possible value is not less than the character credibility threshold, add the first possible value to the empty string sequence, otherwise add all possible values of the character in the CTC character sub-matrix to the empty string sequence to obtain an assigned string sequence;

[0041] Perform permutation and combination on the assigned string sequences, enumerate all string sequences to obtain a set of string sequences.

[0042] An OCR correction device based on CTC loss, including:

[0043] A character recognition module, which is used to extract the text to be recognized, and perform character recognition on the text to be recognized by using an OCR recognition model based on the CTC loss to obtain a CTC decoded character sequence matrix;

[0044] A sub-matrix construction module, which is used to screen out characters that meet the character credibility threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix, and obtain a set of CTC character sub-matrices;

[0045] A string sequence construction module, which is used to traverse the set of CTC character sub-matrices, screen the possible values of characters based on the size relationship between the confidence of each possible value of each character in each CTC character sub-matrix and the character credibility threshold, construct a string sequence, and obtain a set of string sequences;

[0046] An output module, which is used to screen out the optimal string sequence from the set of string sequences and output the optimal string sequence as the OCR correction result; specifically, it is used to traverse each string sequence in the set of string sequences, determine the string sequence that meets the adaptive requirements as the candidate string sequence; use the weighted average of the confidence of the candidate string sequence as the first score, and use the edit distance between the candidate string sequence and the first row elements in the CTC decoded character sequence matrix as the second score; weight the first score and the second score to obtain the final score of the candidate string sequence; determine the optimal string sequence from the candidate string sequences according to the final score; output the optimal string sequence in the character format of the text to be recognized to obtain the OCR correction result.

[0047] The present invention also provides a computer device, which includes a processor and a memory. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the OCR correction method based on CTC loss described in any one of the above.

[0048] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the OCR correction method based on CTC loss described in any one of the above.

[0049] The above OCR correction method, device, computer device, and storage medium based on the CTC loss extract the text to be recognized, use the OCR recognition model based on the CTC loss to perform character recognition on the text to be recognized, and after obtaining the CTC decoded character sequence matrix, screen the characters that meet the character confidence threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix, obtain a set of CTC character sub-matrices, and then traverse the set of CTC character sub-matrices, and screen the possible values of the characters based on the size relationship between the confidence of each possible value of each character in each CTC character sub-matrix and the character confidence threshold, and construct a string sequence to obtain a set of string sequences; finally, screen the optimal string sequence from the set of string sequences, and output the optimal string sequence as the OCR correction result. This method gradually screens to obtain the optimal solution by combining the confidence of each possible value of each character in the CTC decoded character sequence with the threshold, and even for OCR recognition in complex scenarios, high-accuracy recognition results can be obtained, improving the accuracy of OCR recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 FIG. is an application environment diagram of an OCR correction method based on CTC loss in an embodiment.

[0051] Figure 2 FIG. is a flowchart of an OCR correction method based on CTC loss in an embodiment.

[0052] Figure 3 FIG. is a schematic diagram of an image to be detected in an embodiment.

[0053] Figure 4 FIG. is a schematic diagram of the text to be recognized in an embodiment.

[0054] Figure 5 FIG. is a schematic diagram of a CTC decoded character sequence in an embodiment.

[0055] Figure 6 FIG. is a schematic diagram of a CTC character subsequence in an embodiment.

[0056] Figure 7 FIG. is a structural block diagram of an OCR correction device based on CTC loss in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0058] The OCR correction method based on CTC loss provided by the present application can be applied to, for example Figure 1In the application environment shown, the application environment involves the terminal 102 and the server 104. Among them, the terminal 102 communicates with the server 104 through the network. The terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0059] When the terminal 102 receives an OCR recognition instruction, the above OCR correction method based on the CTC loss can be implemented by the terminal 102 alone. Or the terminal 102 can send the OCR recognition instruction to the communicating server 104, and the server 104 implements the above OCR correction method based on the CTC loss. Taking the server 104 as an example, specifically, the server 104 extracts the text to be recognized, uses the OCR recognition model based on the CTC loss to perform character recognition on the text to be recognized, and obtains a CTC decoded character sequence matrix; the server 104 screens the characters that meet the character credibility threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix, and obtains a set of CTC character sub-matrices; the server 104 traverses the set of CTC character sub-matrices, and based on the size relationship between the confidence of each possible value of each character in each CTC character sub-matrix and the character credibility threshold, screens the possible values of the characters to construct a string sequence, and obtains a set of string sequences; the server 104 screens the optimal string sequence from the set of string sequences and outputs the optimal string sequence as the OCR correction result.

[0060] In one embodiment, as Figure 2 shown, a CTC-loss-based OCR correction method is provided. Taking the application of this method to the server as an example, the method includes the following steps:

[0061] Step S201: Extract the text to be recognized, and use the OCR recognition model based on the CTC loss to perform character recognition on the text to be recognized to obtain a CTC decoded character sequence matrix.

[0062] Among them, the text to be recognized is the text that needs to be recognized by OCR. CTC (Connectionist Temporal Classification) loss is a loss function applicable to solving the classification problem of time-series data. The CTC decoded character sequence matrix is constructed from the results of character recognition of the text to be recognized by the OCR recognition model based on the CTC loss.

[0063] In one embodiment, step S201 includes: detecting and intercepting text boxes in the image to be detected to obtain the text to be recognized; using an OCR character recognition model based on CTC loss to perform character recognition on the text to be recognized, obtaining a CTC decoded character sequence and the possible values of each character in the CTC decoded character sequence; constructing a CTC decoded character sequence matrix according to the length of the CTC decoded character sequence and the possible values of each character.

[0064] Specifically, after the server receives an OCR recognition instruction for a certain image to be detected, it first uses a text detection tool to extract the text boxes in the image to be detected, and then intercepts the detected text box area as the text to be recognized. As Figures 3 - 4 shown, Figure 3 a relatively blurred ID card image is provided as a schematic diagram of the image to be detected in this embodiment. Figure 3 The position of the rectangular box is the detected text box. Figure 4 For Figure 3 a schematic diagram of the text to be recognized with blurred text intercepted from the text box area. After the server obtains the text to be recognized, it sends it to the OCR recognition model. The OCR recognition model in this embodiment is a recognition model based on CTC loss. After receiving the input of the text to be recognized, the OCR recognition model based on CTC loss performs character recognition on the text to be recognized, obtaining a CTC decoded character sequence and the possible values of each character in the character sequence. As Figure 5 shown, a schematic diagram of a CTC decoded character sequence is provided. Figure 5 The CTC decoded character sequence shown is obtained by recognizing the text to be recognized shown in Figure 4 However, according to the length of the obtained CTC decoded character sequence and the possible values of each character, a CTC decoded character sequence matrix CTCSep is constructed. The structure of the CTC decoded character sequence matrix CTCSep is as follows:

[0065]

[0066] where n represents the length of the CTC decoded character sequence, that is, the CTC decoded character sequence contains a total of n characters. m represents the possible values of each character, that is, each character has m possible values. In this embodiment, the number of m is selected as the top m with the highest confidence. The specific number of m can be specified by modifying the decoding process of the OCR recognition model to calculate the top m possible values. The larger m is, the higher the accuracy, but the greater the computational load of the model. In this embodiment, usually m is less than 5, and preferably m = 3. It can be seen that α in the CTC decoded character sequence matrix i,jRepresents the i-th possible value of the j-th character, where j = {1, 2, ……, n} and i = {1, 2, ……, m}. Each possible value contains specific character information and the corresponding confidence level. For example Figure 5 The confidence level corresponding to the character "yǒu" in it is "0.9999829469219792".

[0067] Step S202: Screen characters that meet the character credibility threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix, and obtain a set of CTC character sub-matrices.

[0068] Specifically, after the server obtains the CTC decoded character sequence matrix, in order to correct and obtain a more accurate OCR recognition result, a preset character credibility threshold is used to screen out some characters from the CTC decoded character sequence matrix to construct multiple CTC character sub-matrices respectively, and the set of all CTC character sub-matrices is used as the set of CTC character sub-matrices. Among them, the character credibility threshold is a threshold preset for evaluating the credibility of characters, usually given by the OCR recognition model used. Different OCR recognition models give different character credibility thresholds, so the actually used character credibility threshold is determined based on the actually used OCR model. In the OCR recognition model of this embodiment, we consider characters with a threshold greater than 0.99 to be credible characters, so the credibility threshold of this embodiment is set to 0.99.

[0069] Step S203: Traverse the set of CTC character sub-matrices, and screen the possible values of characters based on the size relationship between the confidence level of each character's possible value in each CTC character sub-matrix and the character credibility threshold, and construct a string sequence to obtain a set of string sequences.

[0070] Specifically, after the server obtains the set of CTC character sub-matrices, traverse this set of CTC character sub-matrices, and construct a corresponding string sequence for each CTC character sub-matrix in it. The construction of the string sequence mainly compares the confidence level of each character's possible value in the CTC character sub-matrix with the preset character credibility threshold, and filters out a part of the possible values of characters from the CTC character sub-matrix according to the size relationship to obtain the string sequence. The string sequence is a sequence including the possible values of characters, and then the string sequences corresponding to all CTC character sub-matrices form a set of string sequences.

[0071] In one embodiment, step S203 includes: for each CTC character sub-matrix in the CTC character sub-matrix set, constructing an empty string sequence with a sequence length equal to the number of characters in the CTC character sub-matrix; respectively comparing the confidence of the first possible value of each character in the CTC character sub-matrix with the character confidence threshold; if the confidence corresponding to the first possible value is not less than the character confidence threshold, adding the first possible value to the empty string sequence, otherwise adding all possible values of the character in the CTC character sub-matrix to the empty string sequence to obtain an assigned string sequence; performing permutations and combinations on the assigned string sequence to enumerate all string sequences and obtain a string sequence set.

[0072] Specifically, first determine the number of characters in the CTC character sub-matrix and construct an empty string sequence with a sequence length equal to that number of characters. For example, if there are 6 characters in the CTC character sub-matrix, an empty string sequence of length 6 is constructed. An empty string sequence refers to a sequence without sequence elements.

[0073] Then, since a higher confidence generally indicates a higher character credibility and the more accurate the recognized character is. So when making the comparison, the confidence P(α 1,j (the possible values in the matrix are arranged in ascending order, and the first possible value is the possible value with the highest confidence)) of the first possible value α of each character in the CTC character sub-matrix 1,j is compared with the preset character confidence threshold of 0.99. When P(α 1,j ) is not less than (greater than or equal to) the character confidence threshold of 0.99, it means that the character already meets the requirements of this embodiment for a credible character, and the character is the most accurately recognized character. Then the first possible value α 1,j of the character is directly added to the corresponding sequence position of the empty string sequence, that is, α 1,j is added to the j-th sequence position in the empty string sequence as the j-th sequence element in the empty string sequence. When P(α 1,j ) is less than the character confidence threshold of 0.99, it means that the character does not yet meet the credible requirements of this embodiment as the most accurately recognized character. Therefore, all possible values of the character in the CTC character sub-matrix, that is, α 1,j to α m,j all possible values are added to the j-th sequence position in the empty string sequence as the j-th sequence element in the empty string sequence. At this time, the j-th sequence element includes m possible values.

[0074] After traversing all characters, the assignment of the empty string sequence is completed, and the string sequence after assignment is obtained. Since there may be multiple possible values for the string sequence after assignment, the string sequence after assignment is sorted and combined to enumerate all character sequences. For example, if P(α 1,j ) of 3 characters is less than 0.99, and each character corresponds to m possible values added to the string sequence, then the permutation and combination of the string sequence can enumerate m 3 string sequences. That is to say, since as long as there are characters in the CTC character submatrix that are less than the character confidence threshold, multiple corresponding string sequences can be enumerated through permutation and combination, so a CTC character submatrix may have multiple corresponding string sequences. For each CTC character submatrix in the CTC character submatrix set, the corresponding string sequence is obtained in the above manner. Whether the CTC character submatrix corresponds to only one string sequence or multiple character sequences, all the string sequences corresponding to all CTC character submatrices are combined to obtain a string sequence set.

[0075] Step S204, screen the optimal string sequence from the string sequence set, and output the optimal string sequence as the OCR correction result.

[0076] Specifically, after obtaining the string sequence set, score the string sequences in the string sequence set respectively. Select the optimal string sequence from them based on the score of each string sequence. The optimal string sequence is the optimal recognition result after correction, and it is output as the OCR correction result of this embodiment.

[0077] In one embodiment, step S204 includes: traversing each string sequence in the string sequence set, determining the string sequences that meet the adaptive requirements as candidate string sequences; taking the weighted average of the confidence levels of the candidate string sequences as the first score, and taking the edit distance between the candidate string sequences and the elements in the first row of the CTC decoding character sequence matrix as the second score; weighting the first score and the second score to obtain the final score of the candidate string sequences; determining the optimal string sequence from the candidate string sequences according to the final score; outputting the optimal string sequence in the character format of the text to be recognized to obtain the OCR correction result.

[0078] Among them, the adaptive requirement is the condition that needs to be met set according to the actual business requirements. It can be the regular rule that the field to be corrected needs to meet. The specific rule content is obtained according to prior knowledge or relevant laws. Since this embodiment is applied to the OCR recognition and correction of images, the adaptive requirement can be obtained from the rules of recognizing the characters on the image. For example, Figure 3Taking the ID card picture shown as an example, according to the position characteristics of the ID card picture, the characters of interest at this time are the start and end times of the ID card's validity period. It can be obtained from observation and ID card regulations that "validity period" is a fixed-format character on the ID card, and the variable characters are the start and end time values. Considering that the handwriting on the picture may be blurred and the characters "." and "-" on the validity period may not be recognized due to complex scenarios, but the format is relatively fixed. Therefore, the corresponding regular expression requirement Regex can be set as "Regex = "\d{4}\.?\d{2}\.?\d{2}-?\d{4}\.?\d{2}\.?\d{2}", indicating that the candidate string needs to meet this format requirement.

[0079] Specifically, after obtaining the set of string sequences, the server needs to screen out the optimal string sequence from them as the corrected OCR recognition result. First, the server judges the string sequences in the set of string sequences according to the preset adaptive requirements, directly removes those that do not meet the adaptive requirements, and determines the string sequences that meet the adaptive requirements as candidate string sequences. Then, for each candidate string sequence, calculate the string score, that is, the weighted average of the confidence of each character, as the first score Score1. At the same time, calculate the edit distance between the candidate string sequence and the elements in the first row of the CTC decoding character sequence matrix (α 1,1 , α 1,2 , ……, α 1,n ) as the second score Score2. Finally, calculate the final score of the candidate string sequence according to the first score Score1 and the second score Score2: Score = 0.7 * Score1 + 0.3 * Score1. Select the candidate string sequence with the highest Score as the optimal string sequence. Output the optimal string sequence in a certain format, and the specific format depends on the text to be recognized. For example, in this embodiment, when recognizing the ID card's validity period, it is output in the format YYYY.mm.dd - YYYY.mm.dd of the ID card's validity period in the ID card picture. Just normalize the format of the obtained optimal string sequence according to the relative position relationship as the date (first obtain all the numbers, the first 4 digits are the year, the 4 - 6 digits are the month, the 6 - 8 digits are the day, and connect them with "." in the middle. Then connect with the "-" symbol. Similarly, the 8 - 12 digits are the year, the 12 - 14 digits are the month, the 14 - 16 digits are the day, and connect them with ":" in the middle) and output the final result 2012.07.13 - 2021.07.13.

[0080] The above OCR correction method based on the CTC loss extracts the text to be recognized, uses the OCR recognition model based on the CTC loss to perform character recognition on the text to be recognized, and after obtaining the CTC decoded character sequence matrix, filters out the characters that meet the character confidence threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix, obtaining a set of CTC character sub-matrices. Then, it traverses the set of CTC character sub-matrices, filters the possible values of the characters based on the magnitude relationship between the confidence of each possible value of each character in each CTC character sub-matrix and the character confidence threshold, constructs a string sequence to obtain a set of string sequences; finally, it filters out the optimal string sequence from the set of string sequences and outputs the optimal string sequence as the OCR correction result. This method gradually filters to obtain the optimal solution by combining the confidence of each possible value of each character in the CTC decoded character sequence with the threshold, and can obtain high-accuracy recognition results even for OCR recognition in complex scenarios, improving the accuracy of OCR recognition.

[0081] In one embodiment, step S202 includes: comparing the confidence of the first possible value of each character in the CTC decoded character sequence matrix with the character confidence threshold respectively to determine whether the character is correct; when the character is correct, representing the character with a credible mode character, and when it is incorrect, representing the character with a non-credible mode character, constructing a mode character sequence corresponding to the CTC decoded character sequence matrix; according to the mode character sequence, enumerating all mode character sequences containing credible subsequences to obtain a set of mode character sequences; the credible subsequence is composed of credible mode characters; traversing the set of mode character sequences, filtering characters from the CTC decoded character sequence matrix according to the mode character sequence to construct a CTC character sub-matrix, obtaining a set of CTC character sub-matrices.

[0082] Among them, the credible mode character is used to indicate that the character is a credible character, that is, the probability that the character is a correct character is relatively high, and the non-credible mode character is used to indicate that the character is a non-credible character, that is, the probability that the character is an incorrect character is relatively high. Since the mode character sequence needs to be converted as a binary sequence later, in order to save resources and improve speed, in this embodiment, 1 is preferably used as the credible mode character and 0 as the non-credible mode character.

[0083] Specifically, traverse the CTC decoded character sequence matrix. For the character with P(α 1,j ) not less than the character confidence threshold of 0.99, it indicates that the character is correct and does not need to be modified or deleted, so it is represented by the credible mode character 1. And for the character with P(α 1,j ) less than the character confidence threshold of 0.99, it indicates that the character is incorrect and may need to be modified or deleted, so it is represented by the non-credible mode character 0. Finally, a mode character sequence in the form of 1110101...1 can be obtained, and the length of this mode character sequence is equal to the length of the CTC decoded character sequence. Figure 5Taking the CTC decoded character sequence shown as an example, its corresponding pattern character sequence ModelSeqStr = 111001111011111111111011, where the positions of the characters that are 0 indicate that the highest confidence at that position is less than the character confidence threshold of 0.99, and the positions of the characters that are 1 indicate that the highest confidence at that position is higher than the character confidence threshold of 0.99. The character sequence with all positions being 1 is the credible subsequence of the CTC decoded character sequence matrix. Combining Figure 5 it can be seen that the characters in the low-threshold part just correspond to Figure 4 the blurred parts on the image shown. Then, according to the obtained pattern character sequence, corresponding conversions are performed, and all pattern character sequences containing the credible subsequence are further enumerated to obtain a set of pattern character sequences. Finally, for each pattern character sequence in the set of pattern character sequences, characters are screened from the CTC decoded character matrix according to the pattern characters in the pattern character sequence to construct a CTC character sub-matrix. A corresponding CTC character sub-matrix is constructed for each pattern character sequence, and the collection of all CTC character sub-matrices is the set of CTC character sub-matrices.

[0084] In one embodiment, according to the pattern character sequence, all pattern character sequences containing the credible subsequence are enumerated to obtain a set of pattern character sequences, including: taking the pattern character sequence as a binary character sequence, using the decimal value converted from the binary as the start node, and determining the value according to the length of the character sequence as the end node; traversing from the start node to the end node, performing a bitwise OR operation on each integer with the decimal value and converting it into a binary sequence to obtain a pattern character sequence containing the credible subsequence; combining each pattern character sequence containing the credible subsequence to obtain a set of pattern character sequences.

[0085] Specifically, when enumerating all pattern character sequences, the currently obtained pattern character sequence is taken as a binary character sequence. Since 1 and 0 have been used as pattern characters to represent the pattern character sequence previously, the pattern character sequence can be directly used as a binary character sequence for conversion. According to the binary conversion to a decimal value, this decimal value is denoted as the start node Begin. At the same time, according to the length n of the character sequence, calculate 2 nThe value of -1 is used as the end node End. Then, traverse from Begin to End. For any integer T among them, calculate the bitwise OR V with the decimal value as the start node, that is, V = T | Begin. Further convert V into a binary sequence bitV, so as to obtain a pattern character sequence containing a credible subsequence. The collection of pattern character sequences corresponding to each integer is used as the pattern character sequence set. Taking ModelSeqStr = 111001111011111111111011 as an example, the pattern character sequence set ModelSeqStrCollection = ["111011111111111111111111","111011111111111111111011","111011111011111111111111","111011111011111111111011","111001111111111111111111","111001111111111111111011","111001111011111111111111","111001111011111111111011","111111111111111111111111","111111111111111111111011","111111111011111111111111","111111111011111111111011","111101111111111111111111","111101111111111111111011","111101111011111111111111","111101111011111111111011"] can be obtained, and each pattern character sequence includes a credible subsequence of the CTC decoded character sequence.

[0086] In another embodiment, according to the pattern character sequence, all pattern character sequences containing a trusted subsequence are enumerated to obtain a set of pattern character sequences, including: constructing an empty set of pattern character sequences, and traversing each pattern character in the pattern character sequence; if the currently traversed pattern character is the first one and is a trusted pattern character, add a pattern character sequence with the first character being the trusted pattern character to the set of pattern character sequences; if the currently traversed pattern character is the first one and is an untrusted pattern character, add a pattern character sequence with the first character being the trusted pattern character and a pattern character sequence with the first character being the untrusted pattern character to the set of pattern character sequences; if the currently traversed pattern character is not the first one and is a trusted pattern character, append a trusted pattern character to the end of each existing pattern character sequence in the set of pattern characters; if the currently traversed pattern character is not the first one and is an untrusted pattern character, append a trusted pattern character to the end of each existing pattern character sequence in the set of pattern characters, and copy the existing pattern character sequences, and append an untrusted pattern character to the end of each copied pattern character sequence, doubling the number of pattern character sequences in the set of pattern character sequences.

[0087] Specifically, since enumerating all pattern character sequences containing a trusted subsequence is equivalent to enumerating all pattern character sequences containing the trusted pattern character 1. Therefore, in this embodiment, an empty set of pattern character sequences can be constructed first, and then the currently obtained pattern character sequence ModelSeqStr can be traversed bit by bit. Based on the trustworthiness of each pattern character in the current pattern character sequence ModelSeqStr, multiple pattern character sequences containing trusted subsequences are added to the empty set of pattern character sequences, so as to obtain a set of pattern character sequences containing all trusted subsequences. During the process of traversing and adding pattern characters, different situations are divided according to whether the traversed pattern character is the trusted pattern character 1 or the untrusted pattern character 0 for addition.

[0088] When the k-th pattern character traversed is the first character and is a trusted pattern character 1, since the first character indicates that the set of pattern character sequences is still empty at this time, a pattern character sequence with the first character being the trusted character pattern 1 is directly added to the set of pattern character sequences. When the k-th pattern character traversed is not the first character and is a trusted pattern character 1, since it is no longer the first character, the set of pattern character sequences is not empty at this time. Then, a trusted pattern character 1 is appended to the end of each existing pattern character sequence in the set of pattern character sequences. For example, if there are three pattern character sequences "110", "101", and "111" in the set of pattern character sequences at this time, a "1" is appended to the end of each pattern character sequence. The resulting set of pattern character sequences is "1101", "1011", and "1111". When the k-th pattern character traversed is the first character but is an untrusted pattern character 0, similarly, the set of pattern character sequences is empty at this time. Different from the above, in addition to directly adding a pattern character sequence with the first character being the trusted character pattern 1 to the set of pattern character sequences, a pattern character sequence with the first character being the untrusted pattern character 0 also needs to be added. That is, two pattern character sequences are added to the set of pattern character sequences. One has the first pattern character as the trusted pattern character 1, and the other has the first pattern character as the untrusted pattern character 0. When the k-th pattern character traversed is not the first character but is an untrusted pattern character 0, a trusted pattern character 1 is appended to the end of each existing pattern character sequence in the set of pattern character sequences. At the same time, the existing pattern character sequences are copied, and an untrusted pattern character 0 is appended to the end of the copied pattern character sequences, doubling the number of pattern character sequences in the set of pattern character sequences. For example, if the set of pattern character sequences at this time includes three pattern character sequences "110", "101", and "111", while a "1" is appended to the end of each pattern character sequence respectively, "110", "101", and "111" are copied and a "0" is appended. The resulting set of pattern character sequences is "1101", "1011", "1111", "1100", "1010", and "1110", which is equivalent to appending a "1" and a "0" to the end of each existing pattern character sequence to obtain two sequences. In this embodiment, the set of pattern character sequences is enumerated by the direct addition method through traversal, which can obtain the set of pattern character sequences faster than by converting binary calculation, thereby improving the processing efficiency of calibration.

[0089] In one embodiment, a set of pattern character sequences is traversed, and characters are selected from a CTC decoded character sequence matrix according to the pattern character sequence to construct a CTC character submatrix to obtain a set of CTC character submatrices, including: constructing an empty CTC character submatrix; traversing pattern characters in the pattern character sequence, and if the kth pattern character currently traversed is a credible pattern character, adding the kth column in the CTC decoded character sequence matrix to the empty CTC character submatrix to obtain a CTC character submatrix corresponding to the CTC decoded character sequence matrix; and combining the CTC character submatrices of each pattern character sequence into a set of CTC character submatrices.

[0090] Specifically, first construct an empty CTC character submatrix that does not contain matrix elements, and traverse the pattern character sequence. If the k-th pattern character currently traversed is the trusted pattern character 1, then the matrix elements in the k-th column are added to the empty CTC character submatrix, that is, the matrix elements at the same position in the CTC decoding character sequence matrix are added to the empty CTC character submatrix by column. If the k-th character pattern currently traversed is the untrusted pattern character 0, then skip the character and continue traversing, that is, the matrix elements corresponding to the character will not be selected and added to the empty CTC character submatrix. Until the entire pattern character sequence is traversed, the CTC character submatrix corresponding to the pattern character sequence is obtained. For each pattern character sequence in the pattern character sequence set, the corresponding CTC character submatrix is ​​obtained in this way, thereby obtaining a CTC character submatrix set. As Figure 6 As shown, a schematic diagram of a CTC character sub-matrix corresponding to a CTC character sub-sequence is provided. Figure 6 is based on Figure 5 The CTC character subsequence of the CTC decoded character sequence is shown.

[0091] It should be understood that although Figure 2 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0092] In one embodiment, Figure 7 As shown, an OCR correction device based on CTC loss is provided, comprising:

[0093] A character recognition module, configured to extract the text to be recognized, and perform character recognition on the text to be recognized by using an OCR recognition model based on the CTC loss to obtain a CTC decoded character sequence matrix;

[0094] A sub-matrix construction module, configured to screen characters that meet the character credibility threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix, and obtain a set of CTC character sub-matrices;

[0095] A string sequence construction module, configured to traverse the set of CTC character sub-matrices, screen the possible values of characters based on the magnitude relationship between the confidence of each possible value of each character in each CTC character sub-matrix and the character credibility threshold, construct a string sequence, and obtain a set of string sequences;

[0096] An output module, configured to screen the optimal string sequence from the set of string sequences, and output the optimal string sequence as the OCR correction result.

[0097] In one embodiment, the character recognition module is further configured to detect and intercept text boxes in the image to be detected to obtain the text to be recognized; perform character recognition on the text to be recognized by using an OCR character recognition model based on the CTC loss, obtain the CTC decoded character sequence and the possible values of each character in the CTC decoded character sequence; and construct a CTC decoded character sequence matrix according to the length of the CTC decoded character sequence and the possible values of each character.

[0098] In one embodiment, the sub-matrix construction module is further configured to respectively compare the confidence of the first possible value of each character in the CTC decoded character sequence matrix with the character credibility threshold to determine whether the character is correct; when the character is correct, represent the character with a trusted mode character, and when it is incorrect, represent the character with a non-trusted mode character, and construct a mode character sequence corresponding to the CTC decoded character sequence matrix; according to the mode character sequence, enumerate all mode character sequences including trusted subsequences to obtain a set of mode character sequences; the trusted subsequence is composed of trusted mode characters; traverse the set of mode character sequences, and screen characters from the CTC decoded character sequence matrix according to the mode character sequence to construct a CTC character sub-matrix, and obtain a set of CTC character sub-matrices.

[0099] In one embodiment, the sub-matrix construction module is further configured to use the mode character sequence as a binary character sequence, use the decimal value obtained by converting the binary as the start node, and use the value determined according to the length of the character sequence as the end node; traverse from the start node to the end node, perform a bitwise OR operation on each integer with the decimal value and convert it into a binary sequence to obtain a mode character sequence including a trusted subsequence; combine each mode character sequence including a trusted subsequence to obtain a set of mode character sequences.

[0100] In one embodiment, the sub-matrix construction module is further configured to construct an empty set of pattern character sequences, and traverse each pattern character in the pattern character sequences; if the currently traversed pattern character is the first character and is a trustworthy pattern character, add a pattern character sequence with the first character being a trustworthy pattern character to the set of pattern character sequences; if the currently traversed pattern character is the first character and is an untrustworthy pattern character, add a pattern character sequence with the first character being a trustworthy pattern character and a pattern character sequence with the first character being an untrustworthy pattern character to the set of pattern character sequences; if the currently traversed pattern character is not the first character and is a trustworthy pattern character, append a trustworthy pattern character to the tails of all existing pattern character sequences in the pattern character set; if the currently traversed pattern character is not the first character and is an untrustworthy pattern character, append a trustworthy pattern character to the tails of all existing pattern character sequences in the pattern character set respectively, and copy the existing pattern character sequences, and append an untrustworthy pattern character to the tails of the copied pattern character sequences respectively, doubling the number of pattern character sequences in the pattern character set.

[0101] In one embodiment, the sub-matrix construction module is further configured to construct an empty CTC character sub-matrix; traverse the pattern characters in the pattern character sequences, if the k-th pattern character currently traversed is a trustworthy pattern character, add the k-th column in the CTC decoded character sequence matrix to the empty CTC character sub-matrix to obtain a CTC character sub-matrix corresponding to the CTC decoded character sequence matrix; combine the CTC character sub-matrices of each pattern character sequence into a set of CTC character sub-matrices.

[0102] In one embodiment, the string sequence construction module is further configured to, for each CTC character sub-matrix in the set of CTC character sub-matrices, construct an empty string sequence with the sequence length equal to the number of characters in the CTC character sub-matrix; compare the confidence levels of the first possible values of each character in the CTC character sub-matrix with the character credibility threshold respectively; if the confidence level corresponding to the first possible value is not less than the character credibility threshold, add the first possible value to the empty string sequence, otherwise add all possible values of the character in the CTC character sub-matrix to the empty string sequence to obtain an assigned string sequence; perform permutation and combination on the assigned string sequence, and enumerate all string sequences to obtain a set of string sequences.

[0103] In one embodiment, the output module is further configured to traverse each string sequence in the set of string sequences, determine the string sequences that meet the adaptive requirements as candidate string sequences; take the weighted average confidence of the candidate string sequences as the first score, and take the edit distance between the candidate string sequences and the elements in the first row of the CTC decoding character sequence matrix as the second score; weight the first score and the second score to obtain the final score of the candidate string sequences; determine the optimal string sequence from the candidate string sequences according to the final score; and output the optimal string sequence in the character format of the text to be recognized to obtain the OCR correction result.

[0104] For the specific limitations of the OCR correction device based on the CTC loss, reference can be made to the limitations of the OCR correction method based on the CTC loss in the above text, which will not be elaborated here. Each module in the above OCR correction device based on the CTC loss can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above various OCR correction method embodiments based on the CTC loss can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc.

[0105] In one embodiment, a computer device is provided. The computer device can be a server and includes a processor, a memory, and a network interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an OCR correction method based on the CTC loss. Exemplarily, the computer program can be divided into one or more modules. One or more modules are stored in the memory and are executed by the processor to complete the present invention. One or more modules can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device.

[0106] The so-called processor may be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and circuits.

[0107] The memory can be used to store the computer program and / or modules. By running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory, the processor realizes various functions of the computer device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as hard disks, memory, plug-in hard disks, SmartMedia Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0108] Those skilled in the art can understand that the computer device structure shown in this embodiment is only a partial structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the present invention is applied. The specific computer device may include more or fewer components, or combine certain components, or have different component arrangements.

[0109] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0110] Extract the text to be recognized, and use the OCR recognition model based on the CTC loss to perform character recognition on the text to be recognized, and obtain the CTC decoded character sequence matrix;

[0111] Screen the characters that meet the character credibility threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix, and obtain a set of CTC character sub-matrices;

[0112] Traverse the set of CTC character sub-matrices, screen the possible values of the characters based on the relationship between the confidence of the possible values of each character in each CTC character sub-matrix and the character credibility threshold, construct a string sequence, and obtain a set of string sequences;

[0113] Screen the optimal string sequence from the set of string sequences, and output the optimal string sequence as the OCR correction result.

[0114] In one embodiment, when the processor executes the computer program, the following steps are further implemented: Detect and intercept the text box in the image to be detected to obtain the text to be recognized; Use the OCR character recognition model based on the CTC loss to perform character recognition on the text to be recognized, and obtain the CTC decoded character sequence and the possible values of each character in the CTC decoded character sequence; According to the length of the CTC decoded character sequence and the possible values of each character, construct a CTC decoded character sequence matrix.

[0115] In one embodiment, when the processor executes the computer program, the following steps are further implemented: Compare the confidence of the first possible value of each character in the CTC decoded character sequence matrix with the character credibility threshold respectively to determine whether the character is correct; When the character is correct, represent the character with a credible mode character, and when it is incorrect, represent the character with a non-credible mode character, and construct a mode character sequence corresponding to the CTC decoded character sequence matrix; According to the mode character sequence, enumerate all mode character sequences containing a credible subsequence to obtain a set of mode character sequences; A credible subsequence is composed of credible mode characters; Traverse the set of mode character sequences, and screen characters from the CTC decoded character sequence matrix according to the mode character sequence to construct a CTC character sub-matrix, and obtain a set of CTC character sub-matrices.

[0116] In one embodiment, when the processor executes the computer program, the following steps are further implemented: Take the mode character sequence as a binary character sequence, use the binary converted to a decimal value as the start node, and use the value determined according to the length of the character sequence as the end node; Traverse from the start node to the end node, perform a bitwise OR calculation on each integer with the decimal value and convert it into a binary sequence to obtain a mode character sequence containing a credible subsequence; Combine each mode character sequence containing a credible subsequence to obtain a set of mode character sequences.

[0117] In one embodiment, when the processor executes the computer program, the following steps are further implemented: constructing an empty set of pattern character sequences, and traversing each pattern character in the pattern character sequence; if the currently traversed pattern character is the first character and is a trustworthy pattern character, adding a pattern character sequence with the first character being a trustworthy pattern character to the set of pattern character sequences; if the currently traversed pattern character is the first character and is an untrustworthy pattern character, adding a pattern character sequence with the first character being a trustworthy pattern character and a pattern character sequence with the first character being an untrustworthy pattern character to the set of pattern character sequences; if the currently traversed pattern character is not the first character and is a trustworthy pattern character, appending a trustworthy pattern character to the tail of each existing pattern character sequence in the pattern character set; if the currently traversed pattern character is not the first character and is an untrustworthy pattern character, appending a trustworthy pattern character to the tail of each existing pattern character sequence in the pattern character set, and copying the existing pattern character sequences, and appending an untrustworthy pattern character to the tail of each copied pattern character sequence, so that the number of pattern character sequences in the pattern character set doubles.

[0118] In one embodiment, when the processor executes the computer program, the following steps are further implemented: constructing an empty CTC character sub-matrix; traversing the pattern characters in the pattern character sequence, if the k-th pattern character currently traversed is a trustworthy pattern character, adding the k-th column in the CTC decoded character sequence matrix to the empty CTC character sub-matrix to obtain a CTC character sub-matrix corresponding to the CTC decoded character sequence matrix; combining the CTC character sub-matrices of each pattern character sequence into a set of CTC character sub-matrices.

[0119] In one embodiment, when the processor executes the computer program, the following steps are further implemented: for each CTC character sub-matrix in the set of CTC character sub-matrices, respectively constructing an empty string sequence with the sequence length equal to the number of characters in the CTC character sub-matrix; comparing the confidence levels of the first possible values of each character in the CTC character sub-matrix with the character credibility threshold respectively; if the confidence level corresponding to the first possible value is not less than the character credibility threshold, adding the first possible value to the empty string sequence, otherwise adding all possible values of the character in the CTC character sub-matrix to the empty string sequence to obtain an assigned string sequence; performing permutation and combination on the assigned string sequence to enumerate all string sequences to obtain a set of string sequences.

[0120] In one embodiment, when the processor executes the computer program, the following steps are further implemented: traversing each string sequence in the set of string sequences, determining the string sequences that meet the adaptive requirements as candidate string sequences; taking the weighted average of the confidence levels of the candidate string sequences as the first score, and taking the edit distance between the candidate string sequences and the elements in the first row of the CTC decoding character sequence matrix as the second score; weighting the first score and the second score to obtain the final score of the candidate string sequences; determining the optimal string sequence from the candidate string sequences according to the final score; and outputting the optimal string sequence in the character format of the text to be recognized to obtain the OCR correction result.

[0121] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0122] Extracting the text to be recognized, and using an OCR recognition model based on CTC loss to perform character recognition on the text to be recognized to obtain a CTC decoding character sequence matrix;

[0123] Screening the characters that meet the character confidence threshold from the CTC decoding character sequence matrix to construct a CTC character sub-matrix, and obtaining a set of CTC character sub-matrices;

[0124] Traversing the set of CTC character sub-matrices, screening the possible values of the characters based on the size relationship between the confidence levels of the possible values of each character in each CTC character sub-matrix and the character confidence threshold, and constructing string sequences to obtain a set of string sequences;

[0125] Screening the optimal string sequence from the set of string sequences, and outputting the optimal string sequence as the OCR correction result.

[0126] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented: detecting and intercepting the text boxes in the image to be detected to obtain the text to be recognized; using an OCR character recognition model based on CTC loss to perform character recognition on the text to be recognized to obtain the CTC decoding character sequence and the possible values of each character in the CTC decoding character sequence; and constructing a CTC decoding character sequence matrix according to the length of the CTC decoding character sequence and the possible values of each character.

[0127] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: comparing the confidence levels of the first possible values of each character in the CTC decoded character sequence matrix with a character credibility threshold respectively to determine whether the character is correct; when the character is correct, representing the character with a credible mode character, and when it is incorrect, representing the character with a non-credible mode character, to construct a mode character sequence corresponding to the CTC decoded character sequence matrix; according to the mode character sequence, enumerating all mode character sequences containing a credible subsequence to obtain a set of mode character sequences; the credible subsequence is composed of credible mode characters; traversing the set of mode character sequences, screening characters from the CTC decoded character sequence matrix according to the mode character sequence to construct a CTC character sub-matrix, and obtaining a set of CTC character sub-matrices.

[0128] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: taking the mode character sequence as a binary character sequence, using the binary converted to a decimal value as the start node, and determining a value according to the length of the character sequence as the end node; traversing from the start node to the end node, performing a bitwise OR calculation on each integer with the decimal value and converting it into a binary sequence to obtain a mode character sequence containing a credible subsequence; combining each mode character sequence containing a credible subsequence to obtain a set of mode character sequences.

[0129] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: constructing an empty set of mode character sequences and traversing each mode character in the mode character sequence; if the currently traversed mode character is the first one and is a credible mode character, adding a mode character sequence with the first character being a credible mode character to the set of mode character sequences; if the currently traversed mode character is the first one and is a non-credible mode character, adding a mode character sequence with the first character being a credible mode character and a mode character sequence with the first character being a non-credible mode character to the set of mode character sequences; if the currently traversed mode character is not the first one and is a credible mode character, appending a credible mode character to the tails of all existing mode character sequences in the set of mode characters; if the currently traversed mode character is not the first one and is a non-credible mode character, appending a credible mode character to the tails of all existing mode character sequences in the set of mode characters respectively, and copying the existing mode character sequences, appending a non-credible mode character to the tails of the copied mode character sequences respectively, and the number of mode character sequences in the set of mode characters doubles.

[0130] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: constructing an empty CTC character sub-matrix; traversing the pattern characters in the pattern character sequence, and if the k-th pattern character currently traversed is a credible pattern character, adding the k-th column in the CTC decoded character sequence matrix to the empty CTC character sub-matrix to obtain a CTC character sub-matrix corresponding to the CTC decoded character sequence matrix; combining the CTC character sub-matrices of each pattern character sequence into a CTC character sub-matrix set.

[0131] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: for each CTC character sub-matrix in the CTC character sub-matrix set, respectively constructing an empty string sequence with a sequence length equal to the number of characters in the CTC character sub-matrix; comparing the confidence levels of the first possible values of each character in the CTC character sub-matrix with the character credibility threshold respectively; if the confidence level corresponding to the first possible value is not less than the character credibility threshold, adding the first possible value to the empty string sequence, otherwise adding all possible values of the character in the CTC character sub-matrix to the empty string sequence to obtain an assigned string sequence; performing permutation and combination on the assigned string sequence to enumerate all string sequences to obtain a string sequence set.

[0132] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: traversing each string sequence in the string sequence set, determining the string sequences that meet the adaptive requirements as candidate string sequences; taking the weighted average value of the confidence levels of the candidate string sequences as the first score, and taking the edit distance between the candidate string sequences and the elements in the first row of the CTC decoded character sequence matrix as the second score; weighting the first score and the second score to obtain the final score of the candidate string sequences; determining the optimal string sequence from the candidate string sequences according to the final score; outputting the optimal string sequence in the character format of the text to be recognized to obtain the OCR correction result.

[0133] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0134] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0135] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. An OCR correction method based on CTC loss, characterized in that Including: Extract the text to be recognized, and use an OCR recognition model based on CTC loss to perform character recognition on the text to be recognized to obtain a CTC decoded character sequence matrix; Screen characters that meet the character confidence threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix, and obtain a set of CTC character sub-matrices; Traverse the set of CTC character sub-matrices, and screen the possible values of characters based on the size relationship between the confidence of each possible value of each character in each CTC character sub-matrix and the character confidence threshold, and construct a string sequence to obtain a set of string sequences; Screen the optimal string sequence from the set of string sequences, and output the optimal string sequence as the OCR correction result; specifically including: Traverse each string sequence in the set of string sequences, and determine the string sequence that meets the adaptive requirements as the candidate string sequence; Take the weighted average of the confidence of the candidate string sequence as the first score, and take the edit distance between the candidate string sequence and the first row elements in the CTC decoded character sequence matrix as the second score; Weight the first score and the second score to obtain the final score of the candidate string sequence; Determine the optimal string sequence from the candidate string sequences according to the final score; Output the optimal string sequence in the character format of the text to be recognized to obtain the OCR correction result.

2. The method according to claim 1, characterized in that, The step of extracting the text to be recognized, and using an OCR recognition model based on CTC loss to perform character recognition on the text to be recognized to obtain a CTC decoded character sequence matrix includes: Detect and intercept the text box in the image to be detected to obtain the text to be recognized; Use an OCR character recognition model based on CTC loss to perform character recognition on the text to be recognized, and obtain the CTC decoded character sequence and the possible values of each character in the CTC decoded character sequence; Construct a CTC decoded character sequence matrix according to the length of the CTC decoded character sequence and the possible values of each character.

3. The method according to claim 1, wherein The step of screening characters that meet the character confidence threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix and obtain a set of CTC character sub-matrices includes: Compare the confidence of the first possible value of each character in the CTC decoded character sequence matrix with the character confidence threshold respectively to determine whether the character is correct; When the character is correct, represent the character with a credible mode character, and when it is incorrect, represent the character with a non-credible mode character, and construct a mode character sequence corresponding to the CTC decoded character sequence matrix; According to the mode character sequence, enumerate all mode character sequences containing credible subsequences to obtain a set of mode character sequences; the credible subsequence is composed of credible mode characters; Traverse the set of mode character sequences, and screen characters from the CTC decoded character sequence matrix according to the mode character sequence to construct a CTC character sub-matrix to obtain a set of CTC character sub-matrices.

4. The method according to claim 3, wherein The step of enumerating all mode character sequences containing credible subsequences according to the mode character sequence to obtain a set of mode character sequences includes: Take the pattern character sequence as a binary character sequence, use the decimal value obtained by converting the binary as the start node, and determine the value according to the length of the character sequence as the end node; Traverse from the start node to the end node, perform a bitwise OR operation on each integer with the decimal value and convert it into a binary sequence to obtain a pattern character sequence containing a trusted subsequence; Combine each pattern character sequence containing a trusted subsequence to obtain a set of pattern character sequences.

5. The method according to claim 3, wherein The enumeration of all pattern character sequences containing trusted subsequences according to the pattern character sequence to obtain a set of pattern character sequences includes: Construct an empty set of pattern character sequences and traverse each pattern character in the pattern character sequence; If the currently traversed pattern character is the first one and is a trusted pattern character, add a pattern character sequence with the first character being a trusted pattern character to the set of pattern character sequences; If the currently traversed pattern character is the first one and is an untrusted pattern character, add a pattern character sequence with the first character being a trusted pattern character and a pattern character sequence with the first character being an untrusted pattern character to the set of pattern character sequences; If the currently traversed pattern character is not the first one and is a trusted pattern character, append a trusted pattern character to the end of each existing pattern character sequence in the set of pattern characters; If the currently traversed pattern character is not the first one and is an untrusted pattern character, append a trusted pattern character to the end of each existing pattern character sequence in the set of pattern characters, and copy the existing pattern character sequences, and append an untrusted pattern character to the end of the copied pattern character sequences, doubling the number of pattern character sequences in the set of pattern character sequences.

6. The method according to claim 3, wherein The traversal of the set of pattern character sequences, and the screening of characters from the CTC decoded character sequence matrix according to the pattern character sequence to construct a CTC character sub-matrix to obtain a set of CTC character sub-matrices includes: Construct an empty CTC character sub-matrix; Traverse the pattern characters in the pattern character sequence. If the k-th pattern character currently traversed is a trusted pattern character, add the k-th column of the CTC decoded character sequence matrix to the empty CTC character sub-matrix to obtain a CTC character sub-matrix corresponding to the CTC decoded character sequence matrix; Combine the CTC character sub-matrices of each pattern character sequence into a set of CTC character sub-matrices.

7. The method according to claim 1, wherein The traversal of the set of CTC character sub-matrices, and the screening of character possible values based on the relationship between the confidence of each character possible value in each CTC character sub-matrix and the character credibility threshold to construct a string sequence to obtain a set of string sequences includes: For each CTC character sub-matrix in the set of CTC character sub-matrices, construct an empty string sequence with the sequence length equal to the number of characters in the CTC character sub-matrix; Compare the confidence of the first possible value of each character in the CTC character sub-matrix with the character credibility threshold respectively; When the confidence corresponding to the first possible value is not less than the character credibility threshold, add the first possible value to the empty string sequence; otherwise, add all possible values of the character in the CTC character sub-matrix to the empty string sequence to obtain the assigned string sequence. Perform permutations and combinations on the assigned string sequence, enumerate all string sequences, and obtain a set of string sequences.

8. An OCR correction device based on CTC loss, characterized in that, Including: A character recognition module, configured to extract the text to be recognized, and perform character recognition on the text to be recognized by using an OCR recognition model based on the CTC loss to obtain a CTC decoded character sequence matrix. A sub-matrix construction module, configured to screen characters that meet the character credibility threshold from the CTC decoded character sequence matrix to construct a CTC character sub-matrix, and obtain a set of CTC character sub-matrices. A string sequence construction module, configured to traverse the set of CTC character sub-matrices, screen the possible values of the characters based on the magnitude relationship between the confidence of each possible value of each character in each CTC character sub-matrix and the character credibility threshold, construct a string sequence, and obtain a set of string sequences. An output module, configured to screen the optimal string sequence from the set of string sequences and output the optimal string sequence as the OCR correction result; specifically, configured to traverse each string sequence in the set of string sequences, determine the string sequence that meets the adaptive requirements as the candidate string sequence; take the weighted average of the confidence of the candidate string sequence as the first score, and take the edit distance between the candidate string sequence and the first row elements in the CTC decoded character sequence matrix as the second score. Weight the first score and the second score to obtain the final score of the candidate string sequence. Determine the optimal string sequence from the candidate string sequences according to the final score; output the optimal string sequence in the character format of the text to be recognized to obtain the OCR correction result.

9. A computer device, comprising a processor and a memory, the memory storing a computer program, characterized in that, When the processor is used to execute the computer program, it implements the OCR correction method based on the CTC loss according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the OCR correction method based on the CTC loss according to any one of claims 1-7.

Citation Information

Patent Citations

  • Training data cleaning method and device, computer equipment and storage medium

    CN110377591A

  • Auto-correction of pattern defined strings

    US10963717B1