A method for automatically correcting character recognition results
The character recognition results are filtered and filled through the longest common subsequence method, which solves the problems of inkjet recognition errors, error detection, and missed detection during the production process, and achieves more accurate recognition results and the effect of reducing production costs.
Patent Information
- Application Number
- CN202210257337.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-16
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-03-16
AI Technical Summary
During the production process, due to production line speed jitters and occasional conditions of the inkjet printer, there are problems such as identification errors, error detection, and missed detection of the product inkjet code, and the existing technology is difficult to effectively solve these problems.
The method of the longest common subsequence is adopted, by obtaining the predicted character information and standard character information in the image to be processed, the candidate common subsequence is found, the rationality is traversed and judged, and the predicted character information is filtered and filled with the reasonable common subsequence, and the recognition results are corrected.
It effectively solves the problems of misdetection, misdetection and misdetection in character recognition results, improves the accuracy of recognition results, reduces the probability of normal products being mistakenly eliminated, and reduces production costs.
Smart Images

Figure CN114495115B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of optical character recognition, and particularly relates to a method for automatically correcting character recognition results. Background Art
[0002] For common production scenarios in daily life such as video packaging, pharmaceutical packaging, canned or boxed beverages, etc., the character inkjet content of the products produced on the same day is the same. However, during the production process, due to the jitter of the production line speed and occasional problems of the inkjet printer, the inkjet of the products produced at a certain moment has problems. For example: the inkjet cannot be recognized, the inkjet content is incorrect, the inkjet is missing, etc.
[0003] For the solution to the character detection and recognition results of this type, whether it is traditional image processing or using a convolutional neural network to detect and recognize the characters, a problem will be faced, that is, there are false detections or missed detections for normal products. When the external factors such as the illumination of the production environment change and the imaging is unstable, the probability of false detections or missed detections will be greater. By calculating the IOU value between detection boxes and then screening unreasonable detection results by setting a threshold, the problem that occurs when the detection boxes overlap can be solved, but it is powerless for the discrete boxes of false detections, and it is even more impossible to process the missed detected characters and misrecognized characters. In this case, a considerable number of normal products will be removed as abnormal products, which greatly increases the production cost. The present invention uses the longest common subsequence method to solve the problems of false detections, missed detections, and misrecognition that occur during the detection and recognition results.
[0004] In production, there will also be changes in the size of character inkjet and the spacing between inkjet characters. These changes will make some algorithm processes ineffective, such as judging whether the character size is within the limited range. Benefiting from the use of relative information between characters, the present invention can also solve the problems of size and character spacing changes caused by character inkjet. Summary of the Invention
[0005] In view of the above technical problems, the present invention provides a method for automatically correcting character recognition results.
[0006] The technical solution adopted by the present invention to solve its technical problems is:
[0007] A method for automatically correcting character recognition results, the method comprising the following steps:
[0008] Step S100: Obtain the predicted character information in the image to be processed and the standard character information in the same scene;
[0009] Step S200: Find candidate common subsequences in the predicted character information and the standard character information;
[0010] Step S300: Traverse the candidate common subsequences, and determine whether a candidate common subsequence is a reasonable subsequence. If so, end the traversal;
[0011] Step S400: Use the reasonable common subsequence to filter and fill the predicted character information to obtain the processed predicted character information;
[0012] Step S500: Correct the predicted character information according to the processed predicted character information, and output the correct character information.
[0013] Preferably, in step S100, the predicted character information includes the position information and category information of each predicted character. The position information includes the center point coordinates and the character width and height information of the character, and the category information is the probability that each character belongs to each category.
[0014] Preferably, step S200 includes:
[0015] Step S210: Combine the character category labels of the standard character information and the predicted character information respectively into strings Rs and Rp in the order of first row and then column;
[0016] Step S220: Calculate the length N of the longest common subsequence and the candidate longest common subsequence matrix P of the two strings Rs and Rp. Specifically:
[0017] Assume that the lengths of the strings Rs and Rp are Ls and Lp respectively. Construct a dynamic programming matrix A with Lp + 1 rows and Ls + 1 columns. The first row and the first column of A are initialized to 0, and the values of the remaining positions are calculated according to formula (1):
[0018]
[0019] where A[i, j] is the value of the i-th row and j-th column in the matrix A, A[i - 1, j] is the value of the i - 1-th row and j-th column in the matrix A, A[i, j - 1] is the value of the i-th row and j - 1-th column in the matrix A, S s [i] is the i-th standard character information, S p [j] is the j-th predicted character information, L p +1 represents the total number of rows of the matrix A, L s +1 represents the total number of columns of the matrix A;
[0020] Calculate the value of N according to formula (2). Specifically:
[0021] N = max(A) (2)
[0022] where N is the maximum value of the matrix A;
[0023] If N < 2, it means that the candidate common subsequence is not successfully found, and step S200 ends;
[0024] If N ≥ 2, traverse matrix A. For the element A[i, j] in the i-th row and j-th column of matrix A, if it satisfies equation (3), add a new row to P, and at the same time store the values of A[i, j], i, j, and Ss[i] into the newly added row in P;
[0025]
[0026] Step S230: Delete the unreasonable position matches in P, specifically:
[0027] Traverse each row in P. For the four elements P[i, 0], P[i, 1], P[i, 2], P[i, 3] in the i-th row of P, if they satisfy equation (4), then delete this row from P;
[0028] L p -P[i, 1] < N - P[i, 0] or L s -P[i, 2] < N - P[i, 0] (4)
[0029] Among them, P[i, 1] is the element in the 2nd column of the i-th row in matrix P, L p is the length of Rp, L s is the length of Rs;
[0030] Step S240: Traverse the first column of the candidate longest common subsequence matrix P, count the number of matching positions at each position of the longest common subsequence, and construct all candidate common subsequences Q, including:
[0031] The possible number of matches n at the k-th position of the longest common subsequence with length N (N ≥ 2) k is the number of times k appears in the first column of P, and is calculated according to equation (5), specifically:
[0032] n k =count(P[:, 0] = k) (5)
[0033] Among them, the count function represents statistical calculation on numerical data;
[0034] Then all possible candidate common subsequences Q have a total of M combination ways, and the value of M is calculated according to equation (6), specifically:
[0035]
[0036] Among them, M is the number of all candidate common subsequences Q, N is the maximum value of matrix A, i is the traversal from 1 to N, and ni is the number of choices at the i-th row position of the candidate common subsequence.
[0037] Preferably, in step S300, determining whether the candidate common subsequence is a reasonable subsequence includes:
[0038] Step S310: Take out a longest common subsequence Qm from Q. According to Qm, obtain the center point coordinates Cs corresponding to each character in the standard character information and the center point coordinates Cp corresponding to each character in the predicted character information. Calculate the ratio R[i] of the center point coordinate distances of the two adjacent characters in Qm in the standard character and predicted character sequences, specifically:
[0039]
[0040] where, C s [i] is the center point coordinate of the i-th standard character, C s [i - 1] is the center point coordinate of the (i - 1)-th standard character, C p [i] is the center point coordinate of the i-th predicted character, C p [i - 1] is the center point coordinate of the (i - 1)-th predicted character;
[0041] Step S320: If R[i] satisfies formula (8), then determine that Qm is a reasonable longest common subsequence, specifically:
[0042] count(|R[i] - R[i - 1]| < T v ) < T c (8)
[0043] where, Tv is a preset tolerance threshold parameter, and Tc is a preset tolerance number of times.
[0044] Preferably, step S400 includes:
[0045] Step S410: According to the longest common subsequence and the standard character information, infer the correct relative position information of each character in the predicted character information to obtain the inferred character position information, specifically:
[0046]
[0047]
[0048] Among them, It represents the character index to be inferred, In and Im represent the indices of two characters selected from the longest common subsequence, Cs is the center point coordinate of the character in the standard character information, Cp is the center point coordinate of the character in the predicted character information, Ct is the center point coordinate of the inferred character, Ss is the size of the character in the standard character information, Sp is the size of the character in the predicted character information, and St is the size of the character in the predicted character information;
[0049] When the length N of the longest common subsequence is greater than 2, I n , I m K groups (K≥2) can be selected according to the combination result, and the position C t and the size parameter S t , S t of the finally inferred character are taken as the mean value of the K-group inference values;
[0050] Step S420: For each inferred character Xt(It, Ct, St), traverse the predicted character information. If there is no character information in the predicted character information whose coincidence degree coefficient g iou with the inferred character position satisfies Equation (11), then fill the inferred character X(Ct, St) into the corresponding position in the predicted character information. Specifically:
[0051]
[0052] Among them, D t , D p are the rectangular frames of the inferred character and the predicted character respectively, and T iou is the threshold;
[0053] Step S430: For each character position Xp(Ip, Cp, Sp) in the predicted character information, traverse the set Xt of inferred character information. If no character information in Xt whose coincidence degree coefficient with the predicted character position Xp(Ip, Cp, Sp) satisfies Equation (11) can be found, then delete Xp(Ip, Cp, Sp) from the predicted character information.
[0054] Preferably, step S500 includes:
[0055] S510: For the characters filled in step S420, extract the image region of the filled characters, call the character recognition algorithm to recognize the image region of the filled characters, and at the same time, according to the wildcards corresponding to the positions of the filled characters, select the category label with the highest probability under the corresponding wildcards from the character category labels in step S210 as the category of the filled characters;
[0056] S520: For the characters in the predicted characters other than those corresponding to the longest common subsequence, according to the wildcards corresponding to the padding character positions, select the category label with the highest probability under the corresponding wildcard from the character category labels in step S210 as the category of the padding character;
[0057] S530: Combine the information in S510, S520, and the characters of the longest common subsequence, and output it as the final recognition result.
[0058] The above method for automatically correcting the character recognition result is applicable to post-processing the OCR detection and recognition result, and can effectively solve the problems of character misdetection, character missing detection, and incorrect results caused by character misrecognition. Description of the Drawings
[0059] Figure 1 It is a flowchart of a method for automatically correcting the character recognition result provided by an embodiment of the present invention;
[0060] Figure 2 It is an example diagram of the dynamic programming matrix A in step S220 of the present invention;
[0061] Figure 3 It is a schematic diagram of the initial value of the longest common subsequence in step S220 of the present invention;
[0062] Figure 4 It is a schematic diagram of the result of merging the longest common subsequence by position in step S220 of the present invention;
[0063] Figure 5 It is a schematic diagram of the result of the longest common subsequence after deleting unreasonable matches according to step S230 in step S200 of the present invention. Detailed Embodiment
[0064] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0065] In one embodiment, as Figure 1 shown, a method for automatically correcting the character recognition result, the method includes the following steps:
[0066] Step S100: Obtain the predicted character information in the image to be processed, and the standard character information in the same scene.
[0067] Specifically, character positioning, detection, and recognition algorithms are used to obtain the predicted character information in the image to be processed. Positioning involves locating the target character using traditional template matching. Detection mainly refers to using a deep learning object detection network to locate the target. Common deep learning object detection networks include the YOLO series and the centernet network. The recognition algorithm mainly identifies the detected characters into specific categories. Positioning and detection obtain that there is a character in a certain place, and recognition obtains the information about what category of character it is. The predicted character information includes the position information and category information of each predicted character. The position information includes the center point coordinates of the character and the width and height information of the character. The category information is the probability that each character belongs to each category. Among them, the categories include 10 digits such as 0, 1, 2, etc., 26 English letters such as a, b, c, etc. (case-sensitive), and symbols such as / .
[0068] The same scene means that the entities where the characters are printed are similar, the printed character content is the same, and the size and relative position of the printed characters are approximate. For example, in a food packaging production line, the production date, shelf life, and other characters printed on the same product during a certain period. Further, the method for obtaining the standard character information can use the method of step S100 to obtain the predicted character information for the image to be processed in the same scene above, and then manually review for errors. After manually correcting the errors, save the position and category of the characters as the standard character information.
[0069] The following uses a specific input example to illustrate the effectiveness of the present invention and elaborate on the actual implementation process.
[0070] Suppose there is the following standard character information:
[0071] {10,89,23,112,"J"}
[0072] {26,89,45,111,"N"}
[0073] {75,89,89,111,"4"}
[0074] {92,89,103,109,"8"}
[0075] {105,89,127,112,"3"}
[0076] {129,89,144,112,"7"}
[0077] {146,89,156,109," / "}
[0078] {158,89,171,109,"2"}
[0079] {173,89,184,109,"8"}
[0080] {186,89,196,109," / "}
[0081] {199,89,210,109,"8"}
[0082] {215,89,227,110,"0"}
[0083] {228,96,245,118,"g"}
[0084] {248,96,261,119,"s"}
[0085] {265,96,278,116,"p"}
[0086] {280,96,298,119,"e"}
[0087] And the following predicted character information:
[0088] {11,89,22,112,"J"}
[0089] {106,89,126,112,"3"}
[0090] {130,89,143,112,"7"}
[0091] {147,89,155,109," / "}
[0092] {159,89,170,109,"2"}
[0093] {174,89,183,109,"9"}
[0094] {187,89,195,109,"1"}
[0095] {200,89,209,109,"8"}
[0096] {216,89,226,110,"0"}
[0097] {249,96,260,119,"s"}
[0098] {281,96,297,119,"e"}
[0099] {300,96,320,119,"6"}
[0100] Each of the above pairs of curly brackets contains four values. The first four numbers represent the coordinates of the position rectangle of the character (the upper left and lower right coordinates, in the format of x1, y1, x2, y2. Hereinafter referred to as the detection box for easy understanding). The fifth character represents the category of the character (here, the category corresponding to the maximum recognition result probability). Hereinafter, the standard string R is used s refers to "JN4837 / 28 / 80gspe", which is composed of extracting the character categories from the standard character information. The prediction string R is used p refers to "J37 / 29180se6", which is composed of extracting the characters from the predicted character information
[0101] For the above predicted character information, it includes misdetection (the last prediction result {300, 96, 320, 119, "6"}), misrecognition (the sixth prediction result {174, 89, 183, 109, "9"}, the seventh prediction result {187, 89, 195, 109, "1"}), missed detection (the second, third, fourth, etc. characters of the standard characters), and character box change (the detection result box is smaller than the standard character information box).
[0102] Among them, misdetection means that there are characters in the predicted characters that do not exist in the standard characters. The main manifestation is that the detection box of the predicted character does not intersect with any detection box of the standard character or the overlapping part accounts for a very small proportion. The result of drawing the detection box on the original image is that there are two rectangular boxes, but there is no overlapping part or the overlapping part is very small between the two rectangular boxes. The last predicted character here does not intersect with any character of the standard character, so it is misdetection. Misrecognition means that for the characters at the same position, the predicted character category is inconsistent with the standard character category. The ninth character of the standard character information (the predicted result is character 8) and the sixth character of the predicted character information (the predicted result is character 9). The detection boxes of these two characters basically overlap, which can be seen from the coordinate position values, but the recognition results of the characters are incorrect. Missed detection needs to be judged by the coordinates of the characters. The coordinate difference between the first and second characters of the predicted character information is very large, but the first and second characters of the predicted character can both find corresponding ones in the standard characters. Here, the first predicted character corresponds to the first standard character, and the second predicted character corresponds to the fifth standard character. Then, the second, third, and fourth characters in the standard characters are completely missing in the predicted characters. This is missed detection. The difference from misdetection is that misdetection has a box with a similar position but very little intersection, while missed detection is that the two ends can be well matched and there is no detection box character in the middle. Character box change means that for the predicted character box, the width and height of the detection box are obtained by using the coordinates x2 - x1, y2 - y1. There is a certain difference between this width and height and the width and height of the standard character detection box. In this way, it is said that the predicted character box is too small, too large, or has an offset.
[0103] Step S200: Find a candidate common subsequence in the predicted character information and the standard character information.
[0104] In one embodiment, step S200 includes:
[0105] Step S210: Combine the character category labels of the standard character information and the predicted character information into strings Rs and Rp in the order of first rows and then columns;
[0106] Step S220: Calculate the length N of the longest common subsequence and the candidate longest common subsequence matrix P of the two strings Rs and Rp, specifically:
[0107] Assume the lengths of the strings Rs and Rp are Ls and Lp respectively. Construct a dynamic programming matrix A with Lp + 1 rows and Ls + 1 columns. Initialize the first row and the first column of A to 0, and calculate the values of the remaining positions according to formula (1):
[0108]
[0109] where A[i, j] is the value of the i-th row and the j-th column in the matrix A, A[i - 1, j] is the value of the (i - 1)-th row and the j-th column in the matrix A, A[i, j - 1] is the value of the i-th row and the (j - 1)-th column in the matrix A, S s [i] is the i-th standard character information, S p [j] is the j-th predicted character information, L p +1 represents the total number of rows of the matrix A, L s +1 represents the total number of columns of the matrix A;
[0110] Calculate the value of N according to formula (2), specifically:
[0111] N = max(A) (2)
[0112] where N is the maximum value of the matrix A;
[0113] If N < 2, it means that no candidate common subsequence is successfully found, and step S200 ends;
[0114] If N ≥ 2, traverse the matrix A. For the element A[i, j] in the i-th row and the j-th column of the matrix A, if it satisfies formula (3), add a new row to P, and at the same time store the values of A[i, j], i, j, and Ss[i] into the newly added row in P;
[0115]
[0116] Specifically, the character category label in step S210 is the category label with the highest probability. The order of first rows and then columns means that when processing multiple rows of predicted characters, first form all the detected characters into rows, and then process these individual rows with an algorithm. Attached Figure 2Shows matrix A calculated from the example input according to Equation (1), attached Figure 3 Shows matrix P calculated from the example input according to Equation (3), attached Figure 4 Shows the result of merging the longest common subsequence by position.
[0117] Step S230: Delete the unreasonable position matches in P, specifically:
[0118] Traverse each row in P. For the four elements P[i,0], P[i,1], P[i,2], P[i,3] in the i-th row of P, if they satisfy Equation (4), then delete this row from P;
[0119] L p -P[i,1]<N - P[i,0] or L s -P[i,2]<N - P[i,0] (4)
[0120] Where, P[i,1] is the element in the 2nd column of the i-th row in matrix P, L p is the length of Rp, L s is the length of Rs.
[0121] Specifically, attached Figure 5 Shows matrix P after deleting unreasonable position matches according to Equation (4) for the example input.
[0122] Step S240: Traverse the first column of the candidate longest common subsequence matrix P, count the number of matching positions for each position of the longest common subsequence, and construct all candidate common subsequences Q, including:
[0123] The number of matches n for the k-th position of the longest common subsequence with length N (N≥2) k is the number of times k appears in the first column of P, calculated according to Equation (5), specifically:
[0124] n k =count(P[:,0]=k) (5)
[0125] Where, the count function represents statistical calculation on numerical data;
[0126] Then for all candidate common subsequences Q, there are a total of M combination methods, calculate the value of M according to Equation (6), specifically:
[0127]
[0128] Where, M is the number of all candidate common subsequences Q, N is the maximum value of matrix A, i is the traversal from 1 to N, and ni is the number of choices for the i-th row position of the candidate common subsequence.
[0129] Specifically, the value of ni refers to Appendix Figure Four , Appendix Figure Four where N is 9 (the largest value in the val column). At the three positions where val equals 2, 4, and 6, there are two choices each. At this time, n2, n4, and n6 are all 2, and the others such as n1, n3, etc. are all 1.
[0130] Step S300: Traverse all candidate common subsequences, and determine whether a candidate common subsequence is a reasonable subsequence. If so, end the traversal.
[0131] In one embodiment, determining whether a candidate common subsequence is a reasonable subsequence in step S300 includes:
[0132] Step S310: Take out a longest common subsequence Qm from Q. According to Qm, obtain the center point coordinates Cs of each character in the standard character information and the center point coordinates Cp of each character in the predicted character information. Calculate the center point coordinate distance ratio R[i] of the two characters before and after in Qm in the standard character and predicted character sequences, specifically:
[0133]
[0134] where, C s [i] is the center point coordinate of the i-th standard character, C s [i - 1] is the center point coordinate of the (i - 1)-th standard character, C p [i] is the center point coordinate of the i-th predicted character, C p [i - 1] is the center point coordinate of the (i - 1)-th predicted character.
[0135] Step S320: If R[i] satisfies equation (8), then determine that Qm is a reasonable longest common subsequence, specifically:
[0136] count(|R[i] - R[i - 1]| < T v ) < T c (8)
[0137] where, T v is a preset tolerance threshold parameter, and Tc is a preset tolerance number of times.
[0138] Specifically, equation (8) means that by traversing the vector R, if the absolute value of the difference between two consecutive values is less than Tv and the number of occurrences is less than Tc, it can be considered a reasonable sequence.
[0139] Step S400: Filter and fill the predicted character information using a reasonable common subsequence to obtain the processed predicted character information.
[0140] Specifically, a reasonable longest common subsequence indicates that a partial correspondence between the predicted characters and the standard characters has been found. Based on this correspondence, the predicted character information can be corrected.
[0141] In one embodiment, step S400 includes:
[0142] Step S410: According to the longest common subsequence and the standard character information, infer the correct relative position information of each character in the predicted character information to obtain the inferred character position information. Specifically:
[0143]
[0144]
[0145] Where It represents the index of the character to be inferred, In, Im represent the indices of two characters selected from the longest common subsequence, Cs is the center point coordinate of the character in the standard character information, Cp is the center point coordinate of the character in the predicted character information, Ct is the center point coordinate of the inferred character, Ss is the size of the character in the standard character information, Sp is the size of the character in the predicted character information, and St is the size of the character in the predicted character information;
[0146] When the length N of the longest common subsequence > 2, I n , I m K groups (K ≥ 2) can be selected according to the combination result, and finally the position C of the inferred character t and the size parameter are S t , S t Take the mean of the K groups of inferred values.
[0147] Specifically, calculate the inferred character position information (character center point and width / height) according to equations (9) and (10), where equation (9) needs to calculate the values in the x and y directions, and equation (10) needs to calculate the two dimensions of w and h.
[0148] Step S420: For each inferred character Xt(It, Ct, St), traverse the predicted character information. If there is no character information in the predicted character information whose coincidence degree coefficient giou with the inferred character position satisfies equation (11), then fill the inferred character X(Ct, St) into the corresponding position in the predicted character information. Specifically:
[0149]
[0150] Where D t , Dp Rectangular boxes for the inferred character and the predicted character respectively, T iou is the threshold value;
[0151] Specifically, all the inferred character information constitutes the set X of inferred character information t .
[0152] Step S430: For each character position Xp(Ip, Cp, Sp) in the predicted character information, traverse the set Xt of inferred character information. If no character information in Xt has a coincidence coefficient with the predicted character position Xp(Ip, Cp, Sp) that satisfies Equation (11), then delete Xp(Ip, Cp, Sp) from the predicted character information.
[0153] Step S500: According to the processed predicted character information, correct the predicted character information and output the correct character information.
[0154] Specifically, according to the method in Step S400, the predicted character information can be calibrated to filter out misdetected characters and fill in undetected characters. At this time, the positions of each character in the predicted character information are already completely correct, but the class labels of the characters may be incorrect. There are mainly two reasons for this: First, when the relative positions of the characters are uncertain, in Step S1 during character recognition, only comparison can be made in the global character class library. However, due to the similarity of some characters (such as the number 0 and the capital letter O, the number 8 and the capital letter B, etc.), misrecognition may occur. Second, for the character positions filled in Step S400, the recognition step needs to be performed again to confirm whether there is actually such a character at that position.
[0155] In one embodiment, Step S500 includes:
[0156] S510: For the characters filled in Step S420, extract the image region of the filled characters, call the character recognition algorithm to recognize the image region of the filled characters, and at the same time, according to the wildcards corresponding to the filled character positions, select the class label with the highest probability under the corresponding wildcards from the character class labels in Step S210 as the class of the filled characters;
[0157] S520: For the other characters in the predicted characters except those corresponding to the longest common subsequence, according to the wildcards corresponding to the filled character positions, select the class label with the highest probability under the corresponding wildcards from the character class labels in Step S210 as the class of the filled characters;
[0158] S530: Combine the character information of S510, S520 and the longest common subsequence and output it as the final recognition result.
[0159] Specifically, the wildcards are numbers, letters, Chinese characters, special characters, etc.
[0160] The primary purpose of the algorithm for screening and filling the character detection and recognition results based on string matching is to solve the problem that during the production process, due to the algorithm itself or external environmental changes, there are significant fluctuations in the detection and recognition results of normal products, leading to the situation where normal samples are regarded as abnormal samples and rejected. This method is applicable to post-processing the OCR detection and recognition results and can effectively solve the problems of character misdetection, character missed detection, and incorrect results caused by character misrecognition. This algorithm has the characteristics of a wide application range and high accuracy.
[0161] The above has introduced in detail a method for automatically correcting the character recognition results provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A method for automatically correcting character recognition results, characterized in that, The method includes the following steps: Step S100: Obtain the predicted character information in the image to be processed and the standard character information in the same scene; Step S200: Search for candidate common subsequences in the predicted character information and the standard character information; Step S300: Traverse the candidate common subsequences, and determine whether the candidate common subsequence is a reasonable subsequence. If so, end the traversal; Step S400: Use the reasonable common subsequence to filter and fill the predicted character information to obtain the processed predicted character information; Step S500: Correct the predicted character information according to the processed predicted character information, and output the correct character information; In step S100, the predicted character information includes the position information and category information of each predicted character. The position information includes the center point coordinates and the width and height information of the character, and the category information is the probability that each character belongs to each category; In step S300, determining whether the candidate common subsequence is a reasonable subsequence includes: Step S310: Take out the longest common subsequence Qm from all candidate common subsequences Q. According to Qm, obtain the center point coordinates Cs corresponding to each character in the standard character information and the center point coordinates Cp corresponding to each character in the predicted character information. Calculate the ratio R[i] of the center point coordinate distances of the two characters before and after in the standard character and the predicted character sequences according to the center point coordinates Cs corresponding to each character in the standard character information and the center point coordinates Cp corresponding to each character in the predicted character information, specifically: Among them, C s [i] is the center point coordinate of the i-th standard character, C s [i - 1] is the center point coordinate of the (i - 1)-th standard character, C p [i] is the center point coordinate of the i-th predicted character, C p [i - 1] is the center point coordinate of the (i - 1)-th predicted character; Step S320: If the R[i] satisfies formula (8), then determine that Qm is a reasonable longest common subsequence, specifically: count(|R[i] - R[i - 1]| < T v ) < T c (8) where, Tv is a preset tolerance threshold parameter, and Tc is a preset tolerance number of times.
2. The method according to claim 1, characterized in that, Step S200 includes: Step S210: Form strings Rs and Rp by arranging the character category labels of the standard character information and the predicted character information in the order of first row and then column; Step S220: Calculate the length N of the longest common subsequence and the candidate longest common subsequence matrix P of the two strings Rs and Rp, specifically: Suppose the lengths of the strings Rs and Rp are Ls and Lp respectively. Construct a dynamic programming matrix A with Lp + 1 rows and Ls + 1 columns. The first row and the first column of the A are initialized to 0, and calculate the values of the remaining positions according to formula (1): Among them, A[i,j] is the value of the element in the i-th row and j-th column of matrix A, A[i-1,j] is the value of the element in the (i-1)-th row and j-th column of matrix A, A[i,j-1] is the value of the element in the i-th row and (j-1)-th column of matrix A, S s [i] is the i-th standard character information, S p [j] is the j-th predicted character information, L p +1 represents the total number of rows of matrix A, L s +1 represents the total number of columns of matrix A; Calculate the value of N according to formula (2), specifically: N = max(A)(2) where, N is the maximum value of the matrix A; If N < 2, it means that no candidate common subsequence is successfully found, and step S200 ends; If N ≥ 2, traverse the matrix A. For the element A[i, j] in the i-th row and j-th column of the matrix A, if it satisfies formula (3), add a new row to the P, and at the same time store the values of A[i, j], i, j, and Ss[i] in the newly added row of the P; Step S230: Delete the unreasonable position matches in the P, specifically: Traverse each row in P. For the four elements P[i,0], P[i,1], P[i,2], P[i,3] in the i-th row of P, if they satisfy Equation (4), then delete this row from P; L p -P[i,1] < N - P[i,0] or L s -P[i,2] < N - P[i,0] (4) where P[i, 1] is the element in the second column of the i-th row in matrix P, L p is the length of Rp, and L s is the length of Rs; Step S240: Traverse the first column of the candidate longest common subsequence matrix P, count the number of matching positions at each position of the longest common subsequence, and construct all candidate common subsequences Q, including: The number of possible matches n at the k-th position of the longest common subsequence of length N (N≥2) k is the number of times k appears in the first column of P, calculated according to Equation (5), specifically: n k = count(P[:,0] = k) (5) where the count function represents statistical calculation on numerical data; Then for all possible candidate common subsequences Q, there are a total of M combination ways, and the value of M is calculated according to Equation (6), specifically: Among them, M is the number of all candidate common subsequences Q, N is the maximum value of matrix A, i is the traversal from 1 to N, and n i is the number of selections for the i-th row position of the candidate common subsequence.
3. The method according to claim 2, wherein Step S400 includes: Step S410: According to the longest common subsequence and the standard character information, infer the correct relative position information of each character in the predicted character information to obtain the inferred character position information, specifically: where It represents the index of the character to be inferred, In, Im represent the indices of two characters selected from the longest common subsequence, Cs is the center point coordinate of the character in the standard character information, Cp is the center point coordinate of the character in the predicted character information, Ct is the center point coordinate of the inferred character, Ss is the size of the character in the standard character information, Sp is the size of the character in the predicted character information, and St is the size of the character in the predicted character information; When the length N of the longest common subsequence is greater than 2, I n , I m K groups (K ≥ 2) can be selected according to the combination result, and finally the position C of the character is inferred t and the size parameter S t , S t Take the mean of the inference values of the K groups; Step S420: For each inferred character Xt(It, Ct, St), traverse the predicted character information. If there is no character information in the predicted character information whose coincidence degree coefficient g with the position of the inferred character iou satisfies the formula (11), then fill the inferred character X(Ct, St) into the corresponding position in the predicted character information, specifically: Among them, D t , D p are respectively the rectangular frames of the inferred character and the predicted character, and T iou is the threshold value; Step S430: For each character position Xp(Ip, Cp, Sp) in the predicted character information, traverse the set of inferred character information Xt. If no character information in Xt has a coincidence coefficient with the predicted character position Xp(Ip, Cp, Sp) that satisfies Equation (11), then delete Xp(Ip, Cp, Sp) from the predicted character information.
4. The method according to claim 3, wherein Step S500 includes: S510: For the characters filled in Step S420, extract the image region of the filled characters, call the character recognition algorithm to recognize the image region of the filled characters, and at the same time, according to the wildcards corresponding to the filled character positions, select the category label with the highest probability under the corresponding wildcards from the character category labels in Step S210 as the category of the filled characters; S520: For the characters in the predicted characters other than the characters corresponding to the longest common subsequence, according to the wildcards corresponding to the filled character positions, select the category label with the highest probability under the corresponding wildcards from the character category labels in Step S210 as the category of the filled characters; S530: Combine the information of S510, S520 and the characters in the longest common subsequence, and output it as the final recognition result.
Citation Information
Patent Citations
Information recognition method and device and terminal equipment
CN108345581A