OCR (Optical Character Recognition) curved document correction performance detection method based on multi-modal distance collaborative optimization
Through the multimodal distance collaborative optimization method, the weighted cost matrix is constructed and combined with the Hungarian algorithm, the problem of single evaluation standards after OCR distortion correction is solved, and the accuracy of the OCR recognition effect is achieved and the robustness of the OCR recognition effect is improved.
Patent Information
- Application Number
- CN202510489281.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing OCR technology lacks effective evaluation standards after distortion correction, resulting in a single evaluation, insufficient matching accuracy, poor dynamic adaptability, and sensitive to abnormal data, and the inability to accurately measure the OCR recognition effect.
The multimodal distance collaborative optimization method is adopted to construct a weighted cost matrix, combine geometric distance and text similarity, and use a Hungarian algorithm to achieve optimal matching, and support dynamic parameter adjustment and abnormal data processing to perform performance detection of OCR bending documents.
It realizes accurate quantitative evaluation of OCR recognition effect, improves matching accuracy and robustness, adapts to the needs of different scenarios, significantly improves the matching rate in medical document scenarios and reduces recognition errors.
Smart Images

Figure CN120356228A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and image processing, and particularly to an OCR curved document correction performance detection method based on multi-modal distance collaborative optimization. Especially for images processed by distortion correction (Dewarp), through the comprehensive evaluation of geometric distance and text similarity, it realizes the accurate quantitative analysis of the effect of the OCR distortion correction algorithm. Background Art
[0002] Currently, OCR technology faces certain challenges when processing images after distortion correction (such as distortion correction). Especially after the image is processed by distortion correction, the image size and position may change, resulting in the non-linearity of the corresponding relationship between the OCR recognition result and the original image. Due to the lack of effective evaluation criteria, the existing technology cannot accurately measure the following problems still existing in use after distortion correction: 1. Single evaluation criterion: Traditional methods only evaluate the distortion correction effect through pixel comparison or simple error analysis, ignoring the relevance of text semantics and spatial relationships.
[0003] 2. Insufficient matching accuracy: It does not effectively combine the comprehensive matching of the geometric position of the text box (bbox) (such as the Euclidean distance of the center point, intersection over union) and the text content (such as edit distance).
[0004] 3. Poor dynamic adaptability: The weight coefficient and evaluation threshold are fixed and cannot adapt to the differentiated requirements of different scenarios (such as medical documents, bill recognition).
[0005] 4. Sensitive to abnormal data: Missing text boxes or text errors in OCR recognition will lead to deviation of the matching result, and there is a lack of a robust repair mechanism.
[0006] Therefore, there is an urgent need for an evaluation method that can integrate spatial and text information and support dynamic adjustment. Summary of the Invention
[0007] 1. Technical problems to be solved: Aiming at the problems existing in the prior art, the purpose of the present invention is to provide an OCR curved document correction performance detection method based on multi-modal distance collaborative optimization. By fusing the weighted cost matrix of geometric distance and text similarity, combined with the Hungarian algorithm to achieve the optimal matching, and supporting dynamic parameter adjustment and abnormal data processing, it accurately quantifies the improvement or degradation of the OCR recognition effect by the distortion correction algorithm.
[0008] 2. Technical solutions: To solve the above problems, the present invention adopts the following technical solutions.
[0009] A method for detecting the performance of OCR curved document correction based on multi-modal distance collaborative optimization, comprising the following steps: S1. Construct a cost matrix: For the image after distortion correction, detect the set B of OCR text boxes in the original image and the set D of OCR text boxes in the image after distortion correction, and calculate each original text box b i and the text box d j after distortion correction. The multi-modal matching cost C ij between them is composed of the weighted geometric distance and text similarity distance, and the formula is: , where α and β are weight coefficients, d geom represents the geometric distance between text boxes, and d text represents the edit distance of text content; S2. Optimal matching calculation: Based on the cost matrix, apply the Hungarian algorithm to solve the global optimal matching relationship between the set B of original text boxes and the set D of text boxes after distortion correction; S3. Performance evaluation: According to the optimal matching result, calculate the matching rate, average geometric error, and average text edit distance, and quantitatively evaluate the performance of the distortion correction algorithm in combination with a preset threshold.
[0010] Further improvement lies in that the geometric distance d geom is calculated in the following way: , where α and β are adjustment parameters, d center is the Euclidean distance between the center points of text boxes, and IoU is the intersection over union.
[0011] Further improvement lies in that the text similarity distance d text is calculated by the normalized edit distance, and the formula is: , where EditDistance represents the minimum number of edit operations (insertion, deletion, replacement) of two strings, and after normalization, the range is ensured to be [0,1].
[0012] Further improvement lies in that the weight coefficients α and β are dynamically adjusted according to the actual application scenario, and are specifically implemented through the following steps: (a) Train a regression model based on historical matching data to predict the optimal weight combination; (b) Manually adjust the weight coefficients according to the evaluation requirements input by the user.
[0013] Further improvement lies in that the calculation method of the matching rate in the performance evaluation is: , wherein, X ij is an element of the matching matrix output by the Hungarian algorithm, and m and n are the numbers of text boxes in the original image and the image after distortion correction, respectively.
[0014] A further improvement is that the method further includes a visualization step of displaying the matching relationship between the OCR text boxes of the original image and the image after distortion correction in the form of a heat map or a connection diagram.
[0015] A further improvement is that the preset threshold is set in the following manner: (a) Preset an initial threshold according to the target scenario of the distortion correction algorithm (such as medical document, bill recognition); (b) Dynamically adjust the threshold based on the statistical distribution of the real-time matching results to optimize the evaluation sensitivity.
[0016] A further improvement is that the method further includes an abnormal data processing step: (a) Detect missing text boxes or text errors in the OCR recognition results; (b) After repairing the abnormal data through interpolation or semantic completion algorithms, re-perform the matching calculation.
[0017] 3. Advantageous effects: Adopting the technical solution provided by the present invention, compared with the prior art, it has the following advantageous effects: (1) By comprehensively considering the spatial error and the text error, the present invention overcomes the limitation of the traditional evaluation method that only focuses on a single factor, thereby realizing more accurate performance detection; (2) This method can effectively quantify the improvement or degradation of the distortion correction model in the OCR task, provides an objective performance evaluation criterion, and helps to optimize and adjust the model; (3) The application of the Hungarian algorithm ensures the accuracy of the best match, thereby improving the robustness and accuracy of the detection; (4) By adjusting the α and β coefficients, the evaluation results can flexibly adapt to different application requirements, making this method have strong versatility and scalability; (5) Through experimental verification, in the medical document scenario, the present invention has a 19% increase in the matching rate (from 72% to 91%), a 55.7% reduction in the average geometric error (from 18.5 pixels to 8.2 pixels), and a 57.1% reduction in the average text edit distance (from 0.35 to 0.15), which is significantly better than the traditional single-modal evaluation method.
[0018] It should be noted that the structures not introduced in the present invention are the same as the prior art or can be implemented by the prior art because they do not involve the design key points and improvement directions of the present invention, and will not be elaborated here. Brief Description of the Drawings
[0019] Figure 1 is the original image input to the present invention; Figure 2 is to Figure 1 the effect diagram of the text box after OCR recognition; Figure 3 is to Figure 1 the image after bending correction; Figure 4 is to Figure 3 the effect diagram of the text box after OCR recognition; Figure 5 is the effect diagram of the matching after OCR recognition of the above images respectively and multi-modal detection by the present invention; Figure 6 is the overall flow schematic diagram of the present invention. Detailed Description of the Invention
[0020] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0021] Please refer to Figures 1-6 , the present invention proposes an OCR distortion removal performance detection method based on multi-modal distance collaborative optimization, and the specific steps are as follows: 1. Construct a cost matrix: For the image after distortion correction, by detecting each text box (bbox) in the image and its corresponding OCR recognition result, calculate the Euclidean distance and edit distance of the center point of each bbox.
[0022] 2. Cost calculation: Each element in the cost matrix is determined by two factors: one is the Euclidean distance of the text box position, and the other is the edit distance of the text in the text box. By setting two weight coefficients (α and β), calculate the weighted cost, that is: Cost = α * Euclidean distance + β * edit distance.
[0023] 3. Optimal matching: Use the Hungarian algorithm to calculate the optimal text box matching method through the cost matrix to ensure the optimization of the matching between the text boxes of the distortion correction image and the text boxes in the original image.
[0024] 4. Threshold judgment: Set a threshold for judging the matching quality, and the threshold can be adjusted according to the application scenario, so as to quantitatively evaluate the performance of the distortion correction algorithm.
[0025] Specifically, it is as follows: 1. Representation of OCR detection frames The input image is as Figure 1 shown, which is the original image, and its OCR detection result set is: , where each OCR text box b i is represented by its upper-left and lower-right coordinates and the detected text: , where: (x min , y min ) is the upper-left coordinate of the text box; (x max , y max ) is the lower-right coordinate of the text box; t i is the text content recognized by OCR.
[0026] The result of drawing the coordinate box on the original image is as Figure 2 shown.
[0027] Similarly, the image after distortion correction is as Figure 3 shown, and its OCR detection result set is: , Each text box d after distortion correction j is also represented by its coordinates and text content: , where: (x' min , y' min ) is the upper-left coordinate of the text box after distortion correction; (x' max , y' max ) is the lower-right coordinate of the text box after distortion correction; t' i is the OCR text content after distortion correction.
[0028] Among them, m and n are the numbers of text boxes before and after distortion correction, respectively. Usually, m ≠ n.
[0029] The result of drawing the coordinate box on the original image is as Figure 4 shown.
[0030] 2. Construction of cost matrix We need to construct an m*n-dimensional cost matrix C, where each element C ij represents the original OCR text box b iThe matching cost between the text box d after distortion correction j is as follows. The cost can be determined by multiple factors, and we define it as follows: , where d geom (b i , d j ) represents the geometric distance, measuring the changes in the position and shape of the text box; d text (b i , d j ) represents the text similarity cost, measuring the text differences recognized by OCR; α and β are hyperparameters used to balance the weights of the geometric distance and text similarity.
[0031] 2.1. Geometric distance calculation The geometric distance d geom (b i , d j ) can be calculated by combining the Euclidean distance of the center points and IoU (Intersection over Union): , where: Euclidean distance of the center points: , where (x c , y c ) and (x c ', y c ') are the center points of b i and d j respectively: , IoU calculation:
[0032] where α and β are adjustment parameters used to balance the influence of the two.
[0033] 2.2. Text edit distance The EditDistance screenwriting distance, abbreviated as D, is calculated as follows: , The text similarity cost d text (b i , d j ) is calculated using the above edit distance: , Among them, EditDistance represents the minimum number of edit operations (insertion, deletion, replacement) between two strings. After normalization, it is ensured that the range is [0,1].
[0034] 3. Optimal matching calculation Given the cost matrix C, we use the Hungarian algorithm to solve the optimal matching and define the optimal allocation matrix X: , Where: X ij ∈{0,1}. If b i and d j are matched, then X ij =1, otherwise it is 0.
[0035] Satisfy the unique matching constraint: , The Hungarian algorithm solves this minimum-cost bipartite matching problem, ensuring a globally optimal solution.
[0036] 4. Result analysis After the Hungarian algorithm is solved, the optimal matching relationship (b i , d j ) is obtained, and the following performance metrics can be calculated: a. Matching rate: ,
[0037] b. Average edit distance: ,
[0038] C. Average geometric error: , Through the above calculations, the OCR distortion correction effect can be quantitatively evaluated and can be used to optimize the distortion correction model.
[0039] Visualize the matching of the OCR result of the original image and the OCR result of the image after dewarp distortion correction as Figure 5 shown.
[0040] The overall process involved in the present invention is as Figure 6 shown.
[0041] Example 1: Application of medical document correction This example takes a medical record document as an example to illustrate the specific application of this method in the OCR performance detection after bending correction.
[0042] Parameter settings: Weight coefficients: α = 0.6 (geometric distance weight), β = 0.4 (text similarity weight); Hyperparameters of the Hungarian algorithm: The maximum number of iterations is 1000, and the convergence threshold is 1e-5; The preset matching rate threshold is 85%, the average geometric error threshold is 10 pixels, and the average text edit distance threshold is 0.2.
[0043] Experimental comparison: Test 100 curved medical documents and compare the effects of the traditional method (only matching based on geometric distance) and the multi-modal matching method of the present invention:
[0044] Effect description: By fusing geometric and text features, this method significantly improves the matching accuracy in the medical document scenario, while reducing the OCR recognition error, verifying the effectiveness of multi-modal collaborative optimization.
[0045] Among them, the Hungarian Algorithm is a polynomial-time algorithm for solving the minimum weight matching problem of bipartite graphs, and its time complexity is O(n³), where n is the number of text boxes to be matched. The reasons for the present invention to select this algorithm include: Global optimality: It can ensure the global optimal matching of the text boxes of the original image and the corrected image, avoiding local optimal solutions; Applicability: It is applicable to sparse or dense matching scenarios and is robust to noisy data; Efficiency: In practical applications (n ≤ 1000), the running time of the algorithm can be controlled within milliseconds, meeting the real-time requirements.
[0046] The above-described embodiments only represent certain embodiments of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A performance detection method for OCR curved document correction based on multi-modal distance collaborative optimization, characterized in that Including the following steps: S1. Construct a cost matrix: For the image after distortion correction, detect the set B of OCR text boxes in the original image and the set D of OCR text boxes in the image after distortion correction, and calculate the multimodal matching cost C i between each original text box b j and the text box d after distortion correction ij . The multimodal matching cost is composed of the weighted geometric distance and text similarity distance, and the formula is: , where α and β are weight coefficients, and α + β = 1, d geom represents the geometric distance between text boxes, d text represents the edit distance of text content; S2. Optimal matching calculation: Based on the cost matrix, apply the Hungarian Algorithm to solve the global optimal matching relationship between the set B of original text boxes and the set D of text boxes after distortion correction; S3. Performance evaluation: According to the optimal matching result, calculate the matching rate, average geometric error, and average text edit distance, and quantitatively evaluate the performance of the distortion correction algorithm in combination with a preset threshold.
2. The OCR curved document correction performance detection method based on multi-modal distance collaborative optimization according to claim 1, wherein, The geometric distance d geom is calculated as follows: , where α and β are adjustment parameters, d center is the Euclidean distance of the center point of the text box, and IoU is the intersection over union.
3. The OCR curved document correction performance detection method based on multi-modal distance collaborative optimization according to claim 1, wherein The text similarity distance d text is calculated by the normalized edit distance, and the formula is: , Among them, EditDistance represents the minimum number of editing operations (insertion, deletion, replacement) of two strings, and after normalization, it is ensured that the range is [0,1].
4. A performance detection method for OCR curved document correction based on multi-modal distance collaborative optimization according to claim 1, characterized in that, The weight coefficients α and β are dynamically adjusted according to the actual application scenario, and are specifically implemented through the following steps: (a) Train a regression model based on historical matching data to predict the optimal weight combination; (b) Manually adjust the weight coefficients according to the evaluation requirements input by the user.
5. A performance detection method for OCR curved document correction based on multi-modal distance collaborative optimization according to claim 1, characterized in that The calculation method of the matching rate in the performance evaluation is: , Among them, X ij is the element of the matching matrix output by the Hungarian algorithm, where m and n are the numbers of text boxes in the original image and the image after distortion correction, respectively.
6. A performance detection method for OCR curved document correction based on multi-modal distance collaborative optimization according to claim 1, characterized in that, The method further includes a visualization step of displaying the OCR text box matching relationship between the original image and the image after distortion correction in the form of a heat map or a connection diagram.
7. A performance detection method for OCR curved document correction based on multi-modal distance collaborative optimization according to claim 1, characterized in that The setting of the preset threshold is achieved through the following method: (a) Preset an initial threshold according to the target scenario of the distortion correction algorithm (such as medical document, bill recognition); (b) Dynamically adjust the threshold based on the statistical distribution of the real-time matching result to optimize the evaluation sensitivity.
8. A method for detecting the performance of OCR curved document correction based on multi-modal distance collaborative optimization according to claim 1, characterized in that, The abnormal data processing steps include: (a) Detect missing text boxes or text errors in the OCR recognition result; (b) Repair the coordinates of the missing text box through a linear interpolation algorithm and correct the text error based on a semantic completion model.
Citation Information
Patent Citations
Method for evaluating OCR (optical character recognition) quality and related product
CN115937883A
Map updating method and device, automatic driving control method and device, medium and vehicle
CN115962787A
Performance evaluation method, device and equipment of OCR (Optical Character Recognition) system and readable storage medium
CN116978032A
Evaluation method and device of optical character recognition model and electronic equipment
CN118537868A
Map information updating method and system
CN119293059A