Multi-mode invoice image processing method and device based on vision and deep learning, and medium

Through multimodal angle correction, dynamic morphological operation and multimodal verification, the problems of low preprocessing accuracy, table line breakage and insufficient OCR identification in invoice processing are solved, and efficient and accurate processing of Chinese invoices are achieved.

CN120236296AActive Publication Date: 2025-07-01JIANGSU VOCATION & TECHNICAL COLLEGE OF FINANCE & ECONOMICS

Patent Information

Application Number
CN202510370513.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The prior art has problems in invoice processing such as low preprocessing accuracy, limited repair effect of table line fracture, lack of multimodal verification mechanism, and insufficient recognition ability of OCR for Chinese invoices, which affects the accuracy and reliability of invoice processing.

Method used

The multimodal angle correction technology is used to combine the QR code shielding technology of the ZBar library to dynamically adjust the structural element size of morphological operations, special optimization is carried out based on the PaddleOCRv3 model, and a multimodal verification module is introduced to cross-verify the OCR results, and the identification accuracy is improved through the multi-modal voting mechanism.

Benefits of technology

It improves the preprocessing effect of invoice images, accurately repairs table line breaks, enhances OCR's ability to identify Chinese invoices, ensures the stability and accuracy of key field extraction, and improves the efficiency and accuracy of invoice processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236296A_ABST
    Figure CN120236296A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode invoice image processing method and device based on vision and deep learning, and a medium, and relates to the technical field of image digital processing. The method comprises the following steps: calculating a fusion rotation angle according to an angle of a table edge and an angle of a two-dimensional code positioning frame so as to rotate an invoice image to obtain a first image; inputting the first image into a table detection module, extracting a first table feature, and performing non-maximum suppression and noise filtering processing on the first table feature to obtain a second table feature; performing cell repair processing on the second table feature to obtain a third table feature; constructing an OCR module, and extracting text features from the third table features by using the OCR module; comparing the text features with the two-dimensional code analysis data to judge the consistency; and if yes, extracting key fields from the text features based on a regular expression and a semantic rule. According to the invention, the efficiency and accuracy of invoice processing can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital image processing, and more specifically, to a multi-modal invoice image processing method, device, and medium based on vision and deep learning. Background Art

[0002] There are some major deficiencies in the current invoice processing system, which limit the accuracy and reliability of invoice processing. The following are several key issues:

[0003] (1) Low preprocessing accuracy: In the prior art, the skew correction and noise removal of invoice images are often independent preprocessing steps, lacking the fusion of multi-modal information, resulting in inaccurate results. Only by using the Hough line transform to detect the table line angle, when the invoice edge is blurred or there is carbon copy interference, the calculation error is significant. In addition, the noise processing means are relatively single, and it is difficult to effectively handle invoice images with complex backgrounds, poor scanning quality, or insufficient shooting light, further reducing the reliability of subsequent processing.

[0004] (2) Limited effect in repairing broken table lines: The table lines in invoices may be broken or blurred due to scanning quality or shooting angle problems. The prior art mostly uses traditional morphological operations to repair table lines, but the size of its structural elements is usually fixed (3×3), and it cannot dynamically adapt to the widths and breakage degrees of table lines in different invoices, resulting in over-corrosion. This fixed strategy often has an unsatisfactory repair effect when dealing with severely broken or widely varying-width table lines, affecting the extraction of subsequent cell contents.

[0005] (3) Insufficient recognition ability of OCR for special Chinese characters and handwritten scripts in Chinese invoices: Existing OCR models (such as PaddleOCR, Tesseract, etc.) lack targeted optimization when processing Chinese invoices. Especially when recognizing Chinese printed scripts such as FangSong and KaiTi, and some handwritten scripts, the accuracy rate is relatively low. In addition, the occurrence frequency of special characters (such as "¥", "Tax ID:") in the model training data is relatively low, resulting in a significantly insufficient recognition rate of these key characters, thereby affecting the extraction of key invoice information.

[0006] (4) Lack of multi-modal verification mechanism: The prior art usually only relies on a single OCR data source and lacks a cross-verification mechanism with other data sources such as QR code parsing, resulting in a relatively low extraction accuracy of key fields such as invoice codes and amounts. In addition, traditional methods rely on manual processing of low-confidence OCR results, which is time-consuming. Summary of the Invention

[0007] To solve the above technical problems, the present invention provides a multi-modal invoice image processing method, apparatus and medium based on vision and deep learning. By innovatively introducing deep learning models, OCR technology and image processing technology into the field of invoice processing, this patent aims to provide an accurate, reliable and automated invoice processing method to address the following deficiencies of the existing technologies:

[0008] (1) Reducing the influence of invoice tilt, noise and QR code interference: By utilizing multi-modal angle correction technology and the QR code masking technology of the ZBar library, the present invention improves the influence of factors such as invoice tilt and noise. By combining the Hough line transform with QR code positioning, the horizontal correction of the image is made more reliable. At the same time, the ZBar library masks the QR code area to avoid the interference of the QR code on the rotation angle calculation and subsequent table detection.

[0009] (2) Dynamically repairing the problem of broken table lines: The present invention solves the problem of insufficient repair effect of broken table lines by dynamically adjusting the size of the structural element of morphological operations. According to the width and degree of breakage of the table lines, the sizes of the dilation and erosion structural elements are flexibly adjusted to ensure the accuracy of the repair effect and better restore the complete structure of the table.

[0010] (3) Enhancing the recognition ability of OCR for Chinese invoices: The present invention improves the recognition ability of OCR for Chinese invoices to a certain extent through special optimization based on the PaddleOCR v3 model. It fine-tunes for printed fonts such as FangSong and KaiTi and some handwritten fonts, artificially increases the sample ratio of special characters in the training data, and adjusts the output dimension of the model to make it more adaptable to the complex scenarios of Chinese invoices.

[0011] (4) Multi-modal verification to ensure the accuracy of OCR results: The present invention effectively solves the limitation of single dependence on OCR results by introducing a multi-modal verification module, through cross-verifying the key fields extracted by OCR with the QR code parsing data. In addition, for OCR recognition results with low confidence, a multi-model voting mechanism is introduced. At the same time, a re-scanning mechanism is also adopted to improve the stability and accuracy of key field extraction.

[0012] Therefore, the present invention aims at the deficiencies existing in the prior art in the field of invoice processing, and provides an innovative method to solve the problems such as low preprocessing accuracy, limited repair effect of broken table lines, lack of multi-modal verification mechanism, and insufficient recognition ability of OCR for special characters and handwritten fonts in Chinese invoices in the current Chinese technology invoice processing field. By overcoming the limitations of the prior art, this invention is expected to promote the further development and application of Chinese invoice processing technology.

[0013] In a first aspect, the present invention provides a multi-modal invoice image processing method based on vision and deep learning, and the method includes:

[0014] Obtain an invoice image, and based on the invoice image, calculate the angle of the table edge and the angle of the QR code positioning frame, generate a binary mask to set the pixels in the QR code area in the invoice image to 0 to shield interference, calculate a fusion rotation angle according to the angle of the table edge and the angle of the QR code positioning frame, and rotate the invoice image according to the fusion rotation angle to obtain a first image;

[0015] Construct a detection model, where the detection model is an improved YOLOv5 model obtained by adding two groups of detection frames to the YOLOv5 model, use a training data set to train the detection model to obtain a table detection module, input the first image into the table detection module, extract first table features, and perform non-maximum suppression and noise filtering on the first table features to obtain second table features;

[0016] Perform cell repair processing on the second table features to obtain third table features;

[0017] Construct an OCR module, and use the OCR module to extract text features from the third table features;

[0018] Compare the text features with the QR code parsing data to determine consistency. In the case of determining inconsistency, make an inconsistency mark in the corresponding invoice image;

[0019] In the case of determining consistency, extract key fields from the text features based on regular expressions and semantic rules.

[0020] Further, calculating the fusion rotation angle according to the angle of the table edge and the angle of the QR code positioning frame includes:

[0021] Calculate the average inclination angle of the table line segments in the invoice image whose length exceeds 30% of the image width as the angle of the table edge;

[0022] According to the four corner point coordinates of the QR code positioning frame, calculate the diagonal slope, and calculate the angle of the QR code positioning frame according to the diagonal slope;

[0023] According to the angle of the table edge and the angle of the QR code positioning frame, calculate the fusion rotation angle through the following formula:

[0024] θ = αθ QR +(1 - α)θ Hough

[0025] where θ Hough represents the angle of the table edge, θ QR represents the angle of the QR code positioning frame, and α represents a weight coefficient.

[0026] Further, a loss function is set, and the detection model is trained using the training data set to obtain a table detection module; wherein, the loss function is used to optimize the position, size, and aspect ratio of the detection box, and the loss function is expressed as:

[0027]

[0028] Wherein, represents the loss value, IoU represents the intersection over union of the predicted box and the ground truth box, ρ represents the Euclidean distance between the centers of the predicted box and the ground truth box, C represents the diagonal length of the minimum bounding rectangle, α1 represents the balance parameter, and v represents the aspect ratio consistency parameter.

[0029] Further, non-maximum suppression and noise filtering are performed on the first table feature to obtain a second table feature, including:

[0030] An IoU threshold is set to merge overlapping detection boxes. If the height of the detection box meets a preset condition, re-prediction is triggered; wherein, the preset condition is expressed as:

[0031] H box >0.65H invoice

[0032] Wherein, H box is the height of the detection box, and H invoice is the total height of the invoice image;

[0033] If the first table feature is a nested table, after non-maximum suppression processing, the detection box corresponding to the first table feature is marked as the detection box of the composite table and used as the parent table detection box. The area of the child table detection box is located in the parent table detection box, and segmentation is performed by calculating the edge intersection coordinates of the parent table detection box and the child table detection box, and the mapping of the child table detection box is mapped to the original invoice image coordinate system.

[0034] Further, cell repair processing is performed on the second table feature to obtain a third table feature, including:

[0035] Adaptive morphological operations, closing operations, opening operations, edge filling, and table line repair processing are sequentially performed on the second table feature to obtain a third table feature; wherein:

[0036] The size of the structural element for morphological operations is dynamically adjusted according to the invoice width, and the formula is as follows:

[0037] K size =max(3,[0.02W invoice )

[0038] Wherein, K sizeRepresents the size of the structural element, W invoice Represents the width of the invoice image;

[0039] When performing the closing operation, a square kernel is used to fill the cell gaps caused by stamps or stains; the opening operation is used to eliminate isolated noise points;

[0040] When filling in the edges, a hit-or-miss transform kernel is set. The break points are located by matching the brightness changes in the vertical direction through the hit-or-miss transform kernel, and then a set number of pixels are extended along the vertical direction and the linear interpolation algorithm is used to fill in the line segments;

[0041] In the stage of table line repair, the horizontal lines and vertical lines are respectively merged with adjacent line segments through directional dilation operations; through the thinning algorithm, the thick lines are transformed into skeleton lines with a single pixel width, and the gaps at the intersections of the lines are detected to determine whether the length of the gap is less than or equal to the preset number of pixels. If so, the B-spline curve is applied for smooth connection.

[0042] Furthermore, the OCR module fine-tunes the special Chinese invoice model based on the PaddleOCRv3 framework, improves the recognition accuracy through data augmentation and optimization of the special character library, and performs blank line detection based on the initial output features of the OCR module, eliminates the blank lines, and obtains the text features;

[0043] When performing blank line detection, the blank line determination formula is set as follows:

[0044]

[0045] where cell i represents the i-th cell image in the table row, max(cell i ) represents the maximum gray value of the cell image, min(cell i ) represents the minimum gray value of the cell image; T cn is a threshold designed for complex Chinese strokes;

[0046] When the initial output features satisfy the determination formula, if the peak interval ratio of the cell gray value histogram exceeds 90%, it is determined as a blank line and excluded.

[0047] Furthermore, the text features are compared with the two-dimensional code parsing data to judge the consistency, including:

[0048] Obtain the recognition results of the invoice code and amount fields from the text features output by the OCR module, and at the same time parse the two-dimensional code in the invoice image to obtain the encoded data of the same fields;

[0049] Compare the recognition results of the invoice code and amount fields with the encoded data of the same fields. If they are inconsistent, first use the QR code data and call the API to verify the legality of the invoice code. For the amount field, convert the Arabic numerals recognized by OCR into Chinese capital format through a rule engine, and perform semantic comparison with the capital amount extracted by OCR to determine if they are the same.

[0050] Further, in the case of being judged as consistent, extract key fields from the text features based on regular expressions and semantic rules, including:

[0051] The extraction of the invoice code uses the first regular expression to match 10 to 12 consecutive digits and exclude consecutive repeated characters. The check code matches several digits at the end of the invoice code through the second regular expression; the capital amount field extracts the capital amount text through the character set regular expression.

[0052] The semantic rule engine improves the field extraction priority through context positioning and performs logical verification at the same time to ensure that the last digit of the check code is consistent with the last digit of the invoice code.

[0053] In a second aspect, the present invention provides a multi-modal invoice image processing device based on vision and deep learning. The device includes:

[0054] An image preprocessing unit, configured to obtain an invoice image, and based on the invoice image, calculate the angle of the table edge and the angle of the QR code positioning frame, generate a binary mask to set the pixels in the QR code area of the invoice image to 0 to shield interference, calculate the fusion rotation angle according to the angle of the table edge and the angle of the QR code positioning frame, and rotate the invoice image according to the fusion rotation angle to obtain a first image.

[0055] A table detection unit, configured to build a detection model, where the detection model is an improved YOLOv5 model obtained by adding two groups of detection frames to the YOLOv5 model, train the detection model with a training data set to obtain a table detection module, input the first image into the table detection module, extract the first table feature, and perform non-maximum suppression and noise filtering processing on the first table feature to obtain a second table feature.

[0056] A table repair unit, configured to perform cell repair processing on the second table feature to obtain a third table feature.

[0057] A text feature extraction unit, configured to build an OCR module and use the OCR module to extract text features from the third table feature.

[0058] A consistency comparison unit, configured to compare the text features with the two-dimensional code parsing data to determine consistency, and in the case of determining inconsistency, make an inconsistency mark in the corresponding invoice image;

[0059] A keyword field extraction unit, configured to, in the case of determining consistency, extract keyword fields from the text features based on regular expressions and semantic rules.

[0060] In a third aspect, the present invention provides a readable storage medium storing one or more programs, and the one or more programs can be executed by one or more processors to implement the method as described above.

[0061] The present invention has at least the following beneficial effects:

[0062] (1) Optimize the problems of invoice image tilt, noise, and two-dimensional code interference: Through the multi-modal angle correction technology and the two-dimensional code shielding technology of the ZBar library, this method improves the preprocessing effect of the invoice image. Compared with the traditional method that only relies on the Hough line transform to detect the angle of the table line, the present invention combines the Hough line transform with two-dimensional code positioning to ensure more reliable horizontal correction of the image. At the same time, the two-dimensional code shielding technology of the ZBar library can effectively shield the interference in the two-dimensional code area and avoid its influence on the rotation angle calculation and table detection.

[0063] (2) Improve the repair effect of broken table lines: By dynamically adjusting the size of the structural element of the morphological operation, the table lines are flexibly repaired according to the width and breakage degree of the table lines. According to the width and breakage degree of the table lines, the sizes of the dilation and erosion structural elements are flexibly adjusted to ensure the accuracy of the repair effect and better restore the complete structure of the table.

[0064] (3) Enhance the recognition ability of OCR for special characters and handwritten scripts in Chinese invoices: By fine-tuning based on the PaddleOCRv3 model, this method uses a 5,000-piece Chinese invoice dataset (including FangSong, KaiTi, and some handwritten scripts) for special optimization. The sample ratio of special characters such as "¥" and "Tax ID:" is particularly increased in the training data, significantly improving the model's recognition ability for these characters. In addition, by adjusting the output dimension of the fully connected layer to 5,000 and optimizing the training parameters, the model is more adapted to the complex scenarios of Chinese invoices. Finally, the overall accuracy of OCR recognition has been significantly improved, especially in the recognition of special characters and handwritten scripts.

[0065] (4) Improve the multi-modal verification mechanism: This method introduces a multi-modal verification module. By cross-verifying the keyword fields extracted by OCR with the QR code parsing data, the accuracy of keyword fields such as invoice code and amount is ensured. At the same time, for OCR results with low confidence (confidence lower than 0.9), a multi-model voting mechanism (integrating EasyOCR and Tesseract) is triggered, and the final result is selected through the principle of majority agreement. For the case of missing fields, a re-scanning mechanism is triggered, further improving the reliability and integrity of the data.

[0066] In summary, the present invention effectively solves the deficiencies existing in the prior art in invoice processing. By utilizing an optimized image preprocessing algorithm, a detection model specifically designed for nested tables and small tables, and an improved OCR recognition technology, the efficient parsing of complex invoice structures is achieved, the accuracy of keyword field extraction is improved, and the false detection rate and missed detection rate caused by image quality and format diversity are reduced. At the same time, the present invention also introduces multi-modal data fusion and multi-model verification strategies, further improving the stability and integrity of invoice processing. The present invention greatly improves the efficiency and accuracy of invoice processing, providing an innovative solution for the development and practical application of invoice processing technology. Description of the Drawings

[0067] Figure 1 Shows a schematic diagram of a multi-modal invoice image processing system based on vision and deep learning according to an embodiment of the present invention.

[0068] Figure 2 Shows a flowchart of a multi-modal invoice image processing method based on vision and deep learning according to an embodiment of the present invention;

[0069] Figure 3 Shows a structural diagram of a multi-modal invoice image processing device based on vision and deep learning according to an embodiment of the present invention. Detailed Embodiments

[0070] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be described in detail below in conjunction with the drawings and specific embodiments. The embodiments of the present invention will be further described in detail below in conjunction with the drawings and specific examples, but shall not be construed as a limitation to the present invention. For the various steps described herein, if there is no necessity for a sequential relationship between them, the order in which they are described as examples herein shall not be regarded as a limitation, and those skilled in the art should know that they can be adjusted in order as long as the logic between them is not destroyed and the entire process cannot be realized.

[0071] The embodiment of the present invention provides a multi-modal invoice image processing method based on vision and deep learning. This method can be based on Figure 1Execute using the multi-modal invoice image processing system based on vision and deep learning shown below. The multi-modal invoice image processing system based on vision and deep learning includes six modules. Please combine Figure 2 , corresponding to the six steps of this method respectively. The six modules are an image preprocessing module, a table detection module, a cell repair module, an OCR module, a multi-modal verification module, and a keyword field extraction module. The detailed implementation principle and its effects of this method will be specifically described below in combination with Figure 2 the steps S10 - S60 shown below.

[0072] S10. Obtain the invoice image, and based on the invoice image, calculate the angle of the table edge and the angle of the QR code positioning frame. Generate a binary mask to set the pixels in the QR code area of the invoice image to 0 to shield interference. Calculate the fusion rotation angle according to the angle of the table edge and the angle of the QR code positioning frame, and rotate the invoice image according to the fusion rotation angle to obtain the first image.

[0073] The operations involved in step S10 can be executed by the image preprocessing module. Its purpose is to preprocess the invoice image, including extracting the QR code and shielding the QR code area, as well as rotating the invoice image, so that the preprocessed image (i.e., the first image) can be better processed in the next step. After the processing of step S10, the inclination, noise, and QR code interference of the invoice image can be eliminated, ensuring the efficiency and accuracy of subsequent processing.

[0074] In an exemplary embodiment, the process of preprocessing the invoice image using the image preprocessing module specifically includes:

[0075] S101. Use the ZBar library to scan the invoice image, locate the bounding box of the QR code through the Canny edge detection algorithm (low threshold 50, high threshold 150), generate a binary mask to set the pixels in the QR code area to 0 (black) to shield interference, and set the remaining areas to 1 (white). The edge of the mask uses Gaussian blur (σ = 2) for transition to avoid the sawtooth effect caused by hard cutting. If the QR code cannot be parsed due to damage or blur, the system automatically increases the image resolution to 300 dpi and re-detects it to ensure the integrity of the QR code data.

[0076] S102. In the multi-modal angle correction stage, this stage combines the Hough line transform and the QR code positioning angle, and obtains the fusion rotation angle θ according to the following formula:

[0077] θ = αθ QR +(1 - α)θ Hough

[0078] where θ Hough represents the angle of the table edge detected by the Hough line transform, and θ QRIndicates the angle of the QR code positioning frame. α represents the weight coefficient.

[0079] Specifically, perform Canny edge detection on the invoice image, extract table line segments whose length exceeds 30% of the image width, and calculate their average inclination angle θ Hough ; According to the four corner point coordinates of the QR code positioning frame, calculate the diagonal slope, and derive the QR code inclination angle θ QR ; Finally, fuse the two weighted, and set the weight coefficient α to 0.7, indicating that the angle calculation preferentially relies on the QR code angle to achieve higher accuracy.

[0080] The fused rotation angle θ after substituting the weight α = 0.7 is:

[0081] θ = 0.7θ QR + 0.3θ Hough

[0082] It should be noted that the value of the above weight coefficient α is only an example and does not constitute a limitation to the present invention. In other embodiments, the weight coefficient α can adopt other values, and any value between 0 and 1 can be applied to the calculation formula of the fused rotation angle θ.

[0083] S20. Construct a detection model. The detection model is an improved YOLOv5 model obtained by adding two groups of detection frames to the YOLOv5 model. Use the training data set to train the detection model to obtain a table detection module. Input the first image into the table detection module, extract the first table feature, and perform non-maximum suppression and noise filtering on the first table feature to obtain the second table feature.

[0084] In this embodiment, the purpose of step S20 is to extract the table feature from the first image. The first table feature is the directly extracted table feature, and the second table feature is the table feature after noise reduction processing on the directly extracted table feature. The second table feature contains many useful information in the invoice, and these information will be further processed based on the second table feature in the subsequent steps to extract them.

[0085] In an exemplary embodiment, the table detection module achieves high-precision positioning of nested tables based on an improved YOLOv5 model. The default anchor box sizes of the original YOLOv5 are (10,13), (16,30), and (33,23). To adapt to the dense distribution of small tables in Chinese invoices (such as the commodity details table), the idea of the present invention is to optimize the anchor box sizes and the loss function. Specifically, two new groups of anchor boxes are added on the original basis: (8,10) and (12,18) to improve the detection ability for cancellation targets. The training dataset contains 10,000 Chinese-annotated invoices (including ordinary VAT invoices, electronic invoices, and roll invoices), covering 50 formats. Nested tables (such as a main table with an embedded sub-table) are labeled as the "composite table" category, and the bounding box contains the entire nested area.

[0086] The CIoULoss (Complete-IoU Loss) is used as the loss function to optimize the position, size, and aspect ratio of the detection box. The specific formula is as follows:

[0087]

[0088] Where IoU represents the intersection over union of the predicted box and the ground truth box, ρ represents the Euclidean distance between the centers of the predicted box and the ground truth box, c represents the diagonal length of the minimum bounding rectangle, α1 represents the balance parameter, and ν represents the aspect ratio consistency parameter.

[0089] In the post-processing stage, non-maximum suppression (NMS) is adopted. The IoU threshold is set to 0.5 to merge overlapping detection boxes. If the height of the detection box exceeds 65% of the total height of the invoice, that is, according to the following judgment condition:

[0090] H box >0.65H invoice

[0091] If the above judgment condition is met, it is determined as a noise box (such as the entire invoice being misdetected as a table) and re-detection is triggered. Where H box is the height of the detection box (unit: pixel), and H invoice is the total height of the invoice image (unit: pixel).

[0092] For nested tables, for the detection box marked as "composite table", the detection box of the "composite table" is used as the parent table detection box. The YOLOv5 model is run again inside the parent table detection box to locate the sub-table detection box area, such as the amount details sub-table, etc. Precise segmentation is achieved by calculating the intersection coordinates of the edges of the parent table detection box and the sub-table detection box, and the sub-table coordinates are mapped to the original invoice image coordinate system to ensure the consistency of subsequent cell extraction. Therefore, among the second table features used to represent nested tables, nested cells can be better distinguished, which is conducive to cell extraction and / or cell repair processing.

[0093] S30. Perform cell repair on the second table feature to obtain the third table feature.

[0094] In this embodiment, a cell repair module can be used to perform cell repair on the second table feature. The purpose is to repair the broken table lines in the second table feature, ensure the integrity of the cell structure, and thus ensure the accuracy of subsequent text feature extraction.

[0095] In an exemplary embodiment, the cell repair module repairs the broken table lines through adaptive morphological operations and edge filling techniques to ensure the integrity of the cell structure. The size of the structural element for morphological operations is dynamically adjusted according to the invoice width, and the formula is as follows:

[0096] K size = max(3, [0.02W invoice )

[0097] where K size represents the size of the structural element, W invoice represents the width of the invoice image (unit: pixel). The proportionality coefficient is set to 0.02 to dynamically adapt to invoices with different resolutions, and the minimum size is set to 3 to avoid invalid operations for small invoices.

[0098] Then perform a closing operation to fill the cell gaps caused by stamps or stains by using a square kernel; then perform an opening operation to eliminate isolated noise points (such as signatures or ink marks).

[0099] Next, enter the edge filling stage. Detect the breakage of vertical line segments through a custom hit-miss transform kernel K HMT . The definition of the hit-miss transform kernel K HMT is as follows:

[0100]

[0101] This hit-miss transform kernel K HMT locates the break points by matching the brightness change in the vertical direction (bright in the upper half and dark in the lower half), then extends 3 pixels in the vertical direction and uses a linear interpolation algorithm to fill the line segment;

[0102] In the final table line repair stage, horizontal lines and vertical lines are merged with adjacent line segments through directional dilation operations (using a horizontal kernel [1, 1, 1] and a vertical kernel [1, 1, 1]). The thick lines are transformed into single-pixel-width skeleton lines through the Zhang-Suen thinning algorithm. Then, gaps at the intersections of the lines are detected, and it is judged whether the length of the gap is ≤ 2 pixels. If it is satisfied, B-spline curve smoothing connection is applied to ensure the continuity and accuracy of the table lines. If not, the edge filling operation is performed again.

[0103] S40. Construct an OCR module, and use the OCR module to extract text features from the third table feature.

[0104] In an exemplary embodiment, the specific process of constructing an OCR module and using the OCR module to extract text features from the third table feature is as follows:

[0105] First is the model training stage. The OCR module fine-tunes the special model for Chinese invoices based on the PaddleOCRv3 framework, and improves the recognition accuracy through data augmentation and optimization of the special character library. The training data includes 5,000 Chinese invoice images. By synthesizing printed texts such as FangSong and KaiTi, special characters such as "¥", "tax number", and "capital letters" are covered, and Gaussian noise (σ = 5) and moiré patterns are injected on this basis to simulate complex scenarios. Then, the model is fine-tuned. The output dimension of the last fully connected layer of the model is extended to 5,000 to support the Chinese invoice character set. At the same time, the initial learning rate of the model is set to 0.01, the BatchSize is set to 32, and the Cosine Annealing scheduler is adopted.

[0106] Secondly, enter the blank line detection and confidence calculation stage. The blank line judgment formula is as follows:

[0107]

[0108] where cell i represents the i-th cell image in the table row, max(cell i ) represents the maximum gray value (i.e., the highest brightness value) of the cell image, and min(cell i ) represents the minimum gray value (lowest brightness value) of the cell image. T cn is a threshold designed for complex Chinese strokes. Due to the differences between Chinese and English, the threshold here is set to 85 to avoid misjudging dense strokes as blank lines. If the peak interval (highest frequency gray value) of the cell gray histogram accounts for more than 90%, it is judged as a blank line and the OCR processing is excluded to improve the operation efficiency.

[0109] Calculate the confidence level for the OCR output results using Softmax probability. When the confidence level is lower than the set threshold, the multi-model voting mechanism will be triggered. The specific operation is reflected in the following step S50.

[0110] S50. Compare the text features with the QR code parsing data to determine consistency. In the case of determining inconsistency, make an inconsistency mark in the corresponding invoice image.

[0111] In this embodiment, the processing operations involved in step S50 are executed by the multi-modal verification module. The core task of this multi-modal verification module is to verify the consistency between the key fields extracted by OCR and the QR code parsing data to ensure the accuracy of the basic data.

[0112] In an exemplary embodiment, the multi-modal verification module obtains the recognition results (i.e., text features) and their confidence levels of fields such as invoice code and amount from the OCR module, and at the same time parses the QR code in the invoice image to obtain the encoded data of the same fields, such as invoice code and amount. Compare the extraction results of the OCR module and the data of the QR code. If the two extractions are inconsistent, give priority to the QR code data and call the publicly available API of the tax bureau to verify the legality of the invoice code; for the amount field, convert the Arabic numerals recognized by the OCR module, such as "1234.56", into Chinese capital format, such as "One thousand two hundred and thirty-four yuan and fifty-six cents", and perform a semantic comparison with the capital amount extracted by OCR. If they are inconsistent, mark it as "pending manual review".

[0113] In an exemplary embodiment, multiple different auxiliary OCR models, such as EasyOCR and Tesseract, are integrated in the multi-modal verification module. In the case of calculating the confidence level of the OCR output results, when the confidence level of the output result of the main model is lower than the set threshold T conf = 0.9, the multi-model voting mechanism will be triggered, and the module will integrate the output results of PaddleOCR, EasyOCR, and Tesseract, and select the final recognized content through the principle of majority agreement.

[0114] If a key field is missing, the system will trigger the re-scanning mechanism. Preset the positions of common fields. When a field is missing, such as the verification code is missing and the position is usually in the upper right corner, the system will focus on the local area in the upper right corner of the invoice for re-identification.

[0115] S60. In the case of determining consistency, extract the key fields from the text features based on regular expressions and semantic rules.

[0116] In an exemplary embodiment, step S60 can be performed by a key field module, and the key field extraction module extracts the invoice code, check code and amount capitalization field from the OCR text through basic regular expressions. The invoice code is extracted using the regular expression \d{10,12}, which matches 10 to 12 consecutive digits, such as "123456789012", and excludes consecutive repeated characters, such as 1111111111. The check code matches the 5 digits at the end of the invoice code through the regular expression \d{5}$, such as the last 5 digits of the invoice code "123456789012" "89012"; the amount capitalization field extracts the Chinese amount text through the character set regular expression 壹贰叁肆伍六789拾佰千軟萬亿角分整]+.

[0117] The semantic rule engine uses contextual positioning, such as the text on the right side of "Total (uppercase)", to increase the field extraction priority. At the same time, a logical check is performed to ensure that the last digit of the checksum must be consistent with the last digit of the invoice code. If a field is detected to be missing or formatted incorrectly, it is marked as an exception. Abnormal situations will trigger multi-model voting and rescanning mechanisms. The extraction results are double-verified by regular expressions and semantic rules, stored in the database, and returned to the enterprise ERP system to ensure the accuracy and completeness of financial data.

[0118] In summary, the fusion of vision and deep learning methods can effectively process Chinese invoices. This method comprehensively considers the problems existing in the invoice processing task, including high font complexity, poor model adaptability to Chinese invoice-specific fields, and low recognition accuracy of existing OCR models, and obtains more accurate Chinese invoice detection results. This processing method can improve the speed and accuracy of invoice processing and has good robustness and stability.

[0119] The embodiment of the present invention also provides a multimodal invoice image processing device based on vision and deep learning, such as Figure 3 As shown, the device comprises:

[0120] The image preprocessing unit 301 is configured to obtain an invoice image, and based on the invoice image, calculate the angle of the table edge and the angle of the two-dimensional code positioning frame, generate a binary mask to set the pixels of the two-dimensional code area in the invoice image to 0 to shield interference, calculate the fusion rotation angle according to the angle of the table edge and the angle of the two-dimensional code positioning frame, and rotate the invoice image according to the fusion rotation angle to obtain a first image;

[0121] The table detection unit 302 is configured to construct a detection model, which is an improved YOLOv5 model obtained by adding two sets of detection frames to the YOLOv5 model. The table detection module is obtained by training the detection model with a training data set. The first image is input into the table detection module, and the first table feature is extracted. The non-maximum suppression and noise filtering processes are applied to the first table feature to obtain the second table feature;

[0122] The table repair unit 303 is configured to perform cell repair processing on the second table feature to obtain the third table feature;

[0123] The text feature extraction unit 304 is configured to construct an OCR module and extract text features from the third table feature by using the OCR module;

[0124] The consistency comparison unit 305 is configured to compare the text features with the two-dimensional code parsing data to determine the consistency. In the case of determining inconsistency, an inconsistency mark is made in the corresponding invoice image;

[0125] The keyword field extraction unit 306 is configured to extract keyword fields from the text features based on regular expressions and semantic rules in the case of determining consistency.

[0126] In some embodiments, the image preprocessing unit is further configured to:

[0127] Calculate the average inclination angle of the table line segments whose length exceeds 30% of the image width in the invoice image as the angle of the table edge;

[0128] Calculate the diagonal slope according to the four corner point coordinates of the two-dimensional code positioning frame, and calculate the angle of the two-dimensional code positioning frame according to the diagonal slope;

[0129] According to the angle of the table edge and the angle of the two-dimensional code positioning frame, calculate the fusion rotation angle through the following formula:

[0130] θ = αθ QR +(1 - α)θ Hough

[0131] where θ Hough represents the angle of the table edge, θ QR represents the angle of the two-dimensional code positioning frame, and α represents the weight coefficient.

[0132] In some embodiments, the table detection unit is further configured to set a loss function, and train the detection model with a training data set to obtain a table detection module; wherein, the loss function is used to optimize the position, size and aspect ratio of the detection frame, and the loss function is expressed as:

[0133]

[0134] Among them, represents the loss value, IoU represents the intersection over union of the predicted bounding box and the ground truth bounding box, ρ represents the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box, c represents the diagonal length of the minimum bounding rectangle, α1 represents the balance parameter, and v represents the aspect ratio consistency parameter.

[0135] In some embodiments, the table detection unit is further configured to:

[0136] Set an IoU threshold, merge overlapping detection bounding boxes, and trigger re - prediction if the height of the detection bounding box meets a preset condition; where the preset condition is expressed as:

[0137] H box > 0.65H invoice

[0138] where H box is the height of the detection bounding box, and H invoice is the total height of the invoice image;

[0139] If the first table feature is a nested table, after non - maximum suppression processing, mark the detection bounding box corresponding to the first table feature as the detection bounding box of the composite table, and use it as the parent table detection bounding box. Locate the area of the child table detection bounding box in the parent table detection bounding box, perform segmentation by calculating the edge intersection coordinates of the parent table detection bounding box and the child table detection bounding box, and map the child table detection bounding box to the original invoice image coordinate system.

[0140] In some embodiments, the table repair unit is further configured to:

[0141] Perform adaptive morphological operations, closing operations, opening operations, edge filling, and table line repair processing on the second table feature in sequence to obtain a third table feature; where:

[0142] The size of the structural element for the morphological operation is dynamically adjusted according to the invoice width, and the formula is as follows:

[0143] K size = max(3, [0.02W invoice )

[0144] where K size represents the size of the structural element, and W invoice represents the width of the invoice image;

[0145] When performing the closing operation, use a square kernel to fill the cell gaps caused by stamping or stains; perform the opening operation to eliminate isolated noise points;

[0146] When filling in the edges, a hit-miss transform kernel is set. By using the hit-miss transform kernel to match the luminance change in the vertical direction, the break points are located. Subsequently, a set number of pixels are extended along the vertical direction and the linear interpolation algorithm is used to fill in the line segments.

[0147] In the stage of repairing table lines, the horizontal lines and vertical lines are respectively merged with adjacent line segments through directional dilation operations. Through the thinning algorithm, the thick lines are transformed into skeleton lines with a single-pixel width, and the gaps at the intersections of the lines are detected. It is judged whether the length of the gap is less than or equal to the preset number of pixels. If so, a B-spline curve is applied for smooth connection.

[0148] In some embodiments, the OCR module fine-tunes the special model for Chinese invoices based on the PaddleOCRv3 framework, improves the recognition accuracy through data augmentation and optimization of the special character library. Based on the initial output features of the OCR module, blank line detection is performed, and the blank lines are excluded to obtain text features.

[0149] When performing blank line detection, the blank line determination formula is set as follows:

[0150]

[0151] where cell i represents the i-th cell image in the table row, max(cell i ) represents the maximum gray value of the cell image, min(cell i ) represents the minimum gray value of the cell image; T cn is a threshold designed for complex Chinese strokes.

[0152] When the initial output features satisfy the determination formula, if the proportion of the peak interval in the cell gray histogram exceeds 90%, it is determined as a blank line and excluded.

[0153] According to the output of the OCR module, that is, the text features, Softmax probability calculation is used. Only when the confidence level is lower than the set threshold, the text features are compared with the two-dimensional code parsing data.

[0154] In some embodiments, the consistency comparison unit is further configured to:

[0155] Obtain the recognition results of the invoice code and amount fields from the text features output by the OCR module, and at the same time parse the encoded data of the same fields in the invoice image by scanning the two-dimensional code.

[0156] Compare the recognition results of the invoice code and amount fields with the encoded data of the same fields. If they are inconsistent, first use the QR code data and call the API to verify the legality of the invoice code. For the amount field, use the rule engine to convert the Arabic numerals recognized by OCR into Chinese capital format and perform semantic comparison with the capital amount extracted by OCR to determine if they are the same.

[0157] In some embodiments, the keyword field extraction unit is further configured to:

[0158] The extraction of the invoice code uses the first regular expression to match 10 to 12 consecutive digits and excludes consecutive repeated characters. The check code matches several digits at the end of the invoice code through the second regular expression; the capital amount field extracts the capital amount text through the character set regular expression.

[0159] The semantic rule engine improves the field extraction priority through context positioning and performs logical verification at the same time to ensure that the last digit of the check code is consistent with the last digit of the invoice code.

[0160] It should be noted that the structures of the various multi-modal invoice image processing devices based on vision and deep learning described in this embodiment belong to the same technical concept as the previously described multi-modal invoice image processing method based on vision and deep learning, and achieve the same beneficial effects through the same principle, which will not be elaborated here.

[0161] The embodiment of the present invention also provides a readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method described in any of the above embodiments.

[0162] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of their solutions) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description. Additionally, in the above specific implementation manners, various features can be grouped together to simplify the present invention. This should not be construed as an intention that the features of an invention not claimed are necessary for any claim. On the contrary, the subject matter of the present invention can be less than all the features of a specific embodiment of the invention. Thus, the following claims are incorporated herein as examples or embodiments in the specific implementation manners, where each claim independently serves as a separate embodiment, and considering these embodiments, they can be combined with each other in various combinations or permutations. The scope of the present invention should be determined with reference to the appended claims and the full scope of the equivalent forms empowered by these claims.

Claims

1. A multimodal invoice image processing method based on vision and deep learning, characterized in that: The method comprises: Acquire an invoice image, and based on the invoice image, calculate the angle of the table edge and the angle of the two-dimensional code positioning frame, generate a binary mask to set the pixels of the two-dimensional code area in the invoice image to 0 to shield interference, calculate the fusion rotation angle according to the angle of the table edge and the angle of the two-dimensional code positioning frame, and rotate the invoice image according to the fusion rotation angle to obtain a first image; Constructing a detection model, wherein the detection model is an improved YOLOv5 model obtained by adding two groups of detection frames to the YOLOv5 model, training the detection model using a training data set to obtain a table detection module, inputting the first image into the table detection module, extracting a first table feature, and performing non-maximum suppression and noise filtering on the first table feature to obtain a second table feature; Performing cell repair processing on the second table feature to obtain a third table feature; Constructing an OCR module, and using the OCR module to extract text features from the third table features; Comparing the text features with the QR code parsing data to determine consistency, and if they are determined to be inconsistent, marking the inconsistency in the corresponding invoice image; When it is determined to be consistent, key fields are extracted from the text features based on regular expressions and semantic rules.

2. According to claim 1, the multimodal invoice image processing method based on vision and deep learning is characterized in that: The fusion rotation angle is calculated based on the angle of the table edge and the angle of the QR code positioning frame, including: Calculate the average tilt angle of the table line segments in the invoice image whose length exceeds 30% of the image width as the angle of the table edge; According to the coordinates of the four corner points of the QR code positioning frame, the diagonal slope is calculated, and the angle of the QR code positioning frame is calculated according to the diagonal slope; According to the angle of the table edge and the angle of the QR code positioning frame, the fusion rotation angle is calculated by the following formula: θ = θ QR +1-a)θ Hough Among them, θ Hough Represents the angle of the table edge, θ QR It represents the angle of the QR code positioning frame, and α represents the weight coefficient.

3. The multimodal invoice image processing method based on vision and deep learning according to claim 1 is characterized in that: A loss function is set, and the detection model is trained using a training data set to obtain a table detection module; wherein the loss function is used to optimize the position, size, and aspect ratio of the detection box, and the loss function is expressed as: in, Represents the loss value, IoU represents the intersection over union of the predicted box and the true box, ρ represents the Euclidean distance between the center point of the predicted box and the true box, C represents the diagonal length of the minimum circumscribed rectangle, α1 represents the balance parameter, and v represents the aspect ratio consistency parameter.

4. The multimodal invoice image processing method based on vision and deep learning according to claim 3 is characterized in that: The first table feature is processed by non-maximum suppression and noise filtering to obtain a second table feature, including: Set the IoU threshold, merge the overlapping detection boxes, and trigger re-prediction if the height of the detection box meets the preset conditions; wherein the preset conditions are expressed as: H box >0.65H invoice Among them, H box is the height of the detection box, H invoice is the total height of the invoice image; If the first table feature is a nested table, after non-maximum suppression processing, the detection frame corresponding to the first table feature is marked as the detection frame of the composite table, and is used as the parent table detection frame. The area of ​​the child table detection frame is located in the parent table detection frame, and the parent table detection frame and the child table detection frame are segmented by calculating the edge intersection coordinates of the parent table detection frame and the child table detection frame, and the child table detection frame is mapped to the original invoice image coordinate system.

5. The multimodal invoice image processing method based on vision and deep learning according to claim 1 is characterized in that: Performing cell repair processing on the second table feature to obtain a third table feature includes: The second table feature is subjected to adaptive morphological operation, closing operation, opening operation, edge padding and table line repair processing in sequence to obtain a third table feature; wherein: The size of the structural element of the morphological operation is dynamically adjusted according to the invoice width, and the formula is as follows: K size =max(3,[0.02W invoice ]) Among them, K size Represents the size of the structural element, W invoice Represents the width of the invoice image; When performing the closing operation, the cell gaps caused by stamps or stains are filled by using a square kernel; the opening operation is used to eliminate isolated noise points; When filling the edge, a hit-miss transformation kernel is set, and the breakpoint is located by matching the brightness change in the vertical direction through the hit-miss transformation kernel, and then the set pixels are extended in the vertical direction and the line segment is filled using a linear interpolation algorithm; In the table line repair stage, horizontal lines and vertical lines are merged with adjacent line segments through directional expansion operations respectively; through the refinement algorithm, the thick lines are converted into skeleton lines with a single pixel width, and the gaps at the intersections of the lines are detected to determine whether the gap length is less than or equal to the preset pixels. If so, the B-spline curve is applied to smoothly connect them.

6. The multimodal invoice image processing method based on vision and deep learning according to claim 1 is characterized in that: The OCR module fine-tunes the Chinese invoice-specific model based on the PaddleOCRv3 framework, improves recognition accuracy through data enhancement and special character library optimization, and performs blank line detection and removes blank lines based on the initial output features of the OCR module to obtain text features; When performing blank line detection, set the blank line determination formula as follows: Among them, cell i Represents the i-th cell image in the table row, max(cell i ) represents the maximum grayscale value of the cell image, min(cell i ) represents the minimum grayscale value of the cell image; T cn It is a threshold designed for complex Chinese strokes; When the initial output features satisfy the judgment formula, if the peak interval of the cell grayscale histogram accounts for more than 90%, it is judged as a blank row and excluded; Integrate multiple different auxiliary OCR models. When the confidence of the output result of the OCR module is lower than the set threshold, trigger the multi-model voting mechanism. The multi-model voting mechanism selects the final recognition content as the text feature based on the majority consensus principle of the output results of the integrated multiple auxiliary OCR models.

7. The multimodal invoice image processing method based on vision and deep learning according to claim 1 is characterized in that: Comparing the text features with the QR code parsing data to determine consistency, including: Obtain the recognition results of the invoice code and amount fields from the text features output by the OCR module, and parse the QR code in the invoice image to obtain the coded data of the same fields; Compare the recognition results of the invoice code and amount fields with the encoded data of the same fields. If the two are inconsistent, first use the QR code data and call the API to verify the legitimacy of the invoice code; for the amount field, use the rule engine to convert the Arabic numerals recognized by OCR into Chinese uppercase format, and perform a semantic comparison with the uppercase amount extracted by OCR to determine whether they are the same.

8. The multimodal invoice image processing method based on vision and deep learning according to claim 5 is characterized in that: If the text features are consistent, key fields are extracted from the text features based on regular expressions and semantic rules, including: The first regular expression is used to extract the invoice code, matching 10 to 12 consecutive digits and excluding consecutive repeated characters. The checksum is matched with several digits at the end of the invoice code through the second regular expression. The uppercase amount field is extracted using the character set regular expression. The semantic rule engine increases the priority of field extraction through contextual positioning, and performs logical verification to ensure that the last digit of the verification code is consistent with the last digit of the invoice code.

9. A multimodal invoice image processing device based on vision and deep learning, characterized in that: The device comprises: The image preprocessing unit is configured to obtain an invoice image, and based on the invoice image, calculate the angle of the table edge and the angle of the two-dimensional code positioning frame, generate a binary mask to set the pixels of the two-dimensional code area in the invoice image to 0 to shield interference, calculate a fusion rotation angle according to the angle of the table edge and the angle of the two-dimensional code positioning frame, and rotate the invoice image according to the fusion rotation angle to obtain a first image; A table detection unit is configured to construct a detection model, wherein the detection model is an improved YOLOv5 model obtained by adding two groups of detection frames to the YOLOv5 model, the detection model is trained using a training data set to obtain a table detection module, the first image is input into the table detection module, a first table feature is extracted, and the first table feature is processed by non-maximum suppression and noise filtering to obtain a second table feature; a table repair unit, configured to perform cell repair processing on the second table feature to obtain a third table feature; A text feature extraction unit is configured to construct an OCR module and extract text features from the third table features using the OCR module; A consistency comparison unit is configured to compare the text feature with the QR code parsing data to determine the consistency, and if it is determined to be inconsistent, make an inconsistent mark in the corresponding invoice image; The key field extraction unit is configured to extract the key field from the text features based on regular expressions and semantic rules when the text features are judged to be consistent. 10 . A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, perform the method according to claim 1 .

Citation Information

Patent Citations

  • Method and system for improving electronic invoice identification accuracy

    CN117671713A

  • Picture information structuring method and device based on table lines

    CN117711006A

  • Engineering quality intelligent management method and system based on BIM and AI large models

    CN118095975A

  • Invoice recognition method, computer device and storage medium

    WO2024183298A1

Cited By

  • Logistics express bill automatic identification and bill number extraction method based on rule configuration

    CN121354157A