Intelligent value-added tax invoice information extraction method and device based on coordinate constraint

By using deep learning neural networks for orientation correction and layout cropping, combined with coordinate constraints and arithmetic consistency checks, the problem of accurate extraction of value-added tax invoices in diverse environments was solved, achieving efficient and stable information extraction results.

CN121768027APending Publication Date: 2026-03-31WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511856014.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing automatic VAT invoice recognition technology is not robust enough when faced with image orientation and deformation sensitivity, overlapping fields and noise interference, resulting in low field extraction accuracy and difficulty in adapting to diverse invoice formats and imaging conditions.

Method used

A method based on deep learning neural networks is used for orientation detection and correction, page cropping, optical character recognition, and geometric position constraints based on a normalized coordinate system. Combined with arithmetic consistency verification, key information fields are located and corrected.

Benefits of technology

It significantly improves the accuracy and stability of field extraction, reduces cross-referencing errors, enhances the system's applicability and processing speed, and is suitable for invoice information extraction in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768027A_ABST
    Figure CN121768027A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent value-added tax invoice information extraction method and device based on coordinate constraint, and the method comprises the steps: sequentially carrying out the direction detection and correction and layout cutting preprocessing of a to-be-recognized value-added tax invoice image, and obtaining a face value main image which is correct in direction and is cut off from a redundant background; performing optical character recognition on the ticket face main body image, and outputting a recognition result containing text content and geometric coordinates of the text content in a normalized coordinate system; on the basis of the normalized coordinate system, according to a preset geometric position constraint condition corresponding to the invoice layout, analyzing and positioning each key information field from the identification result; wherein for a plurality of text blocks which are located at adjacent positions due to recognition and splitting, and coordinate spacing of the text blocks meets a merging condition, the text blocks are merged and regarded as the same field content; and performing arithmetic consistency verification on a numerical level on the analyzed and positioned money amount, tax amount and price and tax summation fields, and performing automatic correction on inconsistent fields based on a verification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and intelligent document processing technology, and particularly relates to a technical solution for intelligent extraction of value-added tax invoice information based on coordinate constraints. Background Technology

[0002] Existing automatic VAT invoice recognition technologies primarily rely on keyword matching after OCR text recognition for field extraction. However, this approach suffers from several technical bottlenecks. First, keywords such as "name" and "taxpayer identification number" in invoices often overlap and repeat across different fields (buyer, seller, bank), easily leading to misalignment errors in field location and reducing the accuracy of information extraction. Second, existing solutions are highly sensitive to image orientation and deformation. When invoice images are flipped, tilted, bent, or distorted, the text detection box position often shifts, severely affecting the correct recognition of field areas. Furthermore, scanned documents often contain edge whitespace or random noise, which can be misidentified as valid text areas by OCR, generating redundant detection boxes and interfering with rule-based layout positioning and field matching. These issues collectively result in insufficient robustness of existing automatic data entry solutions under various shooting or scanning environments, making it difficult to adapt to the diverse invoice formats and imaging conditions in actual business operations. Therefore, there is an urgent need for an intelligent invoice information extraction method that combines image coordinate space structure information, has orientation and distortion adaptability, and can effectively remove noise interference, so as to improve the accuracy and stability of the automatic value-added tax invoice entry system. Summary of the Invention

[0003] This invention is mainly aimed at the BeiDou ionospheric random error model. It provides a coordinate-constrained intelligent extraction technology for value-added tax invoice information based on a deep learning neural network model.

[0004] The technical solution of this invention provides a method for intelligent extraction of value-added tax invoice information based on coordinate constraints, comprising: The VAT invoice image to be identified is sequentially subjected to orientation detection and correction, and layout cropping preprocessing to obtain the main image of the invoice with the correct orientation and redundant background cropped out. Optical character recognition is performed on the main image of the ticket, and the recognition result containing the text content and its geometric coordinates in a normalized coordinate system is output. Based on the normalized coordinate system, and according to the preset geometric position constraints corresponding to the invoice layout, each key information field is parsed and located from the recognition result; among them, for multiple text blocks that are adjacent due to recognition splitting and whose coordinate spacing meets the merging condition, they are merged and regarded as the same field content. Perform arithmetic consistency checks on the parsed and located fields of amount, tax amount, and total price including tax, and automatically correct inconsistent fields based on the check results.

[0005] Moreover, the orientation detection and correction is performed by analyzing the overall arrangement orientation of the text in the image and rotating the image to bring it into a standard reading orientation.

[0006] Moreover, the page cropping is achieved by detecting the boundaries of the effective text areas in the image and cropping out the area image that contains only the main body of the invoice, so as to eliminate edge blanks and background interference.

[0007] Furthermore, the optical character recognition outputs the vertex coordinates of the polygonal region corresponding to each recognized text block, and linearly maps all coordinates to a normalized coordinate system in the range of zero to one.

[0008] Furthermore, the key information fields for analytical positioning based on geometric position constraints in a normalized coordinate system specifically include: The search area for the buyer's name and identification number is constrained to the first third of the horizontal direction and the upper half of the vertical direction in the normalized coordinate system. The search area for the seller's name and identification number is constrained to the middle third of the horizontal direction in the normalized coordinate system. For the amount and tax fields, first locate the coordinates of the keywords "amount", "tax", or "total", and then search for the corresponding numerical information in the adjacent extended areas in the horizontal and vertical directions.

[0009] Moreover, the merging condition refers to the process of splicing together the recognized content of multiple text blocks belonging to the same semantic field to form a complete field information when the coordinate distance between them in the normalized coordinate system is less than a dynamic threshold set according to the invoice image size or typical character size.

[0010] Moreover, the automatic correction means that when the sum of the amount and the tax amount is not equal to the total price and tax and the difference exceeds the allowable error, the value of the amount or tax amount field is recalculated and updated based on the total price and tax value.

[0011] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the intelligent extraction method of VAT invoice information based on coordinate constraints as described above.

[0012] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described intelligent extraction method for value-added tax invoice information based on coordinate constraints.

[0013] On the other hand, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the intelligent extraction method for value-added tax invoice information based on coordinate constraints as described above.

[0014] The ticketing recognition scheme based on text recognition and coordinate extraction provided by this invention has the following beneficial effects: (1) Through the implementation of the above technical solutions, the present invention has achieved significant technical effects in practical applications. First, the present invention combines the dual constraint mechanism of spatial prior knowledge and text semantics to judge text content not only based on the text itself, but also based on the positional relationship of the text on the page, effectively solving the cross-referencing problem caused by the traditional method that only relies on keyword matching. Experiments have shown that in real invoice scenarios with multiple formats and multiple shooting angles, the present invention reduces the field cross-referencing rate by more than 75% compared with the traditional method.

[0015] (2) This invention, through orientation detection and page cropping, makes the image area input to OCR more accurate and compact, significantly reducing the misidentified area caused by interference factors such as blank edges and background noise, thereby greatly improving the overall recall rate of OCR. This optimization is particularly advantageous in scenarios such as mobile shooting, poor scanning environment, or large blank space on the ticket.

[0016] (3) The entire process of this invention is based on the spatial structure of the image itself and the text recognition results, so it can be applied to a variety of scenarios, including paper scans and historical photos, and has a wider range of applicability.

[0017] (4) The present invention can be implemented based on the Python programming language, combined with high-performance open source tools such as PaddleOCR. The overall algorithm complexity is low. In a normal 8-core CPU environment, it can complete the extraction and processing of all fields of a value-added tax invoice in less than 0.01 seconds. It has good engineering implementation and embedded deployment capabilities, and is particularly suitable for financial and tax SaaS systems or batch invoice digitization processing needs.

[0018] Therefore, this invention achieves comprehensive performance improvements over existing technologies in terms of field extraction accuracy, applicability, processing speed, and engineering implementation, demonstrating significant practical value and promising industrial application prospects. The solution of this invention is simple and convenient to implement, highly practical, and solves the problems of low practicality and inconvenience in actual application of related technologies. It can improve user experience and has significant market value. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the method flow of an embodiment of the present invention. Detailed Implementation

[0020] The following will further explain the concept, specific structure and technical effects of the present invention in conjunction with the accompanying drawings and embodiments, so as to fully understand the purpose, features and effects of the present invention.

[0021] The example uses open-source OCR (Optical Character Recognition, taking PaddleOCR as an example) combined with page cropping to automatically extract VAT invoice information by locating fields in the coordinate system.

[0022] As shown in Figure 1, this embodiment provides an intelligent extraction method for VAT invoice information based on coordinate constraints and normalized coordinate analysis. The method includes direction detection, layout cropping, text recognition (OCR recognition), text recognition result preprocessing (coordinate normalization and candidate filtering), coordinate rule parsing, and field consistency verification. The text detection and recognition steps output the text content, polygon coordinates, and confidence score of the invoice image, and map the pixel coordinates to a normalized coordinate system. Geometric position constraints and anchor point search windows are set for fields such as buyer, seller, amount, and tax on the normalized coordinates. OCR errors such as keyword separation and character loss are merged and tolerated within coordinate distance thresholds. Simultaneously, arithmetic consistency verification is performed on the amount, tax, and total price including tax, and automatic correction is applied. Experiments show that this method can complete the entire process of an invoice within 0.01 seconds on an 8-core CPU, with a field accuracy improvement of over 15% compared to pure text matching methods. It features low overhead and strong robustness, making it suitable for scenarios such as automatic tax image entry. The specific preferred implementation process is as follows: S1, Direction Detection: The direction detection and correction involves analyzing the overall arrangement direction of the text in the image and rotating and correcting the image to bring it into a standard reading orientation.

[0023] In the specific implementation of this invention, in order to improve the adaptability of invoice images under different shooting or scanning conditions, it is first necessary to perform image orientation detection to ensure the accuracy of subsequent layout analysis and field positioning.

[0024] For orientation detection, this invention can employ existing publicly available technologies, such as the Orientation and Script Detection (OSD) method provided by Tesseract OCR. The input image is analyzed using the pytesseract.image_to_osd interface to determine the overall rotation angle of the image. This method is primarily based on page layout analysis, connected component detection, and text baseline angle statistics. It can output information including page orientation angle, required rotation angle, orientation confidence, and text / script type, thereby determining the correct orientation of the image.

[0025] In this embodiment, a lightweight text orientation classification model or a character aspect ratio heuristic algorithm is used to analyze the VAT invoice image to be identified and output the main orientation angle θ of the image, where the value of θ is in the range of {0°, 90°, 180°, 270°}. If θ≠0°, rotation correction is performed to obtain an image with the correct orientation.

[0026] The text orientation classification model includes a lightweight convolutional neural network based on deep learning. The input is the entire invoice image, and the output is the predicted rotation angle of the image.

[0027] S2, Page Layout Detection and Cropping: The page layout cropping involves detecting the boundaries of valid text areas in the image and cropping out the area image containing only the main body of the invoice to eliminate edge blanks and background interference.

[0028] In the implementation of this invention, in order to improve the accuracy and efficiency of invoice image processing, the input image needs to be first subjected to layout detection and cropping to remove any blank edges, background interference and noise areas that may exist in the image.

[0029] This invention obtains a correct-oriented image of the invoice and removes redundant background by sequentially performing orientation detection and correction, and layout cropping preprocessing on the image of the VAT invoice to be identified. This step performs text region detection on the image processed in step S1, uses a layout analysis model or projection method to determine the invoice area, and removes edge blanks and noise interference areas, thereby obtaining a compact and effective invoice image.

[0030] The specific method for page cropping can be to use the PicoDet-Layout page analysis model or vertical and horizontal projection methods to determine the boundary coordinates of the ticket area and crop out the main body of the ticket.

[0031] Specifically, one feasible approach is based on a traditional projection algorithm combined with adaptive thresholding. This involves analyzing the image horizontally and vertically to detect dense text areas, thereby determining the boundaries of the main body of the ticket and cropping the image to retain the main body while removing unnecessary background. Another approach utilizes existing lightweight deep learning detectors, such as the PicoDet neural network model, to perform object detection on the entire image, directly identifying and outlining the ticket area. This is particularly suitable for scenarios with complex backgrounds or variable shooting environments. By detecting and cropping the layout, the variance of coordinates in the image can be significantly reduced, minimizing interference in subsequent text localization and recognition.

[0032] In a specific implementation of this invention, it is preferred to adopt a geometric boundary calculation method based on text box coordinates. This involves using the coordinate information of the text detection rectangles (dt_polys) output by the text detection model to calculate the minimum bounding rectangle of all text regions, which serves as the boundary of the ticket area. Image cropping is then performed based on this bounding rectangle. This method requires no additional model training, has a fast computation speed, and is suitable for large-scale ticket processing.

[0033] S3, Optical Character Recognition (OCR): Perform optical character recognition on the main image of the ticket and output the recognition result containing the text content and its geometric coordinates in a normalized coordinate system; the optical character recognition outputs the vertex coordinates of the polygonal region corresponding to each recognized text block and linearly maps all coordinates to a normalized coordinate system in the range of zero to one.

[0034] The embodiment performs optical character recognition (OCR) processing on the main image of the ticket and outputs a set of text instances including the geometric coordinates of the text region, the recognition confidence score, and the text content; wherein, the geometric coordinates of the text region are the set of coordinate points of the circumscribed polygon, the confidence score is used to characterize the recognition reliability, and the text content is the recognized character sequence.

[0035] This process consists of two sub-processes: text detection and text recognition. In specific implementation, the optical character recognition method can be implemented based on an existing OCR framework, which can obtain the recognition result of each character and the corresponding polygon coordinate box information. The OCR framework can be implemented through existing text detection and recognition networks, which can directly generate structured recognition results containing coordinate information from images. However, this invention does not limit the specific model structure or implementation method.

[0036] In this embodiment, the system first performs multi-scale convolution feature extraction on the main image of the ticket to obtain a high-dimensional feature map containing text edges and shape information.

[0037] Building upon this foundation, the text detection sub-network performs pixel-by-pixel prediction on the feature map, generating a probability heatmap of each pixel belonging to "text." This allows for the location and separation of text regions within the pixel coordinate system, providing geometric coordinates and a cropping window for subsequent recognition. Two preferred approaches for the specific implementation of detection are: First, the DB (Differentiable Binarization) approach, which involves differentially binarizing the probability heatmap and the learnable threshold map to obtain a binary image, followed by connected component / contour extraction and polygon fitting to obtain the coordinates of the bounding vertices of each text region; second, the EAST (Efficient and Accurate Scene Text Detector) approach, where the network directly regresses the geometric parameters or quadrilaterals of the text, subsequently converting them into polygon bounding vertex coordinates. Both approaches can be combined with threshold filtering and non-maximum suppression (NMS) based on polygon intersection-union ratio (IoU) to remove background noise and overlapping candidates, thus obtaining stable text region coordinates. For each detected region, the system performs affine / perspective correction and line-block cropping (ROI) between detection and recognition, normalizing tilted or distorted small blocks before feeding them into the text recognition sub-network. The text recognition sub-network can employ a CRNN (Convolutional Recurrent Neural Network), which extracts visual features through convolutional layers, models sequence context through bidirectional recurrent units, and decodes and outputs the character sequence and recognition confidence of the region through attention.

[0038] In this embodiment, the character sequence of the k-th text instance is denoted as... (Ordered character sequence), confidence level is Based on the combined detection and recognition results, the system outputs a set of text instance triples. , where k is the instance index and N is the total number of instances; The coordinates of the circumscribed vertices of the polygon for the k-th text region (rearranged in top-left → top-right → bottom-right → bottom-left or clockwise order). Then, the pixel coordinates (x, y) are linearly mapped to the normalized coordinate system [0, 1] × [0, 1], i.e. ,get .in, , The coordinates are pixel coordinates, with the origin at the top left corner of the image, and W and H being the width and height of the image, respectively. For normalized coordinates, it corresponds to linearly mapping pixel coordinates to the normalized coordinate system [0,1]×[0,1].

[0039] In one embodiment, the OCR module is implemented based on the open-source framework PaddleOCR (text detection uses differentiable binarization or efficient and accurate scene text detection, text recognition uses a convolutional recurrent neural network, and the cropped image is processed to output text instance triples). It should be understood that PaddleOCR is only an optional implementation, and this invention is not limited to this specific framework. To make this invention applicable to images of different sizes and resolutions, the coordinate information described above is normalized by scaling the horizontal and vertical coordinates of all text regions to the range of 0 to 1. Through this normalization process, even if the same type of invoice has different sizes or pixel resolutions under different scanning or shooting conditions, the spatial geometric rules can be ensured to have uniform applicability, thereby guaranteeing the accuracy and robustness of field positioning.

[0040] S4, Text Recognition Result Preprocessing: All text content undergoes unified conversion between full-width and half-width characters, that is, full-width numbers, letters, punctuation marks, and spaces in the recognition results are converted to their corresponding half-width forms. Furthermore, this invention also cleans up spaces and zero-width characters in the text, removing invisible characters to improve data consistency and standardization.

[0041] Furthermore, the preprocessing of the identified text content includes: Convert all full-width numbers, full-width letters, and full-width punctuation marks in the recognition results into their corresponding half-width characters; Remove spaces, invisible characters, and zero-width characters from the text.

[0042] Preferably, the preprocessing further includes: After character conversion and cleanup, the taxpayer identification number, invoice code, and invoice number fields are further cleaned up by removing any non-numeric characters and unifying them into a continuous string of half-width numeric characters. The values ​​in the Amount, Tax Amount, and Total Price and Tax fields are uniformly formatted into standard numerical form with two decimal places.

[0043] In the specific implementation of this invention, considering that invoice images may exhibit a mixture of full-width and half-width characters in the text results due to different scanning or shooting environments, or due to the characteristics of OCR recognition algorithms, this is particularly common in key fields such as the unified social credit code, taxpayer identification number, invoice number, and amount. For example, the number "1" may be recognized as the full-width character "1" (Unicode U+FF11) instead of the half-width character "1" (Unicode U+0031), causing field matching or comparison failures and affecting the accuracy of information extraction. Therefore, this invention further proposes that before parsing the text recognition results, all text content undergo a unified conversion process between full-width and half-width characters, that is, converting full-width numbers, letters, punctuation marks, and spaces in the recognition results into their corresponding half-width forms. Furthermore, this invention also cleans up spaces and zero-width characters in the text, removing invisible characters to improve data consistency and standardization.

[0044] S5, Using coordinate rules to process fields: This invention proposes that, based on the normalized coordinate system and according to the preset geometric position constraints corresponding to the invoice layout, each key information field is parsed and located from the recognition result; wherein, for multiple text blocks that are adjacent due to recognition splitting and whose coordinate spacing meets the merging condition, they are merged and regarded as the same field content.

[0045] The key information fields for analytical positioning based on geometric position constraints in a normalized coordinate system specifically include: Based on the conventional layout of the buyer's information on the ticket, the search scope is constrained to the left and upper regions of the normalized coordinate system. Based on the typical layout of the seller's information on the ticket, the search scope is constrained to the central region of the normalized coordinate system. For the amount and tax fields, the search is performed within a specified range around the coordinates of the identified specific keywords.

[0046] This invention further proposes that, based on the OCR recognition results and the standard layout of the VAT invoice, geometric position constraints for different fields are set in a normalized coordinate system, specifically including: (1) The geometric position of the buyer's name and the buyer's taxpayer identification number is located in the upper left part of the ticket, specifically within the left third of the ticket width, and the vertical position is located in the upper half of the ticket height. (2) The seller's name and the seller's taxpayer identification number are located in the upper or lower middle area of ​​the ticket, and the horizontal direction is between one-third and two-thirds of the width of the ticket. (3) The amount and tax fields are searched horizontally based on the “amount” column on the invoice, with the search range being ±2 times the width of the identified “amount” column. At the same time, the vertical search is performed based on the position of the keywords “total” or “total” or “total” in the split, with the range being ±1.5 times the height of the keyword.

[0047] The geometric constraints further include tolerance for the splitting of keywords or phrases during recognition. When the coordinate distance between the split text fragments within the same field is less than a set threshold, the split text fragments are automatically merged into a complete field. That is, the merging condition refers to the situation where, when the coordinate distance between multiple text blocks belonging to the same semantic field in the normalized coordinate system is less than a dynamic threshold set based on the invoice image size or typical character size, the recognized content of these text blocks is merged to form a complete field information.

[0048] To address the relative layout characteristics of VAT invoices in page design, this invention, based on extensive sample analysis, extracts multi-level geometric space rules to limit the search range of each field. Specifically, buyer information is generally located in the upper half of the page, slightly to the left, with x_max less than 2 / 3 of the page width, while seller information is usually located in the upper half of the page, slightly to the right, or the lower half, slightly to the left. Financial information, such as amounts and taxes, is concentrated in specific areas of the lower half of the page. Particularly in the extraction of amount and tax fields, this invention first identifies the column containing the keyword "amount" to determine its horizontal range on the page, and then further searches for corresponding numerical information within specific ranges to its left, right, top, and bottom. This positioning method significantly reduces the misalignment problems caused by traditional keyword-based matching.

[0049] Furthermore, considering that OCR may split keywords into multiple independent detection boxes in practical applications, such as recognizing "tax amount" as two separate text boxes for "tax" and "amount," this invention sets a text segmentation tolerance. That is, it allows these keywords to have a certain distance in the horizontal direction. As long as this distance does not exceed a pre-set threshold, they are still considered as continuous text content in the same field, thereby avoiding information extraction errors caused by keyword splitting.

[0050] S6, Perform field consistency verification: This invention proposes to perform numerical arithmetic consistency verification on the parsed and located amount, tax amount, and total price including tax fields, and automatically correct inconsistent fields based on the verification results. The automatic correction is performed when the sum of the amount and tax amount is not equal to the total price including tax and the difference exceeds the allowable error, by recalculating and updating the value of the amount or tax amount field based on the value of the total price including tax.

[0051] Preferably, the step of performing arithmetic consistency verification and automatic correction on the extracted amount, tax amount, and total price and tax fields includes: Calculate the sum of the amount and the tax, and compare it with the total price including tax. If the absolute difference between the sum of the above values ​​and the total value including tax exceeds a preset threshold, the total value including tax will be used as a benchmark to recalculate and correct the value of the amount or tax so that the three satisfy the arithmetic relationship.

[0052] The example performs an arithmetic consistency check on the amount, tax, and total price and tax values ​​parsed by OCR. When the difference between the sum of the amount and tax and the total price and tax exceeds a set threshold, the total price and tax value is used as the standard, and the fields are automatically corrected. The field consistency verification process further includes normalizing the full-width characters, half-width characters, and spaces in the taxpayer identification number to ensure that the decimal precision of the amount is two digits.

[0053] To further ensure consistency and accuracy of information across fields, this invention employs a field consistency verification mechanism. When the sum of the extracted amount excluding tax and the tax amount differs from the total price and tax (i.e., the price and tax total in lowercase) by a value greater than 0.01, this invention will use the total price and tax field as a benchmark and automatically correct the total amount excluding tax or the tax amount field through reverse calculation. This eliminates data inconsistencies that may be caused by OCR recognition errors or positioning mistakes.

[0054] Ultimately, it can output standardized VAT invoice fields, including invoice code, invoice number, invoice date, buyer's name and identification number, seller's name and identification number, seller's bank and account number, amount, tax amount, and total price including tax.

[0055] Through the implementation of the above process, this invention achieves the following key functions: First, by combining orientation detection and layout cropping preprocessing, problems such as image tilt, distortion, and cluttered backgrounds caused by shooting or scanning are effectively overcome. This preprocessing step ensures that the images processed subsequently have a unified orientation and prominent subjects, significantly reducing recognition errors caused by image deformation or irrelevant noise interference, and laying a reliable image foundation for subsequent high-precision information extraction.

[0056] Secondly, by introducing a geometric constraint parsing mechanism based on normalized coordinates, this method fully utilizes prior spatial layout knowledge of the VAT invoice format. This method not only relies on text semantics but also on the precise position of the text in the normalized coordinate system for field location and matching, thus fundamentally and effectively solving the field misalignment problem caused by keyword repetition in traditional methods. In particular, the included flexible merging mechanism can intelligently identify and merge adjacent text blocks that have been incorrectly split due to OCR recognition, ensuring the integrity and accuracy of the field content and improving the system's robustness to recognition noise.

[0057] Third, by implementing arithmetic consistency checks and automatic corrections between fields, a data error correction defense line is built at the logical level. The system can automatically check the mathematical relationship between amount, tax amount, and total price including tax, and automatically correct inconsistent results. This proactively detects and corrects potential OCR recognition errors or field matching errors before output, further ensuring the logical correctness and reliability of the final extracted information.

[0058] In summary, this invention, through three technological improvements—preprocessing optimization, spatial constraint parsing, and logical verification correction—forms a synergistically enhanced solution, demonstrating significant technical advantages in improving the accuracy, stability, and automation of VAT invoice information extraction.

[0059] To facilitate understanding of the technical effects of this invention, experiments were conducted using the methods described in the above embodiments. The key fields of each invoice included: invoice number, invoice date, issuer, total amount, total tax, total price including tax, seller's name, seller's taxpayer identification number, buyer's name, buyer's taxpayer identification number, and invoice details, totaling 11 fields. Statistics were compiled based on all fields of all experimental electronic invoices, and a comparative experiment was conducted on 500 invoices with a total of 5500 fields. When using a coordinate-based and deep learning-based method for electronic invoice field extraction, the field recognition rate (the proportion of correctly recognized fields out of the total number of fields) was 78.35%, and the recognition accuracy (the proportion of correctly recognized fields) was 91.43%. When using the previous OCR + regular expression matching method for electronic invoice field extraction, the field recognition rate was only 37.40%, and the recognition accuracy was 73.40% (this is because OCR cannot perform character concatenation during the recognition process, and the geometric relationships of tables are difficult for regular expressions to understand). Using all 5500 fields as the statistical base, the method of this invention can correctly extract approximately 3940 fields, while the traditional regular expression matching method can only correctly extract 1510 fields, representing an increase of approximately 1.6 times in the number of correct fields. The overall field accuracy rate (the proportion of correct fields across all fields), calculated by multiplying the field recognition rate by the accuracy, increases from approximately 27.5% in the traditional method to approximately 71.6% in the method of this invention, an improvement of approximately 44 percentage points. Therefore, it is evident that, under the premise of the same OCR recognition results, the electronic invoice field extraction method based on coordinate constraints and deep learning of this invention significantly improves field coverage and recognition accuracy compared to the traditional regular expression matching method, effectively reducing the probability of missed and false recognition of key fields, and enabling more reliable and stable automatic extraction of key information from electronic invoices.

[0060] In specific implementation, the method proposed in the technical solution of this invention can be automatically executed by those skilled in the art using computer software technology. System devices for implementing the method, such as computer-readable storage media storing the corresponding computer program of the technical solution of this invention and computer equipment including the computer program running the corresponding computer program, should also be within the protection scope of this invention.

[0061] It should be noted that the method and system described in this invention are designed for invoice image data with relatively fixed formats, and are particularly suitable for the automatic extraction and structured processing of information in scenarios involving value-added tax invoices, electronic invoices, and other forms of documentation. The description of this invention in the above embodiments is intended to illustrate the technical principles and implementation process of this invention and does not limit the scope of protection of this invention.

[0062] Those skilled in the art should understand that any substitutions, alterations, or equivalent adjustments made to the various steps, modules, algorithms, or system structures of this invention without departing from the core concept of this invention should be considered as included within the scope of protection of this invention. Specifically: Although PaddleOCR is a preferred example of the text recognition module described in this invention, it is not limited to this and can be replaced by other existing or future-developed OCR technologies or text detection and recognition models.

[0063] The orientation detection step described in this invention can be implemented using the Tesseract OSD module, or other orientation discrimination methods based on text line arrangement orientation analysis, character geometric features, machine learning or deep learning models, or a hardware acceleration module can be used to implement orientation detection and image rotation correction, all of which are within the protection scope of this invention.

[0064] The layout detection and cropping steps described in this invention can be implemented using a deep learning-based ticket detection model (such as PicoDet), a projection histogram analysis, an edge detection algorithm, a threshold segmentation algorithm, or a statistical analysis of the outer rectangle of the OCR detection box coordinates. Alternatively, a combination of the above methods can be used, all of which fall within the protection scope of this invention.

[0065] The method of normalizing the coordinates of the text detection results in this invention can be flexibly set according to the specific application scenario. It is not limited to linear scaling, but may also include nonlinear normalization, affine transformation, perspective correction or other geometric transformations, all of which are within the protection scope of this invention.

[0066] The text preprocessing steps described in this invention, such as full-width and half-width character conversion, invisible character removal, and space normalization, can be implemented by software programs, regular expressions, encoding conversion libraries, or deep learning-based text cleaning models, or by hardware circuits, without affecting the essence of this invention.

[0067] Although the step of parsing field coordinates based on geometric constraints described in this invention takes a specific area coordinate range (such as the buyer being located in the left 2 / 3, the amount area being located in the lower half of the page, etc.) as an example, it can also be flexibly set according to the invoice type, printing template, industry standards or user-defined rules, or automatically learn the spatial layout pattern through machine learning, and still fall within the protection scope of this invention.

[0068] The keyword segmentation tolerance processing described in this invention can be achieved by merging based on a distance threshold of text position, or by using semantic association, language models, or graph neural networks for context fusion and cross-frame information aggregation, which are alternative solutions to this invention.

[0069] In the consistency verification logic described in this invention, the judgment on the relationship between the tax amount, the total price excluding tax, and the total price including tax is not limited to a threshold of 0.01. It can also be adaptively adjusted according to the actual scenario or the threshold can be automatically determined by machine learning, which is still within the scope of this invention.

[0070] The present invention uniformly retains two decimal places for the amount field, and the precision of the number of decimal places can also be flexibly adjusted according to national laws and regulations, financial and tax business standards or user needs.

[0071] The method described in this invention can be implemented by software programs, hardware circuits, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), cloud computing services, or a combination of the above technologies.

[0072] Furthermore, the method or system of the present invention can exist in the form of a software product. If the method of the present invention is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. The storage medium may include, but is not limited to: read-only memory (ROM), random access memory (RAM), hard disk, optical disk, flash memory, USB flash drive, portable hard disk, magnetic disk, or any other medium capable of storing program code. The software product includes program instructions for executing the steps of the method of the present invention, causing a computer device (such as a personal computer, server, mobile device, embedded device, or network device, etc.) to execute all or part of the steps of the method of the present invention.

[0073] Furthermore, the logical division of labor and boundaries between the functional modules of the system described in this invention are merely logical divisions made to achieve the functions of this invention. Each functional module can be combined into fewer units according to actual application needs, or a single module can be split into multiple sub-modules, or it can be implemented in a physically distributed manner, or it can be implemented across multiple devices or across networks, all of which are equivalent alternatives to this invention.

[0074] Although the method steps of the present invention are described sequentially, it does not mean that each step must be performed in a strict order. Each step may also be performed in parallel or in another order, as can be understood by those skilled in the art, or some steps may be omitted without affecting the core innovative content of the present invention.

[0075] The following describes the intelligent VAT invoice information extraction electronic device based on coordinate constraints provided in the embodiments of the present invention. The intelligent VAT invoice information extraction electronic device based on coordinate constraints described below can be referred to in correspondence with the intelligent VAT invoice information extraction method based on coordinate constraints described above.

[0076] The electronic device may include a processor, a communications interface, memory, and a communication bus. The processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions from the memory to execute a coordinate-constrained intelligent extraction method for VAT invoice information, primarily including the software processing components described above.

[0077] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0078] In another embodiment, the present invention also provides a computer program product, the computer program product including a computer program, the computer program being stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute the software processing part of the coordinate constraint-based intelligent extraction method for value-added tax invoice information provided by the above methods.

[0079] In another embodiment, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the software processing portion of the coordinate-constrained intelligent extraction method for VAT invoice information provided by the above methods.

[0080] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0081] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligent extraction of value added tax invoice information based on coordinate constraints, characterized in that, The method comprises the following steps: The VAT invoice image to be identified is sequentially subjected to direction detection and correction, layout cutting preprocessing, and a main body image of the invoice is obtained, which is correct in direction and has cut off redundant background; An optical character recognition is performed on the main body image of the invoice, and an identification result containing text content and geometric coordinates thereof in a normalized coordinate system is output; Based on the normalized coordinate system, each key information field is parsed and positioned from the identification result according to a preset geometric position constraint condition corresponding to the layout of the invoice; wherein, for a plurality of text blocks in adjacent positions due to recognition splitting and satisfying a merging condition in terms of coordinate distance, the text blocks are merged to be regarded as the content of the same field; An arithmetic consistency check in the numerical level is performed on the amount, tax and total amount of price and tax fields parsed and positioned, and automatic correction is performed on the inconsistent fields based on the check result.

2. The method of claim 1, wherein: The direction detection and correction is performed by analyzing the overall arrangement direction of the text in the image, and the image is rotated and corrected to be in a standard reading direction.

3. The method of claim 1, wherein: The layout cutting is performed by detecting the boundary of the effective text area in the image, and a region image containing only the main body of the invoice is cut to eliminate edge blank and background interference.

4. The method of claim 1, wherein: The optical character recognition outputs the polygon region vertex coordinates corresponding to each recognized text block, and linearly maps all the coordinates to a normalized coordinate system in the range of zero to one.

5. The method of claim 1, wherein: The geometric position constraint condition based on the normalized coordinate system for parsing and positioning the key information field specifically comprises: The search area of the buyer's name and identification number is constrained in the upper half of the area in the horizontal direction and the vertical direction in the normalized coordinate system. The search area of the seller's name and identification number is constrained in the middle third area in the horizontal direction in the normalized coordinate system. For the amount and tax fields, the coordinates of the keywords "amount", "tax" or "total" are first positioned, and the corresponding numerical information is searched in the adjacent extended area in the horizontal and vertical directions.

6. The method of claim 1, wherein: The merging condition refers to when the coordinate distance of a plurality of text blocks belonging to the same semantic field in the normalized coordinate system is less than a dynamic threshold set according to the size of the invoice image or the typical size of the character, the recognition contents of the text blocks are spliced to form a complete field information.

7. The method of claim 1, wherein: The automatic correction is to recalculate and update the numerical value of the amount or tax field based on the total amount of price and tax when the sum of the amount and tax is not equal to the total amount of price and tax and the difference exceeds the allowed error.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: The processor executes the program to realize the coordinate constraint-based VAT invoice information intelligent extraction method according to any one of claims 1 to 7. 9.A non-transitory computer-readable storage medium having stored thereon a computer program. The computer program is executed by the processor to realize the coordinate constraint-based VAT invoice information intelligent extraction method according to any one of claims 1 to 7.

10. A computer program product comprising a computer program, characterized in that: The computer program is executed by the processor to realize the coordinate constraint-based VAT invoice information intelligent extraction method according to any one of claims 1 to 7.