Financial document processing method, device and electronic equipment

By monitoring updates to the financial business rule library, dynamically reconstructing financial document recognition templates and combining them with cross-reference verification, the problem of rigid recognition logic in financial document processing is solved, achieving efficient and accurate financial data processing.

CN120564213BActive Publication Date: 2025-09-26SUMEC TEXTILE TECH & IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511061691.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-09-26
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Existing OCR technology is unable to flexibly adapt to the frequent and subtle changes in financial rules and document formats in financial document processing, resulting in the rigidification of recognition logic and disconnection from business rules, leading to recognition errors or omissions of key information.

Method used

By monitoring the version updates of the financial business rule library, obtaining the changed target business rules and difference area descriptors, dynamically reconstructing the financial document recognition template, and combining cross-reference relationship verification, adaptive generation and iterative optimization of the template are achieved to ensure recognition accuracy and flexibility.

Benefits of technology

It improves the automation and accuracy of financial document processing, reduces the time cost of template construction and updating, ensures data logical consistency and fault tolerance, and improves processing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564213B_ABST
    Figure CN120564213B_ABST
Patent Text Reader

Abstract

A method, device, and electronic device for processing financial documents relate to the field of data processing. In this method, an original image of a financial document to be processed is obtained; the changed target business rules and key fields and difference area descriptors related to the target business rules are obtained; a corresponding target template is matched according to the key fields; the target template is reconstructed according to the target business rules and difference area descriptors to obtain a first reconstructed template; the financial document is identified based on the first reconstructed template to obtain initial financial data; it is determined whether the initial financial data satisfies a preset cross-reference relationship table; if so, the initial financial data is output; if not, the first reconstructed template is revised twice to obtain a second reconstructed template and corresponding target financial data, and the target financial data is output. Implementing the technical solution provided by this application improves the accuracy of document processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and specifically to a method, device and electronic equipment for processing financial documents. Background Art

[0002] Processing large volumes of documents, such as invoices, expense reports, receipts, and bank slips, is a core and arduous task in financial workflows. Traditional methods rely on manual document entry, verification, and archiving, which is not only inefficient and costly, but also prone to human error in repetitive operations. This makes it difficult to ensure accuracy and timeliness, especially when processing documents from multiple sources, heterogeneous structures, and complex formats.

[0003] To overcome these shortcomings, existing technologies employ automated solutions based on optical character recognition (OCR). These solutions capture document images using an image acquisition device and employ an OCR engine to identify key textual information within the images, such as the amount, date, supplier name, and tax ID number. These solutions then attempt to automatically populate or associate the recognition results with corresponding fields in the financial system, thereby reducing manual intervention and improving processing speed and accuracy. This approach, to a certain extent, alleviates the efficiency and error rates associated with purely manual processing.

[0004] However, in the process of processing financial documents, existing OCR technology solutions still face problems. They are difficult to flexibly adapt to the frequent and subtle changes in financial rules and document formats, resulting in the rigidification of recognition logic and disconnection from business rules. There are many types of documents in the financial field, and their layout structure, key information location, printing specifications, such as invoice supervision stamps, tax classification codes, etc., and related financial and tax rules, such as tax rate changes and reimbursement policy adjustments, are often updated. Existing OCR solutions usually rely on pre-set fixed templates or rule libraries for information positioning and recognition. Once the document format or associated rules change unexpectedly, for example, a new version of the invoice is activated, the position of a specific field is adjusted, or new mandatory information items are added, it is often unable to automatically perceive and adjust the recognition logic, and it is very easy to have large-scale recognition errors or omissions of key information. Time-consuming manual template updates and model retraining must be carried out, which can easily cause a decrease in recognition accuracy during the transition period. Summary of the Invention

[0005] The present application provides a financial document processing method, device and electronic equipment, which improve the accuracy of document processing.

[0006] In a first aspect of the present application, a financial document processing method is provided, the method comprising: obtaining an original image of a financial document to be processed; monitoring the version update status of a financial business rule library, and when a rule change instruction that triggers a preset change condition is detected, obtaining a changed target business rule and key fields and difference area descriptors related to the target business rule, the key fields including changed fields and dependent fields associated with the changed fields in a preset cross-reference relationship table, and the difference area descriptors including field position offset information and newly added field offset information; matching a corresponding target template in a preset reference template library based on the key fields; reconstructing the target template based on the target business rule and the difference area descriptors to obtain a first reconstructed template; identifying the financial document based on the first reconstructed template to obtain initial financial data; determining whether the initial financial data satisfies the preset cross-reference relationship table; if it is determined that the initial financial data satisfies the preset cross-reference relationship table, outputting the initial financial data; if it is determined that the initial financial data does not satisfy the preset cross-reference relationship table, performing a secondary revision on the first reconstructed template to obtain a second reconstructed template and corresponding target financial data, and outputting the target financial data.

[0007] By employing the above technical solution, the original image of the financial document to be processed is obtained, providing a structured data foundation for subsequent template matching and data extraction. By monitoring the version update status of the financial business rule library, the modified target business rules and key fields are obtained in real time, ensuring that the financial document processing method can dynamically adapt to changes in business rules, improving flexibility and adaptability, and reducing the manual correction costs associated with rule changes. The target template is reconstructed using the target business rules and difference area descriptors to obtain a first reconstructed template that matches the current financial document format and business rules. This enables the adaptive generation of financial document recognition templates, reduces the time cost of template construction and updating, and improves the automation level of financial document processing. To prevent the impact of newly added fields on cross-reference relationships, financial documents are identified based on the first reconstructed template, obtaining initial financial data and determining whether it satisfies the preset cross-reference relationship table. By introducing cross-reference relationship verification, preliminary data quality checks are performed simultaneously with data extraction, improving the accuracy and reliability of financial data. For situations where the preset cross-reference table is not met, the first reconstructed template is revised a second time, relocating and identifying key fields within the discrepant area. The template's matching accuracy is then iteratively optimized until the target financial data that satisfies the cross-reference relationship is output. This further improves the fault tolerance of financial document processing and ensures the high quality and consistency of the output data. Through adaptive template reconstruction and cross-reference relationship verification, this technical solution enables intelligent and automated processing of financial documents, significantly improving the efficiency and accuracy of financial document processing.

[0008] Optionally, the target template is reconstructed according to the target business rule and the difference area descriptor to obtain a first reconstructed template, specifically including: calculating the coordinate offset vector of each field according to the field position offset information, and superimposing the coordinate offset vector to the positioning box parameter of the corresponding field in the target template to update the position of the field positioning box; according to the newly added field offset information, combined with the semantic constraints related to the newly added field in the target business rule, determining the candidate positioning area of ​​the newly added field through context analysis, and extracting local features and identifying table lines within the candidate positioning area to generate the positioning box of the newly added field; based on the newly added field verification rules in the target business rules, constructing the association verification logic between fields, and adding the association verification logic to the verification rule set of the target template; integrating the field positioning box, the positioning box of the newly added field and the association verification logic into the target template to generate the first reconstructed template.

[0009] By adopting the above technical solution, by calculating the coordinate offset vector of each field in the difference area descriptor and superimposing it on the positioning box parameters of the corresponding field of the target template, the position of the field positioning box can be quickly updated without redesigning the entire template layout, thereby reducing the computational complexity. For newly added fields, their candidate positioning areas are first determined based on the semantic constraints in the target business rules, and then an accurate positioning box is generated within the area through local feature extraction and table line recognition, thereby maximizing the reuse of the original template layout. At the same time, the verification rules of the newly added fields in the target business rules are merged with the verification rule set of the original template to form a complete associated verification logic. Finally, the updated positioning box, the newly added field positioning box and the associated verification logic are integrated into the target template to obtain the first reconstructed template that can quickly respond to changes in business rules, thereby improving the recognition efficiency of financial documents.

[0010] Optionally, the second correction of the first reconstruction template to obtain the second reconstruction template and the corresponding target financial data specifically includes: determining the abnormal fields in the initial financial data that do not satisfy the cross-reference relationship according to the preset cross-reference relationship table; locating the abnormal area corresponding to the abnormal field in the first reconstruction template, and obtaining the position coordinates of the abnormal area in the financial document image; based on the position coordinates, collecting the context area adjacent to the abnormal area in the financial document image, and performing extended recognition on the context area to obtain an extended recognition result; judging whether there is any influence on the abnormal area according to the extended recognition result. Interference factors affecting field recognition accuracy, the interference factors including field overlap, field breakage, and field tilt; if the interference factors exist, the recognition result of the abnormal area is corrected according to the preset compensation rules to obtain corrected abnormal field data; the corrected abnormal field data is updated to the initial financial data to obtain intermediate financial data; it is determined whether the intermediate financial data satisfies the preset cross-reference relationship table; if the intermediate financial data satisfies the preset cross-reference relationship table, the intermediate financial data is determined as the target financial data, and the first reconstruction template is adjusted according to the extended recognition result to obtain a second reconstruction template.

[0011] By employing this technical solution, anomalous fields in the initial financial data are first identified based on a preset cross-reference table. The anomalous region is located in a first reconstruction template, and its coordinates in the original document image are obtained. Contextual areas adjacent to the anomalous region are then collected for extended recognition to identify interfering factors that may affect the accuracy of anomalous field recognition, such as field overlap, field breakage, and field tilt. This analysis of the localized anomalous region within the overall context can uncover the underlying causes of recognition errors. If interfering factors are present, the anomalous field data is corrected according to preset compensation rules and updated to the initial financial data to generate intermediate financial data. The intermediate financial data is then determined to meet the preset cross-reference table requirements. If so, it is identified as the target financial data, and the first reconstruction template is adjusted based on the extended recognition results to generate a second reconstruction template. If not, the next round of corrections is initiated. This iteratively optimized recognition strategy continuously improves the quality of financial data until it fully complies with business rules.

[0012] Optionally, after determining whether the intermediate financial data satisfies the preset cross-reference relationship table, the method further includes: if the intermediate financial data does not satisfy the preset cross-reference relationship table, identifying the conflicting fields in the intermediate financial data that do not satisfy the preset cross-reference relationship table, and extracting the conflict coordinate area corresponding to the conflicting fields in the first reconstruction template; calculating the theoretical coordinate position of the conflicting fields according to the constraint rules related to the conflicting fields in the preset cross-reference relationship table, and correcting the positioning box of the conflicting fields in the first reconstruction template to the theoretical coordinate position; within the conflict coordinate area, identifying newly added fields through text semantic analysis, and adding the newly added fields to the positioning rules of the first reconstruction template to generate the second reconstruction template, and identifying the financial documents according to the second reconstruction template to obtain the target financial data.

[0013] By adopting the above technical solution, the introduction of new fields or adjustments to existing fields may disrupt the original data association logic, causing the intermediate financial data to fail to meet the constraints of the preset cross-reference table. When the intermediate financial data still fails to meet the preset cross-reference table, the above technical solution first identifies the conflicting fields and extracts the corresponding conflicting coordinate areas, providing a positioning basis for precise error correction. Next, based on the constraints associated with the conflicting fields in the preset cross-reference table, the theoretical coordinate positions of the conflicting fields are calculated, and the positioning boxes of the conflicting fields in the first reconstruction template are corrected to the theoretical positions. This strategy, combining theoretical and actual coordinates, achieves precise adjustment of local areas while maintaining the original template layout, minimizing the impact on the overall structure. Furthermore, considering that field adjustments may introduce new associated fields, this method uses text semantic analysis technology within the conflicting coordinate areas to identify potential new fields and add them to the positioning rules of the first reconstruction template, forming a more complete second reconstruction template. Re-identifying the financial documents based on the second reconstruction template can obtain the target financial data that meets the constraints of the preset cross-reference table. Through continuous iterative optimization of templates and data, the recognition and conversion of financial documents is ultimately achieved.

[0014] Optionally, when a rule change instruction that triggers a preset change condition is detected, the changed target business rule and key fields and difference area descriptors related to the target business rule are obtained, specifically including: detecting the version number of the financial business rule library according to a preset time interval; comparing the detected version number with the locally stored version number to obtain a version number difference value; if the version number difference value is greater than a preset threshold, determining that the version of the financial business rule library has been updated, and triggering the rule change instruction of the preset change condition; when the rule change instruction is detected, obtaining the change log file of the financial business rule library, identifying the change record, and extracting the change record to obtain the changed target business rule, the target business rule including the change content and the change rule; performing semantic analysis on the change content to determine the key fields related to the change rule; performing semantic analysis on the change rule to determine the field position offset information of the key field and the newly added field offset information.

[0015] By adopting the above technical solution, the version number of the financial business rule library is detected according to a preset time interval, and the detected version number is compared with the locally stored version number to obtain a version number difference value. If the version number difference value is greater than a preset threshold, it is determined that the version of the financial business rule library has been updated, and a rule change instruction with a preset change condition is triggered. When a rule change instruction is detected, the change log file of the financial business rule library is obtained, the change record is identified, and the change record is extracted to obtain the changed target business rule, which includes the change content and the change rule. The change content is semantically parsed to determine the key fields related to the change rule; the change rule is semantically parsed to determine the field position offset information of the key field and the new field offset information. This method can automatically detect version changes of the financial business rule library. When the version change reaches a certain level, the changed target business rule, key fields, and difference area descriptors are obtained in a timely manner to provide necessary information for subsequent template reconstruction and difference adaptation. Through automatic detection and acquisition, the tedious process of manually monitoring the rule base version and manually extracting change information is avoided, which improves the timeliness and accuracy of the response to rule changes, ensuring that the financial document recognition process can keep pace with business rule changes and adapt to the latest business needs.

[0016] Optionally, the method further includes: using the target template to identify the original image to obtain basic financial data; based on the target template, performing targeted training on the key fields and the difference area descriptors to generate a difference adaptation model; applying the difference adaptation model to the original image of the financial document, performing targeted identification on the original image, and obtaining a difference identification result; fusing the difference identification result with the basic financial data, updating the changed fields and newly added fields in the basic financial data, and obtaining the initial financial data.

[0017] By employing the above technical solution, the target template is used to identify the original image to obtain basic financial data. This fully leverages the target template's ability to recognize general documents and obtains preliminary structured financial information. Based on the target template, targeted training is performed on key fields and difference region descriptors to generate a difference adaptation model. This specifically improves the difference adaptation model's adaptability to changed and newly added fields, compensating for the shortcomings of the target template. The difference adaptation model is applied to the original image of the financial document for targeted recognition, generating difference recognition results and obtaining specialized identification information for the changed and newly added fields. Finally, the difference recognition results are fused with the basic financial data, updating the changed and newly added fields in the basic financial data to obtain the initial financial data, effectively integrating the difference information with the original information. This phased, task-based recognition and fusion approach fully leverages the general capabilities of the target template and the specialized capabilities of the difference adaptation model. Without losing the original information, the difference regions are efficiently and accurately identified, ultimately resulting in complete, high-quality initial financial data, laying a solid foundation for subsequent financial data cross-checking and processing.

[0018] Optionally, based on the target template, targeted training is performed on the key fields and the difference area descriptors to generate a difference adaptation model, specifically including: extracting corresponding difference image areas in the original image of the financial document according to the key fields and the difference area descriptors to generate a difference training set; loading the pre-trained OCR model corresponding to the target template as the basic model, and freezing the general feature extraction layer parameters in the basic model that are not related to text recognition, decoupling the output layer parameters related to field positioning and the rule layer parameters related to business verification logic in the basic model to obtain a decoupled model; based on the difference training set, the decoupling model is targeted fine-tuned to generate the difference adaptation model.

[0019] By adopting the above technical solution, the corresponding difference image regions are extracted from the original images of financial documents based on key fields and difference region descriptors, generating a difference training set, which provides targeted data support for the training of the difference adaptation model. The pre-trained OCR model corresponding to the target template is loaded as the base model, fully leveraging the pre-trained model's existing capabilities in text recognition and avoiding the huge overhead of training from scratch. The general feature extraction layer parameters unrelated to text recognition in the base model are frozen, preserving the model's ability to extract general document features. The output layer parameters related to field positioning and the rule layer parameters related to business verification logic in the base model are decoupled, allowing them to be optimized separately and improving the flexibility of the model's adaptation to differences. Based on the difference training set, the decoupled model is fine-tuned. By focusing on the difference regions and conducting targeted training on the output layer and rule layer, the model quickly learns the characteristics of key field changes and newly added fields, resulting in a high-quality difference adaptation model and increasing recognition efficiency. This targeted training method, based on differential data and model decoupling, minimizes model adjustments without destroying the generalization capabilities of the original model, enabling it to efficiently adapt to the differences brought about by rule changes and improving the accuracy of the model in identifying changed bills.

[0020] Optionally, based on the difference training set, the decoupling model is subjected to targeted fine-tuning to generate the difference adaptation model, specifically including: inputting the difference training set into the decoupling model to obtain the output layer feature map and the rule layer feature map of the decoupling model; calculating the target loss in the output layer feature map according to the field position offset information and the newly added field offset information, the target loss including the changed field positioning loss and the newly added field positioning loss, the changed field positioning loss is calculated based on the coordinate offset vector of the changed field, and the newly added field positioning loss is calculated based on the local features and table lines within the candidate positioning area; calculating the association loss in the rule layer feature map according to the newly added field verification rule in the target business rule, the association loss is calculated based on the semantic dependency relationship between the newly added field and the associated field; constructing a joint loss function, the joint loss function including the target loss and the association loss, and performing gradient optimization and updating on the output layer parameters and the rule layer parameters of the decoupling model based on the joint loss function to obtain the difference adaptation model.

[0021] By adopting the above technical solution, the differential training set is input into the decoupled model, obtaining the output layer feature map and the rule layer feature map of the decoupled model, which provide the necessary intermediate representation for calculating the target loss and association loss. Based on the field position offset information and the newly added field offset information, the changed field positioning loss and the newly added field positioning loss are calculated in the output layer feature map. The changed field positioning loss is calculated based on the coordinate offset vector of the changed field, enabling the output layer to accurately locate the new position of the changed field. The newly added field positioning loss is calculated based on the local features and table lines within the candidate positioning area, enabling the output layer to predict the newly added field at the correct location. Based on the newly added field verification rules in the target business rules, the association loss is calculated in the rule layer feature map, enabling the rule layer to learn the semantic dependencies between the newly added fields and the associated fields. A joint loss function combining the target loss and the association loss is constructed, and gradient optimization and update of the output layer parameters and the rule layer parameters of the decoupled model are performed based on this loss function, ultimately resulting in a differential adaptation model. This joint training method, which integrates multiple losses, quantifies the degree of model adaptation differences at the level of field location and association verification. Through end-to-end backpropagation learning, the model reduces the difference loss while establishing dependencies between newly added fields and associated fields, resulting in a high degree of adaptation to different regions. Guided by the loss function, the parameters of the output and rule layers are optimally updated, forming a robust and generalizable difference adaptation model, providing a strong guarantee for the accurate identification of altered bills.

[0022] In a second aspect of the present application, a financial document processing device is provided, which includes an original image acquisition module, a difference determination module, a target template matching module, a template reconstruction module, a cross-reference relationship processing module and a secondary reconstruction module, wherein: the original image acquisition module is used to obtain the original image of the financial document to be processed; the difference determination module is used to monitor the version update status of the financial business rule library, and when a rule change instruction that triggers a preset change condition is detected, the changed target business rule and the key fields and difference area descriptors related to the target business rule are obtained, the key fields include the changed fields and the dependent fields associated with the changed fields in the preset cross-reference relationship table, and the difference area descriptors include field position offset information and newly added field offset information; the target template matching module is used to, based on the key fields, A corresponding target template is matched in a preset benchmark template library; the template reconstruction module is used to reconstruct the target template according to the target business rules and the difference area descriptor to obtain a first reconstructed template; the template reconstruction module is also used to identify the financial document based on the first reconstructed template to obtain initial financial data; the cross-reference relationship processing module is used to determine whether the initial financial data satisfies the preset cross-reference relationship table; the cross-reference relationship processing module is also used to output the initial financial data if it is determined that the initial financial data satisfies the preset cross-reference relationship table; the secondary reconstruction module is used to perform secondary correction on the first reconstructed template if it is determined that the initial financial data does not satisfy the preset cross-reference relationship table, to obtain a second reconstructed template and corresponding target financial data, and output the target financial data.

[0023] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs any of the methods described above.

[0024] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions. When the instructions are executed, any one of the methods described above is executed.

[0025] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0026] 1. By dynamically monitoring financial rule updates, generating difference area descriptors related to changed fields, and combining template reconstruction and cross-reference verification, it achieves flexible adaptation to changes in financial document formats and adjustments to business rules. Through automated template adjustment, intelligent verification, and iterative correction, it overcomes the inability of traditional OCR to quickly respond to document format and rule changes, significantly improving recognition accuracy and processing efficiency, reducing manual maintenance costs, and ensuring data logical consistency, thereby mitigating business risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a flowchart of a financial document processing method disclosed in an embodiment of the present application;

[0028] Figure 2 This is another flowchart of a financial document processing method disclosed in an embodiment of the present application;

[0029] Figure 3 This is a module diagram of a financial document processing device disclosed in an embodiment of the present application;

[0030] Figure 4 This is a structural diagram of an electronic device disclosed in an embodiment of the present application.

[0031] Explanation of the accompanying drawings: 301, original image acquisition module; 302, difference determination module; 303, target template matching module; 304, template reconstruction module; 305, cross-reference relationship processing module; 306, secondary reconstruction module; 400, electronic device; 401, processor; 402, communication bus; 403, user interface; 404, network interface; 405, memory. DETAILED DESCRIPTION

[0032] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0033] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.

[0034] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0035] This application provides a financial document processing, refer to Figure 1 , Figure 1 This is a flow diagram of a financial document processing process provided by an embodiment of the present application. The method is applied to a server, which is used to execute a financial document processing program. The server can be a single server, a server cluster consisting of multiple servers, or a cloud computing service center. The server can communicate with the user device via a wired or wireless network. The method includes steps S101 to S108, which are as follows:

[0036] Step S101: Acquire the original image of the financial document to be processed.

[0037] In step S101, the server obtains the original image of the financial document to be processed by communicating with the user device. Specifically, the user can use various types of user devices, such as smart phones, tablet computers, laptops or desktop computers, to establish a connection with the server via a wired or wireless network.

[0038] The user's device can be installed with a dedicated financial document processing client application, or access the web application interface provided by the server through a browser. Regardless of the method used, the user can upload the original image of the financial document to be processed to the server.

[0039] For example, users can use their smartphone's built-in camera to capture paper financial documents, such as invoices and expense reports, and then upload the captured original image to the server through the client application. Alternatively, users can use a scanner to scan paper financial documents into electronic image files, such as JPG and PNG, and then upload the image files to the server through a web browser.

[0040] After receiving the original images uploaded by users, the server performs necessary preprocessing, such as image format conversion and image quality testing, to ensure smooth subsequent processing. For example, the server can convert various uploaded image files into a standardized, easy-to-process format, such as BMP or TIFF. To improve the efficiency and convenience of users uploading original images of financial documents, the server can also provide auxiliary functions. For example, the server can support batch upload, allowing users to select multiple financial document images for upload at once. The server can also provide image preview and editing functions, allowing users to easily check and adjust images before uploading.

[0041] Step S102: Monitor the version update status of the financial business rule library. When a rule change instruction that triggers a preset change condition is detected, obtain the changed target business rule and the key fields and difference area descriptors related to the target business rule. The key fields include the changed fields and the dependent fields associated with the changed fields in the preset cross-reference relationship table. The difference area descriptors include field position offset information and new field offset information.

[0042] In step S102, when a rule change instruction that triggers a preset change condition is detected, the changed target business rule and the key fields and difference area descriptors related to the target business rule are obtained, specifically including: detecting the version number of the financial business rule library according to a preset time interval; comparing the detected version number with the locally stored version number to obtain a version number difference value; if the version number difference value is greater than a preset threshold, it is determined that the version of the financial business rule library has been updated, and the rule change instruction of the preset change condition is triggered; when a rule change instruction is detected, the change log file of the financial business rule library is obtained, the change record is identified, and the change record is extracted to obtain the changed target business rule, which includes the change content and the change rule; performing semantic analysis on the change content to determine the key fields related to the change rule; performing semantic analysis on the change rule to determine the field position offset information of the key field and the new field offset information.

[0043] Specifically, the server pre-sets a time interval, such as 2:00 AM daily, to periodically check the version number of the financial business rule library. The server accesses the version management system of the financial business rule library, obtains the latest version number, and compares it with the locally stored version number. The server then calculates the difference between the two version numbers, which is known as the version number difference value.

[0044] If the version number difference is greater than a preset threshold (e.g., 1), it indicates that the version of the financial business rule library has been updated. The server promptly obtains and processes the changes. At this point, the server determines that the version of the financial business rule library has been updated and triggers a rule change instruction based on the preset change conditions.

[0045] When the server detects a rule change instruction, it begins to retrieve the changed target business rule, along with the associated key fields and difference area descriptors. First, the server accesses the change management module of the financial business rule library to obtain the change log file. This change log file records all changes made to the financial business rule library from the previous version to the current version. The server parses the change log file and identifies the change records. Each change record corresponds to a change to a business rule, including the change content and the change rule. The server extracts the information from the change log and obtains the changed target business rule.

[0046] Next, the server performs semantic analysis on the changes in the target business rules to identify key fields relevant to the changed rules. Key fields consist of two parts: the changed fields, which are the fields affected by the business rule changes; and the dependent fields associated with the changed fields in the pre-set cross-reference table, which are other fields with cross-reference relationships. Cross-reference relationships refer to logical dependencies or constraints between fields, such as Amount field A must equal Quantity field B multiplied by Unit price field C. These relationships are typically defined and maintained in the pre-set cross-reference table. For example, suppose a new version of the invoice processing rule changes the value range of the "Tax Rate" field and adds a new "Amount Excluding Tax" field. "Tax Rate" is a changed field, while fields with cross-reference relationships, such as "Amount Including Tax" and "Tax Amount," are the corresponding dependent fields. After obtaining the target business rules, the server automatically parses these key fields and focuses on them for subsequent document identification and review. The server identifies the changed fields by analyzing keywords and semantic dependencies in the changed content and determines the dependent fields based on the pre-set cross-reference table.

[0047] Finally, the server performs semantic parsing on the change rules within the target business rules to determine the difference area descriptors for the key fields. These difference area descriptors include field position offset information and newly added field offset information, describing the positional changes of key fields within the financial document image. For example, if a change rule includes "Move the invoice number field 50 pixels to the right," the server will parse this rule and determine the position offset of the invoice number field to be "50 pixels right." Alternatively, if a change rule includes "Add a tax rate field below the invoice amount," the server will parse this rule and determine the offset of the newly added tax rate field to be "below" the invoice amount field.

[0048] Step S103: Match the corresponding target template in the preset reference template library according to the key field.

[0049] In step S103, the server first loads a pre-defined reference template library, which contains a large number of standardized document templates, each with its own unique structure and field layout. These templates are created through manual design or automatic generation, covering a variety of common financial document types, such as invoices, receipts, and expense reports.

[0050] Next, the server extracts key information from the key fields, such as field name and field type, to form a set of query conditions. The server then inputs this set of query conditions into the search engine of the preset benchmark template library, which then searches for matches in the template library based on the query conditions.

[0051] The matching search process is divided into two stages: rough matching and exact matching. In the rough matching stage, the search engine quickly scans each template in the template library, screening candidate templates whose field names are identical or similar to those in the query. Similarity can be measured using metrics such as string edit distance and semantic similarity. For example, if the query includes the "invoice number" field, the search engine will find all templates containing fields such as "invoice number," "invoice number," and "document number" as candidate templates.

[0052] During the precise matching phase, the search engine performs a more granular comparison and scoring of candidate templates. Specifically, the search engine extracts the fields in the candidate template that correspond to the query criteria and compares them one by one with the fields in the query criteria. This comparison covers multiple dimensions, including field name, field type, field length, and field position. The search engine calculates a similarity score for each dimension and combines these scores to determine the overall match of the candidate template. The higher the match, the closer the candidate template matches the query criteria and the more likely it is the target template.

[0053] After two phases of matching search, the search engine obtains a list of candidate templates sorted by matching degree. The server selects the candidate template with the highest matching degree as the target template for subsequent financial document recognition and data extraction.

[0054] For example, assume that key fields include "invoice code," "invoice number," "invoice date," "purchaser name," and "purchaser taxpayer identification number." The server uses these fields as query criteria and searches for matches within a pre-set base template library. During the rough matching phase, the search engine finds multiple candidate templates containing these fields, such as sales invoice templates and purchase invoice templates. During the precise matching phase, the search engine finds that the field name, field type, and field position in the sales invoice template most closely match the query criteria, and therefore returns it as the target template to the server.

[0055] Step S104: reconstructing the target template according to the target business rules and the difference region descriptor to obtain a first reconstructed template.

[0056] In step S104, the target template is reconstructed according to the target business rules and the difference area descriptor to obtain a first reconstructed template, which specifically includes: calculating the coordinate offset vector of each field according to the field position offset information, and superimposing the coordinate offset vector to the positioning box parameter of the corresponding field in the target template to update the position of the field positioning box; according to the newly added field offset information, combined with the semantic constraints related to the newly added field in the target business rules, determining the candidate positioning area of ​​the newly added field through context analysis, extracting local features and identifying table lines in the candidate positioning area to generate a positioning box for the newly added field; based on the newly added field verification rules in the target business rules, constructing the association verification logic between fields, and adding the association verification logic to the verification rule set of the target template; integrating the field positioning box, the positioning box of the newly added field and the association verification logic into the target template to generate a first reconstructed template.

[0057] Specifically, the server reconstructs the target template based on the target business rules and the difference area descriptors, generating a first reconstructed template. This reconstruction process primarily involves three aspects: updating the positioning boxes of existing fields, generating positioning boxes for newly added fields, and establishing correlation verification logic between fields. By comprehensively processing these three aspects, the server can generate a new template that matches the actual document and complies with the new business rules.

[0058] First, the server calculates the coordinate offset vector for each field based on the field position offset information in the difference area descriptor. The coordinate offset vector represents the horizontal and vertical displacement of the field and can be expressed in the form of (dx, dy). The server updates the position of the field positioning box by superimposing the offset vector on the positioning box parameters of the corresponding field in the target template. The positioning box is usually described by the coordinates of the upper left corner and the lower right corner, so the process of updating the positioning box is to add the offset vector to these two coordinates respectively. For example, if the original positioning box of a field is ((100, 200), (300, 250)) and the offset vector is (10, 20), the updated positioning box is ((110, 220), (310, 270)).

[0059] Next, the server generates a positioning box for the newly added field recorded in the difference area descriptor. Since the newly added field does not exist in the original template, its location information cannot be directly obtained. To this end, the server first determines candidate positioning areas for the newly added field through contextual analysis based on the semantic constraints associated with the newly added field in the target business rules. Semantic constraints may include the positional relationship of other fields related to the newly added field, as well as the data type and format of the newly added field. For example, if the newly added field is "Tax Rate" and the business rules require that "Tax Rate" be located below the "Amount" field, the server may consider the area below the positioning box of the "Amount" field as a candidate positioning area for "Tax Rate." After determining the candidate positioning area, the server extracts local features within the area and identifies table lines to further determine the specific positioning box for the newly added field. Local features may include text density and font size within the area, while table lines provide a boundary reference for the newly added field. After comprehensively analyzing these features, the server generates a rectangular box within the candidate area, which serves as the positioning box for the newly added field.

[0060] Finally, based on the newly added field validation rules in the target business rules, the server constructs inter-field validation logic and adds it to the validation rule set of the target template. This validation logic verifies data dependencies between different fields, such as ensuring that "Total Amount" equals "Unit Price" multiplied by "Quantity." This validation logic, based on the original template, must be expanded to incorporate the semantic constraints of the newly added fields. For example, if the value of the "Tax Rate" field is constrained to fall within the range of 10% to 16%, corresponding conditional logic must be added to the validation rules. By analyzing the field association conditions in the target business rules, the server constructs a complete set of validation logic covering the newly added fields and integrates it into the existing validation rule set of the target template.

[0061] After completing the above three aspects of processing, the server integrates the updated field positioning box, the positioning box of the newly added field, and the associated verification logic into the target template, and finally generates a first reconstructed template.

[0062] For example, suppose a new "Tax Inclusive Flag" field is added to an invoice template. Business rules require that the "Tax Inclusive Flag" appear after the "Tax Rate" field, with the combination of the two representing tax attributes. The server first analyzes the position of the "Tax Rate" field and determines that the candidate location for the "Tax Inclusive Flag" is to its right. The server then identifies table lines within this area and, based on characteristics such as field length and text density, generates a positioning box for the "Tax Inclusive Flag." Next, based on business rules, the server establishes validation logic linking the "Tax Inclusive Flag" and "Tax Rate," stipulating that their combined values ​​can only be specific forms, such as "Tax Inclusive - 13%" or "Tax Exclusive - 0%." Finally, the server inserts the "Tax Inclusive Flag" positioning box into the corresponding position in the template and adds the associated validation logic to the template's validation rules, creating the first reconstructed template.

[0063] In one possible implementation, the target template is reconstructed according to the target business rule and the difference area descriptor to obtain a first reconstructed template. It can also be as follows: the target business rule is parsed into a rule tree structure, the rule tree structure includes multiple nodes, each node represents a field, and the connection relationship between the nodes represents the association relationship between the fields; the field node that needs to be adjusted is determined according to the difference area descriptor, and the position type of the field node in the rule tree structure is judged; if the field node is a leaf node, the position information of the field node is updated to be the same as the position offset of the corresponding field in the difference area descriptor; if the field node is a non-leaf node, the child node directly connected to the field node is found in the rule tree structure, the relative position of the child node is adjusted according to the field position offset in the difference area descriptor, and the position information of the field node is updated; if the difference area descriptor contains a new field, a new node corresponding to the new field is added to the rule tree structure according to the preset field association rule, and the position information of the new node is determined according to the difference area descriptor; based on the adjusted rule tree structure, the structural layout of the target template is updated to obtain the first reconstructed template.

[0064] Specifically, the server first parses the target business rule into a rule tree structure. This structure, similar to a file directory tree, consists of multiple nodes, each representing a field such as the invoice number or seller name. The connections between nodes represent the relationships between fields, such as subordination and dependency. For example, the "Price and Tax Total" field node might be the parent node of the "Goods Name" and "Amount" field nodes, indicating that the value of "Price and Tax Total" depends on the values ​​of its subordinate fields. The rule tree structure intuitively depicts the logical relationship network between the fields in the target business rule.

[0065] Next, the server determines the field nodes that need adjustment based on the difference area descriptor. Since the difference area descriptor records the structural differences between the target template and the actual document, the server can match the field name, position offset, and other information in the descriptor to identify the affected field nodes in the rule tree and adjust them as needed.

[0066] For field nodes that require adjustment, the server first determines their position type in the rule tree. If the node is a leaf node (i.e., has no child nodes), its position information is directly updated to match the position offset of the corresponding field in the difference area descriptor. This means that the position of the leaf node field in the template will directly inherit the position in the actual document to reflect the actual situation.

[0067] If the node to be adjusted is a non-leaf node, that is, it has subordinate nodes, the adjustment process is relatively complicated. The server first finds the child nodes directly connected to the node in the rule tree, and then adjusts the relative positions of these child nodes according to the field position offset recorded in the difference area descriptor. For example, if the descriptor shows that a child node field is offset 10 pixels to the right compared to the target template, the server will adjust the horizontal position coordinates of the child node accordingly. After adjusting the positions of all child nodes, the server also needs to update the position information of the parent node, that is, the non-leaf node itself, to maintain the relative position relationship with the child node. Through this bottom-up adjustment method, the local layout of the template can be adapted to the characteristics of the actual document while maintaining the overall structure of the rule tree.

[0068] In addition to adjusting the positions of existing nodes, the server also needs to process the newly added fields recorded in the difference area descriptor. For each newly added field, the server first finds the field node with the highest degree of association in the rule tree based on the preset field association rules, and then adds a new child node under this node to represent the newly added field. Association rules can be formulated based on factors such as the semantic similarity of field names and the proximity of fields in the document. For example, if the new field is named "Tax Rate", then according to the association rules, the server may add it to the "Tax Amount" field node, indicating that "Tax Rate" is ancillary information of "Tax Amount". After adding the new node, the server must also determine the specific position coordinates of the new node in the template based on the new field position offset recorded in the difference area descriptor.

[0069] Finally, based on the adjusted rule tree structure, the server updates the overall structural layout of the target template, generating the first reconstructed template. This update process adheres to certain constraints, such as ensuring that the spacing between adjacent fields cannot be too small and that fields cannot exceed the page boundaries, to ensure the visual rationality and aesthetics of the reconstructed template. The server utilizes layout analysis and constraint optimization techniques to generate a new template with a compact layout and coordinated structure while meeting the requirements of the rule tree structure.

[0070] For example, suppose the difference region descriptor indicates that the "Purchaser Name" field in an invoice image is offset 20 pixels downward compared to the target template, and a new "Contract Number" field has been added between the "Purchaser Name" and "Seller Name" fields. The server first locates the "Purchaser Name" node in the rule tree and updates its position to indicate a 20-pixel downward offset. Then, based on the association rule, the server adds a new "Contract Number" child node under the "Purchaser Name" node, determining its location coordinates based on the descriptor. Finally, the server adjusts the position of the "Seller Name" and other related nodes and updates the entire template's layout to conform to the new rule tree structure.

[0071] Step S105: Identify the financial document based on the first reconstruction template to obtain initial financial data.

[0072] In step S105, the server performs block-by-block recognition on the financial document image based on the field positioning frames in the first reconstruction template. Specifically, the server maps the positioning frames of each field in the first reconstruction template onto the document image, forming regions of interest (ROIs). The server then performs OCR (Optical Character Recognition) on the image content within each ROI to extract the text information. Thanks to the more precise positioning and size of the positioning frames in the reconstruction template, the server can more accurately capture the complete image of the field, reducing recognition errors caused by field truncation and offset.

[0073] Step S106: Determine whether the initial financial data satisfies the preset cross-reference relationship table.

[0074] In step S106, the server performs a cross-check on the initial financial data to determine whether it satisfies a preset cross-check relationship table. The preset cross-check relationship table is a set of rules used to verify the inherent logical relationships of financial data. It specifies the identities, inequalities, or other constraints that must be satisfied between various financial fields. The introduction of new fields or adjustments to existing fields may disrupt the original data association logic, causing the intermediate financial data to fail to meet the constraints of the preset cross-check relationship table. By substituting the initial financial data into the cross-check relationship table for verification, the server can identify potential errors, inconsistencies, or anomalies caused by version changes, thereby further improving the reliability and accuracy of the financial data.

[0075] Specifically, the server first extracts the cross-checking rules relevant to the current financial document type from a pre-set cross-checking table. This table is typically stored and organized in a structured format like JSON or XML, making it easier for programs to parse and query. Different financial document types, such as VAT invoices and bank receipts, may require different cross-checking rules. By identifying document type characteristics, such as template IDs and key fields, the server can quickly locate the corresponding cross-checking rules.

[0076] After extracting applicable cross-checking rules, the server substitutes the field values ​​from the initial financial data into the cross-checking rules one by one for calculation and comparison. Each cross-checking rule defines an equality or inequality, with the left and right sides of the equation each consisting of one or more field operations. The server retrieves the corresponding values ​​from the initial financial data based on the field names and substitutes them into the operation. The server then evaluates the left and right side of the equation and determines whether the equality or inequality requirements are met.

[0077] Taking VAT invoices as an example, common cross-checking rules may include:

[0078] Total price and tax = amount of goods or taxable labor or services + tax

[0079] Tax amount = amount of goods or taxable labor or services × tax rate

[0080] Verification code = f(invoice code + invoice number + invoice date + verification code parameter)

[0081] Rules 1 and 2 represent the arithmetic relationships between the total price and tax, the tax amount, and other fields, respectively, while Rule 3 represents the functional relationship between the check code and other fields. The server substitutes the values ​​of the "total price and tax," "goods or taxable services, service amount," and "tax amount" fields from the initial financial data into Rule 1 to verify the equation. It also substitutes the values ​​of the "tax amount," "goods or taxable services, service amount," and "tax rate" fields into Rule 2 to verify the equation. It then substitutes the values ​​of the "invoice code," "invoice number," "invoice date," and "check code parameter" fields into function f in Rule 3 for calculation and compares them with the value of the "check code" field in the initial financial data.

[0082] As the cross-checking rules are checked one by one, the server records the results of each rule, including whether it passes, the difference between the actual value and the expected value, and other information. After checking all applicable cross-checking rules, the server comprehensively determines whether the initial financial data satisfies the preset cross-checking relationship table. If all cross-checking rules pass, the initial financial data is considered to be consistent with the preset cross-checking relationship table and can proceed to subsequent business processing. Conversely, if any rules fail the check, the initial financial data is considered inconsistent or abnormal, requiring further verification and processing.

[0083] For example, for the initial financial data of the above VAT invoice, the server extracts the following applicable cross-checking rules:

[0084] Total price and tax (lowercase) = amount of goods or taxable labor or services + tax

[0085] Tax amount = amount of goods or taxable labor or services × tax rate

[0086] The server then inserts the field values ​​from the initial financial data into the rules for verification. Suppose that after substituting these values, Rule 1 passes, but Rule 2 fails (the actual tax amount is inconsistent with the calculated tax rate). The server then generates an exception record for Rule 2, indicating the inconsistency in the "Tax Amount" field. Based on the results of the two rule checks, the server determines that the initial financial data does not meet the predefined cross-reference table requirements and requires verification and processing.

[0087] Step S107: If it is determined that the initial financial data satisfies the preset cross-reference relationship table, the initial financial data is output.

[0088] In step S107, the server performs final verification and output of the initial financial data. Once the initial financial data satisfies the pre-defined cross-reference table, it signifies that the data has passed multiple checks, including format and logic validation, and possesses a high degree of accuracy and reliability, making it suitable for subsequent business processing. At this point, the server packages and outputs the initial financial data according to the specified format and interface requirements, allowing other business systems or modules to easily access and use the data.

[0089] Step S108: If it is determined that the initial financial data does not satisfy the preset cross-reference relationship table, the first reconstruction template is revised twice to obtain a second reconstruction template and corresponding target financial data, and the target financial data is output.

[0090] In step S108, the first reconstruction template is corrected twice to obtain a second reconstruction template and corresponding target financial data, specifically including: determining abnormal fields in the initial financial data that do not satisfy the cross-checking relationship according to a preset cross-checking relationship table; locating the abnormal area corresponding to the abnormal field in the first reconstruction template, and obtaining the position coordinates of the abnormal area in the financial document image; based on the position coordinates, collecting the context area adjacent to the abnormal area in the financial document image, and performing extended recognition on the context area to obtain an extended recognition result; judging whether there are interference factors that affect the accuracy of abnormal field recognition according to the extended recognition result, the interference factors include field overlap, field breakage, and field tilt; if there are interference factors, correcting the recognition result of the abnormal area according to a preset compensation rule to obtain corrected abnormal field data; updating the corrected abnormal field data to the initial financial data to obtain intermediate financial data; judging whether the intermediate financial data satisfies the preset cross-checking relationship table; if the intermediate financial data satisfies the preset cross-checking relationship table, determining the intermediate financial data as the target financial data, and adjusting the first reconstruction template according to the extended recognition result to obtain a second reconstruction template.

[0091] Specifically, the server performs a secondary verification and correction on the initial financial data. If the initial financial data is determined not to meet the preset cross-reference table, it indicates a logical error or inconsistency in the data, requiring further investigation and resolution. The server then performs a secondary correction on the first reconstructed template, combining the cross-reference table with the original document image to locate the problematic, abnormal fields. The server then attempts to correct them through methods such as extended recognition and data compensation, ultimately producing the corrected second reconstructed template and the corresponding target financial data.

[0092] Specifically, the server first identifies abnormal fields in the initial financial data that do not meet the cross-checking requirements based on a pre-set cross-checking table. These abnormal fields typically exhibit logical discrepancies with other fields, such as incorrect totals in amount fields or incorrect tax calculations. The server then traces back the cross-checking process to identify the specific fields causing the abnormalities and records their names and locations.

[0093] The server then locates the abnormal region corresponding to the abnormal field in the first reconstruction template. The first reconstruction template records the coordinates of each field in the original document image. Based on the name of the abnormal field, the server can quickly locate its location in the image. The server also obtains the coordinates of the circumscribed rectangle of the abnormal region as a reference for subsequent extended recognition.

[0094] Next, the server collects the context area adjacent to the abnormal area in the original document image based on the coordinates of the abnormal area. The context area typically refers to the image area within a certain range surrounding the abnormal area and may contain critical information that affects the recognition of the abnormal field, such as other related fields and table lines. The server uses pre-set context collection rules, such as fixed pixel boundaries and dynamic scaling factors, to crop an appropriately sized image area around the abnormal area.

[0095] The server performs extended recognition on the collected context area to obtain more comprehensive recognition results. Extended recognition combines contextual information and prior knowledge with existing OCR recognition to perform secondary analysis and correction of field content. For example, for an amount field, extended recognition can convert it to numeric format and compare lowercase and uppercase amounts. For a date field, extended recognition can normalize the format and check the date's validity. Extended recognition can compensate for deficiencies in initial recognition and improve the accuracy and completeness of field content.

[0096] After obtaining the extended recognition results, the server analyzes whether there are any interference factors that affect the accuracy of abnormal field recognition. Interference factors include:

[0097] Field overlap: The abnormal field overlaps or is blocked by other fields, resulting in incomplete or incorrect recognition.

[0098] Field breakage: The abnormal field content is broken or incomplete, such as the amount digit is disconnected in the middle, resulting in incoherent recognition.

[0099] Field tilt: The image area where the abnormal field is located appears tilted or deformed, resulting in reduced recognition accuracy.

[0100] The server determines whether these interference factors exist by analyzing the image features of the context area, such as connected areas and edge directions.

[0101] If there are interference factors, the server will correct the recognition results of the abnormal area according to the preset compensation rules. The preset compensation rules are a set of heuristic strategies summarized based on business experience and sample data, which are used to make corrections when the recognition results are not ideal. For example, in the case of field overlap, the compensation rules will infer the complete content of the field based on the pixel density, color distribution and other characteristics of the overlapping part; in the case of field breakage, the compensation rules will splice the broken parts based on the strokes, spacing and other characteristics of the characters at the break; in the case of tilted fields, the compensation rules will perform reverse rotation correction on the image based on the tilt angle and direction. The server restores the true content of the abnormal field by applying the preset compensation rules.

[0102] After obtaining the corrected abnormal field data, the server updates it with the initial financial data, generating intermediate financial data. Intermediate financial data refers to the initial financial data after partial corrections, and the correctness of its logical relationships still requires further verification. Therefore, the server re-enters the preset cross-checking relationship table to determine whether it meets all cross-checking rules.

[0103] If the intermediate financial data passes the cross-check, the server identifies it as the final target financial data and outputs it for subsequent business use. Simultaneously, the server adjusts and optimizes the first reconstruction template based on the expanded recognition results, producing a second reconstruction template. These adjustments primarily include: expanding or shrinking the recognition area for the corresponding field based on the location coordinates of the abnormal field; adding new association rules for field recognition based on contextual information; and setting conditional constraints or default values ​​for field recognition based on the application of compensation rules. These adjustments, based on the first reconstruction template, further improve recognition accuracy and stability, reducing the occurrence of abnormalities.

[0104] For example, suppose the "Tax Amount" field in the initial financial data is misidentified due to printing defects, resulting in a violation of the cross-reference relationship with the "Price and Tax Total" field. After locating the "Tax Amount" field in the document image and collecting its context, the server discovered that the digits on the right side of the field were broken. Using extended recognition, the server corrected the value of the "Tax Amount" field and updated it with the initial financial data. After re-cross-checking, the server confirmed that the corrected data satisfied all cross-reference relationships. The server identified the corrected data as the target financial data and, based on the break discovered during extended recognition, expanded the recognition area of ​​the "Tax Amount" field in the first reconstruction template, generating a second reconstruction template.

[0105] In one possible implementation, after determining whether the intermediate financial data satisfies the preset cross-reference relationship table, the method further includes: if the intermediate financial data does not satisfy the preset cross-reference relationship table, identifying the conflicting fields in the intermediate financial data that do not satisfy the preset cross-reference relationship table, and extracting the conflict coordinate area corresponding to the conflicting fields in the first reconstruction template; calculating the theoretical coordinate position of the conflicting fields according to the constraint rules related to the conflicting fields in the preset cross-reference relationship table, and correcting the positioning box of the conflicting fields in the first reconstruction template to the theoretical coordinate position; within the conflict coordinate area, identifying newly added fields through text semantic analysis, and adding the newly added fields to the positioning rules of the first reconstruction template to generate a second reconstruction template, and identifying the financial documents according to the second reconstruction template to obtain the target financial data.

[0106] Specifically, when the server finds that the intermediate financial data still does not satisfy the preset cross-checking relationship table, it means that after the initial reconstruction of the template, the introduction of new fields or the adjustment of existing fields disrupted the original data association logic, which in turn caused the intermediate financial data to fail to meet the constraints of the preset cross-checking relationship table. There are logical contradictions or conflicts within the data. In order to further locate the problem, the server first identifies the conflicting fields in the intermediate financial data that do not satisfy the preset cross-checking relationship table. Conflicting fields refer to fields that are inconsistent or erroneous during the cross-checking relationship check, such as those that do not match the calculation relationship with other fields or have values ​​outside the reasonable range. The server analyzes the inspection results of each rule in the cross-checking relationship table to find out the specific fields that caused the check to fail, and records the names and location information of these fields.

[0107] After identifying the conflicting fields, the server extracts the conflicting coordinate regions corresponding to these fields in the first reconstructed template. The server matches the conflicting field names, finds their corresponding coordinate information in the first reconstructed template, and extracts a rectangular region containing the field as the conflicting coordinate region. This conflicting coordinate region indicates potential problematic areas in the image and serves as the focus of subsequent analysis and correction.

[0108] Next, the server calculates the theoretical coordinate position of the conflicting field based on the constraints associated with the conflicting field in the pre-set cross-reference table. Constraints are further refinements of the cross-reference table, defining specific positioning and formatting requirements, such as relative position, alignment, and size range between fields. These rules are based on extensive historical data and business experience. By analyzing the constraints associated with the conflicting field, the server calculates the theoretical coordinate position of the field, assuming the constraints are met, and uses this as a reference for correction.

[0109] The server then adjusts the positioning box of the conflicting field in the first reconstructed template to the theoretical coordinate position. The positioning box is a rectangular area that marks the location of the field in the reconstructed template and determines the range of focus of the recognition engine. By adjusting the positioning box to the theoretical coordinate position, the server effectively corrects the target area for recognition, making it more consistent with the expected positional relationship and format requirements.

[0110] However, simply adjusting the positioning box doesn't completely resolve the issue, as errors in conflicting fields may not only be due to positional offsets but may also be related to the content itself. To this end, the server also uses text semantic analysis within the conflicting coordinate area to identify fields that may be missing or incorrect, or that need to be added. Text semantic analysis is a method based on natural language processing technology that analyzes the lexical, grammatical, and semantic features of text to understand its inherent logical relationships and information structure. In the context of financial documents, text semantic analysis can help identify keywords, number formats, contextual relationships, and more, thereby discovering fields that were previously incorrectly identified or that need to be added.

[0111] For example, within the conflicting coordinate area of ​​the "Tax Amount" field, the server discovered, through text semantic analysis, that keywords such as "tax rate" and "amount excluding tax" appeared in this area, but were not reflected in the previous recognition results. This suggests that the error in the "Tax Amount" field may be related to the omission of these newly added fields. Therefore, the server extracted these newly added fields and added them to the positioning rules of the first reconstruction template, generating a more complete second reconstruction template.

[0112] Finally, the server uses the generated second reconstruction template to re-recognize the original financial document image, obtaining the corrected target financial data. Compared to the initial recognition, this new recognition fully utilizes information from cross-reference tables, constraint rules, and text semantics, and optimizes and supplements the position and content of the recognized fields, thereby improving the accuracy and completeness of the results.

[0113] For example, suppose that in the intermediate financial data, the value of the "Total Amount" field is inconsistent with the product of the "Unit Price" and "Quantity" fields, thus being identified as a conflicting field. Within the conflicting coordinate area of ​​the "Total Amount" field, the server, through text semantic analysis, discovers the need to add a "Discount Amount" field. According to the constraints in the pre-set cross-reference table, the "Total Amount" field should be equal to the "Unit Price" multiplied by the "Quantity" minus the "Discount Amount." Therefore, the server adds the "Discount Amount" field to the second reconstruction template and, based on the constraints, calculates the theoretical coordinate position of the "Total Amount" field and adjusts its positioning box. After re-identification, the relationship between the "Total Amount" field value and the other fields is now aligned, resolving the conflict.

[0114] In a possible implementation, after step S108, the method further includes: obtaining user feedback information during the financial document recognition process, the user feedback information including a recognition accuracy score and a recognition abnormality description; when the recognition accuracy score is lower than a preset threshold, extracting key information related to the recognition abnormality description, and locating the abnormal area in the second reconstruction template; based on the key information, performing targeted identification on the abnormal area, and adding the identification result to the target financial data to generate corrected target financial data; and sending the corrected target financial data to the user terminal corresponding to the user feedback information.

[0115] Specifically, after the server completes the financial document recognition process, it pushes the recognition results and a visual analysis page of the recognition process to the corresponding user terminal. The user views the recognition results on the terminal and can provide feedback through interactive functions on the page, including recognition accuracy scores (e.g., a 5-point scale) and descriptions of recognition anomalies (e.g., missing or incorrect recognition of certain fields).

[0116] After receiving user feedback, the server triggers a recognition optimization process. First, the server determines whether the user's recognition accuracy score is below a preset threshold (e.g., 3). If the score reaches the threshold, the recognition result is considered satisfactory and no further optimization is required. If the score is below the threshold, there are issues with the recognition result, requiring targeted optimization.

[0117] Next, the server analyzes the user-entered description of the recognition anomaly and, using natural language processing, extracts key information, such as "buyer name missed" or "invoice amount incorrectly recognized." The server then maps this key information onto the second reconstruction template used in the financial document recognition process. Specifically, the server searches the second reconstruction template for the field corresponding to the key information, locates the coordinate area of ​​that field in the financial document image, and marks it as the anomaly area.

[0118] For example, if a user reports that the "purchaser's name was not recognized", the server will search for the "purchaser's name" field in the second reconstruction template, obtain its coordinate information (such as the upper left corner coordinates and the lower right corner coordinates), and select the area in the original financial document image.

[0119] After identifying the abnormal area, the server conducts targeted recognition. Compared to global recognition, this can utilize more sophisticated and complex algorithms, such as integrating multiple OCR engines and performing multiple iterative recognition cycles, to improve recognition accuracy. Furthermore, the server dynamically adjusts recognition parameters, such as image preprocessing and character segmentation, based on the characteristics of the abnormal area (e.g., font style, background noise), to further optimize recognition results.

[0120] After the targeted identification is completed, the server will add the identification results to the previously generated target financial data, overwriting the original identification results to generate the corrected target financial data. If there are multiple abnormal areas, the server will identify and update the results one by one.

[0121] Finally, the server transmits the corrected target financial data back to the user terminal for review and confirmation. The server also generates a recognition optimization report, detailing the optimization process and results, including the location of the abnormal area, the original recognition result, the optimized recognition result, and quantitative indicators of the recognition confidence improvement. The report is also sent to the user terminal for display.

[0122] In a possible implementation, the method further includes steps S201 to S204, which are as follows:

[0123] Step S201: Use the target template to identify the original image to obtain basic financial data.

[0124] In step S201, the server first loads a target template, which is a template matched against key fields in a preset reference template library. The server then uses this target template to recognize the original image of the financial document. Specifically, the server inputs the original image into the OCR model corresponding to the target template. Through operations such as text detection and recognition, the server extracts the text information from the image and categorizes it into different fields based on the predefined field positions and formats in the target template, generating structured financial data, or basic financial data.

[0125] Step S202: Based on the target template, targeted training is performed on key fields and difference region descriptors to generate a difference adaptation model.

[0126] In step S202, based on the target template, targeted training is performed on key fields and difference area descriptors to generate a difference adaptation model, specifically including: extracting corresponding difference image areas in the original image of the financial document according to the key fields and difference area descriptors to generate a difference training set; loading the pre-trained OCR model corresponding to the target template as the basic model, and freezing the general feature extraction layer parameters that are not related to text recognition in the basic model, decoupling the output layer parameters related to field positioning and the rule layer parameters related to business verification logic in the basic model to obtain a decoupled model; based on the difference training set, the decoupled model is targeted fine-tuned to generate a difference adaptation model.

[0127] Specifically, the server extracts the corresponding difference image area in the original image of the financial document based on the key fields and difference area descriptors to generate a difference training set. Specifically, the server traverses each key field and calculates the target area coordinates of the field in the original image based on its position offset information in the difference area descriptor. Then, the server takes the target area as the center and expands a certain margin around it to extract the difference image area corresponding to the field. For newly added fields, the server locates the position of the invoice amount field in the original image based on its offset information in the difference area descriptor, such as "below the invoice amount", and expands downward to a certain height to extract the difference image area of ​​the newly added field. The server saves all the extracted difference image areas and, together with the corresponding key field labels, forms a difference training set.

[0128] Next, the server loads the pretrained OCR model corresponding to the target template as the base model. This pretrained OCR model consists of a general feature extraction layer, an output layer related to field location, and a rule layer related to business verification logic. The general feature extraction layer extracts general text features from the image, the output layer predicts the location and content of each field based on the text features, and the rule layer performs business verification based on the predicted field content.

[0129] To more efficiently adapt to the difference areas, the server needs to decouple the base model. The server freezes the general feature extraction layer parameters in the base model that are not related to text recognition, retaining only the forward computation function and not performing backpropagation updates. This is done to preserve the general text features already learned by the base model and reduce training overhead. The server then decouples the output layer parameters related to field positioning from the rule layer parameters related to business validation logic in the base model. These decoupled output and rule layer parameters can be independently optimized and updated, allowing for more flexible adaptation to changes in the difference areas. This decoupled model is called a decoupled model.

[0130] Finally, the server performs targeted fine-tuning on the decoupling model based on the difference training set to generate a difference adaptation model. Specifically, the server inputs the difference image area in the difference training set into the decoupling model to obtain the output layer feature map and the rule layer feature map of the decoupling model. Then, the server calculates the target loss in the output layer feature map based on the field position offset information and the newly added field offset information in the difference area descriptor, including the changed field positioning loss and the newly added field positioning loss. The changed field positioning loss is calculated based on the coordinate offset vector of the changed field, so that the output layer can accurately predict the new position of the changed field. The newly added field positioning loss is calculated based on the local features and table lines within the candidate positioning area, so that the output layer can predict the newly added field at the correct position.

[0131] At the same time, the server calculates the association loss in the rule layer feature graph based on the newly added field validation rules in the target business rules. This association loss is calculated based on the semantic dependency between the newly added field and the associated fields, enabling the rule layer to establish association constraints between the newly added field and other fields. The server constructs a joint loss function that takes the weighted sum of the target loss and the association loss as the optimization objective for targeted fine-tuning. Using the backpropagation algorithm, the server performs gradient optimization on the output layer parameters and the rule layer parameters of the decoupled model based on the joint loss function, iteratively updating the parameters until the joint loss function converges.

[0132] After multiple rounds of targeted fine-tuning, the decoupled model's output and rule layers have adapted to the changes in the difference regions, accurately identifying changed and newly added fields and establishing new association constraints. The server uses a parameter fixation strategy to integrate the fine-tuned output and rule layer parameters with the original general feature extraction layer parameters to generate a complete difference adaptation model.

[0133] For example, suppose that before a financial document template was changed, the invoice number field was in the upper left corner and the amount field was in the lower right corner, with no correlation between the two. After the change, the invoice number field was shifted 100 pixels to the right, and a tax rate field was added below the amount field. The tax rate field must satisfy a certain computational relationship with the amount field. The server first extracts the difference image regions corresponding to the invoice number and tax rate fields to generate a difference training set. The server then loads a pretrained OCR model, freezes the common feature extraction layer, and decouples the output and rule layers to create a decoupled model. Next, the server calculates the localization loss for the invoice number and tax rate fields in the output layer, and the correlation loss between the tax rate and amount fields in the rule layer. This joint loss function is constructed and fine-tuned on the decoupled model. Finally, the server solidifies the fine-tuned parameters to generate a difference adaptation model. This model accurately locates the shifted invoice number field and the newly added tax rate field, ensuring that the tax rate and amount fields satisfy a computational relationship.

[0134] In one possible implementation, based on the difference training set, the decoupling model is fine-tuned in a targeted manner to generate a difference adaptation model, specifically including: inputting the difference training set into the decoupling model to obtain the output layer feature map and the rule layer feature map of the decoupling model; calculating the target loss in the output layer feature map based on the field position offset information and the newly added field offset information, the target loss includes the changed field positioning loss and the newly added field positioning loss, the changed field positioning loss is calculated based on the coordinate offset vector of the changed field, and the newly added field positioning loss is calculated based on the local features and table lines within the candidate positioning area; calculating the association loss in the rule layer feature map based on the newly added field verification rules in the target business rules, the association loss is calculated based on the semantic dependency relationship between the newly added field and the associated field; constructing a joint loss function, the joint loss function includes the target loss and the association loss, and gradient optimizing and updating the output layer parameters and the rule layer parameters of the decoupling model based on the joint loss function to obtain a difference adaptation model.

[0135] Specifically, the server inputs the difference training set into the decoupling model. After forward propagation, it obtains the decoupling model's output layer feature map and rule layer feature map. The output layer feature map represents the decoupling model's prediction of the field positions in the difference image region, while the rule layer feature map represents the decoupling model's prediction of the associations between fields.

[0136] Next, the server calculates the target loss in the output layer feature map based on the field position offset information and the newly added field offset information. The target loss consists of two parts: the changed field positioning loss and the newly added field positioning loss.

[0137] For the changed field location loss, the server first calculates the coordinate offset vector for each changed field based on the field position offset information. The server then compares the coordinate offset vector with the predicted coordinates of the corresponding field in the output layer feature map and calculates the distance between the two as the changed field location loss. By minimizing this loss, the output layer of the decoupled model accurately predicts the new location of the changed field.

[0138] For the new field positioning loss, the server first determines the candidate positioning area of ​​the new field in the difference image area based on the new field offset information. Then, the server extracts local features in the candidate positioning area, such as the size and appearance of the field, and identifies the table lines in the area. The server inputs the local features and table lines into a predefined scoring function to calculate the score of the candidate positioning area. The higher the score of the candidate positioning area, the more likely it is to be the correct location of the new field. The server selects the candidate positioning area with the highest score as the predicted location of the new field, and compares it with the actual location of the new field, and calculates the distance between the two as the new field positioning loss. By minimizing this loss, the output layer of the decoupled model can predict the new field at the correct location.

[0139] In addition to the target loss, the server also calculates the association loss in the rule-layer feature graph based on the validation rules for the newly added fields in the target business rules. The association loss reflects the semantic dependency between the newly added fields and the associated fields. The server first extracts the semantic dependency between the newly added fields and the associated fields from the target business rules, such as "tax rate = amount × tax rate value." The server then finds the prediction results for the newly added fields and the associated fields in the rule-layer feature graph and calculates the difference between the two based on the semantic dependency, which is used as the association loss. By minimizing the association loss, the rule layer of the decoupled model learns the validation logic between the newly added fields and the associated fields.

[0140] After calculating the target loss and the associated loss, the server constructs a joint loss function, taking the weighted sum of the two. The joint loss function represents the overall performance of the decoupled model in locating changed fields, predicting newly added fields, and establishing field associations. The server uses the gradient descent algorithm to optimize the output layer parameters and rule layer parameters of the decoupled model, using the joint loss function as the optimization target. Specifically, the server calculates the gradient of the joint loss function with respect to the output layer parameters and the rule layer parameters, and updates the parameters based on the direction and magnitude of the gradient, so that the joint loss function continuously decreases. The server repeats the above process until the joint loss function converges to a smaller value, or the preset number of iterations is reached.

[0141] After multiple rounds of iterative optimization, the decoupling model's output layer can accurately predict the new positions of changed and newly added fields, and the rule layer can correctly establish the validation relationships between newly added and associated fields. This results in a performance-optimized decoupling model, known as the differential adaptation model.

[0142] For example, suppose the difference training set contains an invoice image in which the invoice number field is shifted 100 pixels to the right, and a tax rate field is added below the amount field. The relationship between the tax rate field and the amount field satisfies the following equation: "tax rate = amount × 10%." The server first inputs this invoice image into the decoupling model to obtain an output layer feature map and a rule layer feature map. The server then calculates the coordinate offset vector (100, 0) based on the position offset information of the invoice number field and compares it with the predicted coordinates of the invoice number field in the output layer feature map to calculate the localization loss. Based on the offset information of the tax rate field, "below the amount field," the server extracts local features and table lines in the area below the amount field, calculates the scores of the candidate localization areas, and uses the area with the highest score as the predicted location of the tax rate field. The localization loss is then calculated.

[0143] Next, the server extracts the relationship between the tax rate field and the amount field ("tax rate = amount x 10%) from the target business rule, finds the predicted values ​​for the two in the rule-layer feature graph, and calculates the associated loss. The server then takes a weighted sum of the invoice number field location loss, the tax rate field location loss, and the associated loss to derive a joint loss function. Based on this loss function, the server optimizes the output layer and rule layer parameters of the decoupling model, ultimately resulting in a differential adaptation model.

[0144] Step S203: applying the difference adaptation model to the original image of the financial document, performing targeted recognition on the original image, and obtaining a difference recognition result.

[0145] In step S203, the server applies the generated difference adaptation model to the original image of the financial document, performing targeted recognition on the original image and obtaining a difference recognition result. Unlike the generalized recognition using the target template in step S201, the difference adaptation model performs targeted recognition, focusing on changed and newly added fields. The server inputs the original image into the difference adaptation model, performs text detection and recognition, extracts the text information and location of the changed and newly added fields, and obtains the difference recognition result.

[0146] Step S204: The difference identification result is integrated with the basic financial data, and the changed fields and newly added fields in the basic financial data are updated to obtain the initial financial data.

[0147] In step S204, the server integrates the difference identification results with the base financial data, updating the changed and newly added fields in the base financial data to obtain the initial financial data. Specifically, the server compares the difference identification results with the base financial data, identifies the changed and newly added fields, and updates the corresponding field values ​​in the base financial data with the field values ​​from the difference identification results. Newly added fields are directly added to the base financial data. Through this integration, the server organically combines the incremental information identified by the difference adaptation model with the basic information identified by the target template, resulting in a complete set of initial financial data that adapts to the changed business rules.

[0148] For example, suppose that in the original image of a financial document, the invoice number field has shifted in position, and the tax rate field has been newly added. The server uses the target template to recognize the original image, and the resulting basic financial data lacks the tax rate field, and the invoice number field may be missed. The server generates a differential adaptation model based on the shift in the invoice number field and the addition of the tax rate field. The server uses the differential adaptation model to recognize the original image and obtain the accurate invoice number and tax rate. The server overwrites the invoice number in the basic financial data with the invoice number from the differential recognition result and adds the tax rate to the basic financial data, ultimately obtaining an initial set of financial data containing the correct invoice number and tax rate.

[0149] Reference Figure 3 , the present application also provides a financial document processing device, which is a server, and the server includes: an original image acquisition module 301, a difference determination module 302, a target template matching module 303, a template reconstruction module 304, a cross-reference relationship processing module 305 and a secondary reconstruction module 306, wherein: the original image acquisition module 301 is used to obtain the original image of the financial document to be processed; the difference determination module 302 is used to monitor the version update status of the financial business rule library, and when a rule change instruction that triggers a preset change condition is detected, the changed target business rule and the key fields and difference area descriptors related to the target business rule are obtained, the key fields include the changed fields and the dependent fields associated with the changed fields in the preset cross-reference relationship table, and the difference area descriptors include field position offset information and new field offset information; the target The template matching module 303 is used to match the corresponding target template in the preset benchmark template library according to the key fields; the template reconstruction module 304 is used to reconstruct the target template according to the target business rules and the difference area descriptor to obtain a first reconstructed template; the template reconstruction module 304 is also used to identify the financial document based on the first reconstructed template to obtain initial financial data; the cross-reference relationship processing module 305 is used to determine whether the initial financial data satisfies the preset cross-reference relationship table; the cross-reference relationship processing module 305 is also used to output the initial financial data if it is determined that the initial financial data satisfies the preset cross-reference relationship table; the secondary reconstruction module 306 is used to perform a secondary correction on the first reconstructed template if it is determined that the initial financial data does not satisfy the preset cross-reference relationship table, to obtain a second reconstructed template and the corresponding target financial data, and output the target financial data.

[0150] In one possible implementation, the template reconstruction module 304 reconstructs the target template according to the target business rules and the difference area descriptor to obtain a first reconstructed template, specifically including: the template reconstruction module 304 calculates the coordinate offset vector of each field according to the field position offset information, and superimposes the coordinate offset vector to the positioning box parameter of the corresponding field in the target template to update the position of the field positioning box; the template reconstruction module 304 determines the candidate positioning area of ​​the newly added field through context analysis based on the newly added field offset information and the semantic constraints related to the newly added field in the target business rules, and extracts local features and identifies table lines in the candidate positioning area to generate a positioning box for the newly added field; the template reconstruction module 304 constructs the association verification logic between fields based on the newly added field verification rules in the target business rules, and adds the association verification logic to the verification rule set of the target template; the template reconstruction module 304 integrates the field positioning box, the positioning box of the newly added field and the association verification logic into the target template to generate the first reconstructed template.

[0151] In a possible implementation, the secondary reconstruction module 306 performs a secondary correction on the first reconstruction template to obtain a second reconstruction template and the corresponding target financial data, specifically including: the secondary reconstruction module 306 determines the abnormal fields in the initial financial data that do not satisfy the cross-reference relationship according to the preset cross-reference relationship table; the secondary reconstruction module 306 locates the abnormal area corresponding to the abnormal field in the first reconstruction template, and obtains the position coordinates of the abnormal area in the financial document image; the secondary reconstruction module 306 collects the context area adjacent to the abnormal area in the financial document image based on the position coordinates, and performs extended recognition on the context area to obtain an extended recognition result; the secondary reconstruction module 306 determines whether the abnormal area is correct according to the extended recognition result. There are interference factors that affect the accuracy of abnormal field recognition, including field overlap, field breakage and field tilt; if there are interference factors, the secondary reconstruction module 306 corrects the recognition result of the abnormal area according to the preset compensation rules to obtain corrected abnormal field data; the secondary reconstruction module 306 updates the corrected abnormal field data to the initial financial data to obtain intermediate financial data; the secondary reconstruction module 306 determines whether the intermediate financial data satisfies the preset cross-reference relationship table; if the intermediate financial data satisfies the preset cross-reference relationship table, the secondary reconstruction module 306 determines the intermediate financial data as the target financial data, and adjusts the first reconstruction template according to the extended recognition result to obtain the second reconstruction template.

[0152] In one possible embodiment, after the secondary reconstruction module 306 determines whether the intermediate financial data satisfies the preset cross-reference relationship table, the method further includes: if the intermediate financial data does not satisfy the preset cross-reference relationship table, the secondary reconstruction module 306 identifies the conflicting fields in the intermediate financial data that do not satisfy the preset cross-reference relationship table, and extracts the conflict coordinate area corresponding to the conflicting fields in the first reconstruction template; the secondary reconstruction module 306 calculates the theoretical coordinate position of the conflicting fields according to the constraint rules related to the conflicting fields in the preset cross-reference relationship table, and corrects the positioning box of the conflicting fields in the first reconstruction template to the theoretical coordinate position; within the conflict coordinate area, the secondary reconstruction module 306 identifies the newly added fields through text semantic analysis, and adds the newly added fields to the positioning rules of the first reconstruction template to generate a second reconstruction template, and identifies the financial documents according to the second reconstruction template to obtain the target financial data.

[0153] In one possible implementation, when a rule change instruction that triggers a preset change condition is detected, the template reconstruction module 304 obtains the changed target business rule and key fields and difference area descriptors related to the target business rule, specifically including: the template reconstruction module 304 detects the version number of the financial business rule library according to a preset time interval; the template reconstruction module 304 compares the detected version number with the locally stored version number to obtain a version number difference value; if the version number difference value is greater than a preset threshold, the template reconstruction module 304 determines that the version of the financial business rule library has been updated and triggers the rule change instruction of the preset change condition; when a rule change instruction is detected, the template reconstruction module 304 obtains the change log file of the financial business rule library, identifies the change record, and extracts the change record to obtain the changed target business rule, which includes the change content and the change rule; the template reconstruction module 304 performs semantic analysis on the change content to determine the key fields related to the change rule; and performs semantic analysis on the change rule to determine the field position offset information and the newly added field offset information of the key field.

[0154] In a possible embodiment, the method also includes: the template reconstruction module 304 uses the target template to identify the original image to obtain basic financial data; the template reconstruction module 304 performs targeted training on key fields and difference area descriptors based on the target template to generate a difference adaptation model; the template reconstruction module 304 applies the difference adaptation model to the original image of the financial document, performs targeted identification on the original image, and obtains a difference identification result; the template reconstruction module 304 integrates the difference identification result with the basic financial data, updates the changed fields and newly added fields in the basic financial data, and obtains the initial financial data.

[0155] In one possible implementation, the template reconstruction module 304 performs targeted training on key fields and difference area descriptors based on the target template to generate a difference adaptation model, specifically including: the template reconstruction module 304 extracts corresponding difference image areas in the original image of the financial document according to the key fields and difference area descriptors to generate a difference training set; the template reconstruction module 304 loads the pre-trained OCR model corresponding to the target template as the basic model, and freezes the general feature extraction layer parameters in the basic model that are not related to text recognition, decouples the output layer parameters related to field positioning and the rule layer parameters related to business verification logic in the basic model to obtain a decoupled model; the template reconstruction module 304 performs targeted fine-tuning on the decoupled model based on the difference training set to generate a difference adaptation model.

[0156] In one possible implementation, the template reconstruction module 304 performs targeted fine-tuning on the decoupling model based on the difference training set to generate a difference adaptation model, specifically including: the template reconstruction module 304 inputs the difference training set into the decoupling model to obtain the output layer feature map and the rule layer feature map of the decoupling model; the template reconstruction module 304 calculates the target loss in the output layer feature map according to the field position offset information and the newly added field offset information, the target loss includes the changed field positioning loss and the newly added field positioning loss, the changed field positioning loss is calculated based on the coordinate offset vector of the changed field, and the newly added field positioning loss is calculated based on the local features and table lines within the candidate positioning area; the template reconstruction module 304 calculates the association loss in the rule layer feature map according to the newly added field verification rule in the target business rule, the association loss is calculated based on the semantic dependency relationship between the newly added field and the associated field; the template reconstruction module 304 constructs a joint loss function, the joint loss function includes the target loss and the association loss, and gradient optimizes and updates the output layer parameters and the rule layer parameters of the decoupling model based on the joint loss function to obtain the difference adaptation model.

[0157] It should be noted that the above embodiments provide devices that implement their functions using only the division of the above functional modules as examples. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0158] This application also provides an electronic device. Figure 4 , Figure 4 4 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. The electronic device 400 may include: at least one processor 401, at least one network interface 404, a user interface 403, a memory 405, and at least one communication bus 402.

[0159] The communication bus 402 is used to implement the connection and communication between these components.

[0160] The user interface 403 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 403 may also include a standard wired interface and a wireless interface.

[0161] The network interface 404 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0162] Processor 401 may include one or more processing cores. Using various interfaces and circuits, processor 401 connects to various components within the server. It executes instructions, programs, code sets, or instruction sets stored in memory 405, as well as accesses data stored in memory 405, to perform various server functions and process data. Optionally, processor 401 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). Processor 401 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may also be implemented independently of the processor 401 and implemented as a separate chip.

[0163] Among them, the memory 405 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 405 includes a non-transitory computer-readable storage medium. The memory 405 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 405 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 405 may also optionally be at least one storage device located away from the aforementioned processor 401. Refer to Figure 4 , the memory 405 as a computer storage medium may include an operating system, a network communication module, a user interface module and an application program of a financial document processing method.

[0164] exist Figure 4In the electronic device 400 shown, the user interface 403 is mainly used to provide an input interface for the user and obtain the data input by the user; and the processor 401 can be used to call an application program for storing a financial document processing method in the memory 405. When executed by one or more processors 401, the electronic device 400 executes one or more of the methods described in the above embodiments. It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited to the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application.

[0165] The present application further provides a computer-readable storage medium storing instructions, which, when executed by one or more processors 401 , enable the electronic device 400 to perform one or more of the methods described in the above embodiments.

[0166] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0167] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic, such as the division of units, which is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0168] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0169] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0170] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application. The aforementioned memory includes various media that can store program code, such as USB flash drives, mobile hard drives, magnetic disks, or optical disks.

[0171] The foregoing is merely an exemplary embodiment of the present disclosure and is not intended to limit the scope of the present disclosure. In other words, any equivalent variations and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the disclosure and the practical implications thereof.

[0172] This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not described herein. The description and examples are to be considered as exemplary only, and the scope and spirit of the present disclosure are to be defined by the claims.

Claims

1. A financial document processing method, characterized in that: The method comprises: Obtaining the original image of the financial document to be processed; Monitor the version update status of the financial business rule library. When a rule change instruction that triggers a preset change condition is detected, obtain the changed target business rule and key fields and difference area descriptors related to the target business rule. The key fields include the changed fields and the dependent fields associated with the changed fields in the preset cross-reference relationship table. The difference area descriptors include field position offset information and newly added field offset information. According to the key field, matching the corresponding target template in the preset reference template library; Reconstructing the target template according to the target business rule and the difference area descriptor to obtain a first reconstructed template; Identifying the financial document based on the first reconstruction template to obtain initial financial data; Determining whether the initial financial data satisfies the preset cross-reference relationship table; If it is determined that the initial financial data satisfies the preset cross-reference relationship table, the initial financial data is output; If it is determined that the initial financial data does not satisfy the preset cross-reference relationship table, the first reconstruction template is revised twice to obtain a second reconstruction template and corresponding target financial data, and the target financial data is output.

2. The method according to claim 1, characterized in that The reconstructing the target template according to the target business rule and the difference area descriptor to obtain a first reconstructed template specifically includes: Calculating a coordinate offset vector for each field according to the field position offset information, and adding the coordinate offset vector to a positioning frame parameter of the corresponding field in the target template to update the position of the field positioning frame; Determine a candidate positioning area for the new field based on the new field offset information and in combination with semantic constraints related to the new field in the target business rule through context analysis, extract local features within the candidate positioning area, and identify table lines to generate a positioning frame for the new field; Based on the newly added field verification rules in the target business rules, construct the associated verification logic between the fields, and add the associated verification logic to the verification rule set of the target template; The field positioning box, the positioning box of the newly added field, and the associated verification logic are integrated into the target template to generate the first reconstructed template.

3. The method according to claim 1, characterized in that The second modification of the first reconstruction template to obtain a second reconstruction template and corresponding target financial data specifically includes: According to the preset cross-reference relationship table, determining abnormal fields in the initial financial data that do not satisfy the cross-reference relationship; Locating an abnormal region corresponding to the abnormal field in the first reconstruction template, and obtaining position coordinates of the abnormal region in the financial document image; Based on the position coordinates, collecting a context area adjacent to the abnormal area in the financial document image, and performing extended recognition on the context area to obtain an extended recognition result; According to the extended recognition result, determining whether there are interference factors that affect the accuracy of the abnormal field recognition, the interference factors including field overlap, field breakage, and field tilt; If the interference factor exists, the recognition result of the abnormal area is corrected according to the preset compensation rule to obtain corrected abnormal field data; Updating the corrected abnormal field data into the initial financial data to obtain intermediate financial data; Determining whether the intermediate financial data satisfies the preset cross-reference relationship table; If the intermediate financial data satisfies the preset cross-reference relationship table, the intermediate financial data is determined as target financial data, and the first reconstruction template is adjusted according to the extended recognition result to obtain the second reconstruction template.

4. The method according to claim 3, characterized in that After determining whether the intermediate financial data satisfies the preset cross-reference relationship table, the method further includes: If the intermediate financial data does not satisfy the preset cross-reference relationship table, identifying conflicting fields in the intermediate financial data that do not satisfy the preset cross-reference relationship table, and extracting conflicting coordinate regions corresponding to the conflicting fields in the first reconstruction template; Calculating the theoretical coordinate position of the conflicting field according to the constraint rules related to the conflicting field in the preset cross-reference relationship table, and correcting the positioning frame of the conflicting field in the first reconstruction template to the theoretical coordinate position; In the conflict coordinate area, new fields are determined through text semantic analysis, and the new fields are added to the positioning rules of the first reconstruction template to generate the second reconstruction template. The financial document is identified according to the second reconstruction template to obtain the target financial data.

5. The method according to claim 1, wherein When a rule change instruction that triggers a preset change condition is detected, obtaining the changed target business rule and key fields and difference area descriptors related to the target business rule specifically includes: Detecting the version number of the financial business rule library according to a preset time interval; Compare the detected version number with the locally stored version number to obtain the version number difference value; If the version number difference is greater than a preset threshold, it is determined that the version of the financial business rule library has been updated, and a rule change instruction of the preset change condition is triggered; When the rule change instruction is detected, the change log file of the financial business rule library is obtained, the change record is identified, and the change record is extracted to obtain the changed target business rule, which includes the change content and the change rule; Performing semantic analysis on the change content to determine key fields related to the change rule; The change rule is semantically parsed to determine the field position offset information and the newly added field offset information of the key field.

6. The method according to claim 1, characterized in that The method further comprises: Using the target template to identify the original image to obtain basic financial data; Based on the target template, performing targeted training on the key fields and the difference region descriptors to generate a difference adaptation model; Applying the difference adaptation model to the original image of the financial document, performing targeted recognition on the original image, and obtaining a difference recognition result; The difference identification result is integrated with the basic financial data, and the changed fields and newly added fields in the basic financial data are updated to obtain the initial financial data.

7. The method according to claim 6, characterized in that The step of performing targeted training on the key fields and the difference region descriptors based on the target template to generate a difference adaptation model specifically includes: Extracting corresponding difference image regions from the original image of the financial document according to the key fields and the difference region descriptors to generate a difference training set; Loading the pre-trained OCR model corresponding to the target template as the base model, freezing the general feature extraction layer parameters unrelated to text recognition in the base model, decoupling the output layer parameters related to field positioning and the rule layer parameters related to business verification logic in the base model, and obtaining a decoupled model; Based on the difference training set, the decoupling model is fine-tuned to generate the difference adaptation model.

8. The method according to claim 7, characterized in that The step of fine-tuning the decoupling model based on the difference training set to generate the difference adaptation model specifically includes: Inputting the difference training set into the decoupling model to obtain an output layer feature map and a rule layer feature map of the decoupling model; Calculating a target loss in the output layer feature map according to the field position offset information and the newly added field offset information, wherein the target loss includes a changed field positioning loss and a newly added field positioning loss, wherein the changed field positioning loss is calculated based on a coordinate offset vector of the changed field, and the newly added field positioning loss is calculated based on local features and table lines within a candidate positioning area; According to the newly added field verification rule in the target business rule, the association loss is calculated in the rule layer feature graph, where the association loss is calculated based on the semantic dependency relationship between the newly added field and the associated field; A joint loss function is constructed, where the joint loss function includes the target loss and the association loss, and gradient optimization and update are performed on the output layer parameters and the rule layer parameters of the decoupling model based on the joint loss function to obtain the difference adaptation model.

9. A financial document processing device, characterized in that: The device comprises an original image acquisition module (301), a difference determination module (302), a target template matching module (303), a template reconstruction module (304), a cross-reference relationship processing module (305) and a secondary reconstruction module (306), wherein: The original image acquisition module (301) is used to acquire the original image of the financial document to be processed; The difference determination module (302) is used to monitor the version update status of the financial business rule library. When a rule change instruction that triggers a preset change condition is detected, the module obtains the changed target business rule and key fields and difference area descriptors related to the target business rule. The key fields include the changed fields and the dependent fields associated with the changed fields in the preset cross-reference relationship table. The difference area descriptors include field position offset information and newly added field offset information. The target template matching module (303) is used to match the corresponding target template in a preset reference template library according to the key field; The template reconstruction module (304) is used to reconstruct the target template according to the target business rule and the difference area descriptor to obtain a first reconstructed template; The template reconstruction module (304) is further configured to identify the financial document based on the first reconstruction template to obtain initial financial data; The cross-reference relationship processing module (305) is used to determine whether the initial financial data satisfies the preset cross-reference relationship table; The cross-reference relationship processing module (305) is further configured to output the initial financial data if it is determined that the initial financial data satisfies the preset cross-reference relationship table; The secondary reconstruction module (306) is used to perform secondary correction on the first reconstruction template if it is determined that the initial financial data does not satisfy the preset cross-reference relationship table, to obtain a second reconstruction template and corresponding target financial data, and output the target financial data.

10. An electronic device, characterized in that: The electronic device (400) comprises a processor (401), a memory (405), a user interface (403) and a network interface (404), wherein the memory (405) is used to store instructions, the user interface (403) and the network interface (404) are used to communicate with other devices, and the processor (401) is used to execute the instructions stored in the memory (405) so that the electronic device (400) executes the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Certificate information identification method and device, terminal equipment and storage medium

    CN112348008A

  • Picture reconstruction method based on bill template

    CN112529989A