Method and device for identifying finance and tax bills and storage medium

By introducing a multi-agent system and utilizing pre-defined agents for element extraction and review evaluation, the AI-powered financial and tax document recognition system has overcome the difficulties in recognizing documents with non-standard formats and damage, thus achieving efficient and accurate recognition of financial and tax documents.

CN121434733APending Publication Date: 2026-01-30AISINO CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511596108.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing AI-powered financial and tax recognition systems struggle to accurately extract seller and buyer data, as well as unit price and quantity, when dealing with invoices that are not formatted correctly, have missing information fields, are damaged, or are obscured, resulting in low recognition efficiency and accuracy.

Method used

A pre-defined element extraction agent is introduced to extract key elements. A self-encoded agent performs post-hoc detection. An auditing and evaluation agent queries and verifies data on the external Internet or a dedicated network. External interaction is only initiated when internal detection fails. The recognition accuracy is improved through a multi-agent system.

Benefits of technology

It improves the accuracy and efficiency of tax and financial invoice recognition, reduces the frequency of external interaction of system resources, and can solve problems that are difficult for traditional AI to handle, just like human experts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121434733A_ABST
    Figure CN121434733A_ABST
Patent Text Reader

Abstract

The invention provides a finance and tax bill identification method and device and a storage medium, and the method comprises the steps: extracting key elements of a to-be-identified finance and tax bill through a preset element extraction agent, and generating a structured internal identification bill based on a preset bill format according to the extracted key elements through a self-coding agent, comparing at least one type of element information in the key elements with a preset finance and tax data knowledge database, and performing posterior detection; when the detection fails, a failure notice is sent to the auditing and evaluating agent; and according to the element information in the failure notification, after determining the economic field to which the finance and taxation bill document to be identified belongs by using the auditing and evaluating agent, querying other information corresponding to the element information in the external internet or the private network as verification data, and distinguishing, correcting or confirming the element information which fails in posterior detection in the internal identification bill document, so as to identify the finance and taxation bill document to be identified. And obtaining final identification information. And the identification efficiency and accuracy of the finance and tax bills are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus and storage medium for identifying financial and tax documents. Background Technology

[0002] The financial and tax system is one of the core systems of enterprise management, integrating financial management and tax management functions. With the development of enterprise informatization, how to efficiently and accurately digitize massive amounts of paper or electronic invoices and documents (such as VAT invoices, contract texts, etc.) is a key factor restricting the intelligentization of the financial and tax system.

[0003] While current AI-based tax and financial document recognition systems have made some progress in extracting various types of invoice data, several intractable technical problems remain. For example, when invoices are formatted incorrectly or missing information fields, the system frequently confuses the seller's (name, address, phone number) and buyer's (name, address, phone number) data. Secondly, when invoices display data formats like "AB" (e.g., the amount field is written as "10050.00"), current systems struggle to accurately determine which represents the unit price and which represents the quantity. Furthermore, when invoices are damaged, obscured, or of poor print quality, the system often fails to correctly extract key data such as the product name. To address these bottlenecks in AI recognition, the industry has begun introducing multi-agent systems. However, current agent-based invoice recognition systems still face bottlenecks in efficiency and accuracy when dealing with complex issues such as seller / buyer confusion, unit price / quantity confusion, and damaged invoices. Therefore, a new technical solution is urgently needed to address these problems and further improve the recognition rate of tax and financial documents. The significant reduction in these issues has significantly impacted the system's recognition efficiency and accuracy. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus, and storage medium for identifying financial and tax documents, so as to at least partially solve the above-mentioned problems.

[0005] In a first aspect, embodiments of this application provide a method for identifying tax and financial documents, including: The key elements of the financial and tax invoices to be identified are extracted by a preset element extraction intelligent agent. The key elements include at least the seller data, buyer data, product name, quantity and unit price information. After receiving the key elements through the self-encoding intelligent agent, a structured internal identification invoice is generated based on the preset internal invoice format, and at least one type of element information in the key elements is compared with a preset financial and tax data knowledge database to perform post-testing on the key elements. When the posterior detection fails, a failure notification containing information about the elements that caused the posterior detection failure is sent to the audit and evaluation agent. Based on the information elements in the failure notification, the economic sector to which the tax and financial invoice to be identified belongs is determined using the audit and evaluation agent. Based on the determined economic information, the audit and evaluation agent uses the external Internet or a dedicated network to query other information corresponding to the element information as verification data to distinguish, correct or confirm the element information that failed the post-detection in the internal identification invoice, and obtain the final identification information; wherein, the other information includes at least the main business information of the seller or buyer, or the price information corresponding to the product name.

[0006] Secondly, based on the method for identifying tax and financial documents described in the first aspect of this application, embodiments of this application also provide an apparatus for identifying tax and financial documents, comprising: The extraction module is used to extract key elements of the financial and tax invoices to be identified by a preset element extraction intelligent agent. The key elements include at least the seller data, buyer data, product name, quantity and unit price information. The self-encoding module is used to receive the key elements through the self-encoding intelligent agent, generate a structured internal identification invoice based on a preset internal invoice format, and compare at least one type of element information in the key elements with a preset financial and tax data knowledge database to perform post-testing on the key elements. The evaluation sending module is used to send a failure notification containing element information that caused the failure of the posterior detection to the review and evaluation agent when the posterior detection fails. The review module is used to determine the economic sector to which the tax and financial invoice to be identified belongs based on the information in the failure notification and the review evaluation agent. The confirmation module is used to, based on the determined economic field information, use the audit and evaluation intelligent agent to query other information corresponding to the element information in the external Internet or dedicated network as verification data, distinguish, correct or confirm the element information that failed the post-detection in the internal identification invoice, and obtain the final identification information; wherein, the other information includes at least the main business information of the seller or buyer, or the price information corresponding to the product name.

[0007] Thirdly, embodiments of this application also provide a computer storage medium storing computer-executable instructions, which, when executed, perform any of the methods for identifying financial and tax invoices as described in the first aspect of embodiments of this application.

[0008] Fourthly, embodiments of this application also provide an electronic device, including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction that causes the processor to perform any of the methods for identifying tax invoices as described in the first aspect of the embodiments of this application.

[0009] This application provides a method, apparatus, and storage medium for identifying tax and financial invoices, comprising: extracting key elements of the tax and financial invoice to be identified through a preset element extraction intelligent agent, wherein the key elements include at least seller data, buyer data, product name, quantity, and unit price information; receiving the key elements through a self-encoding intelligent agent, generating a structured internal identification invoice based on a preset internal invoice format, and comparing at least one type of element information from the key elements with a preset tax and financial data knowledge database to perform post-hoc detection on the key elements; and when the post-hoc detection fails, sending a request to an audit and evaluation intelligent agent. The system sends a failure notification containing information about the elements that caused the subsequent detection to fail. Based on the element information in the failure notification, the audit and evaluation agent determines the economic sector to which the tax invoice to be identified belongs. Based on the determined economic sector information, the audit and evaluation agent queries other information corresponding to the element information on the external Internet or a dedicated network as verification data to distinguish, correct, or confirm the element information that failed the subsequent detection in the internally identified invoice, and obtain the final identification information. The other information includes at least the main business information of the seller or buyer, or the price information corresponding to the product name. This application solves the problem by introducing a subsequent detection mechanism and using knowledge database comparison. The audit and evaluation agent only initiates external interaction (at a relatively high cost) when internal detection fails. This greatly reduces the frequency of interaction between the audit and evaluation agent and the outside world, improving the overall system response speed. The audit and evaluation agent has the ability to judge the economic sector and the interaction of information within and outside the sector. This enables it to accurately solve problems that traditional AI cannot handle, such as confusion between sellers and buyers, confusion between unit price and quantity, and poor recognition of corrupted invoices, by leveraging external objective knowledge (such as the company's main business and commodity market prices), just like human experts. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0011] Figure 1 A schematic diagram illustrating the workflow of a method for identifying financial and tax invoices provided in this application embodiment; Figure 2 A schematic diagram of the structure of a device for identifying financial and tax documents provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.

[0013] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.

[0014] Example 1 This application provides a method for identifying financial and tax invoices, such as... Figure 1 As shown, Figure 1 This illustration shows a flowchart of a method for identifying tax and financial documents according to an embodiment of this application, including: Step S101: Extract key elements from the tax invoice to be identified using a preset element extraction agent. These key elements include at least seller data (name, address, telephone number), buyer data (name, address, telephone number), product name, quantity, and unit price information. In this embodiment, the element extraction agent is the system's perception layer, responsible for element extraction from the invoice. It is typically a module combining deep learning-based OCR (Optical Character Recognition) and NER (Named Entity Recognition) models, capable of efficiently and accurately extracting text blocks from the tax invoice image and identifying key fields as key information. This step, by digitizing the tax invoice to be identified, is the first step of the method, aiming to obtain structured basic data from the original invoice image. The agent analyzes the input image and extracts the clearly defined fields (element information) required for subsequent steps. This provides clear and explicit data input for subsequent post-testing and auditing. Furthermore, in the actual application scenarios of this embodiment, key element information may also include amounts, dates, invoice code numbers, etc. The embodiments of this application will not be described in detail here.

[0015] Step S102: After receiving the key elements through the autoencoder agent, a structured internal identification document is generated based on a preset internal document format. At least one type of element information from the key elements is compared with a pre-set financial and tax data knowledge database to perform post-hoc detection on the key elements. In this embodiment, the autoencoder agent is the system's processing and verification layer, primarily responsible for re-encoding the extracted element information. This step in this embodiment adds the function of comparing and judging key information after reconstruction. The reconstructed autoencoder agent in this solution is mainly responsible for formatting, internally validating, and maintaining the knowledge base of the extracted element information data. For example, it can determine the types of various invoices, documents, and contract texts, determine accounting entries and select the correct internal document format, fill in the key information extracted by the element extraction agent into the internal identification document, and fill in the key information extracted by the element extraction agent into the structured internal identification document to be generated. In the process of performing posterior detection, the pre-set financial and tax data knowledge base is generally a verified financial and tax database, which is a highly reliable internal dataset accumulated through long-term use (especially manual review and feedback). This embodiment of the application is limited to checking the consistency of extracted information after information extraction using a database maintained by the self-encoded agent that stores highly reliable business data. This detection, utilizing deterministic knowledge accumulated within the system, enables rapid and low-cost verification of extracted data. This embodiment uses the "supplier" field as an example: the self-encoded agent maintains a "qualified supplier" database. When the element information extracted by the element extraction agent contains a supplier name, the self-encoded agent compares it with this database. If the comparison is successful, the element information is judged to be correct, and the generated structured internal identification document information is considered accurate and compliant; if the comparison fails (e.g., this is a new supplier, or the name is incorrectly identified due to contamination), the self-encoded agent considers the element information "questionable." This posterior detection method greatly improves the coding accuracy of the self-encoding agent and reduces the implementation difficulty of the agent, thus providing an efficient filter for the intermediate processing of this scheme.

[0016] Step S103: When the posterior detection fails, a failure notification containing information about the elements that caused the failure is sent to the review and evaluation agent. This step is crucial for the hierarchical processing implemented in this scheme. The aim is to concentrate system resources (especially the expensive external interaction resources of the review agent) on resolving complex issues. This achieves on-demand problem handling, notifying the review agent only when a comparison fails, and submitting only complex issues that cannot be resolved internally to the review agent for processing. This minimizes the frequency of interaction between the review agent and the outside world, thereby improving the overall system response speed.

[0017] In this embodiment, when the posterior detection fails, the process of sending a failure notification containing information about the elements that caused the posterior detection failure to the review and evaluation agent can be specifically as follows: The system adopts a decentralized communication structure. After the posterior detection fails, the self-encoding agent packages the questionable data (such as ticket images, extracted fields, and reasons for failure) and sends it to the review and evaluation agent through an internal communication protocol. This embodiment here limits the use of a decentralized communication structure; data is not all aggregated in one center, but rather transmitted point-to-point between agents according to task requirements. For example, from extraction agent -> self-encoding agent, self-encoding agent -> review and evaluation agent, and from review and evaluation agent -> self-encoding agent / extraction agent. This communication mechanism integrates reflection, planning, and collaboration into a single system, which is more conducive to the efficient implementation of multiple data identification iterations and is also the foundation for further improving the system performance of this solution.

[0018] Step S104: Based on the element information in the failure notification, the Review and Evaluation Agent determines the economic sector to which the tax invoice to be identified belongs. In this embodiment, the Review and Evaluation Agent relies on its own computing resources and tax data knowledge database to determine the economic sector of the tax invoice to be identified based on information such as the seller, buyer, and product name in the invoice. This step aims to define the scope for subsequent external information queries, narrowing the search range and avoiding ineffective broad searches across the entire internet, thus laying the foundation for accurate verification.

[0019] Step S105: Based on the determined economic field information, the auditing and evaluation agent queries other information corresponding to the element information in the external Internet or dedicated network as verification data. This distinguishes, corrects, or confirms the element information that failed the post-detection in the internal identification document, obtaining the final identification information. The other information includes at least the main business information of the seller or buyer, or the price information corresponding to the product name. The purpose of this step is to obtain external objective evidence to resolve problems that cannot be determined internally (such as seller / buyer confusion, price / quantity confusion). The auditing agent utilizes its interactive capabilities to exchange and collaborate with the outside world, initiating targeted queries to the external Internet or dedicated network (such as business information databases, industry pricing networks) within the corresponding economic field. This resolves problems such as seller and buyer data confusion, or unit price and quantity confusion leading to post-detection failure, significantly improving the system's identification efficiency.

[0020] Optionally, in one implementation of this application embodiment, based on determined economic field information, the audit and evaluation intelligent agent queries other information corresponding to the element information on the external Internet or a dedicated network as verification data to distinguish, correct, or confirm the key information. This includes: using the audit and evaluation intelligent agent to query the main business information and upstream and downstream enterprise information of the enterprises corresponding to the seller data and buyer data within the determined economic field, and / or querying the market price information corresponding to the product name; distinguishing the seller enterprise name and buyer enterprise name of the tax invoice to be identified based on the queried main business information, upstream and downstream information, and product name; and confirming the quantity and unit price information after distinguishing based on the queried price information. This step in this embodiment further defines the preferred specific implementation logic for solving problems such as "seller / buyer confusion". That is, obtaining the business roles (main business, position in the industrial chain) of both parties through external query, and matching and judging them in combination with the goods on the invoice. For example, if a query shows that Company A's main business is "selling cement" and Company B's main business is "construction," and the invoice lists "cement" as the commodity, then Company A is confirmed as the seller and Company B as the buyer. Another example is an invoice with data of "10 * 500" and the commodity as "cement." If a query shows the market price of cement is approximately 500 yuan / ton, then 500 can be identified as the unit price and 10 as the quantity. This embodiment of the application uses commercial substance rather than the invoice's position to more quickly distinguish between the seller and buyer, effectively solving the core problems of seller / buyer confusion and unit price / quantity confusion in existing technologies, thereby improving the system's identification efficiency.

[0021] Optionally, in one implementation of this application embodiment, the method further includes: when the field information of the enterprise name or product name in the extracted key elements is incomplete, performing a fuzzy search on the relevant economic field information to locate the enterprise name or product name. This step further defines the specific technical logic for solving the problem of soiled paper tickets. The already determined economic field is used to further narrow down the scope of the fuzzy search. Even within a limited field, a fuzzy search can quickly locate the soiled name, thereby improving the robustness of the identification and better addressing information extraction defects such as incomplete OCR recognition caused by incomplete or soiled tickets.

[0022] Optionally, in one embodiment of this application, the method further includes: submitting the element information, after correction or confirmation by the review and evaluation agent, to a human for review and confirmation to obtain a review result; using the review result as an agent reward value and feeding it back to the self-encoded agent, for classification and storage or updating of the financial and tax data knowledge database according to the review result, and for guiding the process of extracting key elements after the element extraction agent. This step in this embodiment of the application defines a continuous learning mechanism for a human-machine collaborative loop, establishing a feedback loop for the system's self-evolution. The aim is to ensure that the final data is more accurate and to optimize the system using the human confirmation result. This step uses the manually confirmed element data, considered accurate and error-free label data, and feeds it back to the self-encoded agent to update its internally maintained financial and tax data knowledge database (e.g., adding new suppliers to the qualified list). As the compliance information in the internal underlying knowledge database maintained by the self-encoded agent becomes increasingly complete, the overall system's operating efficiency and recognition accuracy will be continuously and significantly improved, enabling it to break through the traditional recognition rate bottleneck.

[0023] Optionally, in one embodiment of this application, comparing at least one type of element information from the key elements with a pre-set financial and tax data knowledge database includes: using the self-encoded intelligent agent to extract the buyer's name from the key elements; retrieving a subset of key field data corresponding to the extracted buyer's name from the financial and tax data knowledge database; comparing the subset of key field data with the buyer's name; and determining the comparison result as the result of the post-test for the key elements. To address scenarios involving providing financial and tax services to multiple companies, the self-encoded intelligent agent first extracts the buyer's name from the invoice, and then only retrieves the internal invoice format and key field sub-knowledge base (i.e., a subset of field data; for example, a company's product name database) associated with that buyer company. This participates in subsequent post-testing. This comparison, performed only in the retrieved local knowledge base, allows post-testing to only compare with strongly correlated data with a small number of entries, reducing the complexity of global knowledge queries to local query complexity. This solves the problem of low system efficiency caused by the highly dispersed and rapidly increasing number of product names and invoice formats in a multi-customer environment. Furthermore, since the number of entries in the buyer's knowledge base is relatively small and highly relevant in real-world applications, this embodiment of the application further limits this step to a design scheme based on buyer name classification. Based on the buyer's name, all invoice formats within that company are retrieved. Then, a posterior detection method is used to classify and determine the invoice type, thereby generating the correct identified invoice. This allows the posterior detection to operate more accurately and efficiently when the number of data entries is small, further improving the system's operating efficiency and relevance in a multi-tenant environment.

[0024] Optionally, in one embodiment of this application, before the audit and evaluation agent queries other information corresponding to the element information in the external Internet or a dedicated network as verification data to distinguish, correct, or confirm the element information, the method further includes: setting the features perceived by the audit and evaluation agent of the current internal identification document as the state space, combining the prediction results of the internal identification document by a preset deep learning model, and calculating the expected benefits of the generated posterior detection results of the internal identification document as "judged as compliant", "judged as non-compliant", and "marked for manual review" according to the reward function formula of reinforcement learning, wherein the reward function formula is: R =α Accuracy - β Cost of misjudgment Where R is the calculated expected return, and α and β are weighting coefficients set based on the business scenario, used to quantify the returns of different decision actions; The review and evaluation agent selects the action with the highest expected benefit to output, and the action space includes "determine compliance", "determine non-compliance" and "mark for manual review"; If the selected action output is "marked for manual review", the audit and evaluation agent uses its interactive capabilities to exchange and collaborate with the external Internet or dedicated network to query and obtain other information within the defined economic field, iteratively correct the key elements in the internal identification document, and regenerate a new internal identification document.

[0025] In this embodiment, the state space of the intelligent agent controlling the review and evaluation process includes the features of the current data to be identified (such as invoice amount, transaction type, and coding confidence) and historical identification results. The agent combines the prediction results of historical data using a deep learning model (such as a CNN-LSTM hybrid model) and calculates the expected returns for the three actions—"judging compliance," "judging non-compliance," and "marking for manual review"—based on the aforementioned reward function formula. This approach maximizes identification accuracy as the objective function of reinforcement learning while incorporating the cost of misjudgment (β) in the business scenario. The cost of misjudgment is used as a penalty. This shifts the decision-making objective of the review and evaluation agent (selecting the action with the highest expected return) from simple technical accuracy to optimal economic benefit driven by risk management. This ensures that the agent can choose the most economically valuable action when faced with uncertain data, avoiding high-risk, high-cost misjudgments. When the review agent calculates that the action of "marking for human review" has the highest return, the system does not immediately proceed to the human review stage. Instead, it uses its external information interaction capabilities to communicate with the external network environment to obtain additional information to correct the data. The corrected data is then used to regenerate the internal identification slip. This leverages external information to enhance the agent's cognition, attempting to guide suboptimal solutions to optimal solutions without human intervention, thereby overcoming the recognition bottlenecks of traditional agents and further improving the recognition efficiency and autonomous problem-solving capabilities of systems implementing this method.

[0026] Furthermore, the method also includes: introducing a query cost factor γ into the reward function, which is used to quantify the time or API call cost required to perform external information exchange, i.e., the reward function is calculated using the following formula: R = (α) Accuracy) - (β) (Cost of misjudgment) - (γ) (Query cost) This embodiment of the application further quantifies the benefits or penalties gained by the review and evaluation agent after performing an action using the above-described method. This enables the review and evaluation agent to "autonomously decide the strategy for reviewing each key field." For example, it learns that in certain situations (such as severe contamination), the total benefit (considering query costs) of directly marking it as "awaiting manual review" may be higher than attempting a high-cost, low-success-rate "external fuzzy search." This makes the system's decision-making more intelligent and economical. That is, the agent not only learns how to judge correctly, but also how to make judgments more economically, optimizing the system's resource consumption (time and monetary costs) while ensuring recognition accuracy.

[0027] Optionally, in one implementation of an embodiment of this application, the method further includes: optimizing and adjusting the action value function parameters of the review and evaluation agent based on the decision error determined by manual review results using a reinforcement learning algorithm. This step in this embodiment further specifies calculating the decision error of the review agent by collecting the results of manual review (considered accurate). A reward value is calculated using a reinforcement learning algorithm, and the action value function parameters are adjusted sequentially to update the agent's behavioral strategy. This continuous learning based on a feedback mechanism enables the agent's strategy to continuously adapt to real business scenarios and changes in financial and tax rules. It achieves a further closed loop in the feedback of manual confirmation information, fully utilizes system resources, continuously improves and maintains the underlying knowledge database, and ensures continuous improvement in system performance during use.

[0028] Optionally, in one embodiment of this application, the method further includes: evaluating the deep learning model using F1-score as a system performance indicator based on the decision error determined by manual review; and retraining the deep learning model on all data when the determined F1-score is lower than a preset threshold. This step in this embodiment limits the implementation system of this method to periodically use F1-score (the harmonic mean of precision and recall) as a performance evaluation indicator. When the F1-score is lower than a preset threshold (e.g., 0.85), the system automatically triggers full data retraining to monitor the system's generalization ability and robustness, and to address potential concept drift or changes in data distribution. This provides an effective automated operation and maintenance mechanism to ensure recognition performance, ensuring that system performance indicators are continuously and stably maintained at a high level.

[0029] In the practical application environment of this application embodiment, the information extraction agent, the self-encoding agent, and the review and evaluation agent are all agents constructed based on reinforcement learning, and all use "maximizing recognition accuracy" as the objective function, combined with a deep learning model (such as a CNN-LSTM hybrid model) to analyze the prediction results of historical data. They can also calculate the expected return based on the reward function formula "R=α×accuracy-β×misjudgment cost" and select the next method step to be implemented based on this return. This ensures the efficiency and accuracy of the aforementioned steps.

[0030] Optionally, in a preferred implementation of this application embodiment, key elements of the tax invoice to be identified are extracted through a preset element extraction agent. Specifically, this can be achieved by controlling the element extraction agent to employ multi-layer feature extraction and fusion processing when processing the invoice image, and adjusting feature maps of different sizes to a fixed dimension through adaptive average pooling, while combining a dynamic attention mechanism to extract key elements. This application embodiment enhances the model's (element extraction agent's) ability to understand and extract information from tilted, deformed text and complex table structures in this way. Traditional element extraction relies on standard deep convolutional networks, which have poor adaptability to complex image structures (such as tilted text and complex tables) and invoices of different sizes. This application embodiment here introduces advanced image feature processing technology, and adaptive average pooling solves the problem of inconsistent input feature dimensions for invoices of different sizes, ensuring the uniformity of subsequent processing. At the same time, the dynamic attention mechanism guides the model to focus on key text and structures in the image, improving the ability to recognize low-quality or deformed text. The implementation of this technique effectively reduces initial errors caused by insufficient image processing, alleviates the verification burden on subsequent self-encoding agents and audit evaluation agents, optimizes the overall system's resource consumption, and significantly enhances its adaptability to complex invoice scenarios.

[0031] Optionally, in a preferred implementation of this application embodiment, determining the economic sector to which the tax invoice belongs based on the element information in the failure notification using the audit and evaluation agent includes: inputting the extracted "product name" or "seller / buyer name" key fields into a pre-trained semantic classification model (e.g., a text classifier based on the BERT model), wherein the semantic classification model is trained based on massive industry classification standard texts (such as national economic industry classification codes) and enterprise business scope data; determining the classification result (e.g., an industry code or label) of the economic sector to which the invoice belongs using the semantic classification model, and the audit and evaluation agent uses the classification result as the economic sector to which the invoice belongs. This embodiment designs a process based on a deep learning NLP model to perform a quantifiable classification task to determine the economic sector, ensuring the objectivity and robustness of the sector judgment, avoiding errors and omissions based on simple rule judgments, providing a more accurate range for subsequent intra-sector queries, thereby improving the quality of external verification data and query success rate.

[0032] This application provides a method for identifying tax invoices. The method involves using a pre-defined element extraction agent to extract key elements from the tax invoice to be identified. These key elements include at least seller data, buyer data, product name, quantity, and unit price information. After receiving these key elements, a self-encoding agent generates a structured internal identification invoice based on a pre-defined internal invoice format. At least one type of element information from the key elements is compared with a pre-defined tax data knowledge database to perform a post-test. If the post-test fails, a failure notification containing the element information causing the failure is sent to an auditing and evaluation agent. Based on the element information in the failure notification, the auditing and evaluation agent determines the economic sector to which the tax invoice belongs. Based on the determined economic sector information, the auditing and evaluation agent queries other information corresponding to the element information on an external internet or dedicated network as verification data to differentiate, correct, or confirm the element information in the internal identification invoice that failed the post-test, obtaining final identification information. The other information includes at least the main business information of the seller or buyer, or the price information corresponding to the product name. This application introduces a post-hoc detection mechanism, utilizing knowledge database comparison to solve problems. Only when internal detection fails is the (relatively high-cost) auditing agent activated for external interaction. This significantly reduces the frequency of interaction between the auditing agent and the outside world, improving the overall system's response speed. The auditing and evaluation agent possesses the ability to judge information interactions within and outside the economic domain. This allows it to, like a human expert, leverage external objective knowledge (such as a company's main business and commodity market prices) to accurately solve problems that traditional AI cannot handle, such as seller / buyer confusion, unit price / quantity confusion, and identification of corrupted invoices.

[0033] Example 2 Based on the method for identifying tax and financial invoices provided in Embodiment 1 of this application, this embodiment also provides a corresponding apparatus for identifying tax and financial invoices, such as... Figure 2 As shown, Figure 2 A schematic diagram of the structure of a device 20 for identifying tax and financial documents provided in this application embodiment. The device 20 for identifying tax and financial documents includes: The extraction module 201 is used to extract key elements of the financial and tax invoices to be identified through a preset element extraction intelligent agent. The key elements include at least the seller data, buyer data, product name, quantity and unit price information. The self-encoding module 202 is used to receive the key elements through the self-encoding intelligent agent, generate a structured internal identification invoice based on a preset internal invoice format, and compare at least one type of element information in the key elements with a preset financial and tax data knowledge database to perform post-testing on the key elements. The evaluation sending module 203 is used to send a failure notification containing element information that caused the failure of the posterior detection to the review and evaluation agent when the posterior detection fails. The audit module 204 is used to determine the economic sector to which the tax and financial invoice to be identified belongs based on the element information in the failure notification and the audit evaluation agent. The confirmation module 205 is used to, based on the determined economic field information, use the audit and evaluation intelligent agent to query other information corresponding to the element information in the external Internet or dedicated network as verification data, distinguish, correct or confirm the element information that failed the post-detection in the internal identification invoice, and obtain the final identification information; wherein, the other information includes at least the main business information of the seller or buyer, or the price information corresponding to the product name.

[0034] Optionally, in one implementation of this application embodiment, the confirmation module 205 is further configured to use the audit and evaluation intelligent agent to query the main business information and upstream and downstream enterprise information of the enterprises corresponding to the seller data and buyer data within the determined economic field, and / or query the market price information corresponding to the product name; based on the queried main business information, upstream and downstream information and product name, distinguish the seller enterprise name and buyer enterprise name of the tax invoice to be identified; and based on the queried price information, distinguish the quantity and unit price information.

[0035] Optionally, in one embodiment of this application, when the field information of the enterprise name or product name in the extracted key elements is incomplete, the confirmation module 205 is further used to perform a fuzzy search on the relevant economic field information using the audit and evaluation intelligent agent to locate the enterprise name or product name.

[0036] Optionally, in one embodiment of this application, the device 20 further includes: a review module (not shown in the figures), which is used to: submit the element information after correction or confirmation by the audit and evaluation agent to a human for review and confirmation, and obtain a review result; use the review result as an agent reward value, feed it back to the self-encoding agent for classification, storage or updating of the financial and tax data knowledge database according to the review result, and for guiding the process of extracting key elements after the element extraction agent.

[0037] Optionally, in one embodiment of this application, the self-encoding module 202 is further configured to: extract the buyer's name from the key elements using the self-encoding agent; retrieve the key field data subset corresponding to the extracted buyer's name from the financial and tax data knowledge database; compare the key field data subset with the buyer's name, and determine the comparison result as the result of the posterior detection of the key elements.

[0038] Optionally, in one embodiment of this application, the device 20 further includes an iteration module, which is used to distinguish, correct, or confirm the element information before the audit and evaluation agent queries other information corresponding to the element information in the external Internet or a private network as verification data: The review and evaluation agent is configured to perceive the features of the currently identified internal invoice as its state space. Combining the prediction results of the internal invoice from a pre-set deep learning model, and based on the reward function formula of reinforcement learning, the expected rewards for the generated internal invoice's posterior detection results ("qualified", "non-compliant", "marked for manual review") are calculated. The reward function formula is as follows: R =α Accuracy - β Cost of misjudgment Where R is the calculated expected return, and α and β are weighting coefficients set based on the business scenario, used to quantify the returns of different decision actions; The review and evaluation agent selects the action with the highest expected benefit to output, and the action space includes "determine compliance", "determine non-compliance" and "mark for manual review"; If the selected action output is "marked for manual review", the audit and evaluation agent uses its interactive capabilities to exchange and collaborate with the external Internet or dedicated network to query and obtain other information within the defined economic field, iteratively correct the key elements in the internal identification document, and regenerate a new internal identification document.

[0039] Furthermore, the auditing and evaluation agent is also used to introduce a query cost factor γ into the reward function. This cost factor γ is used to quantify the time or API call cost required to perform external information exchange, i.e., the reward function is calculated using the following formula: R = (α) Accuracy) - (β) (Cost of misjudgment) - (γ) (Query cost) This embodiment of the application further quantifies the benefits or penalties gained by the review and evaluation agent after performing an action using the above-described method. This enables the review and evaluation agent to "autonomously decide the strategy for reviewing each key field." For example, it learns that in certain situations (such as severe contamination), the total benefit (considering query costs) of directly marking it as "awaiting manual review" may be higher than attempting a high-cost, low-success-rate "external fuzzy search." This makes the system's decision-making more intelligent and economical. That is, the agent not only learns how to judge correctness but also how to judge correctness most economically, optimizing the system's resource consumption (time and monetary costs) while ensuring recognition accuracy.

[0040] In the practical application environment of this application embodiment, the information extraction agent, the self-encoding agent, and the review and evaluation agent are all agents constructed based on reinforcement learning, and all use "maximizing recognition accuracy" as the objective function, combined with a deep learning model (such as a CNN-LSTM hybrid model) to analyze the prediction results of historical data. They can also calculate the expected return and select the next step based on the "reward function formula R = α × accuracy - β × misjudgment cost". This ensures the efficiency and accuracy of the aforementioned steps.

[0041] Optionally, in a preferred implementation of this application embodiment, the extraction module 201 is further configured to control the element extraction agent to employ multi-layer feature extraction and fusion processing when processing the document image, and to adjust feature maps of different sizes to a fixed dimension through adaptive average pooling, while combining a dynamic attention mechanism to extract key elements. This application embodiment enhances the model's (element extraction agent's) understanding and information extraction capabilities for tilted, deformed text and complex table structures in this way. Traditional element extraction relies on standard deep convolutional networks, which have poor adaptability to complex image structures (such as tilted text and complex tables) and documents of different sizes. This application embodiment here introduces advanced image feature processing technology, and adaptive average pooling solves the problem of inconsistent input feature dimensions for documents of different sizes, ensuring the uniformity of subsequent processing. Simultaneously, the dynamic attention mechanism guides the model to focus on key text and structures in the image, improving the ability to recognize low-quality or deformed text. The implementation of this step effectively reduces initial errors caused by insufficient image processing, alleviates the verification burden of subsequent autoencoder agents and review and evaluation agents, optimizes the overall system resource consumption, and significantly enhances the adaptability to complex document scenarios.

[0042] Optionally, in a preferred implementation of this application embodiment, the review module 204 is further configured to: input the extracted key fields of "product name" or "seller / buyer name" into a pre-trained semantic classification model (e.g., a text classifier based on the BERT model) using the review and evaluation agent, wherein the semantic classification model is trained based on massive industry classification standard texts (such as national economic industry classification codes) and enterprise business scope data; use the semantic classification model to determine the classification result (e.g., an industry code or label) of the economic field to which the invoice belongs, and the review and evaluation agent uses the classification result as the economic field to which it belongs. This embodiment designs a process based on a deep learning NLP model to perform a quantifiable classification task to determine the economic field to which the invoice belongs, ensuring the objectivity and robustness of the field judgment, avoiding errors and omissions based on simple rule judgments, providing a more accurate range for subsequent intra-field queries, thereby improving the quality of external verification data and the query success rate.

[0043] Optionally, in one embodiment of this application, the device 20 further includes an optimization module (not shown in the figures), which is used to: optimize and adjust the action value function parameters of the review and evaluation agent based on the decision error determined by the manual review results through a reinforcement learning algorithm.

[0044] Optionally, in one embodiment of this application, the device 20 further includes a training module (not shown in the figures), which is used to: evaluate the deep learning model using F1-score as a system performance index based on the decision error determined by manual review; and retrain the deep learning model with full data when the determined F1-score is lower than a preset threshold.

[0045] Example 3 This application also provides a storage medium storing a computer program that, when executed by a processor, implements any of the methods for identifying tax and financial documents as described in the foregoing embodiment one of this application. Example 4 This application also provides an electronic device, such as... Figure 3 As shown, Figure 3 This application provides a schematic diagram of the structure of an electronic device 30, which includes: One or more processors 301, communication interface 302, memory 303 and communication bus 304, the processors 301, memory 303 and communication interface 302 communicate with each other through communication bus 304; Memory 303 is used to store one or more programs; When the one or more programs are executed by the one or more processors 301, the one or more processors 301 implement any of the methods for identifying financial and tax invoices as described in Embodiment 1 of this application.

[0046] This application has now described specific embodiments of the subject matter. In some cases, the actions described in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing can be advantageous.

[0047] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system layer onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0048] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0049] The system layers, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0050] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0051] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0052] Those skilled in the art will understand that embodiments of this application can be provided as methods, system-level, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0053] This application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific transactions or implement specific abstract data types. This application can also be practiced in distributed computing environments where transactions are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0054] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system-level embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0055] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for identifying fiscal documents, characterized in that, The method comprises the following steps: extracting key elements of the to-be-identified financial and tax documents by a preset element extraction intelligent agent, wherein the key elements at least include seller data, buyer data, commodity name, quantity and unit price information; generating a structured internal identification document based on a preset internal document format by a self-encoding intelligent agent after receiving the key elements, and comparing at least one type of element information in the key elements with a preset financial and tax data knowledge database to perform a posteriori detection on the key elements; sending a failure notification containing element information causing the posteriori detection failure to an audit and evaluation intelligent agent when the posteriori detection fails; determining an economic field to which the to-be-identified financial and tax documents belong by using the audit and evaluation intelligent agent according to the element information in the failure notification; querying other information corresponding to the element information as verification data from the external Internet or a special network based on the determined economic field information by using the audit and evaluation intelligent agent to distinguish, correct or confirm the element information for which the posteriori detection fails in the internal identification document, and obtaining final identification information; wherein the other information at least includes main information of a seller enterprise or a buyer enterprise, or price information corresponding to the commodity name.

2. The method for identifying a tax document according to claim 1, wherein, The method further comprises the following steps: querying main information of enterprises corresponding to the seller data and the buyer data and upstream and downstream enterprise information, and / or querying market price information corresponding to the commodity name by using the audit and evaluation intelligent agent in the determined economic field; distinguishing the seller enterprise name and the buyer enterprise name of the to-be-identified financial and tax documents according to the main information, the upstream and downstream information and the commodity name, and distinguishing the quantity and the unit price information according to the price information.

3. The method of claim 2, wherein, The method further comprises the following steps:

4. The method of claim 1, wherein, When the field information of the designed enterprise name or commodity name in the extracted key elements is incomplete, performing a fuzzy search on the economic field information to locate the enterprise name or commodity name. The method further comprises the following steps: submitting the element information after being corrected or confirmed by the audit and evaluation intelligent agent to a human for review and confirmation to obtain a review result; 5. The method of claim 1, wherein, using the review result as an intelligent agent reward value to feed back to the self-encoding intelligent agent for classified storage or updating the financial and tax data knowledge database according to the review result, and for guiding the process of extracting key elements by the element extraction intelligent agent. The method further comprises the following steps: extracting the buyer name in the key elements by using the self-encoding intelligent agent; calling a key field data sub-set corresponding to the extracted buyer name from the financial and tax data knowledge database; comparing the key field data sub-set with the buyer name, and determining the comparison result as the result of the posteriori detection on the key elements.

6. The method of claim 1, wherein, Before the method of using the audit evaluation agent to query other information corresponding to the element information in the external Internet or private network as verification data, distinguishing, correcting or confirming the element information, the method further comprises: The audit evaluation agent sets the current internal identification invoice single feature as the state space, combines the prediction result of the preset deep learning model, and calculates the posterior detection result of the generated internal identification invoice single as the expected return of the three types of actions of "determination of compliance", "determination of non-compliance", and "labeling for manual review" according to the reward function formula of reinforcement learning. The reward function formula is: R = α Accuracy - β Misclassification cost Wherein, R is the calculated expected return, and α and β are weight coefficients based on the business scenario, which are used to quantify the return of different decision actions; The audit evaluation agent selects the action output with the highest expected return, and the action space includes "determination of compliance", "determination of non-compliance", and "labeling for manual review"; If the selected action output is "labeling for manual review", the audit evaluation agent uses the interaction capability of information exchange and cooperation with the external Internet or private network to query and obtain the other information in the determined economic field, iteratively corrects the key elements in the internal identification invoice single, and regenerates a new internal identification invoice single.

7. The method of claim 6, wherein, The method further comprises: Based on the decision error determined by manual review results, the action value function parameters of the audit evaluation agent are optimized and adjusted by a reinforcement learning algorithm.

8. The method of claim 6, wherein, The method further comprises: Based on the decision error determined by manual review results, the deep learning model is evaluated by using F1-score as a system performance indicator; When the determined F1-score indicator is lower than the preset threshold, the deep learning model is retrained with full data.

9. A device for identifying tax and financial documents, characterized in that, Comprise: An extraction module for extracting key elements of the to-be-identified financial and tax documents by a preset element extraction agent, the key elements including at least seller data, buyer data, commodity name, quantity, and unit price information; A self-encoding module for receiving the key elements by a self-encoding agent, generating a structured internal identification invoice single based on a preset internal invoice format, and comparing at least one type of element information in the key elements with a preset financial and tax data knowledge database to perform posterior detection on the key elements; A sending evaluation module for sending a failure notification containing element information causing the posterior detection failure to an audit evaluation agent when the posterior detection fails; An audit module for determining the economic field to which the to-be-identified financial and tax documents belong by using the audit evaluation agent according to the element information in the failure notification; The confirmation module is configured to determine economic field information, query other information corresponding to the element information as verification data in an external Internet or a special network by using the audit evaluation agent, distinguish, correct or confirm the element information of the posterior detection failure in the internal identification ticket, and obtain final identification information, wherein the other information at least includes main information of a selling enterprise or a buying enterprise, or price information corresponding to the commodity name.

10. A computer storage medium, characterized in that The computer storage medium stores computer executable instructions, and the computer executable instructions are executed to perform the method for identifying the financial and tax bill according to any one of claims 1-8.