A multi-agent cooperative financial document auditing method and system based on adaptive routing of expense types

CN122841102APending Publication Date: 2026-09-29SHENZHEN RUNDIAN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611042869.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-29

AI Technical Summary

Benefits of technology

本发明通过多模态预处理+视觉语言深层理解、专用智能体路由、跨模态交叉校验+大模型意见生成、RPA闭环的协同架构,实现了对财务单据审核流程的系统性重构。通过构建统一的多模态数据处理链路,将原本分散的影像数据、结构化报账数据及附件信息进行标准化处理,并结合视觉语言模型对影像版式语义的深层理解能力,使审核过程不仅能够获取文本信息,还能够理解单据结构及业务语义,从而有效解决单一模型难以覆盖复杂差异化规则的问题。同时,通过将审核流程拆分为多个协同环节,提升整体处理的稳定性与可扩展性,实现复杂业务场景下的高效自动化处理。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841102A_ABST
    Figure CN122841102A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-agent cooperative financial document auditing method and system based on expense type adaptive routing, it is related to information technology field, comprising the following steps: preprocessing in expense reimbursement management system in expense reimbursement document attachment compressed file and structured accounting data, execute decompression, type classification, format conversion, encoding restoration and table structure processing, obtain image data set and structured accounting data.The application realizes automatic auditing of financial document by multimodal preprocessing, visual language understanding, agent routing and cross-modal verification, solves the problem of complex rules and insufficient semantic understanding.Through task decomposition, the model load is reduced and the accuracy is improved, and the execution stability is improved by combining prearranged process.At the same time, unmanned batch processing is realized, the auditing efficiency is improved and the labor cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, specifically to a multi-agent collaborative financial document review method and system based on cost type adaptive routing. Background Technology

[0002] With the continuous improvement of enterprise financial digitalization, technologies such as Enterprise Resource Planning (ERP), Expense Management System (EMS), and Robotic Process Automation (RPA) are widely used in financial shared service centers to realize business processes such as expense application, approval workflow, and payment processing. Meanwhile, Optical Character Recognition (OCR) technology can extract text information such as invoice number, amount, and date from invoice images. The development of Large Language Model (LLM) and Visual Language Model (VL) has also enabled the system to have a certain semantic understanding capability, allowing for joint analysis of image content and text information. In the document review business of financial shared service centers, it is necessary to review the completeness of attachments, invoice compliance, amount consistency, and business logic of a large number of expense documents. Furthermore, the review rules for different expense types (such as entertainment expenses, travel expenses, and e-commerce invoices) vary significantly, exhibiting characteristics of multimodal data processing and complex rule judgment.

[0003] However, existing technologies still have significant shortcomings: traditional solutions based on RPA+OCR+rule engines can achieve basic automated processing, but OCR can only extract text information and cannot understand the semantics of image layouts, making it difficult to handle non-standard attachments; a single general-purpose large model or a single agent architecture is prone to accumulated intent recognition errors and model illusions when faced with multiple fee types and complex audit rules, leading to unstable audit results; at the same time, existing solutions lack cross-modal consistency verification mechanisms between image data, reimbursement system data, and bank statement data, making it difficult to detect cross-system data inconsistencies in a timely manner. In addition, most solutions rely on user input of natural language and multiple rounds of interaction, which cannot achieve unmanned batch audit closed loops, and rule engines have high maintenance costs and poor scalability when rules change, making it difficult to meet the actual needs of financial shared service centers for high efficiency, high accuracy, and automated processing.

[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-agent collaborative financial document review method and system based on cost type adaptive routing to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a multi-agent collaborative financial document review method based on cost type adaptive routing, comprising the following steps: The compressed files of reimbursement document attachments and structured reimbursement data in the expense reimbursement management system are preprocessed by performing decompression, type classification, format conversion, encoding restoration and table structuring to obtain image data set and structured reimbursement data; The image data set is input into the visual language model for recognition, and the document type, key fields and layout semantic information are obtained, and the structured image recognition results are output. The structured expense report data is parsed to extract basic information, amount information, business information, and attachment index information of the expense report, forming expense report data objects; The document type information in the image recognition results is matched with the expense type information in the expense data object to determine the expense category, and the review task is assigned to the corresponding dedicated review intelligent agent. Within the corresponding dedicated auditing intelligence, differentiated auditing rules are invoked to perform audits on the image recognition results and the reimbursement data objects regarding the completeness of attachments, consistency of amounts, business logic, and compliance, and to output audit results in various dimensions. Based on the review results, cross-modal consistency verification is performed on the image recognition results, reimbursement data objects, and bank receipt data to complete multi-dimensional comparison of amount, account, date, number of people, and name, and obtain the consistency verification results; The review results and consistency verification results are input into the text big language model for analysis, and natural language review comments containing review status and anomaly descriptions are generated. The review comments are structured and packaged and sent back to the expense reimbursement management system for display, thus completing the multi-agent collaborative financial document review based on expense type adaptive routing.

[0007] Before multimodal attachments enter the intelligent processing flow, heterogeneous data needs to be uniformly organized and format-converted to form a standardized input structure. The steps are as follows: Receive compressed attachments and structured expense reimbursement data output from the expense reimbursement management system, and record the association between the attachment source identifier and the expense reimbursement form; Perform decompression on the compressed file, read the internal file directory structure of the compressed package and generate a set of raw files containing file paths and type identifiers; The original file set was MIME type identified and classified into PDF files, XML files, image files, and reimbursement information files according to file extensions and content characteristics, and a classification index was established. The system performs image conversion processing on PDF files, Base64 encoding and parsing on XML files to restore image data, and table parsing on expense report files to convert them into JSON structures, forming a unified set of image data and structured data objects.

[0008] To address the complex layout and semantic information contained in financial images, structured parsing and content recognition are required. The steps are as follows: Receive the pre-processed image data set and establish an association mapping between image numbers and invoices according to file order; Image data is input into a visual language model for multimodal feature extraction, including text region detection, layout structure recognition, and semantic feature encoding. Perform document type determination and field extraction processing, identify the title area, field area and key content area according to the layout structure, and extract the corresponding business field information; The recognition results are structured and organized to form an image recognition result set that includes document type labels, field key-value pairs, confidence information, and layout semantic analysis conclusions.

[0009] Based on the characteristics of structured data in expense reimbursement processes, the following steps are taken to extract and organize the fields of the reimbursement information: Obtain the structured expense reimbursement data output by the expense reimbursement management system and parse the data format to be XML or JSON; Perform field mapping processing on the reimbursement data to extract basic information fields such as reimbursement number, handler, expense type, and application date; Parse the amount-related fields, including the total expense report amount, detailed expense report amount, tax amount and tax rate, and establish numerical relationships between the fields; The business information and attachment index information are integrated to extract payment methods, account information, number of people, and attachment list to form a complete reimbursement data object.

[0010] During the formation of the reimbursement data object, the field structure is further standardized. A unified field code is established for the reimbursement bill number, expense type and application date. The amount field is processed with uniform precision and hierarchical association. A mapping relationship is established between the payment account and the attachment index, forming a structured reimbursement data set with field consistency and association.

[0011] In scenarios where different fee categories correspond to different review logics, task routing and path selection processing need to be completed. The steps are as follows: Receive image recognition results and reimbursement data objects, and establish the field correspondence between the two types of data; Perform a combined matching process on the expense type field and the document type label to construct a set of expense category determination conditions; Based on the matching results, determine the target cost category and generate the corresponding route identification information; The image recognition results, reimbursement data objects, and routing identifiers are encapsulated to form an audit task package and assigned to the corresponding audit agent.

[0012] To address the differences in business rules for various types of expenses, multi-dimensional review and semantic analysis are conducted, following these steps: Receive the audit task package and load the set of differentiated rules for the corresponding fee category; The task data is processed by rule matching, and each item is checked for attachment verification, amount comparison, date verification and field consistency detection. For content involving semantic judgment, a large text language model is invoked for semantic parsing, including contract clause recognition and complex business relationship judgment; The execution results of each rule are summarized to form a multi-dimensional audit result set that includes the rule name, execution status, and exception information.

[0013] To address the inconsistency caused by differences in data sources, a unified comparison and discrepancy identification process is performed, following these steps: Receive image recognition data, expense report data, and bank statement data, and extract the amount, account number, date, and name fields; Perform standardization processing on fields from different sources, including string normalization, date format conversion, and numerical uniformity processing; Conduct multi-dimensional comparison operations, including numerical comparison of amounts, character matching of accounts, and fuzzy matching of names; The comparison results are marked with differences, generating verification result data that includes the difference fields, difference content, and location identifiers.

[0014] During the process of outputting audit results and interacting with the business system, it is necessary to complete result integration and feedback processing. The steps are as follows: Receive the audit result set and consistency verification data, and establish an index relationship for the result data; The review status is categorized and processed, and the data is encapsulated according to a preset field structure to generate a JSON format report; The structured report content is parsed to generate natural language descriptions and associated with corresponding anomalies; The results are returned via API calls, and the review status and detailed information are mapped to the system display interface.

[0015] A multi-agent collaborative financial document review system based on cost type adaptive routing includes an RPA input layer, a workflow engine layer, a dedicated review agent layer, a processing layer, a model layer, and a plugin layer. The RPA input layer is used to perform interactive operations of the financial system through a preset set of skills, which includes acquisition skills, triggering skills, and display skills. Acquisition skills are used to automatically log in to the expense reimbursement management system and obtain expense reimbursement information and image compressed data. Triggering skills are used to call the workflow engine interface to upload image compressed data and structured expense reimbursement data to trigger the review process. Display skills are used to receive the review results and fill in the review information and display the exception items on the expense reimbursement management system page. The workflow engine layer is used to uniformly schedule and orchestrate the review tasks. It is pre-configured with processing pipelines including file decompression nodes, type classification nodes, format conversion nodes, image recognition nodes, fee type routing nodes, rule review nodes, cross-modal verification nodes, result integration nodes, and output nodes to realize the sequential execution and node scheduling of the review process. A dedicated audit agent layer is used to divide multiple independent audit agents according to the fee type. Each audit agent corresponds to a different fee category and has a built-in differentiated audit rule library and audit prompt template to handle audit tasks of different fee types in a targeted manner. The processing layer is used to perform specific data processing operations, including file decompression and classification, data format conversion, image content recognition, rule review, cross-modal consistency verification, and review result integration. The model layer provides intelligent analysis capabilities, including a visual language model and a text-based large language model. The visual language model is used to identify document types and extract key fields from image data, while the text-based large language model is used to perform rule-based reasoning, cross-modal validation logic processing, and generate audit comments. The plugin layer provides basic processing tool support, including PDF to PNG conversion plugin, Base64 encoding / decoding plugin, compressed file decompression plugin, Excel to JSON conversion plugin, and large model interface call plugin, to support data preprocessing and model calling processes. The layers interact with each other through standard interfaces or messaging mechanisms to achieve a complete closed-loop processing flow from data collection, task scheduling, intelligent review to result output.

[0016] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This invention achieves a systematic reconstruction of the financial document review process through a collaborative architecture that integrates multimodal preprocessing, deep visual language understanding, dedicated intelligent agent routing, cross-modal cross-validation, large-scale model opinion generation, and RPA closed-loop processing. By constructing a unified multimodal data processing link, the previously scattered image data, structured expense report data, and attachment information are standardized. Combined with the deep understanding of image layout semantics by a visual language model, the review process can not only acquire textual information but also understand document structure and business semantics, effectively solving the problem that a single model cannot cover complex and differentiated rules. Simultaneously, by breaking down the review process into multiple collaborative stages, the overall processing stability and scalability are improved, enabling efficient automated processing in complex business scenarios.

[0017] Compared to the approach of dynamically generating tool call sequences using a single large model, this invention introduces an adaptive routing mechanism based on cost type. This mechanism splits complex review tasks into multiple dedicated agents based on cost categories, allowing each agent to focus solely on the review rules of its corresponding domain. This significantly reduces the model's cognitive load and inference complexity, minimizing the risk of illusions caused by multi-task mixing. Furthermore, it employs a pre-arranged workflow pipeline instead of dynamic tool orchestration, designing a fixed path for the review process. This ensures a clear and controllable execution order for each processing node, effectively avoiding issues such as chaotic tool call order or missing parameters. Consequently, it maintains stable and high-precision review results across various differentiated business scenarios, including entertainment expenses and travel expenses.

[0018] Compared to traditional OCR-rule engine-based solutions, this invention introduces a visual language model to achieve semantic understanding of image layouts. It not only extracts text but also identifies document types, determines the compliance of the number of signatories, detects consecutive invoice numbers, and parses non-standard attachments, overcoming the limitations of traditional OCR technology that only focuses on character recognition. Furthermore, by constructing a cross-modal consistency verification mechanism, it performs three-way cross-validation of image data, expense report system data, and bank statement data, automatically identifying issues easily overlooked in manual review, such as discrepancies in amounts and account inconsistencies. In addition, through RPA-automated workflow engine triggering, it achieves fully automated processing from data collection to result output, eliminating the need for manual input of natural language or interaction. This meets the unmanned batch initial review needs of financial shared service centers and significantly improves review efficiency. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0020] Figure 1 This is a flowchart of the method of the present invention.

[0021] Figure 2 This is a timing diagram of the interaction between RPA and the intelligent agent in this invention.

[0022] Figure 3 This is the cost type adaptive routing decision graph of the present invention.

[0023] Figure 4 This is a schematic diagram of the cross-modal consistency verification principle of the present invention.

[0024] Figure 5 This is a block diagram of the overall architecture of the present invention. Detailed Implementation

[0025] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0026] This invention provides, for example Figures 1-4 The multi-agent collaborative financial document review method based on cost type adaptive routing is shown below, with the following specific steps: In the automated document review scenario of a financial shared service center, the raw data from the expense reimbursement management system first needs to be standardized. This data is typically obtained by RPA automatically logging into the system and downloading it in batches, specifically including ZIP-packaged attachments and the corresponding structured expense reimbursement data. Due to the complex sources and heterogeneous formats of the attachments, the compressed packages often contain PDF files, XML files, PNG or JPG image files, and Excel or other structured expense reimbursement information files. Therefore, before proceeding to subsequent intelligent recognition and review, it is necessary to build a complete and standardized multimodal attachment preprocessing link to achieve unified expression and high-quality input of different types of data.

[0027] In the specific processing, the ZIP archive is first decompressed. A decompression plugin is used to automatically extract all uploaded ZIP files, preserving the original file hierarchy and index relationships. This ensures accurate tracing of attachment sources and their corresponding relationships during subsequent processing. After decompression, the file type identification stage begins. Based on the MIME type of the files, all attachments are automatically categorized into four main types: PDF files, XML files, image files (PNG / JPG), and expense report files (Excel or other structured data). This standardized classification method allows for matching appropriate processing strategies to different file types, thereby improving overall processing efficiency and accuracy.

[0028] For PDF files, since they are document formats, they cannot be directly input into the visual language model for recognition; therefore, further format conversion is required. A PDF to PNG plugin is used to convert each page of the PDF into a PNG image format. Simultaneously, the PDF scaling factor is set to 2 to ensure that the resolution and clarity of the converted images meet the requirements of the subsequent visual language model recognition, avoiding a decrease in recognition accuracy due to image blur. Regarding XML file processing, considering that some electronic invoices or system-exported data may embed image data in Base64 encoded form, a Base64 encoding / decoding plugin is needed to parse and decode the encoded fields in the XML, restoring them to actual image files, thus incorporating them into a unified image data set for subsequent processing.

[0029] For expense reimbursement documents, especially structured data in Excel format, a dedicated Excel-to-JSON plugin is needed to convert the tabular data into a standardized JSON structure. This process not only preserves the original data fields but also restructures them to ensure efficient integration and fusion with image recognition results in subsequent processes. Simultaneously, the conversion process must ensure data integrity and that field mappings remain intact for accurate verification of critical fields such as amount, date, and account number.

[0030] After completing the above processing operations, all processing results are integrated into a categorized file set. This set mainly includes three parts: first, the set of images to be identified, containing PNG or JPG format image data composed of original image files, PDF conversion results, and Base64 decoded image data; second, structured expense data, including JSON or XML data converted from Excel or original structured files; and third, original attachment index information, used to record the source location, type identifier, and correspondence of each file in the compressed package. In this way, the transformation from raw heterogeneous attachments to a standardized multimodal data set is realized, providing a unified and reliable data foundation for subsequent intelligent image content recognition, expense type determination, and rule review, while ensuring the traceability of the data link and the consistency of the processing process.

[0031] After preprocessing the multimodal attachments, a unified set of images to be identified is formed. This set consists of image data in PNG or JPG format, including the original image files, images converted from PDFs, and image content restored through Base64 decoding. This image data serves as input for subsequent intelligent analysis, entering the image content intelligent recognition stage. The core of this stage lies in using a visual language model to perform deep semantic analysis on the images, enabling them not only to be read for their textual information but also to understand their structure, layout, and business meaning, thereby providing a high-quality structured data foundation for subsequent review.

[0032] In the specific processing, the preprocessed image is first input into the visual language model Qwen3-VL-32B. This model has the ability to simultaneously understand image content and text semantics. Based on a multimodal pre-training framework, it jointly models the layout features, visual structural information, and text semantics in the image, thereby enabling deep semantic analysis of complex financial images. Unlike traditional image recognition models that only focus on pixel features or text recognition, this model establishes a mapping relationship between image regions and semantic concepts through a cross-modal alignment mechanism. This allows the model to not only extract explicit information during the recognition process but also infer implicit business semantics.

[0033] After the image is input into the model, the first step is to perform a document type recognition task. Based on the image's layout features, text content, and overall structure, the model automatically determines the document type to which the image belongs. Specifically, by analyzing information such as the title area, field arrangement, keyword distribution, and format characteristics, it can distinguish different categories of financial documents. Recognizable document types include, but are not limited to, VAT invoices, beverage requisition forms, reception application forms, bank receipts, training notices, hotel bills, contracts, shared waybills, overtime applications, settlement statements, and receiving slips. This method enables automatic classification of complex non-standard attachments, overcoming the technical limitations of traditional OCR in distinguishing document layout types, and laying the foundation for subsequent differentiated processing for different document types.

[0034] After document type identification, key field extraction is further performed. For each identified document type, the model extracts corresponding key business information according to predefined field templates. For invoices, fields such as invoice number, invoice date, amount (including or excluding tax), tax rate, tax amount, buyer's name, seller's name, and specifications can be extracted. For beverage requisition forms, information such as requisition date, requisition amount, signatory's name, and number of signatories can be extracted. For reception application forms, reception date, number of hosts, number of guests, and reception location can be extracted. For bank receipts, payment account, receiving account, transaction amount, and transaction date can be extracted. For hotel bills, guest's name, check-in date, hotel name, and hotel stamp information can be extracted. This type-customized field extraction strategy ensures that key information for different document types is accurately identified and structured for output, while avoiding interference from irrelevant fields and improving data quality.

[0035] Beyond explicit field extraction, image layout semantic understanding is also required. This process focuses not only on textual content but also on analyzing the image's structural layout and regional semantics. For example, by identifying the signature area and counting the number of signatures, it can be determined whether the number of people signing the beverage requisition form is no less than two; by analyzing the continuity of the invoice number area, it can be determined whether the invoices have consecutive numbering features; by detecting the integrity and clarity of image boundaries, it can be determined whether the attached image has missing corners or is blurry; by identifying the invoice markings, it can be determined whether machine-printed invoices have an official seal; and by parsing the invoice details area, it can be determined whether the beverage invoice indicates the specifications and model. These judgments all belong to semantic-level analysis, relying not only on text recognition results but also on a comprehensive understanding of image structure and business rules, thereby achieving deep analysis of unstructured images.

[0036] In this process, the fundamental difference from traditional OCR technology becomes clear. Traditional OCR technologies, such as Tesseract or common commercial OCR engines, are typically based on character segmentation and template matching methods. They can only output the coordinates and text content of characters in an image, and their processing results remain at the level of character recognition. They cannot answer questions such as where the signature area is located, how many people signed, or whether there is a consecutive number relationship between multiple invoices. In contrast, the visual language model Qwen3-VL-32B, through end-to-end multimodal pre-training, establishes a mapping relationship between image layout features and semantic concepts. It can directly output semantic-level conclusions such as document type labels, the number of signatories, and consecutive number anomaly detection. This achieves a leap from text extraction to semantic understanding, breaking through the bottleneck of traditional OCR technology in processing complex financial images.

[0037] After completing the above identification and analysis, the final output is an image recognition result in structured JSON format. This result includes document type information, key-value pairs for each key field, confidence scores for the corresponding fields, and semantic analysis conclusions for the layout. This structured representation not only facilitates automated processing in subsequent workflows but also supports cross-module data flow and the execution of multi-dimensional review logic. Furthermore, the confidence score can be used to assess the reliability of the recognition results, providing a basis for subsequent manual review when necessary. The entire intelligent image content recognition process achieves the transformation from raw image data to high-quality structured semantic data, providing crucial data support for fee type judgment, multi-agent scheduling, and rule review.

[0038] After completing the intelligent recognition of image content, the structured reimbursement data from the expense reimbursement management system needs further analysis to form a data foundation corresponding to the image recognition results. This structured reimbursement data is automatically crawled from the EMS system by RPA using the CollectionSkill. The data format is usually XML or JSON, containing all business information entered into the system for the reimbursement documents. Because this type of data is standardized structured data, it has the characteristics of clearly defined fields and standardized format compared to image data. Therefore, it needs to be transformed into a unified data object through a systematic parsing process to provide support for subsequent expense type determination, rule review, and cross-modal consistency verification.

[0039] During the parsing process, the basic information of the expense report is extracted first. This information is the core identifying data of the entire expense report, including the expense report number, the person in charge, the expense type, and the application date. The expense report number serves as a unique identifier, used throughout the entire review process to establish data association between each processing stage; the person in charge information identifies the entity initiating the business and can be used as an auxiliary judgment basis in some rule validations; the expense type is an important input parameter for subsequent expense classification and agent routing; and the application date is used to compare with the business occurrence date extracted from the image to determine the reasonableness of the business time. By parsing this basic information, the basic framework structure of the expense report can be established.

[0040] After extracting the basic information, the reimbursement amount information is further analyzed. This part mainly includes key fields such as the total reimbursement amount, various detailed amounts, tax amount, and tax rate. The total reimbursement amount reflects the financial scale of the entire transaction and is an important basis for subsequent consistency verification with invoice and bank statement amounts. Various detailed amounts are used to refine the expense composition and can be used in multi-dimensional audits to identify whether there are any split reimbursements or abnormal amounts. Tax amount and tax rate information are used to determine the compliance of invoices and the correctness of tax processing. Through structured analysis of the amount-related fields, an accurate data foundation can be provided for subsequent consistency verification and compliance audits.

[0041] After parsing the monetary information, the business information is further extracted and organized. This information reflects the specific business scenario and execution status, including the date of the reception, the number of guests and hosts, the payment method (e.g., offline payment, wire transfer, or bank transfer), the payment account, and the paperless option (yes or no). The reception date is used to compare with the business time identified in the image to verify the authenticity of the business occurrence time; the number of guests and hosts is important in the review of entertainment expenses and can be used to verify the consistency with the number of people in the reception application form; the payment method determines whether bank receipts and fund flows need to be verified later; the payment account, as key financial information, is used to compare with the account in the bank receipt during cross-modal verification; the paperless option is used to determine whether the attachment format conforms to the system configuration, such as whether it is allowed to submit only electronic invoices without paper attachments. By parsing the business information, the actual execution of the reimbursement business can be fully restored, providing the necessary context for subsequent rule review.

[0042] In addition, the attachment index information needs to be parsed. This information, recorded by the expense reimbursement management system, includes attachment name, attachment type, and attachment quantity. The attachment index information describes the overall situation of the attachments uploaded to the current reimbursement document and is an important basis for judging the completeness of attachments. For example, the number of attachments can be compared with the minimum number of attachments required by the system to identify whether any attachments are missing; the attachment type information can determine whether the necessary document type, such as invoice, contract, or settlement statement, has been uploaded; and the attachment name can assist in verifying the image recognition results and confirming the consistency between the recognition results and the system records. This information plays a crucial bridging role between image data and system data during subsequent review processes.

[0043] After parsing the aforementioned information, all extracted fields are integrated to form a structured expense reimbursement data object. This data object contains all field information extracted from the EMS system, covering basic information, amount information, business information, and attachment index information, while maintaining the logical relationships between fields. This method achieves the transformation from raw XML or JSON data to a standardized structured object, enabling efficient matching and fusion with image recognition results in subsequent processes. Simultaneously, this structured expense reimbursement data object provides crucial input for adaptive routing of expense types, complete data support for differentiated rule review, and a benchmark for cross-modal consistency verification, thus playing a pivotal role in the entire automated expense review process.

[0044] After completing intelligent image content recognition and reimbursement information parsing, image recognition results and structured reimbursement data were obtained. These two types of data describe different aspects of the same reimbursement business from the image semantic layer and the system input layer, respectively. The image recognition results include document type, key fields, and layout semantic analysis conclusions, while the structured reimbursement data includes expense type fields and complete business information. Based on this, an expense type adaptive routing mechanism is needed to accurately assign the current reimbursement document to the corresponding review path, thereby achieving refined processing for different expense categories. This process uses the document type from the image recognition results and the expense type field from the reimbursement data as input, and completes classification and scheduling through a unified decision logic.

[0045] In the specific processing, the routing decision engine first performs joint analysis on the two types of input information. The routing decision engine primarily uses the expense type field in the expense report data, while simultaneously combining it with document type information from the image recognition results to determine the expense category to which the expense report belongs. Since the same expense type may correspond to multiple image document formats in actual business, relying solely on a single field is insufficient for accurate classification. Therefore, a two-dimensional matching approach can significantly improve classification accuracy. For example, when the expense type in the expense report data is expense reimbursement, and the image recognition results include an image of a beverage requisition form, it can be further refined to "private entertainment expenses"; when the expense type is invoice registration, and an e-commerce invoice is identified in the image, it can be determined as an e-commerce invoice registration. This combined judgment method achieves a mapping from coarse-grained expense types to fine-grained business categories, avoiding misjudgments caused by simple classification.

[0046] After determining the expense category, the task routing process is further executed. Based on the determined expense category, the current audit task is adaptively assigned to the corresponding dedicated audit agent. Different expense categories correspond to different audit logics and rule systems; therefore, a clear routing mapping relationship needs to be established to ensure that tasks enter the correct processing path. Specific mapping relationships include: private or official entertainment expenses correspond to an entertainment expense audit agent; bank deductions correspond to a bank deduction audit agent; e-commerce invoice registration corresponds to an e-commerce invoice registration audit agent; general material invoices correspond to general material invoice audit agents; travel expenses correspond to travel expense audit agents; fuel settlement invoices (including freight, miscellaneous fees, and coal prices) correspond to fuel settlement invoice audit agents; account transfer transactions correspond to account transfer audit agents; reversal adjustment requests correspond to reversal adjustment request audit agents; performance-related compensation and business expenditure communication expenses correspond to communication expense audit agents; and transportation expenses correspond to transportation expense audit agents. Through this one-to-one mapping relationship, precise matching of expense categories and audit agents can be achieved, thereby ensuring that each type of business enters the processing unit with the corresponding professional rules.

[0047] During the routing process, the routing results also need to be identified, and the relevant data needs to be packaged uniformly. Specifically, image recognition results, structured accounting data, and routing identification information are integrated to form a complete audit task data package. This data package not only contains the raw data used for auditing but also includes explicit processing path information, enabling subsequent processing to directly call the corresponding intelligent agent to execute the audit logic. Through a unified data encapsulation method, it is ensured that different intelligent agents have a consistent data structure when receiving tasks, and it also facilitates data tracking and result backtracking in subsequent processes.

[0048] Overall, the expense type adaptive routing mechanism not only implements task allocation but also structurally decomposes the review process at the system level. By introducing a routing decision engine, the complex and diverse reimbursement processes are divided into domains based on expense categories, allowing each review agent to handle only the rules within its corresponding domain. This reduces the complexity of individual processing units and improves overall processing efficiency and accuracy. Simultaneously, the joint judgment mechanism based on image recognition results and reimbursement data enhances the robustness of the classification process, avoiding incorrect routing issues caused by inaccurate data sources.

[0049] Furthermore, this mechanism provides a prerequisite for subsequent multi-dimensional differentiated audits. Since the audit rules for different expense categories vary significantly, precise routing in the early stages ensures that subsequent rule execution occurs within the correct context, avoiding rule misuse or logical conflicts. Simultaneously, the unified task data packet structure provides a foundation for subsequent cross-modal consistency verification, enabling comparative analysis of image data, expense reports, and other data sources within the same framework.

[0050] Finally, through the above processing, the output is an audit task package routed to the target dedicated audit agent. This task package contains complete input data and routing information, which can be directly called by subsequent audit stages, thereby achieving a seamless connection between data parsing and intelligent auditing, and providing a stable and efficient scheduling foundation for the entire automated order review process.

[0051] After completing the adaptive routing for different expense types, an audit task package containing image recognition results, structured expense data, and routing identifiers is generated. This task package is then assigned to the corresponding dedicated audit agent for subsequent processing. Upon receiving the audit task package, each dedicated audit agent invokes a pre-built differentiated audit rule library and leverages the semantic reasoning capabilities of the Qwen2.5-72B text large language model to conduct multi-dimensional audit processing for different expense types. Because different expense categories exhibit significant differences in business rules, compliance requirements, and data structures, processing them separately through dedicated audit agents enables a highly targeted and accurate audit process. The following section uses a representative expense type as an example to explain in detail the execution method of multi-dimensional differentiated rule auditing.

[0052] In the scenario of auditing entertainment expenses, the audit process revolves around multiple dimensions, including beverage requisition forms, beverage invoices, reception application forms, and reimbursement assistant indicators. In the beverage requisition form audit dimension, the existence of attachments is first checked by image recognition to confirm the existence of beverage requisition forms, thus ensuring complete business documentation. Next, an amount consistency check is performed, comparing the beverage requisition amount extracted from the image with the amount entered in the reimbursement system to identify any discrepancies. Regarding signature compliance, the number of people in the signature area is counted through semantic analysis of the layout to determine if there are at least two signatures. Regarding unit price compliance, conversion is performed based on the specifications and models identified in the image recognition; for example, baijiu is converted to 500 ml and red wine to 750 ml, calculating the unit price and determining if it exceeds the 500 yuan limit. In the beverage invoice audit dimension, the focus is on the completeness of specifications and models. The invoice details area is identified to confirm whether it contains beverage specifications and models, and the unit price is further calculated based on the specifications and models to determine if it exceeds the limit. In the reception application form review dimension, the main checks are date consistency and number of attendees consistency. The date on the reception application form obtained through image recognition is compared with the reception date entered into the expense reimbursement system. Simultaneously, the number of guests and hosts obtained through image recognition is checked for consistency with the data entered into the expense reimbursement system. In the expense reimbursement assistant indicator review dimension, comprehensive checks are performed on invoice amount accuracy, invoice information completeness, reception date correctness (determining whether the invoice was issued after the business transaction), consecutive invoice detection, overlap between entertainment expenses and travel allowances, and the correctness of the entertainment expense invoice tax rate to ensure that invoices do not contain anomalies such as tax exemption or zero tax rate.

[0053] In bank deduction verification scenarios, the first step is to verify the existence of attachments, confirming the completeness of the voucher by checking for the presence of image attachments. If no attachments are detected, the system further checks whether the paperless option is enabled to confirm whether the transaction complies with paperless processing requirements. Regarding payment type verification, when the payment method is offline payment or wire transfer, the payment account and amount from the bank receipt in the image need to be extracted and compared with the corresponding fields in the reimbursement system. Simultaneously, the paperless option is checked to determine whether it should be disabled, ensuring consistency between the cash flow and system records.

[0054] In travel expense review, the focus is on invoice compliance and business matching. Regarding invoice compliance, for e-invoices, it's necessary to verify the tax rate, tax amount, and whether the buyer's name is a company name. For machine-printed invoices, it's necessary to check if they are complete and bear the official seal, and to determine if the single invoice amount exceeds 100 yuan; if it does, verification results are required. Regarding overtime application matching, when overtime applications are included as attachments, the overtime application information needs to be matched and verified with the invoice time and amount to confirm the reasonableness of the business transaction.

[0055] In e-commerce invoice registration and verification scenarios, three situations are categorized based on different business circumstances: invoice without contract, invoice with contract, and no invoice. In the invoice-without-contract scenario, the main checks are whether a compliant invoice exists in the attachments and whether the paperless option is set correctly. In the invoice-with-contract scenario, it's necessary to check whether the contract terms stipulate a security deposit; if so, further checks are made to ensure that relevant attachments for the security deposit are uploaded, while also verifying the paperless option. In the no-invoice scenario, the application is directly rejected, and a prompt is made to upload the invoice.

[0056] In general material invoice review scenarios, the first step is to review the completeness of attachments, checking whether both the invoice and contract documents exist simultaneously. Next, a review of the warranty deposit or guarantee is conducted to determine if it is necessary to split the warranty deposit or provide a guarantee according to the SRM system requirements. Regarding account information consistency, the receiving account in the invoice or contract is compared with the account information in the expense report system. Finally, regarding contract terms, further verification is conducted to determine if a settlement statement, payment materials, or security deposit are required, and to determine if there are any requirements for splitting the warranty deposit, making payments on aging payments, or providing a guarantee.

[0057] In the scenario of fuel settlement invoice review, the first step is to conduct a classified review, selecting the corresponding review rules based on the expense sub-category (freight, miscellaneous charges, or coal price); regarding the completeness of attachments, check whether there are invoices, contracts, settlement statements, and receiving slips; regarding the supporting materials review, confirm whether auxiliary supporting materials such as railway waybills are provided; regarding the security deposit review, check whether the contract stipulates a security deposit and confirm whether the corresponding attachments are provided.

[0058] In the account transfer verification scenario, the main focus is on the consistency of fund information. For payment information consistency verification, the amount on the expense report is compared with the amount on the bank receipt; for payment account consistency verification, the payment account in the expense report system is verified against the payment account in the bank receipt; for payment method verification, when the payment method is TSS, local fund verification is required to ensure that the fund transfer complies with regulations.

[0059] In the scenario of reviewing reversal-type adjustment applications, the main focus is on verifying the correctness of the payment period. By judging whether the payment period filled in the reimbursement report meets the requirements, the accuracy of the financial processing cycle is ensured.

[0060] In the intelligent agents for reviewing transportation expenses and communication expenses, differentiated rule bases for their respective domains are executed, covering multiple dimensions such as invoice compliance, completeness of attachments, and consistency of business logic, to ensure that different expense categories comply with the corresponding business specifications during the review process.

[0061] When executing the aforementioned rules, a collaborative approach between the rule engine and the large language model is employed. For rules that can be precisely judged through hard coding, such as monetary value comparison and date format validation, the rule engine executes them directly to ensure the determinism and efficiency of the processing results. For complex rules that require semantic understanding, such as contract clause interpretation or comprehensive judgment of invoice compliance, the Qwen2.5-72B large language model is invoked for reasoning and analysis, thereby improving the adaptability to complex business scenarios.

[0062] After the above multi-dimensional differentiated review process, the final output is the execution result of each review dimension. The result includes the review status (passed, failed, or warning), the corresponding review comments, and detailed exception information, which comprehensively reflects the compliance status of the current expense reimbursement documents under various rules and provides basic data support for subsequent cross-modal consistency verification and review comment generation.

[0063] After completing the multi-dimensional differentiated rule review, we have obtained image recognition results, structured reimbursement data, and bank receipt data extracted from scenarios involving fund transfers. These three types of data originate from different information carriers: image recognition results reflect the visual semantic information of invoices and attachments, structured data from the reimbursement system reflects business entry information, and bank-level data reflects the actual flow of funds. Because the three types of data have different sources and different forms of expression, a unified data mapping and cross-validation mechanism is needed to verify the consistency of key fields, thereby identifying potential discrepancies and anomalies and improving the reliability and completeness of the review results.

[0064] In the specific processing, a three-party data mapping table is first established to uniformly model and align fields of data from different sources. Image layer data originates from image recognition results in structured JSON format, primarily extracting information such as amount, payment account, transaction date, reception date, number of people, and hotel name. This data is parsed from the images by a visual language model and possesses strong semantic attributes. Expense layer data originates from the parsing results of the expense reporting system, including fields such as expense amount, payment account, reception date, number of guests, and expense type, reflecting the business data entered into the system. Bank layer data originates from the recognition results of bank receipt images, extracting transaction amounts and payment account information when bank deductions or transfers are involved, used to describe the actual flow of funds. By constructing the three-party data mapping table, data from different sources can be organized under a unified dimension, providing a foundation for subsequent cross-validation.

[0065] After data mapping is completed, the cross-validation rule execution phase begins. For amount consistency verification, the relationship between the image layer amount and the reimbursement layer amount is calculated. The reimbursement layer amount must be less than or equal to the image amount for which an invoice has been issued. If the reimbursement amount exceeds the image amount or exceeds a set threshold, it is considered inconsistent, thus identifying potential over-reimbursement or data entry errors. For account consistency verification, payment account strings are normalized, including removing spaces, standardizing capitalization, and extracting core account segments to eliminate the impact of format differences between different systems. Precise matching is then performed to determine account consistency. For date consistency verification, date strings from different sources are uniformly converted to a standard date format (YYYY-MM-DD), and a difference threshold of one day is set. Comparisons are performed within a reasonable error range to avoid misjudgments due to format or time zone differences. For headcount consistency verification, the headcount values ​​identified by the image layer are compared with the headcount values ​​entered by the reimbursement layer to verify the consistency between the business scale and the actual situation. Regarding name consistency verification, fuzzy matching is performed on text information such as hotel names or supplier names, and similarity is used to determine whether they belong to the same entity, thereby solving the problem of inconsistent name expressions in different data sources. Regarding paperless compliance verification, the consistency between invoice formats (such as OFD, ZIP, XML, etc.) and the paperless options in the expense reimbursement system is checked to determine whether business processes comply with paperless management requirements.

[0066] After executing the above-mentioned verification rules, the verification results need to be marked with discrepancies. When data inconsistencies are detected, specific discrepancy descriptions are generated to clearly indicate the problem. For example, if the payment account number recorded in the expense report system ends in 1234, while the payment account number identified in the bank receipt ends in 5678, a discrepancy description can be generated: the expense report payment account number ends in 1234, and the bank receipt payment account number ends in 5678. Corresponding descriptions are also generated for discrepancies in amount, date deviations, or mismatches in the number of people. This discrepancy marking method not only clearly identifies the location of the problem but also provides clear guidance for subsequent modifications.

[0067] From an overall process perspective, cross-modal consistency verification establishes a unified mapping relationship between image layer data, reimbursement layer data, and bank layer data, and performs cross-validation on multiple key dimensions, achieving a comprehensive check of consistency across different data sources. Compared to single-data-source verification, this multi-source comparison method can effectively uncover inconsistencies hidden between system entries and image vouchers, such as amount discrepancies, account errors, or date anomalies, thereby improving the accuracy and rigor of the audit process. Simultaneously, by introducing techniques such as string normalization, date standardization, and fuzzy matching, the impact of differences in data formats on the verification results can be effectively reduced, improving overall robustness.

[0068] The final output is the cross-modal consistency verification result, which includes the comparison status of each verification dimension, details of differences, and corresponding correction suggestions. The comparison status indicates whether each dimension is consistent, the details of differences describe specific issues, and the correction suggestions provide a reference for subsequent adjustments. This output not only provides important input for the generation of subsequent review comments but also provides intuitive data support for manual review, thus playing a crucial role in the entire automated order review process.

[0069] After completing the multi-dimensional differentiated rule review and cross-modal consistency verification, relatively comprehensive basic audit data has been obtained, including the execution results of each audit dimension, cross-modal consistency verification results, and image recognition results and structured reimbursement data generated in the previous stage. This data reflects the compliance status and potential problems of reimbursement documents from different perspectives, but it still exists in a scattered form, lacking a unified expression and comprehensive judgment. Therefore, it is necessary to integrate and analyze the above information using a large language model to generate unified, clear, and actionable audit opinions, thereby providing a basis for decision-making in subsequent processing.

[0070] In the specific processing, the various input data are first organized in a unified manner. The input content includes the execution results of each audit dimension, cross-modal consistency verification results, image recognition results, and reimbursement data. This data is structured according to a pre-set Prompt template to form complete contextual information and is then input into the Qwen2.5-72B text-based large language model. The Prompt template clearly defines the output format and content requirements for audit opinions, such as stipulating that the output should include the overall audit status, audit conclusions for each dimension, and specific modification suggestions, thereby constraining the structure and standardization of the model's generated results. Simultaneously, by organizing data from different sources in a unified format, the model can understand the relationships between various types of information within the same context, improving inference accuracy.

[0071] After receiving the context information, the model enters the inference and analysis phase. First, the execution results of all review dimensions are summarized to determine if any items fail. The review results for each dimension typically include status information such as pass, fail, or warning. The model performs a comprehensive scan of these statuses to identify key failures and assess their impact on the overall review results. Subsequently, cross-modal consistency verification results are incorporated into the analysis to further determine if there are any inconsistencies across systems, such as amount discrepancies, account mismatches, or date anomalies. This step integrates rule-based review with data consistency verification, resulting in a more comprehensive final judgment.

[0072] After completing the above two types of judgments, the model further generates the overall review status. The overall review status is usually divided into three categories: when all review dimensions pass and there are no cross-modal inconsistencies, it is judged as passed; when there are key non-passing items or serious data inconsistencies, it is judged as failed; when there are missing attachments or correctable issues, it is judged as requiring supplementary attachments. This classification method can compress complex review results into clear decision conclusions, while providing a clear direction for subsequent processing.

[0073] After the overall status is determined, the natural language review opinion generation stage begins. For each review dimension that fails or requires supplementation, the model generates specific natural language modification guidelines. For example, when a discrepancy is detected between the beverage requisition attachments and the beverage requisition amount entered in the requisition form, a prompt can be generated stating that the beverage requisition attachments and the beverage requisition amount entered in the requisition form are inconsistent and should be verified; when there are insufficient signatories, a prompt can be generated stating that the signatures are not approved and should be verified to ensure that there are at least two signatories; when there are consecutive numbering anomalies on invoices, a prompt can be generated stating that consecutive invoices exist and should be rejected. These review opinions not only point out the problem but also provide clear directions for handling it, enabling finance personnel to quickly locate and correct the issue.

[0074] In generating review comments, the large language model not only relies on rule results but also combines image recognition results and contextual information from expense reports to semantically integrate the issues, making the output highly readable and business-adaptable. Compared to traditional rule engines that only output simple flags or error codes, this approach significantly improves the interpretability of review results, making the output more closely aligned with actual business needs.

[0075] Finally, a structured audit report is generated as the output. This report includes the overall audit status, audit details for each dimension, cross-modal validation results, and natural language audit comments. The audit details for each dimension provide a detailed overview of the execution status of each rule, the cross-modal validation results reflect the consistency across different data sources, and the natural language audit comments offer specific modification suggestions. Outputting the report in a structured format not only facilitates subsequent system processing but also makes it easier to display and interact with within the interface.

[0076] Overall, the large-scale model's review opinion generation process realizes the transformation from multi-source review results to unified decision output. Through comprehensive analysis of rule execution results and data consistency information, it forms a review report with a clear structure and semantics, thus completing the key transition from data processing to decision output in the automated order review process.

[0077] After generating the large-scale model's review comments, a structured review report has been obtained. This report integrates the execution results of each review dimension, cross-modal consistency verification results, and natural language review comments, comprehensively reflecting the review conclusions and issues of the current expense reimbursement documents. To achieve systematic output of review results and business closure, this structured review report needs further integration, feedback, and display processing. This will allow the review results to be presented intuitively within the expense reimbursement management system and drive the automatic flow of subsequent business processes.

[0078] In the specific processing, the first step is result integration. The audit report is encapsulated into a standard JSON format using a unified data structure to ensure good compatibility and scalability for data interaction between different systems. The encapsulated JSON data contains several key fields, including: expense report number (identifying the current expense report); expense type (indicating the business category); audit status (indicating the current audit conclusion, with values ​​of "passed," "failed," or "pending supplementation"); details of each audit dimension (recording the execution result of each audit rule, including the dimension name, audit result, and corresponding audit comments); cross-modal validation results (showing the consistency between image data, expense report data, and bank data); image recognition confidence score (reflecting the reliability of the image recognition results); natural language audit comments (providing overall modification suggestions); and a processing timestamp (recording the audit completion time). This structured encapsulation method unifies the expression of scattered audit information, providing a standardized foundation for subsequent data transmission and display.

[0079] After the results are integrated, the process moves to the result feedback stage. The JSON-formatted audit report is sent back to the RPA robot via a standard API call. This interface serves as a data exchange channel, ensuring seamless transfer of audit results from the intelligent audit process to the automated execution stage. Upon receiving the audit report, RPA can perform subsequent operations based on its content, while maintaining data synchronization with the expense reimbursement management system. Using the API for feedback not only improves data transmission efficiency but also ensures consistency and reliability during data transmission.

[0080] After the results are returned, the RPA closed-loop execution phase begins. By invoking the DisplaySkill, the audit results are output to the expense reimbursement management system page, mapping key information from the structured audit report to the system interface. The displayed content includes the overall audit status and detailed audit information for each item, allowing finance personnel to intuitively understand the current document's audit status. Simultaneously, the business process is automatically driven based on the audit status. For example, when the audit status is "approved," the initial review is automatically completed and the process proceeds to the secondary review stage; when the audit status is "failed" or "pending supplementation," anomalies are marked in the system, and the required modifications are displayed. In this way, automatic connection between audit result generation and business process execution is achieved.

[0081] During the presentation, not only can the overall audit conclusions be displayed, but also the audit results for each dimension can be shown in a granular manner. For example, each item that failed the audit can be listed along with its corresponding reason, enabling finance personnel to quickly pinpoint the source of the problem. Simultaneously, combined with natural language audit comments, clear modification guidance can be provided, reducing the cost of human interpretation and improving processing efficiency. Auxiliary information such as image recognition confidence level can also be used to assess the reliability of the recognition results, prompting manual review when necessary.

[0082] Overall, the results integration and closed-loop output process achieves a complete closed loop from the generation of audit data to its display in the business system. Through the collaborative processing of three stages—structured encapsulation, interface feedback, and interface display—the audit results are not only accurately received by the system but also directly drive the execution of business processes, providing clear reference information for finance personnel. Finally, after completing the initial audit process, finance personnel can observe the audit results and various details in the system, thus achieving an effective connection between automated auditing and manual decision-making.

[0083] This invention achieves a systematic reconstruction of the financial document review process through a collaborative architecture that integrates multimodal preprocessing, deep visual language understanding, dedicated intelligent agent routing, cross-modal cross-validation, large-scale model opinion generation, and RPA closed-loop processing. By constructing a unified multimodal data processing link, the previously scattered image data, structured expense report data, and attachment information are standardized. Combined with the deep understanding of image layout semantics by a visual language model, the review process can not only acquire textual information but also understand document structure and business semantics, effectively solving the problem that a single model cannot cover complex and differentiated rules. Simultaneously, by breaking down the review process into multiple collaborative stages, the overall processing stability and scalability are improved, enabling efficient automated processing in complex business scenarios.

[0084] Compared to the approach of dynamically generating tool call sequences using a single large model, this invention introduces an adaptive routing mechanism based on cost type. This mechanism splits complex review tasks into multiple dedicated agents based on cost categories, allowing each agent to focus solely on the review rules of its corresponding domain. This significantly reduces the model's cognitive load and inference complexity, minimizing the risk of illusions caused by multi-task mixing. Furthermore, it employs a pre-arranged workflow pipeline instead of dynamic tool orchestration, designing a fixed path for the review process. This ensures a clear and controllable execution order for each processing node, effectively avoiding issues such as chaotic tool call order or missing parameters. Consequently, it maintains stable and high-precision review results across various differentiated business scenarios, including entertainment expenses and travel expenses.

[0085] Compared to traditional OCR-rule engine-based solutions, this invention introduces a visual language model to achieve semantic understanding of image layouts. It not only extracts text but also identifies document types, determines the compliance of the number of signatories, detects consecutive invoice numbers, and parses non-standard attachments, overcoming the limitations of traditional OCR technology that only focuses on character recognition. Furthermore, by constructing a cross-modal consistency verification mechanism, it performs three-way cross-validation of image data, expense report system data, and bank statement data, automatically identifying issues easily overlooked in manual review, such as discrepancies in amounts and account inconsistencies. In addition, through RPA-automated workflow engine triggering, it achieves fully automated processing from data collection to result output, eliminating the need for manual input of natural language or interaction. This meets the unmanned batch initial review needs of financial shared service centers and significantly improves review efficiency.

[0086] This invention provides, for example Figure 5 The system shown is a multi-agent collaborative financial document review system based on adaptive routing of cost type, comprising an RPA input layer, a workflow engine layer, a dedicated review agent layer, a processing layer, a model layer, and a plugin layer. The RPA input layer is used to perform interactive operations of the financial system through a preset set of skills, which includes acquisition skills, triggering skills, and display skills. Acquisition skills are used to automatically log in to the expense reimbursement management system and obtain expense reimbursement information and image compressed data. Triggering skills are used to call the workflow engine interface to upload image compressed data and structured expense reimbursement data to trigger the review process. Display skills are used to receive the review results and fill in the review information and display the exception items on the expense reimbursement management system page. The workflow engine layer is used to uniformly schedule and orchestrate the review tasks. It is pre-configured with processing pipelines including file decompression nodes, type classification nodes, format conversion nodes, image recognition nodes, fee type routing nodes, rule review nodes, cross-modal verification nodes, result integration nodes, and output nodes to realize the sequential execution and node scheduling of the review process. A dedicated audit agent layer is used to divide multiple independent audit agents according to the fee type. Each audit agent corresponds to a different fee category and has a built-in differentiated audit rule library and audit prompt template to handle audit tasks of different fee types in a targeted manner. The processing layer is used to perform specific data processing operations, including file decompression and classification, data format conversion, image content recognition, rule review, cross-modal consistency verification, and review result integration. The model layer provides intelligent analysis capabilities, including a visual language model and a text-based large language model. The visual language model is used to identify document types and extract key fields from image data, while the text-based large language model is used to perform rule-based reasoning, cross-modal validation logic processing, and generate audit comments. The plugin layer provides basic processing tool support, including PDF to PNG conversion plugin, Base64 encoding / decoding plugin, compressed file decompression plugin, Excel to JSON conversion plugin, and large model interface call plugin, to support data preprocessing and model calling processes. The layers interact with each other through standard interfaces or messaging mechanisms to achieve a complete closed-loop processing flow from data collection, task scheduling, intelligent review to result output.

[0087] This invention provides a method for auditing multi-agent collaborative financial documents based on adaptive routing of cost type. This method is implemented through the aforementioned system for auditing multi-agent collaborative financial documents based on adaptive routing of cost type. For details on the specific methods and processes of the system for auditing multi-agent collaborative financial documents based on adaptive routing of cost type, please refer to the embodiment of the method for auditing multi-agent collaborative financial documents based on adaptive routing of cost type, which will not be repeated here.

[0088] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A multi-agent collaborative financial document review method based on cost type adaptive routing, characterized in that, Includes the following steps: The compressed files of reimbursement document attachments and structured reimbursement data in the expense reimbursement management system are preprocessed by performing decompression, type classification, format conversion, encoding restoration and table structuring to obtain image data set and structured reimbursement data; The image data set is input into the visual language model for recognition, and the document type, key fields and layout semantic information are obtained, and the structured image recognition results are output. The structured expense report data is parsed to extract basic information, amount information, business information, and attachment index information of the expense report, forming an expense report data object; The document type information in the image recognition results is matched with the expense type information in the expense data object to determine the expense category, and the review task is assigned to the corresponding dedicated review intelligent agent. Within the corresponding dedicated auditing intelligence, differentiated auditing rules are invoked to perform audits on the image recognition results and the reimbursement data objects regarding the completeness of attachments, consistency of amounts, business logic, and compliance, and to output audit results in various dimensions. Based on the review results, cross-modal consistency verification is performed on the image recognition results, reimbursement data objects, and bank receipt data to complete multi-dimensional comparison of amount, account, date, number of people, and name, and obtain the consistency verification results; The review results and consistency verification results are input into the text big language model for analysis, and natural language review comments containing review status and anomaly descriptions are generated. The review comments are structured and packaged and sent back to the expense reimbursement management system for display, completing the multi-agent collaborative financial document review based on expense type adaptive routing.

2. The multi-agent collaborative financial document review method based on cost type adaptive routing according to claim 1, characterized in that, Before multimodal attachments enter the intelligent processing flow, heterogeneous data needs to be uniformly organized and format-converted to form a standardized input structure. The steps are as follows: Receive compressed attachments and structured expense reimbursement data output from the expense reimbursement management system, and record the association between the attachment source identifier and the expense reimbursement form; Perform decompression on the compressed file, read the internal file directory structure of the compressed package and generate a set of raw files containing file paths and type identifiers; The original file set was MIME type identified and classified into PDF files, XML files, image files, and reimbursement information files according to file extensions and content characteristics, and a classification index was established. The system performs image conversion processing on PDF files, Base64 encoding and parsing on XML files to restore image data, and table parsing on expense report files to convert them into JSON structures, forming a unified set of image data and structured data objects.

3. The multi-agent collaborative financial document review method based on cost type adaptive routing according to claim 1, characterized in that, To address the complex layout and semantic information contained in financial images, structured parsing and content recognition are required. The steps are as follows: Receive the pre-processed image data set and establish an association mapping between image numbers and invoices according to file order; Image data is input into a visual language model for multimodal feature extraction, including text region detection, layout structure recognition, and semantic feature encoding. Perform document type determination and field extraction processing, identify the title area, field area and key content area according to the layout structure, and extract the corresponding business field information; The recognition results are structured and organized to form an image recognition result set that includes document type labels, field key-value pairs, confidence information, and layout semantic analysis conclusions.

4. The multi-agent collaborative financial document review method based on cost type adaptive routing according to claim 1, characterized in that, Based on the characteristics of structured data in expense reimbursement processes, the following steps are taken to extract and organize the fields of the reimbursement information: Obtain the structured expense reimbursement data output by the expense reimbursement management system and parse the data format to be XML or JSON; Perform field mapping processing on the reimbursement data to extract basic information fields such as reimbursement number, handler, expense type, and application date; Parse the amount-related fields, including the total expense report amount, detailed expense report amount, tax amount and tax rate, and establish numerical relationships between the fields; The business information and attachment index information are integrated to extract payment methods, account information, number of people, and attachment list to form a complete reimbursement data object.

5. A multi-agent collaborative financial document review method based on cost type adaptive routing according to claim 4, characterized in that, During the formation of the reimbursement data object, the field structure is further standardized. A unified field code is established for the reimbursement bill number, expense type and application date. The amount field is processed with uniform precision and hierarchical association. A mapping relationship is established between the payment account and the attachment index, forming a structured reimbursement data set with field consistency and association.

6. The multi-agent collaborative financial document review method based on cost type adaptive routing according to claim 1, characterized in that, In scenarios where different fee categories correspond to different review logics, task routing and path selection processing need to be completed. The steps are as follows: Receive image recognition results and reimbursement data objects, and establish the field correspondence between the two types of data; Perform a combined matching process on the expense type field and the document type label to construct a set of expense category determination conditions; Based on the matching results, determine the target cost category and generate the corresponding route identification information; The image recognition results, reimbursement data objects, and routing identifiers are encapsulated to form an audit task package and assigned to the corresponding audit agent.

7. A multi-agent collaborative financial document review method based on cost type adaptive routing according to claim 1, characterized in that, To address the differences in business rules for various types of expenses, multi-dimensional review and semantic analysis are conducted, following these steps: Receive the audit task package and load the set of differentiated rules for the corresponding fee category; The task data is processed by rule matching, and each item is checked for attachment verification, amount comparison, date verification and field consistency detection. For content involving semantic judgment, a large text language model is invoked for semantic parsing, including contract clause recognition and complex business relationship judgment; The execution results of each rule are summarized to form a multi-dimensional audit result set that includes the rule name, execution status, and exception information.

8. The multi-agent collaborative financial document review method based on cost type adaptive routing according to claim 1, characterized in that, To address the inconsistency caused by differences in data sources, a unified comparison and discrepancy identification process is performed, following these steps: Receive image recognition data, expense report data, and bank statement data, and extract the amount, account number, date, and name fields; Perform standardization processing on fields from different sources, including string normalization, date format conversion, and numerical uniformity processing; Conduct multi-dimensional comparison operations, including numerical comparison of amounts, character matching of accounts, and fuzzy matching of names; The comparison results are marked with differences, generating verification result data that includes the difference fields, difference content, and location identifiers.

9. A multi-agent collaborative financial document review method based on cost type adaptive routing according to claim 1, characterized in that, During the process of outputting audit results and interacting with the business system, it is necessary to complete result integration and feedback processing. The steps are as follows: Receive the audit result set and consistency verification data, and establish an index relationship for the result data; The review status is categorized and processed, and the data is encapsulated according to a preset field structure to generate a JSON format report; The structured report content is parsed to generate natural language descriptions and associated with corresponding anomalies; The results are returned via API calls, and the review status and detailed information are mapped to the system display interface.

10. A multi-agent collaborative financial document review system based on adaptive routing of cost type, used to implement the multi-agent collaborative financial document review method based on adaptive routing of cost type as described in any one of claims 1-9, characterized in that, It includes the RPA input layer, workflow engine layer, dedicated auditing agent layer, processing layer, model layer, and plugin layer: The RPA input layer is used to perform interactive operations of the financial system through a preset set of skills, which includes acquisition skills, triggering skills, and display skills. Acquisition skills are used to automatically log in to the expense reimbursement management system and obtain expense reimbursement information and image compressed data. Triggering skills are used to call the workflow engine interface to upload image compressed data and structured expense reimbursement data to trigger the review process. Display skills are used to receive the review results and fill in the review information and display the exception items on the expense reimbursement management system page. The workflow engine layer is used to uniformly schedule and orchestrate the review tasks. It is pre-configured with processing pipelines including file decompression nodes, type classification nodes, format conversion nodes, image recognition nodes, fee type routing nodes, rule review nodes, cross-modal verification nodes, result integration nodes, and output nodes to realize the sequential execution and node scheduling of the review process. A dedicated audit agent layer is used to divide multiple independent audit agents according to the fee type. Each audit agent corresponds to a different fee category and has a built-in differentiated audit rule library and audit prompt template to handle audit tasks of different fee types in a targeted manner. The processing layer is used to perform specific data processing operations, including file decompression and classification, data format conversion, image content recognition, rule review, cross-modal consistency verification, and review result integration. The model layer provides intelligent analysis capabilities, including a visual language model and a text-based large language model. The visual language model is used to identify document types and extract key fields from image data, while the text-based large language model is used to perform rule-based reasoning, cross-modal validation logic processing, and generate audit comments. The plugin layer provides basic processing tool support, including PDF to PNG conversion plugin, Base64 encoding / decoding plugin, compressed file decompression plugin, Excel to JSON conversion plugin, and large model interface call plugin, to support data preprocessing and model calling processes. The layers interact with each other through standard interfaces or messaging mechanisms to achieve a complete closed-loop processing flow from data collection, task scheduling, intelligent review to result output.