A dynamic auditing and risk control supervising method based on multi-agent
Patent Information
- Application Number
- CN202510868658.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-06-26
AI Technical Summary
现有的OCR技术虽然能够提取图像中的文本,但对文档布局、表格结构等复杂信息的提取精度较低,导致审计过程中信息丢失或误解读;
[0036]1)本发明支持多源数据接入,兼容性强,构建异构数据检索库能够接收来自多种异构数据源的文档,包括Word、PPT、Excel、txt、图片、PDF、结构化数据和网页等格式,突破了传统审计系统在数据格式上的局限性,实现了对多样化文档的全面解析和处理,为动态审计与风险控制提供了完整的数据基础。
Smart Images

Figure CN120851588B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of financial auditing, corporate risk control, and compliance inspection, and specifically to a dynamic auditing and risk control supervision method based on multiple agents. Background Technology
[0002] With the rapid development of information technology, especially the widespread application of big data and artificial intelligence, enterprises and organizations face increasingly complex and massive data sources, including text data, financial data, images, and charts, when conducting audits and risk control. Traditional auditing methods rely on manual labor or simple automated tools, resulting in low efficiency, low accuracy, and difficulty in handling the increasing number of heterogeneous data types and complex data structures. Existing auditing methods typically rely on rule engines and manual analysis, lacking effective processing capabilities for unstructured data, particularly in handling various document formats such as images, PDFs, and Word documents. These methods struggle to achieve efficient risk prediction and automated auditing, leading to high labor costs and error rates during the audit process, and failing to promptly identify potential risks, thus affecting the accuracy and compliance of audit results. Therefore, developing a multi-agent-based dynamic auditing and risk control supervision method that can efficiently process various heterogeneous data, accurately extract document content, automatically identify risk points, and conduct compliance checks in conjunction with regulations has become an urgent need to address the problems of existing technologies.
[0003] Existing audit and risk control methods mainly suffer from the following technical problems:
[0004] 1) Traditional auditing tools are mostly designed for processing structured data (such as database tables), and lack effective parsing and analysis capabilities for unstructured data (such as Word, PPT, PDF, images, etc.). Although existing OCR technology can extract text from images, its accuracy in extracting complex information such as document layout and table structure is low, leading to information loss or misinterpretation during the auditing process;
[0005] 2) Most existing automated auditing systems rely on preset rules and simple text analysis, lacking intelligent risk prediction capabilities and struggling to handle complex financial data or cross-document risk correlation analysis. This makes it difficult for audits to detect potential risks in a timely manner, and the accuracy of audit results cannot be guaranteed.
[0006] 3) Existing document retrieval methods mainly rely on simple keyword matching, lacking deep semantic understanding and vectorized representation, making it difficult to effectively handle contextual relationships and semantic similarity between documents. Especially when faced with massive amounts of documents and complex queries, existing technologies cannot provide sufficiently efficient and accurate retrieval results;
[0007] 4) Existing audit systems often cannot perform comprehensive analysis by combining multiple data sources, and cannot identify and classify risk points from a holistic perspective. In particular, traditional systems lack the capacity to accurately classify multiple types of risks.
[0008] 5) The current audit system lacks deep integration with the content of regulations. The retrieval of regulations and compliance checks during the audit process rely on manual operation, which leads to uncertainty and delay in audit results and makes it difficult to detect compliance issues in a timely manner. Summary of the Invention
[0009] To address the aforementioned issues, this invention proposes a routine monitoring method based on a multi-agent system framework and in conjunction with the actual needs of auditing and risk control, enabling efficient risk identification, dynamic collaboration, real-time feedback, and closed-loop optimization.
[0010] To achieve the above objectives, the present invention adopts the following technical solution:
[0011] This invention discloses a dynamic auditing and risk control monitoring method based on multiple agents, comprising:
[0012] Phase of building a heterogeneous data retrieval library:
[0013] Receive historical case studies, key risk elements, and regulatory documents from various heterogeneous data sources, including unstructured documents and structured data;
[0014] Perform deep parsing on unstructured documents to obtain plain text data;
[0015] Vectorization and embedding are performed on plain text data to generate document index data that supports efficient retrieval, which is then stored in a vector database.
[0016] Structured data is used to generate data table information and field information, and database tables are dynamically created in the relational database and stored in the relational database.
[0017] By combining vector databases and relational databases, a heterogeneous data source document retrieval library is realized, which includes a risk feature library and a regulatory content library.
[0018] Dynamic audit and risk control supervision phase:
[0019] Accepts various text review rules;
[0020] Receive the documents to be audited, execute the heterogeneous data retrieval library process, and obtain unstructured risk elements and structured risk elements from the risk feature library;
[0021] Risk points and risk types are classified based on unstructured and structured risk elements;
[0022] Based on the risk points and risk types, the process engine is invoked to execute automated audit processes and obtain audit anomaly result data.
[0023] Based on the audit anomaly results data, the regulatory content database is searched and enhanced to generate compliance review results data;
[0024] Structured audit results data are obtained based on audit anomaly data and compliance review results data.
[0025] As a further improvement, the deep parsing of unstructured documents described in this invention specifically includes:
[0026] A heterogeneous data source document parsing agent is constructed. An object detection model is used to identify various regions on the page, and the location of each region is predicted by bounding box regression. For the detected table regions, the table structure is parsed to further identify its row and column structure and cell distribution. Each cell is cut into an independent image, and the text content is extracted by OCR technology.
[0027] As a further improvement, the present invention describes the vectorization and embedding of plain text data to obtain document index data that supports efficient retrieval, specifically as follows:
[0028] A retrieval enhancement generation system based on the Elasticsearch vector database is constructed. The retrieval enhancement generation system uses the Tongyi Qianwen 2-1.5B model to vectorize each text fragment to generate an embedding vector, and imports the generated vector and its related metadata into the Elasticsearch vector database.
[0029] As a further improvement, the heterogeneous data retrieval library retrieval process described in this invention is specifically as follows:
[0030] For unstructured documents, deep parsing is first performed to obtain plain text data. Then, the Tongyi Qianwen 2-1.5B model is used for vector embedding generation. Document block vector retrieval based on vector similarity is performed in the vector database to obtain unstructured retrieval results. The Tongyi Qianwen 2.5-32B model is used in conjunction with the SQL example and table field information of the database to construct dynamic SQL queries. Self-verification and correction are performed based on the SQL query results to obtain structured retrieval results.
[0031] As a further improvement, the classification of risk points and risk types based on unstructured risk factors and structured risk factors described in this invention specifically includes:
[0032] An intelligent risk identification agent is constructed, which combines the text data of the documents to be audited with the unstructured and structured risk elements obtained by retrieval, and combines the review rules of the risk category in the text review rules with positive and negative cases in them to perform multi-dimensional cross-validation and automatic classification. Semantic enhancement analysis is performed through the Tongyi Qianwen 2.5-32B model to identify potential risk points and classify risk types.
[0033] As a further improvement, the method of classifying risks based on risk points and risk types and then calling the process engine to execute the automated audit process as described in this invention is as follows:
[0034] An automated auditing agent is built, which performs text review based on the review rules for sensitive words among various review rules. The Tongyi Qianwen 2.5-VL-32B model is used to extract financial data for trend assessment, determine whether there are abnormal trends, and perform data anomaly detection. Error correction review supports the review of Prompt instruction rules based on the Tongyi Qianwen 2.5-32B model. By defining and applying specific error correction rules, the text content is analyzed and text error correction suggestions are returned.
[0035] The beneficial effects of this invention are as follows:
[0036] 1) This invention supports multi-source data access, has strong compatibility, and can build a heterogeneous data retrieval library that can receive documents from various heterogeneous data sources, including Word, PPT, Excel, txt, images, PDF, structured data and web pages, etc. It breaks through the limitations of traditional auditing systems in terms of data formats, realizes comprehensive analysis and processing of diverse documents, and provides a complete data foundation for dynamic auditing and risk control.
[0037] 2) This invention supports high-precision parsing of unstructured data. It constructs a heterogeneous data source document parsing agent and uses object detection models, OCR technology and table structure recognition (TSR) for deep parsing. This technology can accurately extract text content, table structure and nested information. Through grayscale, binarization, object detection and Hough transform technology, it further enhances the parsing ability of complex tables and nested documents, ensuring the accuracy and completeness of data extraction.
[0038] 3) This invention boasts high retrieval efficiency and accuracy. The Elasticsearch-based retrieval enhancement generation system in the heterogeneous data document retrieval agent employs retrieval enhancement generation technology to segment document content into document blocks and achieves efficient retrieval through vector similarity and keyword similarity, greatly improving the efficiency and accuracy of document retrieval. By utilizing dynamic SQL generation technology, it achieves accurate querying of structured data. Combining vector database retrieval results with relational database retrieval results significantly improves the coverage of risk factors.
[0039] 4) This invention provides the ability to annotate and learn historical cases and key risk elements. In the process of constructing a risk feature library, annotating historical cases and key risk elements and storing them in relational databases and Elasticsearch vector databases, the retrieval enhancement generation system is continuously iterated and optimized, making the system more adaptable and scalable.
[0040] 5) This invention conducts in-depth analysis of unstructured and structured risk factors, and automatically identifies and classifies potential risk points using the Tongyi Qianwen 2.5-32B model. This intelligent identification not only improves the accuracy of risk identification, but also continuously optimizes the risk classification model based on historical data and cases, enabling the system to have self-learning and evolution capabilities, thereby better adapting to the complex and ever-changing audit environment.
[0041] 6) This invention, by invoking a process engine, constructs an automated audit process agent capable of automatically executing various audit processes, such as text review, financial data analysis, and error correction review. This automated process not only reduces manual intervention but also significantly improves audit efficiency, especially when processing large amounts of data, enabling the rapid generation of audit anomaly data to ensure the timeliness and accuracy of audit work.
[0042] 7) This invention enhances the retrieval function of the regulatory content database through vectorization. During the compliance review result generation stage, it utilizes the vectorized retrieval and enhanced analysis mechanism of the regulatory content database to quickly retrieve relevant regulations and cases. Combined with the Tongyi Qianwen 2.5-32B model for enhanced analysis, it generates highly reliable compliance review results. This review method not only covers a wide range of regulatory content but also updates in real time based on the latest regulatory developments, ensuring that audit results always comply with current regulatory requirements and effectively reducing compliance risks.
[0043] 8) This invention supports generating data reports based on preset template rules and allows users to dynamically generate customized data reports by inputting query statements. This flexibility not only meets the personalized needs of different users but also allows for real-time updates of report content based on changes in audit results, ensuring the timeliness and accuracy of the reports.
[0044] 9) After generating audit results, this invention can perform data cleaning, transformation, and standardization preprocessing operations to ensure data accuracy and consistency. This processing method not only improves data quality but also provides a reliable foundation for subsequent data analysis and decision-making.
[0045] 10) This invention adopts a modular design, supporting rapid adaptation to new data sources and audit requirements, and possesses strong scalability and compatibility. Whether it's a new data source or new audit requirements, the system can quickly adapt and expand through its modular design, ensuring long-term availability and adaptability.
[0046] 11) This invention employs strict data encryption and access control mechanisms during data processing and storage to ensure data security and privacy. Whether it's sensitive corporate data or personal privacy information, the system provides effective protection against data leakage and misuse.
[0047] 12) This invention supports cross-platform operation, enabling it to run on different operating systems and devices, such as Windows, Linux, and Mac, while also supporting multiple terminal devices such as PCs and mobile devices. This cross-platform and multi-terminal support allows users to access the system anytime, anywhere to perform auditing work, improving the flexibility and convenience of their work.
[0048] 13) This invention provides a wealth of data visualization tools and interactive analysis functions, allowing users to intuitively view audit results and conduct in-depth data analysis through charts, dashboards, and other formats. These visualization and interactive analysis functions not only improve data readability but also help users better understand the meaning behind the data and make more accurate decisions.
[0049] 14) This invention significantly reduces labor and time costs and improves audit efficiency through automated and intelligent audit processes. Simultaneously, by optimizing data processing and storage methods, the system reduces hardware resource consumption, improves overall system performance and resource utilization, and saves enterprises substantial operating costs. Attached Figure Description
[0050] Figure 1 This is a schematic diagram illustrating the process of building a heterogeneous data retrieval library;
[0051] Figure 2 This is a schematic diagram of the retrieval process of a heterogeneous data retrieval database.
[0052] Figure 3 This is a flowchart illustrating a dynamic auditing and risk control supervision method based on multiple agents. Detailed Implementation
[0053] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific implementation examples:
[0054] Figure 1 This is a schematic diagram illustrating the process of building a heterogeneous data retrieval library; Figure 2This is a schematic diagram of the retrieval process of a heterogeneous data retrieval database. Figure 3 This is a flowchart illustrating a multi-agent-based dynamic auditing and risk control monitoring method. According to a specific embodiment of the present invention, a multi-agent-based dynamic auditing and risk control monitoring method is provided, specifically including the following steps:
[0055] I. Phase of Building a Heterogeneous Data Retrieval Library:
[0056] 1) such as Figure 1 The system receives documents from various heterogeneous data sources, including unstructured documents and structured data, containing historical cases, key risk elements, and regulatory content. It uses a standardized data access module that supports batch uploading and real-time access, and performs preliminary classification and storage of various types of data.
[0057] 2) such as Figure 1 The deep parsing module shown performs deep parsing on unstructured documents. It integrates a text parser, an OCR parser, a layout recognition parser, and a TSR (table recognition) parser. For documents that can be converted to images, it uses libraries such as PyPDF2, python-docx, and pptx to convert them into image formats. The images are then converted to grayscale to reduce computational complexity. For color images, the grayscale pixel value I of the RGB three channels can be calculated using the following formula, and empirically optimized coefficients are used to better suppress specific types of noise and highlight the contrast between text areas and the background.
[0058] I=0.299·R+0.587·G+0.114·B
[0059] Binarization is performed using an adaptive thresholding method based on the Otsu algorithm. The Otsu method automatically selects the optimal threshold T by maximizing the inter-class variance, as shown in the formula:
[0060]
[0061] Based on the various regions detected by the object detection model on the page, the location of each region is predicted through bounding box regression. This is used to precisely adjust the predicted bounding box positions of the model, making them closer to the true labeled boxes. The loss function is calculated using the following formula:
[0062]
[0063] Among them B p Given a predicted bounding box, B g For the true bounding box, smooth L1 For smoothing loss function:
[0064]
[0065] For table regions, an object detection-based model is used. The goal is to detect the bounding boxes of the table and further extract its structure, performing table structure parsing, further analyzing the row and column structure and cell distribution of the table, using Hough transform to detect horizontal and vertical lines within the table, grouping the detected lines to determine the row and column positions of the table, generating cell coordinates through the intersections of horizontal and vertical lines, and cropping the image of each cell based on the bounding boxes of the table cells. OCR is used to extract the text content, and the results are post-processed, including: noise reduction, error correction, and cell merging.
[0066] 3) such as Figure 1 The structured data generation process shown in the diagram first utilizes the Tongyi Qianwen 2.5-32B model for deep analysis and semantic recognition of the structured data. This intelligently extracts the inherent logical structure and attribute features of the data, thereby automatically generating accurate data table schema information. This includes, but is not limited to, inferring the optimal table name, field names, assigning appropriate data types (such as integer, string, date, floating-point, etc.) to each field, and potential constraints (such as primary key, NOT NULL, uniqueness, etc.). Based on the generated data table schema information, instructions are sent to the target relational database system through a standard database interface protocol to achieve real-time dynamic creation of the data table without the need for manual pre-definition of the table structure. Finally, each record in the original structured data is mapped and transformed according to the newly created table structure, and efficiently and accurately batch-loaded into the corresponding dynamically generated data table to complete the persistent storage of the data.
[0067] 4) such as Figure 1 The retrieval enhancement generation system shown utilizes the plain text data of unstructured documents obtained through deep parsing. It then uses the Tongyi Qianwen 2-1.5B model to vectorize each text segment, generating embedding vectors. The generated vectors and their associated metadata are imported into Elasticsearch, where the document content is embedded using the Tongyi Qianwen 2-1.5B model. Specifically:
[0068] v d =Embed(d)
[0069] Where V d Embed(d) is a vectorized representation of the document, a high-dimensional semantic vector generated by the General Meaning Questions 2-1.5B model, which is based on vector V in Elasticsearch. d Create a vector retrieval index;
[0070] 5) Import the risk characteristic case data, risk key element data, and regulatory content data into the heterogeneous data retrieval library through steps 3 and 4 to obtain the risk characteristic library and the regulatory content library.
[0071] II. Dynamic Audit and Risk Control Supervision Phase:
[0072] 6) such as Figure 2 The retrieval process of the heterogeneous data retrieval library shown first involves deep parsing to obtain plain text data for unstructured documents. Then, the Tongyi Qianwen 2-1.5B model is used for vectorized embedding. Finally, document block vector retrieval based on vector similarity is performed in the vector database. The vector similarity calculation formula is as follows:
[0073]
[0074] Among them, Sim(v q ,v d The cosine similarity between the query document and the candidate documents is used to rank them. The most similar document blocks are retrieved in the retrieval enhancement generation system to obtain unstructured retrieval results. For structured data, the Tongyi Qianwen 2.5-32B model is used to construct dynamic SQL queries by combining the SQL examples and table field information of the database. Self-verification and correction are performed based on the SQL query results to obtain structured retrieval results.
[0075] 7) Receive various text review rules and corresponding cases, store them in a relational database. Specific fields include review items, review rules, positive and negative cases, review level, and rule category. Use the Tongyi Qianwen 2.5-32B model to rewrite the review rules and generate Prompt instructions. Specific rule categories include: risk, sensitive words, financial review, text correction, etc.
[0076] 8) For example Figure 3 The risk identification agent shown receives the documents to be audited and uses the heterogeneous data retrieval process in step 6 to segment the text data of the documents to be audited into multiple semantic fragments in the risk feature library. These fragments are then compared with document blocks in the risk feature library. Combining vector similarity and keyword matching, the agent selects the risk elements most relevant to the text data of the documents to be audited. Similarity is calculated based on the semantic embedding vectors of the content to ensure that semantically related but differently expressed content can be retrieved. Specific keywords or risk tags are weighted to ensure that core risk information is retrieved, obtaining unstructured risk elements. Using the Tongyi Qianwen 2.5-32B model, an SQL statement is generated, and the SQL query is optimized to dynamically query structured data in the risk feature library to obtain structured risk elements.
[0077] 9) For example Figure 3The risk identification agent shown uses the text data of the document to be audited and the unstructured and structured risk elements obtained by retrieval. It combines the risk-category review rules in the text review rules with positive and negative cases to perform multi-dimensional cross-validation and automatic classification. Semantic enhancement analysis is performed through the Tongyi Qianwen 2.5-32B model. The risk-category review rules define the review rule text. For example, if the transaction amount between the enterprise and related parties is large and lacks commercial rationality or the pricing is unfair, there may be a risk of transfer of benefits. The rule text is rewritten by the model to generate a specific prompt, which can be used to cross-validate the corresponding positive and negative cases in the collaborative review rules. The risk elements, the prompts generated from the rules, and the corresponding positive and negative risk cases are assembled into a complete context and submitted to the model to identify potential risk points and classify risk types.
[0078] 10) such as Figure 3 The various automated audit agents shown are categorized based on risk points and risk types. They invoke the workflow engine for text review, using regular expressions and the Tongyi Qianwen 2.5-32B model for semantic matching to scan the audited document text data according to the review rules for sensitive words. This detects the presence of sensitive words or synonyms, marks the location of sensitive words and records the context for auditors to quickly locate them. Predefined audit domain standard terminology is used to compare documents using terminology specifications, detecting inconsistencies or misuses, marking inconsistencies and providing suggested alternatives. A checklist of key content is defined, and the Tongyi Qianwen 2.5-32B model is used to analyze the document structure, checking for missing key content or related sections. Missing information is marked and suggested for supplementation, generating structured analysis data and yielding text audit anomaly results.
[0079] 11) such as Figure 3 The various automated audit agents shown are categorized based on risk points and risk types. They invoke a process engine to analyze financial data, using the Tongyi Qianwen 2.5-VL-32B model to extract financial data (such as revenue, expenses, and profits) from documents. They plot trend charts for data across multiple consecutive accounting periods, analyze growth or fluctuation trends, mark abnormal trends, and, in conjunction with review rules categorized as financial review in the text review rules, detect potential outliers in the financial data. These outliers are compared with industry averages or historical data, anomaly data points are marked, and potential risk warnings are output. By comparing financial data from different accounting periods, significant deviations in indicators are detected, generating financial audit anomaly results data.
[0080] 12) For example Figure 3The various automated audit agents shown are categorized according to risk points and risk types, and invoke the process engine for error correction review. The text review rules define multiple rule categories for text error correction review. Combining relevant cases within the error correction rules, the Tongyi Qianwen 2.5-32B model detects incorrect citations or data conflicts, verifies the data or facts referenced in the document, marks errors, and provides correct data. It also checks whether the document format conforms to specifications, applies format templates, and generates error correction review result data.
[0081] 13) such as Figure 3 The compliance review process is as follows: Based on various audit anomaly results, the heterogeneous data retrieval process in step 6 is used to search for relevant regulatory clauses in the regulatory content library. The most relevant regulatory content and case information are output, and the Tongyi Qianwen 2.5-32B model is used to combine regulatory content, case information, preset prompts, and the text data of the document to be audited to analyze whether the document violates the law. Detailed compliance review results are generated, including violation clauses, case explanations, and improvement suggestions. The identified risk points are compared with the regulatory content one by one to verify whether the risk points are legal and compliant. For unclear risk points, prompts are generated for auditors to further verify.
[0082] 14) For example Figure 3 The data cleaning feedback is performed as shown. Based on the audit anomaly results data and compliance review results data, the data is cleaned, transformed, and standardized preprocessed before being stored as structured audit results data and then stored in a relational database.
[0083] 15) For example Figure 3 The system provides data feedback, defines common data report template formats, including titles, table layouts, and chart types, uses dynamic placeholders to design templates so that the generated report content can be dynamically replaced, and updates the templates regularly to ensure they meet the latest business needs. It also extracts structured data from audit results and dynamically embeds the data into data reports based on preset templates.
[0084] 16) For example Figure 3 The system receives user-input query statements, supports text-to-sql generative retrieval, loads the relational database schema including table names, field names and their data types, uses the schema information to match the fields described in natural language with the relational database fields, uses the Tongyi Qianwen 2.5-32B model combined with database table information, some data samples and query content to generate corresponding SQL statements, uses an SQL parser to perform syntax checks and multiple validations on the generated SQL, sends the generated SQL statements to the target database execution engine to obtain the search results, and converts the search results into visualization formats (such as tables, charts) or JSON data;
[0085] The multi-agent dynamic auditing and risk control supervision method of this invention can be widely applied in fields such as corporate financial auditing, financial institution risk control, and supply chain compliance supervision, significantly improving audit efficiency, reducing risk costs, and supporting dynamic collaboration and real-time optimization of complex tasks.
[0086] The above are merely preferred embodiments of the present invention. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible variations and modifications to the technical solutions of the present invention using the methods and techniques disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the technical solutions of the present invention. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall still fall within the protection scope of the technical solutions of the present invention.
Claims
1. A dynamic auditing and risk control supervision method based on multiple agents, characterized in that, include: Phase of building a heterogeneous data retrieval library: Receive historical case studies, key risk elements, and regulatory documents from various heterogeneous data sources, including unstructured documents and structured data; Perform deep parsing on unstructured documents to obtain plain text data; Vectorization and embedding are performed on plain text data to generate document index data that supports efficient retrieval, which is then stored in a vector database. Structured data is used to generate data table information and field information, and database tables are dynamically created in the relational database and stored in the relational database. By combining vector databases and relational databases, a heterogeneous data source document retrieval library is realized, which includes a risk feature library and a regulatory content library. Dynamic audit and risk control supervision phase: Accepts various text review rules; Receive the documents to be audited, execute the heterogeneous data retrieval library process, and obtain unstructured risk elements and structured risk elements from the risk feature library; Risk points and risk types are classified based on unstructured and structured risk elements; Based on the risk points and risk types, the process engine is invoked to execute automated audit processes and obtain audit anomaly result data. Based on the audit anomaly results data, the regulatory content database is searched and enhanced to generate compliance review results data; Structured audit results data are obtained based on audit anomaly data and compliance review results data; The aforementioned deep parsing of unstructured documents specifically refers to: A heterogeneous data source document parsing agent is constructed, which uses an object detection model to identify various regions on the page and predicts the location of each region through bounding box regression. For the detected table regions, the table structure is parsed to further identify its row and column structure and cell distribution. Each cell is cut into an independent image and the text content is extracted through OCR technology. The classification of risk points and risk types based on unstructured and structured risk elements is as follows: We construct an intelligent risk identification agent that combines the text data of the documents to be audited with the unstructured and structured risk elements obtained from retrieval, and combines the review rules in the text review rules with the positive and negative cases therein to perform multi-dimensional cross-validation and automatic classification. We then use the Tongyi Qianwen 2.5-32B model to perform semantic enhancement analysis to identify potential risk points and classify risk types. The process of classifying risks based on risk points and risk types and then calling the process engine to execute the automated audit process is as follows: An automated auditing agent is built, which performs text review based on the review rules for sensitive words among various review rules. The Tongyi Qianwen 2.5-VL-32B model is used to extract financial data for trend assessment, determine whether there are abnormal trends, and perform data anomaly detection. Error correction review supports the review of Prompt instruction rules based on the Tongyi Qianwen 2.5-32B model. By defining and applying specific error correction rules, the text content is analyzed and text error correction suggestions are returned.
2. The multi-agent-based dynamic auditing and risk control supervision method according to claim 1, characterized in that, The process of vectorizing and embedding plain text data to generate document index data that supports efficient retrieval specifically involves: A retrieval enhancement generation system based on the Elasticsearch vector database is constructed. The retrieval enhancement generation system uses the Tongyi Qianwen 2-1.5B model to vectorize each text fragment to generate an embedding vector, and imports the generated vector and its related metadata into the Elasticsearch vector database.
3. The multi-agent-based dynamic auditing and risk control supervision method according to claim 1, characterized in that, The process of executing the heterogeneous data retrieval library is as follows: For unstructured documents, deep parsing is first performed to obtain plain text data. Then, the Tongyi Qianwen 2-1.5B model is used for vector embedding generation. Document block vector retrieval based on vector similarity is performed in the vector database to obtain unstructured retrieval results. The Tongyi Qianwen 2.5-32B model is used in conjunction with the SQL examples and table field information of the database to construct dynamic SQL queries. Self-verification and correction are performed based on the SQL query results to obtain structured retrieval results.
Citation Information
Patent Citations
Law supervision clue mining method and system based on full-text retrieval and large model
CN119149569A
Document auditing method, device and equipment based on large model and storage medium
CN119963134A