File auditing method, device, equipment and product
By comparing audit documents using preset comparison conditions and a large language model, a difference details page is generated and user questions are answered, solving the problems of omissions and inefficiency in manual comparison and achieving efficient and accurate auditing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MERCHANTS FINANCE HLDG CO LTD
- Filing Date
- 2025-12-08
- Publication Date
- 2026-04-17
AI Technical Summary
The existing technology of manually comparing audit system documents has problems such as large workload, easy omission of errors, low efficiency and lack of real-time feedback, which makes it difficult to meet the needs of efficient and accurate auditing.
The system judges the audit documents to be compared by setting up comparison conditions, uses a vector knowledge base and a large language model to compare the documents, generates a difference details page, and receives and answers user questions.
It improves the efficiency of document auditing, solves the problems of omissions, errors and inefficiency in manual comparison, provides real-time feedback, and meets the needs of efficient and accurate auditing.
Smart Images

Figure CN121882015A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information security technology, and in particular to a document auditing method, apparatus, equipment, and product. Background Technology
[0002] Audit work covers multiple aspects, including financial audit, internal control audit, and business audit, all of which require various policy documents as core supporting materials. Currently, auditors need to carefully and thoroughly read relevant audit policies, identify and integrate the differences and similarities between different policies, thereby ensuring the smooth progress of subsequent audit work. As a crucial basis for audit work, the accuracy of policy documents and the clarity of their differences directly affect the reliability of audit conclusions. Therefore, auditors must repeatedly review and compare policy documents to obtain key information.
[0003] The existing technology of manually handling system comparison has obvious drawbacks: On the one hand, the system documents are voluminous, and even if there are few differences, the high importance of the documents makes it extremely labor-intensive to manually judge and integrate the similarities and differences, and it is prone to omissions and errors. On the other hand, system documents may be updated rapidly, and auditors need to continuously invest a lot of manpower in document comparison, which not only makes them overwhelmed, but also seriously slows down the work pace, affects the overall audit progress, and further increases the possibility of work errors. In addition, when users have questions about the content of the audit documents, the system cannot provide timely answers to their questions.
[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a document auditing method, apparatus, equipment, and product, which aims to solve the technical problems of large workload, easy omission of errors, low efficiency, and lack of real-time feedback when manually comparing and auditing system documents.
[0006] To achieve the above objectives, this application proposes a document auditing method, which includes: Receive the audit document to be compared and judge it according to the preset comparison conditions; If the audited documents to be compared meet the comparison conditions, the documents are compared using a vector knowledge base and a large language model to obtain the comparison results, and a difference details page is generated based on the comparison results. Receive user questions regarding the comparison results and the difference details page, and respond to the questions.
[0007] In one embodiment, the step of judging the audit document to be compared using preset comparison conditions includes: The content eligibility of the audit document to be compared is determined by the large language model to obtain a first judgment result. If the first judgment result is passed, the number of files to be compared and audited is judged by the comparison quantity threshold. If the number of files in the audit files to be compared meets the comparison number threshold, the file memory of the audit files to be compared is judged by the memory threshold. If the audit file to be compared meets the memory threshold, a judgment result is obtained that the audit file to be compared meets the comparison conditions.
[0008] In one embodiment, before the step of comparing the audited documents using a vector knowledge base and a large language model to obtain the comparison results when the audited documents meet the comparison conditions, the method further includes: Collect historical audit documents; The historical audit files are compared using an initial large language model to obtain the version comparison results; The historical audit files are labeled based on the version comparison results to obtain training samples; The initial large language model is fine-tuned and trained using the training samples to obtain the large language model.
[0009] In one embodiment, the step of comparing documents using a vector knowledge base and a large language model to obtain comparison results, and generating a difference details page based on the comparison results, includes: The audit files to be compared are vectorized and then stored in a vector knowledge base; The audit documents to be compared are classified according to their scenarios to obtain the classification results; Based on the classification results, the audit documents to be compared are compared using the large language model to obtain the comparison results; The original text fragment is obtained by performing a similarity search on the vector knowledge base based on the comparison results. Based on the differences in the comparison results, intelligent grouping and pattern recognition are performed using a clustering algorithm to obtain a details page of the differences.
[0010] In one embodiment, the step of classifying the audit documents to be compared into scenarios and obtaining classification results includes: The redundancy identification algorithm is used to identify the audit file to be compared, obtain redundant information and unstructured data, and remove the redundant information and unstructured data from the audit file to be compared. The audit file to be compared is subjected to data noise reduction, and after the noise reduction is completed, the audit file to be compared is divided into several file sub-units based on the file logical structure and semantic content of the audit file to be compared using a text segmentation algorithm. Based on the aforementioned file sub-units, the audit files to be compared are classified according to scenarios using a machine learning algorithm to obtain classification results.
[0011] In one embodiment, before the step of receiving a user's question about the comparison results and the difference details page, and responding to the question, the method further includes: Based on the comparison results and the difference details page, an initial interactive interface is generated using an interactive visualization framework. Based on the differences in the comparison results, the initial interactive interface is rendered hierarchically to obtain the rendering result; Based on the detailed access to the database and the grouping information of the differences, the report is formatted to obtain the formatting result; A comparison report of the audit documents to be compared is generated based on the rendering and layout results.
[0012] In one embodiment, the step of receiving a user's question about the comparison results and the difference details page, and responding to the question, includes: The user's questions regarding the comparison results and the difference details page are analyzed to obtain the focus information and difference information corresponding to the user's questions; Based on the emphasis information and difference information, the first original text fragment of the audit document to be compared is retrieved. Based on the emphasis information and difference information, historical fragments are extracted through the large language model to obtain the second original text fragment. Based on the information emphasized, a detailed analysis is conducted using the first and second original text fragments to obtain a more detailed report; The detailed report is used to generate response information for the user, and the response information is used to answer the question.
[0013] Furthermore, to achieve the above objectives, this application also proposes a document auditing device, which includes: The judgment module is used to receive the audit file to be compared and judge the audit file to be compared according to preset comparison conditions; The comparison module is used to perform file comparison using a vector knowledge base and a large language model when the audit files to be compared meet the comparison conditions, obtain the comparison results, and generate a difference details page based on the comparison results; The response module is used to receive questions from users regarding the comparison results and the difference details page, and to respond to the questions.
[0014] In addition, to achieve the above objectives, this application also proposes a document auditing device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the document auditing method as described above.
[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the file auditing method described above.
[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the file auditing method described above.
[0017] One or more technical solutions proposed in this application have at least the following technical effects: This application proposes a document auditing method, apparatus, device, and product. It receives documents to be audited and judges them according to preset comparison conditions. If the documents meet the comparison conditions, a document comparison is performed using a vector knowledge base and a large language model to obtain comparison results. A difference details page is generated based on the comparison results. The method also receives and responds to user queries regarding the comparison results and the difference details page. Thus, by judging the received documents to be audited according to preset comparison conditions, performing document comparison using a vector knowledge base and a large language model when the documents meet the comparison conditions, generating a corresponding difference details page based on the comparison results, and finally receiving and responding to user queries regarding the comparison results and the details page, this method solves the problems of high workload, easy omissions of errors, low efficiency, and lack of real-time feedback in manual comparison auditing of institutional documents, thereby improving the efficiency of document auditing. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the first embodiment of the auditing method for the documents in this application. Figure 2 This is a flowchart illustrating the second embodiment of the auditing method for this application document; Figure 3 This is a simplified flowchart illustrating the auditing methodology used in this application. Figure 4 This is a second simplified flowchart illustrating the auditing method for the documents in this application. Figure 5 This is a schematic diagram of the module structure of the document auditing device according to an embodiment of this application; Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the document auditing method in this application embodiment.
[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0024] The main solution of this application embodiment is as follows: The large language model is used to determine the content suitability of the audit file to be compared, resulting in a first determination result; if the first determination result is satisfactory, the number of files in the audit file to be compared is determined using a comparison quantity threshold; if the number of files in the audit file to be compared meets the comparison quantity threshold, a memory threshold is used to determine the file memory usage of the audit file to be compared; if the audit file to be compared meets the memory threshold, a determination result is obtained that the audit file to be compared meets the comparison conditions. Historical audit files are collected; version comparisons are performed on the historical audit files using an initial large language model, resulting in version comparison results; the historical audit files are labeled based on the version comparison results to obtain training samples; the initial large language model is fine-tuned and trained using the training samples to obtain the large language model. The audit files to be compared are vectorized and stored in a vector knowledge base. The audit files are then classified into scenarios to obtain classification results. Based on the classification results, the audit files are compared using the large language model to obtain comparison results. A similarity search is performed on the vector knowledge base using the comparison results to obtain original text fragments. Based on the differences in the comparison results, intelligent grouping and pattern recognition are performed using a clustering algorithm to obtain a difference details page. A redundancy identification algorithm is used to identify redundant information and unstructured data in the audit files, and this redundant information and unstructured data are removed from the audit files. The audit files are then denoised, and after denoising, based on the file's logical structure and semantic content, the audit files are divided into several file sub-units using a text segmentation algorithm. Based on these file sub-units, a machine learning algorithm is used to classify the audit files into scenarios to obtain classification results. Based on the comparison results and the difference details page, an initial interactive interface is generated using an interactive visualization framework. The initial interactive interface is then rendered hierarchically according to the differences in the comparison results to obtain a rendering result. A report is formatted based on the details accessed and the grouping information of the differences to obtain a formatting result. A comparison report for the audited document to be compared is generated using the rendering result and the formatting result. The user's questions regarding the comparison results and the difference details page are parsed to obtain the corresponding emphasis information and difference information. Based on the emphasis information and difference information, a first original text fragment of the audited document to be compared is retrieved. Based on the emphasis information and difference information, historical fragment extraction is performed using the large language model to obtain a second original text fragment. Based on the emphasis information, a detailed analysis is performed using the first and second original text fragments to obtain a detailed report. Response information to the user is generated based on the detailed report, and the user's question is answered using the response information.This invention solves the problems of high workload, easy omissions and errors, low efficiency, and lack of real-time feedback in manual comparison and auditing of system documents. It enables comparison and response of audit documents, improving the efficiency of document auditing. Based on the present invention, considering the problems of numerous system documents, easy omissions in manual processing, and rapid updates to system documents, which consume a lot of manpower and are therefore inefficient, a document auditing method was designed. The effectiveness of the document auditing method of the present invention was verified in the comparison and response of audit documents. Finally, the efficiency of document auditing using the method of the present invention was significantly improved.
[0025] In this embodiment, for ease of description, the document auditing device will be used as the execution subject in the following description.
[0026] Due to the limitations of existing technologies that rely on manual operation for comparing audit policy documents, the efficiency, accuracy, and schedule of audit work are all insufficient. One issue is the workload of manual review: policy documents are complex and of high importance, and even with few differences, the amount of work involved in manually integrating similarities and differences is enormous, easily leading to omissions and errors. Another issue is the problem of dynamic updates and adaptation: policy documents may be updated rapidly, requiring auditors to continuously invest a large amount of manpower in comparisons, which is difficult to handle efficiently and leads to a slowdown in the work pace. Furthermore, there is the issue of ensuring audit schedule; the inefficiency of manual comparison not only slows down the overall audit progress but also further increases the possibility of errors, affecting the reliability of audit conclusions. Therefore, in current multi-domain audit work, the model of relying on manual processing of policy document comparisons can no longer meet the needs of efficient and accurate audits, becoming a key challenge restricting the smooth progress of audit work.
[0027] This application provides a solution that judges the received audit documents to be compared based on preset comparison conditions. If the audit documents meet the comparison conditions, the document comparison is performed using a vector knowledge base and a large language model to obtain the comparison results. Based on the comparison results, a corresponding difference details page is generated. Finally, the solution receives and responds to user queries regarding the comparison results and details access page. This solution solves the problems of high workload, easy omission of errors, low efficiency, and lack of real-time feedback in manual comparison of audit documents, thereby improving the efficiency of document auditing and providing users with better services.
[0028] Based on this, the embodiments of this application provide a document auditing method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the auditing method for this application.
[0029] In this embodiment, the document auditing method includes steps S01 to S03: Step S01: Receive the audit document to be compared and judge it according to the preset comparison conditions; Before describing the solution in this embodiment, it should be clear that auditing work broadly covers multiple core areas such as financial auditing, internal control auditing, and business auditing. Various institutional documents are the key supporting materials for carrying out all auditing work. To ensure the smooth progress of subsequent auditing work and to guarantee the reliability of audit conclusions, auditors must read the relevant institutional documents in detail, comprehensively sort out the similarities and differences between different systems, and complete the integration work. They also need to repeatedly review and compare institutional documents to accurately extract key information. However, the existing technology mainly uses manual methods to handle the comparison of institutional documents, which has significant drawbacks. These drawbacks are that manual comparison of systems is not only labor-intensive and prone to omissions and errors, but also difficult to adapt to the rapid updating needs of institutional documents, resulting in low audit efficiency, hindered progress, and increased risk of work errors.
[0030] Therefore, to address the aforementioned issues, this embodiment, upon receiving the audit document to be compared, judges the document according to preset comparison conditions. These preset comparison conditions include at least file format compatibility, file content integrity, and file version validity. The file format compatibility judgment ensures that the audit document conforms to the system's preset processing standards, such as supporting common audit document formats like PDF, Word, and Excel. If an incompatible format is detected, the system will automatically prompt the user to convert the format or provide a compatible file. The file content integrity judgment verifies whether the file has missing key information, incomplete page numbers, or unclear core clauses, preventing incomplete file content from affecting the accuracy of subsequent comparison analysis. The file version validity judgment primarily targets policy documents with version iterations. The system will filter out the latest version or user-specified version for comparison based on information such as creation time and revision history in the file attributes, preventing the comparison results from losing practical reference value due to the use of outdated versions.
[0031] Step S02: If the audit file to be compared meets the comparison conditions, the file is compared using a vector knowledge base and a large language model to obtain the comparison results, and a difference details page is generated based on the comparison results. If the audit documents to be compared meet the comparison criteria, the corresponding document comparison can be performed using a vector knowledge base and a large language model to obtain the comparison results. Based on the comparison results, a difference details page is generated. The difference details page will present the comparison results in a visual way, such as displaying the content differences between the original document and the target document through column comparison, and using various marking methods such as highlighting, annotation, and underlining to distinguish different types of differences, such as added content, deleted content, modified content, and format adjustments. At the same time, the page will also provide a detailed description of the differences, including the chapter, clause number, specific content, and explanation of the difference type, so that users can quickly locate and understand the differences.
[0032] Step S03: Receive questions from users regarding the comparison results and the difference details page, and respond to the questions.
[0033] After completing the file comparison and obtaining the difference details page, users may still have questions. Therefore, this embodiment also receives questions from users regarding the comparison results and the difference details page, and responds to the questions raised by users. The response process will combine the audit domain expertise stored in the vector knowledge base, the semantic understanding and reasoning ability of the large language model, and the detailed difference data generated during the comparison process to provide users with accurate and targeted answers.
[0034] Specifically, step S01 above, which involves judging the audit document to be compared using preset comparison conditions, includes: Step S011: The content eligibility of the audit document to be compared is judged by the large language model to obtain a first judgment result; Step S012: If the first judgment result is passed, the number of files to be compared and audited is judged by the comparison quantity threshold. Step S013: If the number of files in the audit files to be compared meets the comparison number threshold, the file memory of the audit files to be compared is judged by the memory threshold. Step S014: If the audit file to be compared meets the memory threshold, obtain the judgment result that the audit file to be compared meets the comparison conditions.
[0035] First, the uploaded audit files to be compared are judged for content eligibility using a large language model to obtain the first judgment result. In this embodiment, the large language model will deeply analyze the file content to identify whether the file belongs to the target scope of audit system, standards, etc. At the same time, it will detect whether the file has invalid information such as missing content, garbled characters, and duplication, to ensure that the files entering the subsequent stages have the core value of audit comparison and avoid invalid files occupying system resources.
[0036] Subsequently, if the first judgment result is passed (i.e. the file content is qualified), the file quantity verification stage will begin. The system will automatically detect the number of uploaded files and judge them according to the preset comparison quantity threshold. The threshold can be flexibly set according to the actual audit scenario (e.g., pairwise comparison is set to 2, batch comparison is set to 5-10, etc.).
[0037] If the number of files to be audited meets the comparison threshold, the system will use a memory threshold to determine the memory size of the files. The system will calculate the memory size of each file and the total memory size of the files. The memory threshold is set by taking into account the server's processing capacity and comparison efficiency to avoid system overload or excessive comparison time due to excessively large files. If the number of files is insufficient or the file content is too large during the memory threshold determination, the system will immediately trigger an intelligent prompt mechanism, clearly informing the user of the specific problem through a pop-up window (such as "The current number of uploaded files is 1, at least 2 files need to be uploaded for comparison" or "A certain file exceeds the 50MB memory limit"), and providing targeted adjustment suggestions (such as supplementing uploaded files, compressing or splitting large files) to improve the user's ease of operation.
[0038] Finally, the system will only provide a judgment result that the comparison conditions are met when the audit file to be compared simultaneously meets the three conditions of appropriate content, sufficient quantity, and compliant memory usage, allowing the file to enter the subsequent comparison stage between the vector knowledge base and the large language model.
[0039] More specifically, before step S03 above, which involves receiving a user's question about the comparison results and the difference details page, and responding to the question, the method further includes: Step S0301: Based on the comparison results and the difference details page, generate an initial interactive interface through an interactive visualization framework; Step S0302: Perform hierarchical rendering on the initial interactive interface based on the differences in the comparison results to obtain the rendering result; Step S0303: Based on the details accessed into the database and the grouping information of the differences, the report is formatted to obtain the formatting result; Step S0304: Generate a comparison report of the audit documents to be compared based on the rendering results and layout results.
[0040] Once the number of files meets the requirements and the comparison between the vector knowledge base and the large language model is completed, the system will generate comparison results covering the differences and similarities, and provide an entry point to the comparison results details page. The entry point is designed as a floating button that links to the results list, allowing users to quickly access the details and providing core data support for subsequent report generation.
[0041] Subsequently, based on the generated comparison results and the basic data of the difference details page, an initial interactive interface is generated using an interactive visualization framework (such as ECharts, D3.js, etc.). This interface has preset multi-dimensional display modules, including statistical charts of the distribution of difference points, file directory tree, content preview pane, etc., and supports users to customize the interface layout through drag, zoom and other operations to improve the flexibility of viewing results.
[0042] Users can access the details page of the discrepancies through the details entry on the initial interactive interface. They can view the document differences analyzed by the system. The page uses a left-right column comparison mode, synchronously highlighting the corresponding discrepancies and marking the type of difference (such as text modification, new clauses, or wording adjustments). This provides a clear direction for subsequent interface rendering optimization. The initial interactive interface is then rendered hierarchically based on the discrepancies in the comparison results. The rendering process employs a "hierarchical importance" logic, highlighting key discrepancies affecting audit conclusions (such as changes to core clauses of regulations) with red background and white text, dynamically flashing prompts. General wording differences are marked with yellow background and black text. Similarities are displayed in a collapsed format by default (but can be expanded by user clicks), making the hierarchy of discrepancies clear at a glance and reducing the cost of information filtering for users.
[0043] After the layered rendering is complete, users can continue to ask questions about the differences or document content displayed on the interface (such as "What is the basis for this clause change?" or "What impact does this difference have on the financial audit?"). The system will call the large language model to analyze each difference and obtain a detailed description. These questions and answers will be automatically associated with the corresponding differences as supplementary information for report layout.
[0044] Subsequently, the report was formatted according to the jump logic of the details access entry and the grouping information of the differences (such as grouping by audit areas such as "financial audit terms" and "internal control processes"). The formatting adopted a "general-specific-general" structure, first presenting an overall statistical overview of the differences, then showing the specific differences and related detailed questions and answers in turn by group, and finally adding a summary of the impact assessment of the differences to ensure that the report is logically clear and highlights the key points.
[0045] Finally, a comparison report of the audit documents to be compared is generated by using the interactive interface elements (such as highlighted difference modules and statistical charts) after hierarchical rendering and the structured layout results. In this embodiment, the report supports export in both "interactive version" and "static version" formats. The interactive version retains the zoom and jump functions of the interface for easy online viewing, while the static version uses PDF format and optimizes the layout to adapt to printing needs and meet the usage requirements of different audit scenarios.
[0046] Further, step S03 above, receiving user questions about the comparison results and the difference details page, and responding to the questions, includes: Step S031: Analyze the user's questions about the comparison results and the difference details page to obtain the focus information and difference information corresponding to the user's questions; Step S032: Based on the emphasis information and difference information, retrieve the first original text segment of the audit document to be compared; and extract historical segments through the large language model according to the emphasis information and difference information to obtain the second original text segment. Step S033: Based on the emphasis information, perform detailed analysis using the first original text fragment and the second original text fragment to obtain a detailed report; Step S034: Generate response information for the user based on the detailed report, and respond to the question using the response information.
[0047] The system analyzes user questions on the comparison results and difference details page to obtain the relevant emphasis and difference information. This embodiment uses Natural Language Processing (NLP) technology to extract key dimensions from the questions. The emphasis information includes user focus areas such as "impact analysis," "based on query," and "action suggestions." The difference information is located to the difference point number and content summary associated with the specific question. At the same time, the system will intelligently complete ambiguous questions (e.g., if a user asks "Does this difference have an impact?", it will automatically associate the question with the impact assessment direction of the corresponding difference point), improving the accuracy of question analysis.
[0048] Subsequently, based on the parsed emphasis and difference information, the first original text fragment directly related to the difference point in the audit documents to be compared is retrieved. Then, based on the emphasis and difference information, historical fragments are extracted through the large language model, that is, from the system documents and comparison reports of similar audit projects in the past, the second original text fragments that are similar or related to the current difference point are extracted (such as historical change records of the same clauses, similar difference handling cases, etc.).
[0049] After retrieving the first and second original text snippets, the system will automatically highlight these snippets on the difference details page and generate a jump link. When the user clicks on the snippet reference in the response, they can directly jump to the corresponding position on the details page to intuitively view the original text context and achieve rapid linkage verification between the response content and the original text.
[0050] Next, based on the information highlighted, a detailed analysis is conducted using the first and second original text segments to produce a detailed report. If the highlighted information is "impact assessment," the specific impact of the differences on the audit process, compliance, and risk control is analyzed. If it is "action recommendations," targeted follow-up solutions (such as supplementing verification materials and revising audit procedures) are provided based on case experience from historical segments.
[0051] Furthermore, the detailed report in this embodiment will specifically include an impact assessment module and a recommended action plan module. The impact assessment adopts the form of "risk level, specific impact description" (e.g., "high risk, this clause change may lead to inconsistencies in the accounting of expenses, affecting the accuracy of financial audits"). The recommended action plan clarifies the steps, responsible parties, and timelines, helping users to deeply understand the importance of the differences and the direction of subsequent actions.
[0052] Then, response information is generated for users based on the detailed report, and the questions are answered through the response information. In this embodiment, the response information is presented in the structure of "core conclusions + detailed analysis + excerpt references", and supports multimodal display such as text and charts (such as impact analysis flowcharts). At the same time, a "follow-up question" button is set in the response area to facilitate users to ask further questions about the detailed content.
[0053] After receiving the response information, users can view the detailed conclusions by referring to the highlighted original text snippets and determine if there are still any questions. If there are any questions, users can ask multiple questions through the "Follow-up Questions" button. The system will repeat the above process to generate new responses, thus gradually resolving the questions in depth.
[0054] To ensure the accuracy of the conclusions, users can trigger a second analysis of the current discrepancies and response content through the "Re-verify" function. The system will re-retrieve the original text fragments, verify and refine the analysis logic, and output the verification results to further ensure the reliability of the response conclusions.
[0055] Once the user confirms all conclusions and resolves all questions, they can choose to exit the response system. The system will automatically save the question and response record and link it to the corresponding comparison report for easy review and tracing later.
[0056] This embodiment, through the above-described scheme, specifically receives audit documents to be compared and judges them according to preset comparison conditions; if the audit documents meet the comparison conditions, a document comparison is performed using a vector knowledge base and a large language model to obtain comparison results, and a difference details page is generated based on the comparison results; user queries regarding the comparison results and the difference details page are received and answered. Thus, by judging the received audit documents to be compared according to preset comparison conditions, performing document comparison using a vector knowledge base and a large language model if the audit documents meet the comparison conditions, obtaining comparison results, generating corresponding difference details pages based on the comparison results, and finally receiving and answering user queries regarding the comparison results and details page, this approach solves the problems of high workload, easy omission of errors, low efficiency, and lack of real-time feedback in manual comparison of audit documents, thereby improving the efficiency of document auditing.
[0057] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Before step S02, where the audited documents meet the comparison conditions, and the document comparison is performed using a vector knowledge base and a large language model to obtain the comparison results, the document auditing method further includes steps S0201~S0204: Step S0201: Collect historical audit documents; Step S0202: The historical audit files are compared using the initial large language model to obtain the version comparison results; Step S0203: Based on the version comparison results, the historical audit files are labeled to obtain training samples; Step S0204: Fine-tune the initial large language model using the training samples to obtain the large language model.
[0058] First, historical audit documents are collected, along with historical data, including company policies, laws and regulations, and other documents. These documents are preprocessed to ensure data quality and authenticity, providing foundational data for subsequent training. Then, the historical audit documents are compared using an initial large language model to obtain version comparison results and clarify the differences and relationships between the documents.
[0059] Following the version comparison results, domain experts labeled the historical audit documents to generate training samples. The labeling work focused on accuracy to ensure the high quality of the training samples.
[0060] Finally, the initial large language model was fine-tuned using training samples. P-tuning technique was used to guide the model to generate accurate answers with specific questions, reducing storage and memory usage for each task during training, improving the model's performance on auditing tasks, and ultimately obtaining a large language model adapted to auditing scenarios.
[0061] Specifically, step S02 above, which involves comparing documents using a vector knowledge base and a large language model to obtain comparison results, and generating a difference details page based on the comparison results, includes: Step S021: Vectorize the audit file to be compared and store it in the vector knowledge base; Step S022: Classify the audit documents to be compared into scenarios to obtain classification results; Step S023: Based on the classification results, the audit document to be compared is compared using the large language model to obtain the comparison results; Step S024: Perform a similarity search on the vector knowledge base based on the comparison results to obtain the original text fragment; Step S025: Based on the differences in the comparison results, intelligent grouping and pattern recognition are performed using a clustering algorithm to obtain the difference details page.
[0062] First, the audit files to be compared are vectorized and stored in a vector knowledge base. In this embodiment, a high-quality vector representation is generated by applying a semantic embedding model based on deep learning to the file slices. An efficient vector library is built and an optimized data structure and indexing mechanism are adopted to support approximate nearest neighbor (ANN) search, laying the foundation for subsequent fast retrieval.
[0063] Subsequently, the audit documents to be compared are classified into scenarios to obtain classification results. The classification process combines the semantic understanding capability in the model calling strategy to match suitable analysis models for different audit scenarios (such as financial audit and internal control audit) in advance, thereby improving the accuracy of subsequent comparisons.
[0064] Then, based on the classification results, the audit documents to be compared are compared using a large language model to obtain the comparison results. In this embodiment, a collaborative strategy of precise invocation of general LLM and domain adaptation of customized LLM is adopted. Natural Language Understanding (NLU) technology is used to extract key information and context, and targeted prompts are automatically generated to guide the general LLM to conduct in-depth analysis. At the same time, the finely tuned customized LLM is invoked, which, with its high understanding of professional terms and legal concepts, accurately captures the key elements and subtle differences of domain documents.
[0065] Next, a similarity search is performed on the vector knowledge base based on the comparison results to obtain the original text fragments, accurately locate the original text content corresponding to the differences, and provide a basis for the interpretation of the results.
[0066] Finally, based on the differences in the comparison results, intelligent grouping and pattern recognition are performed through clustering algorithms to obtain a detail page of differences. The clustering uses the density-based DBSCAN algorithm to achieve intelligent grouping and pattern recognition. The detail page integrates in-depth comparison result output and interactive visualization technology, outputting structured content including similarity scores and detailed explanations of differences. It uses difference highlighting and hierarchical display, and supports users to click to jump to the original text. At the same time, it uses natural language generation (NLG) technology to generate a brief description of the differences, presenting the analysis results intuitively.
[0067] More specifically, step S022 above, which involves classifying the audit documents to be compared into scenarios to obtain classification results, includes: Step S0221: The audit file to be compared is identified by a redundancy identification algorithm to obtain redundant information and unstructured data, and the redundant information and unstructured data are removed from the audit file to be compared. Step S0222: Denoise the data of the audit file to be compared, and after the denoising is completed, divide the audit file to be compared into several file sub-units based on the file logical structure and semantic content of the audit file to be compared using a text segmentation algorithm. Step S0223: Based on the aforementioned file sub-units, the audit files to be compared are classified according to scenarios using a machine learning algorithm to obtain classification results.
[0068] First, the redundancy identification algorithm is used to identify the audit files to be compared. This process incorporates data cleaning strategies and develops specialized algorithms to identify and obtain redundant information and unstructured data, such as meaningless filler words and format tags. This redundant information and unstructured data are then removed from the audit files to ensure the conciseness of the file content.
[0069] Subsequently, data denoising was performed on the audit documents to be compared. Using data denoising technology and natural language processing technology, grammatical errors and inconsistent terminology in the documents were intelligently corrected to ensure the consistency and accuracy of the document content. After denoising, based on the file logical structure and semantic content of the audit documents to be compared, a structured slicing scheme was adopted. Advanced text segmentation algorithms were used to divide the audit documents to be compared into several semantically complete file sub-units, laying the foundation for subsequent in-depth analysis.
[0070] Finally, based on several file sub-units, the audit files to be compared are classified according to scenarios using machine learning algorithms. In this embodiment, the documents are automatically classified into two categories, professional fields and non-professional fields, by combining document metadata and content features, and the classification results are obtained to provide decision support for subsequent model calls.
[0071] This embodiment, through the above-described scheme, specifically involves collecting historical audit documents; performing version comparisons on the historical audit documents using an initial large language model to obtain version comparison results; labeling the historical audit documents based on the version comparison results to obtain training samples; and fine-tuning the initial large language model using the training samples to obtain the large language model. Thus, by judging the received audit documents to be compared according to preset comparison conditions, if the audit documents meet the comparison conditions, document comparison is performed using a vector knowledge base and the large language model to obtain comparison results. Based on the comparison results, a corresponding difference details page is generated. Finally, user queries regarding the comparison results and details access page are received and answered. This solves the problems of high workload, easy omission of errors, low efficiency, and lack of real-time feedback that exist in manual comparison of audit system documents, thereby improving the efficiency of document auditing.
[0072] For example, to help understand the implementation process of the document auditing method obtained by combining this embodiment with the above embodiment one, please refer to... Figure 3 as well as Figure 4 , Figure 3 as well as Figure 4 A simplified flowchart of a document auditing method is provided, specifically: As can be seen from the scheme of this application, it mainly includes two main implementation stages: The first step is to compare the audit documents, such as Figure 3 As shown, the process mainly includes two parts: preparation of the institutional comparison model and comparison of production environment documents. In the model preparation phase, the process for preparing the system comparison model is initiated first. Historical data (such as systems, laws and regulations, etc.) is input, and the comparison results of different versions of the documents are output. Experts then label the differences in the results. Training samples are constructed using the processed text combined with the labels. Based on the LLM model, fine-tuning training is completed through P-tuning. The resulting customized model is deployed to the production environment. In the document comparison phase in the production environment, the system documents in the production data are first acquired and preprocessed (cleaning, denoising, and classification). Then, it is determined whether the comparison documents are professional domain documents. If they are non-professional domain documents, a general large model is called to complete the content reading, obtain the user request context, fill in the template to generate prompt words, and then input them into the large model. If they are professional domain documents, a customized model is called to perform complex case judgments. At the same time, the comparison documents are sliced and stored in the vector library. Finally, after the large model outputs the comparison results, the corresponding original text fragments are traced in the vector library and the differences are summarized to obtain the final document comparison results.
[0073] Secondly, it involves responding to user questions, such as... Figure 4 As shown, users first enter the system and upload files to be compared. The system verifies whether the number of files meets the comparison criteria. If not, the system prompts the user to adjust (insufficient number of files or excessive content). If the criteria are met, the system generates comparison results and provides an entry to the details page. Users can enter the details page of the differences. On the details page, clicking on the differences can expand to obtain a detailed description, impact assessment, and suggested action plan. Users can also click on the differences to highlight the content and jump to the corresponding paragraph in the original text. After viewing the conclusions, if users still have questions, they can continue to ask questions to refine the issues and obtain more detailed conclusions (multiple verifications are supported during the process). If the user confirms that there are no questions, they exit the system.
[0074] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the auditing method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0075] This application also provides a document auditing device; please refer to [reference needed]. Figure 5The document auditing device includes: The judgment module is used to receive the audit file to be compared and judge the audit file to be compared according to preset comparison conditions; The comparison module is used to perform file comparison using a vector knowledge base and a large language model when the audit files to be compared meet the comparison conditions, obtain the comparison results, and generate a difference details page based on the comparison results; The response module is used to receive questions from users regarding the comparison results and the difference details page, and to respond to the questions.
[0076] The document auditing device provided in this application, employing the document auditing method described in the above embodiments, can solve the technical problems of high workload, easy omission of errors, low efficiency, and lack of real-time feedback when manually comparing and auditing system documents. Compared with the prior art, the beneficial effects of the document auditing device provided in this application are the same as those of the document auditing method provided in the above embodiments, and other technical features in the document auditing device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0077] This application provides a document auditing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the document auditing method in the first embodiment described above.
[0078] The following is for reference. Figure 6 The diagram illustrates a structural schematic of a document auditing device suitable for implementing embodiments of this application. The document auditing device in this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The document auditing device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0079] like Figure 6As shown, the document auditing device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the document auditing device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the document auditing device to communicate wirelessly or wiredly with other devices to exchange data. While the figures show document auditing devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0080] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0081] The document auditing device provided in this application, employing the document auditing method described in the above embodiments, can solve the technical problems of large workload, easy omission of errors, low efficiency, and lack of real-time feedback when manually comparing and auditing system documents. Compared with the prior art, the beneficial effects of the document auditing device provided in this application are the same as those of the document auditing method provided in the above embodiments, and other technical features of the document auditing device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0082] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0083] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0084] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the file auditing method described in the above embodiments.
[0085] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0086] The aforementioned computer-readable storage medium may be included in the document auditing device; or it may exist independently and not be assembled into the document auditing device.
[0087] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by the file auditing device, the file auditing device causes the following: to receive a file to be compared and audited; to judge the file to be compared and audited according to preset comparison conditions; if the file to be compared and audited meets the comparison conditions, to perform file comparison through a vector knowledge base and a large language model, to obtain comparison results, and to generate a difference details page based on the comparison results; to receive user questions about the comparison results and the difference details page, and to respond to the questions.
[0088] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0089] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0090] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0091] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-described document auditing method. This solves the technical problems of high workload, easy omission of errors, low efficiency, and lack of real-time feedback that exist when manually comparing and auditing system documents. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the document auditing method provided in the above embodiments, and will not be repeated here.
[0092] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the file auditing method described above.
[0093] The computer program product provided in this application can solve the technical problems of large workload, easy omission of errors, low efficiency, and lack of real-time feedback when manually comparing and auditing system documents. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the document auditing method provided in the above embodiments, and will not be repeated here.
[0094] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A document auditing method, characterized in that, The document auditing methods include: Receive the audit document to be compared and judge it according to the preset comparison conditions; If the audited documents to be compared meet the comparison conditions, the documents are compared using a vector knowledge base and a large language model to obtain the comparison results, and a difference details page is generated based on the comparison results. Receive user questions regarding the comparison results and the difference details page, and respond to the questions.
2. The document auditing method as described in claim 1, characterized in that, The step of judging the audit document to be compared using preset comparison conditions includes: The content eligibility of the audit document to be compared is determined by the large language model to obtain a first judgment result. If the first judgment result is passed, the number of files to be compared and audited is judged by the comparison quantity threshold. If the number of files in the audit files to be compared meets the comparison number threshold, the file memory of the audit files to be compared is judged by the memory threshold. If the audit file to be compared meets the memory threshold, a judgment result is obtained that the audit file to be compared meets the comparison conditions.
3. The document auditing method as described in claim 1, characterized in that, Before the step of comparing the audited documents using a vector knowledge base and a large language model to obtain the comparison results when the documents to be compared meet the comparison conditions, the method further includes: Collect historical audit documents; The historical audit files are compared using an initial large language model to obtain the version comparison results; The historical audit files are labeled based on the version comparison results to obtain training samples; The initial large language model is fine-tuned and trained using the training samples to obtain the large language model.
4. The document auditing method as described in claim 1, characterized in that, The steps of comparing documents using a vector knowledge base and a large language model, obtaining comparison results, and generating a difference details page based on the comparison results include: The audit files to be compared are vectorized and then stored in a vector knowledge base; The audit documents to be compared are classified according to their scenarios to obtain the classification results; Based on the classification results, the audit documents to be compared are compared using the large language model to obtain the comparison results; The original text fragment is obtained by performing a similarity search on the vector knowledge base based on the comparison results. Based on the differences in the comparison results, intelligent grouping and pattern recognition are performed using a clustering algorithm to obtain a details page of the differences.
5. The document auditing method as described in claim 4, characterized in that, The step of classifying the audit documents to be compared into scenarios and obtaining the classification results includes: The redundancy identification algorithm is used to identify the audit file to be compared, obtain redundant information and unstructured data, and remove the redundant information and unstructured data from the audit file to be compared. The audit file to be compared is subjected to data noise reduction, and after the noise reduction is completed, the audit file to be compared is divided into several file sub-units based on the file logical structure and semantic content of the audit file to be compared using a text segmentation algorithm. Based on the aforementioned file sub-units, the audit files to be compared are classified according to scenarios using a machine learning algorithm to obtain classification results.
6. The document auditing method as described in claim 1, characterized in that, Before the step of receiving user questions about the comparison results and the difference details page, and responding to the questions, the method further includes: Based on the comparison results and the difference details page, an initial interactive interface is generated using an interactive visualization framework. Based on the differences in the comparison results, the initial interactive interface is rendered hierarchically to obtain the rendering result; Based on the detailed access to the database and the grouping information of the differences, the report is formatted to obtain the formatting result; A comparison report of the audit documents to be compared is generated based on the rendering and layout results.
7. The document auditing method as described in claim 1, characterized in that, The step of receiving user questions regarding the comparison results and the difference details page, and responding to the questions, includes: The user's questions regarding the comparison results and the difference details page are analyzed to obtain the focus information and difference information corresponding to the user's questions; Based on the emphasis information and difference information, the first original text fragment of the audit document to be compared is retrieved. Based on the emphasis information and difference information, historical fragments are extracted through the large language model to obtain the second original text fragment. Based on the information emphasized, a detailed analysis is conducted using the first and second original text fragments to obtain a more detailed report; The detailed report is used to generate response information for the user, and the response information is used to answer the question.
8. A document auditing device, characterized in that, The document auditing device includes: The judgment module is used to receive the audit file to be compared and judge the audit file to be compared according to preset comparison conditions; The comparison module is used to perform file comparison using a vector knowledge base and a large language model when the audit files to be compared meet the comparison conditions, obtain the comparison results, and generate a difference details page based on the comparison results; The response module is used to receive questions from users regarding the comparison results and the difference details page, and to respond to the questions.
9. A document auditing device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the document auditing method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the document auditing method as described in any one of claims 1 to 7.