Document auditing method, device and equipment based on large model and storage medium

Through the document review method based on the big model, intelligent analysis and database retrieval, and fine-tuning the compliance judgment of the big model, the problems of low efficiency and high cost of traditional document review are solved, and the automated, efficient and accurate review of document compliance is realized.

CN119963134APending Publication Date: 2025-05-09SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510107746.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The traditional manual document review method is inefficient, expensive and unstable in the audit quality. Rules-based automatic auditing system is difficult to achieve efficient and accurate automatic compliance audits when processing complex and changeable document content.

Method used

Using a large-model-based document review method, through intelligent document analysis, audit database retrieval and fine-tuning of the large-model compliance judgment, an audit task database and a fine-tuning of the audit model are built to realize automated, efficient and accurate audit of document compliance.

Benefits of technology

It realizes the automation, efficiency and accuracy of document review, improves audit work efficiency and accuracy, and reduces the cost and time of manual review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963134A_ABST
    Figure CN119963134A_ABST
Patent Text Reader

Abstract

The invention discloses a document auditing method, device and equipment based on a large model and a storage medium, and relates to the technical field of deep learning, and the method comprises the following steps: obtaining document auditing data from a plurality of preset data sources, and analyzing the document auditing data; the document auditing data comprises an auditing material for representing a document auditing rule, a historical document auditing case and a standard document; constructing a document auditing database based on the analyzed document auditing data, and constructing a large model fine tuning data set based on the document auditing database; performing fine tuning on a preset document auditing large model by utilizing the large model fine tuning data set to obtain a target document auditing large model; and determining a current document auditing task, and auditing a to-be-audited document of the document auditing task by using the target document auditing large model to generate an auditing report. Automation, high efficiency and accuracy of document auditing are achieved through intelligent document analysis, auditing database retrieval and compliance judgment of a fine adjustment large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a document review method, device, equipment and storage medium based on a large model. Background Art

[0002] With the rapid development of information technology, enterprises or other research units will generate and use a large number of documents in their daily lives, such as technical reports, research reports, operating procedures, contracts, etc. These documents must comply with relevant laws, regulations and industry standards to ensure the legality and standardization of the documents. However, traditional manual review methods have problems such as low efficiency, high cost and unstable review quality. Although the rule-based automatic review system has improved the review efficiency to a certain extent, it still has limitations when dealing with complex and changeable document content, and it is difficult to achieve efficient and accurate automated compliance review. Therefore, how to improve the efficiency of document review is a technical problem to be solved in this field. Summary of the invention

[0003] In view of this, the purpose of the present invention is to provide a document review method, device, equipment and storage medium based on a large model, which can realize the automation, efficiency and accuracy of document review through intelligent document parsing, review database retrieval and fine-tuning of the compliance judgment of the large model. The specific scheme is as follows:

[0004] In a first aspect, the present application provides a document review method based on a large model, comprising:

[0005] Acquire document review materials from a number of preset data sources, and parse the document review materials according to preset document parsing rules; the document review materials include review materials for representing document review rules, historical document review cases, and standard documents;

[0006] Building a document review database based on the parsed document review data, and building a large model fine-tuning data set based on the document review database;

[0007] Using the large model fine-tuning data set to fine-tune the preset document review large model to obtain a fine-tuned target document review large model;

[0008] A current document review task is determined, and the target document review macromodel is used to review the document to be reviewed corresponding to the document review task, so as to generate a review report for the document to be reviewed.

[0009] Optionally, parsing the document review information according to a preset document parsing rule includes:

[0010] Determine the document format of the document review data, and determine the preset document parsing rules corresponding to the document review data in each document format;

[0011] The corresponding document review data is parsed according to the preset document parsing rules.

[0012] Optionally, constructing a large model fine-tuning dataset based on the document review database includes:

[0013] Based on the document review data in the document review database, a risk point identification data set, a risk level assessment data set, a rectification suggestion generation data set and an evidence chain construction data set are constructed.

[0014] Optionally, the fine-tuning of the preset document review big model by using the big model fine-tuning dataset includes:

[0015] The preset document review model is fine-tuned using the risk point identification dataset, the risk level assessment dataset, the rectification suggestion generation dataset and the evidence chain construction dataset through a low-rank adapter.

[0016] Optionally, the step of reviewing the to-be-reviewed document corresponding to the document review task by using the target document review macromodel includes:

[0017] Determine the review materials corresponding to the document review task, and review the to-be-reviewed document corresponding to the document review task based on the review materials using the target document review macromodel;

[0018] Furthermore, the process of determining the review materials corresponding to the document review task includes:

[0019] If supplementary materials for the current document review task are received, the review materials are updated based on the supplementary materials, so that the updated review materials are used to review the to-be-reviewed document corresponding to the document review task.

[0020] Optionally, the using the target document review macromodel to review the document to be reviewed corresponding to the document review task to generate a review report for the document to be reviewed includes:

[0021] Using the target document review model to review the document to be reviewed corresponding to the document review task, and obtaining risk points of the document to be reviewed;

[0022] Performing semantic analysis on the document corresponding to the risk point, and determining the risk level corresponding to the risk point based on the obtained analysis result and the review materials;

[0023] Generate rectification suggestions and audit evidence chains corresponding to the document to be audited based on the risk level of the risk point by using the target document audit big model;

[0024] The audit report of the document to be audited is generated based on the risk points, the risk levels, the rectification suggestions and the audit evidence chain.

[0025] Optionally, after the target document review big model is used to review the document to be reviewed corresponding to the document review task to generate a review report for the document to be reviewed, the method further includes:

[0026] Acquire adjustment opinions sent by the user based on the audit report, so as to generate a target report of the document to be audited based on the audit report according to the adjustment opinions;

[0027] The large model fine-tuning dataset is updated based on the audit report and the target report, so as to continue to fine-tune the target document audit large model using the updated large model fine-tuning dataset.

[0028] In a second aspect, the present application provides a document review device based on a large model, comprising:

[0029] A data parsing module, used to obtain document review data from a number of preset data sources, and parse the document review data according to preset document parsing rules; the document review data includes review materials used to characterize document review rules, historical document review cases and standard documents;

[0030] A data set construction module, used to construct a document review database based on the parsed document review data, and to construct a large model fine-tuning data set based on the document review database;

[0031] A model fine-tuning module, used to fine-tune the preset document review big model using the big model fine-tuning data set to obtain a fine-tuned target document review big model;

[0032] The document review module is used to determine the current document review task, and use the target document review model to review the document to be reviewed corresponding to the document review task, so as to generate a review report for the document to be reviewed.

[0033] In a third aspect, the present application provides an electronic device, comprising a processor and a memory; wherein the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the aforementioned large model-based document review method.

[0034] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the aforementioned large model-based document review method.

[0035] In this application, document review data is first obtained from several preset data sources, and the document review data is parsed according to the preset document parsing rules; the above-mentioned document review data includes review materials, historical document review cases and standard documents used to characterize document review rules; then a document review database is constructed based on the parsed document review data, and a large model fine-tuning data set is constructed based on the document review database, and the preset document review large model is fine-tuned using the large model fine-tuning data set to obtain the fine-tuned target document review large model, and then the current document review task can be determined, and the target document review large model can be used to review the documents to be reviewed corresponding to the document review task to generate an audit report for the documents to be reviewed. In this way, this application can realize intelligent document parsing, use the audit knowledge base retrieval and fine-tune the compliance judgment of the large model by constructing an audit task database and fine-tuning the audit large model, helping users to realize the automation, efficiency and accuracy of document compliance audit, thereby realizing the automation and intelligence of the audit process, and being able to efficiently realize document audit, greatly improving the efficiency and accuracy of audit work. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0037] Figure 1 A flowchart of a document review method based on a large model provided for this application;

[0038] Figure 2 A large model-based document review flow chart provided for this application;

[0039] Figure 3 A specific flowchart of a document review method based on a large model provided for this application;

[0040] Figure 4 A schematic diagram of the structure of a document review device based on a large model provided for this application;

[0041] Figure 5 A structural diagram of an electronic device provided for this application. DETAILED DESCRIPTION

[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0043] The traditional manual review of documents has problems such as low efficiency, high cost and unstable review quality. At the same time, although the rule-based automatic review system has improved the review efficiency to a certain extent, it still has limitations when dealing with complex and changeable document content, and it is difficult to achieve efficient and accurate automated compliance review. The present application discloses a document intelligent review method based on a large model, which is suitable for automated document review in specific industries. It can achieve intelligent document parsing, use audit knowledge base retrieval and fine-tune the compliance judgment of the large model by building an audit task database and fine-tuning the audit large model, helping users to achieve automated, efficient and accurate review of document compliance, thereby greatly improving the efficiency and accuracy of audit work.

[0044] See also Figure 1 As shown, the embodiment of the present invention discloses a document review method based on a large model, including:

[0045] Step S11, obtaining document review data from a number of preset data sources, and parsing the document review data according to preset document parsing rules; the document review data includes review materials for characterizing document review rules, historical document review cases and standard documents.

[0046] The document auditing system corresponding to the document auditing method disclosed in this embodiment includes four core modules: a data collection module, a fine-tuning large model module, an auditing module, and a user feedback module. The data collection module is used for document processing and database construction; the fine-tuning large model module constructs a fine-tuning data set based on the knowledge of the relevant database and then fine-tunes the auditing large model. Specifically, in this embodiment, the data processing module can be used to automatically or manually collect, process, and standardize the relevant information of the audit object to provide high-quality input data support for subsequent document audits. Figure 2As shown, first, the document review materials can be obtained from several preset data sources, and the document review materials can be parsed according to the preset document parsing rules; wherein the above-mentioned document review materials include the review materials used to characterize the document review rules, historical document review cases, standard documents and other document review related materials. Specifically, when collecting the audit object materials, this embodiment can use crawler technology, API (Application Programming Interface) interface calls and the knowledge enhancement capabilities of large models to automatically collect laws and regulations, industry standards and typical compliance documents related to the target field. The collected data from multiple data sources constitute the basic data source for the audit work, which can ensure the comprehensiveness and authority of the audit scope. In addition, in this embodiment, you can also choose to manually upload the historical case data accumulated in the previous audit process of the system, including but not limited to the following: Audit task list: documents describing specific task objectives and audit scope; Typical violation cases: records of past violations and their judgment basis; The final output audit report: including information such as audit conclusions, rectification suggestions and subsequent tracking measures.

[0047] After the document review materials are obtained in this embodiment, the document review materials can be parsed according to the preset document parsing rules. First, the document format of the document review materials can be determined, and the preset document parsing rules corresponding to the document review materials of each document format can be determined, so as to parse the corresponding document review materials according to the preset document parsing rules. That is to say, in this embodiment, when performing document processing, multimodal data parsing can be performed through the data processing module for documents of different formats (text, tables, pictures, etc.), which can specifically include: text parsing: segmented parsing, semantic analysis and context understanding of long text content such as legal clauses and policy descriptions; table parsing: using table recognition algorithms to extract structured information in tables and convert them into machine-readable form; image parsing: using OCR technology and visual models to extract text information or chart content (such as flow charts, statistical charts, etc.) in images, and understand their semantics in combination with the context.

[0048] Step S12: construct a document review database based on the parsed document review data, and construct a large model fine-tuning data set based on the document review database.

[0049] In this embodiment, a document audit database can be constructed based on the parsed document audit data, and a risk point identification data set, a risk level assessment data set, a rectification suggestion generation data set, and an evidence chain construction data set can be constructed based on the document audit data in the document audit database. In this embodiment, when constructing a compliance standard library applicable to document audit model training, the standard library that needs to be constructed specifically includes the following databases: Dynamically updated legal and regulatory library: summarizes and dynamically updates legal and regulatory information related to the target field, covers the latest policy changes and regulatory requirements, and ensures the timeliness and authority of the audit standards; Industry standards and policy library: includes general standards of the target industry, special requirements of subdivided fields, and common policy guidance documents in the industry, providing industry reference for audits; Typical case library: typical compliance cases and violation cases accumulated in the past audit process, after classification, labeling and sorting, form a case library to provide support for the optimization of audit strategies and continuous training of models; Audit methodology library: records theoretical methods and practical experience related to audits, including audit processes, risk assessment methods, compliance judgment standards, etc., to provide theoretical support for system decision-making and analysis.

[0050] It is understandable that when conducting document review based on a large model, the key to fine-tuning the large model is to build a high-quality, targeted fine-tuning dataset to meet the needs of different tasks. Therefore, in this embodiment, it is necessary to first build a fine-tuning dataset, where the construction of the dataset needs to be combined with actual business scenarios and target tasks, specifically including the following categories:

[0051] 1. Risk point identification data set:

[0052] Data sources: violation cases in the typical case database, legal and regulatory provisions, industry standard documents, etc.

[0053] Marking content: Mark the specific locations of potential risk points or violation points in the document, and clarify the corresponding legal and regulatory provisions or industry standard requirements.

[0054] Example:

[0055] Document excerpt: "A company failed to monitor environmental pollutants in accordance with regulations."

[0056] Label content: Risk point - "Failure to monitor in accordance with regulations"; Violation clause - Article 18 of the Environmental Protection Law.

[0057] 2. Risk level assessment data set:

[0058] Data sources: historical audit cases and industry risk assessment standards.

[0059] Note: Different violations or risk points are classified into risk levels (such as high, medium, and low) based on historical data, and the basis for the classification is attached.

[0060] Example:

[0061] Behavior: "Major falsification of financial data."

[0062] Notes: Risk level: High; Basis: “Severely affects the financial authenticity and regulatory compliance of the enterprise.”

[0063] 3. Generate data set for rectification suggestions:

[0064] Data source: historical audit reports and rectification notices.

[0065] Marking content: Mark the corresponding rectification suggestions for each risk point. The suggestions must be specific and feasible.

[0066] Example:

[0067] Risk point: "Failure to conduct employee occupational health examinations in accordance with regulations."

[0068] Note: Corrective suggestions - "Conduct occupational health examinations for all employees every year in accordance with regulations and keep examination records."

[0069] 4. Evidence chain construction data set:

[0070] Data sources: Multimodal data such as documents, images, and tables collected during the audit process.

[0071] Annotation content: For each violation point, organize the chain of evidence supporting its judgment, including relevant document fragments, charts, regulations, etc.

[0072] Example:

[0073] Risk point: "Failure to file annual tax returns."

[0074] Annotated content: Evidence chain - "Tax declaration records missing (system log screenshot) + Article 15 of the Tax Administration Regulations."

[0075] Step S13: fine-tune the preset document review big model using the big model fine-tuning data set to obtain a fine-tuned target document review big model.

[0076] Based on the above steps, in this embodiment, the fine-tuning large model module can be used to fine-tune the preset document review large model through LoRA (Low-Rank Adaptation) using the risk point identification dataset, risk level assessment dataset, rectification suggestion generation dataset and evidence chain construction dataset to obtain the fine-tuned target document review large model.

[0077] And in a specific embodiment, the fine-tuning of the large model can adopt the open source Qwen2.5-72B-Instruct model, and achieve efficient parameter optimization through LoRA (Low-Rank Adaptation) fine-tuning. At the same time, a multi-task fine-tuning strategy is adopted to carry out fine training for the four major tasks of risk point identification, risk level assessment, rectification suggestion generation and evidence chain construction. After fine-tuning, a comprehensive task performance test is carried out in the verification stage. It should be pointed out that after fine-tuning the large model, the model output can be compared and analyzed with the manual review results to verify its accuracy and coverage.

[0078] Step S14: determine the current document review task, and use the target document review model to review the document to be reviewed corresponding to the document review task to generate a review report for the document to be reviewed.

[0079] In this embodiment, the current document review task is first determined, and then the review materials corresponding to the document review task are determined, and the target document review model is used to review the documents to be reviewed corresponding to the document review task based on the review materials to generate a review report for the documents to be reviewed. In addition, in the process of determining the review materials, if the supplementary materials of the current document review task are received, the review materials are updated based on the supplementary materials so that the documents to be reviewed corresponding to the document review task can be reviewed using the updated review materials. It can be understood that in this embodiment, multiple review tasks will be preset in advance when conducting document review, which contain all information related to the review task, including task description, review scope and objectives, and automatically associate multiple review documents related to this task. After that, when reviewing the task, the system will match the laws, regulations, industry standards and other contents related to the review task, and the above contents and matching process can be preset and realized by the document review system disclosed in this embodiment. And it can be understood that if new regulations or policy changes are encountered during the current document review, the user can also update and expand the relevant standards through manual intervention to ensure that the system always maintains support for the latest compliance requirements. In another specific embodiment, newly added relevant materials may also be obtained for review from a predetermined data source for obtaining document review-related materials through a preset interface.

[0080] In this embodiment, document review data is first obtained from several preset data sources, and the document review data is parsed according to preset document parsing rules; the above-mentioned document review data includes review materials, historical document review cases and standard documents used to characterize document review rules; then, a document review database is constructed based on the parsed document review data, and a large model fine-tuning data set is constructed based on the document review database, and the preset document review large model is fine-tuned using the large model fine-tuning data set to obtain a fine-tuned target document review large model, and then the current document review task can be determined, and the target document review large model can be used to review the documents to be reviewed corresponding to the document review task to generate a review report for the documents to be reviewed. In this way, this embodiment first collects relevant information of the audit object through multiple channels through the document audit system, and performs multimodal data analysis on the collected documents, and then builds a legal and regulatory database, an industry standard and policy database, a typical case database and an audit methodology database, and builds a large model fine-tuning data set based on the above database, including a risk point identification data set, a risk level assessment data set, a rectification suggestion generation data set and an evidence chain construction data set, and then optimizes the audit large model through fine-tuning technology (such as LoRA) to adapt it to the professional needs of the audit task, so that the document can be audited using the fine-tuned model. Through the above technical solution, this embodiment can realize intelligent document parsing, use the audit knowledge base retrieval and fine-tune the compliance judgment of the large model by building an audit task database and fine-tuning the audit large model, helping users to realize the automation, efficiency and accuracy of document compliance audit, thereby realizing the automation and intelligence of the audit process, and being able to efficiently realize document audit, greatly improving the efficiency and accuracy of audit work.

[0081] Based on the previous embodiment, it can be seen that the present application can realize intelligent document review by constructing a review task database and fine-tuning the review model. Next, the present embodiment will elaborate on the process of generating a review report after document review. Figure 3 As shown, the embodiment of the present application discloses a specific document review method based on a large model, including:

[0082] Step S21: determine the current document review task, and use the target document review model to review the document to be reviewed corresponding to the document review task to generate a review report for the document to be reviewed.

[0083] In this embodiment, when the target document audit big model is used to audit the document to be audited, the target document audit big model can be used to audit the document to be audited corresponding to the document audit task to obtain the risk points of the document to be audited, and then the document corresponding to the risk point is semantically parsed, and the risk level corresponding to the risk point is determined based on the obtained parsing result and the audit materials, and at the same time, the target document audit big model is used to generate rectification suggestions and audit evidence chains corresponding to the document to be audited based on the risk level of the risk point, thereby generating an audit report for the document to be audited based on the risk points, risk levels, rectification suggestions and audit evidence chains.

[0084] Specifically, in this embodiment, the audit module can be used to complete risk point identification, risk level assessment, rectification suggestion generation, evidence chain construction and audit report generation through the audit big model. That is to say, after the document audit big model starts the current document audit task, the fine-tuned big model can automatically scan the audit documents related to this task, so as to intelligently identify potential risk points or possible violation points in combination with typical case libraries, laws and regulations, industry standards, etc., and compare and analyze the risk points with relevant legal provisions or industry standard contents through segment-by-segment semantic parsing and regulation matching technology and provide a basis. And for each identified risk point in this embodiment, the system can use a multi-dimensional evaluation model to intelligently evaluate the risk level (high, medium, low) from the aspects of impact scope, severity, rectification difficulty, etc., and adjust the evaluation standards in time according to industry dynamics or policy changes to ensure the adaptability of the audit results. On this basis, the system generates targeted and operational rectification suggestions for each risk point, and provides multiple options for users to choose. At the same time, it automatically sorts out a complete chain of evidence, combines document text, tables, pictures and regulatory provisions, and forms a structured chain of evidence display, which is convenient for auditors to verify. Finally, the system automatically generates a structured audit report, including a list of risk points, risk level assessment results, rectification suggestions, evidence chain and quantitative assessment results. It can be understood that the audit report format in this embodiment can be customized according to user needs to fully support document review work.

[0085] Step S22: obtain adjustment suggestions sent by the user based on the audit report, generate a target report of the document to be audited based on the audit report according to the adjustment suggestions, and update the large model fine-tuning dataset based on the audit report and the target report, so as to continue to use the updated large model fine-tuning dataset to fine-tune the target document audit large model.

[0086] In this embodiment, the adjustment opinions sent by the user based on the audit report can be obtained to generate a target report of the document to be audited based on the audit report according to the adjustment opinions, and then the large model fine-tuning data set is updated based on the audit report and the target report, so as to continue to use the updated large model fine-tuning data set to fine-tune the target document audit large model. That is to say, in this embodiment, the user review and feedback results can be incorporated into the system through the user feedback module, and the fine-tuning data set and model performance can be continuously optimized to achieve dynamic iteration of the system. In this way, the audit report generated in this embodiment supports users to review and feedback, and users can verify and correct the identified risk points, risk levels, rectification suggestions and evidence chains to ensure the accuracy and applicability of the final results. Specifically, during the review process of user review, the user's modifications and opinions will be automatically recorded and stored as important feedback data of the system to form new annotation samples, thereby further improving the fine-tuning data set. After that, the system regularly integrates the user review results into the continuous fine-tuning and optimization process of the large model, and continuously improves the accuracy and adaptability of the model in tasks such as risk identification, assessment and suggestion generation, so as to achieve dynamic iteration and intelligent upgrading of the audit process.

[0087] Through the above technical solution, this embodiment can use the fine-tuned model to perform task audits on documents, including risk point identification, risk level assessment, rectification suggestion generation, evidence chain construction and audit report generation, and then incorporate user review and feedback results into the system, continuously optimize the fine-tuning data set and model performance, and realize dynamic iteration of the system. Combined with the previous embodiment, by collecting audit object information, document processing, building a database, building a fine-tuning data set, fine-tuning a large model, task review, and the organic combination of learning and feedback modules, the audit process is automated, intelligent, and dynamically optimized, which can efficiently identify risk points, assess risk levels, generate rectification suggestions, and build a complete evidence chain, greatly improving the efficiency and accuracy of audit work, while continuously optimizing system performance through user feedback.

[0088] See also Figure 4 As shown, the embodiment of the present application also discloses a document review device based on a large model, including:

[0089] The data analysis module 11 is used to obtain document review data from a number of preset data sources and analyze the document review data according to preset document analysis rules; the document review data includes review materials for representing document review rules, historical document review cases and standard documents;

[0090] A data set construction module 12, used to construct a document review database based on the parsed document review data, and to construct a large model fine-tuning data set based on the document review database;

[0091] A model fine-tuning module 13 is used to fine-tune the preset document review big model using the big model fine-tuning data set to obtain a fine-tuned target document review big model;

[0092] The document review module 14 is used to determine the current document review task, and use the target document review macromodel to review the document to be reviewed corresponding to the document review task, so as to generate a review report for the document to be reviewed.

[0093] This embodiment can obtain document review data from several preset data sources, and parse the document review data according to preset document parsing rules; the above-mentioned document review data includes review materials, historical document review cases and standard documents used to characterize document review rules; then, a document review database is constructed based on the parsed document review data, and a large model fine-tuning data set is constructed based on the document review database, and the preset document review large model is fine-tuned using the large model fine-tuning data set to obtain the fine-tuned target document review large model, and then the current document review task can be determined, and the target document review large model can be used to review the documents to be reviewed corresponding to the document review task to generate an audit report for the documents to be reviewed. Through the above-mentioned technical solution, intelligent document parsing, audit knowledge base retrieval and compliance judgment of the fine-tuning large model can be realized by constructing an audit task database and fine-tuning the audit large model, helping users to realize automated, efficient and accurate audit of document compliance, thereby realizing the automation and intelligence of the audit process, and being able to efficiently realize document review, greatly improving the efficiency and accuracy of audit work.

[0094] In some specific embodiments, the data parsing module 11 specifically includes:

[0095] A rule determination unit, used to determine the document format of the document review data, and determine the preset document parsing rules corresponding to the document review data in each document format;

[0096] The document parsing unit is used to parse the corresponding document review data according to the preset document parsing rules.

[0097] In some specific embodiments, the data set construction module 12 specifically includes:

[0098] A data set construction unit is used to construct a risk point identification data set, a risk level assessment data set, a rectification suggestion generation data set and an evidence chain construction data set based on the document review data in the document review database.

[0099] In some specific embodiments, the model fine-tuning module 13 specifically includes:

[0100] A model fine-tuning unit is used to fine-tune the preset document review model by using the risk point identification dataset, the risk level assessment dataset, the rectification suggestion generation dataset and the evidence chain construction dataset respectively through a low-rank adapter.

[0101] In some specific embodiments, the document review module 14 specifically includes:

[0102] A document review submodule, used to determine the review materials corresponding to the document review task, and review the to-be-reviewed document corresponding to the document review task based on the review materials using the target document review macromodel;

[0103] Furthermore, the document review submodule specifically includes:

[0104] The material updating unit is configured to update the review material based on the supplementary material if supplementary material of the current document review task is received, so as to review the to-be-reviewed document corresponding to the document review task using the updated review material.

[0105] In some specific embodiments, the document review module 14 specifically includes:

[0106] A risk point determination unit, configured to use the target document review model to review the document to be reviewed corresponding to the document review task, and obtain risk points of the document to be reviewed;

[0107] A level determination unit, configured to perform semantic analysis on the document corresponding to the risk point, and determine the risk level corresponding to the risk point based on the obtained analysis result and the review material;

[0108] A suggestion generating unit, configured to generate a rectification suggestion and an audit evidence chain corresponding to the document to be audited based on the risk level of the risk point by using the target document audit macro model;

[0109] A report generating unit is used to generate the audit report of the document to be audited based on the risk point, the risk level, the rectification suggestion and the audit evidence chain.

[0110] In some specific embodiments, the document review device based on the big model further includes:

[0111] An opinion acquisition module, used for acquiring adjustment opinions sent by the user based on the audit report, so as to generate a target report of the document to be audited based on the audit report according to the adjustment opinions;

[0112] A data set updating module is used to update the large model fine-tuning data set based on the audit report and the target report, so as to continue to use the updated large model fine-tuning data set to fine-tune the target document audit large model.

[0113] Furthermore, the present application also discloses an electronic device. Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.

[0114] Figure 5 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the document review method based on a large model disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0115] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0116] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0117] The operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including computer programs that can be used to complete the large model-based document review method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.

[0118] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the document review method based on a large model disclosed above is implemented. For the specific steps of the method, reference may be made to the corresponding contents disclosed in the aforementioned embodiments, and no further description will be given here.

[0119] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0120] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0121] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0122] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0123] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A document review method based on a large model, characterized in that: include: Acquire document review materials from a number of preset data sources, and parse the document review materials according to preset document parsing rules; the document review materials include review materials for representing document review rules, historical document review cases, and standard documents; Building a document review database based on the parsed document review data, and building a large model fine-tuning data set based on the document review database; Using the large model fine-tuning data set to fine-tune the preset document review large model to obtain a fine-tuned target document review large model; A current document review task is determined, and the target document review macromodel is used to review the document to be reviewed corresponding to the document review task, so as to generate a review report for the document to be reviewed.

2. The document review method based on a large model according to claim 1 is characterized in that: The parsing of the document review data according to the preset document parsing rules includes: Determine the document format of the document review data, and determine the preset document parsing rules corresponding to the document review data in each document format; The corresponding document review data is parsed according to the preset document parsing rules.

3. The document review method based on a large model according to claim 1 is characterized in that: The step of constructing a large model fine-tuning dataset based on the document review database includes: Based on the document review data in the document review database, a risk point identification data set, a risk level assessment data set, a rectification suggestion generation data set and an evidence chain construction data set are constructed.

4. The document review method based on a large model according to claim 3 is characterized in that: The method of fine-tuning the preset document review model using the large model fine-tuning dataset includes: The preset document review model is fine-tuned using the risk point identification dataset, the risk level assessment dataset, the rectification suggestion generation dataset and the evidence chain construction dataset through a low-rank adapter.

5. The document review method based on a large model according to claim 1 is characterized in that: The step of reviewing the to-be-reviewed document corresponding to the document review task by using the target document review macromodel includes: Determine the review materials corresponding to the document review task, and review the to-be-reviewed document corresponding to the document review task based on the review materials using the target document review macromodel; Furthermore, the process of determining the review materials corresponding to the document review task includes: If supplementary materials for the current document review task are received, the review materials are updated based on the supplementary materials, so that the updated review materials are used to review the to-be-reviewed document corresponding to the document review task.

6. The document review method based on a large model according to claim 1 is characterized in that: The step of using the target document review model to review the document to be reviewed corresponding to the document review task to generate a review report for the document to be reviewed includes: Using the target document review model to review the document to be reviewed corresponding to the document review task, and obtaining risk points of the document to be reviewed; Performing semantic analysis on the document corresponding to the risk point, and determining the risk level corresponding to the risk point based on the obtained analysis result and the review materials; Generate rectification suggestions and audit evidence chains corresponding to the document to be audited based on the risk level of the risk point by using the target document audit big model; The audit report of the document to be audited is generated based on the risk points, the risk levels, the rectification suggestions and the audit evidence chain.

7. The document review method based on a large model according to any one of claims 1 to 6, characterized in that: After the target document review big model is used to review the document to be reviewed corresponding to the document review task to generate a review report for the document to be reviewed, the method further includes: Acquire adjustment opinions sent by the user based on the audit report, so as to generate a target report of the document to be audited based on the audit report according to the adjustment opinions; The large model fine-tuning dataset is updated based on the audit report and the target report, so as to continue to fine-tune the target document audit large model using the updated large model fine-tuning dataset.

8. A document review device based on a large model, characterized in that: include: A data parsing module, used to obtain document review data from a number of preset data sources, and parse the document review data according to preset document parsing rules; the document review data includes review materials used to characterize document review rules, historical document review cases and standard documents; A data set construction module, used to construct a document review database based on the parsed document review data, and to construct a large model fine-tuning data set based on the document review database; A model fine-tuning module, used to fine-tune the preset document review big model using the big model fine-tuning data set to obtain a fine-tuned target document review big model; The document review module is used to determine the current document review task, and use the target document review model to review the document to be reviewed corresponding to the document review task, so as to generate a review report for the document to be reviewed.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory; wherein the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the large model-based document review method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the large model-based document review method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Dynamic audit and risk control supervision method based on multiple Agents

    CN120851588A

  • A dynamic auditing and risk control supervising method based on multi-agent

    CN120851588B

  • Functional security auxiliary authentication method and system based on large language model

    CN121234904A

  • Intelligent patrol auxiliary method and system based on large language model fine tuning

    CN121436166A

  • Financial task auditing method and device based on large model

    CN121685183A