A large model-based intelligent document auditing method

CN122596028APending Publication Date: 2026-08-18BEIJING DEWANG TIMES INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610825160.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0003]现有方案多将文档内容直接输入规则引擎或大模型进行判断,文档审核对象、审核属性和审核规则条件之间缺少稳定的结构化关联,导致同一文档内容在不同审核规则条件下容易出现判断混淆

Benefits of technology

[0059] (1) This invention generates a review semantic triplet through a large model, constructs a triadic rule evidence background, and forms a stable association between the document review object, review attribute, review rule condition and evidence location, thereby reducing confusion in judgment under different review rules and improving the accuracy of review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122596028A_ABST
    Figure CN122596028A_ABST
Patent Text Reader

Abstract

The application discloses an intelligent document auditing method based on a large model, and relates to the technical field of intelligent document processing, and comprises the following steps: acquiring a to-be-audited document, document type information and an auditing rule library, and analyzing to generate a document auditing object set; calling a large model to map auditing attributes, auditing rule conditions, evidence positions and context relationships, and generating an auditing semantic triple set; constructing a triple rule evidence background and writing into an evidence anchor item; constructing a triple closure concept lattice model and configuring three types of closure structures; generating a rule closure output; generating an evidence closure output; triggering a conflict risk item and outputting an auditing report. The application can improve the document auditing accuracy, evidence tracing capability and cross-region risk identification capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent document processing technology, and in particular to an intelligent document review method based on a large model. Background Technology

[0002] In the field of intelligent document processing technology, existing intelligent document review technologies typically first perform layout parsing, text extraction, and field recognition on electronic documents, and then combine keyword matching, rule engines, or large-scale model question-and-answer methods to determine whether the document has missing content, formatting inconsistencies, abnormal clauses, and risky statements. Such solutions can improve the efficiency of manual review and have a certain ability to recognize explicit content such as titles, paragraphs, table fields, and signature areas.

[0003] Existing solutions often directly input document content into rule engines or large models for judgment. This lack of stable, structured relationships between document review objects, review attributes, and review rule conditions leads to confusion in judgments of the same document content under different review rule conditions. While some solutions can provide risk warnings, the lack of effective binding between evidence location, contextual relationships, and review basis makes it difficult to trace review results back to specific clauses, table fields, or signature locations.

[0004] Existing rule-based review methods typically rely on matching individual rules one by one. This makes it difficult to deduce the review attributes a document should meet based on the review rule conditions, and it also makes it difficult to verify the consistency of multiple evidence locations under the same review attribute. When there are content conflicts across paragraphs, tables, attachment description areas, or signature areas in a document, existing technologies often can only identify local field anomalies and cannot match mutually contradictory review attribute combinations, leading to missed reviews, incorrect reviews, and insufficient review report verifiability.

[0005] Therefore, how to provide an intelligent document review method based on a large model is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose an intelligent document review method based on a large model. This invention forms review semantic triples through semantic mapping of the large model, and combines the evidence background of the triple rule with three types of closure processing to complete the identification of document review objects, evidence merging, missing verification and conflict risk triggering. It has the advantages of high review accuracy, clear evidence traceability, comprehensive risk identification and high review efficiency.

[0007] An intelligent document review method based on a large model according to an embodiment of the present invention includes the following steps:

[0008] Retrieve the documents to be reviewed, document type information, and review rule base; perform layout and structure parsing on the documents to be reviewed; and generate a set of document review objects.

[0009] The large model is invoked to process the document review object set and the review rule base, mapping the document review object to the corresponding review attributes, review rule conditions, evidence location and context relationship, and generating a set of review semantic triples;

[0010] Construct the ternary rule evidence background based on the set of audit semantic triples, write the document audit object, audit attribute and audit rule condition into the ternary relation, and write the evidence location and context relationship into the evidence anchor item of the ternary relation;

[0011] Based on the ternary rule evidence background, a ternary closure concept lattice model is constructed, and rule closure operators, evidence closure operators, and proof by contradiction closure structures are configured.

[0012] The rule closure operator is invoked to process the audit rule conditions and audit attributes, deduce the audit attributes that the document audit object should satisfy under the audit rule conditions, compare the audit attributes that should be satisfied with the audit attributes in the audit semantic triple set, and generate the rule closure output;

[0013] The evidence closure operator is invoked to process the rule closure output, the audit semantic triple set, the evidence position and context relationship, and the evidence positions are merged according to the same audit attribute. Consistency check is performed on the merged evidence positions to generate the evidence closure output.

[0014] The system calls upon the output of the rule closure and evidence closure of the counter-evidence closure structure, matches mutually conflicting audit attribute combinations and triggers conflict risk items, and outputs an audit report.

[0015] Optionally, the generation of the document review object set includes the following steps:

[0016] Obtain the document to be reviewed, document type information, and review rule base; read the page content, page number position, and layout coordinates of the document to be reviewed.

[0017] Perform layout parsing on the page content, marking heading levels, paragraph boundaries, clause numbers, table areas, attachment description areas, and signature areas;

[0018] Perform structural analysis on the layout analysis results to establish hierarchical, sequential, and referential relationships between headings, paragraphs, clauses, table fields, attachment description areas, and signature areas;

[0019] Segment the documents to be reviewed according to hierarchical, sequential, and referential relationships, and write the evidence location for the reviewed documents after segmentation;

[0020] Based on document type information and the review rule base, document review objects are filtered to generate a set of document review objects.

[0021] Optionally, the generation of the audit semantic triple set includes the following steps:

[0022] Read the document review objects and evidence locations in the document review object set, extract the text content and object type of the document review objects, and determine the context scope of the document review objects based on hierarchical, sequential, and referential relationships;

[0023] Retrieve the audit rule item that matches the document audit object from the audit rule base based on the document type information;

[0024] Input the document review object, text content, object type, evidence location, context scope, and review rule items into the large model, and the large model will map the document review object into review attributes, review rule conditions, and context relationships;

[0025] The document review object, review attribute, and review rule conditions are used to form a review semantic triple. The evidence location and context relationship are bound to the review semantic triple to generate a set of review semantic triples.

[0026] Optionally, the construction of the ternary rule evidence background includes the following steps:

[0027] Read the set of semantic triples for review to obtain the document review object, review attributes, review rule conditions, evidence location, and contextual relationships;

[0028] Establish a ternary relationship based on the document review object, review attributes, and review rule conditions;

[0029] Write the evidence location into the evidence anchor item corresponding to the document review object in the ternary relation, and write the context relation into the evidence anchor items corresponding to the review attributes and review rule conditions in the ternary relation;

[0030] Merge duplicate ternary relationships according to the document review object, review attribute, and review rule conditions, merge the evidence positions and contextual relationships corresponding to the duplicate ternary relationships, and generate ternary rule evidence background.

[0031] Optionally, the construction of the ternary closure conceptual lattice model includes the following steps:

[0032] Read the ternary relationship and corresponding evidence anchoring items from the ternary rule evidence background, establish concept nodes according to the correspondence between document review object, review attribute and review rule condition, and write document review objects with the same review attribute and review rule condition into the corresponding concept nodes;

[0033] Establish concept grid connection relationships based on the inclusion relationships among document review objects, review attributes, and review rule conditions in the concept nodes;

[0034] Configure rule closure operators based on the audit rule conditions and audit attributes in the audit rule base, configure evidence closure operators based on the evidence position and context relationship in the evidence anchoring item, configure rebuttal closure structures based on the conflicting relationships between audit attributes, and connect rule closure operators, evidence closure operators, and rebuttal closure structures to the concept lattice connection relationship.

[0035] Concept nodes, concept lattice connections, rule closure operators, evidence closure operators, and proof-of-contradiction closure structures are incorporated into the ternary closure concept lattice model.

[0036] Optionally, the rule closure operator, evidence closure operator, and disproving closure structure are called in a progressive order. The rule closure operator processes the audit rule conditions and audit attributes to generate the rule closure output. The evidence closure operator reads the rule closure output and the evidence position and context relationship in the evidence anchoring item to generate the evidence closure output. The disproving closure structure reads the rule closure output and the evidence closure output, matches mutually conflicting audit attribute combinations, and triggers conflict risk items.

[0037] Optionally, the generation of the rule closure output includes the following steps:

[0038] Using the document review object and review rule conditions as indexes, extract the corresponding review attributes from the background of the ternary rule evidence and mark them as review attributes that have appeared.

[0039] Match the audit attributes that should be satisfied according to the audit rule conditions in the audit rule base, and call the rule closure operator to process the audit attributes that should be satisfied and the audit attributes that have already appeared;

[0040] Compare the required audit attributes with the existing audit attributes one by one, and mark the audit attributes that are not covered by the existing audit attributes.

[0041] Write the document review object, review rule conditions, existing review attributes, and review attributes not covered by existing review attributes into the rule closure output.

[0042] Optionally, the generation of the evidence closure output includes the following steps:

[0043] Using the document review object and review rule conditions in the rule closure output as indexes, extract the corresponding review attributes, evidence locations and contextual relationships from the review semantic triple set;

[0044] Evidence locations are grouped according to the same review attribute to form a grouped evidence location;

[0045] Based on the context, perform consistency checks on the audit attributes corresponding to the merged evidence locations and generate consistency check results.

[0046] Write the audit attributes, merged evidence locations, contextual relationships, and consistency verification results into the evidence closure output.

[0047] Optionally, the output of the audit report includes the following steps:

[0048] The code calls the evidence-of-contrast closure structure to receive the rule closure output and the evidence closure output. Based on the document review object, review rule conditions and review attributes, it establishes a correspondence between the review attributes in the rule closure output that are not covered by the already appearing review attributes, the consistency verification results in the evidence closure output and the merged evidence positions.

[0049] Under the same document review object and the same review rule, the review attributes after establishing the corresponding relationship are input into the counter-proof closure structure to match mutually contradictory review attribute combinations;

[0050] The audit attributes that match mutually conflicting audit attribute combinations, the corresponding consistency check results, and the merged evidence locations are marked as conflict risk items;

[0051] The rule closure output, evidence closure output, and conflict risk items are sorted according to the evidence location. The output includes audit attributes not covered by existing audit attributes, consistency verification results, and conflict risk items.

[0052] Optionally, matching the conflicting audit attribute combinations includes the following steps:

[0053] Extract mutually exclusive audit attribute pairs, dependent audit attribute pairs, and sequential constraint audit attribute pairs under the same audit rule condition from the audit rule base. Dependent audit attribute pairs are recorded as directed relationships from preceding audit attributes to subsequent audit attributes.

[0054] The audit attributes that enter the rebuttal closure structure are grouped according to the document audit object, audit rule conditions, and object type, and the audit attributes with evidence positions are arranged according to the evidence position.

[0055] Perform pairwise matching on the grouped audit attributes, matching each audit attribute pair with mutually exclusive audit attribute pairs, and marking mutually conflicting audit attribute combinations when a match is found.

[0056] Dependency validation is performed on the grouped audit attributes. The pre-audit attributes and post-audit attributes in the dependent audit attribute pairs are matched with the rule closure output and the evidence closure output, respectively. When the pre-audit attribute exists and the post-audit attribute is not covered by the existing audit attribute, the mutually conflicting audit attribute combinations are marked.

[0057] Perform order validation on the grouped audit attributes, compare the order of appearance of the audit attribute pairs in the merged evidence positions, and mark the conflicting audit attribute combinations when the order of appearance is inconsistent with the order constraints in the audit rule base.

[0058] The beneficial effects of this invention are:

[0059] (1) This invention generates a review semantic triplet through a large model, constructs a triadic rule evidence background, and forms a stable association between the document review object, review attribute, review rule condition and evidence location, thereby reducing confusion in judgment under different review rules and improving the accuracy of review.

[0060] (2) This invention identifies audit attributes that should be met but have not appeared by using rule closure operators and evidence closure operators, merges the evidence positions corresponding to the same audit attribute, completes consistency verification, and reduces the probability of missed audits and wrong audits.

[0061] (3) This invention uses the closure structure of the counter-evidence to match mutually contradictory audit attribute combinations, triggering conflict risk items, and can discover content conflicts across paragraphs, tables, attachments and signature areas, thereby improving the verifiability of the audit report. Attached Figure Description

[0062] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0063] Figure 1 This is an overall flowchart of an intelligent document review method based on a large model proposed in this invention;

[0064] Figure 2 This is a schematic diagram of the ternary closure concept lattice model in this invention;

[0065] Figure 3 This is a flowchart illustrating the progressive review process for the three types of closures in this invention. Detailed Implementation

[0066] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0067] refer to Figures 1-3 A smart document review method based on a large model includes the following steps:

[0068] Retrieve the documents to be reviewed, document type information, and review rule base; perform layout and structure parsing on the documents to be reviewed; and generate a set of document review objects.

[0069] The large model is invoked to process the document review object set and the review rule base, mapping the document review object to the corresponding review attributes, review rule conditions, evidence location and context relationship, and generating a set of review semantic triples;

[0070] Construct the ternary rule evidence background based on the set of audit semantic triples, write the document audit object, audit attribute and audit rule condition into the ternary relation, and write the evidence location and context relationship into the evidence anchor item of the ternary relation;

[0071] Based on the ternary rule evidence background, a ternary closure concept lattice model is constructed, and rule closure operators, evidence closure operators, and proof by contradiction closure structures are configured.

[0072] The rule closure operator is invoked to process the audit rule conditions and audit attributes, deduce the audit attributes that the document audit object should satisfy under the audit rule conditions, compare the audit attributes that should be satisfied with the audit attributes in the audit semantic triple set, and generate the rule closure output;

[0073] The evidence closure operator is invoked to process the rule closure output, the audit semantic triple set, the evidence position and context relationship, and the evidence positions are merged according to the same audit attribute. Consistency check is performed on the merged evidence positions to generate the evidence closure output.

[0074] The system calls upon the output of the rule closure and evidence closure of the counter-evidence closure structure, matches mutually conflicting audit attribute combinations and triggers conflict risk items, and outputs an audit report.

[0075] In this embodiment, the generation of the document review object set includes the following steps:

[0076] Obtain the document to be reviewed, document type information, and review rule base; read the page content, page number position, and layout coordinates of the document to be reviewed.

[0077] Perform layout parsing on the page content, marking heading levels, paragraph boundaries, clause numbers, table areas, attachment description areas, and signature areas;

[0078] When analyzing layouts, the system identifies heading levels based on font size changes, bolding status, centering position, and numbering format; paragraph boundaries based on line break position, indentation changes, and line spacing changes; clause numbers based on consecutive numbering, Chinese numeral numbering, and Arabic numeral numbering; table areas based on borders, cell boundaries, and field alignment; attachment description areas based on attachment names, attachment numbers, and attachment reference text; and signature areas based on seal images, signature images, and signed text.

[0079] Perform structural analysis on the layout analysis results to establish hierarchical, sequential, and referential relationships between headings, paragraphs, clauses, table fields, attachment description areas, and signature areas;

[0080] During structural analysis, the heading level and paragraph affiliation are written into the hierarchy relationship, the page number position, reading order, and clause number progression relationship are written into the order relationship, and the clause number references, attachment number references, table number references, and table field name references appearing in the text are written into the reference relationship.

[0081] Segment the documents to be reviewed according to hierarchical, sequential, and referential relationships, and write the evidence location for the reviewed documents after segmentation;

[0082] When segmenting documents to be reviewed, the boundaries of the document review objects are defined by the title node, clause number, table field, attachment description area, and signature area; the reference relationships are preserved for text paragraphs, table fields, and attachment description areas that have reference relationships; and an evidence location is written for each segmented document review object, which records the page number position, page coordinates, paragraph boundary, clause number, table cell coordinates, attachment description area position, and signature area position.

[0083] Based on document type information and the review rule base, document review objects are filtered to generate a set of document review objects.

[0084] When filtering document review objects, the corresponding review rule items are read according to the document type information. The title name, clause number, table field name, attachment number, and signature area name are matched with the review rule items. Document review objects that match the review rule items are retained, and decorative content, headers and footers, duplicate page numbers, and blank areas that are not related to the review rule items are deleted, generating a set of document review objects.

[0085] In this embodiment, the generation of the audit semantic triple set includes the following steps:

[0086] Read the document review objects and evidence locations in the document review object set, extract the text content and object type of the document review objects, and determine the context scope of the document review objects based on hierarchical, sequential, and referential relationships;

[0087] When generating the semantic triplet set for review, the text content and evidence location corresponding to the document review object are read from the document review object set. The object type is determined based on heading level, paragraph boundaries, clause number, table area, attachment description area, and signature area. Object types include headings, paragraphs, clauses, table fields, attachment descriptions, and signatures. The context scope is determined according to hierarchical, sequential, and referential relationships. The parent heading, adjacent paragraphs, preceding and following clauses, cited clauses, referenced attachment description areas, and associated table fields are included in the context scope of the corresponding document review object.

[0088] Retrieve the audit rule item that matches the document audit object from the audit rule base based on the document type information;

[0089] When retrieving audit rule items, the audit rule items are read from the audit rule library according to the document type information. The audit object name, clause category, field name, attachment requirements, and signature requirements in the audit rule items are matched with the document audit object to obtain the audit rule item corresponding to the document audit object.

[0090] Input the document review object, text content, object type, evidence location, context scope, and review rule items into the large model, and the large model will map the document review object into review attributes, review rule conditions, and context relationships;

[0091] A large model refers to a pre-trained language model used for semantic understanding and structured mapping of document review objects. It can output review attributes, review rule conditions, and contextual relationships based on the document review object, review rule base, and contextual scope. Before inputting into the large model, the input data is organized in the order of document review object, text content, object type, evidence location, contextual scope, and review rule items. The large model identifies review attributes based on text content, review rule conditions based on review rule items, and contextual relationships between the document review object and adjacent or referenced objects based on the contextual scope.

[0092] The document review object, review attribute, and review rule conditions are used to form a review semantic triple. The evidence location and context relationship are bound to the review semantic triple to generate a set of review semantic triples.

[0093] The audit semantic triple consists of a document audit object, an audit attribute, and an audit rule condition. Evidence location is bound to the corresponding document audit object, used to mark the evidence location corresponding to the audit attribute; contextual relationship is bound to the corresponding audit attribute and audit rule condition, used to mark the source of the audit attribute's association under the current rule condition. After binding, a set of audit semantic triples is obtained.

[0094] In this embodiment, the construction of the ternary rule evidence background includes the following steps:

[0095] Read the set of semantic triples for review to obtain the document review object, review attributes, review rule conditions, evidence location, and contextual relationships;

[0096] Establish a ternary relationship based on the document review object, review attributes, and review rule conditions;

[0097] When establishing a ternary relation, the document review object, review attribute, and review rule condition within the same review semantic triple are read as the writing unit; a ternary relation is written when the document review object, review attribute, and review rule condition all exist.

[0098] Write the evidence location into the evidence anchor item corresponding to the document review object in the ternary relation, and write the context relation into the evidence anchor items corresponding to the review attributes and review rule conditions in the ternary relation;

[0099] When writing the evidence anchoring item, the evidence location is categorized into the document review object according to page number position, page coordinates, clause number, table cell coordinates, attachment description area position, and signature area position; the context relationship is categorized into the review attributes according to hierarchical relationship, sequential relationship, and reference relationship; and the review rule conditions are written into the same evidence anchoring item.

[0100] Merge duplicate ternary relationships according to the document review object, review attribute, and review rule conditions, merge the evidence positions and contextual relationships corresponding to the duplicate ternary relationships, and generate ternary rule evidence background.

[0101] When merging repeated ternary relationships, the merging condition is that the document review object, review attribute, and review rule conditions are completely identical. After merging, only one ternary relationship is retained. The evidence positions are arranged according to the document position order, and the contextual relationships are categorized and merged according to hierarchical relationships, sequential relationships, and referential relationships to form the ternary rule evidence background.

[0102] In this embodiment, the construction of the ternary closure conceptual lattice model includes the following steps:

[0103] Read the ternary relationship and corresponding evidence anchoring items from the ternary rule evidence background, establish concept nodes according to the correspondence between document review object, review attribute and review rule condition, and write document review objects with the same review attribute and review rule condition into the corresponding concept nodes;

[0104] When a concept node is created, the ternary relationship is grouped by the audit attribute and audit rule condition; document audit objects with the same audit attribute and the same audit rule condition are written into the same concept node; the concept node records the audit attribute, audit rule condition, document audit object, and evidence anchor item.

[0105] Establish concept grid connection relationships based on the inclusion relationships among document review objects, review attributes, and review rule conditions in the concept nodes;

[0106] When establishing a concept lattice connection, the document review objects, review attributes, and review rule conditions in different concept nodes are compared; when the document review objects, review attributes, and review rule conditions in one concept node are contained in another concept node, a concept lattice connection is established between the two concept nodes.

[0107] Configure rule closure operators based on the audit rule conditions and audit attributes in the audit rule base, configure evidence closure operators based on the evidence position and context relationship in the evidence anchoring item, configure rebuttal closure structures based on the conflicting relationships between audit attributes, and connect rule closure operators, evidence closure operators, and rebuttal closure structures to the concept lattice connection relationship.

[0108] When configuring the closure structure, the system reads the correspondence between audit rule conditions and audit attributes from the audit rule base, writes the correspondence into the rule closure operator, and connects the rule closure operator to the audit rule conditions and audit attributes in the concept node; it reads the evidence position and context relationship from the evidence anchor item, writes the evidence position and context relationship corresponding to the same audit attribute into the evidence closure operator, and connects the evidence closure operator to the audit attributes and evidence anchor item in the concept node; it reads the conflict relationship between audit attributes from the audit rule base and the audit semantic triple set, writes the conflict relationship into the rebuttal closure structure, and connects the rebuttal closure structure to the output of the rule closure operator and the evidence closure operator.

[0109] Concept nodes, concept lattice connections, rule closure operators, evidence closure operators, and proof-of-contradiction closure structures are incorporated into the ternary closure concept lattice model.

[0110] When the ternary closure concept lattice model is formed, the concept nodes, concept lattice connection relationships, rule closure operators, evidence closure operators, and proof-of-contrast closure structures are written into the same model; the rule closure operators, evidence closure operators, and proof-of-contrast closure structures are processed along the concept lattice connection relationships.

[0111] In this embodiment, the rule closure operator, evidence closure operator, and disproving closure structure are called in a progressive order. The rule closure operator processes the audit rule conditions and audit attributes to generate the rule closure output. The evidence closure operator reads the rule closure output and the evidence position and context relationship in the evidence anchoring item to generate the evidence closure output. The disproving closure structure reads the rule closure output and the evidence closure output, matches mutually conflicting audit attribute combinations, and triggers conflict risk items.

[0112] In this embodiment, the generation of the rule closure output includes the following steps:

[0113] Using the document review object and review rule conditions as indexes, extract the corresponding review attributes from the background of the ternary rule evidence and mark them as review attributes that have appeared.

[0114] When the rule closure processing begins, the document review object and the review rule condition are used as indexes to locate the ternary relationship with the same document review object and the same review rule condition in the ternary rule evidence background. The located review attributes are written into the review attributes that have appeared. When the same index hits several ternary relationships, the review attributes are merged according to the evidence position, and duplicate review attributes are retained once.

[0115] Match the audit attributes that should be satisfied according to the audit rule conditions in the audit rule base, and call the rule closure operator to process the audit attributes that should be satisfied and the audit attributes that have already appeared;

[0116] The audit rule base records the correspondence between audit rule conditions, object types, and audit attributes. During rule closure processing, the corresponding rule item is located by the audit rule condition, and then the object type of the document audit object is matched with the object type in the audit rule item. The matched audit attribute is taken as the audit attribute that should be satisfied. When the same audit rule condition and the same object type correspond to multiple audit attributes, multiple audit attributes are entered into the rule closure operator together and processed with the existing audit attributes.

[0117] Compare the required audit attributes with the existing audit attributes one by one, and mark the audit attributes that are not covered by the existing audit attributes.

[0118] During the item-by-item comparison, name matching, semantic category matching, and audit rule condition matching are performed on the corresponding audit attributes and existing audit attributes. Audit attributes that are successfully matched are considered to be covered by existing audit attributes, and audit attributes that are not successfully matched are recorded as audit attributes that are not covered by existing audit attributes.

[0119] Write the document review object, review rule conditions, existing review attributes, and review attributes not covered by existing review attributes into the rule closure output.

[0120] When writing the rule closure output, the document review object and review rule conditions are recorded to identify the review attributes that have appeared, the review attributes that have not been covered by the review attributes that have appeared, the corresponding ternary relations and the evidence position, forming the rule closure output for the evidence closure operator to call.

[0121] In this embodiment, the generation of the evidence closure output includes the following steps:

[0122] Using the document review object and review rule conditions in the rule closure output as indexes, extract the corresponding review attributes, evidence locations and contextual relationships from the review semantic triple set;

[0123] During evidence closure processing, the document review object and review rule conditions in the rule closure output are used to locate the set of review semantic triples. The review attributes, evidence positions, and contextual relationships of the same document review object and the same review rule conditions are extracted and used as the processing objects of the evidence closure operator.

[0124] Evidence locations are grouped according to the same review attribute to form a grouped evidence location;

[0125] When merging evidence locations, evidence locations are grouped according to the same audit attribute. Duplicate evidence locations are retained once. Different page number locations and different page coordinates are arranged in document location order to form the merged evidence locations.

[0126] Based on the context, perform consistency checks on the audit attributes corresponding to the merged evidence locations and generate consistency check results.

[0127] During consistency verification, the consistency of the same audit attribute in multiple evidence locations is checked based on the context relationship corresponding to the merged evidence locations. When the hierarchical relationship does not correspond, the order relationship is broken, the reference relationship cannot be matched, or the content of multiple evidence locations conflicts, a consistency verification result is formed.

[0128] Write the audit attributes, merged evidence locations, contextual relationships, and consistency verification results into the evidence closure output.

[0129] When writing the evidence closure output, the merged evidence location, context relationship, and consistency verification result are written according to the audit attributes. At the same time, the corresponding document audit object and audit rule conditions are retained for the counter-evidence closure structure to call.

[0130] In this embodiment, the output of the audit report includes the following steps:

[0131] The code calls the evidence-of-contrast closure structure to receive the rule closure output and the evidence closure output. Based on the document review object, review rule conditions and review attributes, it establishes a correspondence between the review attributes in the rule closure output that are not covered by the already appearing review attributes, the consistency verification results in the evidence closure output and the merged evidence positions.

[0132] When establishing a correspondence, the document review object, review rule conditions, and review attribute matching rule closure output and evidence closure output are used. When the document review objects, review rule conditions, and review attributes are the same, and the review attributes belong to the same review attribute category in the review rule base, the review attributes not covered by the already appearing review attributes, the consistency check results, and the merged evidence positions are sent into the rebuttal closure structure.

[0133] Under the same document review object and the same review rule, the review attributes after establishing the corresponding relationship are input into the counter-proof closure structure to match mutually contradictory review attribute combinations;

[0134] When processing the closure structure of the proof of contradiction, the document is grouped according to the review object and review rule conditions, and the review attributes within the same group are combined and matched. During the matching, the mutually exclusive review attribute pairs, dependent review attribute pairs, and order constraint review attribute pairs in the review rule library are called. If the review attribute combination matches any of the mutually exclusive review attribute pairs, dependent review attribute pairs, and order constraint review attribute pairs, it is determined to be a mutually conflicting review attribute combination. If the above review attribute pairs are not matched, no conflict risk item is triggered.

[0135] The audit attributes that match mutually conflicting audit attribute combinations, the corresponding consistency check results, and the merged evidence locations are marked as conflict risk items;

[0136] When marking conflict risk items, the conflicting audit attribute combinations, corresponding consistency verification results, merged evidence locations, and audit rule conditions are written into the same conflict risk item; when there are multiple conflicting audit attribute combinations at the same evidence location, they are recorded separately according to the audit rule conditions.

[0137] The rule closure output, evidence closure output, and conflict risk items are sorted according to the evidence location. The output includes audit attributes not covered by existing audit attributes, consistency verification results, and conflict risk items.

[0138] When outputting the audit report, the evidence is first sorted by location, with the sorting order being page number, page coordinates, clause number, table cell coordinates, attachment description area location, and signature area location in that order. After sorting, the audit report is generated based on the audit attributes not covered by existing audit attributes in the document audit object record, consistency check results, conflict risk items, evidence location, and audit rule conditions.

[0139] In this embodiment, the matching of mutually conflicting audit attribute combinations includes the following steps:

[0140] Extract mutually exclusive audit attribute pairs, dependent audit attribute pairs, and sequential constraint audit attribute pairs under the same audit rule condition from the audit rule base. Dependent audit attribute pairs are recorded as directed relationships from preceding audit attributes to subsequent audit attributes.

[0141] The audit rule base records mutually exclusive audit attribute pairs, dependent audit attribute pairs, and order constraint audit attribute pairs according to the audit rule conditions. Mutually exclusive audit attribute pairs represent two audit attributes that cannot be true simultaneously under the same audit rule condition. Dependent audit attribute pairs represent audit attributes that require the other audit attribute to exist simultaneously if the other audit attribute is true. Order constraint audit attribute pairs represent audit attributes whose order of appearance in the document must satisfy the audit rule condition.

[0142] The audit attributes that enter the rebuttal closure structure are grouped according to the document audit object, audit rule conditions, and object type, and the audit attributes with evidence positions are arranged according to the evidence position.

[0143] Before processing the closure structure of the evidence of rebuttal, the audit attributes are grouped according to the document audit object, audit rule conditions, and object type; the audit attributes with evidence location are arranged according to page number position, page coordinates, clause number, and table cell coordinates, which serve as the basis for subsequent sequential verification.

[0144] Perform pairwise matching on the grouped audit attributes, matching each audit attribute pair with mutually exclusive audit attribute pairs, and marking mutually conflicting audit attribute combinations when a match is found.

[0145] When performing mutually exclusive matching, under the same document review object, the same review rule condition, and the same object type, the review attribute is matched with mutually exclusive review attribute pairs in the review rule library; when a mutually exclusive review attribute pair is matched, it is marked as a combination of mutually conflicting review attributes.

[0146] Dependency validation is performed on the grouped audit attributes. The pre-audit attributes and post-audit attributes in the dependent audit attribute pairs are matched with the rule closure output and the evidence closure output, respectively. When the pre-audit attribute exists and the post-audit attribute is not covered by the existing audit attribute, the mutually conflicting audit attribute combinations are marked.

[0147] Dependency audit attributes are recorded according to a directed relationship. Pre-audit attributes represent audit attributes that trigger associated audit requirements under the same audit rule, while post-audit attributes represent audit attributes that should exist simultaneously after the pre-audit attribute is established. During dependency validation, pre-audit attributes are matched with existing audit attributes, and post-audit attributes are matched with audit attributes not covered by existing audit attributes. When a pre-audit attribute exists but a post-audit attribute is missing, it is marked as a conflicting combination of audit attributes.

[0148] Perform order validation on the grouped audit attributes, compare the order of appearance of the audit attribute pairs in the merged evidence positions, and mark the conflicting audit attribute combinations when the order of appearance is inconsistent with the order constraints in the audit rule base.

[0149] During sequential verification, the order of occurrence of the two audit attributes in the audit attribute pair is constrained according to the position of the merged evidence. If the order of occurrence does not meet the order constraints in the audit rule base, the two audit attributes are marked as a combination of mutually conflicting audit attributes.

[0150] Example 1: To verify the feasibility of this invention in practice, it was applied to a centralized review scenario for business documents. In this scenario, the documents to be reviewed mainly include contract texts, approval materials, policy documents, attachment descriptions, and signature pages. The document source formats are inconsistent, and some documents have issues such as inconsistencies between main text clauses and table fields, mismatched attachment numbers, inconsistent signature subjects with main text subjects, missing related conditions in payment clauses, and conflicting statements in approval materials. Traditional review methods typically rely on manual page-by-page checks or searching for risky content using keywords and fixed rules. While these methods can identify some explicit omissions and formatting anomalies, they are prone to overlooking content consistency across paragraphs, tables, attachment description areas, and signature areas. Although general large-scale model review can understand the semantics of the text, directly generating review opinions can easily lead to problems such as unclear basis locations, ambiguous risk sources, repeated prompts for the same issues, and confusion in judgments under different review rule conditions.

[0151] In this scenario, the documents to be reviewed are first imported into the review platform. The platform performs layout and structure analysis on each document, identifying heading levels, clause numbers, table areas, attachment description areas, and signature areas. It then segments the main text paragraphs, table fields, attachment descriptions, and signature content into document review objects, writing evidence locations for each document review object. Subsequently, the large model does not directly output the final review conclusion. Instead, it combines the review rule base to map the document review objects to review attributes, review rule conditions, evidence locations, and contextual relationships, generating a set of review semantic triples. The review platform then constructs a ternary rule evidence background based on the review semantic triple set, writing the document review object, review attributes, and review rule conditions into the ternary relationship, and the evidence location and contextual relationship into the evidence anchor item. Each review judgment can be traced back to the corresponding original text location.

[0152] In the subsequent review process, the ternary closure concept lattice model forms the connection relationship between concept nodes and concept lattices based on the ternary rule evidence background. The rule closure operator first processes the review rule conditions and review attributes, identifying review attributes that the document review object should satisfy under the current review rule conditions but which are not actually present. For example, when a payment amount exists, the rule closure operator will further check whether the payment period, payment conditions, and liability for breach of contract are associated. The evidence closure operator then merges different evidence locations according to the same review attribute, performing consistency checks on similar information in the main text, tables, attachment descriptions, and signature areas. For example, when the contract subject appears simultaneously on the first page of the main text, in the main text clauses, and in the signature area, the evidence closure operator will merge multiple evidence locations, checking whether the name, authorization relationship, and signing object are consistent. The evidence-by-contrast closure structure finally processes the rule closure output and the evidence closure output, matching mutually exclusive review attribute pairs, dependent review attribute pairs, and sequence constraint review attribute pairs, triggering conflict risk items. For example, if the amount in the main text differs from the amount in the payment form, the delivery period conflicts with the period specified in the attachments, or the signature does not correspond to the signature in the main text, it will be included in the conflict risk item, and the location of the evidence and the basis for the audit will be given in the audit report.

[0153] To verify the effectiveness, a batch of anonymized business documents was selected as the test sample, totaling 480 documents, including main text pages, table pages, attachment description pages, and signature pages. On average, each document contained 18.6 document review objects. The manual review results served as the baseline, identifying 1268 valid risk items, including 384 missing explicit fields, 296 missing rules, 327 inconsistent evidence, and 261 cross-regional conflicts. Comparison methods included manual rule review, standard large-scale model review, and the method of this invention. Manual rule review primarily relied on keywords, field templates, and fixed rules; standard large-scale model review directly output review opinions based on the full text content; the method of this invention employed review semantic triples, triple rule evidence background, and three types of closure progressive review processing. Relevant data is shown in Table 1.

[0154] Table 1 Comparison of Intelligent Document Review Results

[0155] Audit accuracy 82.4% 87.6% 94.8% Risk recall 76.9% 84.1% 93.5% False positive rate 15.7% 12.8% 6.4% Average time to complete 31.5 minutes per part 12.8 minutes per part 6.9 minutes per part Localization accuracy 70.6% 78.3% 96.1%

[0156] As shown in Table 1, the method of this invention outperforms both manual rule review and ordinary large-scale model review in all indicators. Manual rule review achieved an accuracy rate of 82.4% and a risk recall rate of 76.9%, indicating a certain ability to identify issues such as missing fixed fields and formatting problems. However, it lacks sufficient coverage when facing issues such as missing rules, inconsistent evidence, and cross-regional conflicts. Ordinary large-scale model review improved its accuracy rate to 87.6% and its risk recall rate to 84.1%, demonstrating that large-scale models can enhance semantic understanding capabilities. However, due to the lack of stable ternary structure constraints and evidence closure processing, it still suffers from unstable risk basis and inaccurate evidence location.

[0157] The method of this invention achieves an accuracy rate of 94.8% and a risk recall rate of 93.5%, higher than the two comparison methods. The main reason is that this invention does not directly output the review conclusion from the large model, but first forms review semantic triples, and then performs progressive processing through triple rule-based evidence background and triple closure concept lattice models. The rule closure operator can identify review attributes that should be satisfied but have not appeared, the evidence closure operator can merge multiple evidence positions corresponding to the same review attribute and perform consistency checks, and the refutation closure structure can further match mutually contradictory review attribute combinations. Therefore, the identification of missing, inconsistent, and conflicting items is more complete.

[0158] In terms of false positive rate and location accuracy, the false positive rate of the method of this invention is 6.4%, lower than the 15.7% of manual rule review and the 12.8% of ordinary large-scale model review; the location accuracy reaches 96.1%, also higher than the two comparison methods. This indicates that the present invention binds the evidence location, contextual relationship and review attributes through evidence anchoring items, which can reduce unfounded risk warnings, and the risk content in the review report directly corresponds to the specific document location. In terms of average time consumption, the method of this invention is 6.9 minutes / document, lower than the 12.8 minutes / document of ordinary large-scale model review and the 31.5 minutes / document of manual rule review, indicating that the three types of closure progressive processing not only improve the review quality, but also reduce the time spent by humans repeatedly searching for evidence and verifying the source of risk.

[0159] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A smart document review method based on a large model, characterized in that, Includes the following steps: Retrieve the documents to be reviewed, document type information, and review rule base; perform layout and structure parsing on the documents to be reviewed; and generate a set of document review objects. The large model is invoked to process the document review object set and the review rule base, mapping the document review object to the corresponding review attributes, review rule conditions, evidence location and context relationship, and generating a set of review semantic triples; Construct the ternary rule evidence background based on the set of audit semantic triples, write the document audit object, audit attribute and audit rule condition into the ternary relation, and write the evidence location and context relationship into the evidence anchor item of the ternary relation; Based on the ternary rule evidence background, a ternary closure concept lattice model is constructed, and rule closure operators, evidence closure operators, and proof by contradiction closure structures are configured. The rule closure operator is invoked to process the audit rule conditions and audit attributes, deduce the audit attributes that the document audit object should satisfy under the audit rule conditions, compare the audit attributes that should be satisfied with the audit attributes in the audit semantic triple set, and generate the rule closure output; The evidence closure operator is invoked to process the rule closure output, the audit semantic triple set, the evidence position and context relationship, and the evidence positions are merged according to the same audit attribute. Consistency check is performed on the merged evidence positions to generate the evidence closure output. The system calls upon the output of the rule closure and evidence closure of the counter-evidence closure structure, matches mutually conflicting audit attribute combinations and triggers conflict risk items, and outputs an audit report.

2. The intelligent document review method based on a large model according to claim 1, characterized in that, The generation of the document review object set includes the following steps: Obtain the document to be reviewed, document type information, and review rule base; read the page content, page number position, and layout coordinates of the document to be reviewed. Perform layout parsing on the page content, marking heading levels, paragraph boundaries, clause numbers, table areas, attachment description areas, and signature areas; Perform structural analysis on the layout analysis results to establish hierarchical, sequential, and referential relationships between headings, paragraphs, clauses, table fields, attachment description areas, and signature areas; Segment the documents to be reviewed according to hierarchical, sequential, and referential relationships, and write the evidence location for the reviewed documents after segmentation; Based on document type information and the review rule base, document review objects are filtered to generate a set of document review objects.

3. The intelligent document review method based on a large model according to claim 1, characterized in that, The generation of the audit semantic triple set includes the following steps: Read the document review objects and evidence locations in the document review object set, extract the text content and object type of the document review objects, and determine the context scope of the document review objects based on hierarchical, sequential, and referential relationships; Retrieve the audit rule item that matches the document audit object from the audit rule base based on the document type information; Input the document review object, text content, object type, evidence location, context scope, and review rule items into the large model, and the large model will map the document review object into review attributes, review rule conditions, and context relationships; The document review object, review attribute, and review rule conditions are used to form a review semantic triple. The evidence location and context relationship are bound to the review semantic triple to generate a set of review semantic triples.

4. The intelligent document review method based on a large model according to claim 1, characterized in that, The construction of the ternary rule evidence background includes the following steps: Read the set of semantic triples for review to obtain the document review object, review attributes, review rule conditions, evidence location, and contextual relationships; Establish a ternary relationship based on the document review object, review attributes, and review rule conditions; Write the evidence location into the evidence anchor item corresponding to the document review object in the ternary relation, and write the context relation into the evidence anchor items corresponding to the review attributes and review rule conditions in the ternary relation; Merge duplicate ternary relationships according to the document review object, review attribute, and review rule conditions, merge the evidence positions and contextual relationships corresponding to the duplicate ternary relationships, and generate ternary rule evidence background.

5. The intelligent document review method based on a large model according to claim 1, characterized in that, The construction of the ternary closure conceptual lattice model includes the following steps: Read the ternary relationship and corresponding evidence anchoring items from the ternary rule evidence background, establish concept nodes according to the correspondence between document review object, review attribute and review rule condition, and write document review objects with the same review attribute and review rule condition into the corresponding concept nodes; Establish concept grid connection relationships based on the inclusion relationships among document review objects, review attributes, and review rule conditions in the concept nodes; Configure rule closure operators based on the audit rule conditions and audit attributes in the audit rule base, configure evidence closure operators based on the evidence position and context relationship in the evidence anchoring item, configure rebuttal closure structures based on the conflicting relationships between audit attributes, and connect rule closure operators, evidence closure operators, and rebuttal closure structures to the concept lattice connection relationship. Concept nodes, concept lattice connections, rule closure operators, evidence closure operators, and proof-of-contradiction closure structures are incorporated into the ternary closure concept lattice model.

6. The intelligent document review method based on a large model according to claim 5, characterized in that, The rule closure operator, evidence closure operator, and disproving closure structure are called in a progressive order. The rule closure operator processes the audit rule conditions and audit attributes to generate the rule closure output. The evidence closure operator reads the rule closure output and the evidence position and context relationship in the evidence anchor item to generate the evidence closure output. The disproving closure structure reads the rule closure output and the evidence closure output, matches mutually conflicting audit attribute combinations, and triggers conflict risk items.

7. The intelligent document review method based on a large model according to claim 1, characterized in that, The generation of the rule closure output includes the following steps: Using the document review object and review rule conditions as indexes, extract the corresponding review attributes from the background of the ternary rule evidence and mark them as review attributes that have appeared. Match the audit attributes that should be satisfied according to the audit rule conditions in the audit rule base, and call the rule closure operator to process the audit attributes that should be satisfied and the audit attributes that have already appeared; Compare the required audit attributes with the existing audit attributes one by one, and mark the audit attributes that are not covered by the existing audit attributes. Write the document review object, review rule conditions, existing review attributes, and review attributes not covered by existing review attributes into the rule closure output.

8. The intelligent document review method based on a large model according to claim 1, characterized in that, The generation of the evidence closure output includes the following steps: Using the document review object and review rule conditions in the rule closure output as indexes, extract the corresponding review attributes, evidence locations and contextual relationships from the review semantic triple set; Evidence locations are grouped according to the same review attribute to form a grouped evidence location; Based on the context, perform consistency checks on the audit attributes corresponding to the merged evidence locations and generate consistency check results. Write the audit attributes, merged evidence locations, contextual relationships, and consistency verification results into the evidence closure output.

9. The intelligent document review method based on a large model according to claim 1, characterized in that, The output of the audit report includes the following steps: The code calls the evidence-of-contrast closure structure to receive the rule closure output and the evidence closure output. Based on the document review object, review rule conditions and review attributes, it establishes a correspondence between the review attributes in the rule closure output that are not covered by the already appearing review attributes, the consistency verification results in the evidence closure output and the merged evidence positions. Under the same document review object and the same review rule, the review attributes after establishing the corresponding relationship are input into the counter-proof closure structure to match mutually contradictory review attribute combinations; The audit attributes that match mutually conflicting audit attribute combinations, the corresponding consistency check results, and the merged evidence locations are marked as conflict risk items; The rule closure output, evidence closure output, and conflict risk items are sorted according to the evidence location. The output includes audit attributes not covered by existing audit attributes, consistency verification results, and conflict risk items.

10. The intelligent document review method based on a large model according to claim 9, characterized in that, The matching of conflicting audit attribute combinations includes the following steps: Extract mutually exclusive audit attribute pairs, dependent audit attribute pairs, and sequential constraint audit attribute pairs under the same audit rule condition from the audit rule base. Dependent audit attribute pairs are recorded as directed relationships from preceding audit attributes to subsequent audit attributes. The audit attributes that enter the rebuttal closure structure are grouped according to the document audit object, audit rule conditions, and object type, and the audit attributes with evidence positions are arranged according to the evidence position. Perform pairwise matching on the grouped audit attributes, matching each audit attribute pair with mutually exclusive audit attribute pairs, and marking mutually conflicting audit attribute combinations when a match is found. Dependency validation is performed on the grouped audit attributes. The pre-audit attributes and post-audit attributes in the dependent audit attribute pairs are matched with the rule closure output and the evidence closure output, respectively. When the pre-audit attribute exists and the post-audit attribute is not covered by the existing audit attribute, the mutually conflicting audit attribute combinations are marked. Perform order validation on the grouped audit attributes, compare the order of appearance of the audit attribute pairs in the merged evidence positions, and mark the conflicting audit attribute combinations when the order of appearance is inconsistent with the order constraints in the audit rule base.