Knowledge base intelligent deep auditing method and system based on AI large model
Through the intelligent deep audit method of the knowledge base based on the AI big model, the problem of difficulty in identifying risks with manual and traditional tools in highly complex business scenarios is solved, efficient and accurate auditing and compliance judgment are achieved, and the cost of manual review is reduced.
Patent Information
- Application Number
- CN202511186941.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-25
AI Technical Summary
In highly complex business scenarios, manual review and traditional automation tools find it difficult to effectively identify key risks in cross-chapter logical chains, reference chains, and reverse conditional expressions, resulting in long review cycles, many misjudgments and missed judgments, high manual review costs, and an inability to guarantee compliance.
A knowledge base intelligent deep audit method based on AI large models is adopted. By parsing unstructured documents, a set of semantic audit units is constructed, a graph structure based on context dependencies is established, a regulatory knowledge embedding library is introduced, a two-layer risk discrimination function is used to identify candidate problem items, and causal chain backtracing is performed to generate structured audit conclusions.
It automatically completes deep semantic analysis and regulatory adaptation in semi-structured, cross-chapter, and highly cited project documents, significantly improving audit accuracy and efficiency, reducing the burden of manual review, and providing reproducible and traceable intelligent compliance audit capabilities.
Smart Images

Figure CN120746503A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of AI big models, and in particular relates to a method and system for intelligent deep auditing of knowledge bases based on AI big models. Background Art
[0002] In highly complex business scenarios such as bidding, budget review, and government contracts, an organization's internal knowledge base plays a core role in risk management and decision support. As project scale expands and regulations are frequently updated, the documents awaiting review have expanded from structured fields to PDF contracts with meticulously numbered pages and frequent cross-section references. These documents, along with scanned copies, charts, and multiple versions of regulations, create a highly heterogeneous information ecosystem. Manual review relies on personal experience and item-by-item comparisons, which is not only time-consuming but also difficult to ensure consistency. This is especially true when dealing with cross-section logical chains, reference chains, and inverse conditional expressions, making it easy to miss key risks. Traditional automated tools, however, often rely solely on keyword matching or fixed template filtering, failing to understand the causal structure between paragraphs and the applicable boundaries of legal provisions. Furthermore, they are unable to provide reliable judgments when multiple versions of regulations are in parallel and the semantics of clauses are inconsistent. The result is long review cycles, frequent misjudgments and omissions, and high manual review costs, severely hindering the efficiency of bidding and budgeting processes and exposing organizations to uncontrollable potential risks in regulatory compliance. Summary of the Invention
[0003] The purpose of this invention is to propose a knowledge base intelligent deep audit method and system based on AI big model to solve the above problems.
[0004] In order to achieve the above objectives, a first aspect of the present invention provides a knowledge base intelligent in-depth audit method based on an AI large model, the method comprising the following steps: S1. Obtain unstructured files uploaded by users, parse the unstructured documents and extract semantic review units, and construct a semantic review unit set; S2. Construct a context-dependent graph structure based on the semantic review unit set, mark the continuous, neighbor, reference, and reverse logic edge types, and integrate the compliance attention mechanism to update the node representation to obtain a node semantic representation set; S3. Introducing a regulatory knowledge embedding library, combining the graph structure and node semantic representation set design to identify candidate problem items through a two-layer risk discrimination function; S4. Perform causal chain backtracing on the candidate problem items to generate a structured audit conclusion.
[0005] Furthermore, the parsing of unstructured documents includes: The integrated parsing tool PDFPlumber performs structural preprocessing on each unstructured file, including the following key content collection: text content collection, number recognition, title format extraction, and structural position annotation.
[0006] Furthermore, the unstructured document is parsed as follows: Traversing the paragraph set formed by the unstructured document, using a template matching function and a preset semantic template set to perform review unit identification, and generating a corresponding semantic review unit set; wherein the template matching function is used to perform keyword and regular rule identification on each paragraph to generate a hit result; The semantic vector representation of the set of semantic review units is calculated by a lightweight vector encoder to generate a semantic vector representation of the corresponding semantic review unit, which represents the overall semantic expression of its content.
[0007] Furthermore, the lightweight vector encoder is based on the frozen parameter model of the first 4 layers of BERT-base, and only outputs the [CLS] vector as representation.
[0008] Furthermore, the S2 specifically includes: Obtain the fields of the semantic review unit, including: original text content, location of the unstructured document, semantic classification label, and semantic vector representation of the corresponding semantic review unit; Map the semantic review unit into a node set, and map the numbered consecutive edges, semantic neighbor edges, text reference edges, and reverse logic edges into edge sets to construct a graph structure; The adjacent nodes are weighted, and combined with the semantic vector representation of the corresponding semantic review unit, the context embedding representation of the aggregated corresponding semantic review unit is generated through adjacency aggregation, and a node semantic representation set is constructed.
[0009] Furthermore, the numbered continuous edge indicates that if the number difference between any two semantic review units is If they are in the same section, add them; The semantic neighbor edge indicates that if the cosine similarity between any two semantic review units is greater than 0.85, the content is considered to be repeated or extended; The text reference edge indicates that if a referential phrase appears at a position in the unstructured document, the corresponding semantic review unit is hit and located; The reverse logic edge indicates that if the current semantic review unit contains a reverse logic phrase and the content is opposite to the expression direction of another semantic review unit, a reverse annotation edge is established.
[0010] Furthermore, the two-layer risk discrimination function is specifically: The first layer is a regulatory conflict possibility scoring function, which takes the minimum conflict possibility score as the goal and combines the context embedding representation and the embedding vector of the legal provisions in the regulatory knowledge base to generate the minimum conflict possibility score between the current semantic review unit and the regulatory base; The second layer is the structural inconsistency logic residual function, which is used to identify problems caused by reference logic errors or numerical inconsistencies. It generates a structural consistency residual value by combining the node set, edge weights, and context embedding representation directly referenced by the current semantic review unit in the graph structure.
[0011] Furthermore, the combination of the graph structure and the node semantic representation set is designed to identify candidate problem items through a double-layer risk discrimination function, specifically: If the minimum conflict possibility score is greater than the first preset threshold or the structural consistency residual value is greater than the second preset threshold, the system records the current semantic review unit as a problem candidate, marks it as a regulatory conflict candidate or a structural inconsistency candidate, and outputs it to the problem candidate set, where each problem candidate contains the original text, page number, structure number, problem category, reference path node set and legal provision matching number.
[0012] Furthermore, the S4 specifically includes: For each question candidate, perform path backtracking in the graph structure, limiting the maximum backtracking depth to , and retain all text reference edges and reverse logic edges to form causal support paths for building support chains; Extracting nodes in all paths based on the causal support path to generate a structured conclusion object; For each question candidate's structured conclusion object, combined with the adjacent semantic review unit, a structured display weight vector is generated to highlight the most influential node in the path; The final output is a set of structured audit report conclusions; Among them, the structured conclusion object includes: review conclusion label, original paragraph content, page number, node number involved in the causal chain, directed edge sequence in the causal support path, text fragments from nodes in the causal support path, and system recommended operation label.
[0013] In a second aspect of the present invention, a knowledge base intelligent in-depth audit system based on an AI big model is provided, the system comprising: A file acquisition unit is used to acquire unstructured files uploaded by users, parse the unstructured documents, extract semantic review units, and construct a semantic review unit set; An identification unit is used to construct a context-dependent graph structure based on the set of semantic review units, mark the types of continuous, neighbor, reference and reverse logic edges, and integrate the compliance attention mechanism to update the node representation to obtain a set of node semantic representations; A document analysis unit is used to introduce a regulatory knowledge embedding library, combine the graph structure and the node semantic representation set design to identify candidate problem items through a double-layer risk discrimination function; The problem tracing and report generating unit is used to perform causal chain backtracing on the candidate problem items and generate structured audit conclusions.
[0014] The beneficial technical effects of the present invention are at least as follows: To address the above pain points, the present invention proposes an end-to-end deep audit method. First, the semantic audit unit division technology driven by structural rules is used to decompose the unstructured PDF contract into SRU blocks with independent review significance. Then, through graph structure modeling, the numbering continuity, semantic neighbors, reference references and reverse logic are explicitly marked as multi-type edges to generate a computable context dependency network. On this basis, the node representation update is completed by using compliance attention aggregation and content avoidance factors, and the regulatory matching deviation and structural consistency residual are simultaneously measured through a two-layer risk discrimination function to accurately screen out potential high-risk candidates. Finally, combined with the causal chain backtracking algorithm, the trigger path, supporting nodes and legal provisions of each problematic clause are aggregated into a structured audit conclusion, and the key causes are highlighted with an interpretable weight vector, realizing a closed loop from problem discovery, tracing the cause to generating an audit report with evidence. This method does not rely on large-scale annotation or black-box models, and can automatically complete deep semantic analysis and regulatory adaptation in semi-structured, cross-chapter, and highly cited project documents, significantly improving audit accuracy and efficiency, reducing the burden of manual review, and providing governments and enterprises with reproducible, traceable, and intelligent compliance audit capabilities that comply with the latest regulatory requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The present invention is further described with reference to the accompanying drawings. However, the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative effort.
[0016] Figure 1 This is a flow chart of the intelligent deep audit method of the knowledge base based on the AI big model of the present invention. DETAILED DESCRIPTION
[0017] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0018] like Figure 1 As shown, an embodiment of the present invention provides a method for intelligent deep auditing of a knowledge base based on an AI large model, the method comprising: S1. Obtain unstructured files uploaded by users, parse unstructured documents and extract semantic review units, and build a semantic review unit set.
[0019] Specifically, this step targets project-related PDF files with "clear structure and standardized format," such as bidding and procurement documents, project budgets, and government contracts. The goal is to extract Semantic Review Units (SRUs) from these documents, each with independent review value, to support subsequent dependency modeling and review decisions. While these documents may have neat layouts in real-world scenarios, their content is often unstable. For example, the same "payment terms" may appear in different chapters, and their wording may vary depending on writing habits. This necessitates that content segmentation cannot rely solely on paragraph segmentation or title extraction, but rather requires a balance between structural boundaries and semantic recognition.
[0020] The input of this step is a collection of PDF files uploaded by users in the system The system uses the integrated parsing tool PDFPlumber to perform structural preprocessing on each PDF file, including the following key content collection: text content collection (extraction by page and paragraph); number recognition (recognition of structural numbers such as "3.2" and "4.1.5"); title format extraction (recognition of font, font size, bold status, etc.); and structural position annotation (marking the page and line number of the paragraph).
[0021] For example, the following content exists on page 3 of a certain file: 3.2 Bidder Qualification Requirements The bidder shall be an independent legal person registered in the People's Republic of China and shall have no major violation record in the past three years.
[0022] The system recognizes the number "3.2" as a sub-chapter title and splices its subordinate paragraphs into a structural segment object. The system then traverses the paragraph collection formed by the document , using the template matching function With a set of preset semantic templates Conduct audit unit identification: ; in, : No. A set of semantic review units extracted from a PDF document; : A collection of paragraphs after PDF structure parsing, each paragraph contains number, font, content and position; : A set of semantic templates, including keyword phrases, regular expressions, tag names, and matching thresholds; : Structural matching function, which performs keyword and regular rule recognition on each paragraph and generates hit results.
[0023] For example, a paragraph is: Payment terms: 30% of the total contract amount shall be paid within 10 working days after the contract is signed, and the remaining 70% shall be paid after the project is completed and accepted.
[0024] This paragraph matches the keywords and sentence structure of "Payment Terms Template" and is assigned the label "Payment Terms", forming an SRU unit. , and record its original position and hit template.
[0025] In order to ensure the semantic integrity of the extracted content, the system We further calculate its vector representation and evaluate the coherence between its sentences to eliminate review units with mixed structures or fragmented content. The embedding process is as follows: ; in, :Audit unit The semantic vector of , which represents the overall semantic expression of its content; : Lightweight vector encoder, based on the first 4 layers of BERT-base frozen parameter model, only outputs [CLS] vector as representation; : The review unit generated after the structure template is hit, and its text content is input into the encoder.
[0026] The system calculates the average pairwise cosine similarity between the vector and the internal sentences. If it is lower than 0.6, the paragraph is considered to be semantically incoherent. For example, there are mixed sentences such as "See the figure for details" and "Please refer to the attachment", which will be removed or split.
[0027] Finally, the system outputs a set of audit units with unified structure, clear semantics and consistent granularity. , as the basic input for constructing the logical dependency structure diagram in the subsequent steps.
[0028] S2. Construct a context-dependent graph structure based on the semantic review unit set, mark the continuous, neighbor, reference and reverse logic edge types, and integrate the compliance attention mechanism to update the node representation to obtain a node semantic representation set.
[0029] Specifically, this step is to establish contextual dependencies between the review units after completing the extraction of semantic review units (SRU), construct a document structure diagram, and perform context-enhanced modeling on each node to provide a semantically traceable structural foundation for subsequent risk identification and compliance audits. In the "bidding and procurement, project budget, government contract" type documents targeted by the present invention, although the numbering structure is clear, there are a large number of nonlinear relationships between the clauses, such as cross-chapter references, sequential conditional dependencies, logical reversal expressions, etc. The traditional method of establishing logical connections by numbering or adjacent paragraphs cannot meet the needs of "semantic consistency review" and "legal application chain tracking". Therefore, this step innovatively introduces "compliance attention aggregation items" and "context avoidance factors" based on structural mapping.
[0030] The input is the set of semantic review units output by the previous step , where each Contains the following fields: Original text content , the position in the PDF , semantic classification label , and the semantic vector representation generated by the embedding model in step 1 This set is all the basic data for structure mapping and context modeling in this step.
[0031] Furthermore, in order to capture the implicit reference and structural relationship between semantic review units, the present invention uses Build a semantic structure graph for nodes The types of edges in the graph include not only the common numbered continuous edges and semantic neighbor edges , and also innovatively introduced the "reference trigger edge" and "Chapter spans reverse edge" The latter is specifically used to mark the logical transposition relationship that appears in the "converse logical expression of contract conditions" in the document (for example, "no payment if acceptance is not completed"). The edge set is defined as follows: ; in, : The structure graph is constructed by the node set and edge sets composition; : It is composed of all SRU units, namely ; :Numbered structural edges, if and The number difference is If they are in the same section, add them; : semantically similar edge, if cosine similarity , deemed to be a duplication or expansion of the content; :Text citation edge, if Containing reference phrases such as "as mentioned above" and "see above", regular expression hits and locates ; :Reverse logic edge, if If it contains reverse logic phrases such as the sentence pattern of "If not...then...", and the content is opposite to the expression direction of another SRU, a reverse annotation edge is established. is the mapping function.
[0032] Furthermore, after the structural graph is constructed, in order to capture the contextual dependencies of each review unit in the semantic structure graph, the system introduces a structure-aware node representation update strategy. This strategy is based on the adjacency aggregation method of the classic graph neural network and innovatively adds "compliance attention items" and "content avoidance items" to enhance the influence weight of important review nodes and weaken the influence of noise nodes unrelated to the law. The specific formula is as follows: ; in, : Aggregated semantic review unit Contextual embedding representation of ; :node The basic semantic vector calculated in step 1 (dimension is 768); :picture Zhongyu A set of adjacent nodes connected by edges of any type; :from arrive The structural attention weight is assigned according to the semantic relevance + edge type, for example: Edge weight is higher than ; : Content avoidance factor, if If there are non-censored content prompts (such as "see the diagram for details" or "as listed in the attachment"), then ,otherwise ; : is the normalization term, .
[0033] It is understandable that the core innovation of this formula lies in: Instead of using a fixed average, adjacent nodes are weighted to introduce structural awareness. Punitively reduce the weight of SRUs with "non-censored expressions" to effectively eliminate structural noise; Weight calculation is combined with semantics and edge types to enhance the structural influence of "referential reference" and "logical reversal" in compliance review.
[0034] For example: In the "Project Acceptance" section, It means "no payment without acceptance", which refers to (payment terms definition), (Acceptance process), the system passes + Create direction edges and assign Higher attention value , and if Contains a statement such as "see Annex II", Set as , making The aggregation of focuses on the payment content rather than the acceptance details.
[0035] Compared to traditional "vector averaging" or "paragraph alignment" methods, this step significantly improves structural expressiveness and audit logic support. It does not rely on supervised annotation or learnable parameters, but instead expresses the potential dependencies between clauses in complex regulatory documents solely through structural rules and attention control.
[0036] S3. Introduce the regulatory knowledge embedding library, combine the graph structure and node semantic representation set design to identify candidate problem items through a two-layer risk discrimination function.
[0037] Specifically, after the semantic review unit (SRU) has completed the structure mapping and context representation aggregation, this step aims to identify content segments that may have potential risks such as regulatory conflicts, logical contradictions, and structural inconsistencies, mark them as candidate problem items, and use them for subsequent conclusion confirmation. Different from the traditional method based on keyword or rule triggering, this step combines the structure diagram , node context representation and regulatory knowledge embedded in the database By constructing a set of unique computational functions, this system systematically screens content units with "compliance issues." This mechanism is particularly well-suited for complex text systems like project contracts, budget documents, and procurement instructions, which require structured paragraphs, cross-references, and compliance requirements.
[0038] The input is the output of step 2, including the structure diagram , each semantic review unit in the figure Structure-aware representation of , and the legal knowledge base embedding representation .in The clauses from actual regulatory texts (such as the Government Procurement Law and the Bidding Law) are embedded using the same BERT-4 layer frozen model as SRU, and the format and numerical units are cleaned before embedding.
[0039] In actual applications, the present invention has found that certain problems frequently appear in project documents but are extremely difficult to identify using traditional technologies, such as: Surface compliance but semantic reversal (“no payment without acceptance” vs. “payment despite no acceptance criteria”); Inconsistent amount logic ("Payment 30%, 30%, 40%" vs. two installments in the detailed item); The clause is missing but the content is quoted ("see performance requirements" but there is no such clause); References to outdated regulations or irrelevant legal texts.
[0040] To solve such problems, the present invention introduces a candidate identification function with a two-layer discriminant logic. The first layer detects the possibility of regulatory conflict, the second layer evaluates the inconsistency of structural reasoning, and finally merges them into a problem risk score.
[0041] The first layer is the regulatory conflict possibility scoring function, which combines the semantic deviation main term + the article deviation adjustment term + the reference jump structure penalty term, and is constructed as follows: ; in, :express The minimum possible conflict score with the regulatory base, the larger the score, the more suspicious it is; : Generated in step 2 The context-aware semantic representation of , with a dimension of 768; :No. The embedding vector of each legal clause; : Semantic similarity with legal provisions; : Regulation offset adjustment item, indicating whether the law is a "weakly adapted law". If it is an old law or a non-universal law, then ,labeled by manual review or estimated by historical adaptation frequency; : Structural jump adjustment factor, if Cross-section citations exceed 2 levels, or the citation edge contains If there are more than 2 reverse edges, set it to 1.5, otherwise it is 1; : The maximum structure number span referenced (e.g., when referencing "4.3" and its own number is "1.2", it becomes 3.1), indicating its logical span; : The number of adjacent nodes is used to normalize the citation density.
[0042] It is understandable that the innovation of this formula design lies in that it not only considers the minimum distance matching of semantic vectors, but also introduces: Adjustment of legal provisions weight ( ), so that the new law is matched first; Reference path cost factor ( ), suppressing distant low-credibility citations; Number span explicit penalty ( ), suppressing the “pseudo-legitimacy” of cross-chapter citations.
[0043] when (System threshold, usually 0.35), the system determines that this clause has the risk of regulatory adaptation conflict.
[0044] The second layer is the structural inconsistency logic residual function, which is used to identify problems caused by reference logic errors or numerical inconsistencies, such as inconsistencies between the summary number and the sub-item, references to undefined content, etc. This function is based on the "structural consistency residual" between the structural path and the embedding and is defined as follows: ; in, : Structural consistency residual value, used to judge the degree of inference in semantic inconsistency with the content it refers to; :In the structure diagram The set of directly referenced nodes; :node Contextual semantic representation of ; : Edge weight, derived from the structural attention coefficient in step 2 ; : Represents the square of the Euclidean distance between vectors.
[0045] Intuitively speaking, if a clause refers to multiple sub-clauses, but its content vector is significantly different from the aggregated representation of the sub-clauses, then the clause has a semantic aggregation conflict, which may be due to problems such as "contradictory descriptions", "omission of conditions", or "improper merging".
[0046] The final candidate identification logic is: If or , the system records the for , and marked as "regulatory conflict candidates" or "structural inconsistency candidates", and output to the problem candidate set Each candidate item contains fields such as original text, page number, structure number, question category, reference path node set, and legal article matching number.
[0047] This step combines structural diagram aggregation, regulatory embedding comparison, and citation path analysis for the first time to identify "potential logical problem points" that traditional methods cannot handle. It not only improves the discovery rate, but also generates candidate problems with a "structural causal chain", which can provide a traceability basis for subsequent decision-making modules.
[0048] S4. Perform causal chain backtracing on the candidate problem items to generate a structured audit conclusion.
[0049] Specifically, this step has completed the identification of candidate question items in the previous stage, that is, generating a set of candidate questions , each Contains source semantic review unit , triggering reasons (such as 、 ), structure diagram The position in , and its contextual embedding The goal of this step is to trace the structural causal chain of identified problem items, generate clear conclusion labels, and organize structured display fields without introducing any new models or structural judgments, so as to achieve visual, readable, and traceable audit result output.
[0050] The project-related documents processed by this method generally feature stable paragraph numbering, dense structural references, and frequent semantic references. Therefore, conclusion generation cannot simply involve "red-marking" or "labeling"; it must include: the problem paragraph itself, the problem type, the corresponding risk indicator, the causal reference path chain, the set of supporting content nodes, the cited legal provisions (if any), and the final "conclusion suggestion label" (for example, "approved," "risk item," "missing item," "requires manual review," etc.).
[0051] First, the present invention is for each In the structure diagram Execute path backtracking in the process, and limit the maximum backtracking depth to , and keep all and Type edges (i.e., reference edges and reverse logic edges) form causal support paths , used to build the support chain. The path acquisition function is as follows: ; in, : Candidate questions The set of causal support paths; : Trigger the semantic review unit of the candidate; : Reference edge, indicating reference relationships within a paragraph, such as "as mentioned above" and "see previous item"; : Reverse expression edge, used to connect reverse expressions such as "no payment if not accepted"; path depth No more than 3 layers to avoid information drift.
[0052] The system uses this path collection Extract all nodes in the path , used to display supporting content, reference structure and context description. Then generate a structured conclusion object , whose fields include: : Review conclusion label, from Question type label; : original paragraph content, ; : PDF page number, from ; : The node numbers involved in the causal chain, such as ["3.2","3.4"]; : Path structure The sequence of directed edges in ; : A text fragment from a node in the path ; : System recommended operation label, for example: manual review required / modification recommended / can be automatically repaired.
[0053] In order to improve the interpretability of the conclusion, the present invention The conclusion generates a structured display weight vector, which is used to highlight which path nodes have the greatest influence in the display. The formula is as follows: ; in, : Candidate Supporting Nodes in Conclusion The display weight of : Structural edge weight, derived from the structural graph Middle edge weight matrix; :node The risk score of or , normalized; the weight vector is used in the front-end report to show the degree of contribution and highlight the key factors.
[0054] The final output is a set of structured audit report conclusions , which can be displayed on the page through the front-end system, including the main question content, supporting reference fragments, reference path map, recommended action items, etc.
[0055] For example, a conclusion item has the following structure: Main terms: Payment method is 30% + 30%, and the remaining 40% will be paid after the project is completed; Problem type: inconsistent structure; Reference chain: 3.2→3.3 (project acceptance), 3.5 (contract termination clause); Supporting content: Page 3 "Acceptance criteria have not yet been defined", Page 5 "Payment is based on completion of acceptance"; Suggestion: Manual review, there may be a problem with the payment terms not being closed.
[0056] Through this causal path backtracking and structured display mechanism, the system not only locates the problem, but also can "explain how the problem came about", opening up the complete chain of "trigger → attribution → presentation".
[0057] An embodiment of the present invention further provides a knowledge base intelligent in-depth review system based on an AI large model, the system comprising: A file acquisition unit is used to acquire unstructured files uploaded by users, parse the unstructured documents, extract semantic review units, and construct a semantic review unit set; An identification unit is used to construct a context-dependent graph structure based on the set of semantic review units, mark the types of continuous, neighbor, reference and reverse logic edges, and integrate the compliance attention mechanism to update the node representation to obtain a set of node semantic representations; A document analysis unit is used to introduce a regulatory knowledge embedding library, combine the graph structure and the node semantic representation set design to identify candidate problem items through a double-layer risk discrimination function; The problem tracing and report generating unit is used to perform causal chain backtracing on the candidate problem items and generate structured audit conclusions.
[0058] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0059] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a division of logical functions. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0060] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0061] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. The knowledge base intelligent deep audit method based on AI big model is characterized by: The method comprises the following steps: S1. Obtain unstructured files uploaded by users, parse the unstructured documents and extract semantic review units, and construct a semantic review unit set; S2. Construct a context-dependent graph structure based on the semantic review unit set, mark the continuous, neighbor, reference, and reverse logic edge types, and integrate the compliance attention mechanism to update the node representation to obtain a node semantic representation set; S3. Introducing a regulatory knowledge embedding library, combining the graph structure and node semantic representation set design to identify candidate problem items through a two-layer risk discrimination function; S4. Perform causal chain backtracing on the candidate problem items to generate a structured audit conclusion.
2. The method for intelligent deep audit of knowledge base based on AI big model according to claim 1 is characterized in that: The parsing of unstructured documents includes: The integrated parsing tool PDFPlumber performs structural preprocessing on each unstructured file, including the following key content collection: text content collection, number recognition, title format extraction, and structural position annotation.
3. The method for intelligent deep audit of knowledge base based on AI big model according to claim 2 is characterized in that: The unstructured document parsing is performed as follows: Traversing the paragraph set formed by the unstructured document, using a template matching function and a preset semantic template set to perform review unit identification, and generating a corresponding semantic review unit set; The template matching function is used to perform keyword and regular rule recognition on each paragraph to generate a hit result; The semantic vector representation of the set of semantic review units is calculated by a lightweight vector encoder to generate a semantic vector representation of the corresponding semantic review unit, which represents the overall semantic expression of its content.
4. The method for intelligent deep audit of knowledge base based on AI big model according to claim 3 is characterized in that: The lightweight vector encoder is based on the frozen parameter model of the first 4 layers of BERT-base and only outputs the [CLS] vector as representation.
5. The method for intelligent deep audit of knowledge base based on AI big model according to claim 3 is characterized in that: Said S2 specifically includes: Obtain the fields of the semantic review unit, including: original text content, location of the unstructured document, semantic classification label, and semantic vector representation of the corresponding semantic review unit; Map the semantic review unit into a node set, and map the numbered consecutive edges, semantic neighbor edges, text reference edges, and reverse logic edges into edge sets to construct a graph structure; The adjacent nodes are weighted, and combined with the semantic vector representation of the corresponding semantic review unit, the context embedding representation of the aggregated corresponding semantic review unit is generated through adjacency aggregation, and a node semantic representation set is constructed.
6. The method for intelligent deep audit of knowledge base based on AI big model according to claim 5 is characterized in that: The numbered continuous edge means: if the number difference between any two semantic review units is 1 and they are in the same chapter, then add them; The semantic neighbor edge indicates that if the cosine similarity between any two semantic review units is greater than 0.85, the content is considered to be repeated or extended; The text reference edge indicates that if a referential phrase appears at a position in the unstructured document, the corresponding semantic review unit is hit and located; The reverse logic edge indicates that if the current semantic review unit contains a reverse logic phrase and the content is opposite to the expression direction of another semantic review unit, a reverse annotation edge is established.
7. The method for intelligent deep audit of knowledge base based on AI big model according to claim 5 is characterized in that: The two-layer risk discrimination function is specifically: The first layer is a regulatory conflict possibility scoring function, which takes the minimum conflict possibility score as the goal and combines the context embedding representation and the embedding vector of the legal provisions in the regulatory knowledge base to generate the minimum conflict possibility score between the current semantic review unit and the regulatory base; The second layer is the structural inconsistency logic residual function, which is used to identify problems caused by reference logic errors or numerical inconsistencies. It generates a structural consistency residual value by combining the node set, edge weights, and context embedding representation directly referenced by the current semantic review unit in the graph structure.
8. The method for intelligent deep audit of knowledge base based on AI big model according to claim 7 is characterized in that: The design of combining the graph structure and the node semantic representation set to identify candidate problem items through a double-layer risk discrimination function is specifically as follows: If the minimum conflict possibility score is greater than the first preset threshold or the structural consistency residual value is greater than the second preset threshold, the system records the current semantic review unit as a problem candidate, marks it as a regulatory conflict candidate or a structural inconsistency candidate, and outputs it to the problem candidate set, where each problem candidate contains the original text, page number, structure number, problem category, reference path node set and legal provision matching number.
9. The method for intelligent deep audit of knowledge base based on AI big model according to claim 5 is characterized in that: Said S4 specifically includes: For each question candidate, perform path backtracking in the graph structure, limiting the maximum backtracking depth to , and retain all text reference edges and reverse logic edges to form causal support paths for building support chains; Extracting nodes in all paths based on the causal support path to generate a structured conclusion object; For each question candidate's structured conclusion object, combined with the adjacent semantic review unit, a structured display weight vector is generated to highlight the most influential node in the path; The final output is a set of structured audit report conclusions; Among them, the structured conclusion object includes: review conclusion label, original paragraph content, page number, node number involved in the causal chain, directed edge sequence in the causal support path, text fragments from nodes in the causal support path, and system recommended operation label.
10. The knowledge base intelligent deep audit system based on AI big model is characterized by: The system comprises: A file acquisition unit is used to acquire unstructured files uploaded by users, parse the unstructured documents, extract semantic review units, and construct a semantic review unit set; An identification unit is used to construct a context-dependent graph structure based on the set of semantic review units, mark the types of continuous, neighbor, reference and reverse logic edges, and integrate the compliance attention mechanism to update the node representation to obtain a set of node semantic representations; A document analysis unit is used to introduce a regulatory knowledge embedding library, combine the graph structure and the node semantic representation set design to identify candidate problem items through a double-layer risk discrimination function; The problem tracing and report generating unit is used to perform causal chain backtracing on the candidate problem items and generate structured audit conclusions.
Citation Information
Patent Citations
Real estate information verification method and system based on image technology
CN120318840A
Cited By
Explanatable building contract risk review method and system of knowledge graph with mechanism layer
CN121032435A
Intelligent matching method and device, computer equipment and storage medium
CN121144364A
Purchase file checking method based on large language model
CN121304204A
Double-track consistency execution instruction compiling method and system for drawing checking rule overview
CN121788077A
Dual-track consistency enforcement instruction compilation method and system
CN121788077B