An audit path planning and intelligent recommendation method based on a large model and a graph database
Patent Information
- Application Number
- CN202611070234.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-18
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]为了解决上述技术问题,本发明提供一种基于大模型与图数据库的审计路径规划与智能推荐方法,以解决现有技术中因对所有审计场景均执行统一深度检索与生成流程、导致标准审计场景响应效率低下的问题
[0031]本发明采用双层递进处理机制:对于标准审计场景,直接匹配预存模板快速响应;对于复杂或模糊场景,则触发图检索与向量检索的深度融合,结合冲突仲裁机制,从结构化的实体关系和非结构化的历史案例中提取高可信度知识,从而有效解决了传统方法在边界模糊情况下漏判、误判以及单一信息源不可靠的问题;
Smart Images

Figure CN122841103A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of audit informatization and artificial intelligence, specifically a method for audit path planning and intelligent recommendation based on large models and graph databases. Background Technology
[0002] In recent years, knowledge graphs and large language models have been gradually introduced into the field of audit information technology. Knowledge graphs can organize audit entities (such as auditees, legal provisions, and risk points) and their relationships in a graph structure, supporting relational queries and path analysis. Large language models, on the other hand, possess powerful semantic understanding and generation capabilities, capable of extracting information from unstructured text and generating audit recommendations. Existing technologies include using graph databases to store audit knowledge or leveraging vector databases to retrieve historical cases, combined with large language models to generate audit path planning, thus assisting auditors in their work to some extent.
[0003] However, existing solutions typically apply a uniform processing flow to all input when handling user queries: regardless of whether the audit scenario is a common standard type (such as auditing official spending on government vehicles and receptions) or a complex and fuzzy type, they all perform complete steps such as graph database retrieval, vector similarity retrieval, and large language model inference generation. This "one-size-fits-all" approach results in a large amount of computing resources being consumed on simple scenarios, leading to long system response times and low processing efficiency. In particular, when users repeatedly query standard audit scenarios, repeating deep retrieval and generation is unnecessary and reduces the real-time performance of the audit work. Therefore, there is an urgent need for a method that can adaptively switch processing paths according to the type of audit scenario, significantly improving the response efficiency of standard scenarios while ensuring the inference accuracy of complex scenarios. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides an audit path planning and intelligent recommendation method based on a large model and graph database, thereby solving the problem of low response efficiency in standard audit scenarios caused by performing a uniform deep retrieval and generation process on all audit scenarios in the prior art.
[0005] This invention provides an audit path planning and intelligent recommendation method based on large models and graph databases, comprising the following steps:
[0006] Step S1: Construct an audit knowledge graph, which specifically includes: extracting audit entities and relationships between entities from multi-source heterogeneous data, storing the entities and relationships in a graph database in a graph structure, and vectorizing audit case text fragments and storing them in a vector database.
[0007] Step S2: Obtain the audit business description text input by the user;
[0008] Step S3: Use a large language model to perform intent recognition on the audit business description text to determine the audit scenario type;
[0009] Step S4: Based on the audit scenario type, perform graph database retrieval and vector similarity retrieval in parallel to obtain retrieval results, which include subgraph data and case similarity fragments;
[0010] Step S5: Using the search results as context, knowledge reconstruction and reasoning generation are performed through a large language model to output a structured audit path planning scheme;
[0011] The method employs a two-layer progressive processing mechanism: after step S3, if the identified audit scenario type belongs to a preset standard scenario set, the pre-stored audit path template is directly matched as the output, skipping steps S4 and S5; if it does not belong to the standard scenario set, steps S4 and S5 are executed.
[0012] Preferably, the multi-source heterogeneous data includes at least two of the following: historical audit reports, audit ledgers, policy documents, and work prompts.
[0013] Preferably, the structured audit path planning scheme includes the following fields: audit focus, audit key points, recommended audit methods, key data sources, reference system clauses and similar project cases, and supports visual graph display.
[0014] Preferably, the system also includes step S6: dynamic maintenance and quality monitoring of the data source, which specifically includes: during system operation, using the audit business description entered by the user each time, the generated audit path planning scheme, and the user's adoption or feedback results of the scheme as incremental data, and incrementally updating the graph database and the vector database according to a preset time period; at the same time, performing timeliness detection on the stored policy documents, and downgrading or removing policy clauses that have been repealed or updated.
[0015] Preferably, after performing graph database retrieval and vector similarity retrieval in parallel in step S4, the method further includes a heterogeneous information source fusion and conflict arbitration step:
[0016] A consistency check is performed on the subgraph data retrieved from the graph database and the similar fragments of cases retrieved by vector similarity retrieval.
[0017] If the two are consistent in their assessment of key audit entities or risks, the weighted average of the two will be used as the search result.
[0018] If there is a contradiction between the two, conflict arbitration is initiated: the first credibility of the graph database retrieval results and the second credibility of the vector retrieval results are calculated separately. The first credibility is a weighted combination of the average weight value of the edges in the graph database and the grade score of the issuing agency of the corresponding institutional document. If multiple institutional documents are involved, the highest grade score is taken. The weight value of the edges is obtained by normalizing the co-occurrence frequency of the two entities in historical audit data. The grade of the issuing agency is preset to 4, 3, 2, and 1 for national, provincial, municipal, and county levels, respectively. The second credibility is determined based on the decay score of the number of days between the occurrence time of the case and the current date in the vector retrieval results. This decay score is proportional to the reciprocal of the number of days plus 1. The first credibility and the second credibility are compared, and the one with the higher credibility is taken as the final retrieval result. The results that are abandoned are marked as questionable information.
[0019] Preferably, it also includes step S7: closed-loop optimization driven by result feedback, specifically including:
[0020] After outputting the audit path planning scheme in step S5, obtain the user's quality evaluation of the scheme;
[0021] If the quality evaluation is lower than the preset threshold, the search link in step S4 or the generation link in step S5 is located based on the specific information of the quality defect. The search weight or generation parameters are adjusted, and steps S4 and S5 are re-executed until the quality evaluation of the output solution reaches the preset threshold.
[0022] Preferably, the pre-set set of standard scenarios includes at least audit scenarios for official expenses and audit scenarios for engineering projects.
[0023] Preferably, the graph database adopts an attribute graph model, and the vector database adopts an index structure based on cosine similarity.
[0024] Preferably, the incremental update adopts a difference-driven approach: only local regions related to the incremental data in the graph database and the vector database are updated, and the coupling impact range caused by the update is identified, and the affected adjacent nodes or edges are updated synchronously.
[0025] Preferably, in step S5, when the large language model performs knowledge reconstruction and reasoning generation, a historical trend awareness mechanism is also introduced: the user adoption rate of the audit path planning schemes generated in the last three rounds for the same audit scenario type is recorded, and denoted as the _th _ ... Round of adoption rate , No. Round of adoption rate , No. Round of adoption rate ;
[0026] Calculate the first Compared to the first wheel The increase of the wheel ,like ,
[0027] Calculate the first Compared to the first wheel The increase of the wheel ,
[0028] like If the current adjustment direction is saturated, the search strategy or parameters will be switched.
[0029] If the cumulative number of rounds generated for the same audit scenario type is less than 3, the historical trend perception mechanism is skipped.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] This invention employs a two-tiered progressive processing mechanism: for standard audit scenarios, it directly matches pre-stored templates for a quick response; for complex or ambiguous scenarios, it triggers a deep fusion of graph retrieval and vector retrieval, combined with a conflict arbitration mechanism, to extract highly reliable knowledge from structured entity relationships and unstructured historical cases, thereby effectively solving the problems of missed judgments, misjudgments, and unreliable single information sources in traditional methods when boundaries are ambiguous.
[0032] Based on this, the present invention enables the knowledge graph to continuously evolve with the updates of systems and the accumulation of cases through dynamic maintenance of data sources and timeliness detection, avoiding performance degradation caused by relying on outdated models; at the same time, it introduces result feedback-driven closed-loop optimization, which automatically locates defective links and adjusts retrieval strategies or generation parameters based on users' quality evaluation of the output solution, realizing the system's self-correction.
[0033] In addition, the historical trend perception mechanism can predict changes in the marginal returns of the adjustment direction in advance, proactively switch optimization strategies, and accelerate convergence to the optimal audit path.
[0034] In summary, this invention not only significantly improves the accuracy, consistency, and interpretability of audit path planning, but also endows the system with adaptive and efficient iterative capabilities during long-term operation, reduces auditors' reliance on personal experience and historical cases, and ensures the standardization and professionalism of audit work. Attached Figure Description
[0035] Figure 1 This is a flowchart illustrating the audit path planning and intelligent recommendation method based on a large model and graph database in Embodiment 1 of the present invention. Detailed Implementation
[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0037] Example 1: This example provides an audit path planning and intelligent recommendation method based on a large model and graph database. The following is combined with... Figure 1 Detailed explanation.
[0038] Step S1: Construct an audit knowledge graph
[0039] First, audit entities and relationships between entities are extracted from multi-source heterogeneous data. Multi-source heterogeneous data includes, but is not limited to: historical audit reports (PDF or Word format), audit ledgers (Excel or database tables), policy documents (policy and regulation documents), work prompts (internal minutes), etc.; the extraction method uses a large language model (e.g., ChatGLM3-6B) combined with regular expressions and named entity recognition (NER) models.
[0040] Entity types include: audit project, auditing entity, auditee, accounting subject, legal provisions, audit procedures, risk points, etc.
[0041] Relationship types include: belonging, referencing, associating, violating, containing, and preceding.
[0042] After extraction, entities are treated as nodes and relationships as edges, and stored in a graph database (such as Neo4j) using an attribute graph model. At the same time, complete case text fragments from historical audit reports and audit ledgers are cleaned and converted into fixed-dimensional vectors (e.g., 768-dimensional) using a pre-trained embedding model (such as text2vec-large-chinese), stored in a vector database (such as Milvus), and indexed for fast similarity retrieval.
[0043] Step S2: Obtain the audit business description text input by the user.
[0044] Users can input audit business descriptions in natural language through a web interface or API, such as: "Please plan the audit path for whether there are any illegal subcontractings in a certain engineering project."
[0045] Step S3: Large Language Model Intent Recognition
[0046] The user's input text is fed into a large language model, and intent recognition is performed in conjunction with preset prompt templates.
[0047] The large language model output determines the audit scenario type, such as "engineering project audit"; for scenarios that cannot be classified into the above standard scenarios, the output is "other".
[0048] This invention employs a two-layer progressive processing mechanism.
[0049] First-level processing: The system pre-constructs a set of standard scenarios, denoted as... For the scene types identified in step S3 ,like If the template is not found, it will directly match the pre-stored audit path template. For example, the template for "Audit of Official Reception Expenses" includes: audit focus (official reception expenses, official vehicle purchase and operation expenses, and expenses for official overseas trips), audit key points (authenticity of vouchers, compliance with standards), recommended audit methods (verification of budget approvals, random checks of vouchers), key data sources (financial accounting system, invoice system), referenced institutional provisions (Regulations on the Management of Domestic Official Receptions by Party and Government Organs), and similar project cases (Audit of Official Reception Expenses in a certain city in 2023). The process ends after the template is output.
[0050] Second-level processing: If If the object is a fuzzy object or a complex scene that does not belong to the preset standard scene set, then the second layer of processing is triggered, and steps S4 (parallel retrieval) and S5 (large model generation) are executed.
[0051] Step S4: Parallel Search
[0052] Based on the identified scene type Construct graph database query statements and vector retrieval query statements.
[0053] Graph database retrieval: Using the Cypher query language to retrieve data related to... Related entities and their neighborhood subgraphs; for example, if For "Engineering Project Audit", the query path is: The search result is a subgraph. , containing a set of nodes Sum of edges .
[0054] Vector similarity retrieval: based on user-input audit business description text As a query vector, cosine similarity is calculated in the vector database to retrieve the most similar vector. Case fragments ( (Take a value of 5 or 10); the similarity calculation formula is:
[0055]
[0056] in The embedding vector of the query text. For the first The embedding vectors of each case fragment; the retrieval results are denoted as... .
[0057] Heterogeneous Information Source Fusion and Conflict Arbitration
[0058] In obtaining graph retrieval subgraphs Sum of vector search results Then, consistency checks and conflict arbitration are performed.
[0059] First, from Extract key audit entities (For example, the risks involved, regulatory provisions), from Extract the corresponding risk assessment (For example, the conclusion in the case of determining whether there is illegal subcontracting); if the two are consistent in their judgment of key entities or risks (through semantic similarity thresholds) If the judgment is made, the two will be weighted and merged into a retrieval context.
[0060] If a contradiction exists (e.g., the graph database shows a legal provision as valid, while the vector case cites the provision but states it has been repealed), then conflict arbitration is initiated; the first confidence level of the graph database search results is calculated separately. The second confidence level of vector retrieval results .
[0061] First credibility Calculation:
[0062]
[0063] in:
[0064] Subgraph The average weight of all edges; the weight of each edge. Composed of two entities and The co-occurrence frequency in historical audit data is normalized. The historical audit data includes all audit cases, audit ledgers, and entity co-occurrence records in historical audit reports accumulated since the system started operating.
[0065] Specifically, record all entity pairs in historical data. The co-occurrence frequency is Then the normalized weights are:
[0066]
[0067] in and These are the minimum and maximum co-occurrence frequencies of all entity pairs in the historical data, respectively; if no historical data is available, the default value is used. .
[0068] The score is assigned to the issuing agency of the policy document in the sub-graph. The issuing agency is assigned a score based on its administrative level: 4 for national level, 3 for provincial level, 2 for municipal level, and 1 for county level. If there are multiple policy documents in the sub-graph, the highest score is taken.
[0069] and Let be the weighting coefficient, satisfying In this embodiment, , .
[0070] Second credibility Calculation:
[0071] Based on the reciprocal of the number of days from the current date of the most similar case in the vector retrieval results, and considering the similarity of that case:
[0072]
[0073] in , This is the number of days since the date the case occurred; add 1 to prevent division by zero.
[0074] Compare and The results with the highest credibility are selected as the final search results, and the results that are rejected are marked as questionable information and stored in the log for manual review.
[0075] Step S5: Large-scale model knowledge reconstruction and reasoning generation
[0076] Using the search results (subgraph data + case fragments) finally determined in step S4 as context, input into the large language model, and with the help of carefully designed prompt templates, generate a structured audit path planning scheme.
[0077] The system outputs solutions in JSON format or natural language text from large language models. Simultaneously, it supports displaying the generated solutions as visual graphs, i.e., generating force-directed graphs using query results from graph databases (e.g., using ECharts or D3.js), highlighting key path nodes.
[0078] Historical trend perception mechanism:
[0079] In step S5, the system further introduces a historical trend perception mechanism to improve the efficiency of iterative optimization;
[0080] Record the user adoption rate of the most recent 3 rounds (i.e., the most recent 3 audit path planning schemes generated for the same type of scenario);
[0081] When the cumulative number of generated rounds for the same audit scenario type is less than 3, the historical trend awareness mechanism is skipped and the default retrieval strategy is used directly.
[0082] User adoption rate is defined as the percentage of users who adopt each recommendation (such as audit points and methods) in the generated solution, calculated based on the percentage ultimately used after users click "adopt" or manually correct it.
[0083] Record No. Round adoption rate ( First, determine the increase in the previous round (round 2) compared to the previous two rounds (round 1):
[0084]
[0085] when At that time, calculate the increase in the current round (round 3) compared to the previous round:
[0086]
[0087] like If the current adjustment direction is saturated, it means that the benefits of continuing to adjust the retrieval strategy or generation parameters in the same direction have significantly decreased. At this time, the system automatically switches the retrieval strategy, for example, from graph database priority to vector database priority, or adjusts the temperature coefficient in the generation parameters from 0.7 to 0.4, thereby escaping the local optimum.
[0088] This mechanism ensures that the system changes direction in advance when marginal returns decrease, avoiding ineffective iterations and accelerating convergence to the optimal path.
[0089] Dynamic maintenance and quality monitoring of data sources:
[0090] During system operation, step S6 is continuously executed: dynamic maintenance and quality monitoring of data sources.
[0091] Incremental update:
[0092] After each user interaction (entering an audit business description, providing feedback on adoption results), the system records the following information as incremental data in a temporary buffer:
[0093] User-inputted description of audit procedures;
[0094] The system generates an audit path planning scheme;
[0095] User feedback on the adoption of the solution (e.g., which audit points were adopted and which were rejected).
[0096] Incremental updates are triggered according to a preset time period (e.g., 2:00 AM daily); the update adopts a difference-driven approach, that is, only the local area related to the incremental data is updated.
[0097] Specifically, for graph databases: extract newly emerging entities and relationships from incremental data, insert only the newly added nodes and edges, and update the co-occurrence frequency weights of affected nodes; for example, if user feedback confirms a new risk point's association with existing regulations, add an edge to the graph and recalculate the normalized weight of that edge. (Based on the extended historical co-occurrence frequency); at the same time, identify the coupling impact range brought about by the edge update: all adjacent nodes with a path length of no more than 2 are considered to be affected, and the statistical information of these adjacent nodes is updated synchronously (such as recalculating degree centrality).
[0098] For vector databases: Vectorize new audit case text fragments and insert them, then rebuild the index (incrementally if using the HNSW algorithm).
[0099] Timeliness testing:
[0100] The system periodically (e.g., every Monday) checks the timeliness of stored policy documents. The check method involves obtaining the latest list of valid regulations through web crawling or integration with the official regulatory database API. For nodes marked as "reference policy clauses" in the graph database, their version numbers are compared with the latest regulations. If a clause is found to have been repealed or revised, the node's weight is multiplied by a decay factor. (De-weighting) and mark "may be obsolete" in the node attributes; if two consecutive checks confirm obsolescence, the node and its associated edges are moved to the archive area (removal operation) and will no longer be included in the retrieval.
[0101] Results-feedback driven closed-loop optimization:
[0102] Step S7 performs closed-loop optimization after each output audit path planning scheme.
[0103] Users can rate the quality of a solution quantitatively (e.g., 1-5 points) or implicitly through subsequent user behavior (e.g., whether they ask further questions or fully adopt the solution); the system sets a preset threshold. Points (out of 5); if the user rating is lower than Then, based on the specific information of the quality defect, the process is directed to either the retrieval stage in step S4 or the generation stage in step S5:
[0104] The defect is "missing key data source": locate the retrieval stage and increase the weight of key data source related fields in the vector retrieval results (for example, increase the retrieval weight of the keyword "invoice" by 20%).
[0105] The defect is "impractical audit methods": locate the generation stage, adjust the prompt template of the large language model, and require priority to be given to standard audit procedures (for example, add the constraint "only methods that comply with the National Audit Standards are recommended").
[0106] After adjustment, the system re-executes steps S4 and S5 to generate a new solution; this process is repeated until the user rating reaches the threshold or exceeds the maximum number of retries (e.g., 3 times).
[0107] This closed-loop optimization mechanism enables the system to self-correct and gradually improve output quality.
[0108] The embodiments of the present invention are given for the purposes of illustration and description. Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Any changes, modifications, substitutions and variations made by those skilled in the art to the above embodiments within the scope of the present invention should be included within the protection scope of the present invention.
Claims
1. An audit path planning and intelligent recommendation method based on large models and graph databases, characterized in that, Includes the following steps: Step S1: Construct an audit knowledge graph, which specifically includes: extracting audit entities and relationships between entities from multi-source heterogeneous data, storing the entities and relationships in a graph database in a graph structure, and vectorizing audit case text fragments and storing them in a vector database. Step S2: Obtain the audit business description text input by the user; Step S3: Use a large language model to perform intent recognition on the audit business description text to determine the audit scenario type; Step S4: Based on the audit scenario type, perform graph database retrieval and vector similarity retrieval in parallel to obtain retrieval results, which include subgraph data and case similarity fragments; Step S5: Using the search results as context, knowledge reconstruction and reasoning generation are performed through a large language model to output a structured audit path planning scheme; The method employs a two-layer progressive processing mechanism: after step S3, if the identified audit scenario type belongs to a preset standard scenario set, the pre-stored audit path template is directly matched as the output, skipping steps S4 and S5; if it does not belong to the standard scenario set, steps S4 and S5 are executed.
2. The method according to claim 1, characterized in that, The multi-source heterogeneous data includes at least two of the following: historical audit reports, audit ledgers, policy documents, and work notices.
3. The method according to claim 1, characterized in that, The structured audit path planning scheme includes the following fields: audit focus, audit key points, recommended audit methods, key data sources, reference system clauses and similar project cases, and supports visual graph display.
4. The method according to claim 1, characterized in that, It also includes step S6: dynamic maintenance and quality monitoring of data sources, which specifically includes: during system operation, using the audit business description entered by the user each time, the generated audit path planning scheme, and the user's adoption or feedback results of the scheme as incremental data, and incrementally updating the graph database and the vector database according to a preset time period; at the same time, performing timeliness detection on the stored policy documents, and demoting or removing the policy clauses that have been repealed or updated.
5. The method according to claim 1, characterized in that, After performing graph database retrieval and vector similarity retrieval in parallel in step S4, the process also includes heterogeneous information source fusion and conflict arbitration steps: A consistency check is performed on the subgraph data retrieved from the graph database and the similar fragments of cases retrieved by vector similarity retrieval. If the two are consistent in their assessment of key audit entities or risks, the weighted average of the two will be used as the search result. If there is a contradiction between the two, conflict arbitration is initiated: the first credibility of the graph database retrieval results and the second credibility of the vector retrieval results are calculated separately. The first credibility is a weighted combination of the average weight value of the edges in the graph database and the grade score of the issuing agency of the corresponding institutional document. If multiple institutional documents are involved, the highest grade score is taken. The weight value of the edges is obtained by normalizing the co-occurrence frequency of the two entities in historical audit data. The grade of the issuing agency is preset to 4, 3, 2, and 1 for national, provincial, municipal, and county levels, respectively. The second credibility is determined based on the decay score of the number of days between the occurrence time of the case and the current date in the vector retrieval results. This decay score is proportional to the reciprocal of the number of days plus 1. The first credibility and the second credibility are compared, and the one with the higher credibility is taken as the final retrieval result. The results that are abandoned are marked as questionable information.
6. The method according to claim 1, characterized in that, It also includes step S7: closed-loop optimization driven by result feedback, specifically including: After outputting the audit path planning scheme in step S5, obtain the user's quality evaluation of the scheme; If the quality evaluation is lower than the preset threshold, the search link in step S4 or the generation link in step S5 is located based on the specific information of the quality defect. The search weight or generation parameters are adjusted, and steps S4 and S5 are re-executed until the quality evaluation of the output solution reaches the preset threshold.
7. The method according to claim 1, characterized in that, The pre-set set of standard scenarios includes at least audit scenarios for official expenses and audit scenarios for engineering projects.
8. The method according to claim 1, characterized in that, The graph database uses an attribute graph model, and the vector database uses an index structure based on cosine similarity.
9. The method according to claim 4, characterized in that, The incremental update adopts a difference-driven approach: only local regions related to the incremental data in the graph database and the vector database are updated, and the coupling impact range caused by the update is identified, and the affected adjacent nodes or edges are updated synchronously.
10. The method according to claim 1, characterized in that, In step S5, when the large language model performs knowledge reconstruction and reasoning generation, a historical trend awareness mechanism is also introduced: the user adoption rate of the audit path planning schemes generated in the last three rounds for the same audit scenario type is recorded, and denoted as the 1st, 2nd, 3rd, and 4th respectively. Round of adoption rate , No. Round of adoption rate , No. Round of adoption rate ; Calculate the first Compared to the first wheel The increase of the wheel ,like , Calculate the first Compared to the first wheel The increase of the wheel , like If the current adjustment direction is saturated, the search strategy or parameters will be switched. If the cumulative number of rounds generated for the same audit scenario type is less than 3, the historical trend perception mechanism is skipped.