Audit system question and answer method and system based on retrieval enhancement and instruction fine tuning

By applying search enhancement and instruction fine-tuning technologies in the field of audit system, a high-quality instruction data set and fine-tuning large models are built, which solves the problem of low Q&A performance and accuracy of the audit system, and achieves efficient and accurate Q&A capabilities of the audit system.

CN119938828APending Publication Date: 2025-05-06STATE GRID TIANJIN ELECTRIC POWER COMPANY +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411900347.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

It is difficult for the existing technology to quickly and accurately organize instruction data sets in the audit system field, resulting in low performance and accuracy of audit system Q&A.

Method used

Using methods based on search enhancement and instruction fine-tuning, we build a knowledge base, matched search and intelligent voting mechanism through data cleaning, text chunking, and vector embedding models to build a high-quality audit system instruction dataset, and fine-tune the big model.

Benefits of technology

It significantly improves the performance and accuracy of the audit system Q&A, reduces the time and labor costs of data processing and labeling, and enhances the coverage and accuracy of the data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938828A_ABST
    Figure CN119938828A_ABST
Patent Text Reader

Abstract

The invention relates to an audit system question and answer method and system based on retrieval enhancement and instruction fine tuning. The method comprises the following steps: step 1, obtaining a pure audit system data set; 2, generating a related audit instruction question and answer data set, and constructing an audit system knowledge base by using a vector embedding model; step 3, obtaining a corresponding audit system text block; 4, obtaining an audit instruction question and answer data set subjected to preliminary screening; 5, obtaining a final audit instruction question and answer data set; step 6, performing instruction fine tuning on the large model by using the final audit instruction question and answer data set to obtain an audit system large model; and step 7, testing and evaluating the model performance by using the auditing system large model, and obtaining an answer related to the auditing system question of the user according to the auditing system question input by the user. The method has higher practicability, accuracy and stability, and provides reliable support for development of an intelligent auditing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of application of retrieval enhancement and instruction fine-tuning technology to audit system technology, and relates to an audit system question and answer method and system, in particular, an audit system question and answer method and system based on retrieval enhancement and instruction fine-tuning. Background Art

[0002] In the audit system question-and-answer session of the general large-scale model, the general large-scale model often lacks knowledge of the audit system when answering questions about the audit system. The answers it generates are inconsistent with the audit system facts and are completely irrelevant to the audit question. Therefore, designing a large-scale audit system model is very critical.

[0003] The premise for designing an efficient and reliable audit system model is to build a high-quality dataset of audit system instructions. Traditional dataset construction methods rely heavily on manual annotation, a process that is often time-consuming and labor-intensive.

[0004] Specifically, relevant texts need to be collected from a variety of sources (such as audit regulations, policy documents, case studies, and research papers), and then cleaned, screened, and classified by a team of domain experts to ensure the professionalism and accuracy of the data. Experts need to annotate the data in detail based on specific task requirements and clarify the logical relationship between input questions and output answers. This process usually requires repeated discussions and reviews to unify the annotation rules, reduce subjective bias, and ensure that the data coverage is broad enough to cover the multi-dimensional knowledge of the audit system. Due to the high professional requirements in the audit field, multiple experts need to be mobilized to work together, which further increases time and labor costs. The final dataset needs to undergo rigorous quality assessment and multiple rounds of revisions to meet the high standards required for fine-tuning and testing of large models. Although the traditional process is rigorous, it is less efficient.

[0005] To sum up, how to quickly and accurately organize the instruction dataset in the audit system field and significantly improve the performance and accuracy of audit system question answering is an urgent problem that needs to be solved.

[0006] Therefore, in order to solve the above problems, the present invention proposes an audit system question-answering method and system based on retrieval enhancement and instruction fine-tuning.

[0007] After searching, no public documents of the prior art that are identical or similar to the present invention were found. Summary of the Invention

[0008] The purpose of the present invention is to overcome the shortcomings of the existing technology and propose an audit system question and answer method and system based on retrieval enhancement and instruction fine-tuning. By combining retrieval enhancement and instruction fine-tuning technologies, a high-quality audit system field instruction data set is designed, and the large model is fine-tuned to achieve an efficient and high-accuracy audit system question and answer method and system.

[0009] The present invention solves the practical problem by adopting the following technical solutions:

[0010] An audit system question answering method based on retrieval enhancement and instruction fine-tuning includes the following steps:

[0011] Step 1: Clean the acquired audit system text dataset to obtain a pure audit system dataset;

[0012] Step 2: Use the clean audit system dataset to segment the text, use the large model to generate relevant audit instruction question and answer datasets through prompt engineering, and use the vector embedding model to build the audit system knowledge base;

[0013] Step 3: In the audit instruction question-answering dataset, the audit questions are matched with the audit system knowledge base to obtain the corresponding audit system text blocks;

[0014] Step 4: Calculate the similarity between the answers in the audit instruction question-and-answer dataset and the matching retrieved audit system text blocks. If the similarity is greater than the threshold, retain the audit system question-and-answer pair to obtain the preliminarily screened audit instruction question-and-answer dataset.

[0015] Step 5: Use the audit instruction question and answer dataset initially screened in step 4 to conduct an intelligent voting mechanism using the big model. If the big model finds a consensus and the number of votes is greater than half, the audit system question and answer pair is retained to obtain the final audit instruction question and answer dataset.

[0016] Step 6: Use the final audit instruction question-answering dataset to fine-tune the large model to obtain the audit system large model;

[0017] Step 7: Use the audit system model to test and evaluate the model performance, and obtain answers related to the user's audit system questions based on the audit system questions input by the user.

[0018] Moreover, the specific implementation method of step 1 is:

[0019] The obtained audit system text dataset is cleaned, word segmentation is performed, and special characters, tables, and graphs are removed. Finally, the audit system text only needs to retain all the system clauses to obtain a pure audit system dataset.

[0020] Moreover, the specific implementation method of step 2 is:

[0021] Using a clean audit system dataset, we pre-processed and formatted it into unstructured documents, and segmented the text into chunks based on fixed length.

[0022] Input the audit system text dataset D. The specific formula is as follows:

[0023] D={d1,d2,d3,…,d n},i={1,2,3,…n}

[0024] Among them, D is the audit system text dataset, d i It is an independent audit system document.

[0025] Divide the audit system text into blocks and construct the audit system text block dataset B. Set the block size to L. Each text block b j ,satisfy:

[0026] B={b1,b2,b3,…,b m},j={1,2,3,…m},b j =d i [l:k],kl≤L

[0027] Among them, B is the audit system text block dataset, b j It is an independent audit system text block, and l and k are the start and end indexes of the audit system text block.

[0028] Based on the constructed audit system text block dataset B, we design a prompt template P for extracting the audit instruction question and answer dataset. Using the large model, we apply the prompt template to each block content to generate audit instructions and question and answer pairs related to the text.

[0029] The specific formula is as follows:

[0030] (I j ,Q j ,A j )=GPT large (P(b j ))

[0031] Among them, I j is the audit instruction, Q j It is a problem of audit system. j The answer is GPT large is the big model, P is the prompt template, b j It is an independent audit system text block.

[0032] Combine the generated audit system instruction question and answer pairs to obtain the audit instruction question and answer dataset D QA .

[0033] DQA ={(I j ,Q j ,A j )|j={1,2,3,···m}}

[0034] Among them, D QA It is the audit instruction question answering dataset, I j is the audit instruction, Q j It is a problem of audit system. j is the answer.

[0035] For the constructed audit system text block dataset B, semantic vector modeling is performed to build the audit system knowledge base DB audit .

[0036] DB audit =Embed model (B)

[0037] Among them, DB audit Is the audit system knowledge base, Embed model is the vector embedding model, and B is the audit system text block dataset.

[0038] Moreover, the specific implementation method of step 3 is:

[0039] In the audit instruction question-answering dataset, audit questions are processed using vector retrieval technology. First, the questions are converted into fixed-dimensional vector representations using a vector embedding model.

[0040] The specific formula is as follows:

[0041] Q vec =Embed model (Q)

[0042] Among them, Q vec For the vectorized representation of the audit system problem, Embed model is the vector embedding model, and Q is the input audit system problem.

[0043] The text cosine similarity is used to calculate the similarity between the question vector and the text block vector embedded in the audit system knowledge base.

[0044] The specific formula is as follows:

[0045]

[0046] Among them, Sim(Q vec ,DB audit,i ) is the audit system problem vector Q vec And each text block DB in the audit system knowledge base audit,i The similarity value of .

[0047] Next, the calculation results are sorted from high to low by similarity, and the audit system text block that is most relevant to the problem is selected. Finally, the matching text block is output. The specific formula is as follows:

[0048]

[0049] Among them, k match is the most similar audit system text block, argmaxSim is the function that returns the highest similarity, Q vec is the audit system problem vector, DB audit,i It is every text block in the audit system knowledge base.

[0050] Finally, the audit instruction question answering dataset D QA The matching text blocks are joined accordingly.

[0051] Moreover, the specific implementation method of step 4 is:

[0052] The similarity between the answers in the audit instruction question-answering dataset and the matching retrieved audit system text blocks is calculated, and cosine similarity is used as the similarity measurement indicator.

[0053] First, the answer and the retrieved audit system text blocks are vectorized separately. The specific formula is as follows:

[0054] A vec =Embed model (A)

[0055] k matchvec =Embed model (k match )

[0056] Among them, A vec is the vectorized representation of the answer extracted by the large model, k matchvec Embed is the vectorized representation of the retrieved audit system text block. model is the vector embedding model, k match It is the most similar audit system text block.

[0057] Then calculate the cosine similarity between the two. The specific formula is as follows:

[0058]

[0059] Among them, Sim(A vec ,k matchvec ) is the similarity between the answer extracted by the large model and the retrieved audit system text block, A vec is the vectorized representation of the answer extracted by the large model, k matchvecVectorized representation of the audit system text block for retrieval.

[0060] Compare with the set threshold Threshold. If the similarity is greater than the threshold, the audit system question and answer pair is retained, and finally the preliminary screened audit instruction question and answer dataset Filtered D is obtained. QA ;

[0061] Filtered D QA ={D QA |Sim(A vec ,k matchvec )>Threshold}

[0062] Among them, Filtered D QA It is a preliminary screened audit instruction question and answer dataset, Sim(A vec ,k matchvec ) is the similarity between the answer extracted by the large model and the retrieved audit system text block, and Threshold is the set threshold.

[0063] Moreover, the specific implementation method of step 5 is:

[0064] In the preliminary screening of the audit instruction question and answer dataset Filtered D QA In the process, a large model is used to set up multiple voting agents, and high-quality question-answer pairs are screened through the voting mechanism;

[0065] First, for each question-answer pair, the corresponding audit system text block is retrieved and input into the large model for consistency judgment. Each voting agent independently outputs an evaluation result of "consistent" or "inconsistent". The voting results of all models are counted. If the number of "consistent" votes exceeds half of the total number of models, the question-answer pair is retained; otherwise, it is discarded.

[0066] Through this mechanism, we ensure that the retained question-answer pairs are highly consistent in the large model, thereby constructing the final high-quality audit instruction question-answer dataset;

[0067] The specific formula is as follows:

[0068]

[0069] Where n is the number of voting agents, λ is the indicator function, when R = 'consistent', λ = 1, otherwise λ = 0, and Consistency represents the number of consistent votes.

[0070] Repeat the above steps for all audit instruction question-answer pairs to select the final high-quality question-answer dataset FinalD QA , the specific formula is as follows:

[0071] Final D QA ={Filtered D QA |Consustency>n / 2}

[0072] Among them, Final D QA It is the final high-quality question-answering dataset. Consustency>n / 2 means that more than half of the votes are considered valid.

[0073] Moreover, the specific implementation method of step 6 is:

[0074] Using the high-quality audit instruction question answering dataset Final D after final screening QA The specific implementation method for fine-tuning instructions on the large model is as follows: the dataset is divided into a training set and a validation set. The training set is used for model parameter optimization, and the validation set is used for model performance evaluation. Using instruction fine-tuning technology, question-answer pairs are used as fine-tuning samples in input-output format. The general large model is loaded, and an appropriate learning rate and optimizer are set for gradual fine-tuning. During the fine-tuning process, the model learns the semantic understanding of audit instructions and the ability to generate answers. By periodically evaluating accuracy, consistency, and generation quality on the validation set, hyperparameters are adjusted to optimize model performance, ultimately obtaining an audit system large model that can accurately interpret and generate answers related to audit instructions.

[0075] Furthermore, the specific implementation method for step 7 is as follows: user questions are input into the model, and targeted answers are generated through semantic matching and generative answering techniques. Once the answers are complete, rule verification and semantic analysis are used to optimize the answers to ensure logic and accuracy. Finally, a test set is constructed to evaluate model performance using metrics such as accuracy, recall, F1 score, similarity, and exact match, and then compared with human answers. Based on the evaluation results, the model is continuously optimized to enhance its ability to handle complex questions.

[0076] An audit system question answering system based on retrieval enhancement and instruction fine-tuning, including:

[0077] A pure audit system data set acquisition module is used to clean the acquired audit system text data set to obtain a pure audit system data set;

[0078] The audit system knowledge base construction module is used to use the clean audit system dataset to perform text segmentation, use the large model to generate relevant audit instruction question and answer datasets through prompt engineering, and use the vector embedding model to build the audit system knowledge base;

[0079] The audit system text block module is used to match and retrieve the corresponding audit system text blocks in the audit instruction question and answer dataset by using the audit questions and the audit system knowledge base;

[0080] The module for obtaining the preliminary audit instruction question and answer dataset is used to calculate the similarity between the answers in the audit instruction question and answer dataset and the matching retrieved audit system text blocks. If the similarity is greater than the threshold, the audit system question and answer pair is retained to obtain the preliminary screened audit instruction question and answer dataset.

[0081] The final audit instruction question and answer dataset acquisition module uses the preliminary screened audit instruction question and answer dataset to conduct an intelligent voting mechanism using a large model. If the large model finds a consensus and the number of votes is greater than half, the audit system question and answer pair is retained to obtain the final audit instruction question and answer dataset;

[0082] The audit system large model construction module is used to use the final audit instruction question and answer dataset to fine-tune the large model to obtain the audit system large model;

[0083] The answer output module is used to test and evaluate the performance of the audit system model using the audit system model, and obtain answers related to the user's audit system questions based on the audit system questions input by the user.

[0084] Advantages and beneficial effects of the present invention:

[0085] 1. This invention can automatically extract audit system instruction datasets from massive, multi-source audit system text data. This significantly improves the efficiency and quality of data processing and annotation, avoids the tediousness and errors associated with traditional manual annotation, and significantly saves time and labor costs.

[0086] 2. This invention utilizes a two-stage "retrieval-enhanced semantic matching - agent voting" screening mechanism to accurately identify and select high-quality audit system instruction datasets. This screening mechanism combines the advantages of semantic understanding and swarm intelligence, improving the coverage and accuracy of the audit system dataset while also enhancing its robustness and authority.

[0087] 3. This invention leverages a high-quality dataset of audit system instructions to fine-tune the large-scale model's enhanced audit system question-answering capabilities, effectively enhancing the model's question-answering and knowledge generalization capabilities in the audit system domain. Compared to traditional methods, this invention significantly optimizes the large-scale model's performance in complex audit tasks, enhancing its practicality, accuracy, and stability, and providing reliable support for the development of intelligent audit technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] Figure 1 It is the overall processing flow chart of the present invention;

[0089] Figure 2 It is a flow chart of the two-stage "retrieval enhanced semantic matching-intelligent agent voting decision" screening mechanism of steps 4 and 5 of the present invention. DETAILED DESCRIPTION

[0090] The embodiments of the present invention are further described below in conjunction with the accompanying drawings:

[0091] An audit system question answering method based on retrieval enhancement and instruction fine-tuning, such as Figure 1 and Figure 2 As shown, the following steps are included:

[0092] Step 1: Clean the acquired audit system text dataset to obtain a pure audit system text dataset;

[0093] The specific implementation method of step 1 is:

[0094] The obtained audit system text dataset is cleaned, word segmentation is performed, special characters, tables, charts, etc. are removed, and finally the audit system text dataset only needs to retain all the system clauses to obtain a pure audit system text dataset.

[0095] In this embodiment, the acquired audit system text dataset is subjected to comprehensive data cleaning and preprocessing. In the first step, based on the characteristics of the audit field, word segmentation processing is performed using a word segmentation tool suitable for processing audit texts, such as NLTK or the jieba word segmentation tool designed specifically for Chinese, to ensure that audit proper nouns and terms can be accurately identified and segmented. In the second step, special characters, punctuation marks, and irrelevant content, such as web page codes, redundant formatting symbols, figures, and tables, are removed through regular expressions or other string processing methods to extract a pure audit system text dataset. After this step, it is guaranteed that the generated audit system text data has a high degree of standardization and consistency.

[0096] Step 2: The clean audit system text dataset obtained in step 1 is segmented into text blocks, and the large model is used to generate relevant audit instruction question and answer datasets through prompt engineering, and the audit system knowledge base is constructed using the vector embedding model;

[0097] The specific implementation method of step 2 is:

[0098] Using a pure audit system dataset, we pre-processed and formatted it into unstructured documents, then divided the text into blocks based on fixed lengths to avoid semantic fragmentation and ensure the integrity of each block. The specific formula is as follows:

[0099] D={d1,d2,d3,…,d n},i={1,2,3,…n}

[0100] Among them, D is the audit system text dataset, d i It is an independent audit system document.

[0101] Divide the audit system text into blocks and construct the audit system text block dataset B. Set the block size to L. Each text block b j ,satisfy:

[0102] B={b1,b2,b3,…,b m},j={1,2,3,…m},b j =d i [l:k],kl≤L

[0103] Among them, B is the audit system text block dataset, b j It is an independent audit system text block, and l and k are the start and end indexes of the audit system text block.

[0104] Based on the constructed audit system text block dataset B, a prompt template P is designed to extract the audit instruction question and answer dataset. Using a large model (such as ChatGPT), the prompt template is applied to each block content to generate audit instructions and question and answer pairs related to the text;

[0105] The specific formula is as follows:

[0106] (I j ,Q j ,A j )=GPT large (P(b j )),j={1,2,3,…m}

[0107] Among them, I j is the audit instruction, Q j It is a problem of audit system. j The answer is GPT large is the big model, P is the prompt template, b j It is an independent audit system text block.

[0108] Combine the generated audit system instruction question and answer pairs to obtain the audit instruction question and answer dataset D QA .

[0109] D QA ={(I j ,Q j ,A j )|j={1,2,3,···m}}

[0110] Among them, D QA It is the audit instruction question answering dataset, I j is the audit instruction, Q j It is a problem of audit system. j is the answer.

[0111] For the constructed audit system text block dataset B, semantic vector modeling is performed to build the audit system knowledge base DB audit .

[0112] DB audit =Embed model (B)

[0113] Among them, DB audit Is the audit system knowledge base, Embed model is the vector embedding model, and B is the audit system text block dataset.

[0114] In this example, we use a clean audit system dataset to load the raw data. This can be in CSV, JSON, or other structured file formats. If the data is distributed across multiple columns, the required content can be merged into a single column. The formatted text is saved as an unstructured file such as TXT, leaving appropriate white space between each paragraph to ensure the document is neat and readable. Enter the audit system text dataset D. The specific formula is as follows:

[0115] D={d1,d2,d3,…,d n},i={1,2,3,…n}

[0116] Among them, D is the audit system text dataset, each d i It is an independent audit system document.

[0117] The text is segmented according to a fixed length (e.g., 200-300 words). In this process, the text_splitter in the Langchain framework is used for segmentation. When segmenting, simple natural language detection rules such as setting the number of overlapping words in the segmentation are used to ensure the logical coherence and information integrity of each text block. Sequential identifiers or metadata such as block numbers or original text location indexes are added to each audit system text block to facilitate subsequent tracing and analysis. Construct the audit system text block dataset B, set the block size to L, and each text block b j ,satisfy:

[0118] B={b1,b2,b3,…,b m},j={1,2,3,…m},b j =d i [l:k],kl≤L

[0119] Among them, B is the audit system text block dataset, b j It is an independent audit system text block, and l and k are the start and end indexes of the audit system text block.

[0120] For the constructed audit system text block dataset B, we design a prompt template P for extracting the audit instruction question and answer dataset. Using a large model (such as ChatGPT), we apply the prompt template to each block content to generate audit instructions and question and answer pairs related to the text. The specific formula is as follows:

[0121] (I j ,Q j ,A j )=GPT large (P(b j )),j={1,2,3,…m}

[0122] Among them, I j is the audit instruction, Q j It is a problem of audit system. j The answer is GPT large is the big model, P is the prompt template, b j It is an independent audit system text block.

[0123] The prompt template P is as follows:

[0124] ”'

[0125] #01You are an expert in processing audit system question-answering datasets.

[0126] #02Your task is to generate corresponding question-answer pairs based on the audit system content I provide.

[0127] #03The answer should be comprehensive, use more of my information, and have richer content.

[0128] #04 You must generate according to my three question-answer pairs example format:

[0129] """

[0130] {"instruction":"instruction","query":"question","answer":"answer"},{"instruction":"instruction","query":"question","answer":"answer"},{"instruction":"instruction","query":"question","answer":"answer"}

[0131] #05My content is as follows:

[0132] """

[0133] {{Replace this with your content}}

[0134] """

[0135] ”'

[0136] Combine the generated audit system instruction question and answer pairs to obtain the audit instruction question and answer dataset D QA .

[0137] D QA ={(I j ,Q j ,A j )|j={1,2,3,···m}}

[0138] Among them, D QA It is the audit instruction question answering dataset, I j is the audit instruction, Q j It is a problem of audit system. j is the answer.

[0139] For the constructed audit system text block dataset B, a text vector embedding model (such as bge-large-zh) is used to semantically vectorize the text blocks and convert them into high-dimensional vector representations. To improve retrieval efficiency, a vector database (such as FAISS or Milvus) is used to store audit system text vectors, and indexing technology is used to achieve fast query. A matching query interface is established to enable users to quickly retrieve knowledge points related to the audit system through natural language input, thus building an intelligent audit system knowledge base DB. audit .

[0140] DB audit =Embed model (B)

[0141] Among them, DB audit Is the audit system knowledge base, Embed model is the vector embedding model, and B is the audit system text block dataset.

[0142] Step 3: In the audit instruction question-answering dataset, the audit questions are matched with the audit system knowledge base to obtain the corresponding audit system text blocks;

[0143] The specific implementation method of step 3 is:

[0144] In the audit instruction question-answering dataset, audit questions are processed using vector retrieval technology. First, the questions are converted into fixed-dimensional vector representations using a vector embedding model. The specific formula is as follows:

[0145] Q vec =Embed model (Q)

[0146] Among them, Q vec For the vectorized representation of the audit system problem, Embed modelis the vector embedding model, and Q is the input audit system problem.

[0147] The text cosine similarity is used to calculate the similarity between the question vector and the text block vector embedded in the audit system knowledge base. The specific formula is as follows:

[0148]

[0149] Among them, Sim(Q vec ,DB audit,i ) is the audit system problem vector Q vec And each text block DB in the audit system knowledge base audit,i The similarity value of .

[0150] Next, the calculation results are sorted from high to low by similarity, and the audit system text block that is most relevant to the question is selected. Finally, the matching audit system text block is output. The specific formula is as follows:

[0151]

[0152] Among them, k match is the most similar audit system text block, argmaxSim is the function that returns the highest similarity, Q vec is the audit system problem vector, DB audit,i It is every text block in the audit system knowledge base.

[0153] Finally, the audit instruction question answering dataset D QA The matching audit system text blocks are spliced ​​together to pave the way for the next step of screening high-quality audit instruction question and answer data sets.

[0154] In this embodiment, we process audit questions in the audit instruction question-and-answer dataset using vector retrieval technology. We select a suitable vector embedding model (e.g., bge-large-zh) and input the audit system questions into the model to convert them into fixed-dimensional semantic vector representations. We then perform the same processing on all questions in the existing audit system instruction question-and-answer dataset to generate vector representations of standard questions. The specific formula is as follows:

[0155] Q vec =Embed model (Q)

[0156] Among them, Q vec For the vectorized representation of the audit system problem, Embed model is the vector embedding model, and Q is the input audit system problem.

[0157] The text cosine similarity is used to calculate the similarity between the question vector and the text block vector embedded in the audit system knowledge base. The specific formula is as follows:

[0158]

[0159] Among them, Sim(Q vec ,DB audit,i ) is the audit system problem vector Q vec And each text block DB in the audit system knowledge base audit,i The similarity value of .

[0160] Next, the vector retrieval model is used to calculate the similarity between the input audit system question and each audit system text block in the audit system knowledge base. All text blocks are sorted from high to low by similarity value, and the text block with the highest similarity is selected. Finally, the matching text block is output. The specific formula is as follows:

[0161]

[0162] Among them, k match is the most similar audit system text block, argmaxSim is the function that returns the highest similarity, Q vec is the audit system problem vector, DB audit,i It is every text block in the audit system knowledge base.

[0163] Finally, the audit instruction question answering dataset D QA The corresponding text blocks are spliced ​​together to pave the way for the next step of screening high-quality audit instruction question and answer datasets.

[0164] Step 4: Use the answers in the audit instruction question and answer dataset and the matching retrieved audit system text blocks to calculate the similarity. If the similarity is greater than the threshold, retain the audit system question and answer pair to obtain the preliminary screened audit instruction question and answer dataset.

[0165] The specific implementation method of step 4 is:

[0166] The similarity between the answers in the audit instruction question-answering dataset and the matching retrieved audit system text blocks is calculated, and cosine similarity is used as the similarity measurement indicator.

[0167] First, the answer and the retrieved audit system text block are vectorized separately (for example, using embedding models such as Sentence-BERT to generate vector representations). The specific formula is as follows:

[0168] A vec =Embed model (A)

[0169] kmatchvec =Embed model (k match )

[0170] Among them, A vec is the vectorized representation of the answer extracted by the large model, k matchvec Embed is the vectorized representation of the retrieved audit system text block. model is the vector embedding model, k match It is the most similar audit system text block.

[0171] Then calculate the cosine similarity between the two. The specific formula is as follows:

[0172]

[0173] Among them, Sim(A vec ,k matchvec ) is the similarity between the answer extracted by the large model and the retrieved audit system text block, A vec is the vectorized representation of the answer extracted by the large model, k matchvec Vectorized representation of the audit system text block for retrieval.

[0174] Compare with the set threshold Threshold. If the similarity is greater than the threshold, the audit system question and answer pair is retained, and finally the preliminary screened audit instruction question and answer dataset Filtered D is obtained. QA .

[0175] Filtered D QA ={D QA |Sim(A vec ,k matchvec )>Threshold}

[0176] Among them, Filtered D QA It is a preliminary screened audit instruction question and answer dataset, Sim(A vec ,k matchvec ) is the similarity between the answer extracted by the large model and the retrieved audit system text block, and Threshold is the set threshold.

[0177] In this example, the similarity between the answers in the audit instruction question-and-answer dataset and the matching retrieved audit policy text blocks is calculated, using cosine similarity as the similarity metric. Initially, an embedding model (such as Sentence-BERT) is used to vectorize the answers and each audit policy text block, converting them into high-dimensional vector representations. The specific formula is as follows:

[0178] A vec =Embed model (A)

[0179] k matchvec =Embed model (k match )

[0180] Among them, A vec is the vectorized representation of the answer extracted by the large model, k matchvec Embed is the vectorized representation of the retrieved audit system text block. model is the vector embedding model, k match It is the most similar audit system text block.

[0181] Cosine similarity is used as a measurement metric to calculate the similarity between the answer vector and each text block vector. Cosine similarity reflects the degree of proximity between two vectors in vector space by calculating the cosine value of the angle between them. The specific formula is as follows:

[0182]

[0183] Among them, Sim(A vec ,k matchvec ) is the similarity between the answer extracted by the large model and the retrieved audit system text block, A vec is the vectorized representation of the answer extracted by the large model, k matchvec Vectorized representation of the audit system text block for retrieval.

[0184] Sort the text blocks in descending order according to the similarity score, and select the text block with the highest score as the part most relevant to the answer. At the same time, set a threshold (such as 0.85) for comparison. If the similarity is greater than the threshold, the audit system question and answer pair is retained, and finally the preliminary screened audit instruction question and answer dataset Filtered D is obtained. QA .

[0185] Filtered D QA ={D QA |Sim(A vec ,k matchvec )>Threshold}

[0186] Among them, A vec is the vectorized representation of the answer extracted by the large model, k matchvec Embed is the vectorized representation of the retrieved audit system text block. model is the vector embedding model, k match It is the most similar audit system text block.

[0187] Step 5: Use the big model to perform an intelligent voting mechanism on the audit instruction question and answer dataset that was initially screened in step 4. If the big model finds that there is a consensus and the number of votes is greater than half, the audit system question and answer pairs are retained to obtain the final audit instruction question and answer dataset.

[0188] The specific implementation method of step 5 is:

[0189] In the preliminary screening of the audit instruction question and answer dataset Filtered D QA In the process, a large model is used to set up multiple voting agents, and high-quality question-answer pairs are screened through the voting mechanism;

[0190] First, for each question-and-answer pair, the corresponding audit policy text block is retrieved based on its matching. This is then fed into the large model for consistency assessment. Each voting agent independently outputs a "consistent" or "inconsistent" evaluation result. The voting results of all models are tallied. If the number of "consistent" votes exceeds half of the total number of models, the question-and-answer pair is retained; otherwise, it is discarded. This mechanism ensures that the retained question-and-answer pairs are highly consistent within the large model, thereby constructing the final high-quality audit instruction question-and-answer dataset.

[0191] The specific formula is as follows:

[0192]

[0193] Where n is the number of voting agents, λ is the indicator function, when R = 'consistent', λ = 1, otherwise λ = 0, and Consistency represents the number of consistent votes.

[0194] Repeat the above steps for all audit instruction question-answer pairs to select the final high-quality question-answer dataset FinalD QA , the specific formula is as follows:

[0195] Final D QA ={Filtered D QA |Consistency>n / 2}

[0196] Among them, Final D QA It is the final high-quality question-answering dataset. Consistency>n / 2 means that more than half of the votes are considered valid.

[0197] In this embodiment, the audit instruction question and answer dataset Filtered D QAIn this paper, a large model is used to set up multiple voting agents, and a voting mechanism is used to select high-quality question-answer pairs. First, for each question-answer pair, the corresponding audit system text block is retrieved based on its matching. This is then input into the large model for consistency judgment. Each voting agent independently outputs a "consistent" or "inconsistent" evaluation result. The voting results of all models are counted. If the number of "consistent" votes exceeds half of the total number of models, the question-answer pair is retained; otherwise, it is discarded.

[0198] The voting agent setup process: Step 1: Determine the task objective, determine whether two texts are semantically consistent, and give a consistent or inconsistent judgment; Step 2: Design a voting agent model to judge the consistency of the text from five dimensions: semantics, theme, syntax, context, and logic; Step 3: Design a voting mechanism, and only votes greater than half are considered valid.

[0199] The prompts for setting up the agent from a semantic perspective are as follows: Please judge whether the semantics of the following two audit system texts are similar, and answer "consistent" or "inconsistent", and then briefly explain the reasons.

[0200] The prompt for the agent from the perspective of the topic is as follows: Are the topics described in the two audit system texts consistent? The agent should answer "consistent" or "inconsistent" and explain the reasons.

[0201] The prompt for the agent from a syntactic perspective is as follows: From a syntactic perspective, determine whether the following two audit system texts express similar meanings. Answer "consistent" or "inconsistent" and briefly explain the reason.

[0202] The prompt for the agent setting from the contextual perspective is as follows: Do these two audit policy texts convey the same core information? Please answer 'consistent' or 'inconsistent' and explain the reason.

[0203] The prompts for setting up the intelligent agent from a situational perspective are as follows: Please analyze whether the following two audit system texts express similar content logically, and give a judgment of "consistent" or "inconsistent".

[0204] Through this mechanism, we ensure that the retained question-answer pairs are highly consistent in the large model, thereby constructing the final high-quality audit instruction question-answer dataset. The specific formula is as follows:

[0205]

[0206] Where n is the number of voting agents, λ is the indicator function, when R = 'consistent', λ = 1, otherwise λ = 0, and Consistency represents the number of consistent votes.

[0207] Repeat the above steps for all audit instruction question-answer pairs to select the final high-quality question-answer dataset FinalD QA , the specific formula is as follows:

[0208] Final D QA ={Filtered D QA |Consistency>n / 2}

[0209] Among them, Final D QA It is the final high-quality question-answering dataset. Consistency>n / 2 means that more than half of the votes are considered valid.

[0210] Step 6: Use the final audit instruction question and answer dataset to fine-tune the instructions of the large model to obtain the audit system large model.

[0211] The specific implementation method of step 6 is:

[0212] Using the high-quality audit instruction question answering dataset Final D after final screening QA The specific implementation method for fine-tuning instructions on a large model is as follows: the dataset is divided into a training set and a validation set. The training set is used for model parameter optimization, and the validation set is used for model performance evaluation. Using instruction fine-tuning technology, question-answer pairs are used as fine-tuning samples in input-output format. A general large model (such as the ChatGLM series or the Qwen series) is loaded, and an appropriate learning rate and optimizer are set for gradual fine-tuning. During the fine-tuning process, the model learns the semantic understanding of audit instructions and the ability to generate answers. By periodically evaluating the accuracy, consistency, and generation quality on the validation set, hyperparameters are adjusted to optimize model performance, ultimately obtaining an audit system large model that can accurately interpret and generate answers related to audit instructions.

[0213] In this embodiment, to fine-tune the audit instruction large model, the dataset must first be divided into a training set and a validation set, with the training set used for model fine-tuning and the validation set used for performance evaluation. During the fine-tuning process, a question-answer pair format is used, with question-answer samples related to audit instructions as input and output to form the fine-tuning data. A pre-trained general large model (such as ChatGLM or the Qwen series) is loaded, and an appropriate learning rate and optimizer are set. Multiple rounds of training are then performed to gradually optimize the model's parameters to improve its semantic understanding of audit instructions and its ability to generate answers. For example, using LoRA fine-tuning as an example, the parameter settings are: fine-tune for no less than 15 epochs, with a batch size of 8, an initial learning rate of 3e-4, a cosine learning rate scheduler type, a warm-up step size of 0.01, LoRArank set to 64, LoRAalpha set to 16, and LoRAdropout set to 0.05. The maximum length of the input text is 2048.

[0214] During training, the model's performance is monitored by periodically evaluating accuracy, consistency, and generation quality (e.g., BLEU and ROUGE scores) on the validation set. Based on these results, hyperparameters such as the learning rate and batch size are adjusted to avoid overfitting or underfitting, ensuring the model's ability to flexibly generate high-quality answers to audit instructions. Ultimately, this process enables the model to accurately interpret and generate audit-related responses, forming a comprehensive audit system model specifically for the audit system domain.

[0215] Step 7: Use the audit system model to test and evaluate the model performance, and obtain answers related to the user's audit system questions based on the audit system questions input by the user.

[0216] The specific implementation method of step 7 is:

[0217] User questions are fed into the model, which generates targeted answers through semantic matching and generative answering techniques. Once the answers are complete, they are optimized using rule validation and semantic analysis to ensure logic and accuracy. Finally, a test set is constructed to evaluate model performance using metrics such as accuracy, recall, F1 score, similarity, and exact match, and then compared with human-generated answers. Based on these evaluation results, the model is continuously optimized to enhance its ability to handle complex questions.

[0218] An audit system question answering system based on retrieval enhancement and instruction fine-tuning, including:

[0219] A pure audit system data set acquisition module is used to clean the acquired audit system text data set to obtain a pure audit system data set;

[0220] The audit system knowledge base construction module is used to use the clean audit system dataset to perform text segmentation, use the large model to generate relevant audit instruction question and answer datasets through prompt engineering, and use the vector embedding model to build the audit system knowledge base;

[0221] The audit system text block module is used to match and retrieve the corresponding audit system text blocks in the audit instruction question and answer dataset by using the audit questions and the audit system knowledge base;

[0222] The module for obtaining the preliminary audit instruction question and answer dataset is used to calculate the similarity between the answers in the audit instruction question and answer dataset and the matching retrieved audit system text blocks. If the similarity is greater than the threshold, the audit system question and answer pair is retained to obtain the preliminary screened audit instruction question and answer dataset.

[0223] The final audit instruction question and answer dataset acquisition module uses the preliminary screened audit instruction question and answer dataset to conduct an intelligent voting mechanism using a large model. If the large model finds a consensus and the number of votes is greater than half, the audit system question and answer pair is retained to obtain the final audit instruction question and answer dataset;

[0224] The audit system large model construction module is used to use the final audit instruction question and answer dataset to fine-tune the large model to obtain the audit system large model;

[0225] The answer output module is used to test and evaluate the performance of the audit system model using the audit system model, and obtain answers related to the user's audit system questions based on the audit system questions input by the user.

[0226] The working principle of the present invention is:

[0227] The application of retrieval enhancement and instruction fine-tuning technology in the field of audit systems aims to solve the problems of low efficiency in processing massive multi-source data, uneven quality of manual data annotation, time-consuming and labor-intensive, and insufficient model question-answering capabilities in the field of audit systems. First, in view of the huge volume and diverse sources of audit system text data, an automated extraction mechanism was designed. To ensure the high accuracy and authority of the generated dataset, a two-stage "retrieval enhancement semantic matching-agent voting decision" screening mechanism was proposed. The accuracy of data retrieval is improved through semantic matching technology, and the reliable screening of the automatically extracted audit system instruction dataset is achieved by combining the intelligent voting decision mechanism. Based on the high-quality audit system instruction dataset, the large model is fine-tuned and trained to make it more suitable for question-answering scenarios in the field of audit systems.

[0228] It should be emphasized that the embodiments described in the present invention are illustrative rather than restrictive. Therefore, the present invention includes but is not limited to the embodiments described in the specific embodiments. Any other embodiments derived by those skilled in the art based on the technical solutions of the present invention also fall within the scope of protection of the present invention.

Claims

1. An audit system question-answering method based on retrieval enhancement and instruction fine-tuning, characterized by: The following steps are involved: Step 1: Clean the acquired audit system text dataset to obtain a pure audit system dataset; Step 2: Use the pure audit system dataset to segment the text, use the big model to generate relevant audit instruction question and answer datasets through prompt engineering, and use the vector embedding model to build the audit system knowledge base; Step 3: In the audit instruction question and answer dataset, the audit questions are matched with the audit system knowledge base to obtain the corresponding audit system text block; Step 4: Use the answers in the audit instruction question and answer dataset and the matching retrieved audit system text blocks to calculate the similarity. If the similarity is greater than the threshold, retain the audit system question and answer pair to obtain the preliminary screened audit instruction question and answer dataset; Step 5: Use the audit instruction question and answer data set initially screened in step 4 to perform an intelligent voting mechanism using the big model. If the big model believes that there is a consensus and the number of votes is greater than half, the audit system question and answer pair is retained to obtain the final audit instruction question and answer data set; Step 6: Use the final audit instruction question and answer data set to fine-tune the instructions of the large model to obtain the audit system large model; Step 7: Use the audit system model to test and evaluate the model performance, and obtain answers related to the user's audit system questions based on the audit system questions input by the user.

2. According to claim 1, the audit system question-answering method based on retrieval enhancement and instruction fine-tuning is characterized by: The specific implementation method of step 1 is: The acquired audit system text dataset is cleaned, word segmentation is performed, and special characters, tables, and graphs are removed. Finally, the audit system text only needs to retain all the system clauses to obtain a pure audit system dataset.

3. According to claim 1, the audit system question-answering method based on retrieval enhancement and instruction fine-tuning is characterized by: The specific implementation method of step 2 is: Using the clean audit system dataset, we cleaned and formatted it into unstructured documents through preprocessing, and divided the text into blocks according to fixed length; Input the audit system text dataset D. The specific formula is as follows: <h2 style=";text-align:left;direction:ltr">D = {d1,d2,d3,…,d<h2 style=";text-align:left;direction:ltr"> n <h2 style=";text-align:left;direction:ltr">},i = {1,2,3,…n} Among them, D is the audit system text dataset, d i It is an independent audit system document; Divide the audit system text into blocks and construct the audit system text block dataset B. Set the block size to L. Each text block b j ,satisfy: B={b1,b2,b3,…,b m },j={1,2,3,…m},b j =d i [l:k],k-l≤L Among them, B is the audit system text block dataset, b j It is an independent audit system text block, l and k are the start and end indexes of the audit system text block; For the constructed audit system text block dataset B, a prompt template P is designed to extract the audit instruction question and answer dataset. Using the big model, the prompt template is applied to each block content to generate audit instructions and question and answer pairs related to the text. The specific formula is as follows: (I j ,Q j ,A j )6GPT large (P(b). j )) Among them, I j is the audit instruction, Q j It is a problem of audit system. j The answer is GPT large is the big model, P is the prompt template, b j It is an independent audit system text block; Combine the generated audit system instruction question and answer pairs to obtain the audit instruction question and answer dataset D QA ; D QA ={(I j ,Q j ,A j )|j={1,2,3,···m}} Among them, D QA is the audit instruction question answering dataset, I j is the audit instruction, Q j It is a problem of audit system. j is the answer; For the constructed audit system text block dataset B, semantic vector modeling is performed to build the audit system knowledge base DB audit ; DB audit =Embed model (B) Among them, DB audit It is the audit system knowledge base, Embed model is the vector embedding model, and B is the audit system text block dataset.

4. According to claim 1, the audit system question-answering method based on retrieval enhancement and instruction fine-tuning is characterized by: The specific implementation method of step 3 is: In the audit instruction question-answering dataset, the audit questions are processed by vector retrieval technology. First, the questions are converted into fixed-dimensional vector representations through a vector embedding model; The specific formula is as follows: Q vec =Embed model (Q) Among them, Q vec For the vectorized representation of the audit system problem, Embed model is the vector embedding model, Q is the input audit system problem; The similarity between the question vector and the text block vector in the embedded audit system knowledge base is calculated using text cosine similarity; The specific formula is as follows: Among them, Sim(Q vec ,DB audit,i ) is the audit system problem vector Q vec And each text block DB in the audit system knowledge base audit,i Similarity value of ; Next, the calculation results are sorted from high to low according to the similarity, and an audit system text block that is most relevant to the problem is selected, and finally the matching text block is output; the specific formula is as follows: Among them, k match is the most similar audit system text block, argmaxSim is the function that returns the highest similarity, Q vec is the audit system problem vector, DB audit,i It is every block of text in the audit system knowledge base; Finally, the audit instruction question and answer dataset D QA Connect with the matching text blocks.

5. According to claim 1, the audit system question-answering method based on retrieval enhancement and instruction fine-tuning is characterized by: The specific implementation method of step 4 is: The similarity between the answers in the audit instruction question-and-answer dataset and the audit system text blocks that are matched and retrieved is calculated, and the cosine similarity is used as the similarity measurement indicator. First, the answer and the retrieved audit system text blocks are vectorized respectively. The specific formula is as follows: A vec =Embed model (A) k matchvec =Embed model (k match ) Among them, A vec is the vectorized representation of the answer extracted by the large model, k matchvec Embed is the vectorized representation of the retrieved audit system text block. model is the vector embedding model, k match It is the most similar audit system text block; Then calculate the cosine similarity between the two; the specific formula is as follows: Among them, Sim(A vec ,k matchvec ) is the similarity between the answer extracted by the large model and the retrieved audit system text block, A vec is the vectorized representation of the answer extracted by the large model, k matchvec Vectorized representation of the retrieved audit system text block; Compare with the set threshold Threshold; if the similarity is greater than the threshold, the audit system question and answer pair is retained, and finally the preliminary screened audit instruction question and answer dataset Filtered D is obtained. QA ; Filtered D QA ={D QA |Sim(A vec ,k matchvec )>Threshold} Among them, Filtered D QA It is a preliminary screened audit instruction question and answer dataset, Sim(A vec ,k matchvec ) is the similarity between the answer extracted by the large model and the retrieved audit system text block, and Threshold is the set threshold.

6. The audit system question-answering method based on retrieval enhancement and instruction fine-tuning according to claim 1 is characterized by: The specific implementation method of step 5 is: In the preliminary screening of the audit instruction question and answer dataset Filtered D QA In the process, a large model is used to set up multiple voting agents, and high-quality question-answer pairs are selected through the voting mechanism; First, for each question-answer pair, the audit system text block corresponding to the matching retrieval is input into the large model for consistency judgment. Each voting agent independently outputs an evaluation result of "consistent" or "inconsistent". The voting results of all models are counted. If the number of "consistent" votes exceeds half of the total number of models, the question-answer pair is retained. Otherwise, remove; Through this mechanism, it is ensured that the retained question-answer pairs are highly consistent in the large model, thereby constructing the final high-quality audit instruction question-answer dataset; The specific formula is as follows: Where n is the number of voting agents, λ is the indicator function, when R = 'consistency', λ = 1, otherwise λ = 0, and Consistency represents the number of consistent votes; Repeat the above steps for all audit instruction question-answer pairs to select the final high-quality question-answer dataset Final D QA , the specific formula is as follows: Final D QA ={Filtered D QA |Consustency>n / 2} Among them, Final D QA It is the final high-quality question-answering dataset. Consustency>n / 2 means that more than half of the votes are considered valid.

7. The audit system question-answering method based on retrieval enhancement and instruction fine-tuning according to claim 1 is characterized by: The specific implementation method of step 6 is: Using the high-quality audit instruction question answering dataset Final D QA ,The specific implementation method of instruction fine-tuning for the large model is as follows: ,divide the data set into a training set and a validation set, the training set is used for model parameter optimization, and the validation set is used for model performance evaluation; ,adopt instruction fine-tuning technology, take the question-answer pair as the fine-tuning sample in the input-output format, load the general large model, set the appropriate learning rate and optimizer, and perform step-by-step fine-tuning; During the fine-tuning process, the model learns the semantic understanding and answer generation capabilities of audit instructions; by periodically evaluating the accuracy, consistency, and generation quality on the validation set, the hyperparameters are adjusted to optimize the model performance, ultimately obtaining a large audit system model that can accurately interpret and generate answers related to audit instructions.

8. The audit system question-answering method based on retrieval enhancement and instruction fine-tuning according to claim 1 is characterized by: The specific implementation method of step 7 is: input the user question into the model, and generate targeted answers through semantic matching and generative answering technology; after the answer is completed, optimize the answer by rule verification and semantic analysis to ensure logic and accuracy; finally, by constructing a test set, evaluate the model performance with indicators such as accuracy, recall rate, F1 score, similarity, and precise matching, and compare and analyze with manual answers; based on the evaluation results, continuously optimize the model to enhance its ability to handle complex problems.

9. An audit system question-answering system based on retrieval enhancement and instruction fine-tuning, characterized by: include: A pure audit system data set acquisition module is used to clean the acquired audit system text data set to obtain a pure audit system data set; The audit system knowledge base construction module is used to use the pure audit system data set to perform text segmentation, use the large model to generate relevant audit instruction question and answer data sets through prompt engineering, and use the vector embedding model to build the audit system knowledge base; The audit system text block module is used to match and retrieve the audit questions with the audit system knowledge base in the audit instruction question and answer data set to obtain the corresponding audit system text block; A preliminary audit instruction question and answer data set acquisition module is used to calculate the similarity between the answers in the audit instruction question and answer data set and the audit system text blocks that are matched and retrieved. If the similarity is greater than a threshold, the audit system question and answer pairs are retained to obtain a preliminary screened audit instruction question and answer data set; The final audit instruction question and answer data set acquisition module uses the preliminary screened audit instruction question and answer data set to use the big model for intelligent voting mechanism. If the big model believes that it is consistent and the number of votes is greater than half, the audit system question and answer pairs are retained to obtain the final audit instruction question and answer data set; The audit system big model building module is used to use the final audit instruction question and answer data set to fine-tune the big model to obtain the audit system big model; The answer output module is used to test and evaluate the performance of the audit system model using the audit system model, and obtain answers related to the user's audit system questions based on the audit system questions input by the user.

Citation Information

Cited By

  • Large model fine tuning method and system based on power distribution network dispatching

    CN121094117A

  • Auditing interview corpus analysis method and system based on large model

    CN121457479A