Audit dialogue data enhancement method and system based on instruction scoring

By introducing large language models and instruction scoring models in audit tasks, the enhanced instructions are automatically generated and optimized, and the problem of insufficient contextual information in the existing technology is solved, and efficient and accurate audit data enhancement is achieved.

CN120030996APending Publication Date: 2025-05-23STATE GRID TIANJIN ELECTRIC POWER COMPANY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411900345.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In audit tasks, it is difficult for the prior art to automatically generate high-quality text enhancement instructions, resulting in data enhancement quality dependent on manual design, and insufficient context information between different audit tasks, affecting model performance.

Method used

By introducing audit task data sets and large language models, enhancement instructions for audit tasks are automatically generated and selected, and the instruction scoring model is used to optimize the quality of enhanced instructions to ensure the diversity and effectiveness of enhanced data.

Benefits of technology

It realizes automatic generation of high-quality enhanced instructions, improves the efficiency and accuracy of audit data enhancement, reduces performance differences between different tasks, and improves the overall effect of the audit model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030996A_ABST
    Figure CN120030996A_ABST
Patent Text Reader

Abstract

The invention relates to an audit dialogue data enhancement method and system based on instruction scoring, and the method comprises the following steps: 1, confirming an audit task, and obtaining an audit task data set; step 2, constructing a seed enhancement instruction set for the audit task; step 3, outputting an enhanced instruction set for the auditing task; 4, training the basic instruction scoring model to obtain an instruction scoring model; 5, determining a new audit task needing data enhancement, obtaining a new audit task data set corresponding to the new audit task, introducing the instruction scoring model and the enhancement instruction set, and outputting an optimal enhancement instruction of the new audit task; and step 6, introducing an optimal enhancement instruction of a new audit task and a new audit task data set, and performing data enhancement on the new audit task data set by using the large language model to complete audit dialogue data enhancement based on instruction scoring. Effective enhanced data can be provided for specific tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence natural language processing, and relates to an audit dialogue data enhancement method and system, in particular to an audit dialogue data enhancement method and system based on instruction scoring. Background Art

[0002] There are some challenges in using large language models for text data augmentation in auditing tasks. First, the quality of data augmentation usually depends on the design of augmentation instructions, which often need to be manually written by domain experts based on auditing tasks. This not only requires deep expertise, but also may result in inconsistent instruction design, which in turn affects the overall quality of augmented data. Second, text augmentation instructions for auditing tasks are usually written in a task-agnostic manner, that is, lacking contextual information for specific auditing tasks, which may lead to significant differences in the effect of augmented data and model performance between different auditing tasks.

[0003] Therefore, in order to solve the above problems, the present invention proposes an audit dialogue data enhancement method and system based on instruction scoring.

[0004] After searching, no public documents of the prior art identical or similar to the present invention were found. Summary of the invention

[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and propose an audit dialogue data enhancement method and system based on instruction scoring, which can provide effective enhanced data for specific tasks by automatically generating and selecting data enhancement instructions, thereby solving the need for manually designing enhancement instructions and improving the performance differences between different tasks caused by insufficient context information.

[0006] The present invention solves the practical problem by adopting the following technical solutions:

[0007] An audit dialogue data enhancement method based on instruction scoring includes the following steps:

[0008] Step 1: Confirm the audit task and obtain the audit task data set;

[0009] Step 2: Introduce the audit task dataset, and the audit experts build a seed enhancement instruction set for the audit task based on the audit task dataset;

[0010] Step 3: Select a large language model, introduce a seed enhanced instruction set, and use the large language model to output an enhanced instruction set for the audit task;

[0011] Step 4: Introduce audit tasks, audit task data sets, and enhanced instruction sets, train the basic instruction scoring model, and obtain the instruction scoring model;

[0012] Step 5: Confirm the new audit task that needs data enhancement, obtain the new audit task data set corresponding to the new audit task, introduce the instruction scoring model and the enhanced instruction set, and output the optimal enhanced instruction for the new audit task;

[0013] Step 6: Introduce the optimal enhanced instructions for the new audit task and the new audit task dataset, use the large language model to perform data enhancement on the new audit task dataset, and complete the audit dialogue data enhancement based on instruction scoring.

[0014] Moreover, the specific implementation method of step 1 is:

[0015] First, confirm the audit task T, obtain the data set D of the audit task T, and divide the data set D into a training data set and a test data set.

[0016] Moreover, the specific implementation method of step 2 is:

[0017] Audit domain experts construct seed enhancement instructions based on the data set D of audit task T; the initially constructed seed enhancement instructions are reviewed and discussed, and after multiple rounds of review, the edited and optimized seed enhancement instructions of the audit task are finally summarized to form a complete seed enhancement instruction set I_seed for the audit task.

[0018] Moreover, the specific implementation method of step 3 is:

[0019] Select and confirm the large language model, provide the seed enhancement instruction set I_seed obtained in step 2 as input to the large language model, and through prompt engineering, enable the large language model to generate diversified audit task data enhancement instructions; finally, obtain the enhanced instruction set I for the audit task.

[0020] Moreover, the specific implementation steps of step 4 include:

[0021] (1) Introduce the large language model in step 3, combine the training data set of the audit task obtained in step 1 with the audit task enhanced instruction set I as the input of the large language model, and generate new data through prompt engineering for subsequent training of the selected target model.

[0022] (2) Select the basic instruction scoring model, combine each audit task and its corresponding training data with the enhanced instruction set I, and pass them into the basic instruction scoring model as input. The model outputs the "yes" mark logits value of each instruction on each audit task as the predicted score.

[0023] (3) The selected target model is further trained using the newly generated data in step (1). After the training is completed, the target model is evaluated using the test set of the audit task obtained in step 1. The evaluation results are obtained by calculating the accuracy of the classification task or the cosine similarity of the question-answering task, and are used as the true score of the audit task data enhancement instruction set I on each audit task.

[0024] (4) The loss is calculated by combining the predicted score and the true score. The specific formula is as follows:

[0025]

[0026] Among them, for an audit task Ti, r_j represents the actual score of the enhanced instruction Ij in the enhanced instruction set I in the audit task Ti. If the score r_j of the enhanced instruction Ij is the highest, the is_max function returns 1, otherwise is_max returns 0; k_j is the predicted score of the enhanced instruction Ij in the audit task Ti, and σ represents the softmax function, which is used to normalize a set of numerical values.

[0027] (5) Finally, by minimizing the loss function, the model is gradually optimized to obtain the instruction scoring model.

[0028] Moreover, the specific implementation method of step 5 is:

[0029] Confirm the new audit task that needs data enhancement, obtain the new audit task data set corresponding to the new audit task, and after obtaining the instruction scoring model after training through step 4, use the instruction scoring model to score each instruction in the enhanced instruction set I that acts on the new audit task, and select the optimal enhanced instruction I* for the new audit task, which will be used for subsequent data enhancement of the data set corresponding to the new audit task.

[0030] Moreover, the specific implementation method of step 6 is:

[0031] The large language model in step 3 is introduced, and the optimal enhancement instruction I* obtained in step 5 is combined with the data set of the new audit task obtained in step 5 as input to the large language model; the large language model uses its powerful generation ability to combine the optimal enhancement instruction I* for data enhancement, thereby generating a new data set and finally completing the audit dialogue data enhancement based on instruction scoring.

[0032] Advantages and beneficial effects of the present invention:

[0033] 1. The present invention proposes an audit dialogue data enhancement method based on instruction scoring, which can automatically generate and select text enhancement instructions for audit tasks, thereby providing effective enhancement data for each specific audit task. This not only solves the tedious problem of manually designing enhancement instructions, but also reduces the impact of performance differences between tasks by considering the specific needs and contextual information of audit tasks. Through automated instruction generation and selection, data enhancement of audit tasks can more efficiently and accurately meet the needs of different tasks and improve the overall effect of the audit model.

[0034] 2. After the present invention generates data enhancement samples through the audit dialogue data enhancement method based on instruction scoring, the enhanced audit task data and the original audit task training data set are used as the input of the target model to train the target model, and then the training set is tested. The method of the present invention is superior to other existing methods in data enhancement effect. In general, the present invention can automatically generate and select text enhancement instructions, thereby providing effective enhancement data for specific tasks. It not only solves the need for manually designing enhancement instructions, but also improves the performance difference problem between different tasks caused by insufficient context information.

[0035] 3. The present invention utilizes a large language model to automatically generate text data enhancement instructions, uses a scoring model to evaluate and select text data enhancement instructions, and acts on the audit task data set to improve the diversity and effectiveness of text data. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 and Figure 2 It is a flowchart of step 1 of the present invention for constructing an audit task data set and segmenting the audit task data set;

[0037] Figure 3 is a flowchart of the seed instruction set design in step 2 of the present invention;

[0038] Figure 4 is the flowchart of step 3 expansion instruction set of the present invention;

[0039] Figure 5 and Figure 6 This is a flow chart of the training instruction scoring model in step 4 of the present invention;

[0040] Figure 7 It is the flow chart of selecting the optimal enhancement instruction in step 5 of the present invention;

[0041] Figure 8 This is the data enhancement flow chart of step 6 of the present invention. DETAILED DESCRIPTION

[0042] The embodiments of the present invention are further described in detail below with reference to the accompanying drawings:

[0043] An audit dialogue data enhancement method based on instruction scoring includes the following steps:

[0044] Step 1: Confirm the audit task T = {T1, T2, ...Tm, ...Tn}, and obtain the data sets D = {D1, D2, ...Dm, ...Dn} corresponding to the audit tasks T1, T2, ...Tm, ...Tn, respectively. Confirm the audit task and obtain the audit task data set. Figure 1 As shown. The audit task data set D1, D2, ...Dm, ...Dn is divided into a training data set D_train = {D1_train, D2_train, ...Dm_train, ...Dn_train} and a test data set D_test = {D1_test, D2_test, ...Dm_test, ...Dn_test}, and the audit task data set is divided into an audit task training test set and an audit task test data set as shown. Figure 2 shown.

[0045] First, confirm the specific content involved in the audit task T. By studying relevant literature and industry standards, clarify the core tasks and challenges in this field to ensure the pertinence of the subsequent data set D selection. After confirming the audit task field T, the next step is to collect the data set D related to this field. The selection of the data set should take into account its diversity and representativeness to cover different audit scenarios and task requirements.

[0046] (1) For example, the audit task area T is determined, including tasks T1 (audit system questions and answers: answer questions based on the audit system), T2 (audit problem classification: determine the category of audit problems based on audit problems and audit opinions), and Tm (audit examination questions: multiple-choice questions (ABCD), choose answers based on the questions).

[0047] (2) According to the audit task domain T in (1), relevant data sets D are obtained, including data sets D1 of task T1, data sets D2 of task T2, data sets Dm of task Tm, and other task data sets.

[0048] The following is a sample dataset:

[0049] Sample dataset D1 for task T1 (audit system question and answer)

[0050] <human>:Description of audit problem: A company did not pay stamp tax as required when conducting domestic securities transactions, which violated the relevant provisions of the Stamp Tax Law of the People's Republic of China. Please output the relevant audit system regulations that were violated based on the description of the audit problem.

[0051] <bot>Article 1 of the Stamp Tax Law of the People's Republic of China states that "Entities and individuals that issue taxable certificates and conduct securities transactions within the territory of the People's Republic of China shall be taxpayers of stamp tax and shall pay stamp tax in accordance with the provisions of this Law. Entities and individuals that issue taxable certificates outside the territory of the People's Republic of China and use them within the territory of the People's Republic of China shall pay stamp tax in accordance with the provisions of this Law."

[0052] Example of dataset D2 for task T2 (audit question classification)

[0053] <human>:Description of audit problem: In the warehouse, waste materials are not stored in a classified manner, and no special area is set up for management, which makes it impossible to clearly distinguish between waste materials and materials in use, affecting the standardization and effectiveness of material management. Audit opinion: A comprehensive inspection of the warehouse should be carried out immediately, and a special area for waste materials should be set up. The material management ledger should be improved to ensure that all materials are recorded, and physical inventory should be carried out regularly to ensure that the accounts are consistent with the actual situation. Based on the description of the audit problem and the audit opinion, please determine the category of the audit problem from ['Financial asset management', 'Marketing management', 'Three major decisions', 'Project management', 'Internal control management', 'Industrial unit management', 'Human resources management', 'Contract management', 'Procurement management', 'Material management', 'Bidding management', 'Collective enterprise management', 'Official consumption management', 'Major issues'].

[0054] <bot>:Material Management

[0055] Sample dataset Dm for task Tm (audit test questions)

[0056] <human>:Audit question: Among the following items, the one that is not included in the GDP accounting system is: Options: {'A':'A batch of goods exported abroad', 'B':'A piece of equipment imported from abroad by a foreign trade company and sold domestically', 'C':'A commission collected by a broker for the sale of an old house', 'D':'An insurance company receives a household property insurance premium'} Please choose one option from all the options to answer based on the audit question and options, combined with your knowledge.

[0057] <bot>:B: A foreign trade company imported a piece of equipment from abroad and sold it domestically.

[0058] For the collected data set D, each type of audit task data set is divided into training data sets D1_train, D2_train, ...Dm_train, ...Dn_train and test data sets D1_test, D2_test, ...Dm_test, ...Dn_test in a ratio of 8:2.

[0059] Step 2: Based on the audit task dataset D obtained in step 1, combined with the manual design of domain experts, construct the seed enhancement instruction set I_seed of the audit task T. The seed enhancement instruction set I_seed is constructed as follows: Figure 3 shown.

[0060] An example of the seed enhancement instruction set I_seed is as follows.

[0061] "In audit tasks, the diversity of audit data can be enhanced by replacing certain key words with their synonyms while keeping the sentence structure unchanged. This method helps improve the audit model's ability to recognize various expressions. Especially when dealing with different types of audit tasks, by using ready-made synonym databases or word embedding technology, the audit system can better adapt to audit materials from different sources. This method can help improve the accuracy and generalization ability of the audit model in analyzing different types of audit issues."

[0062] Step 3: Select a large language model, use the seed enhanced instruction set I_seed in step 2 as the input of the large language model (LLM), and obtain the enhanced instruction set I (a total of z instructions) of the expanded audit task T through prompt engineering to construct the enhanced instruction set I, such as Figure 4 shown.

[0063] The seed enhancement instruction set I_seed obtained in step 2 is provided as input to the large language model. Through the prompt engineering, the large language model generates diversified audit task data enhancement instructions. Finally, the enhanced instruction set I = {I1, I2, ..., Iz} for the audit task is obtained. The formula is as follows.

[0064] I=LLM(I_seed)

[0065] Among them, LLM is the large language model, I_seed is the seed enhanced instruction set, and I is the enhanced instruction set.

[0066] The following is an example of the enhanced instruction set I generated by the large language model.

[0067] "When conducting audit work, by replacing some keywords with their synonyms without changing the sentence structure, the diversity of audit data can be effectively increased. This practice helps improve the audit model's ability to recognize various expressions, especially when facing different types of audit tasks. By using existing synonym libraries or word vector technology, the audit system can better adapt to audit materials from different sources. This method helps improve the accuracy and generalization ability of the audit model when dealing with a variety of audit problems."

[0068] "In auditing, data diversity can be enhanced by replacing some key terms with their synonyms while keeping the sentence structure unchanged. This strategy helps improve the audit model's ability to recognize various language expressions. Especially when dealing with different types of audit tasks, by using ready-made synonym libraries or word embedding technology, the audit system can better adapt to audit content from different sources. This approach helps improve the accuracy and generalization of the audit model in analyzing different audit issues."

[0069] Step 4: Determine the basic instruction scoring model, and train the instruction scoring model S based on it so that it can score the instructions in the instruction set I more accurately.

[0070] First, the large language model in step 3 is introduced, and the audit task training dataset D_train is combined with the enhanced instruction set I as the input of the large language model. Through prompt engineering, new data is generated for subsequent training of the selected target model, such as Figure 5 The specific expression is as follows.

[0071] D1_by_I1_new_train=LLM(D1_train,I1),..., D1_by_Iz_new_train=LLM(D1_train,Iz) D2_by_I1_new_train=LLM(D2_train,I1),..., D2_by_Iz_new_train=LLM(D2_train,Iz) ...

[0073] Di_by_I1_new_train=LLM(Di_train,I1),..., Di_by_Iz_new_train=LLM(Di_train,Iz) ...

[0075] Dn_by_I1_new_train=LLM(Dn_train,I1),..., Dn_by_Iz_new_train=LLM(Dn_train,Iz)

[0076] Among them, Di_by_Ij_new_train (0 < i < n, 0 < j < z) is obtained by using the training dataset Di_train of the audit task Ti combined with the enhancement instruction Ij as the input of the large language model and performing data enhancement in combination with prompt engineering.

[0077] Next, train the basic instruction scoring model, as Figure 6 shown, and the specific content is as follows.

[0078] First, select a basic instruction scoring model (such as T5), combine each audit task and its corresponding training data with the enhancement instruction set I, and use it as the input to the basic instruction scoring model. The model outputs the "yes" label logits value of each instruction on each audit task as the prediction score.

[0079] Design the following prompt and input it into the basic instruction scoring model.

[0080] "There are the following audit task data samples and data enhancement instructions

[0081] Audit task data samples: {{Randomly select multiple samples from the audit task dataset}}

[0082] Data enhancement instructions: {{A data enhancement instruction}}

[0083] Do you think this data enhancement instruction is suitable for this audit task?"

[0084] Then, select the target model (such as OPT), use D1_by_I1_new_train,..., D1_by_Iz_new_train to train the target model in combination with D1_train respectively. After the training is completed, test the target model through D1_test respectively. By calculating evaluation metrics, such as accuracy for classification problems and cosine similarity for question-and-answer problems, obtain the evaluation results as the true scores of each instruction in the enhancement instruction set I in the audit task T1. Then use D2_by_I1_new_train,..., D2_by_Iz_new_train to train the target model in combination with D2_train respectively. After the training is completed, test the target model through D2_test respectively, and obtain the evaluation results as the true scores of each instruction in the enhancement instruction set I in the audit task T2. Iterate in this way to obtain the true scores of each instruction in the enhancement instruction set I on each audit task.

[0085] After the above steps, the predicted score and the real score of each instruction in the enhanced instruction set I in each audit task are obtained. For each audit task, there are the real score and the predicted score of the instruction in the enhanced instruction set I on each audit task. The loss is calculated by combining the predicted score and the real score. The specific formula is as follows.

[0086]

[0087] Among them, for an audit task Ti, r_j represents the actual score of the enhanced instruction Ij in the enhanced instruction set I in the audit task Ti. If the score r_j of the enhanced instruction Ij is the highest, the is_max function returns 1, otherwise is_max returns 0; k_j is the predicted score of the enhanced instruction Ij in the audit task Ti, and σ represents the softmax function, which is used to normalize a set of numerical values.

[0088] Finally, by minimizing the loss function, the model is gradually optimized to obtain the instruction scoring model S.

[0089] Step 5: Confirm the new audit task T_newTask that needs data enhancement, obtain the new audit task dataset D_newTask corresponding to the new audit task, and obtain the instruction scoring model S after training through step 4. This step uses the instruction scoring model S to score the enhanced instruction set I acting on the new audit task dataset D_newTask, and select the optimal enhanced instruction I* for the new audit task T_newTask, which is used for subsequent data enhancement of the dataset D_newTask corresponding to the new audit task T_newTask, such as Figure 7 The specific formula is as follows.

[0090] I*=argmax(S(I,D_newTask))

[0091] Where I represents the enhanced instruction set, D_newTask is the data set corresponding to the new audit task T_newTask, and S is the instruction scoring model, which is used to score all instructions in the enhanced instruction set I. The argmax function is used to return the instruction with the highest score, and I* is the optimal enhanced instruction in the enhanced instruction set I for the new audit task T_newTask.

[0092] Step 6: Introduce the large language model in step 3, combine the optimal enhancement instruction I* obtained in step 5 with the data set D_newTask of the audit task T_newTask, and pass it as input to the large language model. Generate a new data set D'_newTask through the large language model, such as Figure 8 The specific formula is as follows.

[0093] D'_newTask=LLM(D_newTask,I*)

[0094] Among them, LLM is the large language model, D_newTask is the data set of the audit task T_newTask, I* is the optimal enhancement instruction of the audit task T_newTask, and D'_newTask is the new data set generated by enhancement based on D_newTask.

[0095] Finally, the audit dialogue data enhancement based on instruction scoring was completed.

[0096] It should be emphasized that the embodiments of the present invention are illustrative rather than restrictive. Therefore, the present invention includes but is not limited to the embodiments described in the specific implementation modes. Any other implementation modes derived by those skilled in the art based on the technical solutions of the present invention also fall within the scope of protection of the present invention.< / bot> < / human> < / bot> < / human> < / bot> < / human>

Claims

1. An audit dialogue data enhancement method based on instruction scoring, characterized by: The following steps are involved: Step 1: Confirm the audit task and obtain the audit task data set; Step 2: Introduce the audit task dataset, and the audit experts build a seed enhancement instruction set for the audit task based on the audit task dataset; Step 3: Select a large language model, introduce a seed enhanced instruction set, and use the large language model to output an enhanced instruction set for the audit task; Step 4: Introduce audit tasks, audit task data sets, and enhanced instruction sets, train the basic instruction scoring model, and obtain the instruction scoring model; Step 5: Confirm the new audit task that needs data enhancement, obtain the new audit task data set corresponding to the new audit task, introduce the instruction scoring model and the enhanced instruction set, and output the optimal enhanced instruction for the new audit task; Step 6: Introduce the optimal enhanced instructions for the new audit task and the new audit task dataset, use the large language model to perform data enhancement on the new audit task dataset, and complete the audit dialogue data enhancement based on instruction scoring.

2. According to claim 1, the audit dialogue data enhancement method based on instruction scoring is characterized by: The specific implementation method of step 1 is: First, confirm the audit task T, obtain the data set D of the audit task T, and divide the data set D into a training data set and a test data set.

3. According to claim 1, the audit dialogue data enhancement method based on instruction scoring is characterized in that: The specific implementation method of step 2 is: Audit domain experts construct seed enhancement instructions based on the data set D of audit task T; the initially constructed seed enhancement instructions are reviewed and discussed, and after multiple rounds of review, the edited and optimized seed enhancement instructions of the audit task are finally summarized to form a complete seed enhancement instruction set I_seed for the audit task.

4. According to claim 1, the audit dialogue data enhancement method based on instruction scoring is characterized in that: The specific implementation method of step 3 is: Select and confirm the large language model, provide the seed enhancement instruction set I_seed obtained in step 2 as input to the large language model, and through prompt engineering, enable the large language model to generate diversified audit task data enhancement instructions; finally, obtain the enhanced instruction set I for the audit task.

5. According to claim 1, the audit dialogue data enhancement method based on instruction scoring is characterized in that: The specific implementation steps of step 4 include: (1) Introduce the large language model in step 3, combine the training data set of the audit task obtained in step 1 with the audit task enhanced instruction set I as the input of the large language model, and generate new data through prompt engineering for subsequent training of the selected target model; (2) Select a basic instruction scoring model, combine each audit task and its corresponding training data with the enhanced instruction set I, and pass them into the basic instruction scoring model as input; the model outputs the "yes" mark logits value of each instruction on each audit task as the predicted score; (3) Further use the newly generated data in step (1) to train the selected target model; after the training is completed, use the test set of the audit task obtained in step 1 to evaluate the target model, and obtain the evaluation results by calculating the accuracy of the classification task or the cosine similarity of the question-answering task, and use this as the true score of the audit task data enhancement instruction set I on each audit task; (4) The loss is calculated by combining the predicted score and the true score. The specific formula is as follows: Among them, for an audit task Ti, r_j represents the real score of the enhanced instruction Ij in the enhanced instruction set I in the audit task Ti. If the score r_j of the enhanced instruction Ij is the highest, the is_max function returns 1, otherwise is_max returns 0; k_j is the predicted score of the enhanced instruction Ij in the audit task Ti, σ represents the softmax function, which is used to normalize a set of values; (5) Finally, by minimizing the loss function, the model is gradually optimized to obtain the instruction scoring model.

6. The audit dialogue data enhancement method based on instruction scoring according to claim 1 is characterized by: The specific implementation method of step 5 is: Confirm the new audit task that needs data enhancement, obtain the new audit task data set corresponding to the new audit task, and after obtaining the instruction scoring model after training through step 4, use the instruction scoring model to score each instruction in the enhanced instruction set I that acts on the new audit task, and select the optimal enhanced instruction I* for the new audit task, which will be used for subsequent data enhancement of the data set corresponding to the new audit task.

7. The audit dialogue data enhancement method based on instruction scoring according to claim 1 is characterized by: The specific implementation method of step 6 is: The large language model in step 3 is introduced, and the optimal enhancement instruction I* obtained in step 5 is combined with the data set of the new audit task obtained in step 5 as input to the large language model; the large language model uses its powerful generation ability to combine the optimal enhancement instruction I* for data enhancement, thereby generating a new data set and finally completing the audit dialogue data enhancement based on instruction scoring.

8. An audit dialogue data enhancement system based on instruction scoring, characterized by: include: Audit task dataset module, used to confirm audit tasks and obtain audit task datasets; The seed enhancement instruction set module is used to introduce the audit task data set. Audit experts build the seed enhancement instruction set for the audit task based on the audit task data set. The enhanced instruction set module is used to select a large language model, introduce a seed enhanced instruction set, and use the large language model to output an enhanced instruction set for audit tasks; The instruction scoring model building module is used to introduce audit tasks, audit task data sets, and enhanced instruction sets, train the basic instruction scoring model, and obtain the instruction scoring model; The optimal enhanced instruction output module for new audit tasks is used to confirm new audit tasks that require data enhancement, obtain new audit task data sets corresponding to new audit tasks, introduce instruction scoring models and enhanced instruction sets, and output optimal enhanced instructions for new audit tasks; The audit dialogue data enhancement module is used to introduce the optimal enhancement instructions for new audit tasks and new audit task data sets, use the large language model to perform data enhancement on the new audit task data sets, and complete the audit dialogue data enhancement based on instruction scoring.