Data distillation method oriented to audit field and based on knowledge enhancement and course learning

By integrating audit databases and regulatory policy knowledge into the training of small models, generating thought chains and conducting difficulty-level training, the problems of high cost of large models and insufficient intelligence of small models are solved. This enables the efficient and professional application of small models in the audit field and promotes the intelligentization process of the audit industry.

CN121787557APending Publication Date: 2026-04-03XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, large models are difficult to operate at high cost in the auditing field, making it difficult to meet the enterprise-level low-cost deployment requirements. Small models, on the other hand, lack intelligence, professionalism, and accuracy in complex multi-step reasoning tasks. Existing methods lack business scenario labels and regulatory knowledge guidance, and the reasoning path depends on static triples, resulting in insufficient logical coherence and depth. The transfer and generalization performance of the model in high-difficulty, low-frequency scenarios is unstable.

Method used

By employing a data distillation method based on knowledge enhancement and course learning, audit database data and regulatory policy knowledge are integrated during the training of small models to generate thought chains. The small models are then trained using a difficulty-grading strategy to improve business relevance and professional compliance, ensuring logical consistency and transferability.

Benefits of technology

It significantly improves the reliability and professionalism of small models in complex audit reasoning scenarios, breaks through deployment barriers, and promotes the intelligent transformation and efficient implementation of the audit industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121787557A_ABST
    Figure CN121787557A_ABST
Patent Text Reader

Abstract

The invention discloses an audit field-oriented data distillation method based on knowledge enhancement and course learning. The method comprises the steps of inputting unknown audit data into a trained small model; wherein the trained small model is obtained through training according to the large language model, the fusion data and the thinking chain; the fused data is obtained by performing semantic fusion on audit data in a pre-established audit database and corresponding law and policy knowledge; the auditing database is constructed according to bidding and tendering files, contract performance data and compliance review records; the thinking chain is obtained according to the large language model and the fusion data; and obtaining a reasoning result of the unknown audit data. According to the method, a high-quality and highly-adaptive data basis is provided for small model training; the reliability of the model in a complex audit reasoning scene is effectively improved; the migration capability and the generalization level of the small model are enhanced, and the stability and the high efficiency of the data distillation process are realized; and the bottleneck of the small model in auditing application is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a data distillation method based on knowledge enhancement and curriculum learning for the auditing field. Background Technology

[0002] In recent years, with continuous breakthroughs in Large Language Models (LLMs) and their Chain-of-Thought (CoT) capabilities, intelligent technologies have experienced rapid development and widespread application in knowledge-intensive industries such as finance and auditing. Automated compliance review, risk identification, and case attribution—key business processes driven by large models—are propelling the auditing field towards higher levels of intelligence, automation, and standardization. In particular, the introduction of the chain-of-thought paradigm enables models to complete information integration and decision-making reasoning step-by-step when facing complex audit situations, significantly enhancing the depth of understanding and professional judgment of audit content, thus becoming an important technological trend in the field of intelligent auditing.

[0003] However, in actual deployment, mainstream audit expert models both domestically and internationally generally face significant technical bottlenecks. On the one hand, while large-scale language models with massive parameter sets possess excellent business understanding and reasoning capabilities, their training and inference require enormous computing resources, resulting in high operating costs and making it difficult to meet the low-cost and efficient deployment needs of enterprise-level endpoints on a large scale. On the other hand, while lightweight small models have outstanding advantages in terms of flexibility, scalability, and cost control, they often struggle to achieve high levels of reasoning accuracy and business adaptability in diverse and complex real-world audit scenarios, failing to fully meet industry users' requirements for the professionalism and reliability of intelligent audit systems. Summary of the Invention

[0004] To address the aforementioned problems in the existing technology, this invention provides a data distillation method based on knowledge enhancement and curriculum learning for the auditing field.

[0005] The technical problem to be solved by this invention is achieved through the following technical solution: In a first aspect, the present invention provides a data distillation method based on knowledge enhancement and curriculum learning for the field of auditing, the method comprising: Unknown audit data is input into a trained small model; wherein, the trained small model is obtained by training based on a large language model, fused data, and a thought chain; the fused data is obtained by semantic fusion of audit data in a pre-established audit database and corresponding legal and policy knowledge; the audit database includes bidding documents, contract performance information, and compliance review records; the thought chain is obtained based on the large language model and the fused data; The reasoning results of the unknown audit data are obtained.

[0006] Optionally, the training process of the small model includes: When the thought chain meets the preset requirements, the fused data is standardized to obtain the processed fused data. The difficulty level of each processed fused data is classified according to the preset scoring system to obtain fused data with different difficulty levels. The fused data with the difficulty level is input into the small model and trained using a preset loss function to obtain the trained small model.

[0007] Optionally, before standardizing the fused data to obtain processed fused data when the thought chain meets the preset requirements, the method further includes: S31. Select multiple audit data from the audit database as training samples; wherein each training sample has a corresponding first label, and the first label represents the audit scenario of each training sample; S32. The pre-built classifier is trained using the training samples to obtain the trained classifier; S33. The audit data in the audit database is classified using the trained classifier to obtain classified audit data; wherein, the classified audit data is labeled with a second label, and the second label represents the audit scenario corresponding to each classified audit data. S34. Retrieve relevant legal and policy knowledge related to the categorized audit data based on the second tag, and semantically fuse the categorized audit data with the corresponding legal and policy knowledge to obtain fused data; S35. Input the fused data into the large language model to obtain the thought chain; S36. When the thought chain does not meet the preset requirements, return to S34 to regenerate the thought chain until the preset requirements are met.

[0008] Optionally, prior to S31, the method further includes: Collect raw audit text data; The original audit text data is formatted and standardized to obtain the processed audit data; After reviewing the processed audit data, a corresponding database structure is established to obtain the audit database.

[0009] Optionally, the preset scoring system includes: thought chain length, business complexity, regulatory citation degree, and confidence score.

[0010] Optionally, the scoring method for the length of the thought chain is expressed as follows: ; in, The score represents the length of the thought chain corresponding to the currently processed fused data. This represents the number of reasoning steps in the thought process corresponding to the processed fused data. This is the preset maximum number of reasoning steps.

[0011] Optionally, the preset loss function is expressed as follows: ; in, The preset loss function, This represents the total number of training samples sampled from the fused data of the aforementioned difficulty levels in the current training round. For the first The input data corresponding to each training sample For the first The output label corresponding to each training sample This is the conditional probability output for the small model.

[0012] Secondly, the present invention provides a data distillation apparatus for the auditing field based on knowledge enhancement and course learning, the apparatus comprising: The data reasoning module is used to input unknown audit data into a trained small model; wherein, the trained small model is obtained by training based on fused data and thought chain; the fused data is obtained by semantic fusion of audit data in a pre-established audit database and corresponding legal and policy knowledge; the audit database includes bidding documents, contract performance information and compliance review records; the thought chain is obtained based on a large language model and the fused data; The reasoning results of the unknown audit data are obtained.

[0013] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: In the above technical solution, this invention significantly improves the business relevance and professional compliance of the thought chain by injecting audit scenario classification and relevant legal and policy knowledge into the audit data, providing a high-quality and highly adaptable data foundation for small model training; it ensures the high logic and high consistency of the thought chain through standardized processing, effectively improving the reliability of the model in complex audit reasoning scenarios; it enhances the transferability and generalization level of the small model in high-difficulty, low-frequency audit scenarios by utilizing a difficulty grading strategy, achieving stability and efficiency in the data distillation process; and it effectively breaks through the professional bottlenecks and deployment barriers of small language models in audit applications, promoting the intelligent transformation and efficient implementation of the audit industry.

[0014] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0015] Figure 1 This invention provides a data distillation method based on knowledge enhancement and curriculum learning for the auditing field. Detailed Implementation

[0016] To facilitate understanding of the present invention, a brief description of the prior art and the inventive concept of the present invention will be provided first.

[0017] Existing solutions include: This invention provides a method and system for generating audit rectification opinions based on efficient fine-tuning of a large model. The aim is to generate stable, consistent, and targeted audit rectification opinions by training an audit corpus using a large model. The invention standardizes the audit corpus through a corpus management module and optimizes the model training module to handle noise, enabling the system to effectively process noise in audit data. This maintains the stability and consistency of the model even with noisy data, thus generating accurate and targeted audit rectification opinions that adapt to dynamically changing audit scenarios. This method is particularly suitable for dealing with noisy data and complex audit scenarios, effectively improving audit efficiency and ensuring the stability and consistency of generated rectification opinions. The method includes the following steps: 1. First, acquire and organize the audit corpus, including audit report corpus and audit working paper corpus. This corpus will be used for model pre-training and fine-tuning. 2. Construct a pre-trained model and perform self-supervised pre-training on the model based on the audit report corpus, allowing it to initially learn the basic features of audit language. 3. Supervised fine-tuning of the pre-trained model using audit working paper corpus: By learning from paired data of real audit working papers and rectification opinions, a first version of the rectification opinion generation model is obtained. 4. Based on the audit corpus and samples generated during fine-tuning, a noise-enhanced training dataset is created to simulate noisy data (such as format errors or missing information) and to enhance the model's robustness. 5. With the help of the noise-enhanced training data, the first version of the rectification opinion generation model is further optimized and trained to enable the model to effectively handle noisy audit working papers. 6. The audit working papers to be rectified are input into the optimized audit model to generate corresponding audit rectification opinions, thereby achieving automated and accurate rectification opinion generation.

[0018] This invention provides a data processing method and electronic device that aims to automatically generate audit query statements by utilizing a preset database and a large audit model, thereby improving the efficiency and accuracy of audit data retrieval. The method obtains the user's audit query request, queries relevant audit metadata in the database, and integrates it with the user's query content to generate audit query prompts. Finally, it generates the query statement through the large audit model. This method integrates audit metadata with the user's query content, ensuring more accurate query information while avoiding the tedious process of manually writing query statements, significantly improving the automation level of audit data processing. The method includes the following steps: 1. Obtaining the user's audit query request: Receiving the user's query request and extracting the query content, which may be specific data requirements related to auditing. 2. Querying the database to obtain audit metadata: Based on the query content, querying the preset database to obtain audit metadata related to the query content, such as audit models, table information, or field information. 3. Integrating audit metadata with query content: Integrating the query content with the retrieved audit metadata to generate audit query prompts, making them more accurate and complete. 4. Generate audit query statements: Input the audit query prompt information into the audit model, and the system will automatically generate audit query statements to retrieve the required data.

[0019] This invention presents a multi-path reasoning generation method based on fine-grained knowledge retrieval. The method aims to improve the accuracy and interpretability of the reasoning process through refined knowledge extraction and optimized path generation. The method uses a large language model to perform deep semantic analysis on user queries, extracting relevant entities and relationships. A dual-threshold filtering strategy is used to refine the relevance of triples, ensuring a more accurate knowledge base for the reasoning process. Next, the system constructs and optimizes preliminary reasoning paths, generating multiple highly semantically relevant reasoning paths. Finally, a majority voting mechanism is used to determine the consistent answer, thereby improving the reliability and robustness of the reasoning and making it suitable for complex reasoning tasks and dynamically changing scenarios. The method involves using a large language model to perform deep semantic analysis on user queries, extracting entities and relationships between entities, and generating a set of triples. The steps include: 1. Using a large language model to perform deep semantic analysis on user queries, extracting entities and relationships between entities, and generating a set of triples. 2. Scoring the triples based on context, using a dual-threshold strategy to classify the triples, and selecting triples highly relevant to the question. 3. Construct preliminary reasoning paths using these highly relevant triples, and eliminate redundancy and adjust the order through optimization strategies to make the paths more concise and logically coherent. 4. Sample the optimized reasoning paths using Top-k scoring, and select a set of reasoning paths with high semantic relevance. 5. Finally, divide the reasoning paths into multiple complementary path groups, and select the final consistent answer through a majority voting mechanism.

[0020] The shortcomings of existing technologies include: 1. Large models are difficult to deploy flexibly due to high computing power and operational costs, while small models, although lightweight, lack intelligence in multi-step reasoning such as compliance judgment and risk identification. Existing methods generally rely on large models, but large models, due to their high training and inference costs, cannot meet the requirements of low-cost and efficient deployment. In scenarios requiring enterprise or industry-specific solutions, this high operational cost makes it difficult to popularize large models. While small models have the advantage of lightweight design, they perform poorly in complex multi-stage reasoning tasks, especially in tasks requiring multi-step reasoning such as compliance judgment and risk identification. Their limited intelligence and reasoning ability result in insufficient accuracy and professionalism of the results, making it difficult to effectively support industry applications. 2. Existing methods lack guidance from business scenario labels and regulatory knowledge. The reasoning path relies on static triples, resulting in insufficient logical coherence and depth, and a lack of professionalism in the answers. When dealing with complex audit scenarios, current methods often lack effective utilization of business scenario labels and organic integration of regulatory knowledge, making it difficult to achieve professional guidance in the reasoning process. Meanwhile, the optimization of inference paths still mainly relies on simple triple extraction, depending entirely on the reasoning capabilities of large models and lacking a dynamic adjustment mechanism. This results in a lack of coherence, hierarchy, and professionalism in handling complex logical relationships, leading to a lack of depth in the reasoning process and accuracy in the results, thus causing poor professional performance of the model in actual auditing operations. 3. Existing methods do not stratify audit tasks by difficulty, and the model's transfer and generalization performance in high-difficulty, low-frequency scenarios is unstable, making it difficult to cover key business. Existing methods do not fully consider the differences in the difficulty of audit tasks in different scenarios during their design. Therefore, when facing high-difficulty or low-frequency scenarios, the model often exhibits poor transfer and unstable generalization capabilities. This limitation not only reduces the model's adaptability in complex auditing tasks but also limits its application value in changing business environments. Therefore, this invention proposes a data distillation method based on knowledge enhancement and curriculum learning for the auditing field to solve this technical problem.

[0021] Figure 1 This invention provides a data distillation method for the auditing field based on knowledge enhancement and curriculum learning, as illustrated in this embodiment. Figure 1 As shown, the method includes the following steps: S1. Input the unknown audit data into the trained small model; wherein, the trained small model is obtained by training based on the large language model, fused data and thought chain; the fused data is obtained by semantic fusion of audit data and corresponding legal and policy knowledge in the pre-established audit database; the audit database is constructed based on bidding documents, contract performance information and compliance review records; the thought chain is obtained based on the large language model and fused data; and the reasoning result of the unknown audit data is obtained.

[0022] Optionally, the training process for a small model includes: S31. Select multiple audit data from the audit database as training samples; each training sample has a corresponding first label, which represents the audit scenario of each training sample.

[0023] For example, 600 audit data entries can be selected from the audit database as training samples, and each entry can be labeled for 11 typical audit scenarios (such as: prohibition of joint bidding, performance restrictions in specific administrative regions, requirements for local subsidiaries, setting of pre-selected suppliers, restrictions on ownership enterprises, and requirements for expert fee payments). The first label can be manually labeled.

[0024] S32. Use training samples to train the pre-built classifier to obtain the trained classifier.

[0025] For example, a classifier was built based on the google-bert / bert-base-chinese model. 600 training samples were divided into training and test sets in an 8:2 ratio to ensure a balanced distribution of categories across audit scenarios. The PyTorch framework was used in conjunction with the Hugging Face Transformers tool for training, and the model performance was evaluated on the test set. The final trained classifier achieved a classification accuracy exceeding 99%.

[0026] S33. Use the trained classifier to classify the audit data in the audit database to obtain classified audit data; wherein, the classified audit data is labeled with a second label, which represents the audit scenario corresponding to each classified audit data.

[0027] Understandably, the first step is to manually label some audit data in the audit database with the first label, then train a classifier, and then use the classifier to automatically label all the data to obtain the classified audit data.

[0028] For example, the trained classifier is used to classify the audit data in the audit database, thereby adding a second label before each audit data item to ensure that the downstream data is structured and facilitates knowledge injection and thought chain generation.

[0029] S34. Retrieve relevant legal and policy knowledge based on the second tag and classify the audit data, and semantically fuse the classified audit data with the corresponding legal and policy knowledge to obtain fused data.

[0030] For example, for each categorized audit data point, based on its second tag, the system automatically retrieves highly relevant regulatory and policy knowledge (legal provisions and policy documents, etc.) from the locally constructed policy and regulatory knowledge graph. The retrieval process combines keyword expansion and semantic matching, using a vectorized model (such as a BGE embedding model) to calculate semantic similarity between the audit data and the policy and regulatory knowledge graph, selecting the Top-K most relevant regulatory and policy knowledge. A large model (Qwen2.5-72B) is used as an augmented generation system to semantically fuse the retrieved regulatory and policy knowledge with the audit data, generating scenario-appropriate professional knowledge-based supplementary text—the fused data. This forms a fused data set in the form of a triple of "original content + scenario tag + knowledge injection," providing legal and business contextual support for subsequent inference chain generation.

[0031] S35. Input the fused data into the large language model to obtain the thought chain.

[0032] For example, fused data is used as input and fed into a large model (Deepseek-R1) for chained inference generation. A chained template guides the model to output a multi-step inference process, and the generated results are presented in the format of "...". <think>< / think> "Labeling ensures that the thought chain is easy to extract and analyze."

[0033] S36. If the mind chain fails the review, return to S34 to regenerate the mind chain until the preset requirements are met. S37. When the thought chain passes the review, the fused data is standardized to obtain the processed fused data.

[0034] For example, professional auditors review each thought chain generated by the large model, focusing on logical rationality, accuracy of legal citations, and business compliance. Logical breaks, knowledge omissions, and errors in legal application are manually corrected; if incorrect answers are found, the model is required to re-review and regenerate. This ensures all thought chains meet professional review standards, and the closed-loop process improves data consistency and business rigor. All approved thought chains are standardized into processed fusion data in the form of "problem (scenario + content), thought chain, conclusion" triplets. The output uses a conversational format for easy subsequent training and evaluation, for example: Mind chain (based on) <think>< / think> pack): "role": "user" Enter the question and background information; "role": "assistant" Returns the conclusion and complete thought process.

[0035] The fused data, after processing for non-standard formats or redundant content, is further reorganized to unify fields, formats, and alignment, facilitating subsequent fine-tuning of small models and automatic evaluation. Each sample is assigned a unique ID and index to ensure consistency between training and backtracking.

[0036] S38. According to the preset scoring system, each processed fused data is classified into difficulty levels to obtain fused data with difficulty levels.

[0037] Optionally, the preset scoring system includes: thought chain length, business complexity, regulatory citation degree, and confidence score.

[0038] For example, a large model is used to automatically assess the difficulty of each thought chain, comprehensively considering multiple dimensions such as reasoning chain length, business complexity, level of legal application, and answer confidence. The preset scoring system is as follows: Reasoning chain length (L): The number of logical jumps or causal reasoning steps in the chain. More steps indicate a greater required cognitive span, and this accounts for 30% of the score. The scoring method is linear normalization followed by scoring. ; in, The score represents the length of the thought chain corresponding to the currently processed fused data. This represents the number of reasoning steps in the thought process corresponding to the processed fused data. This is the preset maximum number of reasoning steps.

[0039] Business Complexity (B): Identifies the number of professional audit terms and cross-restriction rules involved in the text. The large model outputs a difficulty level score based on the classification, with the score accounting for 40%.

[0040] Regulatory Citation Ratio (R): This assesses whether the reasoning process involves multiple levels of regulations, such as national, provincial, and local regulations, and evaluates the difficulty of their application in the case. It accounts for 20% of the score.

[0041] Confidence score (C): The confidence assessment of the answer determined by the large model itself. The lower the confidence score, the greater the difficulty. The score accounts for 10%.

[0042] Each processed fused data is divided into multiple difficulty levels to obtain fused data with difficulty levels, which prepares a stratified sample pool for course training. In particular, experts are arranged to review samples with extremely high difficulty or low confidence.

[0043] S39. Input the fused data of difficulty level into the small model and train it with the preset loss function to obtain the trained small model.

[0044] Understandably, the data is categorized by difficulty, employing a learning and training scheme that "starts with easier data and progresses to more difficult data, with each round mixing different levels of difficulty." After categorizing the data by difficulty, it is grouped, with the first group containing the largest proportion of low-difficulty data, decreasing progressively in subsequent rounds. Multiple rounds of progressive fine-tuning are then performed on the small model. Initially, training focuses on low-difficulty fused data, gradually introducing higher-difficulty and more complex samples, dynamically adjusting the learning rate and training pace to promote skill transfer. Specifically, the fused data is categorized by difficulty level... Divided into categories based on difficulty layer( ),Right now In the training... The sampling distribution of the fusion data for each round, categorized by difficulty level, is defined as follows: ; in, Indicates the current training round for the th The sampling ratio of fusion data with varying difficulty levels. The initial training phase is of low difficulty. Larger, more difficult in the later stages The proportion is gradually increasing.

[0045] During training, the inference accuracy and generalization ability of the small model under different difficulties and scenarios are continuously monitored. The course structure and training process are iteratively optimized by combining human evaluation and automated metrics. The overall training consistently uses cross-entropy loss as the core optimization objective, with the preset loss function expressed as follows: ; in, For the preset loss function, This represents the total number of training samples sampled from the fused data with varying difficulty levels in the current training round. For the first The input data corresponding to each training sample For the first The output label corresponding to each training sample The conditional probability output of the small model is continuously adjusted by adjusting the sampling distribution ratio as training progresses. This achieves a learning effect with progressively increasing difficulty.

[0046] This invention significantly improves the business relevance and professional compliance of the thought chain by injecting audit scenario classification and relevant legal and policy knowledge into the audit data, providing a high-quality and highly adaptable data foundation for small model training. Standardized processing ensures the high logic and consistency of the thought chain, effectively improving the model's reliability in complex audit reasoning scenarios. A difficulty grading strategy enhances the small model's transferability and generalization level in high-difficulty, low-frequency audit scenarios, achieving stability and efficiency in the data distillation process. It effectively overcomes the professional bottlenecks and deployment barriers of small language models in audit applications, promoting the intelligent transformation and efficient implementation of the audit industry.

[0047] It should be noted that the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention.

[0048] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0049] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings and the disclosure in carrying out the claimed invention. In the description of the invention, the word "comprising" does not exclude other components or steps, "a" or "an" does not exclude a plurality, and "a plurality" means two or more, unless otherwise explicitly specified. Furthermore, while different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce good results.

[0050] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A data distillation method based on knowledge enhancement and curriculum learning for the auditing field, characterized in that, The method includes: Unknown audit data is input into a trained small model; wherein, the trained small model is obtained by training based on a large language model, fused data, and a thought chain; the fused data is obtained by semantic fusion of audit data in a pre-established audit database and corresponding legal and policy knowledge; the audit database is constructed based on bidding documents, contract performance information, and compliance review records; the thought chain is obtained based on the large language model and the fused data; The reasoning results of the unknown audit data are obtained.

2. The data distillation method for the auditing field based on knowledge enhancement and curriculum learning according to claim 1, characterized in that, The training process of the small model includes: When the thought chain passes the review, the fused data is standardized to obtain the processed fused data; The difficulty level of each processed fused data is classified according to the preset scoring system to obtain fused data with different difficulty levels. The fused data with the difficulty level is input into the small model and trained using a preset loss function to obtain the trained small model.

3. The data distillation method for the auditing field based on knowledge enhancement and curriculum learning according to claim 2, characterized in that, Before standardizing the fused data to obtain processed fused data when the thought chain meets the preset requirements, the method further includes: S31. Select multiple audit data from the audit database as training samples; wherein each training sample has a corresponding first label, and the first label represents the audit scenario of each training sample; S32. The pre-built classifier is trained using the training samples to obtain the trained classifier; S33. The audit data in the audit database is classified using the trained classifier to obtain classified audit data; wherein, the classified audit data is labeled with a second label, and the second label represents the audit scenario corresponding to each classified audit data. S34. Retrieve relevant legal and policy knowledge related to the categorized audit data based on the second tag, and semantically fuse the categorized audit data with the corresponding legal and policy knowledge to obtain fused data; S35. Input the fused data into the large language model to obtain the thought chain; S36. If the thought chain fails the review, return to S34 to regenerate the thought chain until the preset requirements are met.

4. The data distillation method for the auditing field based on knowledge enhancement and curriculum learning according to claim 3, characterized in that, Prior to S31, the method further includes: Collect raw audit text data; wherein, the raw audit text data includes bidding documents, contract performance information, and compliance review records; The original audit text data is formatted and standardized to obtain the processed audit data; After reviewing the processed audit data, a corresponding database structure is established to obtain the audit database.

5. The data distillation method for the auditing field based on knowledge enhancement and curriculum learning according to claim 2, characterized in that, The preset scoring system includes: thought chain length, business complexity, regulatory citation degree, and confidence score.

6. The data distillation method for the auditing field based on knowledge enhancement and curriculum learning according to claim 5, characterized in that, The scoring method for the length of the thought chain is as follows: ; in, The score represents the length of the thought chain corresponding to the currently processed fused data. This represents the number of reasoning steps in the thought process corresponding to the processed fused data. This is the preset maximum number of reasoning steps.

7. The data distillation method for the auditing field based on knowledge enhancement and curriculum learning according to claim 2, characterized in that, The preset loss function is expressed as follows: ; in, The preset loss function, This represents the total number of training samples sampled from the fused data of the aforementioned difficulty levels in the current training round. For the first The input data corresponding to each training sample For the first The output label corresponding to each training sample This is the conditional probability output for the small model.

8. A data distillation device based on knowledge enhancement and curriculum learning for the auditing field, characterized in that, The device includes: The data reasoning module is used to input unknown audit data into a trained small model; wherein, the trained small model is obtained by training based on fused data and thought chain; the fused data is obtained by semantic fusion of audit data in a pre-established audit database and corresponding legal and policy knowledge; the audit database includes bidding documents, contract performance information and compliance review records; the thought chain is obtained based on a large language model and the fused data; The reasoning results of the unknown audit data are obtained.