Training data synthesis method and device based on error extrapolation and inference chain analysis, medium and program product
Through the methods of error extrapolation and inference chain analysis, the inference chain errors of large language models are identified and corrected, and high-quality training data is generated. This solves the problem of insufficient training data quality in existing technologies and improves the model's generalization ability and inference reliability.
Patent Information
- Application Number
- CN202511122425.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-08-12
AI Technical Summary
The quality issues of training data in existing technologies lead to insufficient generalization ability and reasoning reliability of large language models. Existing error correction methods are costly and inefficient, and automated generation methods lack transparency, leading to the spread of erroneous information and unstable model training.
Through error extrapolation and reasoning chain analysis, a small language model is used to generate multiple reasoning chains. The chain with the largest error is identified based on preset error evaluation rules, and then corrected using a large language model. The model's answering ability is improved through iterative fine-tuning to construct a high-confidence, highly interpretable training dataset.
The controllability and interpretability of training data are significantly enhanced, and the generated sample data tends to be stable in terms of confidence, suitable for a variety of task scenarios, and provides a solid foundation for model fine-tuning.
Smart Images

Figure CN120611192A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training data synthesis method, device, medium and program product based on error extrapolation and inference chain analysis. Background Art
[0002] During fine-tuning and boosting of Large Language Models (LLMs), the quality of training data directly determines the model's generalization and reasoning reliability. While existing publicly available datasets (such as ShareGPT and Alpaca) are numerous, they often suffer from issues such as high data noise, incomplete reasoning chains, and frequent labeling errors. These issues can negatively impact model performance during actual training and may even cause the model to fit incorrect logic or facts, impacting generation stability and accuracy.
[0003] To improve the quality of training data, some studies have attempted to introduce manual annotation or expert review to improve data quality. Although this method can correct some errors to a certain extent, it relies on a large amount of manual participation, is costly and inefficient, and cannot meet the needs of large-scale training data construction.
[0004] In addition, some automated augmentation methods have been applied to training data expansion, leveraging the language model's inherent generative capabilities to supplement sample data. However, these methods rely solely on a prompt word scoring mechanism for sample screening, lacking transparent interpretability and making it difficult to track the specific paths and causes of errors. This can lead to the further spread and amplification of erroneous information in the training data, impacting the stability and effectiveness of model training. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present application provides a training data synthesis method, equipment, medium and program product based on error extrapolation and reasoning chain analysis, which is at least used to solve the problem in the existing technology that the training data lacks precise error correction and structured reasoning support, making it difficult to meet the high-quality fine-tuning requirements of large language models.
[0006] In order to achieve the above objectives and other advantages, some embodiments of the present application provide the following aspects: In a first aspect, some embodiments of the present application provide a training data synthesis method based on error extrapolation and inference chain analysis, including: Acquire an initial sample set, wherein the initial sample set includes a plurality of task samples, each of the task samples includes at least one question; Using a small language model to perform multiple sampling reasoning on the questions of each task sample, generating corresponding multiple reasoning chains; Calculating an overall error score for each of the reasoning chains based on a preset error evaluation rule, and determining the reasoning chain with the largest overall error score as the reasoning chain to be corrected; Inputting the to-be-corrected reasoning chain and the question in the corresponding task sample into the large language model to generate a corrected answer; The question and the corrected answer constitute a new task sample, and the new task sample is used to fine-tune some parameters of the small language model to improve the model's answering ability; Repeat the sampling inference, error scoring, answer correction, and model fine-tuning process until the rate of change of the performance indicator of the small language model on the preset task evaluation set is lower than the preset rate of change threshold. Then terminate the above iterative process and output the new task sample generated in the last round as the final task sample. A plurality of the final task samples are combined into a training sample set for performing full parameter fine-tuning training on the small language model.
[0007] In a second aspect, some embodiments of the present application further provide an electronic device, comprising: One or more processors; and a memory storing computer program instructions, wherein when the computer program instructions are executed, the processors execute any one of the above-described training data synthesis methods based on error extrapolation and inference chain analysis.
[0008] In a third aspect, some embodiments of the present application further provide a computer-readable storage medium having stored thereon a computer program and / or instructions, which, when executed by a processor, implements a training data synthesis method based on error extrapolation and inference chain analysis as described above.
[0009] In a fourth aspect, some embodiments of the present application also provide a computer program product, comprising a computer program and / or instructions, which, when executed by a processor, implements a training data synthesis method based on error extrapolation and inference chain analysis as described above.
[0010] Compared with related technologies, the solution provided in the embodiment of the present application constructs an error-driven self-optimizing data generation path, which can identify the logical vulnerabilities or areas of insufficient knowledge coverage exposed by the small language model during the reasoning process from the original task samples, and accordingly screen out the reasoning chain with the weakest logical path and the largest answer deviation. Furthermore, by introducing a large language model with stronger reasoning ability, this type of reasoning chain is repaired in a targeted manner to generate optimized answers with reasonable structure and rigorous semantics. Compared with existing training data expansion or error correction methods, this method not only explicitly models the causal structure in the reasoning process, but also can identify and locate reasoning defects in a more fine-grained manner, and guide the large language model to perform targeted rewriting accordingly, significantly enhancing the controllability and interpretability of the sample generation process. With the execution of iterative correction, the output sample data tends to be stable in terms of confidence, and the semantic expression is accurate, constructing a high-confidence, highly interpretable training data set with good versatility and transferability, suitable for various task scenarios such as mathematical calculations, logical reasoning, and natural question answering, providing a solid data foundation for fine-tuning the target model. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other implementation methods can be obtained based on these drawings without paying any creative work.
[0012] Figure 1 1 is a flow chart of a training data synthesis method based on error extrapolation and inference chain analysis provided in an embodiment of the present application; Figure 2 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0013] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0014] First embodiment The first embodiment of the present application relates to a training data synthesis method based on error extrapolation and inference chain analysis. The method is applied to a terminal system as an example for explanation, hereinafter referred to as "system", and the terminal includes: a mobile terminal, a computer terminal or a similar computing device, etc. Figure 1As shown, the method may include the following steps: Step S1: Obtain an initial sample set, which includes multiple task samples, and each task sample includes at least one question.
[0015] For step S1, specifically, the initial sample set can come from an artificially constructed annotated dataset, a public question-and-answer dataset, real question samples extracted from user logs, or a set of candidate questions automatically generated from a pre-trained model. Each task sample contains at least one question expressed in natural language. In the first round of processing, it can only contain questions without providing corresponding reference answers. This set serves as a seed sample set for the data synthesis process, and its scale can be set to dozens to thousands to guide the subsequent iterative generation process. The initial sample set can come from any existing task type, including but not limited to math problems, logical reasoning problems, question-and-answer problems, etc.
[0016] A task sample can also contain multiple questions expressed in natural language, but they must have the same semantic intent. During the reasoning chain generation process, the system can select any question as a representative input, or input different questions into the small language model in sequence to generate multiple reasoning chains. If each task sample contains multiple questions expressed in natural language, and these questions do not have the same semantic intent, the system first performs question semantic clustering and splitting on the task sample, splitting questions with different semantic intents into multiple sub-task samples. Each sub-task sample is input into the small language model separately, and multiple rounds of sampling reasoning are performed to generate a structured reasoning chain.
[0017] In the actual system, the sample initializer module can complete the standardization processing of the initial seed samples and the construction of a unified input structure, support unified input of multi-task formats, and complete sample standardization processing through a unified format adaptation mechanism, which facilitates subsequent sampling reasoning and chain structure generation.
[0018] Step S2: Use the small language model to perform multiple sampling reasoning on the questions of each task sample to generate corresponding multiple reasoning chains.
[0019] Specifically, for each task sample obtained in step S1, a preset small language model is invoked to perform multiple sampling inference operations to generate corresponding multiple inference chains. Preferably, the small language model can be a lightweight language model with a parameter size of approximately 1B to 14B, such as Qwen-1.5B or LLaMA-7B, which has inference capabilities while keeping computing resources manageable.
[0020] Each reasoning chain must meet the format constraints of structured expression, typically taking the form of "step1→step2→…→stepn→answer," where each step represents an intermediate thought step or a local reasoning conclusion, and answer represents the final output answer. This formatted reasoning chain overall forms a semantic structure that conforms to the Chain-of-Thought (CoT) paradigm. This structured output constraint ensures that the number and semantic position of nodes in the graph structure remain consistent across all reasoning chains. That is, for each reasoning chain generated for the same question, the corresponding graph structure will contain the same number of reasoning nodes, connected in a consistent semantic order, thus ensuring topological alignment between reasoning paths.
[0021] In actual systems, the small model generator can perform multiple rounds of reasoning sampling on input task samples in a unified format, and output a set of structured reasoning chains with intermediate conclusions and final answers.
[0022] Step S3: Calculate the overall error score of each reasoning chain based on the preset error evaluation rules, and determine the reasoning chain with the largest overall error score as the reasoning chain to be corrected.
[0023] For step S3, specifically, for the multiple reasoning chains generated for each task sample in step S2, each reasoning chain is first converted into a corresponding directed acyclic graph (DAG) structure based on the structured modeling mechanism, where each node represents an intermediate reasoning step and the edge represents the dependency relationship between steps.
[0024] Building on this graph structure, a dynamic confidence scoring mechanism is introduced to evaluate the credibility of each node in the inference step. Nodes with confidence scores significantly below a preset threshold are identified as potential error nodes. By combining the hierarchical position and dependency paths in the directed graph, an error propagation path is constructed for each potential error node. Weights are assigned to each affected node in the path (e.g., normalized based on parameters such as node depth and out-degree) to characterize the transmission and amplification of inference errors within the chain. The overall error score of each inference chain is calculated by combining the dynamic confidence scores of the affected nodes on the affected path with their corresponding weights. The chain with the highest error score is identified as the one to be corrected.
[0025] In a real-world system, the Error Extraction and Analyzer module automatically constructs the aforementioned inference chain graph structure, identifies potential error nodes, and calculates error path scores. To improve system response efficiency, this module can also process the inference chains of multiple task samples in parallel and quickly identify high-risk nodes and their impact areas through graph traversal and local normalization mechanisms.
[0026] Step S4: Input the reasoning chain to be corrected and the question in the corresponding task sample into the large language model to generate a corrected answer.
[0027] Specifically, step S4 concatenates the inference chain to be corrected identified in step S3 and its corresponding question text q to form a prompt sequence for input to the large language model, thereby constructing a correction request prompt. This prompt can be encapsulated in a natural language template, such as: "The following is a reasoning process, but it may contain errors. Please correct it and output the final answer: {inference chain step sequence}, question: {q}." The inference chain step sequence is a natural language representation of the intermediate steps in the original reasoning chain, concatenated in logical order. This ensures that the large language model can fully perceive the inference context during processing. To guide the model to retain the existing valid logical structure and correct potential incorrect reasoning links, the temperature parameter and repetition penalty parameter in the generation process can be controlled, and a "keep reasonable part" prompt statement can be added to improve the output quality of the answer. The resulting output is the corrected answer, which will be used to replace the original answer in the task sample in subsequent iterations and enter the next round of inference chain sampling and correction.
[0028] Large language models can include pre-trained models with complex logical understanding and error correction capabilities, such as GPT-4, Claude, and Gemini Pro. These models outperform small language models in parameter scale, training data diversity, knowledge breadth, and reasoning depth, and can more effectively identify logical loopholes, ambiguous expressions, and factual or knowledge errors in reasoning chains. Large language models can be deployed via API calls or locally, with flexible configuration based on inference efficiency and data confidentiality requirements.
[0029] In a real-world system, a correction generator can accept an error chain and a question as input, automatically performing a rewriting and correction process to generate high-quality, structured, logically sound corrections. It supports structured input prompt templates to output stable, accurate, and optimized answers.
[0030] Step S5: The question and the corrected answer are combined into a new task sample, and the new task sample is used to fine-tune some parameters of the small language model to improve the model's answering ability.
[0031] For step S5, specifically, the corrected answer a′ generated by the large language model replaces the answer a in the task sample output by the previous iteration, forming a new task sample with the structure preserved but semantically updated, in the form of {q, a′}.
[0032] Fine-tuning training refers to updating only some parameters in the small language model that are strongly related to output generation without changing the overall architecture of the model. This embodiment preferably uses a low-rank adaptation mechanism (LoRA) for fine-tuning training. The LoRA mechanism introduces two low-rank matrices into some key weight matrices in the language model (such as the weights of the attention layer or feedforward layer in the Transformer decoder). While keeping the original weights frozen, only these two small-scale low-rank adaptation matrices are trained, so as to be suitable for lightweight scenarios where training data is frequently replaced within the iteration cycle. In each round of iteration, the system uses the new task samples generated in that round to update the parameters of the small language model through the LoRA mechanism, so that the model is more focused on filling the reasoning defects or knowledge blind spots shown in its previous answers.
[0033] Step S6: Repeat the sampling inference, error scoring, answer correction and model fine-tuning process until the rate of change of the performance indicator of the small language model on the preset task evaluation set is lower than the preset change rate threshold. Then terminate the above iterative process and output the new task sample generated in the last round as the final task sample.
[0034] Regarding step S6, specifically, after each round of iteration, the system evaluates the current model version based on the preset task evaluation set. The task evaluation set is a specially constructed standardized sample set used to objectively measure the actual performance of the model in the target task scenario. Each sample in the evaluation set contains standard questions, clear reference answers, and structured reasoning paths, and has the ability to be used for chain reasoning verification. The evaluation tasks can cover multiple typical scenarios, for example, mathematical reasoning tasks: samples include algebraic derivations, geometric proofs, and probability calculations, all equipped with standard answers and complete mathematical logic steps; or medical diagnosis tasks: samples are annotated by experts, covering case symptoms, laboratory indicators, and standard diagnostic procedures to evaluate the clinical reasoning ability of the model.
[0035] During the evaluation process, the system extracts at least one performance metric, including F1 score, recall rate, and ROC-AUC value, and calculates the rate of change in this performance metric compared to the previous iteration. When this rate of change falls below a preset threshold (e.g., 1%), the model performance is considered stable and the iteration convergence condition is met.
[0036] In an actual system, the iterative controller module can be responsible for the control logic and loop management in the training data synthesis process. This module is used to coordinate each round of sampling reasoning, error assessment, answer correction and fine-tuning training tasks to ensure that the sample iteration process continues to introduce valid information, and to determine whether to terminate the iteration process based on performance indicators. In terms of performance indicator calculation, the quality filter module can perform performance evaluation on the task samples generated in each round of iteration on the preset task evaluation set, and calculate key performance indicators including F1 score, recall rate, ROC-AUC value, etc., to measure the reasoning accuracy and discrimination ability of the current small language model in the task scenario. This module provides the calculation results to the iterative controller module as the basis for judging the termination of the iteration: when the rate of change of the performance indicator is lower than the preset threshold, it is considered that the model performance is converging and the sample generation process is terminated.
[0037] Step S7: Multiple final task samples are combined into a training sample set for full parameter fine-tuning training of the small language model.
[0038] For step S7, specifically, all task samples of the terminated iterative process are aggregated to form a unified training sample set. This sample set has the characteristics of high semantic accuracy, clear reasoning chain structure, and reasonable logical progression, which can fully cover various typical problems and logical variations in the original task distribution. On this basis, the small language model is fine-tuned with full parameter quantity, that is, all weight parameters of the model are back-propagated and gradient updated based on the final training sample set to maximize the model's overall understanding and generalization capabilities of the target task. Through the above-mentioned full-parameter training stage, the small language model can eventually have the ability to perform chain reasoning tasks robustly and accurately in the target task scenario, providing a high-quality model foundation for downstream system deployment.
[0039] Compared with related technologies, the solution provided in the embodiment of the present application constructs an error-driven self-optimizing data generation path, which can identify the logical vulnerabilities or areas of insufficient knowledge coverage exposed by the small language model during the reasoning process from the original task samples, and accordingly screen out the reasoning chain with the weakest logical path and the largest answer deviation. Furthermore, by introducing a large language model with stronger reasoning ability, this type of reasoning chain is repaired in a targeted manner to generate optimized answers with reasonable structure and rigorous semantics. Compared with existing training data expansion or error correction methods, this method not only explicitly models the causal structure in the reasoning process, but also can identify and locate reasoning defects in a more fine-grained manner, and guide the large language model to perform targeted rewriting accordingly, significantly enhancing the controllability and interpretability of the sample generation process. With the execution of iterative correction, the output sample data tends to be stable in terms of confidence, and the semantic expression is accurate, constructing a high-confidence, highly interpretable training data set with good versatility and transferability, suitable for various task scenarios such as mathematical calculations, logical reasoning, and natural question answering, providing a solid data foundation for fine-tuning the target model.
[0040] Second embodiment The second embodiment of the present application relates to a training data synthesis method based on error extrapolation and inference chain analysis. The second embodiment is an improvement on the first embodiment. Specifically, the improvement is as follows: In the second embodiment of the present application, a specific implementation method for enhancing the diversity of sampled inference chain generation is provided, that is, step S2 can further include the following steps: The same question in the task sample is input into the small language model multiple times to trigger its generation diversity mechanism, generating multiple semantically different but structurally consistent reasoning chains. Each reasoning chain consists of the same number of reasoning steps with the same logical structure. Each reasoning step has a preset logical progressive order, and each reasoning step corresponds to an intermediate reasoning result.
[0041] Specifically, for each question in the task sample, the system feeds the small language model the same text content multiple times, triggering its built-in generative diversity mechanism to generate multiple semantically distinct but structurally consistent reasoning chains. This multi-round sampling reasoning process does not rely on input variations, but rather naturally generates semantic diversity through the language model's sampling strategies (such as temperature control and token selection). During each sampling reasoning process, the small language model uses a step-by-step generation approach based on the input question content, sequentially outputting multiple logically progressive reasoning steps to form a complete reasoning chain. Each reasoning chain consists of multiple intermediate reasoning steps, each of which outputs an intermediate result, demonstrating strong coherence and structural clarity.
[0042] For example, the process of generating an inference chain can be divided into the following stages: Step 1: Problem Analysis and Task Decomposition. The model first performs semantic understanding and structural analysis of the input question, identifying key concepts, implicit conditions, and reasoning objectives to form a basic analysis framework.
[0043] Step 2: Intermediate Calculation and Basic Reasoning: Based on the identified key elements, the model performs preliminary logical deduction or numerical calculation to obtain the first key intermediate result, which serves as the supporting node for subsequent reasoning.
[0044] Step 3...step(n): Deeper Reasoning and Reflection. The model continues to develop a multi-step logical progression based on the generated intermediate conclusions. This involves inferring new knowledge, integrating intermediate conditions, identifying potential contradictions, and making retrospective adjustments. This demonstrates the model's ability to dynamically optimize its own reasoning path.
[0045] Step (n+1): The model logically integrates and summarizes all intermediate steps to generate the final answer to the question, thereby forming a complete and explainable reasoning chain.
[0046] It's easy to see that the solution provided by the embodiments of this application, by setting a different random seed or adjusting the temperature parameter for each input question, encourages the user to construct reasoning chains from different angles and multiple thought paths, each with distinct yet logically coherent styles. The generated reasoning chains use a structured, step-by-step expression with a clear, logically progressive structure, facilitating subsequent error location and chain correction operations, helping to improve the interpretability of training data and the transparency of model reasoning.
[0047] Third embodiment The third embodiment of the present application relates to a training data synthesis method based on error extrapolation and inference chain analysis. The third embodiment is an improvement on the first embodiment. The specific improvement is that: in the third embodiment of the present application, a specific implementation method for identifying inference chain errors based on graph structure modeling and error propagation analysis is provided, that is, step S3 can further include the following steps: Step S301: Construct a directed acyclic graph corresponding to the reasoning chain, where each node in the directed acyclic graph represents an inference step and its intermediate inference result, and each edge represents a causal dependency relationship between two inference steps.
[0048] For each inference chain generated by the small language model in step S2, the system structures it into a directed acyclic graph. In this graph, each node represents an inference step and its corresponding intermediate inference result, and each directed edge represents the dependency relationship between inference steps. For example, if the reasoning of step 2 depends on the output of step 1, a directed edge is established in the graph from step 1 to step 2. This graph structure clearly models the upstream and downstream causal paths between inference steps.
[0049] Step S302: Based on the directed acyclic graph, calculate the overall error score of each reasoning chain. The overall error score is used to evaluate the impact of a potential error node in the reasoning chain on subsequent nodes that depend on it.
[0050] Specifically, step S302 may further include the following steps: Step S3021: For each node in the directed acyclic graph, generate an intermediate inference result of the node through multiple samplings, and calculate the average confidence of the node based on the confidence of each sampling result.
[0051] Each time a sample reasoning is performed on the same problem, the resulting reasoning chain may be slightly different, thus constructing multiple directed acyclic graphs. In these different directed acyclic graphs, if a certain logical node (i.e., a specific intermediate reasoning step) appears in multiple graphs, even though its upstream path or expression details may be slightly different, it can be considered to represent a "logically equivalent" step. After evaluating the confidence values of the reasoning results of this logical node in multiple directed acyclic graphs, taking the average value, we can obtain the average confidence of the node under various contextual perturbations. The calculation formula is:
[0052] in, Representation node The average of confidence levels across multiple rounds of sampling; Represents the node corresponding to a certain reasoning step in the directed acyclic graph; Indicates the number of sampling times; Indicates the Inference chain generated by subsampling; Indicates in In the subsampling, the language model is used to calculate the node The confidence value of the generated result.
[0053] The system can calculate the confidence value using a confidence perception method based on adversarial perturbation evaluation. That is, in the process of generating an inference chain using a small language model, a small noise perturbation (such as Gaussian perturbation or perturbation based on the adversarial gradient direction) is introduced into the most recent hidden layer before outputting the answer, and the output results before and after the perturbation are observed to calculate the confidence value.
[0054] The higher the average value, the more "confident" the model is in the reasoning step, and the more stable the node is under different reasoning paths, which can be used as an indicator of node stability or step credibility.
[0055] Step S3022: Calculate the consistency of the node's answer based on the semantic difference between the intermediate reasoning results generated by multiple samplings.
[0056] For the inference chain generated by multiple sampling, for the same nodes that appear in it , which can extract the corresponding intermediate inference results in different samples By calculating the semantic difference between these results, the answer consistency score of the node is constructed, and the calculation formula is:
[0057] in, Representation node Response consistency score; Indicates that the node The number of answers obtained by multiple sampling; Indicates the Subsampling generated nodes The intermediate reasoning answer; Indicates the The function of the semantic difference between the intermediate answer generated by the m-th sampling can be implemented based on the cosine distance calculation method of semantic embedding; Represents a normalization factor used to normalize the sum of all pairwise differences to an average value between 0 and 1.
[0058] The closer the score is to 1, the more consistent the results generated by the node in different reasoning paths, and the higher the reasoning stability. If the score is low, it means that the node may have potential problems of logical instability or semantic ambiguity, which helps to identify error-prone nodes in the reasoning chain.
[0059] Step S3023: Calculate the node confidence consistency based on the fluctuation range of the confidence score corresponding to each sampling.
[0060] For a node in the reasoning chain , under different sampling paths, the node will correspond to multiple confidence scores output by the language model. If these confidence scores are close in multiple sampling paths, it means that the node reasoning is stable and the credibility is high; otherwise, it means that the node may have logical uncertainty or semantic fluctuations. Node The confidence consistency score can be calculated as follows:
[0061] in, Representation node The confidence consistency score of ; L represents the number of sampling times; Indicates that the language model Nodes in subsampling Confidence output of Representation node The average of the confidence scores across all sampled paths.
[0062] The closer the score value is to 1, the smaller the confidence fluctuation is and the higher the node reasoning stability is; the smaller the value is, the more unstable the reasoning process is at that node, which may be a potential source of error propagation.
[0063] Step S3024: Based on the average confidence, answer consistency and confidence consistency, a dynamic confidence score of the node is generated by fusion.
[0064] Based on the three indicators obtained in the previous steps: average confidence, answer consistency and confidence consistency, for any node in the reasoning graph Calculate its dynamic confidence score. This score is used to measure the stability and credibility of the node under multipath disturbance. Define node The dynamic confidence score of is:
[0065] in, Representation node Dynamic confidence score of Representation node The average confidence level of Representation node consistency of answers; Representation node Confidence consistency.
[0066] The above three indicators are multiplied together to form a comprehensive metric value. The higher the value, the more stable and reliable the node is in the reasoning graph, and the less likely it is to be the source node that causes downstream errors. Conversely, a low value may indicate that the node is a potential source of reasoning errors.
[0067] Step S3025: Summarize the dynamic confidence scores of each node in the reasoning chain as the overall error score of the reasoning chain.
[0068] Specifically, step S3025 may further include the following steps: Step S30251: Based on the dynamic confidence scores of the nodes in the directed acyclic graph, identify the nodes whose dynamic confidence scores are lower than a preset confidence threshold, and determine them as potential error nodes in the reasoning chain; Step S30252: Based on the dependency structure of each potential error node in the directed acyclic graph, construct an error propagation path associated with the potential error node; Step S30253: Assign a weight to each affected node in the error propagation path. The weight is normalized based on the affected node's depth in the graph relative to the potential error node and its out-degree information to reflect the attenuation trend of the error propagation along the path. Step S30254: Based on the dynamic confidence score of each affected node and the corresponding weight, a weighted accumulation operation is performed to obtain an overall error score of the potential error node.
[0069] Specifically, for the directed acyclic graph corresponding to the inference chain, the system traverses all nodes in the graph to determine whether their dynamic confidence scores are lower than the preset confidence threshold. Dynamic confidence score for It is marked as a potential error node For each potential fault node , the system identifies the set of subsequent nodes that depend on the node based on the edge structure of the directed acyclic graph , and construct the error propagation path. For each affected node in the error propagation path , calculate its propagation weight , the weight is calculated as:
[0070] Where α represents the depth diffusion factor, a constant greater than 1, which controls the attenuation or amplification of the error as it propagates backward from the root node. The larger the value, the smaller the weight of nodes deeper in the propagation path, reflecting the assumption that errors have a greater impact closer to the source. Representation node Relative to The depth of the hierarchy in the graph, i.e. the number of edges crossed in the path; Representation node The out-degree of a node is the number of subsequent nodes that directly depend on it.
[0071] The denominator is the normalization factor, that is, all affected nodes are summed according to the same rule to ensure that all weights Satisfying normalization: .
[0072] The weighting mechanism comprehensively considers two factors: the propagation distance of a node from the root error node (controlled by exponential decay); and the node's connectivity or influence within the graph (reflected by its out-degree). This mechanism allows the system to more accurately characterize the impact of a potential error node on different nodes in the entire inference chain, providing a foundation for the subsequent calculation of the potential error node's overall error score.
[0073] Combine the dynamic confidence scores of each affected node in the propagation path The corresponding propagation weight , the system calculates the overall error score of the potential error node, which is calculated as follows:
[0074] in, Indicates that it is identified as a potential error node; Indicates that all dependencies The set of successor nodes of ; express Any affected node in is the weight coefficient, indicating Relative to degree of being affected.
[0075] Step S303: Select the inference chain with the highest overall error score from the multiple inference chains as the inference chain to be corrected.
[0076] For multiple reasoning chains generated by the small language model, the overall error score for each chain is calculated. The overall error scores of all chains are then compared, and the chain with the highest overall error score is selected as the current chain to be corrected. This selection strategy is based on the following considerations: a chain with a higher overall error score indicates a greater scope of potential errors within it and a greater degree of error propagation. Prioritizing correction will significantly improve the accuracy of the final answer and overall system performance.
[0077] It is not difficult to see that in the solution provided by the embodiment of the present application, by introducing an overall error evaluation mechanism for the reasoning chain based on dynamic confidence scoring, it is possible to preferentially identify the logical path with the widest error impact and the highest error propagation degree among multiple candidate reasoning paths, significantly improving the system's ability to locate the cause of complex reasoning failures. This mechanism comprehensively considers multi-dimensional features such as the confidence level, semantic consistency, and confidence stability of the reasoning node, and combines the hierarchical depth and out-degree information of the error node to construct an overall error score to measure the propagation intensity of the error to subsequent nodes, thereby achieving an accurate quantitative evaluation of potential key error nodes. Based on this error scoring strategy, the system can prioritize the correction of the most disruptive reasoning chain, thereby improving the overall accuracy and robustness of the generated reasoning process, avoiding indiscriminate repeated attempts between multiple error paths, and optimizing the screening efficiency of the reasoning chain and the priority sorting of answer corrections.
[0078] It should be noted that the third embodiment of the present application may also be an improvement based on any one or more of the first to second embodiments.
[0079] Fourth embodiment The fourth embodiment of the present application relates to a training data synthesis method based on error extrapolation and inference chain analysis. The fourth embodiment is an improvement on the first embodiment. Specifically, the improvement is that in the fourth embodiment of the present application, a performance change rate-driven iteration termination mechanism is provided, that is, step S6 can further include the following steps: Step S601: After completing each round of model fine-tuning training, use the small language model of the current round to perform test tasks on a preset task evaluation set to obtain the performance indicator score of the current round; Step S602: Compare the performance index score of the current round with the performance index score of the previous round, and calculate the change rate of the performance index; Step S603: Determine whether the rate of change of the performance indicator is less than the preset rate of change threshold. If so, execute step S604: terminate the iterative optimization process of the current task sample, and output the new task sample of the current round as the final task sample; if not, continue to execute the next round of sampling reasoning, error scoring, answer correction and model fine-tuning process.
[0080] Specifically, after completing partial parameter fine-tuning training for the small language model in each round, the updated small language model from the current round is used to perform evaluation tests on a pre-set task evaluation set to obtain the model performance index scores for that round. Performance indicators may include, but are not limited to, F1 score, recall rate, ROC-AUC value, etc. The performance index scores of the current round are compared with those of the previous round, and the rate of change of the performance index is calculated as a basis for measuring the trend of model capability improvement.
[0081] Determine whether the rate of change of the above performance indicators is less than a preset rate of change threshold (for example, set to 0.5%, 1.0%, etc.). If this condition is met, it means that the model performance has stabilized and entered the convergence stage. At this time, execute step S604; otherwise, continue to execute the next round of sampling reasoning, error scoring, answer correction and fine-tuning of the small language model.
[0082] It's easy to see that the solution provided by the present embodiment achieves precise control of the training sample optimization process and adaptive adjustment of data quality by introducing a dynamic termination mechanism based on the rate of change of performance indicators during the data synthesis iteration process. This mechanism effectively avoids the sample redundancy and overfitting risks associated with blind iteration, ensuring that each round of data updates is driven by actual performance gains, thereby improving the efficiency and scientific nature of data optimization.
[0083] It should be noted that the fourth embodiment of the present application may also be an improvement based on any one or more of the first to third embodiments.
[0084] Fifth embodiment The fifth embodiment of the present application relates to a training data synthesis method based on error extrapolation and inference chain analysis. The fifth embodiment is an improvement on the first embodiment. Specifically, the improvement is as follows: In the fifth embodiment of the present application, a specific implementation method for semantic reconstruction for a large language model is provided, that is, step S4 can further include the following steps: Step S401: Express each reasoning step in the reasoning chain to be corrected in series in natural language to construct a reasoning chain input sequence for context understanding; Step S402: Concatenate the inference chain input sequence with the question in the task sample to form an input context for the large language model; Step S403: Control the large language model to retain the valid logical structure in the reasoning chain to be corrected during the generation process, and rewrite the potential erroneous content in the reasoning chain to be corrected to generate a semantically coherent and logically reasonable corrected answer and output it.
[0085] Specifically, each reasoning step may originally be presented in the form of code, expression, templated text, etc. Each reasoning step in the reasoning chain to be corrected is described in series in natural language to form a coherent and readable reasoning chain input sequence, which facilitates the large language model to perform global modeling and semantic understanding of the context.
[0086] After constructing the input sequence for the reasoning chain, it is concatenated with the question in the original task sample to form a unified input context. This context can be used to guide the large language model in understanding the logical relationship between the original question context and the existing reasoning path. After completing the context construction, the system controls the large language model to perform conditional generation on the input, requiring the model to logically rewrite and semantically correct any potential errors while preserving the valid logical structure of the original reasoning chain. This control mechanism can include: prompt templates ("Please indicate whether there are any logical errors in the above reasoning and provide a more reasonable reasoning process"); constraining the model output structure to be consistent with the input format, preserving the naming and numerical association of intermediate variables; and employing a retain-rewrite mechanism to preserve as many intermediate steps as possible and only adjust incorrect paths. This approach not only corrects sequence errors in the original reasoning but also maintains the reusability of some intermediate results, avoiding unnecessary rewriting of already correct parts, thereby effectively improving data sample quality.
[0087] It is not difficult to find that in the solution provided by the embodiment of the present application, the reasoning chain to be corrected is expressed in series in the form of natural language, and is spliced with the task problem to form a unified context input, which can fully stimulate the large language model's ability to understand the global semantics and logical structure. During the generation process, the model is controlled to retain the effective reasoning path and only the potential error links are rewritten in a targeted manner. This not only effectively avoids redundant disturbances to the correct part, but also significantly improves the semantic coherence and logical rationality of the generated answers. Compared with traditional data cleaning and statement replacement methods, this embodiment introduces a structured context-driven correction mechanism, which enables the training samples to achieve high-quality updates of semantic information while maintaining the stability of the original problem representation structure, significantly improving the availability and generalization ability of the training data set.
[0088] It should be noted that the fifth embodiment of the present application may also be an improvement based on any one or more of the first to fourth embodiments.
[0089] The step division of the above various methods is only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process without changing the core design of the algorithm and process are all within the scope of protection of this application.
[0090] In addition, some embodiments of the present application further provide an electronic device. The electronic device may be various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device may also be various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.
[0091] The electronic device includes: one or more processors; and a memory storing computer program instructions, and when the computer program instructions are executed, the processor executes a training data synthesis method based on error extrapolation and inference chain analysis as provided in any one or more of the above embodiments. Figure 2 An exemplary structural diagram of the electronic device is disclosed. The electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, with each device providing some of the necessary operations. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0092] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103 and the output device 1104 may be connected via a bus or other means. Figure 2 The bus connection is taken as an example.
[0093] Input device 1103 can receive input digital or character information and generate key signal input related to user settings and function control of the electronic device. Examples include a touch screen, keypad, mouse, trackpad, touchpad, pointing stick, one or more mouse buttons, trackball, joystick, and other input devices. Output device 1104 may include a display device, auxiliary lighting devices (e.g., LEDs), and tactile feedback devices (e.g., vibration motors). The display device may include, but is not limited to, a liquid crystal display, a light emitting diode display, and a plasma display. In some embodiments, the display device may be a touch screen.
[0094] To provide user interaction, the electronic device may be a computer. The computer includes a display device (e.g., a cathode ray tube or LCD monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse) through which the user can provide input to the computer. Other types of devices may also be used to provide user interaction; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback), and input from the user may be received in any form (e.g., voice input or tactile input).
[0095] In an embodiment of the present application, a computer program / instruction is stored on a computer-readable medium. When executed by a processor, the computer program / instruction implements a training data synthesis method based on error extrapolation and inference chain analysis provided in any one or more of the above-described embodiments. The computer-readable medium may be included in the electronic device described in the above-described embodiments, or it may exist independently and not be incorporated into the device. The computer-readable medium carries one or more computer-readable instructions.
[0096] The memory 1102 can be used as a non-transitory computer-readable storage medium to store non-transitory software programs, non-transitory computer executable programs, and modules. The processor 1101 executes the non-transitory software programs, instructions, and modules stored in the memory 1102 to execute various functional applications and data processing of the server, thereby implementing the program instructions / modules corresponding to the method provided in any one or more of the above embodiments of the present application.
[0097] The memory 1102 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 1102 may optionally include a memory remotely located relative to the processor 1101, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0098] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above. Computer-readable media may be, for example, but not limited to: electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device.
[0099] Computer-readable media includes both permanent and non-permanent, removable and non-removable media, and can be implemented using any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technology, compact discs, digital versatile discs or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0100] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network or a wide area network, or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0101] In the above embodiments, all or part of the steps or functions of the present invention may be implemented using software, hardware, firmware, or any combination thereof. For example, implementation may be achieved using a dedicated integrated circuit, a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application may be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) may be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, a floppy disk, or the like. In addition, some steps or functions of the present application may be implemented using hardware, for example, as a circuit that cooperates with a processor to perform the various steps or functions.
[0102] The computer program product provided in the embodiments of the present application includes one or more computer programs / instructions that, when executed by a processor, fully or partially produce the processes or functions described in accordance with the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).
[0103] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-specific system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0104] The scope of this application is defined by the appended claims rather than the foregoing description and is therefore intended to encompass within this application all changes that come within the meaning and range of equivalents of the claims. Any reference signs in the claims should not be construed as limiting the claims to which they relate. In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in a device claim may also be implemented by one unit or device through software or hardware. Words such as "first" and "second" are only used to distinguish the description and do not indicate any particular order, nor should they be understood as indicating or implying relative importance.
[0105] The above descriptions are merely specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art may easily propose variations or substitutions within the technical scope disclosed in the present application, and such variations or substitutions shall be encompassed within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims, and the above descriptions shall be regarded as exemplary and non-limiting.
Claims
1. A training data synthesis method based on error extrapolation and inference chain analysis, characterized in that: include: Acquire an initial sample set, wherein the initial sample set includes a plurality of task samples, each of the task samples includes at least one question; Using a small language model to perform multiple sampling reasoning on the questions of each task sample, generating corresponding multiple reasoning chains; Calculating an overall error score for each of the reasoning chains based on a preset error evaluation rule, and determining the reasoning chain with the largest overall error score as the reasoning chain to be corrected; Inputting the to-be-corrected reasoning chain and the question in the corresponding task sample into the large language model to generate a corrected answer; The question and the corrected answer constitute a new task sample, and the new task sample is used to fine-tune some parameters of the small language model to improve the model's answering ability; Repeat the sampling inference, error scoring, answer correction, and model fine-tuning process until the rate of change of the performance indicator of the small language model on the preset task evaluation set is lower than the preset rate of change threshold. Then terminate the above iterative process and output the new task sample generated in the last round as the final task sample. A plurality of the final task samples are combined into a training sample set for performing full parameter fine-tuning training on the small language model.
2. The training data synthesis method according to claim 1, characterized in that The step of performing multiple sampling reasoning on the questions of each task sample using the small language model to generate corresponding multiple reasoning chains includes: The same question in the task sample is input into the small language model multiple times to trigger its generation diversity mechanism, thereby generating multiple semantically different but structurally consistent reasoning chains, wherein each of the reasoning chains consists of the same number of reasoning steps with the same logical structure, each reasoning step has a preset logical progressive order, and each reasoning step corresponds to an intermediate reasoning result.
3. The training data synthesis method according to claim 1, characterized in that The step of calculating the overall error score of each of the reasoning chains based on a preset error evaluation rule and determining the reasoning chain with the largest overall error score as the reasoning chain to be corrected includes: Constructing a directed acyclic graph corresponding to the reasoning chain, wherein each node in the directed acyclic graph represents an inference step and its intermediate inference result, and each edge represents a causal dependency relationship between two inference steps; Calculating an overall error score for each of the inference chains based on the directed acyclic graph, wherein the overall error score is used to evaluate the degree of impact of a potential error node in the inference chain on subsequent nodes that depend on it; From the multiple reasoning chains, the reasoning chain with the highest overall error score is selected as the reasoning chain to be corrected.
4. The training data synthesis method according to claim 3, characterized in that: The step of calculating the overall error score of each of the inference chains based on the directed acyclic graph comprises: For each node in the directed acyclic graph, generate an intermediate inference result of the node through multiple samplings, and calculate an average confidence of the node based on the confidence of each sampling result; Calculating the consistency of the answer of the node based on the semantic difference between the intermediate reasoning results generated by multiple samplings; Calculating the confidence consistency of the node based on the fluctuation range of the confidence score corresponding to each sampling; Based on the average confidence, the answer consistency and the confidence consistency, a dynamic confidence score of the node is generated by fusion; The dynamic confidence scores of the nodes in the reasoning chain are aggregated to serve as the overall error score of the reasoning chain.
5. The training data synthesis method according to claim 4, characterized in that: The step of aggregating the dynamic confidence scores of the nodes in the inference chain as the overall error score of the inference chain includes: Based on the dynamic confidence scores of the nodes in the directed acyclic graph, identifying the nodes whose dynamic confidence scores are lower than a preset confidence threshold, so as to determine them as potential error nodes in the reasoning chain; Constructing an error propagation path associated with each potential error node based on a dependency structure of the potential error node in the directed acyclic graph; Assigning a weight to each affected node in the error propagation path, wherein the weight is normalized according to the hierarchical depth and out-degree information of the affected node relative to the potential error node in the graph to reflect the attenuation trend of the error propagation along the path; Based on the dynamic confidence score of each affected node and the corresponding weight, a weighted accumulation operation is performed to obtain the overall error score of the potential error node.
6. The training data synthesis method according to claim 1, characterized in that: The steps of repeatedly performing the sampling inference, error scoring, answer correction, and model fine-tuning processes until the rate of change of the performance indicator of the small language model on the preset task evaluation set is lower than a preset rate of change threshold, terminating the above iterative process, and outputting the new task sample generated in the last round as the final task sample include: After completing model fine-tuning training in each round, the small language model of the current round is used to perform test tasks on the preset task evaluation set to obtain the performance indicator score of the current round; Comparing the performance indicator score of the current round with the performance indicator score of the previous round, and calculating the change rate of the performance indicator; Determine whether the rate of change of the performance indicator is less than the preset change rate threshold. If so, terminate the iterative optimization process of the current task sample and output the new task sample of the current round as the final task sample; if not, continue to execute the next round of sampling reasoning, error scoring, answer correction and model fine-tuning process.
7. The training data synthesis method according to claim 1, characterized in that: The step of inputting the to-be-corrected reasoning chain and the question in the corresponding task sample into the large language model to generate a corrected answer includes: Expressing each reasoning step in the to-be-corrected reasoning chain in series in natural language to construct a reasoning chain input sequence for context understanding; Concatenate the inference chain input sequence with the question in the task sample to form an input context for a large language model; The large language model is controlled to retain the valid logical structure in the to-be-corrected reasoning chain during the generation process, and to rewrite the potential erroneous content in the to-be-corrected reasoning chain to generate and output a semantically coherent and logically reasonable corrected answer.
8. An electronic device, characterized in that: The electronic device comprises: One or more processors; and a memory storing computer program instructions, wherein when the computer program instructions are executed, the processor executes the training data synthesis method based on error extrapolation and inference chain analysis according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program and / or instructions stored thereon, characterized in that: When the computer program and / or instructions are executed by a processor, the training data synthesis method based on error extrapolation and inference chain analysis according to any one of claims 1 to 7 is implemented.
10. A computer program product comprising a computer program and / or instructions, characterized in that When the computer program and / or the instructions are executed by a processor, the training data synthesis method based on error extrapolation and inference chain analysis as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-model collaborative distillation and dynamic fine tuning model training method and system
CN119443313A
Model distillation method and system based on teacher model and situational reasoning
CN119539011A
Natural language question and answer framework, method and device based on self-reflection
CN120011491A
Model distillation method, reply information generation method and device
CN120218182A
Cited By
Classification task-oriented data generation method based on big and small model collaboration
CN120822037A
Trajectory prediction system and method based on multi-modal thinking chain
CN120873507A
Data set construction method oriented to special reasoning model
CN121257759A
A data set construction method for a special reasoning model
CN121257759B
Large language model reasoning enhancement method, system, equipment and medium
CN121457647A