Instruction data enhancement method for improving professional clinical decision-making ability of medical large model

By employing techniques such as multi-path heuristic trajectory generation and reverse instruction deconstruction, the problems of logical breaks and knowledge illusions in the instruction data augmentation process of large medical models have been solved, thereby improving their decision-making ability and robustness in complex medical scenarios.

CN121862446BActive Publication Date: 2026-05-15HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU DIANZI UNIV
Filing Date
2026-03-17
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing large-scale medical models suffer from problems such as logical breaks, mismatch between reasoning trajectories and original instruction constraints, and the tendency to introduce domain knowledge illusions during the instruction data augmentation process, leading to issues with robustness and decision drift in complex medical scenarios.

Method used

By employing multi-path heuristic trajectory generation, semantic clustering and initial trajectory selection, trajectory fragmentation, node local optimization and outward expansion, perplexity-based monotonicity evaluation, and reverse instruction deconstruction, a rigorous medical reasoning chain is constructed to ensure precise alignment and logical consistency between instructions and reasoning trajectories.

Benefits of technology

It enhances the professional clinical decision-making capabilities of the medical big data model, effectively avoids logical illusions, ensures the consistency of the generated data in terms of semantics, logic and knowledge boundaries, and improves the accuracy and reliability of the model in complex medical scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121862446B_ABST
    Figure CN121862446B_ABST
Patent Text Reader

Abstract

The application discloses a kind of instruction data enhancement methods for improving the professional clinical decision-making ability of medical large model, first acquire original medical data, including original instruction and its corresponding answer, and utilize large language model to generate multiple candidate reasoning tracks, and select an initial reasoning track from candidate reasoning track;The initial reasoning track is decomposed into multiple logic nodes, and each logic node is locally optimized and logically expanded to obtain an optimized reasoning track;Based on the optimized reasoning track and its answer, a reconstructed instruction is generated by reverse deduction;Finally, according to the semantic coincidence degree score, the semantic consistency between the original instruction and the reconstructed instruction is evaluated, and the enhanced data triple containing the original instruction or the reconstructed instruction is selected and output according to the evaluation result. The method aims to solve the problems of instruction logic breakage, reasoning track and original instruction constraint mismatch, and the introduction of field knowledge illusion in the enhancement process in the existing data enhancement technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of professional clinical data processing and natural language processing technology, specifically a method for instruction data augmentation to enhance the professional clinical decision-making capabilities of large medical models. Background Technology

[0002] In recent years, artificial intelligence technologies, represented by large language models, have made groundbreaking progress. These models excel in general natural language processing tasks, and their complex reasoning capabilities are becoming a core indicator for measuring model intelligence. Research shows that the performance of large models in specific vertical domains (such as the medical field) is highly dependent on the quality of data during the instruction fine-tuning stage. During this stage, the quality of the instruction data directly determines the upper limit of the model's reasoning ability when handling complex cases. There is a consensus in industry that the "quality" of instruction data is more critical than its "quantity" for improving the model's reasoning ability.

[0003] However, in existing large-scale medical model data engineering, obtaining high-quality instruction data faces severe challenges, mainly in the following aspects:

[0004] 1. The High Barriers to Medical Labeling and the Non-Scalability of Manually Constructed Data: The medical field involves massive amounts of interdisciplinary knowledge, including anatomy, pharmacology, and clinical diagnostics. The construction of authoritative datasets such as MedQA (Medical Knowledge and Clinical Decision Making) and PubMedQA (Biomedical Research) often requires the participation of senior medical experts. This results in high-quality labeled data being not only extremely costly but also having extremely long production cycles, failing to meet the needs of model iteration. Simultaneously, manually constructed data often focuses on common diseases and standardized question answers, lacking sufficient coverage of "long-tail knowledge" such as complex logical reasoning and rare disease diagnosis, limiting the model's ability to handle extremely difficult tasks.

[0005] 2. The "Logical Collapse" and Black-Box Attributes of Existing Automated Augmentation Solutions: To alleviate the problem of data scarcity, the industry commonly adopts automated data augmentation methods based on large models (such as Self-Instruct). However, these methods suffer from serious "model capability traps" in medical scenarios. Most existing technologies drive the model to rewrite instructions through prompt engineering. While this increases text length and modifiers at the semantic level, it fails to construct a rigorous medical reasoning chain within the underlying logical space. The generated complex instructions often exhibit "logical breaks," meaning the reasoning process cannot support the final diagnostic conclusion. Furthermore, the automated augmentation process is essentially a black-box random sampling. When generating reasoning trajectories, the lack of real-time calibration with domain knowledge makes it highly susceptible to "false reasoning"—the model provides the correct answer, but the intermediate reasoning steps are completely incorrect. Once this data enters the fine-tuning stage, it can induce incorrect thinking habits in the model, reducing the reliability of clinical applications.

[0006] 3. Semantic Drift and Knowledge Illusion in Knowledge-Scarce Domains: In knowledge-scarce domains such as smart security and medical monitoring, and rare diagnostic and treatment methods, baseline models are highly susceptible to semantic drift when augmenting data due to insufficient corpus coverage during the pre-training stage. When attempting to simulate complex medical scenarios, models may use fluent language to fabricate "reasonable illusions" that contradict medical common sense. Because medical scenarios have "zero tolerance" for errors, these factual errors hidden in high-quality text are extremely difficult to detect using traditional keyword matching or vector similarity-based filtering methods, posing a significant security risk for industrial deployment.

[0007] 4. Lack of Logical Dimension and Alignment Between Instructions and Inference Trajectories: Current reinforcement paradigms typically follow a one-way mapping from the original question to the final answer, ignoring the intermediate inference trajectories that support medical decisions. This results in models learning only the statistical correlation between clinical features and diagnostic results during fine-tuning, rather than causal logic. In the medical field, minor symptom differences (such as fluctuations in a physiological indicator) often point to completely different inference paths. Due to the lack of deep alignment between "instructions-trajectories-answers," models cannot accurately identify the causal relationship between key clinical features and diagnostic decisions, leading to extremely poor robustness and decision drift when faced with complex and varied inputs from real-world industry applications.

[0008] In summary, existing data augmentation methods struggle to balance the "professionalism," "logical consistency," and "scalability" of medical instruction data. Designing an automated augmentation method specifically for areas with scarce medical knowledge—one that breaks the black-box effect, achieves precise alignment between instructions and reasoning trajectories through reverse logic backtracking, and effectively avoids knowledge illusions—has become a key technological bottleneck in enhancing the professional clinical decision-making capabilities of large-scale medical models and propelling them into practical industrial applications. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings of existing technologies and to provide a method for instruction data augmentation that enhances the professional clinical decision-making capabilities of large-scale medical models. This method aims to address problems in existing data augmentation techniques, such as broken instruction logic, mismatch between inference trajectories and original instruction constraints, and the tendency to introduce domain knowledge illusions during the augmentation process.

[0010] To achieve the above objectives, the technical solution specifically adopted by the present invention is as follows:

[0011] A method for instruction data augmentation to enhance the professional clinical decision-making capabilities of large-scale medical models includes the following steps:

[0012] Step 1: Multi-path heuristic trajectory generation. For the original instruction pair (x, y), N candidate inference trajectories are generated using a large language model under different parameter configurations.

[0013] Step 2, Semantic Clustering and Initial Trajectory Selection: Standardize and centralize the N trajectories generated in Step 1.

[0014] Step 3: Fragment the trajectory. The initial trajectory z0 is broken down into several independent thought nodes according to logical hierarchy or sentence groups. Each node represents a key step in the reasoning process. Define the current trajectory as an ordered set of nodes Z = {n1, n2, ..., n}. k};

[0015] Step 4: Local node optimization and outward expansion, iteratively enhancing each node in the trajectory. Utilizing contextual information (x, y, Z) prev (where Z) prev Generate multiple candidate optimized fragments for the already optimized set of predecessor nodes;

[0016] Step 5: Based on the monotonicity evaluation of perplexity, substitute the candidate optimized segments generated in Step 4 into the original trajectory to construct a temporary complete trajectory z. temp ;

[0017] Step 6: Iteration Termination Determination. Check if the trajectory optimization has reached the preset termination condition. If the indicator reaches the expected threshold, or the number of iterations reaches the upper limit, stop the optimization and output the optimal thought trajectory z. * ;

[0018] Step 7: Reverse instruction deconstruction based on external model. Based on the optimal thinking trajectory and results of Step 6, reversely deduce and reconstruct the original instruction to generate the reconstruction problem x'.

[0019] Step 8: Semantic alignment verification and data reconstruction. Calculate the semantic overlap between the original problem x and the reconstructed problem x'. If the semantic overlap is below a set threshold, it indicates that the original instruction x is too simplistic or contains missing information. In this case, replace the original instruction x with x', which contains more complete constraint information, because x' more accurately reflects z. * The logical constraints inherent in it.

[0020] Preferably, in step 1, the generative diversity and logical reasoning capabilities of the large language model are utilized to construct a set of reasoning paths Z with significant semantic differences for the same input instruction pair (x, y). init The specific steps are as follows:

[0021] Step 1.1: Predict the next token w t At that time, the original log-odds ratio of the model output is l.i The probability distribution after introducing temperature is as follows:

[0022]

[0023] in: For the i-th candidate token, For the context before time t, Let T be the model's original output score for the i-th candidate token, and T be the temperature parameter. Sum all candidate tokens. The "creativity" and "rigor" of the model's generation are controlled by adjusting the Softmax temperature coefficient T during decoding. We employ a hierarchical sampling strategy: setting T=0.2 (low entropy region) focuses on extracting the most robust consensus knowledge within the model, generating logically conservative, high-confidence paths; setting T=0.5 and 0.8 (medium entropy region) balances logical rigor and expressive richness; and setting T=1.2 (high entropy region) encourages the model to make long-range associations, generating unusual reasoning perspectives to uncover complex underlying logic.

[0024] Step 1.2: By explicitly injecting four core logical operators, the model is forced to change its thinking structure: Deductive reasoning starts from general principles and axioms, and derives a specific conclusion y through rigorous logical steps; Inductive reasoning observes the known facts in x, summarizes the potential rules or patterns, and then applies them to the explanation of the answer y; Abductive reasoning starts from the known result y and reverses to find the most reasonable explanation for the preconditions for the occurrence of the instruction x; Analogical reasoning finds a known case A similar to x, and solves the current problem by transferring the reasoning logic of A.

[0025] Step 1.3: For each original pair (x, y), the generator function G iterates through the temperature set T and the heuristic set H. For each pair (T... m H n Given that ∈ T×H, generate a trajectory: The final candidate set generated is Zcand = {z1, z2,..., z}. N}, where N = |T|×|H|×S (S is the number of repeated samples for each configuration).

[0026] Preferably, in step 2, spatial geometry is used to map the thought process of the plain text into a high-dimensional continuous vector space. Outliers are then removed through cluster analysis, and one thought process trajectory is selected as the starting point. The specific steps are as follows:

[0027] Step 2.1: Before vectorization, first process the N trajectories {z1, z2, ..., z...} NPreprocessing is performed to eliminate format interference. Specifically, regularization tools are used to remove redundant modal particles (such as "I think" and "maybe"), while retaining the core predicate logic and converting texts guided by different heuristics into a unified three-part standardized format of "premise-deduction-conclusion".

[0028] Step 2.2: Using the pre-trained deep semantic encoder RoBERTa, the variable-length text trajectory z is processed. i Convert to a fixed-length low-rank eigenvector v i Select the universal value 1024 for the dimension;

[0029] Step 2.3: Use cosine distance to measure the logical deviation between the two trajectories, as shown in the following formula:

[0030]

[0031] in: and Vectors representing the i-th and j-th trajectories. and Let be the magnitude of the vector.

[0032] Step 2.4: To identify the dominant logical paradigm in the trajectory set, the K-Means++ clustering algorithm is used to find K cluster centers C = {c1,...,c...} k The formula for minimizing the sum of squared distances between each data point and the centroid of its cluster is as follows:

[0033]

[0034] Step 2.5: After clustering is complete, calculate the cluster with the most members among all clusters, and calculate the centroid μ of this cluster. main Find the trajectory index i in the original dataset that has the smallest angle with the center vector. * The formula is as follows:

[0035]

[0036] The selected z0 represents the logical consensus reached by the model under multiple sampling distributions. It filters out the illusionary noise caused by excessively high sampling temperatures and avoids overly obscure inference paths.

[0037] Preferably, step 3 transforms the linearly distributed thought trajectory z0 into a structured sequence of logical nodes Z = {n1, n2, ..., n}. k This is to provide precise indices for subsequent node-level optimizations. The specific steps are as follows:

[0038] Step 3.1: Identify explicit operators in the text that represent causality, transition, and progression (such as "because...therefore...", "assuming", "if...then..."). Each closed-loop statement guided by an operator is defined as a logical atom, and the text is segmented based on these logical atoms.

[0039] Step 3.2: Each decomposed node n j Each node is assigned three core attributes, forming a structured tuple: the premise is the input information required for the node's reasoning (from x or the previous node), the core derivation is the specific logical operation performed by the node, and the local conclusion is the intermediate result output by the node, which serves as the input for the next node.

[0040] Step 3.3: Automatically construct a directed acyclic graph between nodes to prevent logical loops or breaks in the optimization process;

[0041] Step 3.4: Use an adaptive merging algorithm to control the granularity of nodes. If the granularity is too large (e.g., a whole segment), the optimization scope will be unclear; if the granularity is too small (e.g., a phrase), it will destroy logical coherence. If a segment only contains factual statements without logical deduction, it is merged with adjacent deduction nodes to ensure that each node is optimizable.

[0042] Preferably, in step 4, the decomposed logic node n j Enhancement is achieved by designing a constrained generation paradigm to guide the model in-depth analysis of individual nodes while maintaining logical chain consistency, thereby increasing the information entropy of the overall trajectory. The specific steps are as follows:

[0043] Step 4.1: Optimize a specific node n j When constructing the input tuple T in = {x, y, Z prev , n j}. Where x is the instruction, y is the answer, and Z is the response. prev This refers to the sequence of nodes that have been optimized before the current node. To avoid the model generating simple semantic restates, non-intertextual constraint factors are explicitly defined in the Prompt. The newly generated candidate fragment n' is required. j With the original fragment n j The text overlap is below the threshold;

[0044] Step 4.2: For each node n j The model will generate M differentiated candidate optimization fragments S. j ={s j,1 , s j,2 , ..., s j,M};

[0045] Step 4.3: Introduce entity coverage measurement based on knowledge graphs or endogenous knowledge bases. If n' j Introducing key intermediate variables that enhance the determinism of the derivation chain is considered a valid extension. While outward expansion is encouraged, the newly generated s j,m The logical continuity constraint must be satisfied. Given the preceding nodes and the current optimization segment, the posterior probability of ultimately deriving the correct answer y should not decrease, i.e.:

[0046]

[0047] Step 4.4: If node n is found during the optimization process... j The information contained is too simplistic to be effectively expanded, so n... j With n j+1 They merge to form a larger logical block.

[0048] Preferably, step 5 involves calculating the perplexity of the generated trajectory using the model and performing a greedy selection among multiple candidate paths to ensure that the optimized trajectory is more robust in terms of semantic distribution. The specific implementation steps are as follows:

[0049] 5.1: In step 4, for node n j Multiple candidate fragments s were generated j,m To assess the impact of these local fragments on the overall logic chain, they need to be re-embedded into the original trajectory. A temporary complete trajectory set Z is constructed. temp For each candidate fragment The assembled sequence is as follows:

[0050]

[0051] Where, {n1, ..., n j-1} represents the optimized predecessor node;

[0052] 5.2: Calculating the sequence using the benchmark model The perplexity level (PPL) measures the difficulty of the model predicting the text sequence. For a text sequence of length L, W = {w1, w2, ..., w...} L} Its perplexity is defined as the exponent of the negative log-likelihood of the sequence, and the calculation formula is as follows:

[0053]

[0054] Where P base The PPL value is based on the basic probability distribution. The lower the PPL value, the stronger the logical coherence of the sequence and the more it conforms to the inherent knowledge distribution of the language model.

[0055] 5.3: Since PPL values ​​tend to be lower for short sentences or simple expressions, a length penalty factor is introduced to encourage outward expansion in step 4. The original PPL is modified using the following formula:

[0056]

[0057] Where L is the total number of tokens. This is used to balance logical depth and conciseness of expression. By adjusting... This can force the system to select candidate nodes that are longer but have higher logical density.

[0058] 5.4: A greedy update strategy is adopted to ensure the monotonicity of trajectory quality during iteration. The current optimal trajectory is defined as z*. The following logical selection is performed:

[0059]

[0060] 5.5: In addition to calculating the PPL of the entire trajectory, the transition probability at the node connection is also calculated. Only when the probability fluctuation at the connection is within a reasonable range (<1) is the node extension considered logically coherent, as shown in the following formula:

[0061] Preferably, step 8 determines the final output data format by quantifying the semantic deviation between the original instruction x and the reconstructed instruction x'. The specific implementation steps are as follows:

[0062] 8.1: To avoid the failure of a single indicator in high-dimensional space, a dual measurement scheme of logical implication and vector space is adopted. A natural language inference model is used to detect whether x and x' are mutually implication. If x' |= x, it means that the reconstructed instruction completely covers the logical scope of the original instruction. An encoding model is used to calculate the cosine similarity between x and x' as a vector space metric.

[0063] 8.2: Based on the calculated overlap score S, perform hierarchical processing: If (High overlap)

[0064] This demonstrates that the inference trajectory z* perfectly matches the original instruction. We retain {x, z*, y} as a standard high-quality sample. If... (Low overlap) indicates that z* introduces incorrect assumptions beyond x. Such data is labeled as noise and directly discarded, or the original instruction x is vague (e.g., "Please analyze the data"), while z* contains more specific analysis steps, resulting in a reconstructed x' that is more accurate than x (e.g., "Please analyze the distribution characteristics of the dataset by combining variance and mean"). If verification confirms that x' has higher logical completeness and can accurately derive y, then x' is used to replace x, constructing an enhanced triple {x', z*, y}.

[0065] This invention has the following characteristics and beneficial effects:

[0066] 1) This invention makes significant improvements in the construction of logical depth and diversity of thought trajectories. Traditional data augmentation methods often rely on simple template rewriting or the generation of a single large-scale black-box model, resulting in shallow and highly homogenized logical thought chains. This invention, by introducing the dual intervention of physical parameters (multi-sampling temperature) and logical operators (deduction, induction, abduction, analogy), constructs a multi-dimensional candidate thought space in the initialization stage. Combined with semantic clustering and centroid anchoring techniques, it can effectively identify and extract the "greatest common divisor of logic" generated by the model, thereby ensuring the robustness and representativeness of the initial trajectory from the source, solving the pain points of high randomness and uncontrollable quality of reasoning paths in existing technologies.

[0067] 2) This invention proposes a monotonic evolution mechanism based on atomic nodes, ensuring the stability of data quality. Addressing the common problems of "semantic redundancy" and "logical degradation" in existing augmentation schemes, this invention decomposes long thought processes into independent logical nodes and implements targeted outward expansion optimization. By introducing perplexity as a quantitative evaluation indicator and combining it with a length normalization factor, a rigorous monotonic non-decreasing update criterion is established. This mechanism mathematically guarantees that each iteration generates a real logical increment, not only improving the richness of the reasoning chain's details but also effectively avoiding the accumulation of logical noise common in traditional augmentation methods.

[0068] 3) This invention proposes a reverse constraint verification method that fundamentally solves the problem of mismatch between instructions and logic. Most existing technologies focus on forward generation (deriving answers from questions), lacking closed-loop verification of the completeness of the reasoning process. This invention proposes a reverse deconstruction technique that traces back from effect to cause, using a high-performance large model to deduce the original instructions from the optimized trajectory and answer. This mechanism acts as a high-precision "logic microscope": on the one hand, it can accurately identify and eliminate inferior trajectories with logical illusions or missing information; on the other hand, for original data with "overly general instructions and lack of detail," it accurately replaces them with reconstructed high-constraint instructions x'. This closed-loop verification ensures that the final output triplet data achieves three-dimensional consistency in semantics, logic, and knowledge boundaries.

[0069] 4) This invention effectively solves the logical illusion problem commonly found in the generation of external general-purpose models in vertical domains. By transforming the data augmentation process from unconstrained heuristic generation to a closed-loop evolution based on model distribution, this invention ensures that, in the field of medical clinical decision-making, the augmented trajectory not only possesses high logical depth but also accurately conforms to the cognitive boundaries of the benchmark model. Experimental results directly demonstrate that this de-illusionization augmentation mode exhibits generalization performance and reliability far exceeding that of directly generated external models with limited corpora. Attached Figure Description

[0070] Figure 1 This is an overall flowchart of the data augmentation method in the embodiments of the present invention.

[0071] Figure 2 This is a flowchart of the multi-path heuristic trajectory generation algorithm in an embodiment of the present invention.

[0072] Figure 3 This is a flowchart of the semantic clustering and initial trajectory selection algorithm in an embodiment of the present invention.

[0073] Figure 4 This is a flowchart illustrating the decomposition of a trajectory into logical nodes in an embodiment of the present invention.

[0074] Figure 5 This is a flowchart illustrating the logic node expansion optimization in an embodiment of the present invention.

[0075] Figure 6 This is a flowchart of the monotonic evaluation algorithm based on perplexity in an embodiment of the present invention.

[0076] Figure 7 This is a flowchart of the problem reconstruction algorithm based on reverse destructuring in an embodiment of the present invention. Detailed Implementation

[0077] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0078] This embodiment provides a method for instruction data augmentation to enhance the professional clinical decision-making capabilities of large medical models, such as... Figure 1 As shown, it includes the following steps:

[0079] Step 1: Multi-path heuristic trajectory generation. For the original instruction pair (x, y), an initial inference path space with high information entropy is constructed by intervening in physical sampling parameters and logical guidance paradigms. For example... Figure 2 As shown, the specific implementation process includes the following steps:

[0080] Step 1-1: The system first constructs a multi-dimensional parameter space, aiming to regulate the output of the large language model from two dimensions: physical sampling distribution and logical construction paradigm. Setting the temperature parameter set:

[0081]

[0082] in: (Set to 0.2): Used to extract the most accurate logical path of the internal weights of the model, ensuring the rigor and high reproducibility of the generated results. (Set to 0.5 and 0.8): Allows the model to select tokens with suboptimal probability distributions, breaking stereotypes and introducing more diverse corpus expressions. (Set to 1.2): Maximizing the uncertainty of the output aims to explore unconventional and imaginative problem-solving approaches, providing a broader sample boundary for subsequent clustering. The system also pre-defines four core logic-guided operators, each corresponding to a specific thought process: the deductive reasoning operator forces the model to follow the path of "general principles → intermediate variables → specific conclusions"; the inductive reasoning operator guides the model to first list the common features of multiple sub-problems and then summarize them to arrive at the final conclusion; the abductive reasoning operator requires the model to start from the known correct answer y and work backward to deduce the necessary pre-existing logical points that led to the result; and the analogical reasoning operator instructs the model to find a case with a semantically similar structure to the current problem x, completing the reasoning through comparison and transfer.

[0083] Step 1-2: For each logical operator H in the above parameter space j This invention designs a specialized instruction encapsulation template. Taking the deductive reasoning operator as an example, the constructed Prompt is as follows: "[Role Setting]: You are a rigorous mathematical logic expert. [Constraints]: You must start from the basic definition or known fact x, and skipping steps is strictly prohibited. Please output in the following format: First state the axiom, then show the derivation formula, and finally verify the logic against the answer y." By filling the original input (x, y) into the above template, the system generates instruction requests for a specific logical dimension.

[0084] Steps 1-3: The system executes a traversal algorithm, sampling each instruction pair (x, y) using a T×H Cartesian product combination. The core gain of this parallel generation strategy lies in its forced introduction of "redundancy" and "complementarity" at the source of data generation. Even if the LLM produces a logical illusion under a single path, paths under other parameter combinations still have a high probability of capturing the correct reasoning logic.

[0085] Step 2, Semantic Clustering and Initial Trajectory Selection: From the redundant and noisy sample set generated by high-temperature sampling, statistical geometric methods are used to identify the logical greatest common divisor, thereby locking in the most robust initial trajectory. For example... Figure 3 As shown, the specific implementation process includes the following steps:

[0086] Step 2-1: The system will process the preprocessed normalized trajectory set Z norm = {z1, z2, ..., z n The input is fed into the pre-trained encoding model E. In this embodiment, the encoding model is preferably Sentence-BERT or other semantic representation models based on the Transformer architecture (such as RoBERTa, SimCSE, etc.). Each trajectory z is then extracted using feature extraction. i It is mapped to its corresponding dense vector space with a dimension of 1024.

[0087] Step 2-2: To extract common logic from the massive number of candidates, a clustering algorithm is used on the vector set V = {v1, v2,..., v...} n The clusters are divided into sub-clusters. The K-Means++ algorithm is used for iterative processing. Compared to traditional K-Means, K-Means++ employs heuristic probabilistic selection when initializing cluster centers, effectively avoiding convergence to local optima and improving the stability of the clustering results. Considering the varying trajectories under different tasks, this embodiment introduces the elbow rule. The system calculates the sum of squared errors (SSE) for different K values, searching for the inflection point where the rate of SSE decline significantly slows down, thereby dynamically determining the optimal number of clusters K.

[0088] Steps 2-3: After clustering is complete, the system performs statistical analysis on the number of members in each cluster. The cluster with the largest number of members is identified and defined as the main logical cluster. After determining the main logical cluster, a real sample that best represents the logical center is extracted from it as the starting point for subsequent iterations.

[0089] Steps 2-4: Calculate the arithmetic mean of all vectors within the principal cluster to obtain the theoretical centroid of the cluster. Since the theoretical centroid may not correspond to any existing text sample in the vector space, to ensure the interpretability and validity of the trajectory, this application performs a nearest neighbor search within the main cluster. The original text corresponding to the vector closest to the centroid is found as the optimal initial inference trajectory z0.

[0090] Step 3: The trajectory is decomposed into logical nodes, transforming the unstructured continuous text trajectory z0 into discrete, independently optimizable atomic units, thus providing precise operation objects for subsequent optimization. For example... Figure 4 As shown, the specific implementation process includes the following steps:

[0091] Step 3-1: Call the Sentence-BERT segmenter to perform preliminary decomposition of the initial optimal trajectory z0. During the decomposition process, irrelevant escape characters and redundant spaces are removed, and punctuation marks are standardized to ensure that each basic sentence group s j They are relatively independent in terms of semantics.

[0092] Step 3-2: To identify true logical turning points in the text, a dual-track detection strategy combining deep semantics and explicit operators is employed. This involves obtaining adjacent sentence groups s. j With s j+1 semantic vector v j With v j+1 The cosine similarity between the two is calculated. A multi-dimensional dictionary of logical operators is pre-defined, including sequence operators, causal operators, and transition operators. Once the above keywords are matched in the sentence group sequence, even if the semantic similarity is high, the system will still identify them as the starting point of the logical structure.

[0093] Step 3-3: After obtaining all candidate splitting points, the system executes the physical splitting logic, aggregating the sequence S into a logical node set Z = {n1, n2, ..., n}. k Each atomic node n i It is defined as the smallest optimizable unit, whose internal structure is indivisible, but it supports overall replacement, rewriting or deletion operations.

[0094] Steps 3-4: Construct a directed acyclic graph where each vertex is an atomic node n. i Directed edge e ij Represents node n i It is node n j The preceding logical dependencies. Record the input context required for the execution of each node (the conclusions of the preceding nodes and known variables), if subsequent steps n j n was cited i The derivation results in the code will be automatically connected.

[0095] Step 4: Logic Node Expansion and Optimization. This involves in-depth analysis of specific logic nodes after discretization, introducing external logic increments to address the common problems of "logic scarcity" or "simple repetition" in generated samples. For example... Figure 5 As shown, the specific implementation process includes the following steps:

[0096] Step 4-1: The system is for node n to be optimized. j A multi-dimensional optimization context was constructed: global constraints include the original task instruction x and the expected final goal / reference answer y; local dependencies are extracted from the logical dependency graph n. j The set of predecessor nodes Z prev= {n1, ..., n j-1 The above information is encapsulated according to a template of task objective + historical reasoning chain + current unit to be optimized, so that the large language model can clearly perceive the current logical progress and optimization boundary during generation, avoiding isolated reasoning out of context.

[0097] Step 4-2: Based on the assembled context, the system calls a higher-performance language model as an enhancer for node n. j Perform multi-path concurrent expansion. Sampling is performed using a non-zero temperature coefficient to stimulate diversity in generated content, resulting in M ​​candidate optimization fragments. A logical depth strategy is explicitly injected into the prompts, requiring the model to perform "boundary condition discussion," "proof by contradiction verification," or "computational step decomposition" for the current node.

[0098] Step 4-3: To prevent large language models from falling into the trap of semantic stagnation, a mandatory physical isolation mechanism is introduced. Calculate candidate segments s. j,m With the original node n j The degree of text overlap between the segments. Preferably, Jaccard similarity or N-gram overlap rate is used as the evaluation index. An overlap threshold of 0.4 is set; segments exceeding this threshold are considered unoriginal repetitions and are discarded directly.

[0099] Step 4-4: After overlap screening, the system further evaluates the value of candidate segments to ensure that the introduced content has a substantial logical contribution. This involves checking whether the segment contains key entities or terms related to the task that were not included in the original node; using regular expressions or syntax tree analysis to identify whether the segment supplements necessary derivation formulas or intermediate calculation steps; and identifying whether logical judgments for special cases such as "if...then..." have been added. Only segments that meet any of the above gain indicators are stored in the valid candidate library as data for subsequent logical reorganization and trajectory synthesis.

[0100] Step 5: A perplexity-based monotonic evaluation algorithm is used to quantify the quality of the optimized candidate trajectories using objective mathematical indicators and to perform monotonicity filtering to ensure that the trajectory evolution process always moves in the direction of increasing logical rigor. For example... Figure 6 As shown, the specific implementation process includes the following steps:

[0101] Step 5-1: Use the baseline model itself as the benchmark judge to evaluate the complete thought process z after assembly. temp The text describes the language probability distribution characteristics. It uses perplexity as the core metric; a lower perplexity value generally indicates that the text's logical coherence better matches the distribution of pre-trained knowledge, resulting in higher semantic fluency and logical coherence.

[0102] Step 5-2: Considering the tendency of large language models to favor short sentences during computation, a length normalization mechanism is introduced to fairly evaluate deep logical chains. The formula is in step 5.3 of the invention content. This is the length penalty factor, with a value of 0.5.

[0103] Step 5-3: Strictly compare the scores of the newly generated candidate trajectories with the currently recorded best trajectory scores. Update the scores of those that are better than the best trajectory. If, after one round of expansion and optimization, the scores of all M candidate trajectories are not better than the best trajectory score, the system executes the rejection update policy.

[0104] Step 6: Based on the problem reconstruction algorithm of reverse deconstruction, establish a strong closed-loop alignment relationship between the instruction and the optimized trajectory, and dynamically reshape the original instruction according to the alignment quality. For example... Figure 7 As shown, the specific implementation process includes the following steps:

[0105] Step 6-1: Call GPT-4 as the inverse destructor. The system only inputs the final optimized trajectory z* and the corresponding conclusion y into it, without providing the original instructions x.

[0106] Step 6-2: After obtaining the reconstructed instruction x', the system compares it with the original instruction x in multiple dimensions to quantify the degree of consistency between the two. Vector semantic similarity (S sim A similarity score of 0.8 is required to be considered similar. Logical implication detection is performed based on a Cross-Encoder architecture. entail ).

[0107] Step 6-3: The system aligns the scores based on the overall score. The final data processing decision is executed. If S > 0.85, the current optimized trajectory z* is determined to perfectly match the original intention, and (x, z*, y) is stored in the high-quality fine-tuning library. Otherwise, the original instruction x is determined to have an unclear description problem, and the system replaces the original instruction x with the reconstructed instruction x', using (x', z*, y) as the final training sample.

[0108] The effects of the present invention will be further explained below with reference to simulation experiments.

[0109] Experimental setup

[0110] The simulation experiments of this invention used MedQA and PubMedQA datasets, which are medical text question answering datasets and biomedical question answering datasets, respectively (only the classification task was used). The classification accuracy was used as the final evaluation metric. The widely used open-source model llama2-7b-chat was employed, and the large model was fine-tuned using the llamafactory framework. The fine-tuned model was then used to directly infer results on the dataset, and comparisons were made to determine the effectiveness of our augmentation method.

[0111] Results Analysis

[0112] To verify the effectiveness of the data augmentation technology proposed in this invention in practical applications, this embodiment conducted a comparative experiment on inference accuracy on medical professional test sets (MedQA and PubMedQA). The accuracy metric directly reflects the model's depth of understanding of clinical knowledge and the precision of its inference logic when processing medical clinical cases. The experimental results are shown in the table below:

[0113]

[0114] As shown in the table above, introducing initial trajectory selection directly improved MedQA accuracy by 2.00%. Trajectories generated directly from large models often contain a large amount of invalid random sampling noise. Through designed semantic clustering and initial trajectory selection, the system successfully selected a better thinking trajectory. This proves that in data augmentation, selection is more important than generation, ensuring that the iteration starting point is located in a high-probability correct logical region. After selecting the initial trajectory, executing optimized thinking trajectories further brought gains of 2.26% (MedQA) and 0.75% (PubMedQA). This step, through physical partitioning and outward expansion of logical nodes, solved the common problem of local logical collapse in long-chain reasoning. The final problem reconstruction stage, even with high optimization, still brought about an improvement of approximately 0.8% - 1.0% in performance for both tasks. This verifies that instruction reconstruction corrects potential ambiguities in the original instructions, making the gradient update direction more consistent during fine-tuning.

[0115] The simulation experiments above demonstrate that the reverse thinking trajectory of this invention ensures that the generated data not only contains the answers but also a rigorous, closed-loop clinical reasoning chain. This provides traceable and auditable reasoning logic for industrial-grade clinical decision support systems, effectively alleviating the "illusion" problem of large models in complex medical scenarios and significantly improving the security of assisted diagnostic systems in practical applications. For scenarios with abundant data scarcity in the medical field (such as rare diseases and emerging diagnostic technologies), this invention generates high-quality instructions through logical backtracking and reconstruction techniques, effectively filling the gaps in long-tail knowledge in existing industrial datasets. Performance on knowledge-intensive tasks such as MedQA and PubMedQA proves that this solution achieves precise enhancement of the capabilities of vertical domain models under data-scarce conditions, possessing strong industrial application value.

[0116] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for augmenting instruction data to enhance the professional clinical decision-making capabilities of large-scale medical models, characterized in that, Includes the following steps: S1. Obtain raw medical data, which includes raw instruction x and its corresponding answer y. Then, construct raw instruction pair (x, y), use the raw instruction pair (x, y) as input to generate multiple candidate reasoning trajectories using a large language model, and select an initial reasoning trajectory from the candidate reasoning trajectories. The method for generating multiple candidate inference trajectories using a large language model is as follows: Step 1.1: By adjusting the temperature parameter T in the large language model decoding process, the diversity and rigor of the generated text are controlled. A low temperature value is set to generate robust logical paths, a medium temperature value is set to balance logical rigor and expressive richness, and a high temperature value is set to stimulate unusual reasoning perspectives. Step 1.2: By injecting a preset logical guidance operator into the large language model, the model is forced to generate a reasoning trajectory according to the specified thought structure; Step 1.3: Traverse different combinations of temperature parameters and logical guidance operators, sample the same original instruction pair (x, y) multiple times, and generate a candidate set containing N candidate inference trajectories; The method for obtaining the initial inference trajectory is as follows: Step 1.4: Perform text preprocessing on the N candidate reasoning trajectories, unify the format, and retain the core logic; Step 1.5: Use a pre-trained semantic coding model to convert each text trajectory into a fixed-length feature vector; Step 1.6: Based on the feature vector, use a clustering algorithm to perform cluster analysis on the N candidate inference trajectories and identify the main logical cluster with the largest number of members; Step 1.7: Calculate the centroid of the main logic cluster and select the candidate inference trajectory with the highest similarity to the centroid vector as the initial inference trajectory; S2. Decompose the initial inference trajectory into multiple logical nodes, and perform local optimization and logical expansion on each logical node to obtain an optimized inference trajectory z. * ; Step S2 specifically includes: Step 2.1: Decompose the initial trajectory z0 into K independent logical nodes according to the logical hierarchy to form an ordered node set Z, where each node contains the preconditions, core derivation and local conclusion; Step 2.2: For each logical node in the node set Z, generate multiple candidate optimization fragments based on the original instruction x, the answer y, and the optimized set of preceding nodes; Step 2.3: Substitute each generated candidate optimization segment into the original node set Z to construct a temporary complete trajectory, and calculate the perplexity of the temporary complete trajectory based on the benchmark model. Select the candidate optimization segment with the best perplexity score to update the current node. Step 2.4: Repeat steps 2.1 to 2.3 to iteratively optimize the logical nodes in the node set Z until the preset termination condition is met, and output the optimal inference trajectory z*. S3. Based on the optimized inference trajectory z * And its answer y, and reverse the derivation to generate a reconstruction instruction x'; S4. Evaluate the semantic consistency between the original instruction x and the reconstructed instruction x' based on the semantic overlap score, and select an enhanced data triple containing the original instruction x or the reconstructed instruction x' based on the evaluation result.

2. The method according to claim 1, characterized in that, The logical guiding operator includes at least one of deductive reasoning, inductive reasoning, abductive reasoning, and analogical reasoning.

3. The method according to claim 1, characterized in that, In step 2.1, the decomposition according to logical hierarchy specifically includes: identifying logical connectors in the text that represent causal, transitional, or progressive relationships, and segmenting the text by taking the closed-loop statement guided by each logical connector as a logical atom.

4. The method according to claim 1, characterized in that, In step S2.3, the confusion score is obtained by length normalization of the confusion value, and the confusion score is a negative correlation function between the confusion value and the length of the trajectory text.

5. The method according to claim 1, characterized in that, In step S3, the optimized inference trajectory z* and the answer y are input by calling an external large language model to generate the reconstruction instruction x'.

6. The method according to claim 1, characterized in that, Step S4 specifically includes: S4.1 Calculate the semantic overlap score between the original instruction x and the reconstructed instruction x'; S4.2 If the semantic overlap score is higher than the preset threshold, output the triple {x, z, y}; if it is lower than the preset threshold, output the triple {x', z, y}.

7. The method according to claim 6, characterized in that, The semantic overlap score is calculated using a dual-metric scheme, which includes: using a natural language reasoning model to detect the logical implication relationship between the original instruction x and the reconstructed instruction x', and using an encoding model to calculate the cosine similarity between the two in the vector space.

8. The method according to claim 7, characterized in that, The logical implication relationship is determined by a natural language reasoning model, and the vector space similarity is obtained by calculating cosine similarity using a semantic encoding model.

9. The method according to claim 1, characterized in that, The enhanced instruction data generated by the method is used to fine-tune the large language model in the medical field to improve its reasoning ability and accuracy in clinical decision-making tasks.