Language model reasoning optimization method driven by logic mistake recognition and repair
By using a logic fallacy identification and repair-driven approach, the model is divided into fallacy misconception and repair verification stages. Robust argument chains are generated and closed-loop iterative optimization is performed. This solves the problems of incoherent reasoning and frequent fallacies in LLMs in complex logical reasoning tasks, and improves the model's logical fallacy handling ability and reasoning stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HAINAN FENGQI YUNHANG INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-15
AI Technical Summary
Existing large-scale language models (LLMs) lack a deep understanding of logical fallacies and the ability to systematically identify and correct them when dealing with multi-step logical reasoning, abstract thinking, and cross-domain knowledge integration. This leads to incoherent reasoning and frequent logical fallacies, making them unsuitable for high-reliability scenarios.
By employing a logic fallacy identification and repair-driven approach, the model is broken down into fallacy misconception and repair verification stages. A three-element hierarchical analysis is conducted to generate a robust argument chain. Combined with expert heuristic scoring and automated indicator verification, the optimal reasoning path is iteratively optimized in a closed loop to improve the model's fallacy handling capabilities.
It significantly reduces the probability of logical fallacies, improves the logical rigor, fluency, and stability of the reasoning chain, and ensures that the model maintains high coherence and robustness in complex logical reasoning tasks.
Smart Images

Figure CN122047477A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of natural language processing technology, and more specifically, it relates to a language model reasoning optimization method driven by logical fallacy identification and repair. Background Technology
[0002] In recent years, large-scale language models (LLMs) have made significant progress in the fields of natural language understanding and generation. With their powerful semantic capture and language expression capabilities, they have been widely applied to various intelligent tasks such as reading comprehension, decision support, and content generation. As a core cognitive ability supporting various intelligent tasks, logical reasoning is a key indicator for measuring the intelligence level of a model. With the emergence of instruction-based fine-tuning models such as GPT-4 and DeepSeek, the performance of LLMs on basic reasoning tasks has been improved, further expanding their application boundaries.
[0003] However, LLMs still face significant limitations when dealing with complex tasks such as multi-step logical reasoning, abstract thinking, and cross-domain knowledge integration. These models often fall into structural confusion and incoherent reasoning, and even exhibit various logical fallacies, such as incorrect generalizations, subjective assumptions, and error dilemmas. Some models may get stuck in a vicious cycle due to logical fallacies during deep reasoning; even if they recognize the reasoning deviation and attempt to correct it, they may still repeat the same fallacy, making it difficult to form a reasonable reasoning chain. The core reason for this problem is that existing LLMs lack a deep understanding of logical fallacies and the ability to systematically identify and correct them, making it impossible to learn from erroneous reasoning and optimize their own reasoning paths.
[0004] Existing research largely focuses on deductive reasoning, natural language reasoning, and multiple-choice reasoning tasks, emphasizing only the evaluation of model performance in specific scenarios while neglecting the crucial aspect of "learning from erroneous reasoning." Furthermore, while existing datasets like LFUD cover logical fallacies, they lack structured fallacy deconstructions and correction processes, hindering models from systematically mastering fallacy identification and correction methods. In addition, traditional training frameworks fail to integrate fallacy reasoning chains with effective correction processes and lack comparative learning mechanisms, preventing models from fully absorbing lessons learned from erroneous reasoning and limiting the improvement of their reasoning abilities.
[0005] In practical applications, the insufficient logical reasoning ability of LLMs seriously affects their applicability in scenarios with high reliability requirements. Therefore, there is an urgent need for a technical solution that enables the model to systematically identify and analyze logical fallacies and learn from erroneous reasoning to optimize the reasoning path, so as to solve the problems of incoherent reasoning and frequent logical fallacies in existing LLMs and improve the accuracy and rationality of the model in complex logical reasoning tasks. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a language model reasoning optimization method driven by logical fallacy identification and repair, thereby resolving the technical issues of incoherent reasoning and frequent logical fallacies in traditional LLMs.
[0007] The purpose and effectiveness of the language model reasoning optimization method driven by logical fallacy identification and repair of the present invention are achieved by the following specific technical means: A language model reasoning optimization method driven by logical fallacy identification and repair includes the following steps: S1: Based on the original logical fallacy data, perform fallacy misconstruction and repair process operations, and perform hierarchical analysis of the three elements of fallacy on the original logical fallacy data to obtain the three-element analysis results. S2: Based on the analysis results of the three elements, perform argumentation and reasoning operations to obtain a robust argumentation chain. Based on the obtained robust argumentation chain, perform the first fine-tuning operation of the large language model to obtain a language model with preliminary error handling capabilities. S3: Based on a language model with preliminary error handling capabilities, candidate reasoning chains are obtained. Based on the candidate reasoning chains, expert heuristic scoring and automated indicator verification are performed to obtain the optimal reasoning path. S4: Obtain the original question based on the original logical fallacy data, and obtain the optimized inference chain by performing closed-loop iteration of the inference chain based on the optimal inference path and the original question; S5: Perform secondary fine-tuning of the large-scale language model based on the optimized inference chain to obtain a language model with optimized inference capabilities.
[0008] According to a preferred embodiment, the step of performing fallacy misconception and repair procedures based on the original logical fallacy data, and obtaining the three-element analysis results by performing a hierarchical analysis of the original logical fallacy data, includes: The three elements are represented as flawed premises, invalid reasoning patterns, and flawed conclusions. The process of understanding and correcting logical fallacies is broken down into a fallacy deconstruction stage and a correction and verification stage. In the fallacy deconstruction stage, the type of each fallacy instance is named, and a hierarchical decomposition is carried out in the order of defective premise, invalid reasoning mode, and conclusion loophole to obtain the analysis results of the three elements of fallacy.
[0009] According to a preferred embodiment, the step of performing argumentation and reasoning operations based on the three-element analysis results to obtain a robust argument chain, and then performing an initial fine-tuning operation on a large language model based on the obtained robust argument chain to obtain a language model with preliminary error handling capabilities, includes: Based on the results of the three-element analysis, counterfactual conditions are generated to obtain a preliminary inference chain. Based on the initial inference chain, a multi-level refutation path construction operation is performed to obtain multiple refutation paths; A robust argument chain is obtained by generating an effective argument chain based on multiple rebuttal paths; The first fine-tuning of a large language model is performed based on a robust argument chain to obtain a language model with preliminary error handling capabilities.
[0010] According to a preferred embodiment, the counterfactual condition generation operation based on the three-factor analysis results obtains a preliminary reasoning chain, and the multi-level refutation path construction operation based on the preliminary reasoning chain obtains multiple refutation paths, including: Based on the three-element analysis results, the factual statements, reasoning rules and conclusions in the original reasoning chain are extracted. A chain-like prompt is designed to drive the large language model to output contradictory hypothetical scenarios and candidate answers. Counterfactual hypotheses are constructed by reversing, replacing or referencing similar cases in the knowledge base of key facts to obtain the initial reasoning chain. For each counterfactual hypothesis, a multi-level deduction is performed to construct a refutation path. For each hypothesis, retrieval enhancement technology is used to obtain relevant facts or cases, and subsequent conclusions are simulated and deduced along the causal / logic chain. Multiple refutation paths are generated and the self-consistency of the original conclusion is verified, thus obtaining multiple refutation paths. The construction of the refutation path is represented by refuting the core flaw through real-world examples.
[0011] According to a preferred embodiment, the step of generating a valid argument chain based on multiple rebuttal paths to obtain a robust argument chain, and then performing an initial fine-tuning operation on a large language model based on the robust argument chain to obtain a language model with preliminary error handling capabilities, includes: The evaluation path is integrated according to a predefined reasoning paradigm to meet structural constraints and align factual knowledge. By optimizing the target selection path, the merits of counterfactual answers and original answers are compared and the original arguments are revised to obtain a robust argument chain. The LFUD-R dataset, built upon robust argument chains, is used to train large-scale language models, enabling them to master error identification, deconstruction, and repair methods, thus obtaining language models with preliminary error handling capabilities.
[0012] According to a preferred embodiment, the step of obtaining candidate reasoning chains based on a language model with preliminary error handling capabilities, performing expert heuristic scoring and automated metric verification based on the candidate reasoning chains, and obtaining the optimal reasoning path includes: Based on a language model with preliminary error handling capabilities, a multi-candidate thought chain generation operation is performed, generating three different candidate reasoning chains for each question. Based on the three different candidate reasoning chains, an expert heuristic scoring operation is performed, scoring from 1 to 5 points according to four dimensions: logical completeness, clarity of thought path, redundancy control, and anti-interference ability, with each dimension having a weight of 40%, 30%, 15%, and 15%, respectively, and then weighted summed to obtain the expert heuristic aggregate score. Automated indicator verification is performed based on three different candidate inference chains. The BLEU-4 and ROUGE-L indicators are used to quantitatively evaluate the coherence and information redundancy of the inference chains and obtain the automated indicator aggregation score. Based on expert-heuristic aggregated scores and automated index aggregated scores, the optimal inference chain is selected. The final score is calculated by weighting and summing the scores according to preset weights, and the inference chain with the highest score is selected as the optimal inference path.
[0013] According to a preferred embodiment, the step of calculating the final score by weighted summation according to preset weights includes: The final score is calculated as follows: Final Score = 0.6 × Expert Heuristic Aggregated Score + 0.4 × Automated Metric Aggregated Score.
[0014] According to a preferred embodiment, the step of obtaining the original question based on the original logical fallacy data and obtaining an optimized inference chain by performing closed-loop iteration of the inference chain based on the optimal inference path and the original question includes: The original question is obtained based on the original logical fallacy data. The optimal reasoning path and the original question are input into a language model with preliminary fallacy processing capabilities. The optimized reasoning chain is obtained through closed-loop iterative optimization in four aspects: multi-round evaluation and adjustment mechanism, scoring fusion and feedback strategy, dynamic correction and re-reasoning, and convergence judgment unit.
[0015] According to a preferred embodiment, the closed-loop iterative optimization process can be performed multiple times until the inference chain reaches a predefined quality benchmark.
[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. By breaking down the understanding and correction of logical fallacies into two stages—fallacy deconstruction and correction verification—the fallacy type is precisely named in the fallacy deconstruction stage, and a hierarchical analysis is conducted in the order of flawed premises, invalid reasoning patterns, and conclusion loopholes, allowing the model to clearly grasp the core components of the fallacy. Then, through the generation of counterfactual conditions, the construction of multi-level refutation paths, and the generation of effective argument chains, a robust argument chain is formed. The model is then fine-tuned for the first time on the LFUD-R dataset constructed based on this argument chain, enabling the model system to master the methods of fallacy identification, deconstruction, and correction, and no longer blindly reason, thus reducing the probability of logical fallacies occurring.
[0017] 2. Based on a model with preliminary fallacy handling capabilities, three different candidate reasoning chains are generated, covering more logical perspectives. Expert heuristic scoring is used to comprehensively evaluate the reasoning chains from four dimensions: logical completeness, clarity of thought process, redundancy control, and anti-interference ability. At the same time, the coherence and redundancy are quantitatively verified by combining BLEU-4 and ROUGE-L automated indicators. Finally, the optimal reasoning path is selected by calculating the final score according to scientific weights. This ensures the logical rigor of the reasoning chain while taking into account the fluency and conciseness of the expression, avoiding logical jumps or redundant information interference in the reasoning process, and making the reasoning path more coherent and persuasive.
[0018] 3. The optimal reasoning path and the original problem are re-inputted into the model. Closed-loop iteration is carried out through multiple rounds of evaluation and adjustment, score fusion and feedback, dynamic correction and re-reasoning, and convergence judgment, until the reasoning chain reaches the predefined quality benchmark, continuously correcting minor deviations in the reasoning process. On this basis, a second fine-tuning is performed to solidify the optimized reasoning experience into its own capabilities. It can not only maintain low error and high coherence reasoning performance in similar logic problems, but also generalize to logic reasoning tasks in different scenarios, improving the stability and robustness of the model's reasoning. This solves the problems of unstable reasoning quality and difficulty in adapting to complex scenarios in traditional LLMs. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the steps of a language model reasoning optimization method driven by logical fallacy identification and repair according to the present invention. Figure 2 This is a flowchart illustrating the core idea of constructing the LFUD-R dataset in this invention; Figure 3 This is a flowchart illustrating the core idea of the dynamic selection and optimization framework in this invention; Figure 4 This is a schematic diagram illustrating the principle of LFUD-R dataset construction and DSO-CoT framework in this invention. Detailed Implementation
[0020] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate the technical solutions of the present invention, but should not be used to limit the scope of protection of the present invention.
[0021] Example:
[0022] As attached Figure 1 As shown: This invention provides a language model inference optimization method driven by logical fallacy identification and repair, comprising the following steps: S1: Based on the original logical fallacy data, perform fallacy misconstruction and repair process operations, and perform hierarchical analysis of the three elements of fallacy on the original logical fallacy data to obtain the three-element analysis results. In this embodiment, the three elements are represented as flawed premise, invalid reasoning pattern, and conclusion loophole. The process of understanding and repairing logical fallacies is broken down into a fallacy deconstruction stage and a repair verification stage. In the fallacy deconstruction stage, the type of each fallacy instance is named. For example, "Because Xiaoming is a student, he must be very poor" is named the fallacy of composition, and "To conclude that a conclusion is correct based solely on an authoritative statement" is named the fallacy of appeal to authority. Then, a hierarchical decomposition is carried out in the order of flawed premise, invalid reasoning pattern, and conclusion loophole. For the fallacy "All electronic products need electricity, electronic products are part of my family, therefore everything in my family needs electricity," the flawed premise is equating "electronic products" with "all family members," the invalid reasoning pattern is an invalid extrapolation from "subset" to "whole," and the conclusion loophole is ignoring the existence of non-electronic products. The analysis results of the three elements of fallacy are obtained through such decomposition. Its purpose is to allow the model to deeply understand the internal mechanism of fallacy, rather than just identifying the surface type, and to provide a clear target for subsequent repair operations. The analysis results of the three elements of fallacy are obtained by carrying out a hierarchical decomposition in the order of flawed premise, invalid reasoning pattern, and conclusion loophole.
[0023] S2: Based on the analysis results of the three elements, perform argumentation and reasoning operations to obtain a robust argumentation chain. Based on the obtained robust argumentation chain, perform the first fine-tuning operation of the large language model to obtain a language model with preliminary error handling capabilities. In this embodiment, a counterfactual condition generation operation is performed based on the three-element analysis results to obtain a preliminary reasoning chain. Based on the preliminary reasoning chain, a multi-level refutation path construction operation is performed to obtain multiple refutation paths. Specifically, S21: Based on the three-element analysis results, extract the factual statements, reasoning rules and conclusions in the original reasoning chain, design chain prompts to drive the large language model to output contradictory hypothetical scenarios and candidate answers, construct counterfactual hypotheses by reversing, replacing or referencing similar cases in the knowledge base, and obtain the preliminary reasoning chain; For example, to address the fallacy that "all household items require electricity," we can construct the hypothesis that "if there are non-electronic products in the household," thus obtaining a preliminary chain of reasoning. This operation breaks the limitations of the original flawed logic and provides a starting point for refuting the fallacy.
[0024] S22: Perform multi-level deduction for each counterfactual hypothesis to construct a refutation path. For each hypothesis, use retrieval enhancement technology to obtain relevant facts or cases, simulate and deduce subsequent conclusions along the causal / logic chain, generate multiple refutation paths in the form of "if...then..." and verify the self-consistency of the original conclusion. This process can comprehensively reveal the defects of the original reasoning from different perspectives.
[0025] S23: Constructing a refutation path means refuting the core flaws through real-world examples. For example, using a real-world scenario such as "family members also include humans, furniture, and other non-electronic products that do not require electricity" to refute the flawed premise of the above fallacy, thereby enhancing the persuasiveness of the refutation.
[0026] S24: Generate an effective chain of arguments based on multiple refutation paths, integrate and evaluate these paths according to a predefined reasoning paradigm, ensure that each path meets structural constraints and aligns with factual knowledge to the greatest extent, filter paths by minimizing contradictions and maximizing causal consistency, compare the relative merits of counterfactual answers and original answers, and revise the original argument. For example, revising "everything in the house needs electricity" to "if all items in the house are electronic products, then everything needs electricity" creates a robust chain of arguments, effectively correcting the original fallacy.
[0027] S25: Perform the first fine-tuning operation of a large language model based on robust argument chains. Construct the LFUD-R dataset based on robust argument chains of all fallacies, and use this dataset to train the large language model so that it can systematically master the identification, deconstruction and repair methods of 12 typical fallacies, and obtain a language model with preliminary fallacy processing capabilities, laying the foundation for subsequent optimization of thought chains.
[0028] S3: Based on a language model with preliminary error handling capabilities, candidate reasoning chains are obtained. Based on the candidate reasoning chains, expert heuristic scoring and automated indicator verification are performed to obtain the optimal reasoning path. In this embodiment, S31: Based on a language model with preliminary error handling capabilities, a multi-candidate thought chain generation operation is performed. Carefully designed prompts guide the model to reason from different perspectives such as deduction, induction, and abduction, generating three different candidate reasoning chains for each problem—for example, in a math word problem, one chain derives the result step-by-step from known conditions (deduction), one chain summarizes the rules through multiple small examples and then applies them (induction), and one chain assumes the result is true and reverses the conditions (abduction). This approach strikes a balance between diversity and computational efficiency. Subsequently, based on the three different candidate reasoning chains, an expert heuristic scoring operation is performed, evaluating the chain from four dimensions: logical completeness, clarity of thought path, redundancy control, and anti-interference. The logical completeness assessment... The evaluation assesses whether the logical steps are complete and without jumps (5 points indicates all key steps are complete and closely connected, 1 point indicates missing core steps); the clarity of the thought process is assessed by whether the logical flow is easy to understand and the structure is reasonable (5 points indicates clear hierarchy and intuitive process, 1 point indicates chaotic steps and vague expression); redundancy control is assessed by whether there are unnecessary repetitions or lengthy passages (5 points indicates concise and efficient with no superfluous content, 1 point indicates excessive redundancy and inefficient reasoning); and the resistance to interference is assessed by whether it can resist interference from irrelevant information (5 points indicates accurate extraction of key information and focus on core logic, 1 point indicates reliance on irrelevant information and reasoning deviating from the topic). Each dimension is weighted at 40%, 30%, 15%, and 15% respectively, scored from 1 to 5 points, and the scores are weighted and summed to obtain an expert-heuristic aggregated score.
[0029] S32: Automated indicator verification is performed based on three different candidate inference chains. BLEU-4 and ROUGE-L indicators are used for quantitative evaluation. BLEU-4 reflects the fluency of the inference chain by measuring the accuracy of n-grams, while ROUGE-L measures the content coverage by measuring the longest common subsequence. The two are combined to quantitatively verify the coherence and information redundancy of the inference chain and obtain the automated indicator aggregation score.
[0030] S33: Based on the expert heuristic aggregation score and the automated indicator aggregation score, perform the optimal reasoning chain screening operation, calculate the final score by weighting and summing according to the preset weights. The final score is calculated as follows: Final Score = 0.6 × Expert Heuristic Aggregation Score + 0.4 × Automated Indicator Aggregation Score. Select the reasoning chain with the highest score as the optimal reasoning path. The purpose of this operation is to take into account the subtle differences in subjective judgment and the quantitative standards of objective indicators, and to ensure that the selected reasoning chain is logically sound, clear and concise.
[0031] S4: Obtain the original question based on the original logical fallacy data, and obtain the optimized inference chain by performing closed-loop iteration of the inference chain based on the optimal inference path and the original question; In this embodiment, the corresponding original questions are extracted from the original logical fallacy data. The optimal reasoning path and the original questions are input into a language model with preliminary fallacy processing capabilities. Closed-loop iterative optimization is carried out through four aspects: multi-round evaluation and adjustment mechanism, score fusion and feedback strategy, dynamic correction and re-reasoning, and convergence determination unit. The multi-round evaluation and adjustment mechanism is used to check the quality of the reasoning chain after each iteration. The score fusion and feedback strategy feeds the new evaluation results back to the model to guide the adjustment direction. The dynamic correction and re-reasoning guides the model to supplement key steps for dimensions with insufficient scores. The convergence determination unit monitors whether the quality of the reasoning chain meets the standard. The closed-loop iterative optimization process can be performed multiple times until the reasoning chain reaches the predefined quality benchmark and the optimized reasoning chain is obtained. This process can continuously correct reasoning deviations and continuously improve the rationality and reliability of the reasoning chain.
[0032] S5: Based on the optimized reasoning chain, a large-scale language model is fine-tuned, which enables the model to deeply integrate the experience of error identification and repair with the method of constructing a high-quality reasoning chain. This allows it to be generalized to more logical reasoning tasks in different scenarios, ultimately obtaining a language model with optimized reasoning ability. Its role is to significantly improve the stability and coherence of the model in complex logical reasoning tasks and reduce the occurrence of logical fallacies.
[0033] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A language model reasoning optimization method driven by logical fallacy identification and repair, characterized in that: It includes the following steps: S1: Based on the original logical fallacy data, perform fallacy misconstruction and repair process operations, and perform hierarchical analysis of the three elements of fallacy on the original logical fallacy data to obtain the three-element analysis results. S2: Based on the analysis results of the three elements, perform argumentation and reasoning operations to obtain a robust argumentation chain. Based on the obtained robust argumentation chain, perform the first fine-tuning operation of the large language model to obtain a language model with preliminary error handling capabilities. S3: Based on a language model with preliminary error handling capabilities, candidate reasoning chains are obtained. Based on the candidate reasoning chains, expert heuristic scoring and automated indicator verification are performed to obtain the optimal reasoning path. S4: Obtain the original question based on the original logical fallacy data, and obtain the optimized inference chain by performing closed-loop iteration of the inference chain based on the optimal inference path and the original question; S5: Perform secondary fine-tuning of the large-scale language model based on the optimized inference chain to obtain a language model with optimized inference capabilities.
2. The language model reasoning optimization method driven by logical fallacy identification and repair according to claim 1, characterized in that, The process of deconstructing and repairing fallacies based on the original logical fallacy data involves performing a hierarchical analysis of the three elements of fallacy in the original logical fallacy data to obtain the analysis results, including: The three elements are represented as flawed premises, invalid reasoning patterns, and flawed conclusions. The process of understanding and correcting logical fallacies is broken down into a fallacy deconstruction stage and a correction and verification stage. In the fallacy deconstruction stage, the type of each fallacy instance is named, and a hierarchical decomposition is carried out in the order of defective premise, invalid reasoning mode, and conclusion loophole to obtain the analysis results of the three elements of fallacy.
3. The language model reasoning optimization method driven by logical fallacy identification and repair according to claim 1, characterized in that, The process of performing argumentation and reasoning based on the three-element analysis results to obtain a robust argument chain, and then performing initial fine-tuning of a large language model based on the robust argument chain to obtain a language model with preliminary error handling capabilities, includes: Based on the results of the three-element analysis, counterfactual conditions are generated to obtain a preliminary inference chain. Based on the initial inference chain, a multi-level refutation path construction operation is performed to obtain multiple refutation paths; A robust argument chain is obtained by generating an effective argument chain based on multiple rebuttal paths; The first fine-tuning of a large language model is performed based on a robust argument chain to obtain a language model with preliminary error handling capabilities.
4. The language model reasoning optimization method driven by logical fallacy identification and repair according to claim 3, characterized in that, The counterfactual condition generation operation based on the three-element analysis results obtains a preliminary reasoning chain. Based on this preliminary reasoning chain, a multi-level refutation path construction operation is performed to obtain multiple refutation paths, including: Based on the three-element analysis results, the factual statements, reasoning rules and conclusions in the original reasoning chain are extracted. A chain-like prompt is designed to drive the large language model to output contradictory hypothetical scenarios and candidate answers. Counterfactual hypotheses are constructed by reversing, replacing or referencing similar cases in the knowledge base of key facts to obtain the initial reasoning chain. For each counterfactual hypothesis, a multi-level deduction is performed to construct a refutation path. For each hypothesis, retrieval enhancement technology is used to obtain relevant facts or cases, and subsequent conclusions are simulated and deduced along the causal / logic chain. Multiple refutation paths are generated and the self-consistency of the original conclusion is verified, thus obtaining multiple refutation paths. The construction of the refutation path is represented by refuting the core flaw through real-world examples.
5. The language model reasoning optimization method driven by logical fallacy identification and repair according to claim 3, characterized in that, The process of generating a robust argument chain based on multiple rebuttal paths, followed by initial fine-tuning of a large language model based on this robust argument chain, yields a language model with preliminary error handling capabilities. This includes: The evaluation path is integrated according to a predefined reasoning paradigm to meet structural constraints and align factual knowledge. By optimizing the target selection path, the merits of counterfactual answers and original answers are compared and the original arguments are revised to obtain a robust argument chain. The LFUD-R dataset, built upon robust argument chains, is used to train large-scale language models, enabling them to master error identification, deconstruction, and repair methods, thus obtaining language models with preliminary error handling capabilities.
6. The language model reasoning optimization method driven by logical fallacy identification and repair according to claim 1, characterized in that, The process involves obtaining candidate reasoning chains based on a language model with preliminary error handling capabilities, performing expert heuristic scoring and automated metric verification based on these candidate reasoning chains, and obtaining the optimal reasoning path, including: Based on a language model with preliminary error handling capabilities, a multi-candidate thought chain generation operation is performed, generating three different candidate reasoning chains for each question. Based on the three different candidate reasoning chains, an expert heuristic scoring operation is performed, scoring from 1 to 5 points according to four dimensions: logical completeness, clarity of thought path, redundancy control, and anti-interference ability, with each dimension having a weight of 40%, 30%, 15%, and 15%, respectively, and then weighted summed to obtain the expert heuristic aggregate score. Automated indicator verification is performed based on three different candidate inference chains. The BLEU-4 and ROUGE-L indicators are used to quantitatively evaluate the coherence and information redundancy of the inference chains and obtain the automated indicator aggregation score. Based on expert-heuristic aggregated scores and automated index aggregated scores, the optimal inference chain is selected. The final score is calculated by weighting and summing the scores according to preset weights, and the inference chain with the highest score is selected as the optimal inference path.
7. The language model reasoning optimization method driven by logical fallacy identification and repair according to claim 6, characterized in that, The calculation of the final score by weighted summation according to preset weights includes: The final score is calculated as follows: Final Score = 0.6 × Expert Heuristic Aggregated Score + 0.4 × Automated Metric Aggregated Score.
8. The language model reasoning optimization method driven by logical fallacy identification and repair according to claim 1, characterized in that, The process of obtaining the original question based on the original logical fallacy data, and obtaining the optimized inference chain through closed-loop iteration between the optimal inference path and the original question, includes: The original question is obtained based on the original logical fallacy data. The optimal reasoning path and the original question are input into a language model with preliminary fallacy processing capabilities. The optimized reasoning chain is obtained through closed-loop iterative optimization in four aspects: multi-round evaluation and adjustment mechanism, scoring fusion and feedback strategy, dynamic correction and re-reasoning, and convergence judgment unit.
9. The language model reasoning optimization method driven by logical fallacy identification and repair according to claim 8, characterized in that, The closed-loop iterative optimization process can be performed multiple times until the inference chain reaches a predefined quality benchmark.