Counterargument generation model, model training and inference method, evaluation criteria based on large model

By fine-tuning the large language model with argumentative instructions and training it with human preference data, combined with thought chain instructions, the problems of small scale and inaccurate evaluation in the counterargument generation model were solved, achieving high-quality counterargument generation and promoting the development of debate.

CN117407589BActive Publication Date: 2026-01-02FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311418867.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2026-01-02
Estimated Expiration
2043-10-30

AI Technical Summary

Technical Problem

In existing technologies, counter-argument generation models are small in scale, lack reasoning ability, have inaccurate evaluation metrics, and are difficult to generate high-quality counter-arguments.

Method used

By introducing argumentative instructions to fine-tune the large language model with low-rank instructions, and using human preference data to train a scoring model, error detection and rebuttal generation are carried out by combining instruction guidance in the form of thought chains, and high-quality counterarguments are selected using a scoring model.

Benefits of technology

It improves the adaptability of large language models to the task of generating counterarguments, generating high-quality, logically sound, and targeted counterarguments suitable for debate and discussion processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117407589B_ABST
    Figure CN117407589B_ABST
Patent Text Reader

Abstract

The application aims to provide a model and evaluation method for a sentence-level counterargument generation task, an adaptive model for the task, and a new evaluation standard, which comprises: in the training stage, fine-tuning a large language model with argumentation class instructions, and training a scoring model with human preference data; in the reasoning stage, guiding the large language model to detect errors in the input through instructions in the form of thought chains; further generating counterarguments based on error types with the help of the large language model; finally, scoring the candidate outputs with the pre-trained scoring model, and obtaining the system output after screening. The method constructs a pipeline reasoning framework of thought chain prompt-counterargument generation-screening, which can effectively identify the main errors in the argument and generate counterarguments with logic and pertinence.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer, in particular to a sentence-level counterargument generation model, a model training and reasoning method, and an evaluation standard based on a large language model. BACKGROUND

[0002] Debate behavior is an important manifestation of human high-level intelligence. The counterargument generation task aims to automatically generate arguments that are opposite in stance. Researchers have tried to index external knowledge bases, multi-step reasoning, and other methods to solve this task, and have completed a series of work. Counterargument generation aims to automatically generate refutations to a given statement or claim. This task plays a crucial role in various applications, such as persuasive writing, debate assistance, and opinion mining. For example, it can be widely used in various social media software to help users better integrate their own opinions and summarize others' statements.

[0003] Current research methods involve models of generally small size, while it has been proven that large-scale large language models can exhibit reasoning capabilities that small models do not possess. This capability naturally fits the counterargument generation task. At the same time, most current research uses n-gram-related metrics as automated evaluation indicators, ignoring the fact that counterargument generation task results have diversity, so more accurate evaluation indicators are needed. SUMMARY

[0004] The method provided by the present application provides a model and evaluation method for the sentence-level counterargument generation task to adapt to the model and new evaluation standard of the task. In the training phase, the large language model is fine-tuned with low-rank instructions by introducing debate instructions to improve the adaptability of the large language model to the counterargument generation task. In addition, a scoring model is trained using human preference data to evaluate the quality of the generated counterarguments.

[0005] In one embodiment, the training phase of the counterargument generation model and evaluation method provided by the present application fine-tunes the large language model with low-rank instructions using debate instructions. By introducing debate instructions as additional training signals, the adaptability of the large language model to the counterargument generation task can be improved. This instruction fine-tuning method can guide the model to generate more accurate and targeted counterarguments. Specifically, during the fine-tuning process, by incorporating debate instructions as part of the model input and training with the original corpus, the model can better understand and process the semantics and logic related to counterarguments.

[0006] In one embodiment, the training phase of the counterargument generation model and evaluation method provided by the present specification uses human preference data to train a scoring model. Using human preference data, a scoring model can be trained to evaluate the quality of generated counterarguments. By comparing with reference counterarguments, the scoring model can calculate scores to determine the pros and cons of generated counterarguments. Such an evaluation method can provide objective evaluation criteria for the generation of counterarguments, improving the reliability of model generation quality. Specifically, in the training process of the scoring model, first, a manually annotated dataset containing reference counterarguments and generated counterarguments is prepared, and then the scoring model is trained according to the annotated data to enable it to score generated counterarguments.

[0007] The method provided by the present specification provides a sentence-level counterargument generation model and evaluation method in the reasoning phase. In the reasoning phase, the large language model is guided by the instruction in the form of thought chain to detect errors in the input to help the model focus on error types and problems. This instruction in the form of thought chain can provide targeted guidance to enable the large language model to better identify and correct major errors in the argument. Specifically, in the reasoning phase, the input argument is compared and analyzed with the instruction according to the predefined thought chain instruction to determine the possible error types, such as logical errors, factual errors, etc. Through this instruction-guided approach, the large language model can more accurately determine the errors in the argument to support the generation of reasonable counterarguments.

[0008] In one embodiment, the reasoning phase of the counterargument generation model and evaluation method provided by the present specification guides the large language model to detect errors in the input through instructions in the form of thought chain. The design of the instruction in the form of thought chain can guide the model to focus on error types and problems to provide guidance for generating counterarguments. Specifically, the thought chain instruction can include prompts and questions for different error types, such as requiring the model to detect logical errors, insufficient data support, insufficient evidence, etc. Through such thought chain instructions, the large language model can better understand and locate errors in the argument to provide guidance for generating targeted counterarguments.

[0009] In one embodiment, the reasoning stage of the counterargument generation model and evaluation method provided by the present specification generates counterarguments based on error types with the help of a large language model. According to the detected error types, a large language model is used to generate counterarguments with logic and pertinence. The large language model can generate high-quality counterarguments by learning language rules and patterns in large-scale corpus, further improving the generation ability and accuracy of the model. Specifically, in the process of generating counterarguments, the large language model can use a pre-trained large language model or a generation model to generate reasonable counterarguments related to errors according to error types. Such a generation process can take advantage of the semantic understanding and generation ability of the large language model to support the generation of high-quality counterarguments.

[0010] In one embodiment, the reasoning stage of the counterargument generation model and evaluation method provided by the present specification uses a pre-trained scoring model to score candidate outputs. The generated candidate counterarguments are input into the scoring model for scoring and evaluation. The scoring model evaluates the quality of the counterarguments according to certain standards and rules, and selects the best system output. This scoring mechanism can provide objective evaluation and screening for generated counterarguments, ensuring the quality and reliability of system output. Specifically, in the application process of the scoring model, by defining appropriate scoring standards and indicators, the generated counterarguments are scored according to the scoring model, and the scoring results are sorted and screened, so as to obtain the best quality counterargument output.

[0011] As can be seen from the technical solutions provided by the embodiments of the present specification, the present specification provides a model and evaluation method for sentence-level counterargument generation task, which can generate high-quality, logical, and targeted counterarguments by introducing argumentation instructions, human preference data, and thinking chain guidance strategies. This method has a wide application prospect in debates and discussions, and can provide valuable support and reference for debaters, promoting the development and depth of debates. The results of the experiment also prove that the model training and reasoning method provided by the present specification can effectively learn and complete the counterargument generation task. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present specification, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0013] Figure 1 is a description of a sentence-level counterargument generation task provided by the present specification;

[0014] Figure 2 is a model framework structure and reasoning process diagram provided by the present specification. DETAILED DESCRIPTION

[0015] The technical solutions in the embodiments of the present specification will be described clearly and completely in combination with the drawings in the embodiments of the present specification. Obviously, the described embodiments are only part of the embodiments of the present specification, not all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0016] Referring to Figure 1 and Figure 2 As shown in the present specification, a model for a sentence-level counterargument generation task is provided, which is applied to processing topic-argument pair input to realize counterargument output. The method comprises:

[0017] Obtaining a topic-argument;

[0018] Combining the topic, the argument and the thought chain error prompted in the form of a thought chain obtained according to the topic-argument in a large language model, thereby generating a plurality of candidate arguments related to the thought chain error;

[0019] Scoring and ranking a plurality of candidate arguments by a scoring model, thereby screening the optimal counterargument.

[0020] Specifically, taking the topic-argument as input, for a variety of common error types in argumentation, a large language model in the form of a thought chain is used to obtain a series of errors. Through a pre-designed instruction template, the topic, the argument and the extracted thought chain error are combined to guide the large language model to output a series of candidate arguments. Through the form of a thought chain, the large language model can reason and analyze on the input topic and argument, thereby generating candidate arguments related to the error;

[0021] Wherein, after generating the candidate arguments, a scoring model trained based on human preference corpus is used to score the candidate arguments. The scoring model can sort and evaluate the candidate arguments according to the rules and standards of human preference. By training the scoring model, we can score the candidate arguments according to the quality and effect of argumentation;

[0022] Finally, the optimal counterargument is screened from the scoring model. These optimal counterarguments perform well in quality, logic and effectiveness, and can better refute the original argument and support the opposite view. Through this screening and scoring process, we can provide high-quality counterarguments to promote meaningful and insightful debates and discussions.

[0023] The present specification provides a counterargument generation framework for detecting original argument errors and automatically generating counterarguments. The method can include the following steps.

[0024] Step S10: Input the original argument and use the thought chain form of prompt instructions to guide the large language model to detect error types. The large language model will identify and label the error types in the argument, such as logical errors, factual errors, or grammatical errors, by analyzing the argument. At the same time, the large language model will generate natural language feedback, which will be combined with the original argument. The prompt can be understood as using "instructions" to "prompt" a specific "large language model".

[0025] Step S11: Input the original argument and the natural language feedback obtained in S10 into the system. In this step, we use the thought chain form of instructions again, which is used as model input to guide the large language model to refute based on the error types detected in S10. The large language model provides corresponding refutation points and arguments according to the types and properties of errors, in order to correct the errors or defects in the original argument.

[0026] Step S12: Select the optimal argument from a series of candidate arguments as the output of the system. In order to achieve this goal, a pre-trained scoring model is used to evaluate and rank the candidate arguments. This scoring model may consider a series of factors, such as logical coherence, evidence support, grammatical correctness, etc. This scoring model can find the most appropriate and most convincing argument as the final output of the system.

[0027] First, the external retrieval as a knowledge base has a serious problem, that is, it is highly dependent on the accuracy and precision of the knowledge base itself and the retrieval method used. Some current fuzzy matching retrieval methods have the risk of retrieving suboptimal information. This may be due to incomplete, incorrect or ambiguous information in the knowledge base, or due to limitations and deficiencies of the retrieval algorithm. Second, recent research has confirmed the ability of large large language models, which can memorize and generate a large amount of factual knowledge. These models can understand and remember facts about various fields of the world by training on a large corpus. They not only can accurately present these facts, but also can understand semantic relationships and contextual information, thus providing more comprehensive and accurate knowledge. Therefore, the framework uses large language models for error detection and knowledge output, which can help generate more fluent and logical counterarguments.

[0028] The embodiments of the present specification propose better utilization of large language models as knowledge bases for error detection. The training of the model goes through two stages, namely base large language model training and scoring model training. In the base large language model training stage, the model learns to better respond to logical instructions through instruction fine-tuning; in the scoring model training stage, another large language model is fine-tuned through human preference ranking data.

[0029] In the present embodiment, referring to Figure 2 , the input text is completed through multiple thought chain instructions to generate counterarguments; and further filtered through a scoring model to obtain the final result. As Figure 2 , three candidate arguments are filtered through the scoring model to obtain the optimal argument "not all birds can fly, for example, penguins cannot fly".

[0030] In the present embodiment, for the input (one argument), first, the large language model is guided by the prompts (instructions) of the three error forms (fact error, logical error, and cognitive bias) in the form of thought chains. Specifically, a multi-step reasoning process is constructed as an example input into the large language model, requiring the large language model to perform error detection by imitating the example. When implemented by a computer program, first, the guidance instructions and the example input are combined into a new text and input into the large language model. The large language model will attempt to insert the corresponding type of error form in the new input according to the requirements of the guidance instructions. Finally, the results of error detection are obtained from the model output.

[0031] In the present embodiment, after obtaining the error, the large language model is further guided to generate a targeted counterargument using the error, and each instruction generates an error and then generates a counterargument. When implemented by a computer program, first, the guidance instructions for generating a counterargument are constructed according to the error type output by the model. For example, if the model detects a fact error, the instruction "please provide an argument that corrects the aforementioned fact error in the following argument" can be constructed. The new counterargument guidance instruction and the original argument input are combined into a new text and interacted with the large language model again. The model will generate a counterargument as output according to the new guidance instruction.

[0032] In the present embodiment, each counterargument is spliced with the original input and input into a scorer (scoring model) with a BERT structure as the main framework, and the highest scoring counterargument is selected as the system output based on the output of the scorer (scoring model).

[0033] The present application also provides a training method of a large language model, comprising the following steps:

[0034] Obtaining a plurality of seed instructions;

[0035] By means of self-guiding, the seed instructions are used as starting points to enable the large language model to autonomously generate argumentation instructions for different debate scenarios and techniques.

[0036] The base model is fine-tuned for instructions by using a low-rank fine-tuning method.

[0037] Among them, the seed instructions can be a series of instructions with strong logicality and high relevance to debate techniques. These seed instructions are instructive and can guide the large language model to autonomously generate a complete set of argumentation instructions without supervision.

[0038] By means of self-guiding, the seed instructions are used as starting points to enable the large language model to autonomously generate other related argumentation instructions. This self-guiding method can enable the model to gradually expand and generate new instructions from existing knowledge to cover a wider range of debate scenarios and techniques.

[0039] The base model is fine-tuned for instructions by using a low-rank fine-tuning method. By fine-tuning the generated argumentation instruction set, the model can better understand and apply these instructions, improving its adaptability and execution effect for debate tasks.

[0040] The present application also provides a training method of a large language model, comprising the following steps:

[0041] The obtained multiple seed instructions are used as an instruction set: a series of argumentation technique-related instructions are artificially designed as an argumentative seed instruction set, including instructions to point out the errors of the original argument, provide arguments to support the original argument, etc.

[0042] Based on the self-guiding instruction generation template, four seed instructions are randomly extracted from the instruction set and filled into the template to form a complete self-guiding instruction, and new argumentative instructions are generated through the self-guiding instruction, and the final argumentative instruction set is obtained by iteration: a self-guiding instruction generation template is designed, and multiple instructions are randomly extracted from the seed instruction set and filled into the template to form a complete self-guiding instruction, and new argumentative instructions are generated with the help of the self-guiding instruction, and the process is repeated to obtain the final argumentative instruction set.

[0043] Based on the obtained final argumentative instruction set, the base model is fine-tuned for instructions in a self-regressive manner by using a low-rank instruction fine-tuning method, thereby updating the large language model after fine-tuning: based on the argumentative instruction set, the base model is fine-tuned for instructions in a self-regressive manner by using a low-rank instruction fine-tuning method.

[0044] Among them, a training method of a scoring model for counter-arguments is applied to the scenario of scoring counter-arguments given an original argument and a topic, and the method comprises:

[0045] Based on the artificially generated annotation data in the segment, the human preference ranking data is constructed according to the pre-set rules;

[0046] Based on these ranking data, a List-Wise algorithm is used as the optimization target of the ranking algorithm to supervise the training of the large language model. This ranking algorithm can evaluate and rank the generation results of the large language model according to the human preference ranking results;

[0047] During the training process, the given ranking score is used as the label, and by comparing the model-generated arguments with the human-preference-ordered reference results, the model's parameters and learning algorithm are optimized to better generate ranking results that meet human preferences.

[0048] Among them, in the step of constructing human preference ranking data, it includes:

[0049] From the original sentence-paragraph pair forum exchange corpus, sentence pairs are extracted. When operating, the paragraph is divided into several sentences according to the period, question mark, etc. The annotator marks whether the divided sentence constitutes a refutation relationship with the original sentence;

[0050] For each original sentence, 4 sentences are constructed to form ranking data. They are respectively the selected sentence of the same language pair, the unselected sentence of the same language pair, the safe reply with no information content, and the sentence of the different language pair.

[0051] Among them, in the step of training the scoring model, it includes:

[0052] During the training of the scoring model, the four sentences are assigned an increasing ranking score in order, which is 1, 2, 3 and 4 respectively;

[0053] The four sentences are respectively concatenated with the original argument to form the input of the model. The model is trained using a supervised training method based on the ranking score.

[0054] The specification also provides a method for fine-tuning a large language model for argumentation instructions, which includes the following steps.

[0055] Step S20: Starting from 10 argumentation-related instructions, extract 5 instructions as examples each time, and require ChatGPT to generate new instructions. If the new instructions exceed 0.8 in word overlap (represented by the ROUGE-1 index) with the old instructions, the new instructions are not considered. Finally, an argumentation instruction set is obtained.

[0056] Step S21: Train the large language model using the low-rank fine-tuning strategy from the argument instruction set. The low-rank instruction fine-tuning strategy is a technique generally known to those skilled in the art and is not described here. The method described in this specification is a low-rank fine-tuning of the query, key, value, and linear transformation matrix in LLaMA (Transformer structure).

[0057] The present specification also provides a scoring model training method specially suitable for counterargument generation tasks, which comprises the following steps.

[0058] Step S30: First, construct human preference ranking data. Based on the dialogue artificial annotation record, the data can be divided into the following four categories: the first category is the sentence in the same dialogue that is confirmed to be a counterargument; the second category is the sentence in the same dialogue that is confirmed to be not a counterargument; the third category is a conservative reply (lacks information, such as just raising objections "I disagree"); the fourth category is the sentence in different dialogues. In this way, human preference ranking data composed of four sentences is constructed for each input.

[0059] Step S31: Then, fine-tune the BERT-base model using the List-Wise method to enable it to learn the argument quality ranking task and thus obtain scoring ability. Specifically, the loss function used in training is the cross-entropy loss function:

[0060]

[0061] where s i represents the ranking score of the i-th sentence (the aforementioned four categories are 1, 2, 3, and 4 points, respectively), x represents the original argument, y i represents the i-th sentence.

[0062] In the reasoning phase, the sentence that maximizes the following probability is selected as the output:

[0063]

[0064] During the experiment, when training the large language model using the instruction adjustment, LLaMA-7b was used as the base model, the learning rate was set to 3x10 -4 , the batch size was set to 256, the gradient accumulation step was set to 16, and both a and r of the low-rank fine-tuning method were set to 16. The model was trained for 5 epochs on 4 NVIDIA RTX3090 GPUs. When training the scoring model, BERT-base was used as the base model, the learning rate was set to 1x10 -5, batch size is set to 64, and trained for 2 epochs on an NVIDIA RTX3090 GPU. AdamW is used as the optimizer for training both models. Among them, BaseModel refers to the source of the model we obtain. Here, we use the "instruction template obtained instruction (instruction data)" to train and fine-tune the parameters of a "base model", and then obtain a specific large language model for subsequent instruction templates (or instructions). The scoring model is another model independent of the previous ones. This model is a general large language model, not a large language model. We use "human preference ranking data" to train this scoring model.

[0065] During the experiment, various comparative models were tried as baselines, and the effectiveness of the framework was verified on automatic indicators (BLEU, ROUGE, METEOR), large language model indicators (ChatGPT Eval, Arg-Judge).

[0066] Table 1: Experimental results on the ArgTersely dataset

[0067]

[0068] It can be clearly observed that compared with the baseline model, the framework performs well on both automatic indicators and large language model indicators. Each component in this framework can more accurately identify errors and provide effective refutations. Compared with non-instruction fine-tuning models, instruction fine-tuning has achieved better performance in improving performance. In addition, large language model indicators are more consistent and reliable than automatic indicators.

[0069] In terms of automatic indicators, the framework has made significant progress in the tasks of automatic detection and evaluation, as shown by the results of various evaluation indicators. This means that the framework can more accurately detect and mark errors, providing more reliable and accurate automatic indicator results. At the same time, large language model indicators provide a more comprehensive and consistent evaluation method. Compared with traditional automatic indicators, these large language model indicators can better capture the subtle differences in language expression and the accuracy of logical reasoning. Through the evaluation of large language model indicators, we can better evaluate the performance of the framework in terms of semantic understanding and expression.

[0070] Table 2: Results of component ablation experiments

[0071]

[0072] From the results of the above component ablation experiments, it can be found that using argumentation instructions during training is very helpful to the model. This proves that the method of generating argumentation instructions is effective.

[0073] Instruction tuning is important. Simply generating based on LLaMA large language model will affect performance, while instruction tuning can help the model better adapt to argumentation scenarios, respond to instructions, and reason out correct refutations from chain-of-thought instructions.

[0074] Chain-of-thought instructions can produce higher-quality refutations compared to ordinary single-step instructions. This is because chain-of-thought instructions can give logical chains and multi-step reasoning processes, thereby improving output quality.

[0075] Multi-error templates can improve the quality of generated refutations. Multi-error templates can help large language models find potential errors from multiple angles, thereby generating more diverse candidate sentences.

[0076] The scoring model component plays a crucial role in the system. It can select high-quality arguments from candidate sentences, while random selection cannot achieve similar results.

[0077] The above-described apparatuses or modules, etc. can be implemented by computer chips or entities, or by products with certain functions. For the convenience of description, the above apparatuses are described as various modules respectively. Of course, in the implementation of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be combined or integrated into another system, or some features can be ignored or not executed.

[0078] As a person skilled in the art knows, in addition to implementing the controller in the form of pure computer readable program code, the controller can also be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps to achieve the same functions. Therefore, such a controller can be considered as a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0079] The application can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform particular tasks or implement particular abstract data types. The application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.

[0080] From the description of the above embodiments, those skilled in the art can clearly understand that the application can be implemented by means of software and necessary universal hardware platforms. Based on such an understanding, the technical solutions of the application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments of the application.

[0081] The various embodiments in the specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. The application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, etc.

[0082] Although the application is described through embodiments, those skilled in the art know that the application has many modifications and changes without departing from the spirit of the application, and it is hoped that the appended claims include these modifications and changes without departing from the application.

Claims

1. A method for generating sentence-level counterarguments, the method comprising: receiving a text; identifying a claim in the text; identifying a counterargument to the claim; and generating a sentence-level counterargument to the claim. The method is applied to processing topic-original argument pair input to realize counter-argument output, and the method comprises: Obtaining a topic-original argument; The topic and the original argument obtained according to the topic-original argument and the thought chain error prompted in the form of a thought chain are combined in a large language model to generate a plurality of candidate arguments related to the thought chain error, including: filling the obtained topic and original argument into each instruction template corresponding to a plurality of error types to form a complete argument instruction, obtaining the word sequence of the instruction through a word segmentation device, and inputting the word sequence into the large language model to generate a plurality of candidate arguments; The plurality of candidate arguments are scored and sorted by a scoring model to filter out the optimal counter-argument, including: separating a single candidate argument from the original argument by a special character "[SEP]" to form a text, and obtaining the corresponding word sequence by segmenting the text; The word sequence obtained by segmentation is input into the model to obtain word embedding, and the word embedding is subjected to average pooling, regularization and linear transformation to obtain the final score. The scoring method for generating counter-arguments includes the following steps: Based on the manually generated annotation data in the text segment, the scoring model constructs the sorting data in accordance with the pre-set rules to conform to the human preference, and the step of "constructing the sorting data in accordance with the human preference" includes: extracting sentence pairs from the original sentence-paragraph pair forum exchange corpus, wherein the paragraph is divided into a plurality of sentences according to the whole sentence delimiter, and the annotators label whether the divided sentences constitute a refutation relationship with the original sentence; for each original sentence, a plurality of sentences are constructed to form sorting data, wherein the order is a certain sentence of the same language pair that is selected, a certain sentence of the same language pair that is not selected, a safe reply with no information content, and a certain sentence of a different language pair; in the process of training the scoring model, the plurality of sentences are assigned an increasing sorting score in order; the plurality of sentences are spliced with the original argument to form the input of the model; the scoring model is trained using a supervised training method based on the sorting score; Based on the sorting data, the scoring model uses the List-Wise algorithm as the optimization target of the sorting algorithm to supervise the training of the large language model; wherein, in the training process of the scoring model, the given sorting score is used as a label, the generated argument of the model is compared with the reference result of the human preference sorting, the parameters and learning algorithm of the model are optimized to make it better generate the sorting result conforming to the human preference.

2. The counterargument generation method according to claim 1, characterized in that, The plurality of error types include at least one of a factual error, a logical error and a cognitive bias.

3. The counterargument generation method according to claim 1, characterized by, The step of "filling the obtained topic and original argument into each instruction template corresponding to a plurality of error types to form a complete instruction, obtaining the word sequence of the instruction through a word segmentation device, and inputting the word sequence into the large language model to generate a plurality of candidate arguments" includes the following steps: setting the error type of the hypothetical argument, finding the argument defect by using the hypothetical error type, and proposing a counter-argument based on the argument defect.

4. A training method of a large language model using the method for generating a sentence-level counterargument according to claim 1, characterized in that, The method comprises the following steps: Obtaining a plurality of seed instructions; By means of self-guiding, the seed instructions are used as a starting point to enable the large language model to autonomously generate argumentation instructions for various debate scenarios and skills; The low-rank fine-tuning method is adopted to fine-tune the argumentation instructions in the base model.

5. A training method of a large language model using the method for generating a sentence-level counterargument according to claim 1, characterized in that, The method comprises the following steps: The obtained multiple seed instructions are taken as an instruction set; Based on the self-guiding instruction generation template, four seed instructions are randomly extracted from the instruction set and filled into the template to form a complete self-guiding instruction, a new argumentation instruction is generated through the self-guiding instruction, and the final argumentation instruction set is obtained through iteration; Based on the obtained final argumentation instruction set, the base model is fine-tuned in a self-recurrent manner by using the low-rank instruction fine-tuning method.