Question and answer generation method and system based on big language model preference alignment

By generating and filtering the interpretations of the preferred interpretations of large language models, the pre-trained language model is trained to obtain a biased interpretation generator, which solves the interpretation deviation of large language models and the problems of computing resources and privacy and security, and realizes an efficient and secure question-and-answer generation method.

CN119938863AActive Publication Date: 2025-05-06XI AN JIAOTONG UNIV

Patent Information

Application Number
CN202510105661.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

When existing large language models deal with problems with similar semantics but different expressions, they may output incorrect answers, resulting in interpretation bias, and traditional white box tuning methods are difficult to meet computing resources and privacy security needs.

Method used

By obtaining multiple original questions used to train the target large language model, generating their corresponding multiple interpretations, and inputting these interpretations into the target large language model, judging the preferences of the interpretation based on the correctness of the reply answer, it is divided into a set of partial interpretations and a set of non-binding interpretations. These interpretations are used to train the pre-trained language model to obtain a generator with partial interpretations, which is used to generate the definitions preferred by the target large language model.

Benefits of technology

The interpretation preferred by generating the target large language model is achieved, the interpretation deviation problem is solved, and the high computing resource requirements and privacy leakage security risks of traditional white box tuning methods are avoided, and the computing cost and privacy security needs of large language models are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938863A_ABST
    Figure CN119938863A_ABST
Patent Text Reader

Abstract

The invention discloses a question and answer generation method and system based on big language model preference alignment, and relates to the technical field of artificial intelligence natural language processing, and the method comprises the following steps: splicing each original question with a specified prompt, inputting a splicing instruction into a paraphrase generator, and generating a plurality of paraphrases corresponding to the original question; inputting each original question and the corresponding paraphrases into a target large language model, and dividing the paraphrases into a biased paraphrasing set and a non-biased paraphrasing set; and training the pre-training language model by taking the unbiased paraphrasing set as input and the biased paraphrasing set as output to obtain a generator with biased paraphrasing. The whole process required for problem preference alignment is a black box optimization process, any parameter of the target large language model does not need to be obtained and changed, and the problems that a traditional white box method is high in computing resource requirement and has privacy disclosure and potential safety hazards can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence natural language processing, and in particular to a question and answer generation method and system based on large language model preference alignment. Background Art

[0002] Large Language Model (LLM) refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. Large language models can handle a variety of natural language tasks, such as text classification, question answering, and dialogue, and are an important path to artificial intelligence. With the continuous growth of parameter scale and the continuous development of pre-training methods, the overall performance of LLM has become more and more powerful. However, the rapid development of LLM has also brought some problems. First, the surge in parameter scale has led to a significant increase in the computing resources required to adjust them. At present, some studies have proposed methods to optimize the performance of LLM with lower computing requirements. They are mainly white-box adjustment methods and require access to model parameters. However, the security and privacy issues around LLM are increasingly valued. Therefore, some LLM designers choose to provide services to users through API interfaces instead of exposing the complete model. This trend makes it increasingly challenging to apply white-box adjustments.

[0003] Paraphrases are texts that convey the same meaning using different words or structures, which reflects the complexity and diversity of human language. Given a sentence, the goal of paraphrase generation is to produce a paraphrase that is different from the original sentence but maintains the original meaning. Many natural language processing use paraphrase generation to complete downstream tasks, such as question-answering systems, semantic parsing, dialogue systems, and machine translation. In the era of large language models, end-to-end paraphrasing using language models ensures accuracy and diversity while maintaining flexibility and ease of use. Therefore, this method has become the mainstream choice for paraphrase generation.

[0004] There is a problem of interpretation bias in existing large language models. When an original question is input into the large language model, it is able to answer the original question correctly, but when the original question is asked in a slightly different but semantically similar way, the large language model may output the wrong answer. This can be explained by the fact that the large language model not only learns the knowledge itself from the corpus during pre-training, but also learns the expression patterns related to specific knowledge. Existing studies regard this problem as a problem of the robustness of large language models to question interpretation, and propose a full-parameter fine-tuning method to enhance the robustness of large language models. However, this method is white-box and requires adjusting all parameters of the large language model, which makes it difficult to meet the computational cost and privacy security requirements of LLM. Summary of the invention

[0005] The present invention provides a question-answer generation method and system based on preference alignment of a large language model, which solves the problem that the existing methods need to adjust all parameters of the model when solving the problem of interpretation bias in a large language model, which makes it difficult to meet the computational cost and privacy security requirements of LLM.

[0006] In a first aspect, the present invention provides a question-answer generation method based on large language model preference alignment, comprising the following steps:

[0007] Acquire multiple original questions for training the target large language model, and concatenate each original question with a specified prompt to obtain a concatenation instruction, and input the concatenation instruction into a paraphrase generator to generate multiple paraphrases corresponding to the original question;

[0008] Input each original question and the corresponding multiple interpretations into the target large language model, obtain multiple reply answers of the target large language model, judge whether the multiple interpretations of each original question are preferred by the target large language model according to the correctness of the reply answers, and divide the multiple interpretations into a biased interpretation set and a non-biased interpretation set according to the preference results of the multiple interpretations;

[0009] The pre-trained language model is trained using the unbiased interpretation set as input and the biased interpretation set as output to obtain a biased interpretation generator.

[0010] The question to be aligned with preference is input into the biased interpretation generator to obtain the preference-aligned interpretation, and the preference-aligned interpretation is input into the target large language model to obtain the corresponding reply answer.

[0011] Preferably, the specified prompt includes multiple interpretation types.

[0012] Preferably, judging whether the multiple interpretations of each original question are preferred by the target large language model according to the correctness of the reply answer comprises the following steps:

[0013] The reply answer is compared with the target answer for correctness. If it is correct, it is considered that the interpretation corresponding to the reply answer is favored by the target large language model.

[0014] If it is wrong, it is considered that the interpretation corresponding to the reply answer is not favored by the target large language model.

[0015] Preferably, before inputting each original question and the corresponding multiple interpretations into the target large language model, low-quality interpretations among the multiple interpretations need to be filtered through an interpretation filtering mechanism.

[0016] Preferably, the parameter scale of the pre-trained language model is smaller than that of the target large language model.

[0017] Preferably, the obtaining of a plurality of original questions for training the target large language model comprises the following steps:

[0018] Obtaining an original data set for training a biased interpretation generator; the original data set is a data set specified by a user, and the task type and field of the data can be selected by the user according to training requirements;

[0019] Convert each raw data in the original dataset into the original question and the target answer in question-answer format.

[0020] In a second aspect, the present invention provides a question-answer generation system based on large language model preference alignment, comprising:

[0021] An acquisition module is used to acquire multiple original questions for training a target large language model, and concatenate each original question with a specified prompt to obtain a concatenation instruction, and input the concatenation instruction into a paraphrase generator to generate multiple paraphrases corresponding to the original question;

[0022] A classification module is used to input each original question and the corresponding multiple interpretations into the target large language model, obtain multiple reply answers of the target large language model, judge whether the multiple interpretations of each original question are preferred by the target large language model according to the correctness of the reply answers, and divide the multiple interpretations into a biased interpretation set and a non-biased interpretation set according to the preference results of the multiple interpretations;

[0023] A generation module, which is used to train the pre-trained language model with the unbiased interpretation set as input and the biased interpretation set as output to obtain a biased interpretation generator;

[0024] The interpretation module is used to input the questions to be aligned with preference into the biased interpretation generator to obtain the interpretation after preference alignment, and input the interpretation after preference alignment into the target large language model to obtain the corresponding reply answer.

[0025] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned question and answer generation method based on large language model preference alignment when executing the program.

[0026] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned question and answer generation method based on large language model preference alignment.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] The present invention constructs a biased interpretation generator, which can generate interpretations preferred by a target large language model. During the construction process, each original question used to train the target large language model is first spliced ​​with a specified prompt to obtain a splicing instruction, and the splicing instruction is input into the interpretation generator to generate multiple interpretations corresponding to the original question. The present invention proposes a prompt construction method based on interpretation type, which enhances the diversity of interpretations by adding a specified interpretation type to the prompt input to the interpretation generator. Then, multiple interpretations are input into the target large language model, and the preference of the target large language model divides the multiple interpretations into a biased interpretation set and an unbiased interpretation set. Finally, the pre-trained language model is trained by the unbiased interpretation set and the biased interpretation set to obtain a biased interpretation generator. Compared with the traditional white-box tuning method, the present invention regards the interpretation deviation in LLM as a problem of alignment with the preference of the target large language model. The entire process required for the present invention to achieve question preference alignment is a black box optimization process, that is, the entire process does not require obtaining or changing any parameters of the target large language model, which can solve the problems of high computing resource requirements, privacy leakage and security risks in traditional white box methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0030] Figure 1 A flowchart of a question-answer generation method based on large language model preference alignment according to the present invention;

[0031] Figure 2 It is a schematic diagram of the reasoning process of the present invention;

[0032] Figure 3 It is a flow chart of the automatic construction method of the training set of the present invention. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0034] The present invention provides a question-answer generation method based on preference alignment of a large language model, specifically a question-answer generation method based on preference alignment of a large language model based on interpretation, wherein the interpretation is a question in which the original question conveys the same meaning in different words or structures, and aims to use the interpretation to solve the preference alignment problem of the large language model in the form of black-box tuning, and enhance the model performance by interpreting the question as an expression preferred by the model. The present invention trains a biased interpretation generator to learn the preference of the large language model for the expression method. During the reasoning process, the biased interpretation generator generates interpretations that are aligned with the preferences of the large language model, and then inputs them into the large language model. The training of the biased interpretation generator adopts a sequence-to-sequence (seq2seq) method. Taking into account the diversity of interpretation types and the aspects of semantic preservation, the present invention proposes a prompt-based method for automatically constructing the final training set. Reference Figure 1-Figure 3 , specifically including the following steps:

[0035] Step 1: Get the original dataset for training the biased explanation generator, and organize the original question and target answer of each original data in the original dataset in the question-answer format.

[0036] Figure 1 The original query in Figure 3 The initial training set in is multiple original questions. Figure 1 Query 1-query n in are all original questions, Figure 3 Question 1, Question 2, and Question n (Question n) are all original questions. For example, the original question is: "Where will the XXVI Summer Olympic Games be held?", and the target answer is: "Atlanta, USA".

[0037] In this implementation, the original data set is a data set specified by the user, and the task type and field of the data can be selected by the user according to training requirements.

[0038] The present invention selects six public datasets from three tasks of open domain question answering, common sense reasoning and mathematical application problems: WebQSP, CompQ, CWQ, SiQA, CSQA and SVAMP, and constructs the original dataset through these six datasets.

[0039] Step 2: Concatenate each original question in step 1 with the specified prompt to obtain a concatenation instruction, input the concatenation instruction into the interpretation generator, and generate multiple interpretations corresponding to the original question, that is, Figure 1 The query candidate set in Figure 1 Q1-a, Q1-b and Q1-c are multiple interpretations of query 1, and Qna, Qnb and Qnc are multiple interpretations of query n.

[0040] The specified prompt contains the specified interpretation type. The purpose of this step is to include enough interpretation types in the training data of the biased interpretation generator to increase the chance of generating effective interpretations. For example, please use {modality change} to interpret this sentence: {Where is the venue for the 26th Summer Olympic Games}". Splicing adds the specified interpretation type "modality change" and the sentence to be interpreted to the template. For example Figure 3 Paraphrase thissentence using{Type lnstruction}:Question1, that is, use {type structure} to explain this sentence: Question 1.

[0041] For the interpretation types used in prompt construction, the present invention selects the classification standard of interpretation type detection PTD, which includes 7 first-level interpretation types and 27 second-level interpretation types. The second-level classification is to further subdivide each category in the first-level classification.

[0042] The interpretation generator used in the present invention can be a dedicated interpretation generator, such as the open source text editing model CoEdit, etc., or it can be a large language model with interpretation capabilities.

[0043] In this embodiment, the present invention includes a meaning filtering mechanism, which can filter out the low-quality meanings generated in step 2, and only the filtered meanings will be passed to the target large language model to judge their biased characteristics. The purpose of this step is to ensure the semantic consistency before and after the question interpretation, so as to prevent low-quality interpretations from entering the training set and affecting the training of the biased interpretation generator. Filtering is specifically a process of scoring the interpretation pairs through the machine evaluation indicators generated by the text and filtering according to the threshold, which specifically includes:

[0044] (1) Combine the original question with the corresponding multiple interpretations to obtain multiple interpretation pairs, and use the machine evaluation indicators generated by the text to score the multiple interpretation pairs. The optional evaluation indicators include BLUERT, BARTscore, and UniTE. Figure 3 Scores: 0.23, 0.95, 0.68, 0.63, and 0.84.

[0045] (2) Based on the score of each interpretation pair and a pre-set threshold, the interpretations that do not meet the interpretation quality requirements are deleted to obtain multiple filtered interpretations. Figure 3 The meaning of and

[0046] Step 3: Pass the original question obtained in step 2 and the filtered multiple interpretations as input to the target large language model (knowledge base) to obtain the response answer of the target large language model. According to the correctness of the response answer, it is judged whether the multiple interpretations of each original question are preferred by the target large language model.

[0047] Figure 1 Answer 1-answer n are the target answers to queries 1-query n, answers A 1-a, A 1-b and A1-c are the response answers to Q1-a, Q 1-b and Q1-c, and answers An-a, Anb and Anc are the response answers to Q na, Q nb and Q nc.

[0048] The reply answer is judged for correctness against the target answer of the original question. If it is correct, it is considered that the interpretation corresponding to the reply answer is favored by the target large language model; if it is wrong, it is considered that the interpretation corresponding to the reply answer is not favored by the target large language model.

[0049] Step 4: Divide the homologous interpretation sets (all interpretations of the same original question) into two groups according to the bias obtained in step 3: biased interpretation sets and unbiased interpretation sets, which serve as the final training set for the biased interpretation generator. Cartesian product is performed on the biased interpretation set and the unbiased interpretation to obtain the interpretation sentence pair set, such as Figure 3 In and Figure 1 In which query 1-a and query 1-b are the biased and unbiased interpretations in sentence pair 1, label 1 is the corresponding label of sentence pair 1, Qnb and Qnc are the biased and unbiased interpretations in sentence pair n, and label n is the corresponding label of sentence pair n.

[0050] In this embodiment, a data augmentation strategy is adopted when constructing the final training set of the biased paraphrase generator, that is, a certain number of identical sentence pairs are used to enhance the training set, and the parameter λ is used to control the ratio of the biased paraphrase set to the non-biased paraphrase set. The purpose of this step is to allow the biased paraphrase generator to understand that it can skip questions that have been aligned with the model's preferences.

[0051] Step 5: Use the final training set to fine-tune the specified pre-trained language model in a sequence-to-sequence manner with biased interpretation, and obtain a biased interpretation generator that learns the expression preference of the target large language model. The biased interpretation generator consists of a bidirectional encoder and an autoregressive encoder.

[0052] In this embodiment, the parameter scale of the pre-trained language model is much smaller than that of the target large language model, and it only requires a small amount of training data for a given task. Optional pre-trained language models include Flan-T5-Large, etc.

[0053] The Flan-T5-Large model is trained using a sequence-to-sequence training method, and the trained Flan-T5-Large model is used as a biased interpretation generator.

[0054] The unbiased query in the sentence pair is fed into the encoder of the model as an input sequence. The encoder processes the input sequence and converts it into a set of hidden states that contain the semantic information of the input sequence. The decoder receives the hidden states of the encoder and generates an output sequence based on this information. During the training process, the output sequence generated by the decoder is compared with the target sequence (i.e., the biased query in the sentence pair). During training, the decoder uses a teacher forcing strategy, that is, the real target sequence is used as the input of the decoder instead of the output generated by the decoder itself. This can accelerate the convergence of training. During training, cross entropy is used as the loss function. This loss function calculates the difference between the output sequence generated by the decoder and the target sequence. The smaller the loss value, the closer the output sequence of the model is to the target sequence. Through the back-propagation algorithm, the parameters (weights and biases) in the model are adjusted according to the gradient information of the loss value to minimize the loss function.

[0055] Figure 2 For the reasoning process, the original query is What structure is made from DNA and proteinmolecules coiled together? The original query is input into the biased interpretation generator, and the potential optimal query is obtained: What is the composition of the structure formed by intertwined DNA and protein molecules? The potential optimal query is input into the target large language model, and the query result is: chromatin.

[0056] Step 6: Given the question to be aligned, use the biased interpretation generator trained in step 5 to interpret the question to be aligned, obtain the biased aligned question (biased interpretation), and then pass the biased aligned question to the target large language model to obtain the reply answer.

[0057] For example, the MPT-instruct-7B model gives different answers to different questions with the same semantics under different interpretations. They include:

[0058] Example 1

[0059] Q: When wildlife reproduce we often refer to what comes out as what?

[0060] P:What is the term used to describe the outcome of wildlife reproduction? (Answer: What is the term used to describe the outcome of wildlife reproduction?)

[0061] Example 2:

[0062] Q:While washing clothes they became what when caught on the sharpobject?

[0063] (Question: What happens to clothes when they get caught in sharp objects while washing?)

[0064] P:What was the outcome of the clothes getting caught on the sharp object while washing?

[0065] Example 3:

[0066] Q: What is the least dangerous radioactive decay?

[0067] P: Among radioactive decay types, which one poses the least risk?

[0068] In this embodiment, the entire process required to achieve question preference alignment is a black box optimization process, that is, the entire process does not require obtaining or changing any parameters of the target large language model.

[0069] In essence, the method of the present invention is a process of learning a model's preference for expression and incorporating this preference into questions to guide the model to produce correct answers. This is similar to prompt learning, in which prompts are adjusted to adapt to the expression that the model is accustomed to during pre-training. However, prompt learning involves incorporating preference information into prompts that are unrelated to the question. It should be noted that prompts are usually task- or domain-specific, which means that the preference information obtained through prompt learning is saved in a fixed format in the prompts of a specific task or domain. However, questions are not specific to any domain or task, they are constantly changing. This requires dynamically incorporating the model's preference information into the question. The biased interpretation generator can adaptively incorporate preferences in the interpretation process by learning the model's preference for expression under different semantics, and its adaptability can be specifically reflected in its ability to apply different types of interpretations to different questions. The interpretation type represents the lexical variable operated in the interpretation generation process.

[0070] Table 1 shows the performance of the present invention and existing tuning methods on six data sets of three tasks, where the existing tuning methods include LORA (Low-Rank Adaptation), Manual Prompt (manual prompt tuning), UPRISE (Universal Prompt Retrieval for Improving Zero-Shot Evaluation) and EPR (Efficient Prompt Retrieval). Table 2 shows the performance of the present invention and different large language models on six data sets of three tasks, where the large language models used include TinyLlama-1.3B, Bloom-7B, RedPajama-instruct-3B, MPT-instruct-7B and MPT-instruct-30B.

[0071] Table 1 Performance of the proposed method and existing optimization methods on six datasets of three tasks

[0072]

[0073] Table 2 Performance of different large language models on six datasets of three tasks

[0074]

[0075] Experimental results show that this technical solution can significantly improve the performance of large models in open-domain question answering, common sense reasoning and mathematical application problems. By adding model preference information to the original problem through the conversion of the expression method, accurate and reliable question answering services can be provided to large model users.

[0076] This invention breaks away from the traditional white-box tuning method, regards the interpretation deviation in LLM as a problem of model preference alignment, and proposes a question-answer generation method based on interpretation and large language model preference alignment. It can solve the problems of high computing resource requirements, privacy leakage and security risks of traditional white-box methods, thereby adapting to the needs of rapid development and iterative updates of enterprises, giving full play to the effectiveness of large language models, accurately and reliably obtaining the required knowledge and information, thereby improving production efficiency, and having good economic and social benefits.

[0077] Based on the same concept, the present invention also provides a question-answer generation system based on large language model preference alignment, including an acquisition module, a splicing module, a classification module, a generation module and an interpretation module.

[0078] The acquisition module is used to obtain multiple original questions for training the target large language model, and splice each original question with a specified prompt to obtain a splicing instruction, and input the splicing instruction into the interpretation generator to generate multiple interpretations corresponding to the original question.

[0079] The classification module is used to input each original question and the corresponding multiple interpretations into the target large language model, obtain multiple reply answers of the target large language model, judge whether the multiple interpretations of each original question are preferred by the target large language model according to the correctness of the reply answers, and divide the multiple interpretations into biased interpretation sets and non-biased interpretation sets according to the preference results of the multiple interpretations.

[0080] The generation module is used to train the pre-trained language model by taking the unbiased interpretation set as input and the biased interpretation set as output to obtain a biased interpretation generator.

[0081] The interpretation module is used to input the questions to be aligned with preference into the biased interpretation generator to obtain the interpretation after preference alignment, and input the interpretation after preference alignment into the target large language model to obtain the corresponding reply answer.

[0082] The present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the above-mentioned question and answer generation method based on large language model preference alignment is implemented.

[0083] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned question and answer generation method based on large language model preference alignment is implemented.

[0084] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0085] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A question-answer generation method based on preference alignment of a large language model, characterized in that: The following steps are involved: Acquire multiple original questions for training the target large language model, and concatenate each original question with a specified prompt to obtain a concatenation instruction, and input the concatenation instruction into a paraphrase generator to generate multiple paraphrases corresponding to the original question; Input each original question and the corresponding multiple interpretations into the target large language model, obtain multiple reply answers of the target large language model, judge whether the multiple interpretations of each original question are preferred by the target large language model according to the correctness of the reply answers, and divide the multiple interpretations into a biased interpretation set and a non-biased interpretation set according to the preference results of the multiple interpretations; The pre-trained language model is trained using the unbiased interpretation set as input and the biased interpretation set as output to obtain a biased interpretation generator. The question to be aligned with preference is input into the biased interpretation generator to obtain the preference-aligned interpretation, and the preference-aligned interpretation is input into the target large language model to obtain the corresponding reply answer.

2. The question-answer generation method based on large language model preference alignment according to claim 1, characterized in that: The specified prompt includes multiple interpretation types.

3. The question-answer generation method based on large language model preference alignment as claimed in claim 1, characterized in that: The step of judging whether the multiple interpretations of each original question are preferred by the target large language model according to the correctness of the reply answer includes the following steps: The reply answer is compared with the target answer for correctness. If it is correct, it is considered that the interpretation corresponding to the reply answer is favored by the target large language model. If it is wrong, it is considered that the interpretation corresponding to the reply answer is not favored by the target large language model.

4. The question-answer generation method based on large language model preference alignment according to claim 1, characterized in that: Before inputting each original question and the corresponding multiple interpretations into the target large language model, low-quality interpretations among the multiple interpretations need to be filtered through an interpretation filtering mechanism, which specifically includes the following steps: Combine the original question with the corresponding multiple interpretations to obtain multiple interpretation pairs; Scoring multiple paraphrase pairs using evaluation metrics; The score of each interpretation pair is compared with the set threshold, and the interpretations corresponding to the interpretation pairs whose scores are less than the set threshold are deleted.

5. The question-answer generation method based on large language model preference alignment according to claim 1, characterized in that: The parameter scale of the pre-trained language model is smaller than that of the target large language model.

6. The question-answer generation method based on large language model preference alignment according to claim 1, characterized in that: The step of obtaining a plurality of original questions for training a target large language model comprises the following steps: Obtaining an original data set for training a biased interpretation generator; the original data set is a data set specified by a user, and the task type and field of the data can be selected by the user according to training requirements; Convert each raw data in the original dataset into the original question and the target answer in question-answer format.

7. A question-answer generation system based on preference alignment of a large language model, characterized in that: include: An acquisition module is used to acquire multiple original questions for training a target large language model, and concatenate each original question with a specified prompt to obtain a concatenation instruction, and input the concatenation instruction into a paraphrase generator to generate multiple paraphrases corresponding to the original question; A classification module is used to input each original question and the corresponding multiple interpretations into the target large language model, obtain multiple reply answers of the target large language model, judge whether the multiple interpretations of each original question are preferred by the target large language model according to the correctness of the reply answers, and divide the multiple interpretations into a biased interpretation set and a non-biased interpretation set according to the preference results of the multiple interpretations; A generation module, which is used to train the pre-trained language model with the unbiased interpretation set as input and the biased interpretation set as output to obtain a biased interpretation generator; The interpretation module is used to input the questions to be aligned with preference into the biased interpretation generator to obtain the interpretation after preference alignment, and input the interpretation after preference alignment into the target large language model to obtain the corresponding reply answer.

8. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for generating questions and answers based on preference alignment of a large language model as described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the question and answer generation method based on large language model preference alignment as described in any one of claims 1 to 6 above is implemented.

Citation Information

Patent Citations

  • Open domain natural language reasoning question-answering system and method driven by large language model

    CN116932708A

  • Text generation-oriented black box knowledge distillation method and system for multi-step collaborative prompt learning

    CN117057414A

  • Adaptive retrieval enhanced large language model construction and question answering method and system, storage medium and program product

    CN119248916A

  • Data processing method and device and storage medium

    CN119293187A

  • Preference alignment training method and system of large language model, medium and electronic equipment

    CN119336899A

Cited By

  • Language model comparison method and device based on preference

    CN120493943A

  • A language model comparison method and device based on preferences

    CN120493943B