Post-training method, device, electronic device and storage medium for question-and-answer large model
By constructing a logical split inference data set and post-training of the Q&A big model, the problem of insufficient inference ability of the Q&A big model in the existing technology is solved, and the logical inference ability of the model is improved.
Patent Information
- Application Number
- CN202510287593.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-12
AI Technical Summary
In the prior art, the reasoning ability of the question-and-answer model is insufficient, resulting in inaccurate answers. Although the RLHF and DPO methods can make the model's answers close to human preferences, they fail to substantially improve the model's reasoning ability.
By obtaining the inference data obtained by logically splitting the original question and its original answer, constructing the inference data set, and post-training the question-and-answer model using this data set, the model can learn to infer the input problems step by step.
It effectively improves the logical reasoning ability of the Q&A big model, so that the model can answer questions more accurately.
Smart Images

Figure CN119808964B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular, to a post-training method, device, electronic device, and storage medium for a large question-and-answer model. Background Art
[0002] In the prior art, the model is often made to answer more accurately through post-training. For example, typical methods are the RLHF (Reinforcement Learning from Human Feedback) method and the DPO (Direct Preference Optimization) method. However, the essence of both the RLHF method and the DPO method is to enable the model to answer questions according to human preferences, that is, to make the answer content of the model close to human preferences. However, neither of these two methods substantially enhances the reasoning ability of the model.
[0003] Therefore, how to increase the reasoning ability of the model to make the model answer more accurately is an urgent problem to be solved. Summary of the Invention
[0004] The purpose of the present invention is to provide a post-training method, device, electronic device, and storage medium for a large question-and-answer model to improve the problems existing in the prior art.
[0005] The embodiments of the present invention can be implemented as follows:
[0006] In a first aspect, the present invention provides a post-training method for a large question-and-answer model, including:
[0007] Obtaining a plurality of inference data, where the inference data is obtained by logically splitting the original question and its original answer;
[0008] Constructing an inference data set based on all the inference data;
[0009] Performing post-training processing on the large question-and-answer model by using the inference data set; the inference data set is used to guide the large question-and-answer model to learn to perform step-by-step reasoning on the input question.
[0010] Optionally, the inference data set includes a first subset, a second subset, a third subset, and a fourth subset; the step of constructing an inference data set based on all the inference data includes:
[0011] For each piece of the inference data, based on the inference data, obtaining at least one first sample by using a progressive inference formula construction rule; the first subset includes all the first samples corresponding to each piece of the inference data;
[0012] For each piece of the inference data, based on the inference data, at least one second sample is obtained by using a progressive decomposition construction rule, and the second subset includes all the second samples corresponding to each piece of the inference data;
[0013] For each piece of the inference data, based on the inference data, at least one third sample is obtained by using an intermediate truncation construction rule, and the third subset includes all the third samples corresponding to each piece of the inference data;
[0014] For each piece of the inference data, based on the inference data, at least one fourth sample is obtained by using a step-by-step inference construction rule, and the fourth subset includes all the fourth samples corresponding to each piece of the inference data.
[0015] Optionally, the inference data includes an original question, N sub-questions with progressive logical levels split from the original question, and sub-answers for each sub-question; the step of obtaining at least one first sample by using a progressive inference construction rule based on the inference data includes:
[0016] The first guiding text, the original question in the inference data, and the first sub-question are concatenated to obtain the first first sample, and the label of the first first sample is set as the first sub-answer in the inference data; the first guiding text is used to instruct the question-answering large model to infer the answer to the last sub-question in the input content based on the input content;
[0017] The first guiding text, the original question in the inference data, the nth sub-question, and each sub-question and its sub-answer before the nth sub-question are concatenated to obtain the nth first sample, and the label of the nth first sample is set as the nth sub-answer in the inference data, 。
[0018] Optionally, the inference data includes an original question, N sub-questions with progressive logical levels split from the original question, and sub-answers for each sub-question; the step of obtaining at least one second sample by using a progressive decomposition construction rule based on the inference data includes:
[0019] The second guiding text, the original question in the inference data, and the first sub-question are concatenated to obtain the first second sample, and the label of the first second sample is set as the second sub-answer in the inference data; the second guiding text is used to instruct the question-answering large model to disassemble the next sub-question of the last sub-question in the input content from the original question of the input content;
[0020] Concatenate the second guiding text, the original question in the inference data, the m-th sub-question, and each sub-question before it to obtain the m-th second sample, and set the label of the m-th second sample as the (m + 1)-th sub-question in the inference data. 。
[0021] Optionally, the inference data includes an original question, N sub-questions with a progressive logical hierarchy split from the original question, and sub-answers for each sub-question; the step of obtaining at least one third sample based on the inference data using the intermediate truncation construction rule includes:
[0022] Concatenate the third guiding text, each other sub-question and its sub-answer in the inference data except the first sub-question and the first sub-answer to obtain the first third sample, and set the label of the first third sample as the first sub-question and the first sub-answer in the inference data; the third guiding text is used to instruct the Q&A large model to infer the missing intermediate sub-questions and their intermediate sub-answers in the input content based on the input content.
[0023] Concatenate the third guiding text, each other sub-question and its sub-answer in the inference data except the m-th sub-question and the m-th sub-answer to obtain the m-th third sample, and set the label of the m-th third sample as the m-th sub-question and the m-th sub-answer in the inference data. 。
[0024] Optionally, the inference data includes an original question, N sub-questions with a progressive logical hierarchy split from the original question, and sub-answers for each sub-question; the step of obtaining at least one fourth sample based on the inference data using the step-by-step inference construction rule includes:
[0025] Concatenate the fourth guiding text and the original question in the inference data to obtain the first fourth sample, and set the label of the first fourth sample as each sub-question and its sub-answer in the inference data; the fourth guiding text is used to instruct the Q&A large model to infer subsequent sub-questions and sub-answers step by step based on the input content.
[0026] Concatenate the fourth guiding text, the original question in the inference data, each sub-question and its sub-answer before the n-th sub-question to obtain the n-th fourth sample, and set the label of the n-th fourth sample as each sub-question and its sub-answer after the (n - 1)-th sub-question in the inference data. 。
[0027] Optionally, the inference data set includes a first subset, a second subset, a third subset, and a fourth subset.
[0028] The steps of post-training the question-answering large model using the inference dataset include:
[0029] Respectively using the first subset, the second subset, the third subset, and the fourth subset to perform post-training processing on the question-answering large model;
[0030] Among them, the first subset is used to guide the question-answering large model to learn to answer the input question step by step, the second subset is used to guide the question-answering large model to learn to logically disassemble the input question, the third subset is used to guide the question-answering large model to learn to complete the missing intermediate process based on step-by-step reasoning, and the fourth subset is used to guide the question-answering large model to learn to answer step by step after logically disassembling the input question.
[0031] In a second aspect, the present invention also provides a post-training device for a question-answering large model, including:
[0032] A data acquisition module, configured to acquire a number of inference data, where the inference data is obtained by logically splitting the original question and its original answer;
[0033] A construction module, configured to construct an inference dataset based on all the inference data;
[0034] A post-training module, configured to perform post-training processing on the question-answering large model using the inference dataset; the inference dataset is used to guide the question-answering large model to learn to perform step-by-step reasoning on the input question.
[0035] In a third aspect, the present invention also provides an electronic device, including: a memory and a processor, where the memory stores a software program, and when the electronic device runs, the processor executes the software program to implement the post-training method of the question-answering large model as described in the first aspect.
[0036] In a fourth aspect, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the post-training method of the question-answering large model as described in the first aspect.
[0037] Compared with the prior art, the embodiments of the present invention provide a post-training method, device, electronic device, and storage medium for a question-answering large model. First, a number of inference data are acquired, and then an inference dataset is constructed based on all the inference data; finally, the inference dataset is used to perform post-training processing on the question-answering large model; since each inference data is obtained by logically splitting the original question and its original answer, the inference dataset can be used to guide the question-answering large model to learn to perform step-by-step reasoning on the input question. The present invention enables the question-answering large model to learn to perform step-by-step reasoning on questions, thereby fundamentally improving the logical reasoning ability of the model. Description of the Drawings
[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as a limitation of the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0039] Figure 1 It is one of the schematic flowcharts of a post-training method for a question-and-answer large model provided by an embodiment of the present invention.
[0040] Figure 2 It is another schematic flowchart of a post-training method for a question-and-answer large model provided by an embodiment of the present invention.
[0041] Figure 3 It is a schematic structural diagram of a post-training device for a question-and-answer large model provided by an embodiment of the present invention.
[0042] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0044] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0045] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0046] In addition, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0047] It should be noted that, without conflict, the features in the embodiments of the present invention can be combined with each other.
[0048] In the prior art, in order to make the content answered by the model more accurate, there are the following two methods:
[0049] The first is to enhance the reasoning ability of the model through post-training, so that the answers of the model are more accurate. Typical methods are the RLHF method and the DPO method. Among them, the implementation principle of the RLHF method is as follows: First, it is necessary to feed data into the model after instruction fine-tuning. After obtaining the answers to the data, human beings label the preferred data, such as labeling the data (x, ylose, ywin); then use the labeled preferred data to train a reward model modelaward; in order to make the distribution of the trained model not differ too much from the distribution of the original model, a reference model modelrefer is also needed; therefore, three models are required during the training process. The DPO method directly fine-tunes the model on the preference dataset and does not require a reward model and a reference model, which is a simplified form of the RLHF method.
[0050] The second is to increase the reasoning time to make the model's answers more accurate. The most common way is to let the model link to an external knowledge base or search engine, so that the model uses the obtained external knowledge as a reference to answer questions. When linking to an external knowledge base, the BGE (Bidirectional Generative Embedding) module is needed to segment and vectorize the content of the external knowledge base in advance, so that the model can efficiently retrieve and utilize this knowledge.
[0051] For the first method above, in essence, it only makes the model answer questions according to human preferences, and does not substantially enhance the reasoning ability of the model. Moreover, the standard RLHF method requires three models simultaneously during training, which has higher requirements for hardware.
[0052] For the second method above, it only provides more knowledge for the model to refer to, and also does not substantially increase the reasoning ability of the model. And the reasoning process needs to access the bge module, which increases the additional reasoning time.
[0053] Therefore, how to make the model answer more accurately by increasing the reasoning ability of the model is an urgent problem to be solved.
[0054] Based on the discovery of the above technical problems, the inventor has put forward the following technical solutions through creative labor to solve or improve the above problems. It should be noted that the defects existing in the above solutions in the prior art are all the results obtained by the inventor through practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the embodiments of the present application below for the above problems should be the contributions made by the inventor to this application during the invention creation process, and should not be understood as the technical content known to those skilled in the art.
[0055] An embodiment of the present invention provides a post-training method for a question-and-answer large model, which can obtain inference data obtained by logically splitting the original question and its original answer, construct an inference data set by using several such inference data, and thus use the inference data set to perform post-training on the question-and-answer large model, so that the question-and-answer large model gradually learns to perform step-by-step reasoning on questions during the post-training process, thereby fundamentally improving the logical reasoning ability of the model.
[0056] Please refer to Figure 1 , Figure 1 FIG. is a schematic flowchart of a post-training method for a question-and-answer large model provided by an embodiment of the present invention. The execution subject of this method can be, but is not limited to, computing devices such as smart phones, personal notebooks, personal computers, and servers. The post-training method for the question-and-answer large model includes the following steps:
[0057] S101. Obtain a number of inference data.
[0058] In this embodiment, each piece of inference data can be obtained by logically splitting an original question and its original answer.
[0059] S102. Construct an inference data set based on all the inference data.
[0060] S103. Use the inference data set to perform post-training processing on the question-and-answer large model.
[0061] In this embodiment, the inference data set is used to guide the question-and-answer large model to learn to perform step-by-step reasoning on the input question. Therefore, in the post-training stage, the question-and-answer large model can gradually learn to perform step-by-step logical reasoning on the input question.
[0062] The post-training method for the question-and-answer large model according to the embodiment of the present invention first obtains a number of inference data, then constructs an inference data set based on all the inference data; finally, uses the inference data set to perform post-training processing on the question-and-answer large model; since each piece of inference data is obtained by logically splitting the original question and its original answer, the inference data set can be used to guide the question-and-answer large model to learn to perform step-by-step reasoning on the input question. The present invention enables the question-and-answer large model to learn to perform step-by-step reasoning on questions, thereby fundamentally improving the logical reasoning ability of the model.
[0063] Optionally, a number of original questions and their original answers can be collected, and then each original question can be logically split into multiple sub-questions and their sub-answers by means of manual splitting or using an additional logical splitting model. Among them, if a logical splitting model is used, it is necessary to manually verify the sub-questions and their sub-answers to ensure that there are no logical errors.
[0064] Therefore, each piece of the inference data may include an original question, N sub-questions with a progressive logical hierarchy split from the original question, and sub-answers for each sub-question, where the Nth sub-answer in the inference data is the original answer to the original question.
[0065] In an optional example, assume the original question is: What is the result of 1 + 3×9 + 9÷2, and the original answer is 32.5. The following 4 sub-questions and their sub-answers can be obtained through splitting:
[0066] The first sub-question Q1: What is the result of 3×9, and the first sub-answer A1: 27;
[0067] The second sub-question Q2: What is the result of 9÷2, and the second sub-answer A2: 4.5;
[0068] The third sub-question Q3: What is the result of 1 + 27, and the third sub-answer A3: 28;
[0069] The fourth sub-question Q4: What is the result of 28 + 4.5, and the fourth sub-answer A4: 32.5.
[0070] This example is only a splitting example of a simple logical reasoning problem and is not limited here.
[0071] In an optional implementation, if the process of the Q&A large model learning to perform logical reasoning is split, then it is hoped that the Q&A large model can learn from shallow to deep: disassemble questions progressively, reason about sub-questions progressively, and perform step-by-step logical reasoning completely. And in some scenarios, the logical questions that need to be answered by the Q&A large model may be: presenting a complete question and some sub-questions of the complete question, and it is necessary to deduce the missing intermediate questions and their answers. Therefore, the inference data set constructed in the present invention may include four subsets: a first subset, a second subset, a third subset, and a fourth subset, and these four subsets are constructed in four different ways respectively.
[0072] Please refer to Figure 2 , that is, the process of "constructing an inference data set based on all the inference data" in the above step S102 may include the following sub-steps S1021~S1024.
[0073] S1021. For each piece of the inference data, based on the inference data, obtain at least one first sample by using a progressive inference construction rule.
[0074] In this embodiment, the first subset includes all the first samples corresponding to each piece of the inference data. That is, each piece of inference data corresponds to at least one first sample in the first subset.
[0075] Among them, the first subset is obtained through the first construction method, and the first construction method adopts a progressive reasoning construction rule. Specifically, for each reasoning sample, the process of "obtaining at least one first sample based on the reasoning data by using the progressive reasoning construction rule" in step S1021 may include the following sub-steps S10211 to S10212:
[0076] S10211. Concatenate the first guiding text, the original question and the first sub-question in the reasoning data to obtain the first first sample, and set the label of the first first sample as the first sub-answer in the reasoning data;
[0077] S10212. Concatenate the first guiding text, the original question, the nth sub-question and each sub-question and its sub-answer before the nth sub-question in the reasoning data to obtain the nth first sample, and set the label of the nth first sample as the nth sub-answer in the reasoning data. 。
[0078] In this embodiment, the preset first guiding text can be used to instruct the Q&A large model to infer the answer to the last sub-question in the input content based on the input content. The actual text content of the first guiding text is not limited here.
[0079] S1022. For each piece of the reasoning data, obtain at least one second sample based on the reasoning data by using the progressive decomposition construction rule.
[0080] In this embodiment, the second subset includes all second samples corresponding to each piece of the reasoning data. That is, each piece of reasoning data corresponds to at least one second sample in the second subset.
[0081] Among them, the second subset is obtained through the second construction method, and the second construction method adopts a progressive decomposition construction rule. Specifically, for each reasoning sample, the process of "obtaining at least one second sample based on the reasoning data by using the progressive decomposition construction rule" in step S1022 may include the following sub-steps S10221 to S10222:
[0082] S10221. Concatenate the second guiding text, the original question and the first sub-question in the reasoning data to obtain the first second sample, and set the label of the first second sample as the second sub-answer in the reasoning data;
[0083] S10222. Concatenate the second guiding text, the original question, the mth sub-question and each sub-question before it in the reasoning data to obtain the mth second sample, and set the label of the mth second sample as the (m + 1)th sub-question in the reasoning data. 。
[0084] In this embodiment, the preset second guiding text can be used to instruct the question-and-answer large model to disassemble the next sub-question of the last sub-question in the input content from the original question of the input content based on the input content. The actual text content of the second guiding text is not limited herein.
[0085] S1023. For each piece of the inference data, based on the inference data, use the intermediate interception construction rule to obtain at least one third sample.
[0086] In this embodiment, the third subset includes all the third samples corresponding to each piece of the inference data. That is, each piece of inference data corresponds to at least one third sample in the third subset.
[0087] Among them, the third subset is obtained by the third construction method. The third construction method adopts the intermediate interception construction rule. Specifically, for each inference sample, the process of "based on the inference data, use the intermediate interception construction rule to obtain at least one third sample" in step S1023 may include the following sub-steps S10231 to S10232:
[0088] S10231. Concatenate the third guiding text with each other sub-question and other sub-answer in the inference data except for the first sub-question and the first sub-answer to obtain the first third sample, and set the label of the first third sample as the first sub-question and the first sub-answer in the inference data;
[0089] S10232. Concatenate the third guiding text with each other sub-question and other sub-answer in the inference data except for the m-th sub-question and the m-th sub-answer to obtain the m-th third sample, and set the label of the m-th third sample as the m-th sub-question and the m-th sub-answer in the inference data. 。
[0090] In this embodiment, the third guiding text can be used to instruct the question-and-answer large model to infer the missing intermediate sub-question and its intermediate sub-answer in the input content based on the input content. The actual text content of the third guiding text is not limited herein.
[0091] S1024. For each piece of the inference data, based on the inference data, use the step-by-step inference construction rule to obtain at least one fourth sample.
[0092] In this embodiment, the fourth subset includes all the fourth samples corresponding to each piece of the inference data. That is, each piece of inference data corresponds to at least one fourth sample in the fourth subset.
[0093] Among them, the fourth subset is obtained by the fourth construction method, and the fourth construction method adopts a step-by-step reasoning construction rule. Specifically, for each reasoning sample, the process of "obtaining at least one fourth sample based on the reasoning data by using the step-by-step reasoning construction rule" in step S1024 may include the following sub-steps S10241 to S10242:
[0094] S10241. Concatenate the fourth guiding text and the original question in the reasoning data to obtain the first fourth sample, and set the label of the first fourth sample as each sub-question and its sub-answer in the reasoning data;
[0095] S10242. Concatenate the fourth guiding text, the original question in the reasoning data, each sub-question and its sub-answer before the nth sub-question to obtain the nth fourth sample, and set the label of the nth fourth sample as each sub-question and its sub-answer after the (n - 1)th sub-question in the reasoning data. 。
[0096] In this embodiment, the preset fourth guiding text can be used to instruct the question-answering large model to step by step reason out subsequent sub-questions and sub-answers based on the input content. Wherein, the actual text content of the fourth guiding text is not limited here.
[0097] In an optional example, assume that a reasoning data includes an original question, 5 sub-questions (Q1 to Q5) and 5 sub-answers (A1 to A5) corresponding to the 5 sub-questions one by one. Then, the reasoning data is processed by using four construction methods respectively, and the obtained samples are as follows:
[0098] (1) By using the first construction method, 5 first samples as shown in Table 1 below can be obtained, where prompt1 represents the first guiding text:
[0099] Table 1
[0100]
[0101] (2) By using the second construction method, 4 second samples as shown in Table 2 below can be obtained, where prompt2 represents the second guiding text:
[0102] Table 2
[0103]
[0104] (3) By using the third construction method, 4 third samples as shown in Table 3 below can be obtained, where prompt3 represents the third guiding text:
[0105] Table 3
[0106]
[0107] (4) Using the fourth construction method, five fourth samples as shown in Table 4 below can be obtained, where prompt4 represents the fourth guiding text:
[0108] Table 4
[0109]
[0110] The above examples are only for illustration, and the embodiments of the present invention do not limit the number of sub-questions in a piece of inference data.
[0111] Based on the four subsets in the inference dataset, the sub-steps of step S103 may include S1031 or S1032:
[0112] S1031. Respectively use the first subset, the second subset, the third subset, and the fourth subset to perform post-training processing on the question-answering large model.
[0113] S1032. Sequentially use the second subset, the first subset, the fourth subset, and the third subset to perform post-training processing on the question-answering large model.
[0114] Combining the content of the four tables in the above examples, it can be seen that the first subset can be used to guide the question-answering large model to learn to answer the input question step by step, the second subset can be used to guide the question-answering large model to learn to logically disassemble the input question, the third subset can be used to guide the question-answering large model to learn to complete the missing intermediate process based on the step-by-step reasoning method, and the fourth subset can be used to guide the question-answering large model to learn to answer the input question step by step after logically disassembling it.
[0115] Therefore, the 4 subsets can be used to perform post-training processing on the question-answering large model in any order. It is also possible to sequentially use the second subset, the first subset, the fourth subset, and the third subset to perform post-training processing on the question-answering large model, so that the question-answering large model can gradually learn: logical disassembly, step-by-step answering, step-by-step answering after logical disassembly, and completing the missing intermediate process based on the step-by-step reasoning method.
[0116] Optionally, during the post-training processing, methods such as supervised fine-tuning (SFT) and instruction tuning can be adopted. Among them, the specific process of post-training is prior art and will not be elaborated here.
[0117] It should be noted that the execution order of each step in the above method embodiments is not limited by the figures shown, and the execution order of each step shall be subject to the actual application situation.
[0118] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0119] The present invention constructs four subsets through four construction methods respectively to perform post-training processing on the question-answering large model, thereby essentially enhancing the model's reasoning ability;
[0120] In the present invention, the first subset is used to guide the question-answering large model to learn to answer the input question step by step, so that the question-answering large model learns to answer the newly split sub-question based on the original question and the sub-questions and their sub-answers obtained from the previous split;
[0121] In the present invention, the second subset can be used to guide the question-answering large model to learn to logically disassemble the input question, so that the question-answering large model learns to continue to split the next sub-question based on the original question and the split sub-questions;
[0122] In the present invention, the third subset can be used to guide the question-answering large model to learn to complete the missing intermediate process based on the step-by-step reasoning method, so that the question-answering large model learns to answer the reasoning questions with missing intermediate steps through step-by-step reasoning;
[0123] In the present invention, the fourth subset can be used to guide the question-answering large model to learn to answer step by step after logically disassembling the input question, so that the question-answering large model learns to perform complete step-by-step logical reasoning on the input question.
[0124] In order to execute the corresponding steps in the above method embodiments and all possible implementation manners, the following provides an implementation manner of a post-training device for a question-answering large model.
[0125] Please refer to Figure 3 , Figure 3 which shows a schematic structural diagram of a post-training device for a question-answering large model provided by an embodiment of the present invention. The post-training device 200 for the question-answering large model includes: a data acquisition module 210, a construction module 220, and a post-training module 230.
[0126] The data acquisition module 210 is used to acquire a plurality of inference data, and the inference data is obtained by logically disassembling the original question and its original answer;
[0127] The construction module 220 is used to construct an inference data set based on all the inference data;
[0128] A post-training module 230 for post-training the question-answering large model using the inference dataset; the inference dataset is used to guide the question-answering large model to learn to perform step-by-step reasoning on the input question.
[0129] Optionally, the inference dataset includes a first subset, a second subset, a third subset, and a fourth subset. In the process of constructing the inference dataset based on all the inference data, the construction module 220 can specifically be used to: for each piece of the inference data, based on the inference data, obtain at least one first sample using the progressive reasoning construction rule; the first subset includes all the first samples corresponding to each piece of the inference data; for each piece of the inference data, based on the inference data, obtain at least one second sample using the progressive decomposition construction rule, and the second subset includes all the second samples corresponding to each piece of the inference data; for each piece of the inference data, based on the inference data, obtain at least one third sample using the middle intercept construction rule, and the third subset includes all the third samples corresponding to each piece of the inference data; for each piece of the inference data, based on the inference data, obtain at least one fourth sample using the step-by-step reasoning construction rule, and the fourth subset includes all the fourth samples corresponding to each piece of the inference data.
[0130] Optionally, the inference data includes an original question, N sub-questions with progressive logical levels split from the original question, and sub-answers for each sub-question. In the process of obtaining at least one first sample using the progressive reasoning construction rule based on the inference data, the construction module 220 can specifically be used to: splice a first guiding text, the original question in the inference data, and the first sub-question to obtain a first first sample, and set the label of the first first sample to the first sub-answer in the inference data; the first guiding text is used to instruct the question-answering large model to infer the answer to the last sub-question in the input content based on the input content; splice the first guiding text, the original question in the inference data, the nth sub-question, and each sub-question and its sub-answer before the nth sub-question to obtain the nth first sample, and set the label of the nth first sample to the nth sub-answer in the inference data. 。
[0131] Optionally, the inference data includes an original question, N sub-questions with progressive logic levels split from the original question, and sub-answers for each sub-question. In the process of the construction module 220 obtaining at least one second sample based on the inference data and using the progressive decomposition construction rule, it can specifically be used to: splice a second guiding text, the original question in the inference data, and the first sub-question to obtain a first second sample, and set the label of the first second sample as the second sub-answer in the inference data; the second guiding text is used to instruct the question-and-answer large model to disassemble the next sub-question of the last sub-question in the input content from the original question of the input content based on the input content; splice the second guiding text, the original question in the inference data, the m-th sub-question, and each sub-question before it to obtain the m-th second sample, and set the label of the m-th second sample as the (m + 1)-th sub-question in the inference data, 。
[0132] Optionally, the inference data includes an original question, N sub-questions with progressive logic levels split from the original question, and sub-answers for each sub-question. In the process of the construction module 220 obtaining at least one third sample based on the inference data and using the middle interception construction rule, it can specifically be used to: splice a third guiding text, each other sub-question and other sub-answer in the inference data except the first sub-question and the first sub-answer to obtain a first third sample, and set the label of the first third sample as the first sub-question and the first sub-answer in the inference data; the third guiding text is used to instruct the question-and-answer large model to infer the missing middle sub-question and its middle sub-answer in the input content based on the input content; splice the third guiding text, each other sub-question and other sub-answer in the inference data except the m-th sub-question and the m-th sub-answer to obtain the m-th third sample, and set the label of the m-th third sample as the m-th sub-question and the m-th sub-answer in the inference data, 。
[0133] Optionally, the inference data includes an original question, N sub-questions with progressive logical levels split from the original question, and sub-answers for each sub-question. In the process of using the step-by-step inference construction rule to obtain at least one fourth sample based on the inference data, the construction module 220 can specifically be used to: splice the fourth guiding text and the original question in the inference data to obtain the first fourth sample, and set the label of the first fourth sample as each sub-question and its sub-answer in the inference data; the fourth guiding text is used to instruct the Q&A large model to step by step infer subsequent sub-questions and sub-answers based on the input content; splice the fourth guiding text, the original question in the inference data, each sub-question and its sub-answer before the nth sub-question to obtain the nth fourth sample, and set the label of the nth fourth sample as each sub-question and its sub-answer after the (n - 1)th sub-question in the inference data. 。
[0134] Optionally, the inference data set includes a first subset, a second subset, a third subset, and a fourth subset. In the process of using the inference data set to perform post-training processing on the Q&A large model, the post-training module 230 can specifically be used to: respectively use the first subset, the second subset, the third subset, and the fourth subset to perform post-training processing on the Q&A large model; wherein, the first subset is used to guide the Q&A large model to learn to step by step answer the input question, the second subset is used to guide the Q&A large model to learn to logically disassemble the input question, the third subset is used to guide the Q&A large model to learn to complete the missing intermediate process based on the step-by-step inference method, and the fourth subset is used to guide the Q&A large model to learn to step by step answer the input question after logically disassembling it.
[0135] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described post-training device 200 of the Q&A large model can refer to the corresponding process in the foregoing method embodiment, and will not be elaborated herein.
[0136] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 300 includes a processor 310, a memory 320, and a bus 330. The processor 310 is connected to the memory 320 through the bus 330.
[0137] The memory 320 can be used to store software programs, for example, the software program corresponding to the post-training device 200 of the question-and-answer large model provided by the embodiments of the present invention. The processor 310 executes various functional applications and data processing by running the software program stored in the memory 320, so as to implement the post-training method of the question-and-answer large model provided by the embodiments of the present invention.
[0138] Among them, the memory 320 can be but is not limited to: RAM (Random Access Memory), ROM (Read Only Memory), FLASH (Flash Memory), PROM (Programmable Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electric Erasable Programmable Read-Only Memory), etc.
[0139] The processor 310 can be an integrated circuit chip with signal processing capabilities. The processor 310 can be a general-purpose processor, including: CPU (Central Processing Unit), NP (Network Processor), SoC (System on Chip), etc.; it can also be: DSP (Digital Signal Processing), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0140] It can be understood that Figure 4 the structure shown is only for illustration, and the electronic device 300 may also include more or fewer components than Figure 4 those shown, or have a different configuration from Figure 4 that shown. Figure 4 Each component shown can be implemented by hardware, software, or a combination thereof.
[0141] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the post-training method of the question-and-answer large model disclosed in the above embodiment. The computer-readable storage medium may be, but is not limited to: various media such as USB flash drives, external hard drives, ROM, RAM, PROM, EPROM, EEPROM, FLASH disks or optical discs that can store program codes.
[0142] In summary, the embodiments of the present invention provide a post-training method, device, electronic device and storage medium for a question-and-answer large model. First, a number of inference data are obtained, and then an inference data set is constructed based on all the inference data; finally, the inference data set is used to perform post-training processing on the question-and-answer large model; since each piece of inference data is obtained by logically splitting the original question and its original answer, the inference data set can be used to guide the question-and-answer large model to learn to perform step-by-step reasoning on the input question. The present invention enables the question-and-answer large model to learn to perform step-by-step reasoning on questions, thereby fundamentally improving the logical reasoning ability of the model.
[0143] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A post-training method for a large question-answering model, characterized in that: include: obtaining some inference data; For each piece of the inference data, based on the inference data, at least one first sample is obtained by using a progressive inference formula construction rule, wherein the first subset includes all first samples corresponding to each piece of the inference data; For each piece of the inference data, based on the inference data, at least one second sample is obtained by using a progressive decomposition construction rule, and the second subset includes all the second samples corresponding to each piece of the inference data; For each of the inference data, based on the inference data, at least one third sample is obtained by using an intermediate interception construction rule, and the third subset includes all the third samples corresponding to each of the inference data; For each piece of the reasoning data, based on the reasoning data, at least one fourth sample is obtained by using a step-by-step reasoning construction rule, wherein the fourth subset includes all fourth samples corresponding to each piece of the reasoning data; The question-answering big model is post-trained using an inference data set, the inference data set is used to guide the question-answering big model to learn to perform step-by-step inference on input questions, the inference data set includes the first subset, the second subset, the third subset and the fourth subset, the first subset is used to guide the question-answering big model to learn to answer the input questions step by step, the second subset is used to guide the question-answering big model to learn to logically decompose the input questions, the third subset is used to guide the question-answering big model to learn to complete the missing intermediate process based on step-by-step reasoning, and the fourth subset is used to guide the question-answering big model to learn to logically decompose the input questions and then answer them step by step; The inference data includes the original question, N sub-questions with progressive logical levels obtained by splitting the original question, and sub-answers to each sub-question, and the step of obtaining at least one third sample based on the inference data by using an intermediate interception construction rule includes: The third guiding text, each other sub-question and other sub-answers in the inference data except the first sub-question and the first sub-answer are concatenated to obtain a first third sample, and the label of the first third sample is set to the first sub-question and the first sub-answer in the inference data, and the third guiding text is used to instruct the question-answering large model to infer the missing intermediate sub-questions and their intermediate sub-answers in the input content based on the input content; The third guide text, each other sub-question and other sub-answer in the reasoning data except the m-th sub-question and the m-th sub-answer are concatenated to obtain the m-th third sample, and the label of the m-th third sample is set to the m-th sub-question and the m-th sub-answer in the reasoning data. .
2. The method according to claim 1, characterized in that The step of obtaining at least one first sample by constructing a rule based on the inference data using a progressive inference formula includes: The first guiding text, the original question in the inference data, and the first sub-question are concatenated to obtain a first first sample, and the label of the first first sample is set to the first sub-answer in the inference data, wherein the first guiding text is used to instruct the question-answering model to infer the answer to the last sub-question in the input content based on the input content; The first guide text, the original question in the reasoning data, the nth sub-question, and each sub-question and its sub-answer before the nth sub-question are concatenated to obtain the nth first sample, and the label of the nth first sample is set to the nth sub-answer in the reasoning data. .
3. The method according to claim 1, characterized in that The step of obtaining at least one second sample based on the inference data by using a progressive decomposition construction rule comprises: The second guide text, the original question in the inference data, and the first sub-question are concatenated to obtain a first second sample, and the label of the first second sample is set to the second sub-answer in the inference data, wherein the second guide text is used to instruct the question-answering model to decompose the next sub-question of the last sub-question in the input content from the original question of the input content based on the input content; The second guide text, the original question in the reasoning data, the mth sub-question and each of the previous sub-questions are concatenated to obtain the mth second sample, and the label of the mth second sample is set to the m+1th sub-question in the reasoning data. .
4. The method according to claim 1, characterized in that: The step of obtaining at least one fourth sample based on the inference data by constructing a rule using a step-by-step inference formula comprises: The fourth guide text and the original question in the reasoning data are concatenated to obtain a first fourth sample, and a label of the first fourth sample is set to each sub-question and its sub-answer in the reasoning data, wherein the fourth guide text is used to instruct the question-answering model to infer subsequent sub-questions and sub-answers step by step based on the input content; The fourth guide text, the original question in the reasoning data, each sub-question before the nth sub-question and its sub-answer are concatenated to obtain the nth fourth sample, and the label of the nth fourth sample is set to each sub-question after the n-1th sub-question in the reasoning data and its sub-answer, .
5. The method according to claim 1, characterized in that The step of post-training the question-answering large model using the inference data set includes: The first subset, the second subset, the third subset and the fourth subset are respectively used to perform post-training processing on the question-answering large model.
6. A post-training device for a large question-answering model, characterized in that: include: A data acquisition module, used to acquire a number of inference data, wherein the inference data is obtained by logically splitting the original question and its original answer; Construction modules for: For each piece of the inference data, based on the inference data, at least one first sample is obtained by using a progressive inference formula construction rule, wherein the first subset includes all first samples corresponding to each piece of the inference data; For each piece of the inference data, based on the inference data, at least one second sample is obtained by using a progressive decomposition construction rule, and the second subset includes all the second samples corresponding to each piece of the inference data; For each of the inference data, based on the inference data, at least one third sample is obtained by using an intermediate interception construction rule, and the third subset includes all the third samples corresponding to each of the inference data; For each piece of the reasoning data, based on the reasoning data, at least one fourth sample is obtained by using a step-by-step reasoning construction rule, wherein the fourth subset includes all fourth samples corresponding to each piece of the reasoning data; a post-training module, for performing post-training processing on the question-answering big model using an inference data set, the inference data set being used to guide the question-answering big model to learn to perform step-by-step inference on input questions, the inference data set comprising the first subset, the second subset, the third subset and the fourth subset, the first subset being used to guide the question-answering big model to learn to answer the input questions step by step, the second subset being used to guide the question-answering big model to learn to logically decompose the input questions, the third subset being used to guide the question-answering big model to learn to complete the missing intermediate process based on step-by-step reasoning, and the fourth subset being used to guide the question-answering big model to learn to logically decompose the input questions and then answer them step by step; The inference data includes an original question, N sub-questions with progressive logical levels obtained by splitting the original question, and a sub-answer to each sub-question. The construction module is used to obtain at least one third sample based on the inference data by using an intermediate interception construction rule, specifically for: The third guiding text, each other sub-question and other sub-answers in the inference data except the first sub-question and the first sub-answer are concatenated to obtain a first third sample, and the label of the first third sample is set to the first sub-question and the first sub-answer in the inference data, and the third guiding text is used to instruct the question-answering large model to infer the missing intermediate sub-questions and their intermediate sub-answers in the input content based on the input content; The third guide text, each other sub-question and other sub-answer in the reasoning data except the m-th sub-question and the m-th sub-answer are concatenated to obtain the m-th third sample, and the label of the m-th third sample is set to the m-th sub-question and the m-th sub-answer in the reasoning data. .
7. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a software program, and when the electronic device is running, the processor executes the software program to implement the post-training method of the question-answering large model as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the post-training method of the question-answering large model described in any one of claims 1-5.
Citation Information
Patent Citations
Large model training method, device and equipment and question answering method, device and equipment based on large model
CN118469019A
Intelligent answering method, system and device for multi-step reasoning problem and medium
CN118674056A
Campus spoofing prevention and control management method, device, system and equipment and storage medium
CN118964561A