Question and answer model training method and question and answer method and device based on large model
By training the Q&A model in stages, the preset prompt word data set is used to gradually meet multiple training goals, solving the problem that existing models are difficult to learn multiple goals at the same time, and improving the model's adaptability and human-computer interaction experience.
Patent Information
- Application Number
- CN202510157456.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-16
AI Technical Summary
It is difficult for existing artificial intelligence models to learn multiple goals at the same time during training, which affects the human-computer interaction question-and-answer effect of the model and reduces the user's human-computer interaction experience.
By obtaining the preset prompt word dataset, traversing the dataset, and training the trained question-and-answer model in stages, so that the model gradually meets multiple training goals.
It improves the adaptability and flexibility of model training, so that the model can better learn and adapt to different goals, and improves the experience and question-and-answer effect of human-computer interaction.
Smart Images

Figure CN120012938A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of human-computer interaction in the field of artificial intelligence, and in particular to a training method for a question-answering model, and a question-answering method and device based on a large model. Background Art
[0002] In the actual application of human-computer interaction, many tasks cannot be completed by a single goal. For example, in natural language processing tasks, artificial intelligence models may need to meet multiple goals such as accuracy, consistency, and semantic richness at the same time.
[0003] Current artificial intelligence models find it difficult to learn multiple objectives during training, which affects the model's human-computer interaction question-and-answer effect and reduces the user's human-computer interaction experience. Summary of the invention
[0004] The present invention provides a training method for a question-answering model, and a question-answering method and device based on a large model.
[0005] According to a first aspect of the present disclosure, a method for training a question-answering model is provided, comprising:
[0006] Acquire a preset prompt word data set; wherein the preset prompt word data set includes prompt word information, the prompt word information corresponds to a training target one by one, and the training target represents the requirement for the output data of the model;
[0007] By traversing the preset prompt word data set, the question-answering model to be trained is trained at least according to the currently traversed prompt word information to obtain the current question-answering model; wherein the question-answering model meets the training target corresponding to the traversed prompt word information;
[0008] In response to determining that the traversal of the preset prompt word data set is completed, determining that the current question-answering model is a trained question-answering model.
[0009] According to a second aspect of the present disclosure, a large model-based question-answering method is provided, comprising:
[0010] Receive question information input by the user;
[0011] The question information is input into the question-answering model, and based on the model prompt words, the reply information corresponding to the question information is obtained; wherein the question-answering model representation weights are any one of the trained question-answering models 1 to 11, and the model prompt words are used to guide the question-answering model to generate reply information.
[0012] According to a third aspect of the present disclosure, a training device for a question-answering model is provided, comprising:
[0013] An acquisition unit, configured to acquire a preset prompt word data set; wherein the preset prompt word data set includes prompt word information, the prompt word information corresponds to a training target one by one, and the training target represents a requirement for output data of the model;
[0014] A training unit, configured to train the question-answering model to be trained by traversing the preset prompt word data set, at least according to the prompt word information currently traversed, to obtain a current question-answering model; wherein the question-answering model satisfies the training goal corresponding to the prompt word information that has been traversed;
[0015] A determination unit is used to determine that the current question-answering model is a trained question-answering model in response to determining that the traversal of the preset prompt word data set is completed.
[0016] According to a fourth aspect of the present disclosure, a question-answering device based on a large model is provided, comprising:
[0017] A receiving unit, used for receiving question information input by a user;
[0018] A reply unit is used to input the question information into the question-answering model, and obtain reply information corresponding to the question information based on the model prompt words; wherein the question-answering model represents the trained question-answering model described in the third aspect, and the model prompt words are used to guide the question-answering model to generate reply information.
[0019] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:
[0020] at least one processor; and
[0021] a memory communicatively coupled to the at least one processor;
[0022] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to the first aspect and the second aspect of the present disclosure.
[0023] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method according to the first and second aspects of the present disclosure.
[0024] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the steps of the method described in the first aspect and the second aspect of the present disclosure when executed by a processor.
[0025] According to the technology disclosed in the present invention, the adaptability and flexibility of model training are improved, so as to better learn different objectives.
[0026] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0028] Figure 1 It is a flowchart of a training method for a question-answering model provided according to an embodiment of the present disclosure;
[0029] Figure 2 is a schematic diagram of phased training of a question-answering model provided according to an embodiment of the present disclosure;
[0030] Figure 3 It is a flowchart of a training method for a question-answering model provided according to an embodiment of the present disclosure;
[0031] Figure 4 It is a flowchart of a question-answering method based on a large model provided according to an embodiment of the present disclosure;
[0032] Figure 5 It is a structural block diagram of a training device for a question-answering model provided according to an embodiment of the present disclosure;
[0033] Figure 6 It is a structural block diagram of a training device for a question-answering model provided according to an embodiment of the present disclosure;
[0034] Figure 7 is a structural block diagram of a question-answering device based on a large model provided according to an embodiment of the present disclosure;
[0035] Figure 8 is a block diagram of an electronic device for implementing the training method of the question-answering model and the question-answering method based on a large model according to an embodiment of the present disclosure;
[0036] Fig. 9 It is a block diagram of an electronic device used to implement the training method of the question-answering model and the question-answering method based on a large model of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0037] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0038] In the practical application of modern artificial intelligence, many tasks cannot be quantified and optimized by a single goal. This is because these tasks usually involve multi-dimensional standards and goals. For example, in natural language processing tasks, the model may need to meet multiple goals such as accuracy, consistency, and semantic richness at the same time. However, current human feedback reinforcement learning methods, such as DPO (Distributional Policy Optimization), show certain limitations in supporting multi-objective optimization. Specifically, it is difficult for DPO to effectively learn the rewards and penalties of multiple objectives at the same time, resulting in limited adaptability and generalization capabilities in complex environments.
[0039] In the same period of research, such as multi-objective reinforcement learning human feedback, attempts were made to express different preferences by weighting multiple reward functions. However, this approach relies on the definition of a clear reward function for each goal, which is difficult to achieve in many application scenarios and will affect the effectiveness of the model without a clear reward function.
[0040] The present invention provides a question-answering model training method, a large-model-based question-answering method and a device, which are applied to the human-computer interaction field in the field of artificial intelligence to improve the flexibility of model training and enhance the human-computer interaction experience.
[0041] It should be noted that the model in this embodiment is not a model for a specific user and cannot reflect the personal information of a specific user. It should be noted that the data in this embodiment comes from a public data set.
[0042] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0043] In order to enable readers to more deeply understand the implementation principle of the present disclosure, the following Figure 1-Figure 9 The embodiment is further refined.
[0044] Figure 1 FIG. 1 is a flow chart of a method for training a question-answering model according to an embodiment of the present disclosure, and the method can be performed by a training device for a question-answering model. Figure 1 As shown, the method comprises the following steps:
[0045] S101, obtaining a preset prompt word data set; wherein the preset prompt word data set includes prompt word information, the prompt word information corresponds to a training target one by one, and the training target represents the requirement for the output data of the model.
[0046] Exemplarily, multiple prompts are pre-set, and a prompt word data set is composed of multiple prompts. That is, the prompt word data set includes multiple prompts, that is, prompt word information. Each prompt corresponds to a training target. The training target is the requirement that the output data of the model is expected to meet during model training, that is, the ability that the model is expected to have. For example, the training target can be accuracy, then the output data of the model is expected to be accurate; the training target can be semantic richness, then the semantics of the output data of the model is expected to be rich.
[0047] The prompt word information can be text data of preset content, which can be used to guide the model to generate output data. The content of the prompt word information corresponding to different training targets can be different. For example, the prompt word information can be "You are an intelligent agent with XXX ability". Among them, "XXX" represents the training target, and different content can be replaced according to different training targets.
[0048] Preset multiple prompt word information to build a prompt word data set. When training the model, you can directly obtain the prompt word data set, or adjust or design the prompt word information at any time during model training. The model in this embodiment can be a large model for human-computer interaction. For example, the user can input question information to the model, and the model can output reply information based on the question information to achieve human-computer interaction with the user. In this embodiment, the model structure is not specifically limited.
[0049] In this embodiment, in the SFT (Supervised Fine-Tuning) stage, different prompts can be designed for different styles and goals, so that the model can recognize and adapt to the needs of different training goals during training, ensuring that the model can perceive diverse target requirements at an early stage.
[0050] Design specific role-playing prompts for each type of training target. These prompts serve as signals to guide the model to learn specific capabilities. The prompts for the training targets represent the role that the model needs to play. For example, the role played by the model is an expert who is specifically responsible for answering three-view questions. Each role has its own prompt, that is, each training target has its own prompt. For example, for data that requires reasoning ability, the prompt may involve "please reason about the answers to the following questions", while for data that requires value alignment, the prompt may guide the model to consider "answering the following questions from an ethical perspective." By designing role-playing prompts, the need to rely on external classifiers in traditional methods is avoided, the complexity and running time of the model are reduced, and the model effect is improved.
[0051] S102, by traversing the preset prompt word data set, at least according to the currently traversed prompt word information, the question-answering model to be trained is trained to obtain the current question-answering model; wherein the question-answering model meets the training objectives corresponding to the traversed prompt word information.
[0052] Exemplarily, the order of the prompt word information in the prompt word data set may not be limited. For example, each prompt word information may be randomly stored in the prompt word data set. The training order of each training target may also be preset, and the order of the prompt word information in the prompt word data set is determined according to the training order of the training target. For example, the accuracy of the model is trained first, then the consistency of the model is trained, and then the semantic richness of the model is trained. Then, the first prompt word information in the prompt word data set is the prompt word information corresponding to the accuracy, the second prompt word information is the prompt word information corresponding to the consistency, and the third prompt word information is the prompt word information corresponding to the semantic richness.
[0053] Traverse the prompt word data set, determine each prompt word information from the prompt word data set in turn, and each time a prompt word information is traversed, train the question-answering model to be trained according to the currently traversed prompt word information, so that the output data of the question-answering model to be trained can meet the training target corresponding to the prompt word information. The model obtained after training the question-answering model to be trained is determined as the current question-answering model, that is, the current question-answering model is a model that meets the training target corresponding to the currently traversed prompt word information. Each prompt word information can correspond to a current question-answering model after training.
[0054] The question-answering model to be trained is a model that satisfies the training objectives corresponding to the traversed prompt word information, that is, the question-answering model to be trained for each prompt word information is the current question-answering model of the previous prompt word information. For example, when traversing to the second prompt word information in the prompt word data set, for the second prompt word information, the question-answering model to be trained should meet the training objectives corresponding to the first prompt word information, that is, the current question-answering model corresponding to the first prompt word information can be trained; when traversing to the third prompt word information in the prompt word data set, for the third prompt word information, the question-answering model to be trained should meet the training objectives corresponding to the first two prompt word information, that is, the current question-answering model corresponding to the second prompt word information can be trained.
[0055] An initial model may be pre-set as the question-answering model to be trained for the first prompt word information. That is, when the first prompt word information is traversed, the preset initial model may be trained to obtain the current question-answering model corresponding to the first prompt word information.
[0056] S103: In response to determining that the traversal of the preset prompt word data set is completed, determining that the current question-answering model is a trained question-answering model.
[0057] Exemplarily, if it is determined that the traversal of the prompt word data set is completed, the current question-answering model corresponding to the last prompt word information can be determined as the final question-answering model, that is, the model training process is completed. The traversal completion can mean that the model training has been completed based on all the prompt word information. The input data of the question-answering model can be the user's question information, and the output data can be the reply information for the question information, that is, the question-answering model can be used to output reply information based on the question information input by the user.
[0058] Figure 2 Schematic diagram of the phased training of the question-answering model. Figure 2 An initial model is preset in the algorithm. The initial model is trained for the training target of the first prompt word information to obtain a model that meets the first training target. Then, the model that meets the first training target is trained for the training target of the second prompt word information to obtain a model that meets the first and second training targets. Then, the model that meets the first two training targets is trained for the training target of the third prompt word information to obtain a model that meets the first three training targets. Continue to train for subsequent prompt word information until a model that meets all training targets is obtained, that is, the final question-answering model.
[0059] In the disclosed embodiment, a separate prompt word information is preset for each training target of the model to obtain a prompt word data set. The prompt word information can be used to perform role-playing for the training target. The prompt word is used to perform role-playing, that is, the training target that the model needs to complete is explained by the prompt word, so as to achieve role differentiation of different training targets. The prompt word data set is traversed, and model training is performed for each training target according to the prompt word information traversed each time, so that the question-answering model completed each time can meet the requirements of all previous training targets. That is, model training is performed for each training target in turn, and the model trained for the previous training target is used as the reference model for the next training target, so as to achieve staged training. When the last prompt word information is traversed, the final question-answering model is obtained. In the disclosed embodiment, through role-playing and staged training, it is possible to better adapt to different task requirements and goals, improve the adaptability and flexibility of model training, and improve the accuracy of responses.
[0060] Figure 3 A flowchart of a method for training a question-answering model provided in an embodiment of the present disclosure.
[0061] In this embodiment, by traversing a preset prompt word data set, the question-answering model to be trained is trained at least according to the currently traversed prompt word information to obtain the current question-answering model, including: according to the currently traversed prompt word information and the preset data set to be trained, the question-answering model to be trained is trained to obtain the current question-answering model; wherein the preset data set to be trained includes input data to be trained and data labels of the input data to be trained, and the data labels represent the expected output data corresponding to the input data to be trained.
[0062] This embodiment is based on the above embodiment. Figure 3 As shown, the method comprises the following steps:
[0063] S301, obtaining a preset prompt word data set; wherein the preset prompt word data set includes prompt word information, the prompt word information corresponds to a training target one by one, and the training target represents the requirement for the output data of the model.
[0064] Exemplarily, this step may refer to the above-mentioned step S101 and will not be described in detail.
[0065] S302: By traversing the preset prompt word data set, the question-answering model to be trained is trained according to the currently traversed prompt word information and the preset data set to be trained to obtain the current question-answering model.
[0066] Exemplarily, a data set to be trained is pre-set, and the data set to be trained may include input data to be trained and data labels, and the data labels represent the expected output data corresponding to the input data to be trained, that is, the data that the model should output according to the data to be trained. The input data to be trained is various question information collected in advance, and each input data to be trained corresponds to a data label.
[0067] Traverse the prompt word data set to determine the currently traversed prompt word information and the question-answering model to be trained. Input the data set to be trained into the question-answering model to be trained, and train the question-answering model to be trained according to the currently traversed prompt word information. For example, the input data to be trained and the currently traversed prompt word information can be input into the question-answering model to be trained, and the question-answering model to be trained calculates the input data to be trained according to the prompt word information, and iteratively trains the question-answering model to be trained according to the calculation result and the expected output data, and uses the model obtained after training as the current question-answering model.
[0068] In this embodiment, for each training target, the data set to be trained and the corresponding prompt word information can be input into the corresponding question-answering model to be trained for training, so that the model can output results with different requirements for the same or different input data, thereby enabling the model to learn and strengthen in a targeted manner during the training process, thereby improving the training accuracy of the model.
[0069] In this embodiment, the question-answering model to be trained is trained according to the currently traversed prompt word information and the preset data set to be trained to obtain the current question-answering model, including: obtaining the predicted output data corresponding to the input data to be trained according to the currently traversed prompt word information and the preset data set to be trained; wherein the predicted output data represents the data output by the question-answering model to be trained according to the input data to be trained; according to the expected output data corresponding to the input data to be trained and the corresponding predicted output data, obtaining the recognition accuracy of the question-answering model to be trained; wherein the recognition accuracy represents the probability that the question-answering model to be trained meets the training objectives; according to the recognition accuracy, the question-answering model to be trained is trained to obtain the current question-answering model.
[0070] Specifically, the training data set is input into the question-answering model to be trained. The question-answering model to be trained obtains predicted output data according to the training input data in the training data set and the currently traversed prompt word information. The output data is the data actually output by the model corresponding to the training input data, that is, the answer to the training input data predicted by the model. Each training input data can correspond to an expected output data and a predicted output data.
[0071] For each input data to be trained, the recognition accuracy of the question-answering model to be trained can be determined based on the expected output data and predicted output data of these input data to be trained. The recognition accuracy characterizes the degree to which the question-answering model to be trained meets the training objectives, that is, determines whether the question-answering model to be trained can output correct responses. The higher the recognition accuracy, the higher the probability that the question-answering model to be trained meets the training objectives, that is, the question-answering model to be trained is more likely to output correct responses. Each expected output data corresponds to a predicted output data, and the recognition accuracy of the question-answering model to be trained can be determined based on the gap between the expected output data and the corresponding predicted output data. For example, the larger the gap between the expected output data and the corresponding predicted output data, the smaller the recognition accuracy.
[0072] According to the recognition accuracy, the question-answering model to be trained is iterated for multiple times until the current question-answering model is obtained. For example, it can be determined whether the recognition accuracy is greater than a preset accuracy threshold. If not, the model parameters can be adjusted to recalculate a new recognition accuracy until the calculated recognition accuracy is greater than the preset accuracy threshold.
[0073] The beneficial effect of this setting is that each input data to be trained in the data set corresponds to the actual output data of a model as the predicted output data. According to the predicted output data and the expected output data, it can be determined whether the model meets the standard, so as to carry out training and improve the training accuracy of the model.
[0074] In this embodiment, the recognition accuracy of the question-answering model to be trained is obtained based on the expected output data corresponding to the input data to be trained and the corresponding predicted output data, including: determining the recognition result of the input data to be trained based on the expected output data corresponding to the input data to be trained and the corresponding predicted output data; wherein the recognition result represents whether the expected output data and the predicted output data are consistent; determining the recognition accuracy of the question-answering model to be trained based on the recognition results of each input data to be trained.
[0075] Specifically, for each input data to be trained, the recognition result corresponding to the input data to be trained is determined according to the expected output data and the predicted output data of the input data to be trained. The recognition result represents whether the expected output data and the predicted output data of the input data to be trained are consistent. Two identifiers can be preset, indicating that the expected output data and the predicted output data are consistent, and the expected output data and the predicted output data are inconsistent. For example, if the recognition result is true, it means that the expected output data and the predicted output data are consistent; if the recognition result is false, it means that the expected output data and the predicted output data are inconsistent. In other words, it is determined whether the expected output data and the predicted output data of the input data to be trained are consistent. If so, the recognition result is determined to be true; if not, the recognition result is determined to be false.
[0076] Determine the recognition results of each input data to be trained, and determine the recognition accuracy of the question-answering model to be trained based on the recognition results of the input data to be trained. For example, the more the number of true recognition results is, the higher the recognition accuracy is.
[0077] The beneficial effect of such a setting is that the recognition accuracy is determined by combining the recognition results of all the input data to be trained, thereby improving the accuracy of determining the recognition accuracy, thereby improving the accuracy and efficiency of model training.
[0078] In this embodiment, the recognition accuracy of the question-answering model to be trained is determined based on the recognition results of each input data to be trained, including: in response to the recognition result of the input data to be trained indicating that the expected output data and the predicted output data are consistent, the input data to be trained is determined as a positive sample; based on the total number of input data to be trained and the number of positive samples in the training data set, the recognition accuracy of the question-answering model to be trained is determined.
[0079] Specifically, it is determined whether the expected output data and the predicted output data of the input data to be trained are consistent. If so, the input data to be trained is determined as a positive sample; if not, the input data to be trained is determined as a negative sample. That is, the positive sample representation model has learned the input data to be trained, and the negative sample representation model has not learned the input data to be trained.
[0080] Determine the total number of input data to be trained in the training data set and the number of positive samples in the training data set. Determine the recognition accuracy of the question-answering model to be trained based on the total number of input data to be trained and the number of positive samples in the training data set. For example, the number of positive samples can be divided by the total number of input data to be trained, and the ratio obtained is the recognition accuracy.
[0081] The beneficial effect of this setting is that by determining the positive and negative samples, the recognition accuracy of the model can be determined, thereby determining the current effect of the model and improving the efficiency and accuracy of model training.
[0082] In this embodiment, the question-answering model to be trained is trained according to the recognition accuracy to obtain the current question-answering model, including: in response to the recognition accuracy being less than a preset accuracy threshold, updating the preset loss function according to the recognition accuracy to obtain the objective function; according to the objective function, back-propagation training is performed on the question-answering model to be trained to obtain the current question-answering model.
[0083] Specifically, a correctness threshold is set in advance, and the recognition accuracy is compared with the correctness threshold. If the recognition accuracy is equal to or greater than the preset correctness threshold, it means that the current model can meet the requirements of the training objective, and the model training is completed. The current question-answering model is obtained, and training for the next training objective can continue.
[0084] If the recognition accuracy is less than the preset accuracy threshold, the model can also be trained for the current training target. A loss function is preset, and the preset loss function is updated according to the recognition accuracy, and the updated loss function is determined as the objective function. For example, the loss function is a DPO function, and the parameters in the DPO loss function can be automatically adjusted and updated according to the recognition accuracy to obtain the objective function. According to the objective function, the loss value of the model is calculated, and the question-answering model to be trained is back-propagated according to the loss value to obtain the current question-answering model. In this embodiment, the back-propagation process is not specifically limited.
[0085] The beneficial effect of this setting is that for different training targets, the function parameters can be automatically adjusted according to the recognition accuracy corresponding to the training target. This adaptive adjustment mechanism can improve the performance of the model on certain training targets without affecting the capabilities of other targets, thereby improving the model effect.
[0086] In this embodiment, the preset loss function is updated according to the recognition accuracy to obtain the objective function, including: determining the parameter value of the hyperparameter in the preset loss function according to the recognition accuracy; wherein the hyperparameter is used to balance the difference between the question-answering model to be trained and the current question-answering model; and determining the objective function according to the parameter value of the hyperparameter.
[0087] Specifically, there is a hyperparameter β in the DPO loss function. According to the recognition accuracy, the β value in the DPO loss function can be adjusted, and the loss function after the adjusted β value is determined as the objective function.
[0088] The β in the DPO loss function is a hyperparameter that controls the influence of the KL (Kullback Leibler) divergence and is used to balance the difference between the current model and the reference model. A smaller β value makes the model deviate more from the reference model during training, and actively learns and updates. A larger β value encourages the model to refer to the original model more conservatively during training to ensure stability.
[0089] The parameter value range of β can be preset, for example, the parameter value range is 0.1 to 0.5. According to the recognition accuracy, the parameter value of β can be adjusted within the parameter value range. For example, the recognition accuracy can be used as a new β value, or the recognition accuracy can be calculated by other methods such as normal distribution or normalization to convert the recognition accuracy into a new β value. In the case of a low recognition accuracy, it means that the question-answering model to be trained has identified more errors, and a lower β value can be used to achieve strong intervention, thereby improving the model effect.
[0090] The beneficial effect of this setting is that, through the recognition accuracy, targeted parameter adjustments can be made for different training objectives without maintaining a static setting of the β value, thereby intervening and improving the performance of the model on certain training objectives, increasing the flexibility of model training, and improving the adaptability and performance of the model.
[0091] In this embodiment, the objective function is determined according to the parameter value of the hyperparameter, including: in response to the recognition result of the input data to be trained representing that the expected output data and the predicted output data are consistent, the input data to be trained is determined as a positive sample, and in response to the recognition result of the input data to be trained representing that the expected output data and the predicted output data are inconsistent, the input data to be trained is determined as a negative sample; based on the positive samples and negative samples in the data set to be trained, a plurality of sample pairs are obtained; wherein each sample pair includes a positive sample and a negative sample; and based on the parameter value of the hyperparameter and the plurality of sample pairs, the objective function is obtained.
[0092] Specifically, positive samples and negative samples are determined from the data set to be trained. Positive samples refer to input data to be trained whose expected output data and predicted output data are consistent, and negative samples refer to input data to be trained whose expected output data and predicted output data are inconsistent. Based on the positive samples and negative samples in the data set to be trained, multiple sample pairs can be generated. Each sample pair includes a positive sample and a negative sample. The positive samples and negative samples in the sample pair can be in the form of triplets, and each triplet can include one data to be trained, one predicted output data, and one recognition result. For example, the input data to be trained is represented as query, the predicted output data is represented as answer, and the recognition result is true or false. The positive sample in the sample pair can be represented as [query, answer, true], and the negative sample can be represented as [query, answer, false]. Positive samples in different sample pairs may be repeated.
[0093] The DPO loss function can include β hyperparameters and positive and negative samples. The parameter value of β and sample pairs are added to the DPO loss function to obtain the objective function. The DPO loss function can be expressed as:
[0094]
[0095] Among them, L DPO Characterizes the loss value of the DPO loss function; σ is the sigmoid function; β is the hyperparameter; x is the input data to be trained; y w is the predicted output data in the positive sample data; y l is the predicted output data in the negative sample data; π θ (y w |x) represents the cumulative probability of a good answer generated by the current model given the input x; π ref (y w |x) represents the cumulative probability of a good answer generated by the reference model, i.e. the question-answering model to be trained, given the input x; π θ (y l |x) represents the cumulative probability of a bad answer generated by the current model given an input x; π ref (y l |x) represents the cumulative probability of a bad answer generated by the reference model given an input x.
[0096] The beneficial effect of this setting is that for each training target, different objective functions can be generated during model training, achieving targeted optimization of multiple targets and improving model performance.
[0097] In this embodiment, the question-answering model to be trained is trained according to the currently traversed prompt word information and the preset data set to be trained to obtain the current question-answering model, including: in response to the first prompt word information in the currently traversed preset prompt word data set, the first prompt word information and the data set to be trained are input into the preset initial model for training to obtain the current question-answering model.
[0098] Specifically, an initial model is pre-set as the question-answering model to be trained for the first training target. That is, when the first prompt word information in the prompt word data set is traversed, the initial model is used for training. The data set to be trained can be input into the initial model, and based on the first prompt word information, the predicted output data corresponding to each input data to be trained is obtained. According to the expected output data corresponding to each input data to be trained and the corresponding predicted output data, the recognition accuracy of the initial model is obtained. According to the recognition accuracy, the initial model is trained to obtain the current question-answering model corresponding to the first prompt word information. Then continue to traverse and perform model training for subsequent training targets.
[0099] The beneficial effect of this setting is that, by presetting an initial model, it is ensured that when the first prompt word information is traversed, the model optimization can be performed on the first training target, so that subsequent optimization can be performed while meeting the first training target, thereby achieving phased optimization of different training targets and improving model performance.
[0100] S303: In response to determining that the traversal of the preset prompt word data set is completed, determine that the current question-answering model is a trained question-answering model.
[0101] Exemplarily, this step may refer to the above-mentioned step S103 and will not be described in detail.
[0102] In this embodiment, the method further includes: in response to the preset prompt word data set not being traversed completely, determining the current question-answering model as the question-answering model to be trained for the next traversed prompt word information.
[0103] Specifically, each time a prompt word information is traversed, the model is trained for the training target. When the training of the training target is completed, it is determined whether the prompt word data set has been traversed. If not, the next prompt word information is traversed, and the current question-answering model obtained from the previous training is used as the question-answering model to be trained for the newly traversed prompt word information. The new prompt word information is continuously traversed until the prompt word data set is traversed, and the current question-answering model of the last prompt word information is obtained.
[0104] The beneficial effect of this setting is that the current question-answering model for each prompt word information can be used as the question-answering model to be trained for the next prompt word information, thereby training each training objective in stages, allowing the model to gradually meet each training objective, achieving multi-objective optimization of the model, and improving the performance and adaptability of the model in complex environments.
[0105] In this embodiment, it also includes: obtaining model prompt words of the trained question-answering model according to the prompt word information corresponding to each training target; wherein the model prompt words represent the prompt word information corresponding to all training targets.
[0106] Specifically, each training target corresponds to a prompt word information. During the model training process, the prompt word information can be adjusted at any time. For example, the text content of the prompt word information can be changed, thereby speeding up the model training process.
[0107] After all training targets are trained, the prompt word information corresponding to each training target is obtained. Combined with the content of these prompt word information, a total prompt word is obtained as the model prompt word of the question-answering model. The model prompt word can characterize all training targets, that is, the question-answering model can output output data that meets different training targets based on the model prompt word. For example, there are three training targets, and the prompt word information is "You are an intelligent agent with A ability", "You are an intelligent agent with B ability", and "You are an intelligent agent with C ability". Based on these three prompt words, the model prompt word obtained can be "You are an intelligent agent with A, B, and C abilities".
[0108] The beneficial effect of this setting is that the overall prompt word of the question-answering model can be obtained based on the prompt word information of each training target, so that the question-answering model can obtain output data that meets different requirements based on a total prompt word, thereby improving the model's capabilities and enhancing the user experience.
[0109] In the disclosed embodiment, a separate prompt word information is preset for each training target of the model to obtain a prompt word data set. The prompt word information can be used to perform role-playing for the training targets, thereby achieving role differentiation for different training targets. The prompt word data set is traversed, and model training is performed for each training target based on the prompt word information traversed each time, so that the question-answering model completed each time can meet the requirements of all previous training targets. That is, model training is performed for each training target in turn, and the model trained for the previous training target is used as a reference model for the next training target, thereby achieving phased training. When the last prompt word information is traversed, the final question-answering model is obtained. In the disclosed embodiment, through role-playing and phased training, it is possible to better adapt to different task requirements and goals, improve the adaptability and flexibility of model training, and enhance user satisfaction.
[0110] Figure 4 FIG. 1 is a flow chart of a question-answering method based on a large model provided according to an embodiment of the present disclosure. The method can be executed by a question-answering device based on a large model. Figure 4 As shown, the method comprises the following steps:
[0111] S401: receiving question information input by a user.
[0112] Exemplarily, question-answering models can be applied in various fields. For example, in the field of natural language processing, question-answering models can be used in tasks such as generating multi-style texts or adapting to multiple language styles. In the field of autonomous driving and intelligent control, question-answering models can be used in complex systems that need to optimize safety, efficiency, and user comfort at the same time. In the field of personalized recommendation systems, question-answering models can improve the personalization, accuracy, and user satisfaction of recommendation systems through multi-objective optimization. In the field of robotics and intelligent agents, question-answering models can be applied in intelligent agents that need to adapt to multiple tasks and environments.
[0113] The user can input question information to the question-answering model according to the needs, for example, the user can input question information in the form of text or voice, etc. Through the question-answering model, the user's question information can be received in real time.
[0114] S402: Input the question information into the question-answering model, and obtain reply information corresponding to the question information based on the model prompt words; wherein the model prompt words are used to guide the question-answering model to generate reply information.
[0115] For example, the model category of the question-answering model is a large model, and the large model may correspond to a model prompt word, and the model prompt word may represent multiple training targets, and the question-answering model may be guided to generate reply information through the model prompt word. The question information is input into the question-answering model, and based on the model prompt word, the capabilities required by the user may be determined, that is, in the absence of a classifier, it may be determined what goals the user hopes the reply information output by the model can meet, thereby obtaining the reply information corresponding to the capability.
[0116] When the model is being trained, it calculates each input data to be trained for different training goals, so as to learn what goals different input data correspond to. When the model is actually used, it can determine what goal the user's question information belongs to, and then obtain the reply information that meets the goal. For example, in the intelligent question-and-answer scenario, for questions that require reasoning, you can call the prompt words of reasoning ability to provide more accurate answers.
[0117] In the disclosed embodiment, by adopting a question-answering model and model prompt words, reply information that meets different goals can be output according to user needs, thereby improving the accuracy and efficiency of determining the reply information and enhancing the user's human-computer interaction experience.
[0118] Figure 5 A structural block diagram of a training device for a question-answering model provided in an embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Figure 5 The training device 500 of the question-answering model includes: an acquisition unit 501, a training unit 502 and a determination unit 503.
[0119] The acquisition unit 501 is used to acquire a preset prompt word data set; wherein the preset prompt word data set includes a plurality of prompt word information, the prompt word information corresponds to a training target one by one, and the training target represents the requirement for the output data of the model;
[0120] The training unit 502 is used to train the question-answering model to be trained by traversing the preset prompt word data set at least according to the prompt word information currently traversed, so as to obtain the current question-answering model; wherein the question-answering model meets the training target corresponding to the prompt word information that has been traversed;
[0121] The determination unit 503 is used to determine that the current question-answering model is a trained question-answering model in response to determining that the traversal of the preset prompt word data set is completed.
[0122] Figure 6 A structural block diagram of a training device for a question-answering model provided in an embodiment of the present disclosure, such as Figure 6 As shown, the training device 600 of the question-answering model includes an acquisition unit 601, a training unit 602 and a determination unit 603, wherein the training unit 602 includes a data set acquisition module 6021 and a model training module 6022.
[0123] The data set acquisition module 6021 is used to acquire a preset data set to be trained; wherein the preset data set to be trained includes input data to be trained and data labels, and the data labels represent expected output data corresponding to the input data to be trained;
[0124] The model training module 6022 is used to train the question-answering model to be trained according to the currently traversed prompt word information and the preset data set to be trained to obtain the current question-answering model.
[0125] In one example, the model training module 6022 includes:
[0126] The data prediction submodule is used to obtain the predicted output data corresponding to the input data to be trained according to the currently traversed prompt word information and the preset data set to be trained; wherein the predicted output data represents the data output by the question-answering model to be trained according to the input data to be trained;
[0127] The accuracy determination submodule is used to obtain the recognition accuracy of the question-answering model to be trained according to the expected output data and the corresponding predicted output data corresponding to the input data to be trained; wherein the recognition accuracy represents the probability that the question-answering model to be trained meets the training objectives;
[0128] The model training submodule is used to train the question-answering model to be trained according to the recognition accuracy to obtain the current question-answering model.
[0129] In one example, the accuracy determination submodule is specifically used for:
[0130] Determine the recognition result of the input data to be trained according to the expected output data and the corresponding predicted output data corresponding to the input data to be trained; wherein the recognition result indicates whether the expected output data and the predicted output data are consistent;
[0131] According to the recognition results of each input data to be trained, the recognition accuracy of the question-answering model to be trained is determined.
[0132] In one example, the accuracy determination submodule is specifically used for:
[0133] In response to the recognition result of the input data to be trained indicating that the expected output data and the predicted output data are consistent, determining the input data to be trained as a positive sample;
[0134] The recognition accuracy of the question-answering model to be trained is determined according to the total number of input data to be trained and the number of positive samples in the data set to be trained.
[0135] In one example, the model training submodule is specifically used to:
[0136] In response to the recognition accuracy being less than a preset accuracy threshold, updating a preset loss function according to the recognition accuracy to obtain an objective function;
[0137] According to the objective function, back-propagation training is performed on the question-answering model to be trained to obtain the current question-answering model.
[0138] In one example, the model training submodule is specifically used to:
[0139] Determine the parameter value of the hyperparameter in the preset loss function according to the recognition accuracy; wherein the hyperparameter is used to balance the difference between the question-answering model to be trained and the current question-answering model;
[0140] The objective function is determined according to the parameter value of the hyperparameter.
[0141] In one example, the model training submodule is specifically used to:
[0142] In response to the recognition result of the input data to be trained indicating that the expected output data and the predicted output data are consistent, the input data to be trained is determined as a positive sample; and in response to the recognition result of the input data to be trained indicating that the expected output data and the predicted output data are inconsistent, the input data to be trained is determined as a negative sample;
[0143] According to the positive samples and negative samples in the training data set, a plurality of sample pairs are obtained; wherein each sample pair includes a positive sample and a negative sample;
[0144] The objective function is obtained according to the parameter value of the hyperparameter and the multiple sample pairs.
[0145] In one example, the model training module 6022 includes:
[0146] The initial traversal submodule is used to respond to the first prompt word information in the preset prompt word data set currently traversed, input the first prompt word information and the data set to be trained into the preset initial model for training, and obtain the current question-answering model.
[0147] In one example, it also includes:
[0148] The traversal unit is used to determine the current question-answering model as the question-answering model to be trained for the next traversed prompt word information in response to the preset prompt word data set not being traversed completely.
[0149] In one example, it also includes:
[0150] The prompt word synthesis unit is used to obtain the model prompt words of the trained question-answering model according to the prompt word information corresponding to each training target; wherein the model prompt words represent the prompt word information corresponding to all training targets.
[0151] Figure 7 The structural block diagram of a large model-based question-answering device provided by the embodiment of the present disclosure. For the convenience of explanation, only the part related to the embodiment of the present disclosure is shown. Figure 7 The question-answering device 700 based on the large model includes: a receiving unit 701 and a replying unit 702.
[0152] The receiving unit 701 is used to receive question information input by the user;
[0153] The reply unit 702 is used to input the question information into the question-answering model, and obtain reply information corresponding to the question information based on the model prompt words; wherein the question-answering model represents the trained question-answering model described in the above embodiment, and the model prompt words are used to guide the question-answering model to generate reply information.
[0154] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device.
[0155] Figure 8 A structural block diagram of an electronic device provided in an embodiment of the present disclosure, such as Figure 8 As shown, the electronic device 800 includes: at least one processor 802; and a memory 801 communicatively connected to the at least one processor 802; wherein the memory stores instructions executable by the at least one processor 802, and the instructions are executed by the at least one processor 802 so that the at least one processor 802 can execute the training method of the question-answering model and the question-answering method based on a large model of the present disclosure.
[0156] The electronic device 800 further includes a receiver 803 and a transmitter 804. The receiver 803 is used to receive instructions and data sent by other devices, and the transmitter 804 is used to send instructions and data to external devices.
[0157] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0158] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which includes: a computer program, the computer program is stored in a readable storage medium, at least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes the solution provided by any of the above embodiments.
[0159] Fig. 9 A schematic block diagram of an example electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0160] like Fig. 9As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0161] A number of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0162] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the training method of the question-answering model and the question-answering method based on the large model. For example, in some embodiments, the training method of the question-answering model and the question-answering method based on the large model may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the training method of the question-answering model and the question-answering method based on the large model described above may be executed. Alternatively, in other embodiments, the computing unit 901 may be configured in any other appropriate manner (e.g., by means of firmware) to execute the question-answering model training method and the large-model-based question-answering method.
[0163] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0164] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0165] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0166] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0167] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0168] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0169] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0170] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A training method for a question-answering model, comprising: Acquire a preset prompt word data set; wherein the preset prompt word data set includes prompt word information, the prompt word information corresponds to a training target one by one, and the training target represents the requirement for the output data of the model; By traversing the preset prompt word data set, the question-answering model to be trained is trained at least according to the currently traversed prompt word information to obtain the current question-answering model; wherein the question-answering model meets the training target corresponding to the traversed prompt word information; In response to determining that the traversal of the preset prompt word data set is completed, determining that the current question-answering model is a trained question-answering model.
2. The method according to claim 1, wherein: The step of traversing the preset prompt word data set and training the question-answering model to be trained at least according to the currently traversed prompt word information to obtain the current question-answering model includes: According to the currently traversed prompt word information and the preset data set to be trained, the question-answering model to be trained is trained to obtain the current question-answering model; The preset data set to be trained includes input data to be trained and data labels of the input data to be trained, and the data labels represent expected output data corresponding to the input data to be trained.
3. The method according to claim 2, wherein: The step of training the question-answering model to be trained according to the currently traversed prompt word information and the preset data set to be trained to obtain the current question-answering model includes: According to the currently traversed prompt word information and the preset data set to be trained, the predicted output data corresponding to the input data to be trained is obtained; wherein the predicted output data represents the data output by the question-answering model to be trained according to the input data to be trained; According to the expected output data corresponding to the input data to be trained and the corresponding predicted output data, the recognition accuracy of the question-answering model to be trained is obtained; wherein the recognition accuracy represents the probability that the question-answering model to be trained meets the training objectives; The question-answering model to be trained is trained according to the recognition accuracy to obtain the current question-answering model.
4. The method according to claim 3, wherein: The obtaining the recognition accuracy of the question-answering model to be trained according to the expected output data and the corresponding predicted output data corresponding to the input data to be trained includes: Determine the recognition result of the input data to be trained according to the expected output data and the corresponding predicted output data corresponding to the input data to be trained; wherein the recognition result indicates whether the expected output data and the predicted output data are consistent; According to the recognition results of each input data to be trained, the recognition accuracy of the question-answering model to be trained is determined.
5. The method according to claim 4, wherein: Determining the recognition accuracy of the question-answering model to be trained according to the recognition results of each input data to be trained includes: In response to the recognition result of the input data to be trained indicating that the expected output data and the predicted output data are consistent, determining the input data to be trained as a positive sample; The recognition accuracy of the question-answering model to be trained is determined according to the total number of input data to be trained and the number of positive samples in the data set to be trained.
6. The method according to any one of claims 3 to 5, wherein: The step of training the question-answering model to be trained according to the recognition accuracy to obtain the current question-answering model includes: In response to the recognition accuracy being less than a preset accuracy threshold, updating a preset loss function according to the recognition accuracy to obtain an objective function; According to the objective function, back-propagation training is performed on the question-answering model to be trained to obtain the current question-answering model.
7. The method according to claim 6, wherein: The step of updating the preset loss function according to the recognition accuracy to obtain the objective function includes: Determine the parameter value of the hyperparameter in the preset loss function according to the recognition accuracy; wherein the hyperparameter is used to balance the difference between the question-answering model to be trained and the current question-answering model; The objective function is determined according to the parameter value of the hyperparameter.
8. The method according to claim 7, wherein: Determining the objective function according to the parameter value of the hyperparameter includes: In response to the recognition result of the input data to be trained indicating that the expected output data and the predicted output data are consistent, the input data to be trained is determined as a positive sample; and in response to the recognition result of the input data to be trained indicating that the expected output data and the predicted output data are inconsistent, the input data to be trained is determined as a negative sample; According to the positive samples and negative samples in the training data set, a plurality of sample pairs are obtained; wherein each sample pair includes a positive sample and a negative sample; The objective function is obtained according to the parameter value of the hyperparameter and the multiple sample pairs.
9. The method according to any one of claims 2 to 8, wherein: The step of training the question-answering model to be trained according to the currently traversed prompt word information and the preset data set to be trained to obtain the current question-answering model includes: In response to the first prompt word information in the preset prompt word data set currently traversed, the first prompt word information and the data set to be trained are input into a preset initial model for training to obtain the current question-answering model.
10. The method according to any one of claims 1 to 9, further comprising: In response to the preset prompt word data set not being traversed completely, the current question-answering model is determined as the question-answering model to be trained for the next traversed prompt word information.
11. The method according to any one of claims 1 to 10, further comprising: According to the prompt word information corresponding to each training target, the model prompt words of the trained question-answering model are obtained; wherein the model prompt words represent the prompt word information corresponding to all training targets.
12. A question answering method based on a large model, comprising: Receive question information input by the user; The question information is input into the question-answering model, and based on the model prompt words, the reply information corresponding to the question information is obtained; wherein the question-answering model representation weights are any one of the trained question-answering models 1 to 11, and the model prompt words are used to guide the question-answering model to generate reply information.
13. A training device for a question-answering model, comprising: An acquisition unit, configured to acquire a preset prompt word data set; wherein the preset prompt word data set includes prompt word information, the prompt word information corresponds to a training target one by one, and the training target represents a requirement for output data of the model; A training unit, configured to train the question-answering model to be trained by traversing the preset prompt word data set, at least according to the prompt word information currently traversed, to obtain a current question-answering model; wherein the question-answering model satisfies the training goal corresponding to the prompt word information that has been traversed; A determination unit is used to determine that the current question-answering model is a trained question-answering model in response to determining that the traversal of the preset prompt word data set is completed.
14. A question-answering device based on a large model, comprising: A receiving unit, used for receiving question information input by a user; A reply unit is used to input the question information into the question-answering model, and obtain reply information corresponding to the question information based on the model prompt words; wherein the question-answering model represents the trained question-answering model described in right 13, and the model prompt words are used to guide the question-answering model to generate reply information.
15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.
17. A computer program product, wherein: The invention comprises a computer program, which implements the steps of the method according to any one of claims 1 to 12 when being executed by a processor.
Citation Information
Cited By
Prompt word reconstruction model training and retrieval enhancement generation method, device and equipment
CN120723861A
Prompt word reconstruction model training, retrieval enhancement generation method, device and equipment
CN120723861B