Method, device, medium and program product for training question answering system
By introducing a new architecture of variable language model, reward model and policy gradient update module, we solve the problem of answering in intelligent question-and-answer services that does not meet human preferences and lack targetedness, and generate diverse answers that are more in line with user needs, improving the user experience.
Patent Information
- Application Number
- CN202410055105.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-07-18
AI Technical Summary
The answers generated by existing smart Q&A services may not meet human preferences and values, and are not targeted, resulting in too general and boring answers.
Using a new architecture based on reinforcement learning and variational inference, including a variational language model, reward model and policy gradient update module, we can generate targeted answers that meet human preferences and values by generating multiple answers and evaluating and updating model parameters using human feedback mechanisms.
The user experience of the intelligent question-and-answer system has been improved, and the generated answers are more in line with human preferences and value, while having higher targeting and diversity.
Smart Images

Figure CN120336450A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to a method, an electronic device, a medium, and a computer program product for training a question-answering system. Background Art
[0002] With the development of computer technology, intelligent question-answering services such as chatbots and virtual assistants have become more and more intelligent. However, there are two problems with existing intelligent question-answering services. On the one hand, answers that do not conform to human preferences and values may be generated. On the other hand, the answers generated by intelligent question-answering services trained based on language models often appear too general and boring, lacking pertinence to the questions.
[0003] In order to improve the user experience of intelligent question-answering services, there is an urgent need for a method that can both generate answers that meet human preferences and values and, on the other hand, make the answers more targeted to meet different user needs. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method, an electronic device, a medium, and a computer program product for training a question-answering system. The technical solution provided by the present disclosure can generate answers that not only conform to human values and preferences but also are more targeted to the questions. Thereby, the user experience of the intelligent question-answering system is improved.
[0005] In a first aspect of the present disclosure, a method for training a question-answering system is provided. The method includes determining a distribution of latent variables in a variational language model based on queries in a training dataset. The method further includes generating, using the variational language model, multiple answers to the queries based on multiple latent variables randomly sampled from the distribution. The method further includes determining, using a reward model, reward scores for the multiple answers. The method further includes updating the variational language model based on the query and the best answer among the multiple answers having the highest reward score.
[0006] In a second aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory coupled to the at least one processor and having instructions stored thereon that, when executed by the at least one processor, cause the electronic device to implement the method according to the first aspect of the present disclosure.
[0007] In a third aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, having machine-executable instructions stored thereon that, when executed by a machine, cause the machine to implement the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect of the present disclosure, there is provided a computer program product tangibly stored on a non-transitory computer-readable medium and including machine-executable instructions which, when executed, cause a machine to implement the method according to the first aspect of the present disclosure.
[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent, wherein in the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0011] Figure 1 FIG. shows a schematic block diagram of an environment for training a question-answering system according to an embodiment of the present disclosure;
[0012] Figure 2 FIG. shows a schematic flowchart of a method for updating a variational language model according to an embodiment of the present disclosure;
[0013] Figure 3 FIG. shows a schematic flowchart of a method for generating multiple answers using a variational language model according to an embodiment of the present disclosure;
[0014] Figure 4 FIG. shows a schematic flowchart of a method for updating model parameters according to an embodiment of the present disclosure; and
[0015] Figure 5 FIG. shows a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not used to limit the scope of protection of the present disclosure.
[0017] In the description of the embodiments of the present disclosure, the term "including" and its like terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0018] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.
[0019] A "neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and it generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. The neural networks used in deep learning applications usually include many hidden layers, thus increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.
[0020] Generally, machine learning can roughly include three stages, namely, the training stage, the testing stage, and the usage stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously iteratively updated until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the mapping from input to output) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, so as to determine the performance of the model. In some implementations, the testing stage can be omitted. In the usage stage, the model can be used to process the actual input based on the parameter values obtained from training and determine the corresponding output.
[0021] As mentioned above, there are two problems with intelligent question-and-answer services. On the one hand, answers that do not conform to human preferences and values may be generated. On the other hand, the answers generated by intelligent question-and-answer services based on language model training often appear too general and boring, and lack pertinence. Since existing intelligent question-and-answer services are generally trained based on language models, and language models cannot obtain features such as user preferences and value choices, the generated answers may not conform to human values and preferences. In addition, existing language models cannot generate diverse answers and select the most appropriate answers to output to users, resulting in low quality answers. For example, when a user initiates a query and asks the intelligent question-and-answer service "Can technology A be applied to industry B?", the answers generated by the intelligent question-and-answer service may be "What is technology A specifically?" and "What is industry B specifically?". It is unable to give a clear answer to the user, and the answer content lacks pertinence.
[0022] At least to solve the above and other potential problems, an embodiment of the present disclosure provides a method for training a question-answering system. In this method, a new architecture based on reinforcement learning and variational inference is proposed, which includes three parts. The first part is a variational language model, through which a diverse answer is generated. The second part is a reward model, through which a human feedback mechanism is introduced to achieve quality assessment of different answers generated by the variational language model and select and output the answer with the highest reward score. The third part is a policy gradient update module, which implements the update of the variational language model parameters based on maximizing the reward score of the answer. In this way, the answer output by the question-answering system can be made consistent with human values and preferences. At the same time, the generated answer is more targeted, which can provide users with more helpful answers and improve the user experience of the intelligent question-answering service. The embodiments of the present disclosure will be further described in detail in conjunction with the accompanying drawings.
[0023] Figure 1 FIG. 1 is a schematic block diagram of an environment for training a question answering system according to an embodiment of the present disclosure. Figure 1As shown, the environment 100 includes a training data set 104 and a question answering system 102 required for training the variational language model 106. In the question answering system 102, there are also three parts, namely, the first part is the variational language model 106, the second part is the reward model 108, and the third part is the model update module 110. In some embodiments, the training data set 104 can be the Persona-Chat data set or the CoQA data set. First, the variational language model 106 is pre-trained using the training data set 104. Secondly, after the pre-training is completed, multiple answers output by the variational language model 106 are obtained based on new inputs. In order to evaluate the quality of the answers generated by the variational language model 106, the reward model 108 is used to generate a reward value or a reward score for each answer. In some embodiments, the generated reward score is a scalar. The answer with the highest quality is selected according to the level of the reward score. Finally, according to the reward score situation, the model update module 110 is used to update the parameters of the variational language model 106.
[0024] In some embodiments, the variational language model 106 can generate multiple answers based on a user query and a latent variable. In statistics, a latent variable is opposite to an observed variable and refers to an unobservable random variable that can be inferred from the observed data based on a mathematical model. By using the latent variable, the uncertainty and variability in the answer generation process can be measured. In some embodiments, the uncertainty and variability can be the tone, style of the user query, and the content of the answer. In some embodiments, the variational language model 106 can infer the tone, style of the user query, and the content of the answer from the text input by the user.
[0025] In some embodiments, the variational language model 106 can be a conditional variational autoencoder. The conditional variational autoencoder includes two parts, the first part is the encoder part, and the second part is the decoder part. Through the encoder, the latent variable of the query can be obtained based on the user query. After obtaining the latent variable, through the decoder part, the answer to the query can be obtained based on the user query and the latent variable. Specifically, first, the encoder in the variational autoencoder will convert the distribution of all feature information carried by the input user query into a Gaussian distribution, and this Gaussian distribution can be regarded as the latent variable. Then, the latent variable, the user query, and the generated answer characters are input into the decoder to obtain the answer to the query.
[0026] In some embodiments, the reward model 108 can be a neural network model based on human feedback. The human feedback can be explicit or implicitly indicated. For example, explicit human feedback can be that humans directly rate or score the answers generated by the variational language model 106. Implicitly indicated human feedback can be indirect signals reflecting human preferences and values, such as user click-through rate, dwell time, or sentiment analysis. In some embodiments, when the user query and the corresponding generated answer are input into the reward model 108, the reward model 108 can output a reward score that measures the quality of the answer.
[0027] In some embodiments, the model update module 110 can be an evaluation algorithm. This algorithm can optimize the parameters of the variational language model 106 based on the user query, the corresponding answer, and the reward score, so that the answer output by the question-answering system 102 meets the user's needs.
[0028] Figure 2 A flowchart showing a method for updating a variational language model according to an embodiment of the present disclosure is shown. In some embodiments, the method 200 can be implemented by, for example, Figure 1 the question-answering system 102 shown. It should be understood that the method 200 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this regard.
[0029] As Figure 2 shown, at block 202, based on the queries in the training dataset, determine the distribution of the latent variables in the variational language model 106. In some embodiments, it can be assumed that the latent variables are a multivariate Gaussian distribution. Based on the given queries in the training dataset, use the variational language model 106 to generate the mean and variance of the queries, and determine the probability distribution of the latent variables (i.e., a multivariate Gaussian distribution) based on the mean and variance of the above queries.
[0030] At block 204, based on a plurality of latent variables randomly sampled from the distribution, use the variational language model to generate a plurality of answers for the query. In some embodiments, given a query, a value is randomly sampled from the latent variable distribution. Based on this value, use the variational language model 106 to generate an answer for the query. The above steps are implemented multiple times to obtain a plurality of randomly sampled latent variables, and a plurality of answers are generated based on the plurality of randomly sampled latent variables.
[0031] At block 206, use the reward model 108 to determine the reward scores for the plurality of answers. In some embodiments, the reward model evaluates the plurality of answers based on human feedback. The evaluation can be a rating or a score, or can be based on indirect evaluation metrics such as click-through rate and dwell time, which are not limited herein.
[0032] At block 208, the variational language model 106 is updated based on the query and the best response among multiple responses having the highest reward score. In some embodiments, the variational language model 106 is updated based on a given query, a response corresponding to the given query, and the reward score of the response.
[0033] Figure 3 A flow schematic diagram of a method for generating multiple responses using a variational language model according to an embodiment of the present disclosure is shown. Figure 3 These are the specific implementation steps of block 204, which will be described in detail below in conjunction with Figure 3 the process of generating multiple responses based on a given query.
[0034] As Figure 3 shown, the method flow 300 for generating multiple responses is Figure 2 a specific implementation of block 204 in, and this method can be executed on, for example, Figure 1 the variational language model 106 in or any suitable computing device or server.
[0035] To generate multiple responses, a variational language model needs to be pre-trained first. Specifically, at block 302, a joint distribution is defined:
[0036] p θ (a, z|q) = p θ (a|z, q)p θ (z|q) (1)
[0037] where q is a user query, a is a response generated for the query, z is a latent variable, and θ are the model parameters of the variational language model. This joint distribution formula can be regarded as a variant of Bayesian estimation. Based on this formula, the probability distribution of the latent variable z can be obtained given q. Subsequently, given z and q, a for the query can be obtained. Then, a joint distribution can be obtained based on z, q, and a.
[0038] At block 304, the probability distribution of the latent variable z is defined as a prior distribution. In some embodiments, the latent variable z can be considered a multivariate Gaussian distribution, which can be a matrix with diagonal covariance:
[0039]
[0040] where μ θ (q) and σ θ (q) are the mean and variance of the multivariate Gaussian distribution, respectively. In some embodiments, the given q is input into the encoder of the variational language model to obtain the mean and variance of the multivariate Gaussian distribution, respectively.
[0041] At block 306, a conditional distribution is defined, and an answer is generated based on this conditional distribution. Specifically, based on the generated answer characters, latent variables, and the given query, the current answer character is generated.
[0042]
[0043] Where T is the character length of the answer. For example, if the answer content is "A technology can be applied to the B industry", then the length of T is 11. at refers to the current answer character. In some embodiments, the decoder in the variational language model can generate the current answer character based on the generated answer characters, latent variable z, and the given query q, and then, based on the generated answer characters, form a complete answer. In some embodiments, the decoder in the variational language model can be implemented by any form of autoregressive model. For example, it can be a recurrent neural network (RNN) or a Transformer neural network.
[0044] At block 308, the variational language model is pre-trained. In some embodiments, the variational language model can be trained using the method of maximizing the evidence lower bound (ELBO):
[0045]
[0046] Where φ is the model parameter of the inference network, and q φ (z|q,a) is a probability distribution introduced based on the method of variational inference, used to approximately estimate the intractable posterior probability p θ (z|q,a), that is:
[0047]
[0048] Where the mean μ φ (q,a) and variance σ φ (q,a) are generated by using the given query q and the generated answer a as inputs and inputting them into the inference network. In some embodiments, the model parameters of the inference network can be the same as the model parameters of the variational language model, or the best model parameters can be obtained by pre-training the inference network.
[0049] Returning to the formula (4) for maximizing the evidence lower bound, the similarity between the two distributions of q φ (z|q,a) and p θ (z|q) is measured using the KL divergence. Maximizing the evidence lower bound is to maximize the log-likelihood of the generated answer. In other words, based on q φ (z|q,a), maximize the expected value of log p θ (a|z,q). The variational language model is pre-trained by using the training method of maximizing the evidence lower bound.
[0050] After the pre-training is completed, backpropagation is performed on the variational language model to update the model weights. At block 310, the following reparameterization method is used to backpropagate and update the model weights.
[0051] z = μ υ (q, a) + σ φ (q, a) ⊙ ∈(6)
[0052] where ∈ ∼ N(0, I) is a random noise vector, and ⊙ represents the element-wise product of two matrices. By using this method, the latent variable can be transformed into a differentiable form, and then the model weights are updated by backpropagation using the chain rule.
[0053] After the above pre-training process is completed, at block 312, the variational language model after the pre-training is used to generate multiple answers. For a given query q, a value is first randomly sampled from the latent variable distribution.
[0054]
[0055] Specifically, a latent variable is randomly sampled from the latent variable distribution (i.e., a multivariate Gaussian distribution), and this latent variable is used as the input to generate the current answer character.
[0056]
[0057] Specifically, the current answer character is generated based on the generated answer characters, a value randomly sampled from the latent variable distribution, and the given query q. Then, a complete answer is obtained based on the generated answer characters. In some embodiments, diverse answers can be generated based on beam search with a diversity penalty function. In some embodiments, a value can continue to be sampled from the latent variable distribution, and the current answer character is generated based on this value, the given query q, and the generated answer characters, and then a new answer is generated. Subsequently, beam search with a diversity penalty function is used to generate diverse answers. By continuously implementing the above method, multiple different answers with diversity can be generated.
[0058] In some embodiments, a reward model can be used to select the answer with the highest quality. The reward model can be regarded as a classifier, that is, taking the given query and the generated answer as the model input, and outputting a reward score to measure the quality of the answer.
[0059] Defining the query as q and the answer as a, the reward model can map q and a to a scalar:
[0060] r ψ(a|q) = f ψ (q, a) (9)
[0061] where ψ are the model parameters of the reward model, and f ψ can be a neural network, for example, it can be a multi-layer perceptron (MLP) or a Transformer network. This neural network can use human feedback data for supervised learning. In some embodiments, the human feedback data can be collected online or offline, which is not limited herein. The data collected online can be the data that users feedback to the system in real time when using the question and answer system online. The data collected offline can be the data that human annotators score the collected question and answer data in advance. In some embodiments, the reward model can be self-supervised trained using data that indirectly indicates human scores, and the data can be, for example, the click-through rate, dwell time, or sentiment analysis of users.
[0062] After the training of the reward model is completed, the reward model can score multiple answers generated by the variational language model and select the answer with the highest reward score.
[0063]
[0064] In some embodiments, the best answer a selected by the reward model * can be used in the model update module to update the parameters of the variational language model, which will be elaborated below in combination with Figure 4 for specific elaboration.
[0065] Figure 4 shows a flowchart of a method for updating model parameters according to an embodiment of the present disclosure. Method 400 is the specific implementation steps of block 208. The following will be combined with Figure 4 to describe in detail the process of updating the variational language model based on the query and the best answer with the highest reward score among multiple answers.
[0066] In block 402, an optimization objective is defined. In some embodiments, the policy gradient algorithm is used to update the model parameters of the variational language model and the model parameters of the reward model. The objective of the policy gradient algorithm is to maximize the reward score output by the reward model, that is:
[0067]
[0068] where, after generating an answer a based on a given query q, the expectation of the reward score generated by the reward model is obtained. In block 404, based on formula (11), the model parameters are updated:
[0069]
[0070]
[0071] Among them, and are respectively the updated gradients of the model parameters. In block 406, estimate the updated gradients:
[0072]
[0073]
[0074] In some embodiments, the parameter updates of the variational language model and the reward model can be updated using the method of gradient ascent. The q in formulas (14) and (15) (n) ,a (n) can be a pair of query and corresponding answer randomly sampled from the variational language model.
[0075] In block 408, use a baseline function. Specifically, in order to reduce the variance of the gradient estimation, a baseline function can be used to estimate the expected value of a given query:
[0076]
[0077] In some embodiments, another different language model can be used to generate an answer a for a randomly given query q. In some embodiments, the variational language model can also be used to randomly generate an answer a for the given query q, which is not limited here. In some embodiments, the above given query q and the generated answer a can be used as the input of formula (4) to train the variational language model. In some embodiments, formula (16) can share model parameters with the reward model or obtain a model parameter through pre-training.
[0078] After obtaining the randomly given query q and the answer a obtained for this query, use the variational language model to generate another answer a. In some embodiments, the given query q can be the same as the query q when obtaining the baseline function value, or another query q can be given additionally. Based on this query q, obtain a new reward score.
[0079] In block 410, subtract the baseline function value to obtain the advantage function. Subtract the obtained baseline function value from the reward score to obtain the advantage function value, that is:
[0080] A ψ (a|q) = r ψ (a|q) - b ψ (q) (17)
[0081] Formula (17) is the advantage function, which measures the difference between the continuously generated answer and the reference answer, and uses this difference to evaluate the quality of the answer. In some embodiments, an advantage function value is obtained by subtracting the reference function value from the reward score of the answer generated by the variational language model. Substitute this advantage function value into Formula (11), and then update the model parameters of the variational language model and the model parameters of the reward model respectively based on the gradient update formula.
[0082] At block 412, the model parameters are updated using the advantage function. That is:
[0083]
[0084]
[0085] where γ in Formula (19) is the learning rate, and the last term is used to make the reference function and the reward function consistent.
[0086] The above reference Figures 1 to 4 describes exemplary embodiments of the present disclosure. Compared with existing solutions, the method for training a question-and-answer system by the user of the present disclosure can generate more diverse and targeted answers, avoiding the generation of general and boring answers, and improving the user interaction experience. On the other hand, the answers generated by the method of the present disclosure are more in line with human values and preferences, achieving "AI alignment". Specifically, the present disclosure proposes a brand-new architecture, including a variational language model, a reward model, and a policy gradient update module. By introducing latent variables, the variational language model can generate diverse answers. By introducing a human feedback mechanism, the reward model can screen out answers that meet human values and preferences, realizing the quality evaluation of different answers generated by the variational language model and selecting and outputting the answer with the highest reward score. The policy gradient update module updates the parameters of the variational language model based on maximizing the reward score of the answer, improving the performance and robustness of the model.
[0087] Figure 5 shows a block diagram of an electronic device 500 according to some embodiments of the present disclosure. As Figure 5As shown, device 500 includes a central processing unit (CPU) 502 and a graphics processing unit (GPU) 504, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 506 or computer program instructions loaded from a storage unit 518 into a random access memory (RAM) 508. In the RAM 508, various programs and data required for the operation of device 500 can also be stored. The CPU 502, GPU 504, ROM 506, and RAM 508 are connected to each other via a bus 510. An input / output (I / O) interface 512 is also connected to the bus 510. Although not shown in Figure 5 , device 500 may also include a coprocessor.
[0088] Multiple components in device 500 are connected to the I / O interface 512, including: an input unit 514, such as a keyboard, a mouse, etc.; an output unit 516, such as various types of displays, speakers, etc.; a storage unit 518, such as a magnetic disk, an optical disc, etc.; and a communication unit 520, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 520 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0089] Each of the methods or processes described above can be executed by the CPU 502 and the GPU 504. For example, in some embodiments, the method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 518. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 500 via the ROM 506 and / or the communication unit 520. When the computer program is loaded and executed by the CPU 502 and the GPU 504, one or more steps or actions in the methods or processes described above can be executed.
[0090] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0091] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0092] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0093] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, voice input, or tactile input).
[0094] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0095] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other.
[0096] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0097] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for training a question-answering system, the question-answering system including a variational language model, the method comprising: Determining a distribution of latent variables in the variational language model based on queries in a training dataset; Generating, using the variational language model, multiple answers to the query based on multiple latent variables randomly sampled from the distribution; Determining reward scores for the multiple answers using a reward model; And Updating the variational language model based on the query and the best answer having the highest reward score among the multiple answers.
2. The method according to claim 1, wherein The latent variables conform to a Gaussian distribution, and determining the distribution of latent variables in the variational language model includes: By inputting the query into the variational language model, obtaining the mean and variance of the Gaussian distribution from the variational language model.
3. The method according to claim 1 or 2, further comprising performing supervised pre-training on the variational language model, the pre-training comprising: Defining a prior distribution of the latent variables with respect to the query and a conditional distribution of the answer with respect to the query and the latent variables; Defining a joint distribution of the answer and the latent variables with respect to the query based on the prior distribution and the conditional distribution; Determining a posterior probability of the latent variables with respect to the query and the answer; And Training the variational language model by maximizing the evidence lower bound (ELBO), wherein the ELBO is calculated based on the joint distribution, the prior distribution, the conditional distribution, and the posterior probability.
4. The method according to claim 3, wherein the posterior probability conforms to a Gaussian distribution, and determining the posterior probability of the latent variables with respect to the query and the answer includes: By inputting the query and the answer into a variational inference network, determining the mean and variance of the Gaussian distribution of the posterior probability, wherein the variational inference network is included in the variational language model or is a separate network.
5. The method according to claim 1, wherein generating, using the variational language model, multiple answers to the query includes: Obtaining one latent variable by randomly sampling from the distribution; Based on the one latent variable, the answer characters already generated by the variational language model, and the query, using the variational language model to generate the current answer character to obtain one answer among the multiple answers.
6. The method according to claim 1, wherein generating, using the variational language model, multiple answers to the query includes: Using beam search with a diversity penalty function to generate the multiple answers to the query.
7. The method according to claim 1, the reward model being trained by reinforcement learning based on human feedback, wherein the human feedback measures the quality of the answer.
8. The method according to claim 1, wherein updating the variational language model comprises: Performing the following actions at least once: Updating the variational language model based on a first optimization objective, the first optimization objective being defined as the expectation of a reward function of the reward model for the query and the best answer; And Updating the variational language model based on a second optimization objective, the second optimization objective being defined as the expectation of the difference between the reward function and a baseline function.
9. The method according to claim 8, wherein the baseline function is used to estimate the expected reward score of the query, and the baseline function is implemented by sharing parameters with the reward model or a separate network model.
10. An electronic device, the electronic device comprising: at least one processor; and a memory coupled to the at least one processor and having instructions stored thereon, the instructions causing the electronic device to perform actions when executed by the at least one processor, the actions including: determining a distribution of latent variables in a variational language model of a question answering system based on queries in a training dataset; generating, based on a plurality of latent variables randomly sampled from the distribution, a plurality of answers to the query using the variational language model; determining reward scores for the plurality of answers using a reward model; and updating the variational language model based on the query and the best answer among the plurality of answers having the highest reward score.
11. The electronic device according to claim 10, wherein, The latent variables conform to a Gaussian distribution, and determining the distribution of the latent variables in the variational language model includes: inputting the query into the variational language model to obtain the mean and variance of the Gaussian distribution from the variational language model.
12. The electronic device according to claim 10 or 11, wherein the actions further include performing supervised pre-training on the variational language model, the pre-training including: defining a prior distribution of the latent variables with respect to the query and a conditional distribution of the answer with respect to the query and the latent variables; defining a joint distribution of the answer and the latent variables with respect to the query based on the prior distribution and the conditional distribution; determining a posterior probability of the latent variables with respect to the query and the answer; and and training the variational language model by maximizing the evidence lower bound (ELBO), wherein the ELBO is calculated based on the joint distribution, the prior distribution, the conditional distribution, and the posterior probability.
13. The electronic device according to claim 12, wherein the posterior probability conforms to a Gaussian distribution, and determining the posterior probability of the latent variables with respect to the query and the answer includes: inputting the query and the answer into a variational inference network to determine the mean and variance of the Gaussian distribution of the posterior probability, wherein the variational inference network is included in the variational language model or is a separate network.
14. The electronic device according to claim 10, wherein generating a plurality of answers to the query using the variational language model includes: obtaining one latent variable by randomly sampling from the distribution; generating a current answer character using the variational language model based on the one latent variable, the answer characters already generated by the variational language model, and the query to obtain one answer among the plurality of answers.
15. The electronic device according to claim 10, wherein generating a plurality of answers to the query using the variational language model includes: using beam search with a diversity penalty function to generate the plurality of answers to the query.
16. The electronic device according to claim 10, wherein the reward model is trained by reinforcement learning based on human feedback, and the human feedback measures the quality of the answer.
17. The electronic device according to claim 10, wherein updating the variational language model includes: Perform the following actions at least once: Update the variational language model based on a first optimization objective, the first optimization objective being defined as the expectation of the reward function of the reward model for the query and the best answer; And Update the variational language model based on a second optimization objective, the second optimization objective being defined as the expectation of the difference between the reward function and a baseline function.
18. The electronic device according to claim 17, wherein the baseline function is used to estimate the expected reward score of the query, and the baseline function is implemented by sharing parameters with the reward model or by a separate network model.
19. A non-transitory computer-readable storage medium having machine-executable instructions stored thereon, the machine-executable instructions, when executed by a machine, cause the machine to: Determine the distribution of latent variables in a variational language model of a question-and-answer system based on queries in a training dataset; Generate multiple answers to the query using the variational language model based on a plurality of latent variables randomly sampled from the distribution; Determine reward scores for the multiple answers using a reward model; And Update the variational language model based on the query and the best answer having the highest reward score among the multiple answers.
20. A computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including machine-executable instructions, the machine-executable instructions, when executed, cause the machine to: Determine the distribution of latent variables in a variational language model of a question-and-answer system based on queries in a training dataset; Generate multiple answers to the query using the variational language model based on a plurality of latent variables randomly sampled from the distribution; Determine reward scores for the multiple answers using a reward model; And Update the variational language model based on the query and the best answer having the highest reward score among the multiple answers.