A model training method, device and system

By building and training a dialogue model including generative models, inferred models and retrieval models, the insufficient performance of existing knowledge-based task dialogue systems under complex query conditions is solved, and more efficient knowledge fusion and query capabilities are achieved.

CN116910205BActive Publication Date: 2025-05-13RES INST OF CHINA MOBILE COMM GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310752628.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-25
Publication Date
2025-05-13
Estimated Expiration
2043-06-25

AI Technical Summary

Technical Problem

The existing knowledge-based task dialogue system has poor performance when query conditions are complex, making it difficult to effectively utilize the query capabilities of the database or knowledge base.

Method used

By building the dialogue model to be trained, including generative models, inferred models and retrieval models, supervised pre-training and semi-supervised training methods, the model's query ability and knowledge fusion ability are improved.

Benefits of technology

It improves the performance of the knowledge-based task dialogue system under complex query conditions, enhances the system's query capabilities for databases and knowledge bases, and achieves more efficient knowledge fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116910205B_ABST
    Figure CN116910205B_ABST
Patent Text Reader

Abstract

The present invention provides a model training method, device and system, the method comprising constructing a dialogue model to be trained, the dialogue model to be trained comprising a generation model, an inference model and a retrieval model; performing supervised pre-training on the inference model and the retrieval model in the dialogue model to be trained based on labeled first data, the pre-trained inference model is used to obtain latent variable data based on user input data and system response data in the first data, the pre-trained retrieval model is used to retrieve query results from a database based on the user input data in the first data; based on training samples formed by the first data and unlabeled second data, semi-supervised training is performed on the pre-trained dialogue model to obtain a trained dialogue model, the generation model in the trained dialogue model is used to generate dialogue action data and system response data. The ability of the model to combine knowledge in task-based dialogue tasks is improved, and it is more suitable for knowledge-based tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent dialogue technology, and in particular to a model training method, device and system. Background Art

[0002] Task-oriented Dialog System refers to a dialog system that has clear task settings in the process of interacting with users, such as booking a restaurant, booking a flight, etc. Knowledge-based dialogue is often based on databases, knowledge graphs or documents to improve the dialog system's ability to utilize knowledge. Therefore, knowledge-based dialogue is introduced into task dialogue, and a knowledge-based task dialogue system is constructed to improve the knowledge and reliability of system responses.

[0003] Currently, in knowledge-based dialogues, queries to databases or knowledge bases are generally completed simply through application programming interface (API) calls. However, the query conditions in task dialogues are usually more complex, which greatly limits the query capabilities of existing knowledge-based task dialogue systems for databases and knowledge bases.

[0004] It can be seen that in the existing technology, the knowledge-based task dialogue system has the problem of poor performance. Summary of the invention

[0005] The embodiments of the present invention provide a model training method, device and system to solve the problem of poor performance of knowledge-based task dialogue systems in the prior art.

[0006] To solve the above technical problems, the present invention is achieved as follows:

[0007] In a first aspect, an embodiment of the present invention provides a model training method, the method comprising:

[0008] Constructing a dialogue model to be trained, wherein the dialogue model to be trained includes a generation model, an inference model, and a retrieval model;

[0009] Based on the labeled first data, supervised pre-training is performed on the inference model and the retrieval model in the dialogue model to be trained, wherein the pre-trained inference model is used to obtain latent variable data according to the user input data and the system response data in the first data, and the pre-trained retrieval model is used to retrieve the query result from the database according to the user input data in the first data;

[0010] Based on the training samples formed by the first data and the unlabeled second data, semi-supervised training is performed on the pre-trained dialogue model to obtain a trained dialogue model, wherein the generation model in the trained dialogue model is used to generate dialogue action data and system response data.

[0011] Optionally, during the semi-supervised training of the pre-trained dialogue model, model parameters of the generation model are updated based on the training samples.

[0012] Optionally, in the process of performing semi-supervised training on the pre-trained dialogue model, when the first data is used as a training sample, the latent variable data is a first label spliced ​​according to the query result and dialogue action data in the first data.

[0013] Optionally, in the process of performing semi-supervised training on the pre-trained dialogue model, when the second data is used as a training sample, the latent variable data is sampled according to an iterative importance sampling method, and the sampling result is determined as the second label.

[0014] Optionally, the training sample includes T rounds of dialogue, the T rounds of dialogue include historical dialogue data from the 1st round to the T-1th round and dialogue data of the Tth round, each round of dialogue data includes user input data and corresponding system response data, and T is a positive integer greater than 1.

[0015] In a second aspect, an embodiment of the present invention provides a model training device, the device comprising:

[0016] A construction module, used to construct a dialogue model to be trained, wherein the dialogue model to be trained includes a generation model, an inference model, and a retrieval model;

[0017] A pre-training module, used to perform supervised pre-training on the inference model and the retrieval model in the dialogue model to be trained based on the labeled first data, wherein the pre-trained inference model is used to obtain latent variable data according to the user input data and the system response data in the first data, and the pre-trained retrieval model is used to retrieve the query result from the database according to the user input data in the first data;

[0018] A training module is used to perform semi-supervised training on a pre-trained dialogue model based on training samples formed by the first data and unlabeled second data to obtain a trained dialogue model, wherein the generation model in the trained dialogue model is used to generate dialogue action data and system response data.

[0019] Optionally, during the semi-supervised training of the pre-trained dialogue model, model parameters of the generation model are updated based on the training samples.

[0020] Optionally, in the process of performing semi-supervised training on the pre-trained dialogue model, when the first data is used as a training sample, the latent variable data is a first label spliced ​​according to the query result and dialogue action data in the first data.

[0021] Optionally, in the process of performing semi-supervised training on the pre-trained dialogue model, when the second data is used as a training sample, the latent variable data is sampled according to an iterative importance sampling method, and the sampling result is determined as the second label.

[0022] Optionally, the training sample includes T rounds of dialogue, the T rounds of dialogue include historical dialogue data from the 1st round to the T-1th round and dialogue data of the Tth round, each round of dialogue data includes user input data and corresponding system response data, and T is a positive integer greater than 1.

[0023] In a third aspect, an embodiment of the present invention provides a knowledge-based task dialogue system, the system comprising a dialogue model obtained according to the above-mentioned model training method.

[0024] In a fourth aspect, an embodiment of the present invention provides an electronic device, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program implements the steps of the above-mentioned model training method when executed by the processor.

[0025] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the model training method as described above are implemented.

[0026] In an embodiment of the present invention, based on the construction of a generation model and an inference model, knowledge-based dialogue modeling is introduced into task-oriented dialogues, and a retrieval enhancement method is applied to task-based dialogues. The knowledge integration capability of the model in task-based dialogue tasks is improved based on the retrieval model, making the trained dialogue model more suitable for knowledge-based tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0028] Figure 1 is a flow chart of a model training method provided by an embodiment of the present invention;

[0029] Figure 2 is a schematic diagram of the structure of a latent variable dialogue model provided by an embodiment of the present invention;

[0030] Figure 3 is a flow chart of a retrieval enhancement method used in a dialogue system provided by an embodiment of the present invention;

[0031] Figure 4 is a schematic diagram of the structure of a retrieval model provided by an embodiment of the present invention;

[0032] Figure 5 It is a schematic diagram of the structure of the generation model and the inference model provided by the embodiment of the present invention;

[0033] Figure 6 It is a structural schematic diagram of a model training device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0035] See also Figure 1 , Figure 1 is a flow chart of a model training method provided by an embodiment of the present invention, such as Figure 1 As shown, the method comprises the following steps:

[0036] Step 101: construct a dialogue model to be trained, wherein the dialogue model to be trained includes a generation model, an inference model, and a retrieval model;

[0037] In this embodiment, the knowledge-based task dialogue system is constructed based on the retrieval enhancement method. When constructing the dialogue model to be trained, the knowledge-based dialogue modeling is introduced into the task-oriented dialogue. Figure 2 The latent variable dialogue model shown is implemented by constructing a model with parameters θ and φ. The generated model can be recorded as p θ gen , the inference model can be recorded as q φ , the retrieval model can be recorded as p θ ret , by using the retrieval model to find relevant query results from the database (Knowledge Base, KB), rather than integrating knowledge information into the task-based dialogue system through database query methods. This improves the ability of the knowledge-based task dialogue system to integrate knowledge.

[0038] Step 102: Performing supervised pre-training on the inference model and the retrieval model in the dialogue model to be trained based on the labeled first data, wherein the pre-trained inference model is used to obtain latent variable data according to the user input data and the system response data in the first data, and the pre-trained retrieval model is used to retrieve the query result from the database according to the user input data in the first data;

[0039] First, the inference model q is used on the first labeled data φ and the retrieval model p θ ret Conduct supervised pre-training. Experimental results show that for the generation model p θ gen Pre-training will limit the model from achieving better performance. Therefore, in this embodiment, the generation model p is not used. θ gen Pre-training is performed to improve the performance of the knowledge-based task dialogue system. Then, semi-supervised training is performed based on the joint random approximation method through step 103.

[0040] Step 103: Based on the training samples formed by the first data and the unlabeled second data, semi-supervised training is performed on the pre-trained dialogue model to obtain a trained dialogue model, wherein the generation model in the trained dialogue model is used to generate dialogue action data and system response data.

[0041] Random sampling is performed from the labeled first data and the unlabeled second data to form labeled or unlabeled training samples, and semi-supervised training is performed in combination with the Joint Stochastic Approximation (JSA) method. When the query results generated based on the retrieval model have high recall, the query results are sent as conditional input to the generation model, and then the generation model is used to generate inferred dialogue action data and system response data.

[0042] In this implementation, based on the construction of generation models and inference models, knowledge-based dialogue modeling is introduced into task-oriented dialogues, and the retrieval enhancement method is applied to task-based dialogues. The knowledge integration capability of the model in task-based dialogue tasks is improved based on the retrieval model, making the trained dialogue model more suitable for knowledge-based tasks.

[0043] Among them, by constructing a model with parameter θ, a latent variable dialogue model is obtained, which includes a generation model p θ gen and the retrieval model p θ ret Take a T-round dialogue as an example, where T is a positive integer greater than 1. t Indicates user input data, a tIndicates dialogue action data, kb t Indicates that the query results are retrieved from the database, r t Indicates the system reply data. The subscript t indicates the tth round, and the value range of t is 1 to T. The database KB is composed of a series of slot-value pairs, and the query result kb t The retrieval model p θ ret Obtained from the database. Figure 3 As shown, for example, the user inputs data u t It could be "Can I find a cheaper package?" The retrieval model can find some knowledge fragments that are most relevant to the current conversation from the database / knowledge graph / document, and use these knowledge fragments as query results. t , kb t It can include "Name: Tianyi 4G package price 8 yuan total traffic 100M", and splice it with the context of the conversation to obtain the conversation action data a t , a t It can include "[notification][inquiry]", and use the generative model to generate and infer the system response data including dialogue action data and system response data. t , r t It could be "OK, I found a Tianyi 4G package for you, the price is 8 yuan, including 100M of traffic, are you satisfied?"

[0044] Among them, an entry sv in the database can be represented by a slot-value pair consisting of a slot name and a slot value. i ,i=1,2,…,N, where N is a positive integer, to improve the convenience of retrieval and align with the traditional task-based dialogue system. The retrieval model aims to find the knowledge required for the current round from the KB (containing N slot-value pairs). in:

[0045]

[0046] Assume that sv t i Does it appear in kb t are independent. In this way, we can transform the joint probability p θ (kb t |c t ,u t )(given database KB) is split into sv t i ,i=1,2,…,N pairs are in the KB:

[0047]

[0048] The probability p in the above formula is θ ret (kb t i |c t ,u t ) represents sv t i Appeared in kb t The probability in can be realized by a Bert model, such as Figure 4 As shown, c t ⊕u t ⊕sv i As input, let the model perform a binary classification task to determine whether it is a related knowledge fragment.

[0049] Specifically, after calculating the joint probability p θ ret (kb t |c t ,u t ), we can use the retrieval enhancement generation method (RAG) for model training. When training the model, there are kb t Annotation, based on p θ (kb t ,a t ,r t |c t ,u t )=p θ ret (kb t |c t ,u t ) θ gen (a t ,r t |c t ,u t ,kb t ) respectively optimize p θ ret (kb t |c t ,u t ) and p θ gen (a t ,r t |c t ,u t ,kb t ). In the test, kb t Unknown, based on the following formula:

[0050]

[0051] Perform a beam search to generate a t ,r t In practice, experiments show that based on the retrieval model p θ ret (kb t |c t ,u t ) generated kb t When the recall is high, the kb t As conditional input, it is fed into the generative model p θ gen (a t ,r t |c t ,u t ,kb t ) has a better result. Therefore, in practice, kbt p θ ret (kb t |c t ,u t ) θ gen (a t ,r t |c t ,u t ,kb t ) is to make an approximation to p θ ret (kb t |c t ,u t ) sets a uniform threshold for each item and infers kb t , and then use the generative model to generate and infer a t ,r t .

[0052] In an optional example, the training sample includes T rounds of dialogue, the T rounds of dialogue include historical dialogue data from the 1st round to the T-1th round and dialogue data from the Tth round, each round of dialogue data includes user input data and corresponding system response data, and T is a positive integer greater than 1. Define the historical dialogue data from the 1st round to the T-1th round as c t ={u1,r1,……,u t-1 ,,r t-1}, each round of dialogue data includes user input data u t And the corresponding system reply data r t .

[0053] Among them, the knowledge-based task dialogue system can be implemented by a model with a parameter θ, and the probability can be decomposed into each round:

[0054]

[0055] Based on the above modeling, the model can be implemented as follows: θ (h t ,r t |c t ,u t ) can be further instantiated using the following formula and decomposed into retrieval probability and generation probability:

[0056] p θ (h t ,r t |c t ,u t )

[0057] =p θ (kb t ,a t ,r t |c t ,u t )

[0058] =p θ ret (kb t |c t ,u t ) θ gen (a t ,r t |c t ,u t ,kb t )

[0059] Retrieval probability p θ ret (kb t |c t ,u t ), that is, the retrieval model p θ ret It can be implemented based on the pre-trained language model (Bert); generate probability p θ gen (a t ,r t |c t ,u t ,kb t ), that is, generating model p θ gen This can be achieved by a pre-trained language model (GPT2). Multiple rounds of historical conversation data are included in the training samples to improve the conversation performance after model training.

[0060] Introducing the inference model qφ (h t |c t ,u t ,r t ), to approximate the true posterior distribution, i.e., the posterior model p θ (h 1:T |u 1:T ,r 1:T ), so that unsupervised training can be performed on unlabeled second data. This model can be decomposed with the following probability:

[0061]

[0062] The conditional probability q φ (h 1:T |u 1:T ,r 1:T ) is instantiated using the following formula: φ (kb t ,a t |c t ,u t ,r t ). In the experiment, the posterior model is also implemented by a pre-trained language model (GPT2), so that kb can be obtained without KB. t , without relying on an external database KB.

[0063] In an optional example, in the process of semi-supervised training of the pre-trained dialogue model, when the first data is used as a training sample, since the hidden state, that is, the hidden variable data, is known at this time, supervised training can be performed in rounds, and the hidden variable data h is defined as t ={kb t ,a t} is to put the query result kb t and dialogue action data a t The concatenated result is used as the first label to perform semi-supervised training on the pre-trained dialogue model, where the generation probability p θ (a t ,r t |c t ,u t ,kb t ) and the inferred probability q φ (kb t ,a t |c t ,u t ,r t ) are optimized according to the maximum likelihood method, such as Figure 5 shown.

[0064] In this implementation, the semi-supervised training can be further performed by the joint stochastic approximation (JSA) method, and the overall semi-supervised knowledge-based task dialogue system is called JSA-KGTOD. In this system, for the labeled first data, the method of maximizing the joint conditional likelihood can be directly used to optimize the following objective function: log p θ (r 1:T ,h 1:T |u 1:T ).

[0065] In another optional example, in the process of performing semi-supervised training on the pre-trained dialogue model, when using the second data x i , i = 1, 2, ..., n as training samples, the conditional marginal likelihood log p of the data can be optimized according to the Monte Carlo method θ (r 1:T |u 1:T ), and the inner KL divergence between the generated model and the inferred model The latent variable data is sampled according to the round-by-round Metropolis importance sampling method, and the sampling result is determined as the second label to perform semi-supervised training on the pre-trained dialogue model, where n is a positive integer. The following sampling is performed to obtain the latent variable data h 1:T :

[0066] For the tth round according to q φ (h t |c t ,u t ,r t ) Sampling with distribution h t , ;

[0067] Sample ε from a uniform distribution in the interval [0,1] and use ε to accept or reject:

[0068]

[0069] The specific calculation method of the importance ratio in the above formula is as follows:

[0070]

[0071] When we obtain a hidden variable sample h through the above sampling 1:T After that, the model parameters are updated. The hidden state h corresponding to the second label determined according to the sampling result is 1:T As the true label, logp can be calculated through supervised training θ (r 1:T ,h 1:T |u 1:T) and logq φ (h 1:T |u 1:T ,r 1:T ) with respect to the gradient of parameters θ and φ.

[0072] Among them, the joint random estimation algorithm can be used to update the model parameters. Monte Carlo sampling can be used to sample κ from 1, 2, ..., n to select the sample point x (κ) And hidden variables When updating parameters:

[0073] Use the gradient descent method with the gradient: Update the parameters θ.

[0074] Use the gradient descent method with the gradient: Update the parameter φ.

[0075] In a specific embodiment, based on the JSA-KGTOD system, the inference model q is used on the labeled first data. φ and the retrieval model p θ ret Perform supervised pre-training without generating a model p θ gen Pre-training is performed because experimental results show that for the generative model p θ gen Pre-training will reduce model performance.

[0076] Then, randomly sample the labeled first data and the unlabeled second data to form training samples, and perform semi-supervised training on the dialogue model to obtain a trained dialogue model. Among them, for the labeled first data, the hidden state h at this time 1:T is the first label given, and for the unlabeled second data, the Metropolis importance sampling method is used for the hidden state h 1:T Sampling is performed and the sampling result is regarded as the real second label, and the parameters θ and φ are optimized. The subsequent gradient calculation and model parameter update are the same for labeled data and unlabeled data.

[0077] Among them, for unlabeled data, since KB cannot be obtained, the retrieval probability We approximate it as kb t i The experiment shows that this approximation is correct for p without KB. θ ret (kb t ∣c t ,u t ) is a better way to estimate .

[0078] In an optional embodiment, in the process of semi-supervised training of the pre-trained dialogue model, since there is no KB in the unlabeled data, for the training samples, in the semi-supervised training based on JSA, the model p of the generated model is θ gen The parameters are updated without updating the retrieval model p θ ret , has better stability and higher data utilization, thereby improving model performance.

[0079] In this way, the semi-supervised training is performed by combining the Joint Stochastic Approximation (JSA) method. The retrieval enhancement module is integrated into the modeling. The retrieval enhancement method greatly improves the ability of the task-based dialogue system to combine knowledge. The query results generated based on the retrieval model are sent as conditional input to the generation model, and then the generation model is used to generate inferred generation dialogue action data and system response data, which can be better applied to knowledge-based tasks.

[0080] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a model training device provided by an embodiment of the present invention, such as Figure 6 As shown, the model training device 600 includes:

[0081] A construction module 601 is used to construct a dialogue model to be trained, wherein the dialogue model to be trained includes a generation model, an inference model and a retrieval model;

[0082] A pre-training module 602 is used to perform supervised pre-training on the inference model and the retrieval model in the dialogue model to be trained based on the labeled first data, wherein the pre-trained inference model is used to obtain latent variable data according to the user input data and the system response data in the first data, and the pre-trained retrieval model is used to retrieve the query result from the database according to the user input data in the first data;

[0083] The training module 603 is used to perform semi-supervised training on the pre-trained dialogue model based on the training samples formed by the first data and the unlabeled second data to obtain a trained dialogue model, wherein the generation model in the trained dialogue model is used to generate dialogue action data and system response data.

[0084] Optionally, during the semi-supervised training of the pre-trained dialogue model, model parameters of the generation model are updated based on the training samples.

[0085] Optionally, in the process of performing semi-supervised training on the pre-trained dialogue model, when the first data is used as a training sample, the latent variable data is a first label spliced ​​according to the query result and dialogue action data in the first data.

[0086] Optionally, in the process of performing semi-supervised training on the pre-trained dialogue model, when the second data is used as a training sample, the latent variable data is sampled according to an iterative importance sampling method, and the sampling result is determined as the second label.

[0087] Optionally, the training sample includes T rounds of dialogue, the T rounds of dialogue include historical dialogue data from the 1st round to the T-1th round and dialogue data of the Tth round, each round of dialogue data includes user input data and corresponding system response data, and T is a positive integer greater than 1.

[0088] It should be noted that the model training device 600 is capable of implementing each process of each embodiment of the above-mentioned model training method, the technical features correspond one to one, and can achieve the same technical effect. To avoid repetition, it will not be described here.

[0089] An embodiment of the present invention provides a knowledge-based task dialogue system, which includes a dialogue model obtained according to the above-mentioned model training method.

[0090] It should be noted that the knowledge-based task dialogue system includes a dialogue model obtained according to the above-mentioned model training method, and can implement each process of each embodiment of the above-mentioned model training method. The technical features correspond one to one and can achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0091] An embodiment of the present invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, the various processes of the above-mentioned model training method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0092] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above-mentioned model training method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0093] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present invention is not limited to performing functions in the order discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0094] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0095] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.

Claims

1. A model training method, characterized in that: The method comprises: Construct a dialogue model to be trained, wherein the dialogue model to be trained includes a generation model, an inference model, and a retrieval model; wherein, when constructing the dialogue model to be trained, knowledge-based dialogue modeling is introduced into the task-oriented dialogue, and the model is a latent variable dialogue model, which is implemented by constructing a model with parameters θ and φ, and the generation model is denoted by p θ gen , the inference model is denoted as q φ , the retrieval model is denoted as p θ ret ; and find the query results from the database through the retrieval model; wherein the knowledge-based task dialogue system is implemented by a model with a parameter θ, and the probability is decomposed into each round: Among them, p θ (h t ,r t |c t ,u t ) is instantiated with the following formula and decomposed into retrieval probability and generation probability: p θ (h t ,r t |c t ,u t ) =p θ (kb t ,a t ,r t |c t ,u t ) =p θ ret (kb t |c t ,u t )p θ gen (a t ,r t |c t ,u t ,kb t ) Inference model q φ (h t |c t ,u t ,r t ) is decomposed with the following probability: The conditional probability q φ (h 1:T |u 1:T ,r 1:T ) is instantiated using the following formula: φ (kb t ,a t |c t ,u t ,r t ) is the inferred probability; Based on the labeled first data, supervised pre-training is performed on the inference model and the retrieval model in the dialogue model to be trained, wherein the pre-trained inference model is used to obtain latent variable data according to the user input data and the system response data in the first data, and the pre-trained retrieval model is used to retrieve the query result from the database according to the user input data in the first data; Based on the training samples formed by the first data and the unlabeled second data, semi-supervised training is performed on the pre-trained dialogue model to obtain a trained dialogue model, wherein the generation model in the trained dialogue model is used to generate dialogue action data and system response data; In the process of semi-supervised training of the pre-trained dialogue model, when the first data is used as a training sample, the latent variable data is a first label formed by splicing the query result and the dialogue action data in the first data; the generation probability p θ (a t ,r t |c t ,u t ,kb t ) and the inference probability q φ (kb t ,a t |c t ,u t ,r t ) are optimized according to the maximum likelihood method; In the process of performing semi-supervised training on the pre-trained dialogue model, when the second data is used as a training sample, the latent variable data is sampled according to an iterative importance sampling method, and the sampling result is determined as a second label.

2. The method according to claim 1, characterized in that In the process of performing semi-supervised training on the pre-trained dialogue model, the model parameters of the generation model are updated based on the training samples.

3. The method according to claim 1, characterized in that The training sample includes T rounds of dialogue, and the T rounds of dialogue include historical dialogue data from the 1st round to the T-1th round and dialogue data of the Tth round. Each round of dialogue data includes user input data and corresponding system response data, and T is a positive integer greater than 1.

4. A model training device, characterized in that: The device comprises: A construction module is used to construct a dialogue model to be trained, wherein the dialogue model to be trained includes a generation model, an inference model, and a retrieval model; wherein, when constructing the dialogue model to be trained, knowledge-based dialogue modeling is introduced into the task-oriented dialogue, and the model is modeled as a latent variable dialogue model, which is implemented by constructing a model with parameters θ and φ, and the generation model is denoted by p θ gen , the inference model is denoted as q φ , the retrieval model is denoted as p θ ret ; and find the query results from the database through the retrieval model; wherein the knowledge-based task dialogue system is implemented by a model with a parameter θ, and the probability is decomposed into each round: Among them, p θ (h t ,r t |c t ,u t ) is instantiated with the following formula and decomposed into retrieval probability and generation probability: p θ (h t ,r t |c t ,u t ) =p θ (kb t ,a t ,r t |c t ,u t ) =p θ ret (kb t |c t ,u t )p θ gen (a t ,r t |c t ,u t ,kb t ) Inference model q φ (h t |c t ,u t ,r t ) is decomposed with the following probability: The conditional probability q φ (h 1:T |u 1:T ,r 1:T ) is instantiated using the following formula: φ (kb t ,a t |c t ,u t ,r t ) is the inferred probability; A pre-training module, used to perform supervised pre-training on the inference model and the retrieval model in the dialogue model to be trained based on the labeled first data, wherein the pre-trained inference model is used to obtain latent variable data according to the user input data and the system response data in the first data, and the pre-trained retrieval model is used to retrieve the query result from the database according to the user input data in the first data; A training module, configured to perform semi-supervised training on the pre-trained dialogue model based on training samples formed by the first data and the unlabeled second data, to obtain a trained dialogue model, wherein the generation model in the trained dialogue model is used to generate dialogue action data and system response data; In the process of semi-supervised training of the pre-trained dialogue model, when the first data is used as a training sample, the latent variable data is a first label formed by splicing the query result and the dialogue action data in the first data; the generation probability p θ (a t ,r t |c t ,u t ,kb t ) and the inference probability q φ (kb t ,a t |c t ,u t ,r t ) are optimized according to the maximum likelihood method; In the process of performing semi-supervised training on the pre-trained dialogue model, when the second data is used as a training sample, the latent variable data is sampled according to an iterative importance sampling method, and the sampling result is determined as a second label.

5. The device according to claim 4, characterized in that In the process of performing semi-supervised training on the pre-trained dialogue model, the model parameters of the generation model are updated based on the training samples.

6. The device according to claim 4, characterized in that The training sample includes T rounds of dialogue, and the T rounds of dialogue include historical dialogue data from the 1st round to the T-1th round and dialogue data of the Tth round. Each round of dialogue data includes user input data and corresponding system response data, and T is a positive integer greater than 1.

7. A knowledge-based task dialogue system, characterized in that: The system includes a dialogue model obtained by the model training method according to any one of claims 1 to 3.

8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the model training method as described in any one of claims 1 to 3.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the model training method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Semi-supervised learning method, system, device, storage medium and semantic parsing method

    CN112464645A

  • Training method of dialogue generation model, dialogue method, equipment and storage medium

    CN112559706A