Gating fusion-based identity consistency dialogue text generation method

By adopting a gated fusion-based encoder model in the dialogue system, the current dialogue information and identity information are encoded and processed, and the problem of inconsistent reply information and user identity information in the existing technology is solved, achieving a high-quality dialogue experience.

CN120179791APending Publication Date: 2025-06-20NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510629751.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When generating replies, it is difficult for existing dialogue systems to accurately ensure that they are consistent with the user's identity information, causing the user to be out of the dialogue context and affect the dialogue experience.

Method used

Using an encoder model based on gated fusion, the current dialogue information and identity information are encoded and processed through the bidirectional converter model and the bidirectional gating cycle unit, joint information is obtained, and initial reply information is generated through the decoder, and finally the generator is ensured that the reply information is consistent with the identity information.

Benefits of technology

It realizes the accurate generation of reply information consistent with user identity information in the dialogue context, improving dialogue experience and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179791A_ABST
    Figure CN120179791A_ABST
Patent Text Reader

Abstract

The invention discloses an identity consistency dialogue text generation method based on gating fusion, which comprises the following steps of: independently encoding current dialogue information and identity information in spliced information by using a bidirectional converter model in an encoder, and obtaining joint information representation based on encoded identity information representation and current dialogue information representation; the encoder comprises a bidirectional converter model, a bidirectional gating circulation unit and a gating fusion layer. Through joint information representation, a multi-head attention layer and a feedforward neural network of the bidirectional transducer model obtain a first hidden layer vector; processing the joint information representation through forward and reverse gating circulation units of the bidirectional gating circulation unit to obtain a second hidden layer vector and a third hidden layer vector, and fusing the first hidden layer vector, the second hidden layer vector and the third hidden layer vector; the fused hidden layer vector is decoded through the decoder, the decoded reply information and identity information are input into the generator to obtain final reply information, and the reply information consistent with the identity information of the user can be generated in the dialogue context.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text generation research, and relates to, but is not limited to, a method for generating identity-consistent dialogue texts based on gated fusion. Background Art

[0002] With the rapid development of technology and the continuous expansion of application fields, dialogue systems have been widely used in various scenarios: virtual voice assistants, emotional companion robots, intelligent customer service systems, and vehicle-human interaction systems. To meet the above scenario requirements, various enterprises have invested a large amount of research in the field of dialogue systems. These intelligent dialogue customer services classify and answer users' questions through trained models, handle some common problems, and save a large amount of resources and manpower. In this context, the research on dialogue systems has seen a booming development, attracting a large amount of investment from the industrial community and extensive attention from the academic community. Although some progress has been made in dialogue systems based on generative models, there is a problem that the generated responses of the models are inconsistent with the role identity information, manifested as the generated text changing the semantics of the predefined role identity or conflicting with it. Even a small amount of inconsistent content may cause users to break away from the current dialogue context, lose interest in further dialogue, and affect the dialogue experience. For example, if the system first claims to be an engineer and then says in a subsequent dialogue that it is in junior high school, this will confuse users and reduce their interest in further dialogue.

[0003] In related technologies, a memory-augmented architecture has been proposed to utilize the character information in the context, and a neural variational autoencoder model has been introduced to generate diverse and sustainable dialogues, or a character-based dialogue generator has been proposed to support the dialogue between two character-based chatbots by modeling each other's roles. However, in the above methods, when the model processes complex text sequences, due to its limited encoder feature extraction ability, it is often easy to lose the long-term dependencies within the text, resulting in insufficient feature information obtained by the decoder; and although the model uses a large amount of data for training, the generated responses contain role information to a certain extent, but cannot explicitly reflect the role identity information corresponding to the previous context of the dialogue query, and cannot generate responses related to the role identity in the dialogue context, etc.

[0004] Therefore, how to accurately generate response information consistent with the user's identity information in the dialogue context has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a method for generating identity-consistent dialogue texts based on gated fusion, which at least solves the problem that the related technology cannot accurately generate response information consistent with the user's identity information in the dialogue context.

[0006] According to the first aspect of the embodiments of the present invention, a method for generating identity-consistent dialogue text based on gated fusion is provided, including: Separately encoding the current dialogue information and identity information in the spliced information by using a bidirectional transducer model in the encoder to obtain an identity information representation and a current dialogue information representation; and obtaining a joint information representation based on the identity information representation and the current dialogue information representation; the encoder includes the bidirectional transducer model, a bidirectional gated recurrent unit, and a gated fusion layer; Obtaining a first hidden layer vector based on the joint information representation, a first multi-head attention layer, and a first feed-forward neural network in the bidirectional transducer model; Processing the joint information representation respectively by a forward gated recurrent unit and a backward gated recurrent unit in the bidirectional gated recurrent unit to obtain a second hidden layer vector and a third hidden layer vector, and obtaining a fused hidden layer vector based on the first hidden layer vector, the second hidden layer vector, and the third hidden layer vector; Decoding the hidden layer vector by a decoder to obtain an initial reply message; and inputting the initial reply message and the identity information into a generator to obtain a final reply message.

[0007] According to the second aspect of the embodiments of the present invention, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete mutual communication through the communication bus; the memory is used for storing at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the method described in the first aspect.

[0008] According to the third aspect of the embodiments of the present invention, a computer storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in the first aspect or the second aspect.

[0009] The solution provided by the embodiments of the present invention uses a bidirectional transducer model in the encoder to separately encode the current conversation information and identity information in the spliced information, obtaining an identity information representation and a current conversation information representation; and obtaining a joint information representation based on the identity information representation and the current conversation information representation; the encoder includes the bidirectional transducer model, a bidirectional gated recurrent unit, and a gated fusion layer; obtaining a first hidden layer vector based on the joint information representation, the first multi-head attention layer, and the first feed-forward neural network in the bidirectional transducer model; processing the joint information representation through the forward gated recurrent unit and the backward gated recurrent unit in the bidirectional gated recurrent unit respectively to obtain a second hidden layer vector and a third hidden layer vector, and obtaining a fused hidden layer vector based on the first hidden layer vector, the second hidden layer vector, and the third hidden layer vector; decoding the hidden layer vector through a decoder to obtain an initial reply information; and inputting the initial reply information and the identity information into a generator to obtain a final reply information. In this process, the spliced information consists of the current conversation information and the identity information, and splicing can fuse two pieces of information from different sources or types together, thereby providing a richer data representation. Using the bidirectional transducer model and the bidirectional gated recurrent unit in the encoder to process the joint information representation respectively improves the feature extraction ability of the encoder and further provides sufficient feature information for the decoder. Finally, the initial reply information and the identity information are input into the generator for consistency understanding and reply text generation processing to obtain the final reply information, enabling the accurate generation of reply information consistent with the user's identity information in the conversation context. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where: Figure 1 is a flowchart showing a method for generating identity-consistent dialogue text based on gated fusion provided by an embodiment of the present invention Figure 1 ; Figure 2 is a flowchart showing a feature fusion provided by an embodiment of the present invention; Figure 3 is a schematic diagram showing the decoding effect of a decoder provided by an embodiment of the present invention; Figure 4 is a flowchart showing a method for generating identity-consistent dialogue text based on gated fusion provided by an embodiment of the present invention Figure 2 ; Figure 5Schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0011] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, rather than all, of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0012] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0013] It should be noted that the terms "first / second / third" involved in the embodiments of the present invention are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.

[0014] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as that generally understood by those of ordinary skill in the art in the field to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0015] Figure 1 Flow schematic of a method for generating identity-consistent dialogue text based on gated fusion provided by an embodiment of the present invention Figure 1 A method for generating identity-consistent dialogue text based on gated fusion provided by an embodiment of the present invention can be executed by an electronic device, which can be, for example, a computer, a server, etc.

[0016] As Figure 1 shown, a method for generating identity-consistent dialogue text based on gated fusion includes: S101. Use the bidirectional transducer model in the encoder to separately encode the current conversation information and identity information in the spliced information to obtain an identity information representation and a current conversation information representation; and obtain a joint information representation based on the identity information representation and the current conversation information representation. The encoder includes a bidirectional transducer model, a bidirectional gated recurrent unit, and a gated fusion layer.

[0017] In an embodiment of the present invention, the encoder includes a bidirectional transducer model, a bidirectional gated recurrent unit (Bidirectional Gated Recurrent Unit, BiGRU), and a gated fusion layer. The bidirectional gated recurrent unit includes a forward gated recurrent unit and a backward gated recurrent unit. The spliced information consists of the current conversation information q, the identity information p, and an interval identifier. The current conversation information can be the information newly input by the user in the conversation system, and the identity information can be the identity information corresponding to the user. Inputting the spliced information into the encoder, the encoding layer in the bidirectional transducer model in the encoder can separately encode the current conversation information and the identity information in the spliced information to obtain an identity information representation and a current conversation information representation, and obtain a joint information representation through the identity message representation and the current conversation information representation.

[0018] In an embodiment of the present invention, a sub-task in the conversation system is mainly studied, that is, open-domain single-round conversation generation, that is, for the current conversation information input by the user, the conversation system only generates the final reply information corresponding to the current conversation information.

[0019] Exemplarily, input the spliced information into the encoder, and the bidirectional transducer model in the encoder separately encodes the current conversation information and the identity information in the spliced information to obtain an identity information representation and a current conversation information representation , and the joint information representation is , where m is the length of the identity information and n is the length of the current conversation information.

[0020] S102. Obtain a first hidden layer vector based on the joint information representation, the first multi-head attention layer in the bidirectional transducer model, and the first feed-forward neural network.

[0021] In an embodiment of the present invention, input the joint information representation into the first multi-head attention layer in the bidirectional transducer model. The inputs of the Q, K, and V matrices in each self-attention are all in the form of an embedding vector (i.e., the joint information representation) as the input, and use the scaled dot product method to calculate Attention. The calculation formulas are as shown in the following (1) and (2): ; ; First, Attention is the attention calculation formula, softmax is the formula for converting a numerical vector into a probability distribution, Q, K, and V are the input matrices, and d k is the dimension of K, and k T is the transpose matrix of matrix K. The Q, K, and V vector matrices are subjected to multiple linear transformations to generate multiple attention heads. Each self-attention head is mapped to a different subspace for the corresponding self-attention mechanism, so as to capture information in different representation subspaces and different dimensions. Then, each self-attention head corresponds to a group of weight matrices W Q , W K , W V to perform self-attention calculation to obtain the self-attention matrix. Secondly, the self-attention matrices are concatenated and matrix multiplication is performed with the additional weight matrix W O , and finally the result obtained from the calculation is used as the input of the output layer. The specific calculation is as shown in (3), (4), and (5) below: ; ; ; Among them, respectively represent weight matrices, , and are learnable training parameters; is the attention calculation result of the th head; is the number of self-attention heads, Z is the hidden layer vector after weighted summation, Multihead represents the multi-head attention calculation formula, and concat represents the attention head concatenation formula.

[0022] The multiple attention features are concatenated and linearly projected to obtain the attention feature matrix; then residual connection is performed, and at the same time, linear standard normalization processing is carried out, and the processed matrix is passed into the feed-forward neural network (FNN) composed of two fully connected networks and activated with the activation function ReLU; the hidden layer vector of each layer is as shown in formula (6) below: (6); Among them: is the hidden vector of the th layer in the encoder, k = 0, 1,..., N - 1, is the encoded hidden layer representation , that is, the hidden variable of the th layer, that is, the first hidden layer vector obtained.

[0023] Among them, when k is 0, the joint information representation is used as , and then after being processed by the first layer, 。

[0024] S103. Process the joint information representation through the forward gated recurrent unit and the backward gated recurrent unit in the bidirectional gated recurrent unit respectively to obtain a second hidden layer vector and a third hidden layer vector, and obtain a fused hidden layer vector based on the first hidden layer vector, the second hidden layer vector and the third hidden layer vector.

[0025] In an embodiment of the present invention, the word vector sequence after the joint information representation is processed by the word embedding layer is respectively used as the input of the forward gated recurrent unit (Gated Recurrent Unit, GRU) and the backward gated recurrent unit. Taking the forward GRU as an example, the second hidden layer vector of the hidden layer at this moment is calculated according to the following formula (7) : (7); Where: is the hidden layer vector output at the th moment, and

[0026] is the word vector input at the th moment (i.e., the joint information representation). 。

[0027] Further, the first hidden layer vector and the second hidden layer vector are concatenated to obtain a bidirectional hidden layer vector , such as: , and then the concatenated hidden layer vector is processed by the self-attention layer of the BiGRU model to obtain a hidden layer vector . In the hidden layer vector, is the maximum length of the input. Finally, the hidden layer vector and the hidden layer vector are input into the gated fusion layer to retain effective feature information and remove redundant features. Finally, the two feature vectors are integrated into one. The feature fusion formula is shown in the following formulas (8) and (9): (8); (9); In the above formula, is the parameter of the fully connected layer, is the activation function, is the first hidden layer vector output by the bidirectional transformer model at the th moment, and The hidden layer vector output at time BiGRU, is the hidden layer vector at time t after fusion, is the gating signal, and its value range is .

[0028] In the embodiments of the present invention, the joint information representation is respectively input into the forward GRU and the backward GRU of the bidirectional gated recurrent unit to obtain and , where . Combine and by concatenation to obtain , where .

[0029] As shown in Figure 2 , Figure 2 is a schematic flow diagram of feature fusion provided by the embodiments of the present invention. In Figure 2 , it contains two inputs. Add through ⊕, and pass the added result through an activation function σ. Here, σ usually refers to the Sigmoid function, which maps the input value to the interval (0, 1). The activated result is multiplied by through a multiplication operation (⊙). At the same time and are multiplied by through a multiplication operation (⊙). Finally, the sum operation is performed on the multiplication results of the two to output . Among them, and maintain a relatively balanced state. For the dimensional information contained in , the gating unit will selectively forget, and the weight corresponding to the forgotten part decreases. Correspondingly, the weight in will increase.

[0030] S104. Decode the hidden layer vector through the decoder to obtain the initial response information; and input the initial response information and the identity information into the generator to obtain the final response information.

[0031] In some embodiments of the present invention, the decoder in the Transformer architecture is initialized with the trained parameters of the bidirectional transducer model. When generating the initial response information, the decoder receives partial information of the hidden layer vector as input and predicts the next word. This process is autoregressive, that is, only one word is generated each time and added to the generated sequence, and then this process is repeated until the complete initial response information is generated. At the same time, the masking technology is used to ensure the autoregressive nature during the generation process.

[0032] Further, the decoding module adopts a cross-attention mechanism to generate an initial reply message based on the context and predefined identity information. Among them, the specific decoding formula is as follows (10): ; Among them, is the fused hidden layer vector (i.e., in formula (9)); is the (i + 1)-th context hidden layer vector obtained through the cross-attention layer, is the i-th context hidden layer vector obtained through the cross-attention layer; the hidden layer in the decoder is the same as that in the encoder, both are layers, and the representation output by the last hidden layer, that is, the initial reply message .

[0033] Further, after obtaining the initial reply message, the initial reply message and the identity information are input into the trained generator to obtain the final reply message. Among them, the generator is composed of an encoder.

[0034] In the embodiment of the present invention, as Figure 3 shown, Figure 3 is a schematic diagram of the decoding effect of a decoder provided by an embodiment of the present invention. In Figure 3 , the decoder is initialized with the parameters of the bidirectional transformer model, and the word vectors after the current time step are masked. The input M of the decoder becomes the target reply vector that increases step by step from left to right, and the initial reply message is generated in an autoregressive manner.

[0035] In the embodiment of the present invention, as Figure 4 shown, Figure 4 is a schematic flow chart of a method for generating identity-consistent dialogue text based on gated fusion provided by an embodiment of the present invention Figure 2 . In Figure 4 , the identity information and the current dialogue information are respectively input into the embedding layer to obtain the word vectors after the transformation of p and q. The p and q are concatenated to obtain the concatenated information. The concatenated information is respectively input into the bidirectional transformer model and the bidirectional gated recurrent unit in the encoder. The bidirectional transformer model outputs the hidden layer vector H, and the bidirectional gated recurrent unit outputs the hidden layer vector C. The H and C are respectively input into the gated fusion layer in the encoder for fusion to obtain the fused hidden layer vector M. The M is input into the decoder and processed through the cross-attention layer in the decoder to obtain the hidden layer vector of the network layer, until the hidden layer vector output by the last N-th network layer, that is, the initial reply message , the generator includes two stages: consistent understanding and response text generation. Input and p into the generator, and finally obtain the final response information through formulas (12) and (13).

[0036] It can be understood that in the embodiments of the present invention, the bidirectional transformer model in the encoder is used to separately encode the current conversation information and identity information in the spliced information to obtain the identity information representation and the current conversation information representation; and a joint information representation is obtained based on the identity information representation and the current conversation information representation; the encoder includes the bidirectional transformer model, the bidirectional gated recurrent unit and the gated fusion layer; a first hidden layer vector is obtained based on the joint information representation, the first multi-head attention layer and the first feed-forward neural network in the bidirectional transformer model; the forward gated recurrent unit and the backward gated recurrent unit in the bidirectional gated recurrent unit are respectively used to process the joint information representation to obtain a second hidden layer vector and a third hidden layer vector, and a fused hidden layer vector is obtained based on the first hidden layer vector, the second hidden layer vector and the third hidden layer vector; the decoder decodes the hidden layer vector to obtain the initial response information; and the initial response information and the identity information are input into the generator to obtain the final response information. In this process, the spliced information is composed of the current conversation information and the identity information, and splicing can fuse two pieces of information from different sources or types together, thereby providing a richer data representation. The bidirectional transformer model and the bidirectional gated recurrent unit in the encoder are respectively used to process the joint information representation, which improves the feature extraction ability of the encoder and further provides sufficient feature information for the decoder. Finally, the initial response information and the identity information are input into the generator for consistent understanding and response text generation processing to obtain the final response information, so that a response information consistent with the user's identity information can be accurately generated in the conversation context.

[0037] In some embodiments of the present invention, before S101, S10 to S11 are further included, which are described through the following steps.

[0038] S10: Insert an interval identifier after the current conversation information input by the user to obtain the inserted information.

[0039] S11: Splice the inserted information and the identity information input by the user to obtain the spliced information.

[0040] Exemplarily, as shown in the following formula (11): ; In the above formula, is the identity information; is the interval identifier; is the current conversation information; is the length of the identity information and the current conversation information.

[0041] In the above formula, first insert an interval identifier after the current conversation information to obtain the information after insertion, and splice the information after insertion with the identity information input by the user to obtain the spliced information, that is, the input information.

[0042] In some embodiments of the present invention, S104 can be implemented through S1041 to S1042, and the following steps are used for illustration.

[0043] S1041. Obtain a third hidden layer vector based on the identity information, the second multi-head attention layer in the generator, and the second feed-forward neural network.

[0044] S1042. Obtain the final reply information based on the third hidden layer vector, the initial reply information, the third multi-head attention layer in the generator, and the third feed-forward neural network.

[0045] In some embodiments of the present invention, the generator is used to perform understanding reasoning and secondary generation on the decoded initial reply information. Specifically, the second multi-head attention layer and the second feed-forward neural network in the generator are used to process the identity information to obtain a third hidden layer vector. The third multi-head attention layer and the third feed-forward neural network in the generator are used to process the third hidden layer vector and the initial reply information to obtain the final reply information. As shown in the following formulas (12) and (13): ; ; In the above formula, is the i-th hidden layer vector obtained through the third multi-head attention mechanism, is the (i + 1)-th hidden layer vector obtained through the third multi-head attention mechanism, is the (i + 1)-th response generation output result of each network layer. The two multi-head attention layers in the consistency understanding generator share parameters to reduce the mutual interference between two different tasks. The output of the last layer of the generator , is a set of n , and finally obtains the final reply information after consistent understanding through a linear output layer.

[0046] The embodiments of the present invention use natural language reasoning on a non-dialogue reasoning data set to help the hidden layer establish a dependence on the identity information, further improving the semantic consistency of the final reply information.

[0047] In some embodiments of the present invention, the training process of the generator can be implemented through S201 to S204, and the following steps are used for illustration.

[0048] S201. Collect a first training data set with semantically positive correlation and a second training data set with semantically negative correlation; the first training data set includes first premise samples and first hypothesis samples, and the second training data set includes second premise samples and second hypothesis samples.

[0049] In some embodiments of the present invention, a first training data set and a second training data set are collected. The first training data set is a data set with semantically positive correlation, and the second training data set is a data set with semantically negative correlation. Among them, the first training data set contains first premise samples and first hypothesis samples, and the second training data set includes second premise samples and second hypothesis samples, as shown in the following formula (14): (14); In the above, is the first training data set with semantically positive correlation, is the second training data set with semantically negative correlation, is the th premise sample in the first training data set; is the th premise sample in the second training data set; is the th hypothesis sample that is semantically positively correlated with the premise; is the th hypothesis sample that is semantically contradictory to the premise.

[0050] S202. Obtain the maximum likelihood loss based on the probability that there is a relationship between the first premise sample and the first hypothesis sample.

[0051] In some embodiments of the present invention, for the training of data with positive semantic correlation, the present invention uses the maximum likelihood loss, as shown in the following formula (15): (15); In the above formula, is the maximum likelihood loss, is the and probability that there is a relationship, is the sequence of all elements before the th position, is the hypothesis sample, is the premise sample, is the sequence.

[0052] S203. Obtain the non - likelihood loss based on the probability that there is a relationship between the second premise sample and the second hypothesis sample.

[0053] In some embodiments of the present invention, for data with negative semantic correlation, a non-likelihood loss is used to minimize the possibility of semantic contradiction in the generated text, as shown in the following (16): (16); In the above formula, is the non-likelihood loss, is and the probability of the relationship, is all elements before the -th position in the sequence.

[0054] S204. Obtain the first total loss of the consistency inference subtask through the maximum likelihood loss and the non-likelihood loss; and obtain the generator based on the first total loss.

[0055] In some embodiments of the present invention, the maximum likelihood loss and the non-likelihood loss are summed to obtain the first total loss of the consistency inference subtask, as shown in the following (17): (17); In the above formula, is the first total loss.

[0056] Furthermore, the generator is trained based on the first total loss to obtain a trained generator.

[0057] In some embodiments of the present invention, S204 can be implemented through S2041 to S2043, and the description is as follows through the following steps.

[0058] S2041. Obtain the first negative log-likelihood loss based on the initial reply information training sample, the identity information training sample, the current dialogue information training sample, and the first loss calculation formula.

[0059] In some embodiments of the present invention, the encoder to be trained reads the identity information training sample and the current dialogue information training sample, and finally the decoder to be trained outputs the initial reply information training sample. The decoder to be trained hopes that the initial reply information training sample approaches the target reply information training sample. Therefore, the first loss calculation formula is as shown in the following formula (18): (18); In the above formula, is the -th generated initial reply information training sample, is and the probability of the relationship, is the identity information training sample, For the training samples of the current dialogue information, For all elements before the i-th position in the target sequence.

[0060] S2042. Obtain the second negative log-likelihood loss based on the initial response information training samples, identity information training samples, and the second loss calculation formula.

[0061] S2043. Obtain the generator based on the first total loss, the first negative log-likelihood loss, and the second negative log-likelihood loss.

[0062] In some embodiments of the present invention, the consistency understanding generator is also trained using the NLL loss function. Inputting the identity information training samples and the initial response information training samples output by the decoder, and hoping to predict the target response information training samples. Therefore, the second loss calculation formula is as shown in formula (19) below: (19); In the above formula, is the initial response information training sample; is and the probability of the relationship between

[0063] Furthermore, train the generator to be trained based on the first total loss, the first negative log-likelihood loss, and the second negative log-likelihood loss to obtain the trained generator.

[0064] In some embodiments of the present invention, S2043 can be implemented through S3201 and S302, and the specific description is as follows through the following steps.

[0065] S301. Sum the first negative log-likelihood loss and the second negative log-likelihood loss to obtain the second total loss of the response generation task.

[0066] In some embodiments of the present invention, as shown in formula (20) below: (20); In the above formula, is the second total loss.

[0067] S302. Sum the first total loss and the second total loss to obtain the target total loss, and train the generator to be trained based on the target total loss to obtain the generator.

[0068] In some embodiments of the present invention, as shown in formula (21) below: (21); In the above formula, is the target total loss.

[0069] In an embodiment of the present invention, to verify the effectiveness of an identity-consistent dialogue text generation method based on gated fusion, the performance of the model in generating text was compared, and the results are shown in Table 1 and Table 2. Among them, ICGF is the identity-consistent dialogue text generation model based on gated fusion proposed in the present invention, and other models are comparison models. Among them, ppl is the perplexity, which measures the uncertainty of the model in predicting samples. The lower the value, the higher the confidence of the model in generating text, and the smaller the value, the better. is the unigram diversity, which calculates the proportion of unique unigrams (single words) in the generated text among all words, reflecting the lexical richness, and the larger the value, the better; Dist.2 is the bigram diversity, which calculates the proportion of unique bigrams (combinations of two adjacent words) in the generated text, reflecting the phrase diversity, and the larger the value, the better; is the consistency score, and the larger the value, the better, is the personalization gain, and the larger the value, the better.

[0070] Table 1 Performance of the model on the PersonChat dataset ; Table 2 Performance of the model on the PersonalDialog dataset ; Table 1 and Table 2 respectively show the results of the present invention and other technical models on two experimental data. The model proposed in this chapter achieved the top two performance in the consistency index, demonstrating the effectiveness of adding constraints in the process of fusing role identity information.

[0071] Because this chapter uses a gated fusion layer and uses a consistency understanding generator to perform secondary understanding and generation on the initial response generated by the model, the model achieved the optimal performance in terms of the consistency of the generated response.

[0072] Compared with other models, the index increased by 46.16%, the index increased by 3.77%, the index increased by 12.89%, the index increased by 4.04%, the index increased by 60.04%.

[0073] Referring to Figure 5 , a schematic structural diagram of an electronic device according to an embodiment of the present invention is shown. The specific implementation of the electronic device is not limited in the specific embodiments of the present invention.

[0074] As Figure 5As shown in the figure, the electronic device may include: a processor 502, a communications interface 504, a memory 506, and a communication bus 508.

[0075] Among them: The processor 502, the communications interface 504, and the memory 506 communicate with each other through the communication bus 508.

[0076] The communications interface 504 is used to communicate with other electronic devices or servers.

[0077] The processor 502 is used to execute the program 510, and specifically can execute the relevant steps in the above method embodiments.

[0078] Specifically, the program 510 may include program code, and the program code includes computer operation instructions.

[0079] The processor 502 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0080] The memory 506 is used to store the program 510. The memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0081] The program 510 is specifically used to cause the processor 502 to execute the operations corresponding to the methods described in the above method embodiments.

[0082] For the specific implementation of each step in the program 510, reference may be made to the corresponding steps and descriptions in the corresponding units in the above method embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.

[0083] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present invention can be split into more components / steps, or two or more components / steps or partial operations of the components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.

[0084] The method according to an embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as software or computer code stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or can be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a RAM, a ROM, a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.

[0085] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.

[0086] The above embodiments are only used to illustrate the embodiments of the present invention, rather than to limit the embodiments of the present invention. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims.

Claims

1. A method for generating identity-consistent dialogue text based on gated fusion, characterized in that: include: The current conversation information and the identity information in the concatenated information are separately encoded using a bidirectional transformer model in the encoder to obtain the identity information representation and the current conversation information representation; and obtaining a joint information representation based on the identity information representation and the current conversation information representation; the encoder includes the bidirectional transformer model, the bidirectional gated recurrent unit and the gated fusion layer; Obtaining a first hidden layer vector based on the joint information representation, a first multi-head attention layer and a first feedforward neural network in the bidirectional transformer model; The joint information representation is processed by a forward gated recurrent unit and a reverse gated recurrent unit in the bidirectional gated recurrent unit to obtain a second hidden layer vector and a third hidden layer vector, and a fused hidden layer vector is obtained based on the first hidden layer vector, the second hidden layer vector and the third hidden layer vector; The hidden layer vector is decoded by a decoder to obtain initial response information; and the initial response information and the identity information are input into a generator to obtain final response information.

2. The method according to claim 1, characterized in that Before the bidirectional transformer model in the encoder is used to separately encode the current conversation information and the identity information in the spliced ​​information to obtain the identity information representation and the current conversation information representation, the method further includes: Inserting an interval identifier after the current conversation information input by the user to obtain the inserted information; The inserted information and the identity information input by the user are spliced ​​to obtain the spliced ​​information.

3. The method according to claim 1, characterized in that The step of inputting the initial reply information and the identity information into a generator to obtain final reply information includes: Acquire a fourth hidden layer vector based on the identity information, a second multi-head attention layer and a second feedforward neural network in the generator; The final reply information is obtained based on the fourth hidden layer vector, the initial reply information, the third multi-head attention layer in the generator, and the third feedforward neural network.

4. The method according to claim 1, characterized in that The generator is trained through the following process: Collecting a first training data set with positive semantic correlation and a second training data set with negative semantic correlation; the first training data set includes a first premise sample and a first hypothesis sample, and the second training data set includes a second premise sample and a second hypothesis sample; Obtaining a maximum likelihood loss based on a probability that a relationship exists between the first premise sample and the first hypothesis sample; Obtaining a non-likelihood loss based on a probability that a relationship exists between the second premise sample and the second hypothesis sample; Obtain a first total loss of the consistency reasoning subtask through the maximum likelihood loss and the non-likelihood loss; and obtain the generator based on the first total loss.

5. The method according to claim 4, characterized in that The obtaining the generator based on the first total loss comprises: Obtaining a first negative log-likelihood loss based on the initial reply information training sample, the identity information training sample, the current dialogue information training sample and the first loss calculation formula; Obtaining a second negative log-likelihood loss based on the initial reply information training sample, the identity information training sample and a second loss calculation formula; The generator is obtained based on the first total loss, the first negative log-likelihood loss, and the second negative log-likelihood loss.

6. The method according to claim 5, characterized in that The obtaining the generator based on the first total loss, the first negative log-likelihood loss, and the second negative log-likelihood loss comprises: The first negative log-likelihood loss and the second negative log-likelihood loss are summed to obtain a second total loss for the response generation task; The first total loss and the second total loss are summed to obtain a target total loss, and the generator to be trained is trained based on the target total loss to obtain the generator.