Conversation generation method based on complex interaction factor modeling
By introducing complex interaction factor modeling and collaborative graph joint modeling technology into the dialogue generation method, the problem of failure to effectively consider multiple interaction factors in the existing technology is solved, and a more human-style and context-related dialogue response is achieved.
Patent Information
- Application Number
- CN202411989926.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art fails to effectively consider the mutual influence between multiple interaction factors when generating dialogues, resulting in a lack of human style and contextual relevance of the generated response.
A dialogue generation method based on complex interactive factors is adopted to extract representations of historical dialogues, personality, emotions and topics through an encoder, and a collaborative graph is used to jointly model the temporal dynamics of each factor in each round of dialogue to generate a more human-style response.
By taking into account multiple interaction factors comprehensively, the generated response is more semantic, appropriate and consistent in topics, improving the quality and human style of dialogue generation.
Smart Images

Figure CN119990308A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence technology, and in particular to a dialogue generation method based on complex interactive factor modeling. Background Art
[0002] Open domain dialogue systems aim to achieve highly realistic human-computer interactions. They not only play an increasingly important role in daily life, but are also widely used in industrial production, such as question-answering systems. In order to improve the quality of generated dialogues, more and more factors are integrated into dialogue systems. Recent studies have explored multiple aspects, such as the speaker's emotions, the consistency of character settings, and the diversity of topics, to better simulate human dialogues. These studies have significantly promoted the development of human-computer dialogue systems.
[0003] However, most existing works only focus on the individual effects of different interaction factors on response generation, which is incomplete because the interaction factors between humans and chatbots influence each other. If the speaker's personality is not taken into account, when speaker A shares negative news about speaker B's idol, speaker B may respond positively and continue the conversation. However, if speaker B's personality as a fan is taken into account, they are more likely to show negative emotions and try to end or change the topic of the conversation. Obviously, people's responses are affected by multiple interaction factors, so jointly modeling these factors is the key to generating high-quality responses. Summary of the invention
[0004] The technical problem to be solved by the present invention is to provide a dialogue generation method based on complex interaction factor modeling in view of the defects in the prior art.
[0005] The technical solution adopted by the present invention to solve the technical problem is: a dialogue generation method based on complex interactive factor modeling, comprising the following steps:
[0006] 1) Given a set of predefined personalities P and historical conversation sentences X;
[0007] 2) Use the encoder to encode the historical conversation sentences and predefined personalities to obtain various representations in the embedding space, including: personality representation sequence P pre , historical dialogue representation x, historical emotion representation sequence E his and historical topic representation sequence T his ; The details are as follows:
[0008] 2.1) For each personality p i ∈P, encode it separately to get the personality representation sequence P pre ={p1,p2,…,p n};
[0009] 2.2) For each utterance x in the historical dialogue i ∈X is encoded to get the sentence representation
[0010] 2.3) X enc Input into the bidirectional RNN to extract semantic information and obtain the contained semantic information
[0011] X s =RNN semantic (X enc );
[0012] 2.4) Calculate X s The weight of each semantic representation in , and perform weighted summation to obtain the historical dialogue representation x:
[0013]
[0014] The score function is as follows:
[0015]
[0016] Among them, W1 and W2 are trainable linear layers, and b1 is the bias term;
[0017] 2.5) X enc Input into another bidirectional RNN to extract the specific historical sentiment sequence representation E his ={e1,e2,…,e m}:
[0018] E his =RNN emotion (X enc );
[0019] 2.6) X enc Input into another bidirectional RNN to extract a specific historical topic sequence representation T his ={t1,t2,…,t m}:
[0020] T his =RNN topic (X enc );
[0021] 3) Joint modeling of interaction factors;
[0022] Interaction factors influence each other and change over time, and collaboration graphs are used to jointly model the temporal dynamics of each factor in each round of dialogue;
[0023] 3.1) Representing sequence P according to personality pre, historical dialogue representation x, historical emotion representation sequence E his and historical topic representation sequence T his , get the personality node p i , Topic history node t i , emotional history node e i and semantic node x;
[0024] 3.2) Establish personality state node P, emotion state node E and topic state node T;
[0025] 3.3) Establish fusion node;
[0026] 3.4) Obtain the dependency relationship between nodes;
[0027] 3.4.1) The personality displayed by the speaker in the conversation should be consistent with his own personality and highly relevant to the current conversation. Therefore, the personality state node P in the collaboration graph is based on the personality node P pre and semantic node x, that is, P(P|x,P pre ):
[0028] P = TransEnc(x,P pre ),
[0029] Among them, TransEnc is a single-layer Transformer encoder;
[0030] 3.4.2) Update the topic state node T using the following formula:
[0031] f t =x+g t ·P+(1-g t )·f,
[0032] T=TransEnc(f t ,T his ),
[0033]
[0034] Where f = TransEnc(x,W3[f e ′ ;f t ′ ]+b3);
[0035] f e ′ =TransEnc(T his ,E his ),
[0036] f t ′=TransEnc(E his ,T his ),
[0037] Among them, f t It is the aggregation result of semantic, emotional, thematic and personality information on the subject factor. t is the aggregation weight of the above information, P is the updated personality status node, f is a fusion node used to represent the preliminary aggregation result of emotion and topic information, and f e ′ and f t ′ It is the state of emotion and subject information after they influence each other; is a trainable linear layer used to map the concatenation results of multiple factors into the target latent space. is the bias term of the mapping;
[0038] 3.4.3) Update the emotional state node E using the following formula:
[0039] f e =x+g e ·p+(1-g e )·f,
[0040] E=TransEnc(f e ,E his ),
[0041]
[0042] Among them, f e It is the aggregation result of semantic, emotional, thematic and personality information on emotional factors. e is the aggregation weight of the above information;
[0043] 4) Inject the interaction factors into the response generation process to generate the target response Y.
[0044] 4.1) Splice the historical dialogue into the decoder to generate the hidden state of the original response without any factor-related information:
[0045] h ori,i =Decoder(E ori,<i ,E χ )
[0046] Among them, E ori,<i is the word embedding vector generated before time step i, E χ is the embedding vector corresponding to the historical dialogue;
[0047] 4.2) The original hidden state h ori,iFuse with each signal s separately to obtain the adapted hidden state h s,i :
[0048] w s =σ(W s [h ori,i ;s]+b s );
[0049] h s,i =w s ·s+(1-w s )·h ori,i ;
[0050] Among them, the signal s∈[P,T,E], w s is the fusion weight, W s is a trainable linear layer used to map the concatenation results of multiple factors into the target latent space, b s is the bias term of the mapping;
[0051] 4.3) Get the vocabulary y of the tth token t Distribution P(y t |y <t ,x);
[0052] P(y t |y <t ,χ)=softmax(Wh adp,t );
[0053] Among them, χ is the historical dialogue, h adp,t are the three adapted hidden states h at time step t P,t ,h T,t ,h E,t The mean of h, W is a trainable linear layer used to transform h adp,t Mapping to the vocabulary;
[0054] 4.4) According to vocabulary y t The target response Y is obtained by distribution.
[0055] According to the above scheme, in step 2.5), during the training process, e i is input into a linear layer and then subjected to a softmax operation to generate the sentiment category distribution P emo (e i ), and then by minimizing the sentiment category distribution P emo With the true label The cross entropy loss between them is used to optimize the model:
[0056]
[0057] According to the above scheme, in step 2.6), during the training process, t i is input into a linear layer and then subjected to a softmax operation to generate the sentiment category distribution P top (t i ), and then by minimizing the sentiment category distribution P top With the true label The cross entropy loss between them is used to optimize the model:
[0058]
[0059] According to the above scheme, in step 3.4), in order to ensure that the updated nodes contain corresponding emotion and topic signals during the training process, the topic state node T and the emotion state node E are input into the linear layer to obtain the corresponding probability distribution P emo (E) and P top (T), and through the cross entropy loss function and Constrain it:
[0060]
[0061] According to the above scheme, the cross entropy loss function is used in step 4.4) to constrain the generation process:
[0062]
[0063] The beneficial effects produced by the present invention are:
[0064] 1. The present invention generates dialogues by modeling complex multiple interactive factors, integrating the correlation and synergy between multiple factors to generate more human-style responses;
[0065] 2. The present invention provides a time dynamic method for jointly modeling each factor based on a collaboration graph to obtain multiple interactive factor signals. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0067] Figure 1 is a method flow chart of an embodiment of the present invention;
[0068] Figure 2 It is a schematic diagram of an embodiment of the present invention. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0070] The goal of the present invention is to dynamically perceive interactive factors such as emotions and topics in a conversation, integrate them into the generation process together with the speaker's personality, and generate responses that are semantically reasonable, emotionally appropriate, and consistent with the topic.
[0071] like Figure 1 and Figure 2 As shown, a dialogue generation method based on complex interaction factor modeling includes the following steps:
[0072] 1) Given a set of predefined personalities P and historical conversation sentences X;
[0073] 2) Use the encoder to encode the historical conversation sentences and predefined personalities to obtain various representations in the embedding space, including: personality representation sequence P pre , historical dialogue representation x, historical emotion representation sequence E his and historical topic representation sequence T his ; The details are as follows:
[0074] 2.1) For each personality p i ∈P, encode it separately to get the personality representation sequence P pre ={p1,p2,…,p n};
[0075] 2.2) For each utterance x in the historical dialogue i ∈X is encoded to get the sentence representation
[0076] 2.3) X enc Input into the bidirectional RNN to extract semantic information and obtain the contained semantic information
[0077] X s =RNN semantic (X enc );
[0078] 2.4) Calculate X s The weight of each semantic representation in , and perform weighted summation to obtain the historical dialogue representation x:
[0079]
[0080] The score function is as follows:
[0081]
[0082] Among them, W1 and W2 are trainable linear layers, and b1 is the bias term;
[0083] 2.5) X enc Input into another bidirectional RNN to extract the specific historical sentiment sequence representation E his ={e1,e2,…,e m}:
[0084] E his =RNN emotion (X enc );
[0085] 2.6) X enc Input into another bidirectional RNN to extract a specific historical topic sequence representation T his ={t1,t2,…,t m}:
[0086] T his =RNN topic (X enc );
[0087] During the training process, e i is input into a linear layer and then subjected to a softmax operation to generate the sentiment category distribution P emo (e i ), and then by minimizing the sentiment category distribution P emo With the true label The cross entropy loss between them is used to optimize the model:
[0088]
[0089] t i is input into a linear layer and then subjected to a softmax operation to generate the sentiment category distribution P top (t i ), and then by minimizing the sentiment category distribution P top With the true label The cross entropy loss between them is used to optimize the model:
[0090]
[0091] 3) Jointly model personality, emotion, and topic interaction factors;
[0092] Interaction factors influence each other and change over time, and collaboration graphs are used to jointly model the temporal dynamics of each factor in each round of dialogue;
[0093] The collaboration diagram includes the following three types of nodes;
[0094] 3.1) Representing sequence P according to personality pre , historical dialogue representation x, historical emotion representation sequence E his and historical topic representation sequence T his , get the individual node p in the collaboration graph i , Topic history node t i , emotional history node e i and semantic node x;
[0095] 3.2) Establish personality state node P, emotion state node E and topic state node T in the collaboration graph;
[0096] 3.3) Establish fusion node;
[0097] 3.4) Obtain the dependency relationship between nodes;
[0098] 3.4.1) The personality displayed by the speaker in the conversation should be consistent with his own personality and highly relevant to the current conversation. Therefore, the personality state node P in the collaboration graph is based on the personality node P pre and semantic node x, that is, P(P|x,P pre ):
[0099] P = TransEnc(x,P pre ),
[0100] TransEnc is a single-layer Transformer encoder; we use a single-layer Transformer encoder as a factor fusion module to handle the dependencies between different factors. The factor fusion module takes two factors as input: the first factor is used as the query in the attention mechanism, and the second factor is used as both the key and the value.
[0101] 3.4.2) Update the topic state node T using the following formula:
[0102] f t =x+g t ·P+(1-g t )·f,
[0103] T=TransEnc(f t ,T his ),
[0104]
[0105] Where f = TransEnc(x,W3[f e ′ ;f t ′]+b3);
[0106] f e ′ =TransEnc(T his ,E his ),
[0107] f t ′ =TransEnc(E his ,T his ),
[0108] Among them, f t It is the aggregation result of semantic, emotional, thematic and personality information on the subject factor. t is the aggregation weight of the above information, P is the updated personality status node, f is a fusion node used to represent the preliminary aggregation result of emotion and topic information, and f e ′ and f t ′ It is the state of emotion and subject information after they influence each other; is a trainable linear layer used to map the concatenation results of multiple factors into the target latent space. is the bias term of the mapping;
[0109] 3.4.3) Update the emotional state node E using the following formula:
[0110] f e =x+g e ·p+(1-g e )·f,
[0111] E=TransEnc(f e ,E his ),
[0112]
[0113] Among them, f e It is the aggregation result of semantic, emotional, thematic and personality information on emotional factors. e is the aggregation weight of the above information;
[0114] According to the above relationship, we can obtain a collaboration diagram, such as Figure 2 shown.
[0115] To ensure that the updated nodes contain the corresponding sentiment and topic signals, we input them into the linear layer to obtain the corresponding probability distribution P emo (E) and P top (T), and constrain them through the cross entropy loss function:
[0116]
[0117] 4) Inject the interaction factor related signals into the response generation process to generate the target response Y.
[0118] 4.1) Splice the historical dialogue into the decoder to generate the hidden state of the original response without any factor-related information:
[0119] h ori,i =Decoder(E ori,<i ,E x )
[0120] Among them, E ori,<i is the word embedding vector generated before time step i, E x is the embedding vector corresponding to the historical dialogue, obtained through step 2);
[0121] 4.2) In order to make the generated response more consistent with the speaker’s situation, we adopt a unified and scalable method to modify the original hidden state of the response by taking into account various factors related to the signal. ori,i Fuse with each signal s separately to obtain the adapted hidden state h s,i :
[0122] w s =σ(W s [h ori,i ;s]+b s );
[0123] h s,i =w s ·s+(1-w s )·h ori,i ;
[0124] Among them, the signal s∈[P,T,E], w s is the fusion weight, W s is a trainable linear layer used to map the concatenation results of multiple factors into the target latent space, b s is the bias term of the mapping,
[0125] 4.3) Get the vocabulary y of the tth token t Distribution P(y t |y <t ,x);
[0126] P(y t |y <t ,x)=softmax(Wh adp,t );
[0127] Among them, x is the historical dialogue, hadp,t are the three adapted hidden states h at time step t P,t ,h T,t ,h E,t The mean of h, W is a trainable linear layer used to transform h adp,t Mapping to the vocabulary;
[0128] 4.4) According to vocabulary y t The target response Y is obtained by distribution.
[0129] Use the cross entropy loss function to constrain the generation process:
[0130]
[0131] Finally, the model parameters are optimized using the following joint optimization objective:
[0132]
[0133] Effect verification:
[0134] We evaluate all methods on the Persona-Chat dataset and the Synthetic-Persona-Chat dataset. The results are shown in Table 1.
[0135] Table 1 Index comparison
[0136]
[0137] On the Synthetic-Persona-Chat and Persona-Chat datasets, the present invention achieves the best results in most metrics, demonstrating the effectiveness of the model. Thanks to the comprehensive consideration of multiple interaction factors during the generation process, the present invention is able to generate high-quality responses with low perplexity. In addition, the improvement in the Distinct metric highlights the advantage of CoMIF in generating more informative responses.
[0138] It is worth noting that compared with the base model GPT-2, the invention incorporates additional factors when generating responses, resulting in more diverse and contextually relevant outputs with almost the same perplexity. This shows that the performance improvement of the invention is due to considering interaction factors rather than just optimizing the model structure.
[0139] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all these improvements and changes should fall within the scope of protection of the appended claims of the present invention.
Claims
1. A method for generating dialogue based on modeling of complex interactive factors, characterized in that: The following steps are involved: 1) Given a set of predefined personalities P and historical conversation sentences X; 2) Use the encoder to encode the historical conversation sentences and predefined personalities to obtain various representations in the embedding space, including: personality representation sequence P pre , historical dialogue representation x, historical emotion representation sequence E his and historical topic representation sequence T his ; 3) Joint modeling of interaction factors; Interaction factors influence each other and change over time, and collaboration graphs are used to jointly model the temporal dynamics of each factor in each round of dialogue; 3.1) Representing sequence P according to personality pre , historical dialogue representation x, historical emotion representation sequence E his and historical topic representation sequence T his , get the personality node p i , Topic history node t i , emotional history node e i and semantic node x; 3.2) Establish personality state node P, emotion state node E and topic state node T; 3.3) Establish fusion node; 3.4) Obtain the dependency relationship between nodes; 4) Inject the interaction factors into the response generation process to generate the target response Y.
2. The method for generating a dialogue based on complex interaction factor modeling according to claim 1, characterized in that: The specific steps of step 2) are as follows: 2.1) For each personality p i ∈P, encode it and get the personality representation sequence P pre ={p1, p2, ..., p n }; 2.2) For each utterance x in the historical dialogue i ∈X is encoded to get the sentence representation 2.3) For X enc Extract semantic information and obtain the contained semantic information 2.4) Calculate X s The weight of each semantic representation in , and perform weighted summation to obtain the historical dialogue representation x: 2.5) X enc Input into a bidirectional RNN to extract the specific historical sentiment sequence representation E his ={e1, e2, ..., e m }: HAVE BEEN his =RNN emotion (X enc ); 2.6) X enc Input into another bidirectional RNN to extract a specific historical topic sequence representation T his ={t1, t2, ..., t m }: T his =RNN topic (X enc )。 3. The method for generating a dialogue based on complex interaction factor modeling according to claim 2, characterized in that: In step 2.3), X enc Input into the bidirectional RNN to extract semantic information; X s =RNN semantic (X enc ); Among them, RNN semantic It is a bidirectional two-layer RNN model for modeling semantic information.
4. The method for generating a dialogue based on complex interaction factor modeling according to claim 2, characterized in that: In step 2.3), calculate X s The weight of each semantic representation in , and perform weighted summation to obtain the historical dialogue representation x: The score function is as follows: Among them, W1 and W2 are trainable linear layers, and b1 is the bias term.
5. The method for generating a dialogue based on complex interaction factor modeling according to claim 1, characterized in that: In step 3.4), the details are as follows: 3.4.1) The individual state node P in the collaboration graph is based on the individual node P pre and update the semantic node x, that is, P(P| x , P pre ): P=TransEnc(x,P pre ), Among them, TransEnc is a single-layer Transformer encoder; 3.4.2) Update the topic state node T using the following formula: f t =x+g t ·P+(1-g t )·f, T=TransEnc(f t ,T his ), g t =σ(W gt [x;P;f]+b gt ); Where f = TransEnc(x, W3[f′ e ; f′ t ]+b3); f′ e =TransEnc(T his ,E his ), f′ t =TransEnc(E his ,T his ), Among them, f t It is the aggregation result of semantic, emotional, thematic and personality information on the subject factor. t is the aggregation weight of the above information, P is the updated personality status node, f is a fusion node used to represent the preliminary aggregation result of emotion and topic information, and f′ e and f′ t It is the state of emotion and subject information after they influence each other; is a trainable linear layer used to map the concatenation results of multiple factors into the target latent space. is the bias term of the mapping; 3.4.3) Update the emotional state node E using the following formula: f e =x+g e ·p+(1-g e )·f, E=TransEnc(f e ,E his ), g e =σ(W ge [x;P;f]+b gte ); Among them, f e It is the aggregation result of semantic, emotional, thematic and personality information on emotional factors. e is the aggregation weight of the above information.
6. The method for generating a dialogue based on complex interaction factor modeling according to claim 1, characterized in that: In step 4), the details are as follows: 4.1) Splice the historical dialogue into the decoder to generate the hidden state of the original response without any factor-related information: Among them, E ori,<i is the word embedding vector generated before time step i, is the embedding vector corresponding to the historical dialogue; 4.2) The original hidden state h ori,i Fuse with each signal s separately to obtain the adapted hidden state h s,i : w s =σ(W s [h ori,i ;s]+b s ); h s,i =w s ·s+(1-w s )·h ori,i ; Among them, the signal s∈[P, T, E], w s is the fusion weight, W s is a trainable linear layer used to map the concatenation results of multiple factors into the target latent space, b s is the bias term of the mapping; 4.3) Get the vocabulary y of the tth token t distributed in, For historical dialogue, h adp,t are the three adapted hidden states h at time step t P,t ,h T,t ,h E,t The mean of h, W is a trainable linear layer used to transform h adp,t Mapping to the vocabulary; 4.4) According to vocabulary y t The target response Y is obtained by distribution.
7. The method for generating a dialogue based on complex interaction factor modeling according to claim 2, characterized in that: In step 2.5), during the training process, e i is input into a linear layer and then subjected to a softmax operation to generate the sentiment category distribution P emo (e i ), and then by minimizing the sentiment category distribution P emo With the true label The cross entropy loss between them is used to optimize the model: In step 2.6), during the training process, t i is input into a linear layer and then subjected to a softmax operation to generate the sentiment category distribution P top (t i ), and then by minimizing the sentiment category distribution P top With the true label The cross entropy loss between them is used to optimize the model:
8. The method for generating a dialogue based on complex interaction factor modeling according to claim 2, characterized in that: In the step 4.4), a cross entropy loss function is used to constrain the generation process:
9. An electronic device, characterized in that: include: one or more processors; as well as a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are caused to perform the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.