An AI digital human-based interactive response generation method and system

By coding the current dialogue text and historical dialogue text of AI digital people, more accurate and natural responses are generated, and the problem of unreasonable expression of AI digital people in interaction is solved and the user experience is improved.

CN119719317BActive Publication Date: 2025-07-04JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510221226.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-04
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

Existing AI digital people cannot fully reflect personalized emotional expressions in interactive replies, resulting in poor user experience.

Method used

By obtaining the user's current dialogue text and historical dialogue text, converting it into word vectors, determining emotional characteristics and auxiliary knowledge characteristics, and coding processing, constructing and solving equations to generate final replies, improving the accuracy of emotion recognition and anthropomorphism of replies.

Benefits of technology

It improves the accuracy and speed of AI digital people's emotional recognition and reply generation, and enhances its sense of reality and nature.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719317B_ABST
    Figure CN119719317B_ABST
Patent Text Reader

Abstract

The present invention provides an interactive response generation method and system based on an AI digital human. The method includes obtaining the current conversation text and historical conversation text of the user, and converting the current conversation text into a current word vector; determining an emotional feature based on the current word vector and the historical conversation text; obtaining external common sense information and converting the external common sense information into an auxiliary knowledge graph, and determining an auxiliary knowledge feature based on the auxiliary knowledge graph, the current word vector and the historical conversation text; respectively encoding the emotional feature and the auxiliary knowledge feature to obtain an emotional encoding feature and a knowledge encoding feature, and determining the final response of the AI digital human based on the emotional encoding feature and the knowledge encoding feature. The present invention can improve the accuracy and speed of emotion recognition, and at the same time can also improve the realism, naturalness and anthropomorphism of the AI digital human.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of AI interaction, and particularly relates to an interaction response generation method and system based on an AI digital human. Background Art

[0002] An AI digital human refers to a realistic virtual character created through technologies such as artificial intelligence technology, virtual reality technology, and computer graphics. It is used to simulate aspects such as human language, behavior, emotion, and thinking, making it appear to have a real appearance and behavior. AI digital humans can be applied in fields such as games, animations, movies, virtual reality, customer service, education, and healthcare.

[0003] In the actual usage process, interaction response is an important use of AI digital humans. When a user interacts with an AI digital human, it is usually necessary for the AI digital human to fully understand the user's emotion or mood in combination with the context of the interaction content and the current input content, so as to make the response content of the AI digital human more anthropomorphic. However, for existing AI digital humans, they usually have problems such as unreasonable expression of emotion or mood and inability to fully reflect personalization during the conversation, thereby affecting the actual experience of users. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides an interaction response generation method and system based on an AI digital human to solve the technical problems in the prior art.

[0005] On the one hand, the present invention provides the following technical solution. An interaction response generation method based on an AI digital human includes:

[0006] Obtain the current conversation text and historical conversation text of the user, and convert the current conversation text into a current word vector;

[0007] Determine an emotion feature based on the current word vector and the historical conversation text;

[0008] Obtain external common sense information and convert the external common sense information into an auxiliary knowledge graph, and determine an auxiliary knowledge feature based on the auxiliary knowledge graph, the current word vector, and the historical conversation text;

[0009] Encode the emotion feature and the auxiliary knowledge feature respectively to obtain an emotion encoding feature and a knowledge encoding feature, and determine the final response of the AI digital human based on the emotion encoding feature and the knowledge encoding feature;

[0010] The step of encoding the emotion feature and the auxiliary knowledge feature respectively to obtain an emotion encoding feature and a knowledge encoding feature includes:

[0011] Encode the emotional feature and the auxiliary knowledge feature respectively to obtain an emotional embedding and an auxiliary knowledge embedding ;

[0012] Based on the emotional embedding and the auxiliary knowledge embedding construct a first solution equation and a second solution equation:

[0013] ; ;

[0014] wherein, , respectively represent an emotional coding feature and a knowledge coding feature, , respectively represent the infinitesimal change amounts of the emotional embedding and the auxiliary knowledge embedding , respectively represent the weight vectors corresponding to the emotional embedding and the auxiliary knowledge embedding , respectively represent the bias vectors corresponding to the emotional embedding and the auxiliary knowledge embedding ;

[0015] Solve the first solution equation and the second solution equation respectively to obtain the emotional coding feature and the knowledge coding feature .

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention first obtains the current conversation text and the historical conversation text of the user, and converts the current conversation text into a current word vector; then determines the emotional feature based on the current word vector and the historical conversation text; obtains the external common sense information and converts the external common sense information into an auxiliary knowledge graph, and determines the auxiliary knowledge feature based on the auxiliary knowledge graph, the current word vector and the historical conversation text; then encodes the emotional feature and the auxiliary knowledge feature respectively to obtain the emotional coding feature and the knowledge coding feature, and determines the final reply of the AI digital human based on the emotional coding feature and the knowledge coding feature. When recognizing the emotion of the user, the present invention can enhance the emotion in the current conversation text, thereby improving the accuracy and speed of emotion recognition. At the same time, by determining the auxiliary knowledge feature, the reply of the AI digital human can be guided and assisted by knowledge, avoiding the situation of inaccurate or chaotic reply expression. Finally, the final reply is generated by combining the emotional feature and the auxiliary knowledge feature, thereby improving the realism, naturalness and anthropomorphism of the AI digital human.

[0017] Preferably, the step of determining the emotional feature based on the current word vector and the historical conversation text includes:

[0018] Obtain an emotion dataset and construct an emotion knowledge graph according to the emotion dataset. Expand the historical dialogue text into a historical word sequence and construct the historical word sequence into a historical knowledge graph according to the construction method of the emotion knowledge graph;

[0019] Extract keywords from the current word vector and determine the current relevant knowledge graph corresponding to each keyword;

[0020] Determine semantic features based on the current relevant knowledge graph;

[0021] Extract personalized information from the historical dialogue text, form a personalized matrix according to the quantity and sequence length of the personalized information, and input the personalized information into a trained preset prediction model for prediction to obtain predicted personalized information;

[0022] Symbolically process the historical dialogue text and the predicted personalized information respectively to obtain processed text and processed information;

[0023] Concatenate the processed text and the processed information to obtain a first concatenated feature, and perform encoding processing on the first concatenated feature and the personalized matrix respectively to obtain a first encoded feature and a second encoded feature;

[0024] Determine emotion features based on the first encoded feature, the second encoded feature, and the semantic features.

[0025] Preferably, the step of extracting keywords from the current word vector and determining the current relevant knowledge graph corresponding to each keyword includes:

[0026] Extract keywords from the current word vector to obtain a keyword group, and extract implicit triples in the emotion knowledge graph that include at least one keyword in the keyword group to obtain a triple set;

[0027] Determine a number of implicit emotion subgraphs based on the triple set, convert the number of implicit emotion subgraphs into natural sentences, and convert the current dialogue text into a first word embedding;

[0028] Convert the natural sentence into a second word embedding and calculate the similarity between the first word embedding and the second word embedding;

[0029] Retain the implicit emotion subgraphs corresponding to a similarity greater than a preset similarity to obtain the current relevant knowledge graph corresponding to each keyword 。

[0030] Preferably, the step of determining semantic features based on the current relevant knowledge graph includes:​

[0031] Based on the current relevant knowledge graph Determine the key sentences :

[0032] ;

[0033] In the formula, represents the concatenation operation, represents the -th keyword embedding;

[0034] Input the key sentences into a preset sentence encoding model for encoding to obtain an encoding matrix , and determine the first pooling feature and the second pooling feature :

[0035] ;

[0036] ;

[0037] In the formula, represents the element in the -th row and -th column of the encoding matrix;

[0038] Based on the historical knowledge graph, the first pooling feature , and the second pooling feature determine the emotion vector :

[0039] ;

[0040] ;

[0041] In the formula, represents the global feature, represents layer normalization, represents the vertex in the historical knowledge graph, represents 's adjacent matrix, represents the -th self-attention head, represents the weight of the -th self-attention head, represents self-attention heads forming a self-attention layer;

[0042] Based on the emotion vector calculate the semantic feature :

[0043] ;

[0044] In the formula, Represent the first and second weights respectively, Represent the first and second deviations respectively, Indicates the error value.

[0045] Preferably, the first coding feature, the second coding feature, and the semantic feature The steps to characterize emotions include:

[0046] Performing average pooling processing on the second coding feature and concatenating the second coding feature after the average pooling processing with the first coding feature to obtain a second concatenated feature, and sequentially performing encoding and decoding processing on the second concatenated feature to obtain a personalized feature;

[0047] The personalized feature and the semantic feature Connect to get personalized semantic features , based on the personalized semantic features Identify emotional sentence features :

[0048] ;

[0049] ;

[0050] In the formula, , Respectively represent The hidden state corresponding to the layer decoder, , , Represent the first, second, and third learnable parameter matrices respectively, For the dimension, For fusion operation, Indicates Keywords;

[0051] The emotional sentence features Input the trained preset emotion recognition model for recognition to obtain emotional features.

[0052] Preferably, the step of determining auxiliary knowledge features based on the auxiliary knowledge graph, the current word vector and the historical conversation text includes:

[0053] Extracting an auxiliary knowledge subgraph corresponding to the historical conversation text from the auxiliary knowledge graph and using it as a reference subgraph;

[0054] Each node in the reference sub - graph is input into a graph convolutional neural network for node embedding layer calculation and relationship embedding layer calculation to obtain node embeddings and relationship embeddings :

[0055] ;

[0056] ;

[0057] wherein, is an activation function, represents the neighbor sub - graph of node in the reference sub - graph, , , respectively represent the first, second, and third learnable weight matrices of the th layer in the graph convolutional neural network, respectively represent the node embeddings of the th layer, respectively represent the relationship embeddings of the th layer;

[0058] Based on the node embeddings and relationship embeddings calculate the evaluation scores of each node in the reference sub - graph, and select the node with the maximum evaluation score for access connection for each node in the reference sub - graph to obtain an updated reference sub - graph:

[0059] ;

[0060] wherein, is an adjustment coefficient, represents the evaluation score of the node that has been accessed and connected, is a non - linear activation function, represents the node embedding of node in the th layer, represents the fourth learnable weight matrix of the th layer in the graph convolutional neural network, represents the hidden state of the th layer in the graph convolutional neural network at the th time step;

[0061] Extract the first semantic embedding of the nodes in the historical knowledge graph corresponding to the historical dialogue text, the second semantic embedding of the node connection edges, and extract the third semantic embedding of the nodes in the updated reference sub - graph 4th semantic embedding of node connection edges Calculate the final knowledge feature based on the 1st semantic embedding, 2nd semantic embedding, 3rd semantic embedding, and 4th semantic embedding :

[0062] ;

[0063] In the formula, is a connection operation;

[0064] Determine the auxiliary knowledge feature based on the said final knowledge feature : :

[0065] ;

[0066] In the formula, is probability calculation, represents a set in is the updated reference subgraph, is the historical dialogue text.

[0067] Preferably, the step of determining the final response of the AI digital human based on the emotion coding feature and the knowledge coding feature includes:

[0068] Encode the current word vector to obtain the current word representation, convert the triples in the auxiliary knowledge graph into natural sentences and encode the natural sentences to obtain the natural sentence representation;

[0069] Based on the current word representation and the natural sentence representation Determine the emotion dialogue representation :

[0070] ;

[0071] In the formula, is a connection operation, is a linear weight matrix;

[0072] Extract historical keywords from the historical dialogue text and obtain the historical keyword embeddings of the historical keywords, encode the historical keyword embeddings to obtain historical key features, calculate the similarity between the historical key features and each triple in the auxiliary knowledge graph, sort each triple in the auxiliary knowledge graph in descending order according to the similarity, and select the top several triples as the sorted knowledge group;

[0073] Determine whether any historical keyword is the same as the tail entity of the triple in the permutation knowledge group. If any historical keyword is the same as the tail entity of the triple in the permutation knowledge group, store the historical keyword, the corresponding word of the triple, and the tail entity in the final keyword set, and calculate the attention weight of the final keyword set to obtain the key attention weight vector ;

[0074] Based on the key attention weight vector , emotional dialogue representation Determine the final reply of the AI digital human

[0075] Preferably, the step of determining the final reply of the AI digital human based on the key attention weight vector , emotional dialogue representation includes:

[0076] Based on the key attention weight vector , emotional dialogue representation Calculate the output state :

[0077] ;

[0078] In the formula, represents the transformer decoder, is a linear transformation, is the embedding corresponding to the current word vector, is the preset token embedding, is the generated token embedding, is the concatenation operation;

[0079] Based on the output state Determine the word probability distribution , and generate the final reply of the AI digital human based on the word probability distribution:

[0080] ;

[0081] In the formula, represents the pointer-generator network, , respectively represent the emotional encoding feature and the knowledge encoding feature

[0082] In a second aspect, the present invention provides the following technical solution, an interactive reply generation system based on an AI digital human, the system adopts the above-mentioned interactive reply generation method based on an AI digital human, and the system includes:

[0083] A conversion module, configured to obtain the current conversation text and historical conversation text of the user, and convert the current conversation text into a current word vector;

[0084] An emotion module, configured to determine an emotion feature based on the current word vector and the historical conversation text;

[0085] An auxiliary module, configured to obtain external common sense information and convert the external common sense information into an auxiliary knowledge graph, and determine an auxiliary knowledge feature based on the auxiliary knowledge graph, the current word vector, and the historical conversation text;

[0086] A reply module, configured to encode the emotion feature and the auxiliary knowledge feature respectively to obtain an emotion encoding feature and a knowledge encoding feature, and determine the final reply of the AI digital human based on the emotion encoding feature and the knowledge encoding feature.

[0087] In a third aspect, the present invention provides the following technical solution. A computer includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the interactive reply generation method based on the AI digital human as described above is implemented.

[0088] In a fourth aspect, the present invention provides the following technical solution. A storage medium stores a computer program, and when the computer program is executed by a processor, the interactive reply generation method based on the AI digital human as described above is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0090] Figure 1 It is a flowchart of an interactive reply generation method based on an AI digital human provided in Embodiment 1 of the present invention;

[0091] Figure 2 It is a structural block diagram of an interactive reply generation system based on an AI digital human provided in Embodiment 2 of the present invention;

[0092] Figure 3 It is a schematic hardware structure diagram of a computer provided in another embodiment of the present invention.

[0093] The following will further describe the embodiments of the present invention with reference to the drawings. DETAILED DESCRIPTION OF THE INVENTION

[0094] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the embodiments of the present invention, and should not be construed as limiting the present invention.

[0095] In the description of the embodiments of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the embodiments of the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present invention.

[0096] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, the meaning of "a plurality" is two or more, unless otherwise specifically defined.

[0097] In the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed", etc. should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific circumstances.

[0098] Embodiment 1

[0099] In Embodiment 1 of the present invention, as Figure 1 shown, an interactive response generation method based on an AI digital human includes:

[0100] S1. Obtain the current conversation text and historical conversation text of the user, and convert the current conversation text into a current word vector;

[0101] Specifically, the current conversation text is the content input by the user in the current round, and the historical conversation text is the content input by the user in the previous several rounds and the content replied by the AI digital human. At the same time, the current conversation text can be converted into a current word vector through the BERT model in the prior art.

[0102] S2. Determine the emotional feature based on the current word vector and the historical conversation text;

[0103] Among them, the step S2 includes:

[0104] S21. Obtain an emotion dataset and construct an emotion knowledge graph according to the emotion dataset. Expand the historical conversation text into a historical word sequence and construct the historical word sequence into a historical knowledge graph according to the construction method of the emotion knowledge graph;

[0105] Specifically, the emotion dataset here can be interactively constructed through existing common sense knowledge graphs and emotion lexicons.

[0106] S22. Extract the keywords in the current word vector and determine the current relevant knowledge graph corresponding to each keyword;

[0107] Among them, the step S22 includes:

[0108] S221. Extract the keywords in the current word vector to obtain a keyword group. Extract the implicit triples in the emotion knowledge graph that include at least one keyword in the keyword group to obtain a triple set;

[0109] Specifically, the keywords here include verbs, nouns, adjectives, adverbs, etc. And for the emotion graph, it includes several triples, and there are corresponding relationships between the triples. At the same time, for the triples, there will be some triples that are not easy to express symbolically, which are the implicit triples in this step. In the subsequent steps, the matching with the current conversation text can be realized by introducing the implicit triples.

[0110] S222. Determine several implicit emotion subgraphs based on the triple set, convert the several implicit emotion subgraphs into natural sentences, and convert the current conversation text into a first word embedding;

[0111] Specifically, the natural sentence here is a sentence in natural language, and the first word embedding can be realized by the BERT model.

[0112] S223. Convert the natural sentence into a second word embedding and calculate the similarity between the first word embedding and the second word embedding;

[0113] Specifically, the second word embedding can also be realized by the BERT model, and the similarity here is specifically the cosine similarity.

[0114] S224. Retain the implicit emotion subgraphs corresponding to the similarity greater than the preset similarity to obtain the current relevant knowledge graph corresponding to each keyword ; ;

[0115] Specifically, this step is a process of matching and screening. When the similarity is greater than the preset similarity, it indicates that the implicit emotion subgraph is relevant to the current dialogue text. When the similarity is not greater than the preset similarity, it indicates that it is not relevant to the current dialogue text.

[0116] S23. Determine semantic features based on the current relevant knowledge graph;

[0117] Among them, the step S23 includes:

[0118] S231. Based on the current relevant knowledge graph determine key sentences :

[0119] ;

[0120] In the formula, represents the concatenation operation, represents the th keyword embedding.

[0121] S232. Input the key sentences into a preset sentence encoding model for encoding to obtain an encoding matrix , and determine the first pooling feature and the second pooling feature :

[0122] ;

[0123] ;

[0124] In the formula, represents the element in the th row and th column of the encoding matrix;

[0125] Specifically, the preset sentence encoding model here is a bidirectional LSTM model.

[0126] S233. Determine the emotion vector based on the historical knowledge graph, the first pooling feature , and the second pooling feature :

[0127] ;

[0128] ;

[0129] In the formula, represents the global feature, represents layer normalization, Represents a vertex in the historical knowledge graph, represents the adjacent matrix of, represents the th self-attention head, represents the weight of the th self-attention head, represents the self-attention layer composed of self-attention heads.

[0130] S234. Calculate the semantic feature based on the emotion vector :

[0131] ;

[0132] In the formula, respectively represent the first and second weights, respectively represent the first and second biases, represents the error value.

[0133] S24. Extract personalized information from the historical dialogue text, form a personalized matrix according to the quantity and sequence length of the personalized information, and input the personalized information into the trained preset prediction model for prediction to obtain predicted personalized information;

[0134] Specifically, the personalized information here includes hobbies, age, gender information, etc. contained in each dialogue text, and the preset prediction model is specifically the SPP model.

[0135] S25. Symbolically process the historical dialogue text and the predicted personalized information respectively to obtain processed text and processed information.

[0136] S26. Concatenate the processed text and the processed information to obtain a first concatenated feature, and perform encoding processing on the first concatenated feature and the personalized matrix respectively to obtain a first encoded feature and a second encoded feature;

[0137] Specifically, the encoding process here can be implemented by a transformer encoder.

[0138] S27. Determine the emotion feature based on the first encoded feature, the second encoded feature, and the semantic feature;

[0139] Among them, the step S27 includes:

[0140] S271. Perform average pooling on the second encoded feature and concatenate the second encoded feature after average pooling with the first encoded feature to obtain a second concatenated feature, and perform encoding and decoding on the second concatenated feature in sequence to obtain a personalized feature;

[0141] Specifically, in step S271, perform average pooling on the second encoded feature in the 0th dimension, and then concatenate the second encoded feature after average pooling with the first encoded feature in the 1st dimension to obtain a second concatenated feature. The encoding and decoding processes here can be implemented by a Transformer encoder and a Transformer decoder.

[0142] S272. Connect the personalized feature with the semantic feature to obtain a personalized semantic feature , and determine an emotional sentence feature based on the personalized semantic feature :

[0143] ;

[0144] ;

[0145] wherein , respectively represent the hidden states corresponding to the th layer of the decoder, , , respectively represent the first, second, and third learnable parameter matrices, is the dimension, is the fusion operation, represents the th keyword;

[0146] Specifically, in each Transformer network, the self-attention mechanism is used to aggregate the output from the previous layer. Therefore, the first, second, and third learnable parameter matrices can be updated and learned during the aggregation process.

[0147] S273. Input the emotional sentence feature into a trained preset emotion recognition model for recognition to obtain an emotion feature;

[0148] Specifically, the preset emotion recognition model here is the DuaIGATS model;

[0149] It should be noted that in the above step S2, the corresponding emotion in the current dialogue text is enhanced through a series of processing procedures, and then the implicit emotion in the sentence is recognized by the model to improve the recognition accuracy and speed.

[0150] S3. Obtain external common sense information and convert the external common sense information into an auxiliary knowledge graph, and determine auxiliary knowledge features based on the auxiliary knowledge graph, the current word vector, and the historical dialogue text;

[0151] Among them, the auxiliary knowledge graph here is specifically a common sense knowledge graph.

[0152] Among them, the step S3 includes:

[0153] S31. Extract an auxiliary knowledge subgraph corresponding to the historical dialogue text from the auxiliary knowledge graph and use it as a reference subgraph.

[0154] S32. Input each node in the reference subgraph into a graph convolutional neural network for node embedding layer calculation and relationship embedding layer calculation to obtain node embedding and relationship embedding :

[0155] ;

[0156] ;

[0157] In the formula, is an activation function, represents the neighbor subgraph of node in the reference subgraph, , , respectively represent the first, second, and third learnable weight matrices of the th layer in the graph convolutional neural network, respectively represent the node embeddings of the th layer, respectively represent the relationship embeddings of the th layer;

[0158] Specifically, the neighbor subgraph here is a graph composed of the neighbor nodes of node and the connection relationships between the neighbor nodes. At the same time, the initial node embedding here is the embedding corresponding to the node in the reference subgraph, and the initial relationship embedding is the embedding corresponding to the connection edge of the node in the reference subgraph.

[0159] S33. Calculate the evaluation score of each node in the reference subgraph based on the node embedding and the relationship embedding , select the evaluation score for each node in the reference sub-graph for the node with the largest value for access connection to obtain an updated reference sub-graph:

[0160] ;

[0161] In the formula, is an adjustment coefficient, represents the evaluation score of the node that has been accessed and connected , is a non-linear activation function, represents the node at the layer node embedding, represents the fourth learnable weight matrix in the layer of the graph convolutional neural network, represents the layer in the graph convolutional neural network at the th time step hidden state;

[0162] Specifically, the above process is a problem of node connection search. According to the evaluation score of each node, the next node to be connected by the current node is determined. Repeating the above process can obtain a new node connection relationship and distribution relationship, that is, an updated reference sub-graph.

[0163] S34. Extract the first semantic embedding of the nodes in the historical knowledge graph corresponding to the historical dialogue text , the second semantic embedding of the node connection edges and extract the third semantic embedding of the nodes in the updated reference sub-graph , the fourth semantic embedding of the node connection edges , and calculate the final knowledge feature based on the first semantic embedding, the second semantic embedding, the third semantic embedding, and the fourth semantic embedding :

[0164] ;

[0165] In the formula, is a connection operation;

[0166] Specifically, after the above traversal connection of the nodes, the connection process can be regarded as a search action of the nodes. Through different search actions, the relationship between the nodes and the connection edges and the relationship between the semantic embeddings and the nodes and the connection edges are reflected. And through the above search, different knowledge features can be selected as auxiliary knowledge features.

[0167] S35. Determine the auxiliary knowledge feature based on the final knowledge feature :

[0168] ;

[0169] Wherein, is for probability calculation, represents a set in is the updated reference subgraph, is the historical dialogue text;

[0170] Specifically, the purpose of step S35 is to select the subgraph with the highest probability from the searched subgraphs as the auxiliary knowledge feature.

[0171] S4. Encode the emotion feature and the auxiliary knowledge feature respectively to obtain an emotion encoding feature and a knowledge encoding feature, and determine the final response of the AI digital human based on the emotion encoding feature and the knowledge encoding feature;

[0172] Among them, the steps of encoding the emotion feature and the auxiliary knowledge feature respectively to obtain an emotion encoding feature and a knowledge encoding feature include:

[0173] S411. Encode the emotion feature and the auxiliary knowledge feature respectively to obtain an emotion embedding and an auxiliary knowledge embedding ;

[0174] Specifically, the encoding process can be implemented by a transformer encoder.

[0175] S412. Based on the emotion embedding and the auxiliary knowledge embedding construct a first solution equation and a second solution equation:

[0176] ; ;

[0177] Wherein, , respectively represent the emotion encoding feature and the knowledge encoding feature, , respectively represent the emotion embedding , the auxiliary knowledge embedding a small change amount of, respectively represent the emotion embedding , the auxiliary knowledge embedding corresponding weight vectors, respectively represent the emotion embedding , the auxiliary knowledge embedding corresponding bias vectors.

[0178] S413. Solve the first solution equation and the second solution equation respectively to obtain the emotion coding features and the knowledge coding features .

[0179] Among them, the step of determining the final reply of the AI digital human based on the emotion coding features and the knowledge coding features includes:

[0180] S421. Encode the current word vector to obtain the current word representation, convert the triples in the auxiliary knowledge graph into natural sentences and encode the natural sentences to obtain the natural sentence representation.

[0181] S422. Based on the current word representation and the natural sentence representation determine the emotion dialogue representation :

[0182] ;

[0183] In the formula, is the connection operation, is the linear weight matrix.

[0184] S423. Extract historical keywords from the historical dialogue text and obtain the historical keyword embeddings of the historical keywords, encode the historical keyword embeddings to obtain historical key features, calculate the similarity between the historical key features and each triple in the auxiliary knowledge graph, and sort each triple in the auxiliary knowledge graph in descending order according to the similarity and select the top several triples as the sorted knowledge group;

[0185] Specifically, the similarity here is the cosine similarity.

[0186] S424. Determine whether any historical keyword is the same as the tail entity of the triple in the sorted knowledge group. If any historical keyword is the same as the tail entity of the triple in the sorted knowledge group, store the historical keyword, the word corresponding to the triple, and the tail entity in the final keyword set, and calculate the attention weight of the final keyword set to obtain the key attention weight vector ;

[0187] Specifically, if any historical keyword is the same as the tail entity of the triple in the sorted knowledge group, there is a unique path between the historical keyword and the word corresponding to the triple, and the historical keyword, the word corresponding to the triple, and the tail entity are stored in the final keyword set. Repeat the above process to obtain all the word paths, that is, obtain the final keyword set.

[0188] S425. Determine the final response of the AI digital human based on the key attention weight vector and the emotional dialogue representation ;

[0189] The step S425 includes:

[0190] S4251. Calculate the output state based on the key attention weight vector and the emotional dialogue representation :

[0191] ;

[0192] In the formula, represents the transformer decoder, is a linear transformation, is the embedding corresponding to the current word vector, is the preset token embedding, is the generated token embedding, is the concatenation operation.

[0193] S4252. Determine the word probability distribution based on the output state , and generate the final response of the AI digital human based on the word probability distribution:

[0194] ;

[0195] In the formula, represents the pointer-generator network, and respectively represent the emotional encoding feature and the knowledge encoding feature;

[0196] Specifically, after generating the word probability distribution, the token can also be used as the guidance of the pointer, and the final response can be generated according to the probability.

[0197] The interactive response generation method based on an AI digital human provided in the first embodiment of the present invention first obtains the current conversation text and the historical conversation text of the user, and converts the current conversation text into a current word vector; then determines the emotional feature based on the current word vector and the historical conversation text; obtains external common sense information and converts the external common sense information into an auxiliary knowledge graph, and determines the auxiliary knowledge feature based on the auxiliary knowledge graph, the current word vector and the historical conversation text; then encodes the emotional feature and the auxiliary knowledge feature respectively to obtain an emotional encoding feature and a knowledge encoding feature, and determines the final response of the AI digital human based on the emotional encoding feature and the knowledge encoding feature. The present invention can enhance the emotion in the current conversation text when identifying the emotion of the user, thereby improving the accuracy and speed of emotion recognition. At the same time, by determining the auxiliary knowledge feature, the response of the AI digital human can be guided and assisted by knowledge, avoiding inaccurate or chaotic reply expressions. Finally, the final response is generated by combining the emotional feature and the auxiliary knowledge feature, thereby improving the realism, naturalness and anthropomorphism of the AI digital human.

[0198] Embodiment 2

[0199] As Figure 2 shown, the second embodiment of the present invention provides an interactive response generation system based on an AI digital human. The system includes:

[0200] Conversion module 1, configured to obtain the current conversation text and the historical conversation text of the user, and convert the current conversation text into a current word vector;

[0201] Emotion module 2, configured to determine the emotional feature based on the current word vector and the historical conversation text;

[0202] Auxiliary module 3, configured to obtain external common sense information and convert the external common sense information into an auxiliary knowledge graph, and determine the auxiliary knowledge feature based on the auxiliary knowledge graph, the current word vector and the historical conversation text;

[0203] Reply module 4, configured to encode the emotional feature and the auxiliary knowledge feature respectively to obtain an emotional encoding feature and a knowledge encoding feature, and determine the final response of the AI digital human based on the emotional encoding feature and the knowledge encoding feature;

[0204] The emotion module 2 includes:

[0205] Expansion sub-module, configured to obtain an emotion data set and construct an emotion knowledge graph according to the emotion data set, expand the historical conversation text into a historical word sequence, and construct the historical word sequence into a historical knowledge graph according to the construction method of the emotion knowledge graph;

[0206] An extraction sub-module, configured to extract keywords from the current word vector and determine the current relevant knowledge graph corresponding to each keyword;

[0207] A semantic sub-module, configured to determine semantic features based on the current relevant knowledge graph;

[0208] A personalization sub-module, configured to extract personalization information from the historical conversation text, form a personalization matrix according to the quantity and sequence length of the personalization information, and input the personalization information into a trained preset prediction model for prediction to obtain predicted personalization information;

[0209] A symbolization sub-module, configured to perform symbolization processing on the historical conversation text and the predicted personalization information respectively to obtain processed text and processed information;

[0210] A splicing and encoding sub-module, configured to splice the processed text and the processed information to obtain a first splicing feature, and perform encoding processing on the first splicing feature and the personalization matrix respectively to obtain a first encoding feature and a second encoding feature;

[0211] An emotion sub-module, configured to determine emotion features based on the first encoding feature, the second encoding feature, and the semantic features.

[0212] The extraction sub-module includes:

[0213] A triple unit, configured to extract keywords from the current word vector to obtain a keyword group, and extract implicit triples in the emotion knowledge graph that include at least one keyword in the keyword group to obtain a triple set;

[0214] An embedding unit, configured to determine a plurality of implicit emotion sub-graphs based on the triple set, convert the plurality of implicit emotion sub-graphs into natural sentences, and convert the current conversation text into a first word embedding;

[0215] A similarity unit, configured to convert the natural sentence into a second word embedding and calculate the similarity between the first word embedding and the second word embedding;

[0216] A relevant graph unit, configured to retain the implicit emotion sub-graphs corresponding to a similarity greater than a preset similarity to obtain the current relevant knowledge graph corresponding to each keyword .

[0217] The semantic sub-module includes:

[0218] A statement unit, configured to determine key statements based on the current relevant knowledge graph :

[0219] ​​ ;

[0220] In the formula, represents the connection operation, represents the embedding of the

[0221] Pooling unit, for inputting the key sentence into a preset sentence encoding model for encoding to obtain an encoding matrix , and determining a first pooling feature and a second pooling feature based on the encoding matrix:

[0222] ;

[0223] ;

[0224] In the formula, represents the element in the th row and the th column of the encoding matrix;

[0225] Emotion vector unit, for determining an emotion vector based on the historical knowledge graph, the first pooling feature and the second pooling feature :

[0226] ;

[0227] ;

[0228] In the formula, represents the global feature, represents layer normalization, represents the vertex in the historical knowledge graph, represents adjacent matrix of represents the th self-attention head, represents the weight of the th self-attention head, represents self-attention layer composed of

[0229] Semantic feature unit, for calculating a semantic feature based on the emotion vector :

[0230] ;

[0231] In the formula, respectively represent the first and second weights respectively represent the first and second biases represents the error value

[0232] The emotion sub-module includes:

[0233] An average pooling unit, configured to perform average pooling processing on the second encoded feature and splice the second encoded feature after the average pooling processing with the first encoded feature to obtain a second spliced feature, and perform encoding and decoding processing on the second spliced feature in sequence to obtain a personalized feature;

[0234] A statement feature unit, configured to connect the personalized feature with the semantic feature to obtain a personalized semantic feature , and determine the emotion statement feature based on the personalized semantic feature : :

[0235] ;

[0236] ;

[0237] In the formula, , respectively represent the hidden states corresponding to the layer decoder, , , respectively represent the first, second, and third learnable parameter matrices, is the dimension, is the fusion operation, represents the th keyword;

[0238] An identification unit, configured to input the emotion statement feature into a trained preset emotion recognition model for recognition to obtain an emotion feature

[0239] The auxiliary module 3 includes:

[0240] A reference sub-module, configured to extract an auxiliary knowledge sub-graph corresponding to the historical conversation text from the auxiliary knowledge graph and use it as a reference sub-graph;

[0241] A layer computing sub-module, configured to input each node in the reference sub-graph into a graph convolutional neural network for node embedding layer computing and relationship embedding layer computing to obtain a node embedding and a relationship embedding :

[0242] ;

[0243] ;

[0244] In the formula, is the activation function, represents the neighbor subgraph of node in the reference subgraph, , , respectively represent the first, second, and third learnable weight matrices of the th layer in the graph convolutional neural network, respectively represent the node embeddings of the th layer, respectively represent the relationship embeddings of the th layer;

[0245] The scoring sub-module is used to calculate the evaluation score of each node in the reference subgraph based on the node embedding and the relationship embedding , select the node with the maximum evaluation score for access connection among each node in the reference subgraph to obtain an updated reference subgraph:

[0246] ;

[0247] In the formula, is the adjustment coefficient, represents the evaluation score of the node that has been accessed and connected, is the non-linear activation function, represents node 's node embedding in the th layer, represents the fourth learnable weight matrix of the th layer in the graph convolutional neural network, represents the hidden state of the th layer in the graph convolutional neural network at the th time step;

[0248] The semantic embedding sub-module is used to extract the first semantic embedding of the nodes in the historical knowledge graph corresponding to the historical dialogue text, the second semantic embedding of the node connection edges, and extract the third semantic embedding of the nodes in the updated reference subgraph, the fourth semantic embedding of the node connection edges, and calculate the final knowledge feature based on the first semantic embedding, the second semantic embedding, the third semantic embedding, and the fourth semantic embedding:

[0249] ;

[0250] In the formula, is a connection operation;

[0251] A knowledge feature sub-module, used to determine auxiliary knowledge features based on the final knowledge feature : :

[0252] ;

[0253] In the formula, is probability calculation, represents a set in is an updated reference sub-graph, is the historical dialogue text.

[0254] The reply module 4 includes:

[0255] An embedding encoding sub-module, used to encode the emotion feature and the auxiliary knowledge feature respectively to obtain an emotion embedding and an auxiliary knowledge embedding ;

[0256] An equation sub-module, used to construct a first solution equation and a second solution equation based on the emotion embedding and the auxiliary knowledge embedding :

[0257] ; ;

[0258] In the formula, , respectively represent emotion encoding features and knowledge encoding features, , respectively represent the tiny change amounts of the emotion embedding and the auxiliary knowledge embedding , respectively represent the weight vectors corresponding to the emotion embedding and the auxiliary knowledge embedding , respectively represent the bias vectors corresponding to the emotion embedding and the auxiliary knowledge embedding ;

[0259] A solution sub-module, used to solve the first solution equation and the second solution equation respectively to obtain an emotion encoding feature and a knowledge encoding feature ;

[0260] The reply module 4 further includes:

[0261] A representation sub-module, configured to encode the current word vector to obtain a current word representation, convert the triples in the auxiliary knowledge graph into natural sentences and encode the natural sentences to obtain natural sentence representations;

[0262] A dialogue sub-module, configured to determine an emotional dialogue representation based on the current word representation and the natural sentence representation : :

[0263] ;

[0264] wherein, is a concatenation operation, is a linear weight matrix;

[0265] A permutation sub-module, configured to extract historical keywords from the historical dialogue text and obtain historical keyword embeddings of the historical keywords, encode the historical keyword embeddings to obtain historical key features, calculate the similarity between the historical key features and each triple in the auxiliary knowledge graph, and sort each triple in the auxiliary knowledge graph in descending order according to the similarity and select the top several triples as a permutation knowledge group;

[0266] A judgment sub-module, configured to judge whether any historical keyword is the same as the tail entity of the triple in the permutation knowledge group. If any historical keyword is the same as the tail entity of the triple in the permutation knowledge group, store the historical keyword, the word corresponding to the triple, and the tail entity into the final keyword set, and calculate the attention weight of the final keyword set to obtain a key attention weight vector ;

[0267] A reply output sub-module, configured to determine the final reply of the AI digital human based on the key attention weight vector and the emotional dialogue representation ;

[0268] The reply output sub-module includes:

[0269] A state unit, configured to calculate and output a state based on the key attention weight vector and the emotional dialogue representation : :

[0270] ;

[0271] wherein, represents a transformer decoder, is a linear transformation, is the embedding corresponding to the current word vector, is the preset marker embedding, is the generated marker embedding, is a concatenation operation;

[0272] A reply output unit, for determining a word probability distribution based on the output state and generating the final reply of the AI digital human based on the word probability distribution:

[0273] ;

[0274] In the formula, represents a pointer generation network, , respectively represent an emotion coding feature and a knowledge coding feature.

[0275] In some other embodiments of the present invention, the present invention embodiment provides the following technical solution. A computer includes a memory 102, a processor 101, and a computer program stored on the memory 102 and executable on the processor 101. When the processor 101 executes the computer program, it implements an interactive reply generation method based on an AI digital human as described above.

[0276] Specifically, the above-mentioned processor 101 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC for short), or may be configured as one or more integrated circuits implementing the embodiments of the present invention.

[0277] ​Among them, the memory 102 may include a mass storage for data or instructions. By way of example and not limitation, the memory 102 may include a hard disk drive (HDD), a floppy disk drive, a solid state drive (SSD), a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 102 may include removable or non-removable (or fixed) media. In a suitable case, the memory 102 may be inside or outside the data processing device. In a specific embodiment, the memory 102 is a non-volatile memory. In a specific embodiment, the memory 102 includes a read-only memory (ROM) and a random access memory (RAM). In a suitable case, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or a flash memory, or a combination of two or more of these. In a suitable case, the RAM may be a static random-access memory (SRAM) or a dynamic random access memory (DRAM), where the DRAM may be a fast page mode dynamic random access memory (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random-access memory (SDRAM), etc.

[0278] The memory 102 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 101.

[0279] The processor 101 reads and executes the computer program instructions stored in the memory 102 to implement the above-mentioned AI digital human-based interactive response generation method.

[0280] In some of these embodiments, the computer may further include a communication interface 103 and a bus 100. Among them, as Figure 3 shown, the processor 101, the memory 102, and the communication interface 103 are connected through the bus 100 to complete communication with each other.

[0281] The communication interface 103 is used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present invention. The communication interface 103 can also implement data communication with other components, such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations, etc.

[0282] Bus 100 includes hardware, software, or both, and couples components of a computer device to each other. Bus 100 includes, but is not limited to, at least one of the following: Data Bus, Address Bus, Control Bus, Expansion Bus, Local Bus. By way of example and not limitation, Bus 100 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a Memory Bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable bus or a combination of two or more of these. In suitable cases, Bus 100 may include one or more buses. Although embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.

[0283] The computer may execute the interactive response generation method based on an AI digital human of the present invention based on an interactive response generation system based on an AI digital human obtained, so as to implement the interactive response generation based on an AI digital human.

[0284] In still some other embodiments of the present invention, in combination with the above-mentioned interactive response generation method based on an AI digital human, embodiments of the present invention provide the following technical solution: a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned interactive response generation method based on an AI digital human is implemented.

[0285] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0286] More specific examples (a non-exhaustive list) of the readable medium include the following: an electrical connection portion (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0287] It should be understood that the various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.

[0288] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as falling within the scope described in this specification.

[0289] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. An interactive response generation method based on an AI digital human, characterized in that, Including: Obtain the current conversation text and historical conversation text of the user, and convert the current conversation text into a current word vector; Determine an emotional feature based on the current word vector and the historical conversation text; Obtain external common sense information and convert the external common sense information into an auxiliary knowledge graph, and determine an auxiliary knowledge feature based on the auxiliary knowledge graph, the current word vector, and the historical conversation text; Encode the emotional feature and the auxiliary knowledge feature respectively to obtain an emotional encoding feature and a knowledge encoding feature, and determine the final reply of the AI digital human based on the emotional encoding feature and the knowledge encoding feature; The step of encoding the emotional feature and the auxiliary knowledge feature respectively to obtain an emotional encoding feature and a knowledge encoding feature includes: Encode the emotional feature and the auxiliary knowledge feature respectively to obtain an emotion embedding and an auxiliary knowledge embedding ; Based on emotion embedding and auxiliary knowledge embedding Construct the first solution equation and the second solution equation: ; ; Wherein, and respectively represent the emotion coding feature and the knowledge coding feature, and respectively represent the tiny change amount of the emotion embedding and the auxiliary knowledge embedding ; and respectively represent the weight vectors corresponding to the emotion embedding and the auxiliary knowledge embedding ; and respectively represent the bias vectors corresponding to the emotion embedding Solve the first solution equation and the second solution equation respectively to obtain the emotion coding features and the knowledge coding features .

2. The interactive response generation method based on an AI digital human according to claim 1, wherein The step of determining an emotional feature based on the current word vector and the historical conversation text includes: Obtain an emotion dataset and construct an emotion knowledge graph according to the emotion dataset, expand the historical conversation text into a historical word sequence, and construct the historical word sequence into a historical knowledge graph according to the construction method of the emotion knowledge graph; Extract the keywords in the current word vector, and determine the corresponding current relevant knowledge graph for each keyword; Determine a semantic feature based on the current relevant knowledge graph; Extract personalized information from the historical conversation text, form a personalized matrix according to the quantity and sequence length of the personalized information, and input the personalized information into a trained preset prediction model for prediction to obtain predicted personalized information; Perform symbolic processing on the historical conversation text and the predicted personalized information respectively to obtain processed text and processed information; Concatenate the processed text and the processed information to obtain a first concatenated feature, and perform encoding processing on the first concatenated feature and the personalized matrix respectively to obtain a first encoding feature and a second encoding feature; Determine an emotional feature based on the first encoding feature, the second encoding feature, and the semantic feature.

3. The interactive response generation method based on an AI digital human according to claim 2, wherein The step of extracting the keywords in the current word vector and determining the corresponding current relevant knowledge graph for each keyword includes: Extract the keywords in the current word vector to obtain a keyword group, and extract the implicit triples in the emotion knowledge graph that include at least one keyword in the keyword group to obtain a triple set; Determine a number of implicit emotion subgraphs based on the triple set, convert the number of implicit emotion subgraphs into natural sentences, and convert the current conversation text into a first word embedding; Convert the natural sentence into a second word embedding and calculate the similarity between the first word embedding and the second word embedding; Retain the implicit emotion subgraphs whose similarity is greater than the corresponding preset similarity to obtain the current relevant knowledge graph corresponding to each keyword corresponding to the current relevant knowledge graph .

4. The interactive reply generation method based on an AI digital human according to claim 2, wherein The step of determining a semantic feature based on the current relevant knowledge graph includes: Based on the current relevant knowledge graph Determine the key sentences : ; In the formula, represents a connection operation, represents the embedding of the Input the key statement into a preset statement encoding model for encoding to obtain an encoding matrix , and determine a first pooling feature based on the encoding matrix and a second pooling feature : ; ; In the formula, represents the element in the th row and th column of the encoding matrix; Based on the historical knowledge graph and the first pooled feature and the second pooled feature determine the emotion vector : ; ; Wherein, represents the global feature, represents layer normalization, represents the vertex in the historical knowledge graph, represents the adjacent matrix of represents the th self-attention head, represents the weight of the th self-attention head, represents the self-attention layer composed of Based on the emotion vector Calculate semantic features : ; In the formula, respectively represent the first and second weights, respectively represent the first and second deviations, represents the error value.

5. The interactive response generation method based on an AI digital human according to claim 2, characterized in that The steps of determining the emotion feature based on the first coding feature, the second coding feature, and the semantic feature include: Perform average pooling processing on the second encoding feature, concatenate the second encoding feature after average pooling processing with the first encoding feature to obtain a second concatenated feature, and perform encoding and decoding processing on the second concatenated feature in sequence to obtain a personalized feature; Connect the personalized feature with the semantic feature to obtain a personalized semantic feature , and determine the emotional statement feature based on the personalized semantic feature : ; ; Wherein, and respectively represent the hidden states corresponding to the -th layer decoder, , and respectively represent the first, second, and third learnable parameter matrices, is the dimension, is the fusion operation, represents the -th keyword; Input the emotional statement feature into the trained preset emotional recognition model for recognition to obtain the emotional feature.

6. A method for generating interactive responses based on an AI digital human according to claim 1, characterized in that The step of determining an auxiliary knowledge feature based on the auxiliary knowledge graph, the current word vector, and the historical conversation text includes: Extract the auxiliary knowledge subgraph corresponding to the historical dialogue text from the auxiliary knowledge graph and use it as the benchmark subgraph; Each node in the reference subgraph is input into a graph convolutional neural network for node embedding layer calculation and relationship embedding layer calculation to obtain node embeddings and relationship embeddings : ; ; wherein, is the activation function, represents the neighbor subgraph of node in the reference subgraph, , , respectively represent the first, second, and third learnable weight matrices of the th layer in the graph convolutional neural network, respectively represent the node embeddings of the th layer, respectively represent the relationship embeddings of the th layer; Based on node embedding and relationship embedding Calculate the evaluation score of each node in the reference subgraph , and select the node with the maximum evaluation score for each node in the reference subgraph for access connection to obtain an updated reference subgraph: ; Wherein, is an adjustment coefficient, represents the evaluation score of the node that has been accessed and connected, is a non-linear activation function, represents the node at the layer's node embedding, represents the fourth learnable weight matrix of the layer in the graph convolutional neural network, represents the hidden state of the layer in the graph convolutional neural network at the th time step; Extract the first semantic embedding of the nodes in the historical knowledge graph corresponding to the historical dialogue text and the second semantic embedding of the node connection edges and extract the third semantic embedding of the nodes in the updated benchmark subgraph and the fourth semantic embedding of the node connection edges , and calculate the final knowledge features based on the first semantic embedding, the second semantic embedding, the third semantic embedding, and the fourth semantic embedding : ; In the formula, is a connection operation; Based on the final knowledge features Determine the auxiliary knowledge features : ; In the formula, is for probability calculation, represents a set in is the updated reference subgraph, is the historical dialogue text.

7. A method for generating interactive responses based on an AI digital human according to claim 1, characterized in that, The step of determining the final response of the AI digital human based on the emotion coding feature and the knowledge coding feature includes: Encode the current word vector to obtain a current word representation, convert the triples in the auxiliary knowledge graph into natural sentences and encode the natural sentences to obtain natural sentence representations; Based on the current word representation and the natural sentence representation determine the emotional dialogue representation : ; In the formula, is a connection operation, is a linear weight matrix; Extract historical keywords from the historical dialogue text and obtain historical keyword embeddings of the historical keywords, encode the historical keyword embeddings to obtain historical key features, calculate the similarity between the historical key features and each triple in the auxiliary knowledge graph, and arrange each triple in the auxiliary knowledge graph in descending order according to the similarity and select the first several triples as the arranged knowledge group; Determine whether any historical keyword is the same as the tail entity of the triple in the permutation knowledge group. If any historical keyword is the same as the tail entity of the triple in the permutation knowledge group, then store the historical keyword, the corresponding word of the triple, and the tail entity in the final keyword set, and calculate the attention weight for the final keyword set to obtain the key attention weight vector ; Based on the key attention weight vector , the emotional dialogue representation to determine the final response of the AI digital human.

8. A method for generating interactive responses based on an AI digital human according to claim 7, characterized in that, Based on the key attention weight vector , and the emotional dialogue representation The steps for determining the final response of the AI digital human include: Based on the key attention weight vector , the emotional dialogue representation is used to calculate the output state : ; In the formula, represents the transformer decoder, is a linear transformation, is the embedding corresponding to the current word vector, is the preset token embedding, is the generated token embedding, is a concatenation operation; Based on the output status Determine the word probability distribution , and generate the final response of the AI digital human based on the word probability distribution: ; In the formula, represents the pointer generation network, , respectively represent the emotion coding feature and the knowledge coding feature.

9. An interactive response generation system based on an AI digital human, the system adopts the interactive response generation method based on an AI digital human as described in claim 1, characterized in that, The system includes: A conversion module for obtaining the current dialogue text and the historical dialogue text of the user and converting the current dialogue text into a current word vector; An emotion module for determining emotion features based on the current word vector and the historical dialogue text; An auxiliary module for obtaining external common sense information and converting the external common sense information into an auxiliary knowledge graph, and determining auxiliary knowledge features based on the auxiliary knowledge graph, the current word vector and the historical dialogue text; A reply module for encoding the emotion feature and the auxiliary knowledge feature respectively to obtain an emotion coding feature and a knowledge coding feature, and determining the final reply of the AI digital human based on the emotion coding feature and the knowledge coding feature.

Citation Information

Patent Citations

  • Emotion-sharing dialogue generation method based on context awareness and emotion reasoning

    CN117892736A

  • Multi-round dialogue semantic analysis method and system based on long short-term memory network

    WO2021042543A1