Data processing method for dialogue content generation, virtual dialogue, dialogue content
Patent Information
- Application Number
- CN202310120641.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-18
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-01-18
AI Technical Summary
[0059]根据本说明书实施例的第十一方面,提供了一种计算机程序,其中,当所述计算机程序在计算机中执行时,令计算机执行上述第一方面或者第二方面或者第三方面或者第四方面所提供方法的步骤。
Smart Images

Figure CN116136869B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a method for generating dialogue content. One or more embodiments of this specification also relate to a virtual dialogue method, a method for processing dialogue content data, a dialogue content generation apparatus, a virtual dialogue apparatus, a dialogue content data processing apparatus, a computing device, a computer-readable storage medium, and a computer program. Background Technology
[0002] With the development of internet technology, people's lifestyles have undergone significant changes. Users can now interact with chatbots online to address their daily needs, such as online chatting, online bill payment, and online education. Therefore, how chatbots can generate accurate responses has gradually become a key research focus.
[0003] Currently, responses are typically generated based on sentiment analysis and rules: first, the sentiment in the user's request is analyzed, and then corresponding responses are given based on the identification results; if the sentiment is negative, general reassurance is provided. However, the above methods fail to demonstrate genuine empathy in the generated responses, resulting in poor accuracy. Therefore, a solution that can demonstrate empathy and high accuracy is urgently needed. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a method for generating dialogue content. One or more embodiments of this specification also relate to a virtual dialogue method, a data processing method for dialogue content, a dialogue content generation apparatus, a virtual dialogue apparatus, a data processing apparatus for dialogue content, a computing device, a computer-readable storage medium, and a computer program, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a method for generating dialogue content is provided, comprising:
[0006] Obtain emotional conversation content;
[0007] Extract at least one emotional keyword from the content of the emotional dialogue;
[0008] From a pre-constructed sentiment association graph, a target subgraph containing at least one sentiment keyword is determined, wherein the sentiment association graph is constructed based on the association relationships between multiple sample sentiment keywords;
[0009] Based on at least one sentiment keyword and a target subgraph, predict the target keywords in the response content;
[0010] Generate dialogue response content based on the emotional dialogue content and target keywords.
[0011] According to a second aspect of the embodiments of this specification, a virtual dialogue method is provided, comprising:
[0012] Receive dialogue requests sent from the front end, where the dialogue requests carry emotional dialogue text;
[0013] Extract at least one emotional keyword from the emotional dialogue text;
[0014] From a pre-constructed sentiment association graph, a target subgraph containing at least one sentiment keyword is determined, wherein the sentiment association graph is constructed based on the association relationships between multiple sample sentiment keywords;
[0015] Based on at least one sentiment keyword and a target subgraph, predict the target keyword in the response text;
[0016] Generate dialogue response text based on the emotional dialogue text and target keywords;
[0017] Send the dialogue reply text to the front end so that the front end can display the dialogue reply text.
[0018] According to a third aspect of the embodiments of this specification, a method for processing dialogue content is provided, applied to a cloud-side device, the method comprising:
[0019] Obtain the first sample set, which includes multiple training emotional dialogue contents, multiple training emotional words, and the training emotional dialogue contents carry response word tags;
[0020] Feature extraction is performed on the training emotional words corresponding to each training emotional dialogue content to obtain the training emotional features corresponding to each training emotional dialogue content.
[0021] The training sentiment features and the training sentiment association graph corresponding to each training sentiment dialogue content are input into the graph attention model to obtain the predicted response words corresponding to each training sentiment dialogue content. The training sentiment association graph is constructed based on the association relationship between multiple training sentiment words.
[0022] The graph attention model is trained based on the response word tags and predicted response words to obtain the model parameters of the trained graph attention model;
[0023] Send the trained graph attention model parameters to the edge device.
[0024] According to a fourth aspect of the embodiments of this specification, a method for processing dialogue content is provided, applied to a cloud-side device, the method comprising:
[0025] Obtain a second sample set, which includes multiple training emotional dialogue contents and training response words corresponding to each training emotional dialogue content. The training emotional dialogue contents carry response content tags.
[0026] Multiple training emotional dialogue contents and the corresponding training response words are input into the encoder of the response generation model to obtain the predicted encoding representation;
[0027] The predicted encoding is input into the decoder of the response generation model to obtain the predicted response content corresponding to the trained emotional dialogue content;
[0028] The response generation model is trained based on the predicted response content and response content tags to obtain the model parameters of the trained response generation model;
[0029] Send the trained response to the edge device to generate the model parameters of the model.
[0030] According to a fifth aspect of the embodiments of this specification, a dialogue content generation apparatus is provided, comprising:
[0031] The first acquisition module is configured to acquire emotional dialogue content;
[0032] The first extraction module is configured to extract at least one emotional keyword from the emotional dialogue content;
[0033] The first determining module is configured to determine a target subgraph containing at least one sentiment keyword from a pre-constructed sentiment association graph, wherein the sentiment association graph is constructed based on the association relationships between multiple sample sentiment keywords;
[0034] The first prediction module is configured to predict target keywords in the response content based on at least one sentiment keyword and a target subgraph.
[0035] The first generation module is configured to generate dialogue response content based on the emotional dialogue content and target keywords.
[0036] According to a sixth aspect of the embodiments of this specification, a virtual dialogue device is provided, comprising:
[0037] The first receiving module is configured to receive dialogue requests sent by the front end, wherein the dialogue requests carry emotional dialogue text.
[0038] The second extraction module is configured to extract at least one sentiment keyword from the sentiment dialogue text.
[0039] The second determining module is configured to determine a target subgraph containing at least one sentiment keyword from a pre-constructed sentiment association graph, wherein the sentiment association graph is constructed based on the association relationships between multiple sample sentiment keywords;
[0040] The second prediction module is configured to predict target keywords in the response text based on at least one sentiment keyword and a target subgraph.
[0041] The second generation module is configured to generate dialogue response text based on the emotional dialogue text and target keywords;
[0042] The first sending module is configured to send the dialogue reply text to the front end so that the front end can display the dialogue reply text.
[0043] According to a seventh aspect of the embodiments of this specification, a data processing apparatus for dialogue content is provided, applied to a cloud-side device, the apparatus comprising:
[0044] The second acquisition module is configured to acquire a first sample set, wherein the first sample set includes multiple training emotional dialogue contents, the training emotional dialogue contents include multiple training emotional words, and the training emotional dialogue contents carry response word tags;
[0045] The third extraction module is configured to extract features from the training emotional words corresponding to each training emotional dialogue content, and obtain the training emotional features corresponding to each training emotional dialogue content.
[0046] The first input module is configured to input the training sentiment features and the training sentiment association graph corresponding to each training sentiment dialogue content into the graph attention model to obtain the predicted response words corresponding to each training sentiment dialogue content. The training sentiment association graph is constructed based on the association relationship between multiple training sentiment words.
[0047] The first training module is configured to train the graph attention model based on the response word tags and the predicted response words, and obtain the model parameters of the trained graph attention model.
[0048] The second sending module is configured to send the model parameters of the trained graph attention model to the end device.
[0049] According to an eighth aspect of the embodiments of this specification, a data processing apparatus for dialogue content is provided, applied to a cloud-side device, the apparatus comprising:
[0050] The third acquisition module is configured to acquire a second sample set, which includes multiple training emotional dialogue contents and training response words corresponding to each training emotional dialogue content. The training emotional dialogue contents carry response content tags.
[0051] The second input module is configured to input multiple training emotional dialogue contents and the training response words corresponding to each training emotional dialogue content into the encoder of the response generation model to obtain the predicted encoded representation;
[0052] The third input module is configured to input the predicted response content into the decoder of the response generation model to obtain the predicted response content corresponding to the trained emotional dialogue content;
[0053] The second training module is configured to train the response generation model based on the predicted response content and the response content tags, and obtain the model parameters of the trained response generation model.
[0054] The third sending module is configured to send the model parameters of the trained response generation model to the end device.
[0055] According to a ninth aspect of the embodiments of this specification, a computing device is provided, comprising:
[0056] Memory and processor;
[0057] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method provided in the first, second, third, or fourth aspect described above.
[0058] According to a tenth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of the methods provided in the first, second, third, or fourth aspects described above.
[0059] According to an eleventh aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the method provided in the first, second, third, or fourth aspect described above.
[0060] This specification provides a dialogue content generation method according to one embodiment, which involves: acquiring emotional dialogue content; extracting at least one emotional keyword from the emotional dialogue content; determining a target subgraph containing at least one emotional keyword from a pre-constructed emotional association graph, wherein the emotional association graph is constructed based on the association relationships between multiple sample emotional keywords; predicting target keywords in the response content based on at least one emotional keyword and the target subgraph; and generating dialogue response content based on the emotional dialogue content and the target keyword. By constructing an emotional association graph, a deeper understanding of the specific keywords that trigger emotions in the dialogue can be achieved, thereby predicting target keywords through emotional keywords, effectively demonstrating empathy in the dialogue response content, and thus realizing empathetic dialogue and improving the accuracy of dialogue content generation. Attached Figure Description
[0061] Figure 1 This is a framework diagram of a dialogue content generation system provided in one embodiment of this specification;
[0062] Figure 2 This is a framework diagram of another dialogue content generation system provided in one embodiment of this specification;
[0063] Figure 3 This is a flowchart of a dialogue content generation method provided in one embodiment of this specification;
[0064] Figure 4 This is a flowchart illustrating a virtual dialogue method provided in one embodiment of this specification;
[0065] Figure 5 This is a flowchart illustrating a data processing method for dialogue content provided in one embodiment of this specification;
[0066] Figure 6 This is a flowchart of another method for processing dialogue content provided in one embodiment of this specification;
[0067] Figure 7 This is a flowchart illustrating the processing steps of a dialogue content generation method provided in one embodiment of this specification.
[0068] Figure 8 This is a flowchart illustrating the construction of an emotion association graph in a dialogue content generation method provided in one embodiment of this specification;
[0069] Figure 9 This is a schematic diagram of a virtual dialogue interface provided in one embodiment of this specification;
[0070] Figure 10 This is a schematic diagram of the structure of a dialogue content generation device provided in one embodiment of this specification;
[0071] Figure 11 This is a schematic diagram of the structure of a virtual dialogue device provided in one embodiment of this specification;
[0072] Figure 12 This is a schematic diagram of the structure of a data processing device for dialogue content provided in one embodiment of this specification;
[0073] Figure 13 This is a schematic diagram of the structure of another data processing device for dialogue content provided in one embodiment of this specification;
[0074] Figure 14 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0075] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0076] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0077] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0078] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0079] Empathic dialogue: Empathy, also known as compassion, refers to the ability to imagine oneself in the other person's situation, to truly understand their feelings, and to respond appropriately during a conversation.
[0080] Emotional Reasons: After identifying the user's emotions, find out the specific reasons behind those emotions.
[0081] In the intelligent customer service industry, most chatbots focus solely on developing IQ, neglecting the importance of EQ, resulting in poor conversational accuracy and a subpar user experience. Therefore, creating a warm, empathetic, and accurate chatbot is both crucial and challenging.
[0082] Currently, most emotional responses are based solely on sentiment analysis and rules. For example, they first analyze the emotional tendency in a user's request, then provide corresponding responses based on the identification results; if the emotion is negative, a general reassurance is given. For instance, if a user's emotional tendency is sadness due to failing an exam, a general response like "Don't be sad, tomorrow will be better" is given. However, these methods fail to specifically understand the underlying cause of the user's emotion (failing the exam), only providing generic responses. They cannot demonstrate genuine empathy in the generated responses, resulting in low accuracy and a poor user experience.
[0083] To address the aforementioned issues, this specification proposes an empathic dialogue content generation scheme based on an emotion association graph. First, it identifies emotional keywords in the user-input emotional dialogue content and constructs an emotion association graph based on multiple sample dialogues. Second, using the emotion association graph, it predicts target keywords to be included in the response based on the identified emotional keywords in the user's request. Finally, it generates a final empathic response based on the user-input emotional dialogue content and the predicted target keywords. This allows for a better understanding of the specific reasons triggering the user's emotions and enables the robot to better demonstrate empathy in its responses, thereby creating a more accurate dialogue experience. For example, if the emotional tendency is sadness due to failing an exam, the empathic response would be, "Don't be sad, your grades have always been excellent."
[0084] Specifically, the process involves: acquiring emotional dialogue content; extracting at least one emotional keyword from the emotional dialogue content; determining a target subgraph containing at least one emotional keyword from a pre-constructed emotional association graph, wherein the emotional association graph is constructed based on the association relationships between multiple sample emotional keywords; predicting target keywords in the response content based on at least one emotional keyword and the target subgraph; and generating dialogue response content based on the emotional dialogue content and the target keyword. By constructing an emotional association graph, a deeper understanding of the specific keywords that trigger emotions in the dialogue can be achieved, thereby predicting target keywords through emotional keywords, effectively demonstrating empathy in the dialogue response content, and ultimately realizing empathetic dialogue, thus improving the accuracy of dialogue content generation.
[0085] This specification provides a method for generating dialogue content, and also relates to a virtual dialogue method, a method for processing dialogue content data, a dialogue content generation apparatus, a virtual dialogue apparatus, a dialogue content data processing apparatus, a computing device, a computer-readable storage medium, and a computer program, which will be described in detail in the following embodiments.
[0086] See Figure 1 , Figure 1 A framework diagram of a dialogue content generation system provided in one embodiment of this specification is shown, wherein the dialogue content generation system includes a server 100 and a client 200;
[0087] Client 200: Sends emotional dialogue content to server 100;
[0088] Server 100: Extract at least one sentiment keyword from the sentiment dialogue content; determine a target subgraph containing at least one sentiment keyword from a pre-constructed sentiment association graph, wherein the sentiment association graph is constructed based on the association relationships between multiple sample sentiment keywords; predict the target keyword in the response content based on at least one sentiment keyword and the target subgraph; generate dialogue response content based on the sentiment dialogue content and the target keyword; and send the dialogue response content to client 200.
[0089] Client 200: Receives the dialogue response content sent by server 100.
[0090] The method described in this specification involves: acquiring emotional dialogue content; extracting at least one emotional keyword from the emotional dialogue content; determining a target subgraph containing at least one emotional keyword from a pre-constructed emotional association graph, wherein the emotional association graph is constructed based on the association relationships between multiple sample emotional keywords; predicting target keywords in the response content based on at least one emotional keyword and the target subgraph; and generating dialogue response content based on the emotional dialogue content and the target keyword. By constructing an emotional association graph, a deep understanding of the specific keywords that trigger emotions in the dialogue can be achieved, thereby predicting target keywords through emotional keywords, effectively demonstrating empathy in the dialogue response content, and thus realizing empathetic dialogue and improving the accuracy of dialogue content generation.
[0091] See Figure 2 , Figure 2 This specification illustrates a framework diagram of another dialogue content generation system according to an embodiment. The system may include a server 100 and multiple clients 200. The multiple clients 200 can establish communication connections through the server 100. In a dialogue content generation scenario, the server 100 provides dialogue content generation services between the multiple clients 200. Each client 200 can act as a sender or receiver, achieving real-time communication through the server 100.
[0092] Users can interact with server 100 through client 200 to receive data sent by other clients 200, or send data to other clients 200, etc. In the scenario of generating dialogue content, users can publish data streams to server 100 through client 200, server 100 can generate dialogue response content based on the data stream, and push the dialogue response content to other clients that have established communication.
[0093] In this system, client 200 and server 100 establish a connection via a network. The network provides the medium for communication between the client and server. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by client 200 may need to undergo encoding, transcoding, compression, or other processing before being published to server 100.
[0094] Client 200 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. Client 200 can be developed based on the software development kit (SDK) provided by the server, such as a real-time communication (RTC) SDK. Client 200 can be deployed on electronic devices, requiring the device to run or certain apps on the device to function. Electronic devices may have displays and support information browsing, such as personal mobile terminals like smartphones, tablets, and personal computers. Various other types of applications can also be configured on electronic devices, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platforms.
[0095] Server 100 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 100 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0096] It is worth noting that the dialogue content generation method provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the dialogue content generation method provided in the embodiments of this specification. In other embodiments, the dialogue content generation method provided in the embodiments of this specification may also be executed jointly by the client and the server.
[0097] See Figure 3 , Figure 3 This specification shows a flowchart of a dialogue content generation method according to an embodiment, which specifically includes the following steps:
[0098] Step 302: Obtain the emotional dialogue content.
[0099] In one or more embodiments of this specification, emotional dialogue content can be acquired, analyzed, and processed. When a user makes emotional statements, the specific reasons that trigger the user's emotions can be accurately understood, and empathetic dialogue responses can be generated accurately, thereby improving the accuracy of dialogue content generation and user experience.
[0100] Specifically, emotional dialogue content refers to the user's dialogue content, such as: "I failed my exam today." In the embodiments of this specification, emotional dialogue content includes, but is not limited to, emotional dialogue voice, emotional dialogue image, and emotional dialogue text, which are selected according to the actual situation, and this embodiment of the specification does not impose any limitations on this.
[0101] In practical applications, there are various ways to obtain emotional dialogue content, and the specific method should be selected according to the actual situation. This specification does not impose any limitations on these methods in its embodiments. One possible implementation of this specification involves manually inputting the emotional dialogue content. Another possible implementation involves reading the emotional dialogue content from other data acquisition devices or databases.
[0102] Step 304: Extract at least one emotional keyword from the emotional dialogue content.
[0103] In one or more embodiments of this specification, after obtaining the emotional dialogue content, the emotional dialogue content can be further analyzed to extract at least one emotional keyword from the emotional dialogue content.
[0104] Specifically, emotional keywords in emotional dialogue content refer to keywords related to the user's emotional state during the dialogue, including but not limited to emotional expression words and emotional cause words. Emotional expression words are words that express the user's emotions during the dialogue, such as "happy," "sad," "grief," and "joyful." Emotional cause words are words that trigger the user's emotions during the dialogue. For example, if "failing the exam" is the cause of "sadness," then "failing the exam" is an emotional cause word.
[0105] In practical applications, there are multiple ways to extract at least one emotional keyword from emotional dialogue content. The specific method should be selected according to the actual situation. This specification does not limit the specific methods used in this embodiment.
[0106] In the first possible implementation of this specification, the emotional dialogue content can be matched with a pre-built emotional keyword template to extract at least one emotional keyword from the emotional dialogue content, wherein the emotional keyword template includes multiple emotional keywords.
[0107] In the second possible implementation of this specification, a word extraction model can be used to analyze the emotional dialogue content and extract at least one emotional keyword from the emotional dialogue content. That is, the above extraction of at least one emotional keyword from the emotional dialogue content can include the following steps:
[0108] The emotional dialogue content is input into the word extraction model, and after processing by the word extraction model, at least one emotional keyword in the emotional dialogue content is obtained.
[0109] Specifically, the word extraction model is a machine learning model, trained based on multiple training dialogues and the corresponding sentiment keyword tags for each dialogue. The word extraction model can be a recurrent neural network (RNN), a long short-term memory neural network (LSTM), etc., selected according to the specific circumstances; this specification does not impose any limitations on this selection.
[0110] For example, if the emotional dialogue content "I messed up the exam yesterday" is input into the word extraction model, after processing by the word extraction model, the emotional keywords in the emotional dialogue content "I messed up the exam yesterday" are "exam" and "messed up".
[0111] By applying the solution of the embodiments in this specification, the emotional dialogue content is input into the word extraction model. After processing by the word extraction model, at least one emotional keyword in the emotional dialogue content is obtained, which improves the efficiency and accuracy of obtaining emotional keywords.
[0112] In the third possible implementation of this specification, the emotional dialogue content can be input into a pre-trained text extraction model. After processing by the text extraction model, emotional fragments in the emotional dialogue content can be obtained. Furthermore, the emotional fragments can be input into a word extraction model. After processing by the word extraction model, at least one emotional keyword in the emotional dialogue content can be obtained.
[0113] Step 306: From the pre-constructed sentiment association graph, determine the target subgraph containing at least one sentiment keyword, wherein the sentiment association graph is constructed based on the association relationship between multiple sample sentiment keywords.
[0114] In one or more embodiments of this specification, after obtaining the emotional dialogue content and extracting at least one emotional keyword from the emotional dialogue content, a target subgraph containing at least one emotional keyword can be further determined from a pre-constructed emotional association graph.
[0115] Specifically, the sentiment association graph, also known as the sentiment cause transition graph, constructs a corresponding transition matrix graph using sentiment keywords in each sentence as nodes and contextual adjacency relationships as edges.
[0116] In practical applications, there are multiple ways to determine the target subgraph containing at least one sentiment keyword from a pre-constructed sentiment association graph. The specific method to be selected depends on the actual situation, and the embodiments in this specification do not impose any limitations on this.
[0117] In one possible implementation of this specification, a pre-constructed sentiment association graph and at least one sentiment keyword can be input into a graph segmentation model to obtain a target subgraph. The graph segmentation model is trained based on sample subgraphs in multiple sample association graphs and sample keywords corresponding to each sample subgraph.
[0118] In another possible implementation of this specification, target nodes with edges to sentiment keywords can be found, and target subgraphs can be determined based on the target nodes and sentiment keywords. That is, determining the target subgraph containing at least one sentiment keyword from the pre-constructed sentiment association graph can include the following steps:
[0119] In the sentiment association graph, starting from the sentiment keyword, we find the target node that has an edge with the sentiment keyword;
[0120] By using sentiment keywords and target nodes as nodes in the target subgraph, the target subgraph is obtained by segmenting from the sentiment association graph.
[0121] Specifically, the target nodes that have an edge relationship with the sentiment keyword include the first target node that has a direct edge relationship with the sentiment keyword, and may also include the second target node that has an edge relationship with the first target node. Furthermore, it may also include the third target node that has an edge relationship with the second target node, and so on. The target nodes are selected according to the actual situation, and the embodiments in this specification do not limit them in any way.
[0122] It should be noted that, firstly, the nodes corresponding to the sentiment keywords can be found in the sentiment association graph. Then, starting from the sentiment keywords, the target nodes that have edges with the sentiment keywords are found. The nodes corresponding to the sentiment keywords and the target nodes are then used as nodes in the target subgraph to divide the sentiment association graph and obtain the target subgraph. Each node in the target subgraph is the node corresponding to the sentiment keyword and the target node.
[0123] By applying the scheme of the embodiments in this specification, in the sentiment association graph, starting from the sentiment keyword, the target node with an edge to the sentiment keyword is found; the sentiment keyword and the target node are used as nodes of the target subgraph, and the target subgraph is obtained by segmenting from the sentiment association graph, which improves the accuracy of obtaining the target subgraph.
[0124] In the embodiments of this specification, there are multiple methods for obtaining the sentiment association graph, and the specific method selected depends on the actual situation. This specification does not impose any limitations on this method. In one possible implementation, the sentiment association graph can be read from other data acquisition devices or a database.
[0125] In another possible implementation of this specification, a sample dialogue set can be obtained, and a sentiment association graph can be constructed based on the sample dialogue set. That is, before determining the target subgraph containing at least one sentiment keyword from the pre-constructed sentiment association graph, the following steps may also be included:
[0126] Obtain a sample dialogue set, which includes multiple sample dialogue contents;
[0127] Sentiment recognition is performed on multiple sample dialogue contents to obtain sentiment sample fragments from the multiple sample dialogue contents.
[0128] Keyword identification is performed on sentiment sample fragments to determine sentiment sample words in multiple sample dialogue contents;
[0129] Based on the dialogue relationships of multiple sample dialogues, determine the association between sentiment sample words;
[0130] Using sentiment sample words as nodes and association relationships as edges, construct a sentiment association graph.
[0131] Specifically, an emotional sample fragment refers to a content segment that includes emotional information. For example, if the sample dialogue content is "My girlfriend broke up with me," the emotional sample fragment in the sample dialogue content is "My girlfriend broke up with me." Emotional sample words refer to sample words in the emotional sample fragment that are related to the emotion in the dialogue, including but not limited to sample emotional expression words and sample emotional reason words.
[0132] In the embodiments of this specification, there are multiple ways to obtain the sample dialogue set. One possible implementation is to receive a large amount of sample dialogue content manually input to construct the sample dialogue set. Another possible implementation is to read the sample dialogue set from other data acquisition devices or databases.
[0133] Furthermore, after obtaining the sample dialogue set, emotion recognition can be performed on the content of each sample dialogue in the sample dialogue set to obtain emotional sample fragments in each sample dialogue content.
[0134] There are various ways to perform sentiment recognition on multiple sample dialogues and obtain sentiment sample fragments from them. The specific method should be selected according to the actual situation, and this specification does not limit the specific method. The specific implementation of "performing keyword recognition on sentiment sample fragments to determine sentiment sample words in multiple sample dialogues" is the same as the implementation of "extracting at least one sentiment keyword from sentiment dialogue content" described above, and will not be repeated in this specification.
[0135] In one possible implementation of this specification, sample dialogue content and sample fragment template can be matched to determine the emotional sample fragments in the sample dialogue content, wherein the sample fragment template includes multiple emotional sample fragments.
[0136] In another possible implementation of this specification, a text extraction model can be used to obtain sentiment sample fragments from the sample dialogue content. That is, the above-mentioned sentiment recognition of multiple sample dialogue contents to obtain sentiment sample fragments from multiple sample dialogue contents may include the following steps:
[0137] The sample dialogue content is input into the text extraction model. After processing by the text extraction model, sentiment sample fragments from multiple sample dialogue contents are obtained. The text extraction model is trained based on multiple training dialogue contents and the sentiment tags corresponding to each training dialogue content.
[0138] Specifically, text extraction models refer to machine learning models with text extraction capabilities, such as the SpanBERT pre-trained model. SpanBERT is an extension of the pre-trained model (BERT). Compared to BERT, SpanBERT makes the following changes: instead of masking random features (tokens) like BERT, SpanBERT masks continuous spans. SpanBERT predicts the entire masked span by training the representation of the span's boundaries, rather than individual tokens. It introduces a span-boundary objective (SPO) to encourage the model to store span-level information in the token representations of its boundaries, thereby achieving better results during the fine-tuning stage.
[0139] By applying the scheme of the embodiments of this specification, the sample dialogue content is input into the text extraction model. After processing by the text extraction model, sentiment sample fragments from multiple sample dialogue contents are obtained. The text extraction model is trained based on multiple training dialogue contents and the sentiment tags corresponding to each training dialogue content, thus achieving efficient and accurate acquisition of sentiment sample fragments.
[0140] It should be noted that after identifying the sentiment sample words in multiple sample dialogues, the relationships between the sentiment sample words can be determined based on the dialogue relationships between the multiple sample dialogues. A sentiment association graph can be constructed with sentiment sample words as nodes and relationships as edges.
[0141] Specifically, when determining the relationship between sentiment sample words based on the dialogue relationship of multiple sample dialogue contents, if sentiment sample word B appears in the next sentence of the sample dialogue content of sentiment sample word A, it indicates that there is a relationship between sentiment sample word B and sentiment sample word A. When constructing the graph, a directed edge can be added from sentiment sample word A to sentiment sample word B.
[0142] By applying the scheme of the embodiments in this specification, a sample dialogue set is obtained, wherein the sample dialogue set includes multiple sample dialogue contents; sentiment recognition is performed on the multiple sample dialogue contents to obtain sentiment sample fragments in the multiple sample dialogue contents; keyword recognition is performed on the sentiment sample fragments to determine sentiment sample words in the multiple sample dialogue contents; the association relationship between sentiment sample words is determined according to the dialogue relationship of the multiple sample dialogue contents; with sentiment sample words as nodes and association relationships as edges, a sentiment association graph is constructed to prepare for subsequent determination of target subgraphs containing at least one sentiment keyword, so that dialogue response content can be generated efficiently and accurately.
[0143] Step 308: Based on at least one sentiment keyword and a target subgraph, predict the target keywords in the response content.
[0144] In one or more embodiments of this specification, after obtaining emotional dialogue content, extracting at least one emotional keyword from the emotional dialogue content, and determining a target subgraph containing at least one emotional keyword from a pre-constructed emotional association graph, the target keyword in the response content can be predicted based on the at least one emotional keyword and the target subgraph.
[0145] Specifically, target keywords refer to the keywords in the response content of the generated emotional dialogue. There can be one or more target keywords, depending on the content of the emotional dialogue. This specification does not limit the specific target keywords in this embodiment. For example, if the emotional dialogue content is "My girlfriend broke up with me, and I am very sad," the target keywords would be "love each other" and "together."
[0146] In practical applications, there are multiple ways to predict target keywords in response content based on at least one sentiment keyword and a target subgraph. The specific method to be selected depends on the actual situation, and the embodiments in this specification do not impose any limitations on this.
[0147] In one possible implementation of this specification, a node that has a direct edge with the sentiment keyword can be randomly selected from the target subgraph, and that node can be used as the target keyword.
[0148] In another possible implementation of this specification, a graph attention model can be used to generate target keywords corresponding to the emotional dialogue content. That is, the above-mentioned prediction of target keywords in the response content based on at least one emotional keyword and a target subgraph may include the following steps:
[0149] Obtain the contextual dialogue content of the emotional dialogue;
[0150] Extract at least one contextual sentiment keyword from the contextual dialogue content;
[0151] Based on contextual sentiment keywords and at least one sentiment keyword, determine the sentiment sequence features;
[0152] By inputting the sentiment sequence features and the target subgraph into the graph attention model, the target keywords corresponding to the sentiment dialogue content are obtained.
[0153] Specifically, contextual dialogue content refers to dialogue content that has a contextual relationship with emotional dialogue content. It should be noted that there are multiple ways to obtain the contextual dialogue content of emotional dialogue content, and the specific method should be selected according to the actual situation. This specification does not limit this method in any way. In one possible implementation of this specification, historical dialogue content of emotional dialogue content can be obtained, and contextual dialogue content can be filtered from the historical dialogue content. In another possible implementation of this specification, the contextual dialogue content corresponding to the emotional dialogue content can be read from other data acquisition devices or databases.
[0154] Furthermore, the specific implementation method of "extracting at least one contextual sentiment keyword from the contextual dialogue content" is the same as that of "extracting at least one sentiment keyword from the sentimental dialogue content", and will not be described again in the embodiments of this specification.
[0155] By applying the scheme of the embodiments in this specification, the contextual dialogue content of the emotional dialogue content is obtained; at least one contextual emotional keyword is extracted from the contextual dialogue content; based on the contextual emotional keyword and at least one emotional keyword, the emotional sequence features are determined; the emotional sequence features and the target subgraph are input into a graph attention model to obtain the target keywords corresponding to the emotional dialogue content, thereby realizing the comprehensive consideration of the contextual dialogue content and further improving the accuracy of the target keywords.
[0156] In practical applications, there are multiple ways to determine the features of an emotional sequence based on contextual emotional keywords and at least one emotional keyword. The specific method to be selected depends on the actual situation, and the embodiments in this specification do not impose any limitations on this.
[0157] In one possible implementation of this specification, contextual sentiment keywords and at least one sentiment keyword can be mapped to the same feature space to obtain sentiment sequence features.
[0158] In another possible implementation of this specification, a gated loop unit can be used to process the contextual sentiment keywords and at least one sentiment keyword to obtain sentiment sequence features. That is, the above-mentioned determination of sentiment sequence features based on contextual sentiment keywords and at least one sentiment keyword may include the following steps:
[0159] The contextual sentiment keywords and at least one sentiment keyword are input into a gated recurrent unit. After processing by the gated recurrent unit, the sentiment sequence features are obtained.
[0160] Specifically, the Gated Recurrent Unit (GRU) uses update gates and reset gates. These two gating vectors determine which information is ultimately used as the output of the GRU. The special feature of these gating mechanisms is that they can preserve information from long-term sequences without it being cleared over time or removed because it is irrelevant to the prediction. The inputs to the reset and update gates in the GRU are the hidden states from the previous time step in the input domain of the current time step, and the output is calculated by a fully connected layer with the sigmoid activation function. The sigmoid function transforms the values of elements between 0 and 1. Therefore, the value range of each element in the reset and update gates is [0,1].
[0161] In practical applications, the computational logic in the gated loop unit is shown in the following formula:
[0162] hs i =GRU(hs) i-1 CS i ), i∈[1,n] (1)
[0163]
[0164]
[0165]
[0166] Where i is the sentence sequence in the context dialogue, j is the sequence of sentiment keywords within sentence sequence i, and hs i It is the latent state representation of sentence i incorporating sentiment keywords, that is, sentiment sequence features, cs i It is the sentiment keyword representation of sentence i, derived from the sentiment keyword representation in sentence i. Weighted summation yields α ij It is the weight of sentiment keywords, β ij It is represented by the implicit state of the preceding sentence in the context of the dialogue. With sentiment keywords The calculated weight score, W3 is a learnable parameter.
[0167] By applying the scheme of the embodiments of this specification, contextual sentiment keywords and at least one sentiment keyword are input into a gated loop unit. After processing by the gated loop unit, sentiment sequence features are obtained, thereby improving the accuracy of sentiment sequence features and further improving the accuracy of target keywords.
[0168] Furthermore, the process of inputting the sentiment sequence features and the target subgraph into the graph attention model to obtain the target keywords corresponding to the sentiment dialogue content can include the following steps:
[0169] Input the sentiment sequence features and the target subgraph into the graph attention model to determine the graph attention features;
[0170] Based on graph attention characteristics, determine the subgraph metrics corresponding to the target subgraph;
[0171] Based on the subgraph indicators and the target nodes in the target subgraph, determine the target keywords corresponding to the emotional dialogue content.
[0172] It should be noted that the method of inputting the sentiment sequence features and the target subgraph into the graph attention model to determine the graph attention features is as shown in the following formulas (5)-(7). The method of determining the subgraph index corresponding to the target subgraph based on the graph attention features is as shown in the following formula (8):
[0173] β j =(W4[hdc t ;hs n ]) T ·(W5g i (5)
[0174]
[0175]
[0176]
[0177] Where, α j It is a subplot indicator; β j For graph attention features, graph attention features are derived from sentiment sequence features hs n Sentiment keyword features hdc t Sentiment keyword features related to graph attention aggregation g i The calculated emotional keyword features (hdc) are as follows: t It is to use emotional keywords Enter TRS dec Received; TRS dec It is a decoder based on the Transformer architecture; W5 is a learnable parameter; the sentiment keyword feature g is generated by graph attention aggregation. i Based on the keyword transition probability α jk and emotional keywords Yes; the Transformer is an encoder-decoder model. The Transformer encoder consists of six stacked encoding layers with identical structures, but they do not share weights. The decoder also consists of six stacked decoding layers with identical structures, but they also do not share weights. Each encoding layer is divided into two sub-layers: the first is a self-attention layer, which helps the encoding layer consider other words while encoding a specific word; the second is a feedforward neural network layer. Each decoding layer also has these two layers, but it also has an attention layer to help the decoding layer focus on relevant parts of the input sentence for the encoder.
[0178] Furthermore, when determining the target keywords corresponding to the emotional dialogue content based on the subgraph index and the target nodes in the target subgraph, the subgraph index α can be used... j With concept word transition probability α jk Multiply the results to calculate the node index of each target node in the target subgraph. If the node index is greater than the preset threshold t, then the target node is used as the target keyword and incorporated into the generated dialogue response content.
[0179] By applying the scheme of the embodiments in this specification, the emotional sequence features and the target subgraph are input into the graph attention model to determine the graph attention features; based on the graph attention features, the subgraph index corresponding to the target subgraph is determined; based on the subgraph index and the target node in the target subgraph, the target keywords corresponding to the emotional dialogue content are determined, and the target keywords are accurately generated, further improving the accuracy of the dialogue response content.
[0180] Step 310: Generate dialogue response content based on the emotional dialogue content and target keywords.
[0181] In one or more embodiments of this specification, emotional dialogue content is obtained; at least one emotional keyword is extracted from the emotional dialogue content; a target subgraph containing at least one emotional keyword is determined from a pre-constructed emotional association graph; and after predicting the target keyword in the response content based on the at least one emotional keyword and the target subgraph, dialogue response content can be generated based on the emotional dialogue content and the target keyword.
[0182] Specifically, the dialogue response content refers to the response content corresponding to the emotional dialogue content generated by the aforementioned dialogue content generation method. The dialogue response content can take various forms, including but not limited to dialogue response text, dialogue response audio, dialogue response animation, etc., and the specific form should be selected according to the actual situation. This specification does not impose any limitations on this aspect in the embodiments.
[0183] The method described in this specification involves: acquiring emotional dialogue content; extracting at least one emotional keyword from the emotional dialogue content; determining a target subgraph containing at least one emotional keyword from a pre-constructed emotional association graph, wherein the emotional association graph is constructed based on the association relationships between multiple sample emotional keywords; predicting target keywords in the response content based on at least one emotional keyword and the target subgraph; and generating dialogue response content based on the emotional dialogue content and the target keyword. By constructing an emotional association graph, a deep understanding of the specific keywords that trigger emotions in the dialogue can be achieved, thereby predicting target keywords through emotional keywords, effectively demonstrating empathy in the dialogue response content, and thus realizing empathetic dialogue and improving the accuracy of dialogue content generation.
[0184] In practical applications, there are various ways to generate dialogue responses based on the emotional dialogue content and target keywords. The specific method to be selected depends on the actual situation, and the embodiments in this specification do not impose any limitations on this.
[0185] In one possible implementation of this specification, a dialogue response template corresponding to the target keyword can be determined based on the emotional dialogue content and the target keyword. The target keyword and the dialogue response template can then be merged to generate dialogue response content corresponding to the emotional dialogue content.
[0186] For example, assuming the target keywords are "love" and "together", and the corresponding dialogue response template is determined to be "Don't worry, **we will definitely ***", then by merging the template keywords and the dialogue response template, the dialogue response content "Don't worry, love will definitely be together" can be obtained.
[0187] In another possible implementation of this specification, a response generation model can be used to generate dialogue response content corresponding to the emotional dialogue content. That is, the above-mentioned generation of dialogue response content based on emotional dialogue content and target keywords may include the following steps:
[0188] The target keywords and the contextual dialogue content of the emotional dialogue are input into the cross-attention layer of the response generation model to obtain the fused feature sequence;
[0189] The fused feature sequence is input into the decoder of the response generation model. The fused feature sequence is processed using a copy mechanism to obtain the dialogue response content corresponding to the emotional dialogue content.
[0190] It's important to note that the copy mechanism learns from numerical sequences, with the ultimate goal of ensuring the output sequence matches the input sequence. For example, if the input is [a,b,c,d,e], the output will also be [a,b,c,d,e]. The copy mechanism is used to determine if the model is running correctly and has acquired basic learning capabilities.
[0191] First, we fuse the dialogue context and sentiment concept words in the predicted response using a Transformer-based context encoder. Second, we generate the final response using a Transformer-based response decoder and a copy mechanism, as detailed below:
[0192]
[0193] H dec =Trs dec (H ctx (10)
[0194] P(w)=A h ⊙P copy ·M src +(1-P copy )P gw (w) (11)
[0195] P copy =Sigmoid(W8·H) dec (12)
[0196] P gw (w) = Softmax(W9·H) dec (13)
[0197] Among them, H ctx It is a hidden state representation that contains the context of the dialogue. It is a text output that includes contextual dialogue content and predicted sentiment keywords, TRS dec It is a decoder based on the Transformer model, H dec It is the hidden state representation output after decoding, and P(ω) is the probability of generating a certain word during decoding. copy It is directly from M src According to A in the sentiment keyword matrix h The probability of selecting a word and copying its generation is (1-P) copy ) is the probability P of the word generated by decoding using the model. gw (w), P copy It is a probability between 0 and 1, W9 is a learnable parameter, and Softmax and Sigmoid are both processing functions.
[0198] By applying the scheme of the embodiments in this specification, the target keywords and the contextual dialogue content of the emotional dialogue content are input into the cross-attention layer of the response generation model to obtain a fused feature sequence; the fused feature sequence is then input into the decoder of the response generation model, and the fused feature sequence is processed using a copy mechanism to obtain the dialogue response content corresponding to the emotional dialogue content, thereby enabling the efficient and accurate generation of dialogue response content.
[0199] The following is in conjunction with the appendix Figure 4 Taking the dialogue content generation method provided in this specification in the application of virtual dialogue as an example, the dialogue content generation method will be further explained. Among other things, Figure 4 A flowchart of a virtual dialogue method provided in one embodiment of this specification is shown, which specifically includes the following steps:
[0200] Step 402: Receive a dialogue request sent by the front end, wherein the dialogue request carries emotional dialogue text.
[0201] Step 404: Extract at least one emotional keyword from the emotional dialogue text.
[0202] Step 406: From the pre-constructed sentiment association graph, determine the target subgraph containing at least one sentiment keyword, wherein the sentiment association graph is constructed based on the association relationship between multiple sample sentiment keywords.
[0203] Step 408: Based on at least one sentiment keyword and a target subgraph, predict the target keyword in the response text.
[0204] Step 410: Generate dialogue response text based on the emotional dialogue text and target keywords.
[0205] Step 412: Send the dialogue reply text to the front end so that the front end can display the dialogue reply text.
[0206] It should be noted that the specific implementation methods of steps 404, 406, 408, and 410 are the same as those of steps 304, 306, 308, and 310 above, and will not be described again in the embodiments of this specification.
[0207] The scheme implemented in this specification involves receiving a dialogue request sent by a front-end, wherein the dialogue request carries emotional dialogue text; extracting at least one emotional keyword from the emotional dialogue text; determining a target subgraph containing at least one emotional keyword from a pre-constructed emotional association graph, wherein the emotional association graph is constructed based on the association relationships between multiple sample emotional keywords; predicting target keywords in the response text based on at least one emotional keyword and the target subgraph; generating dialogue response text based on the emotional dialogue text and the target keyword; and sending the dialogue response text to the front-end for display. By constructing an emotional association graph, a deep understanding of the specific keywords that evoke dialogue emotions can be achieved, thereby predicting target keywords through emotional keywords, effectively demonstrating empathy in the dialogue response text, and thus realizing empathetic dialogue and improving the accuracy of dialogue text generation.
[0208] See Figure 5 , Figure 5 This specification illustrates a flowchart of a data processing method for dialogue content according to an embodiment of the present invention. The data processing method for dialogue content is applied to a cloud-side device and specifically includes the following steps:
[0209] Step 502: Obtain the first sample set, wherein the first sample set includes multiple training sentiment dialogue contents, the training sentiment dialogue contents include multiple training sentiment words, and the training sentiment dialogue contents carry response word tags;
[0210] Step 504: Extract features from the training sentiment words corresponding to each training sentiment dialogue content to obtain the training sentiment features corresponding to each training sentiment dialogue content.
[0211] Step 506: Input the training sentiment features and the training sentiment association graph corresponding to each training sentiment dialogue content into the graph attention model to obtain the predicted response words corresponding to each training sentiment dialogue content. The training sentiment association graph is constructed based on the association relationship between multiple training sentiment words.
[0212] Step 508: Train the graph attention model based on the response word tags and predicted response words to obtain the model parameters of the trained graph attention model;
[0213] Step 510: Send the trained graph attention model parameters to the edge device.
[0214] In practical applications, there are multiple ways to obtain the first sample set. It can be formed by manually inputting a large amount of training emotional dialogue content, or it can be formed by reading a large amount of training emotional dialogue content from other data acquisition devices or databases. The specific method of obtaining the first sample set is selected according to the actual situation, and the embodiments in this specification do not limit it in any way.
[0215] It should be noted that the implementation methods of steps 504 and 506 are the same as those described above for "determining emotional sequence features based on contextual emotional keywords and at least one emotional keyword; inputting the emotional sequence features and the target subgraph into the graph attention model to obtain the target keywords corresponding to the emotional dialogue content," and will not be repeated in the embodiments of this specification. After the cloud-side device sends the model parameters of the trained graph attention model to the edge device, the edge device can construct a graph attention model based on the model parameters of the graph attention model, thereby generating dialogue content using the graph attention model.
[0216] Furthermore, when training the graph attention model based on the response word labels and predicted response words, a first loss value can be calculated based on the response word labels and predicted response words. Based on the first loss value, the model parameters of the graph attention model are adjusted, and the process returns to the step of extracting features from the training sentiment words corresponding to each training sentiment dialogue content to obtain the training sentiment features corresponding to each training sentiment dialogue content. When the first preset stopping condition is met, the model parameters of the trained graph attention model are obtained.
[0217] In one possible implementation of this specification, the first preset stopping condition includes a first loss value being less than or equal to a first preset threshold. The first preset threshold is specifically selected based on actual circumstances, and this specification does not impose any limitations on it. The training sentiment features and the training sentiment association graph corresponding to each training sentiment dialogue content are input into the graph attention model to obtain the predicted response words corresponding to each training sentiment dialogue content. After obtaining the predicted response words, a first loss value is calculated based on the predicted response words and response word labels, and the first loss value is compared with the first preset threshold.
[0218] Specifically, if the first loss value is greater than the first preset threshold, it indicates that the difference between the response word label and the predicted response word is large, and the graph attention model has poor prediction ability for the response word. At this time, the model parameters of the graph attention model can be adjusted, and the steps of extracting features from the training sentiment words corresponding to each training sentiment dialogue content can be returned to obtain the training sentiment features corresponding to each training sentiment dialogue content. The graph attention model can continue to be trained until the first loss value is less than or equal to the first preset threshold, indicating that the difference between the response word label and the predicted response word is small, and the first preset stopping condition is reached, and the model parameters of the graph attention model that has been trained are obtained.
[0219] In another possible implementation of this specification, in addition to comparing the magnitude of the first loss value and the first preset threshold, the number of iterations can also be used to determine whether the current graph attention model has been trained.
[0220] Specifically, if the first loss value is greater than the first preset threshold, the parameters of the graph attention model are adjusted, and the process returns to the step of extracting features from the training emotional words corresponding to each training emotional dialogue content to obtain the training emotional features corresponding to each training emotional dialogue content. The graph attention model is then trained again, and the iteration stops when the first preset number of iterations is reached, resulting in a fully trained graph attention model. The first preset number of iterations is selected according to the actual situation, and this embodiment does not limit it in any way.
[0221] In practical applications, there are many functions for calculating the first loss value, such as the cross-entropy loss function, L1 norm loss function, maximum loss function, mean squared error loss function, logarithmic loss function, etc. The specific function to be selected depends on the actual situation, and the embodiments in this specification do not impose any limitations on this. Preferably, the cross-entropy loss function can be used to calculate the first loss value. By using the cross-entropy loss function to calculate the cross-entropy between the response word label and the predicted response word as the first loss value, the efficiency of calculating the first loss value is improved, thereby improving the training efficiency of the graph attention model.
[0222] By applying the scheme of the embodiments of this specification, the training emotional features and the training emotional association graph corresponding to each training emotional dialogue content are input into the graph attention model to obtain the predicted response words corresponding to each training emotional dialogue content. The graph attention model is trained according to the response word labels and the predicted response words to obtain the model parameters of the trained graph attention model. By continuously adjusting the parameters of the graph attention model, the final graph attention model is made more accurate.
[0223] See Figure 6 , Figure 6 This specification illustrates a flowchart of another method for processing dialogue content data according to an embodiment of the present invention. This method is applied to a cloud-side device and specifically includes the following steps:
[0224] Step 602: Obtain the second sample set, which includes multiple training emotional dialogue contents and training response words corresponding to each training emotional dialogue content. The training emotional dialogue contents carry response content tags.
[0225] Step 604: Input multiple training sentiment dialogue contents and the training response words corresponding to each training sentiment dialogue content into the encoder of the response generation model to obtain the predicted encoding representation;
[0226] Step 606: Input the predicted encoding representation into the decoder of the response generation model to obtain the predicted response content corresponding to the trained emotional dialogue content;
[0227] Step 608: Train the response generation model based on the predicted response content and response content tags to obtain the model parameters of the trained response generation model;
[0228] Step 610: Send the trained response generation model parameters to the edge device.
[0229] In practical applications, there are multiple ways to obtain the second sample set. It can be formed by manually inputting a large amount of training emotional dialogue content, or it can be formed by reading a large amount of training emotional dialogue content from other data acquisition devices or databases. The specific method of obtaining the second sample set is selected according to the actual situation, and the embodiments in this specification do not limit it in any way.
[0230] It should be noted that the implementation methods of steps 604 and 606 are the same as those described above: "Inputting the target keywords and the contextual dialogue content of the emotional dialogue into the cross-attention layer of the response generation model to obtain the fused feature sequence; inputting the fused feature sequence into the decoder of the response generation model, and using the copy mechanism to process the fused feature sequence to obtain the emotional response content corresponding to the emotional dialogue content." Therefore, these embodiments in this specification will not be described in detail again. After the cloud-side device sends the model parameters of the trained response generation model to the edge device, the edge device can construct a response generation model based on the model parameters, thereby generating dialogue content using the response generation model.
[0231] Furthermore, when training the response generation model based on the predicted response content and response content labels, a second loss value can be calculated based on the predicted response content and response content labels. Based on the second loss value, the model parameters of the response generation model are adjusted, and the process returns to the step of inputting multiple training emotional dialogue contents and the training response words corresponding to each training emotional dialogue content into the encoder of the response generation model to obtain the predicted encoding representation. When the second preset stopping condition is met, the model parameters of the trained response generation model are obtained.
[0232] In one possible implementation of this specification, the second preset stopping condition includes a second loss value being less than or equal to a second preset threshold. The second preset threshold is specifically selected based on actual circumstances, and this specification does not impose any limitations on it. The predicted encoded representation is input into the decoder of the response generation model to obtain the predicted response content corresponding to the training emotional dialogue content. After obtaining the predicted response content, a second loss value is calculated based on the predicted response content and the response content label, and the second loss value is compared with the second preset threshold.
[0233] Specifically, if the second loss value is greater than the second preset threshold, it indicates that the difference between the predicted response content and the response content label is large, and the response generation model has poor predictive ability for the response content. At this time, the model parameters of the response generation model can be adjusted, and the step of inputting multiple training emotional dialogue contents and the training response words corresponding to each training emotional dialogue content into the encoder of the response generation model can be returned to continue training the response generation model until the second loss value is less than or equal to the second preset threshold, indicating that the difference between the predicted response content and the response content label is small, and the second preset stopping condition is reached, thus obtaining the model parameters of the completed response generation model.
[0234] In another possible implementation of this specification, in addition to comparing the relationship between the second loss value and the second preset threshold, the number of iterations can also be used to determine whether the current response generation model has been trained.
[0235] Specifically, if the second loss value is greater than the second preset threshold, the parameters of the response generation model are adjusted, and the process returns to the step of inputting multiple training emotional dialogue contents and the training response words corresponding to each training emotional dialogue content into the encoder of the response generation model, and the response generation model continues to be trained. When the second preset number of iterations is reached, the iteration stops, and the trained response generation model is obtained. The second preset number of iterations is selected according to the actual situation, and this specification embodiment does not limit it in any way.
[0236] In practical applications, there are many functions for calculating the second loss value, such as the cross-entropy loss function, L1 norm loss function, maximum loss function, mean squared error loss function, logarithmic loss function, etc. The specific function to be selected depends on the actual situation, and the embodiments in this specification do not impose any limitations on this. Preferably, the cross-entropy loss function can be used to calculate the second loss value. By using the cross-entropy loss function to calculate the cross-entropy between the predicted response content and the response content label as the second loss value, the efficiency of calculating the second loss value is improved, thereby improving the training efficiency of the response generation model.
[0237] Applying the scheme of the embodiments in this specification, multiple training emotional dialogue contents and the corresponding training response words are input into the encoder of the response generation model to obtain a predictive encoding representation; the predictive encoding representation is input into the decoder of the response generation model to obtain the predicted response content corresponding to the training emotional dialogue contents; the response generation model is trained based on the predicted response content and the response content labels to obtain the model parameters of the trained response generation model; by continuously adjusting the parameters of the response generation model, the final response generation model becomes more accurate.
[0238] See Figure 7 , Figure 7A flowchart illustrating the processing procedure of a dialogue content generation method provided in one embodiment of this specification is shown.
[0239] Construction of the sentiment graph: The required sentiment graph is constructed based on a massive corpus of dialogues. First, sentiment sample fragments are extracted from multiple sample dialogues using SpanBERT. Second, a keyword discovery algorithm is used to identify keywords in the sentiment sample fragments, determining the sentiment sample words in the multiple sample dialogues. Finally, based on the adjacency characteristics of the context of the sample dialogues, with sentiment sample words as nodes, if keyword B appears in the next sentence after keyword A, a directed edge is added from keyword A to keyword B, thus constructing the sentiment graph.
[0240] Prediction of target keywords in response content: First, the user inputs emotional dialogue content, and at least one emotional keyword is extracted from the emotional dialogue content; correspondingly, the contextual emotional keywords are extracted from the contextual dialogue content of the emotional dialogue content, and sequence modeling is performed through a gated recurrent unit to obtain emotional sequence features; second, the corresponding target subgraph is obtained through graph retrieval, the emotional sequence features and the target subgraph are fused, and the target keywords corresponding to the emotional dialogue content are predicted through a graph attention model.
[0241] The generation of dialogue response content is as follows: First, the contextual dialogue content and target keywords of the emotional dialogue content are fused through a Transformer-based context encoder; second, the emotional response content is generated through a Transformer-based response decoder and copy mechanism.
[0242] The solution applied in the embodiments of this specification extracts emotional keywords that evoke user emotions from the user's dialogue history and explicitly models the causes of emotions based on contextual adjacency, obtaining an emotional association graph. By understanding the specific reasons that trigger user emotions, empathy can be better achieved. By predicting emotional keywords in responses through the emotional association graph, empathy can be effectively demonstrated in responses, thereby achieving empathetic dialogue. In other words, after incorporating the emotional association graph, the responses generated by the embodiments of this specification have significantly improved empathy, greatly enhancing the accuracy of dialogue and the user's dialogue experience, thus creating a warm and emotionally resonant chatbot.
[0243] See Figure 8 , Figure 8 This document illustrates a flowchart of the construction process for a sentiment graph in a dialogue content generation method according to an embodiment of this specification. The sample dialogue content is obtained as follows:
[0244] User: I messed up my exam yesterday.
[0245] Machine: What's wrong? Your grades have always been excellent.
[0246] User: My girlfriend broke up with me, and I'm very sad.
[0247] Machine: Don't worry, two people who love each other will definitely be together.
[0248] The process involves performing sentiment recognition on the sample dialogue content to obtain sentiment sample fragments; identifying keywords in the sentiment sample fragments to determine sentiment sample words in the sample dialogue content; determining the association relationships between sentiment sample words based on the dialogue relationships in the sample dialogue content; and constructing a sentiment association graph with sentiment sample words as nodes and association relationships as edges. The sentiment association graph includes nodes such as "excellent", "mess up", "break up", "together", "girlfriend", "love", "grades", and "exam".
[0249] By applying the solutions in the embodiments of this specification, emotional keywords that evoke user emotions in the user's dialogue history are extracted, and the emotional causes are explicitly modeled based on the contextual adjacency characteristics to obtain an emotional association graph. This allows for a deep understanding of the specific reasons that evoke user emotions, effectively demonstrating empathy in responses, and thus achieving empathetic dialogue.
[0250] See Figure 9 , Figure 9 This diagram illustrates a virtual dialogue interface according to an embodiment of this specification. The virtual dialogue interface includes an emotional dialogue content input box, an "OK" control, a "Cancel" control, and a dialogue response content display box. The user inputs emotional dialogue content through the emotional dialogue content input box displayed on the front end and clicks the "OK" control. The server extracts at least one emotional keyword from the emotional dialogue content; it determines a target subgraph containing at least one emotional keyword from a pre-constructed emotional association graph, wherein the emotional association graph is constructed based on the association relationships between multiple sample emotional keywords; based on at least one emotional keyword and the target subgraph, it predicts the target keyword in the response content; based on the emotional dialogue content and the target keyword, it generates dialogue response content and sends it to the front end so that the front end displays the dialogue response content in the dialogue response content display box. It should be noted that the user can operate the controls in any way, including clicking, double-clicking, touching, mouse hovering, swiping, long-pressing, voice control, or shaking, etc., and the specific method selected depends on the actual situation. This embodiment of the specification does not limit this in any way.
[0251] The method described in this specification involves: acquiring emotional dialogue content; extracting at least one emotional keyword from the emotional dialogue content; determining a target subgraph containing at least one emotional keyword from a pre-constructed emotional association graph, wherein the emotional association graph is constructed based on the association relationships between multiple sample emotional keywords; predicting target keywords in the response content based on at least one emotional keyword and the target subgraph; and generating dialogue response content based on the emotional dialogue content and the target keyword. By constructing an emotional association graph, a deep understanding of the specific keywords that trigger emotions in the dialogue can be achieved, thereby predicting target keywords through emotional keywords, effectively demonstrating empathy in the dialogue response content, and thus realizing empathetic dialogue and improving the accuracy of dialogue content generation.
[0252] It should be noted that the emotional dialogue content, emotional association graph, sample dialogue set, dialogue request, emotional dialogue text, first sample set, second sample set, and other information and data involved in the above method embodiments are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0253] Corresponding to the above-described embodiments of the dialogue content generation method, this specification also provides embodiments of the dialogue content generation apparatus. Figure 10 A schematic diagram of a dialogue content generation apparatus according to one embodiment of this specification is shown. Figure 10 As shown, the device includes:
[0254] The first acquisition module 1002 is configured to acquire emotional dialogue content;
[0255] The first extraction module 1004 is configured to extract at least one emotional keyword from the emotional dialogue content;
[0256] The first determining module 1008 is configured to determine a target subgraph containing at least one sentiment keyword from a pre-constructed sentiment association graph, wherein the sentiment association graph is constructed based on the association relationship between multiple sample sentiment keywords;
[0257] The first prediction module 1010 is configured to predict target keywords in the response content based on at least one sentiment keyword and a target subgraph.
[0258] The first generation module 1012 is configured to generate dialogue response content based on the emotional dialogue content and target keywords.
[0259] Optionally, the device further includes: a construction module configured to acquire a sample dialogue set, wherein the sample dialogue set includes multiple sample dialogue contents; perform sentiment recognition on the multiple sample dialogue contents to obtain sentiment sample fragments in the multiple sample dialogue contents; perform keyword recognition on the sentiment sample fragments to determine sentiment sample words in the multiple sample dialogue contents; determine the association relationship between sentiment sample words based on the dialogue relationship of the multiple sample dialogue contents; and construct a sentiment association graph with sentiment sample words as nodes and association relationships as edges.
[0260] Optionally, the building module is further configured to input sample dialogue content into a text extraction model, and after processing by the text extraction model, obtain sentiment sample fragments from multiple sample dialogue contents, wherein the text extraction model is trained based on multiple training dialogue contents and the sentiment tags corresponding to each training dialogue content.
[0261] Optionally, the first acquisition module 1002 is further configured to input the emotional dialogue content into the word extraction model, and after processing by the word extraction model, obtain at least one emotional keyword in the emotional dialogue content.
[0262] Optionally, it is further configured to, in the sentiment association graph, starting from the sentiment keyword, find the target node that has an edge with the sentiment keyword; and use the sentiment keyword and the target node as nodes of the target subgraph to segment the sentiment association graph to obtain the target subgraph.
[0263] Optionally, the first prediction module 1010 is further configured to: acquire the contextual dialogue content of the emotional dialogue content; extract at least one contextual sentiment keyword from the contextual dialogue content; determine the sentiment sequence features based on the contextual sentiment keyword and at least one sentiment keyword; and input the sentiment sequence features and the target subgraph into a graph attention model to obtain the target keyword corresponding to the emotional dialogue content.
[0264] Optionally, the first prediction module 1010 is further configured to input contextual sentiment keywords and at least one sentiment keyword into a gated recurrent unit, and obtain sentiment sequence features through processing by the gated recurrent unit.
[0265] Optionally, the first prediction module 1010 is further configured to input the sentiment sequence features and the target subgraph into the graph attention model to determine the graph attention features; based on the graph attention features, determine the subgraph index corresponding to the target subgraph; and based on the subgraph index and the target node in the target subgraph, determine the target keywords corresponding to the sentiment dialogue content.
[0266] Optionally, the first generation module 1012 is further configured to input the target keywords and the contextual dialogue content of the emotional dialogue content into the cross-attention layer of the response generation model to obtain a fused feature sequence; input the fused feature sequence into the decoder of the response generation model, and use a copy mechanism to process the fused feature sequence to obtain the dialogue response content corresponding to the emotional dialogue content.
[0267] The method described in this specification involves: acquiring emotional dialogue content; extracting at least one emotional keyword from the emotional dialogue content; determining a target subgraph containing at least one emotional keyword from a pre-constructed emotional association graph, wherein the emotional association graph is constructed based on the association relationships between multiple sample emotional keywords; predicting target keywords in the response content based on at least one emotional keyword and the target subgraph; and generating dialogue response content based on the emotional dialogue content and the target keyword. By constructing an emotional association graph, a deep understanding of the specific keywords that trigger emotions in the dialogue can be achieved, thereby predicting target keywords through emotional keywords, effectively demonstrating empathy in the dialogue response content, and thus realizing empathetic dialogue and improving the accuracy of dialogue content generation.
[0268] The above is an illustrative scheme of a dialogue content generation device according to this embodiment. It should be noted that the technical solution of this dialogue content generation device and the technical solution of the dialogue content generation method described above belong to the same concept. For details not described in detail in the technical solution of the dialogue content generation device, please refer to the description of the technical solution of the dialogue content generation method described above.
[0269] Corresponding to the above-described virtual dialogue method embodiments, this specification also provides virtual dialogue device embodiments. Figure 11 A schematic diagram of the structure of a virtual dialogue device according to one embodiment of this specification is shown. Figure 11 As shown, the device includes:
[0270] The first receiving module 1102 is configured to receive a dialogue request sent by the front end, wherein the dialogue request carries emotional dialogue text.
[0271] The second extraction module 1104 is configured to extract at least one sentiment keyword from the sentiment dialogue text;
[0272] The second determining module 1106 is configured to determine a target subgraph containing at least one sentiment keyword from a pre-constructed sentiment association graph, wherein the sentiment association graph is constructed based on the association relationships between multiple sample sentiment keywords;
[0273] The second prediction module 1108 is configured to predict target keywords in the response text based on at least one sentiment keyword and a target subgraph.
[0274] The second generation module 1110 is configured to generate dialogue response text based on the emotional dialogue text and target keywords;
[0275] The first sending module 1112 is configured to send the dialogue reply text to the front end so that the front end can display the dialogue reply text.
[0276] The scheme implemented in this specification involves receiving a dialogue request sent by a front-end, wherein the dialogue request carries emotional dialogue text; extracting at least one emotional keyword from the emotional dialogue text; determining a target subgraph containing at least one emotional keyword from a pre-constructed emotional association graph, wherein the emotional association graph is constructed based on the association relationships between multiple sample emotional keywords; predicting target keywords in the response text based on at least one emotional keyword and the target subgraph; generating dialogue response text based on the emotional dialogue text and the target keyword; and sending the dialogue response text to the front-end for display. By constructing an emotional association graph, a deep understanding of the specific keywords that evoke dialogue emotions can be achieved, thereby predicting target keywords through emotional keywords, effectively demonstrating empathy in the dialogue response text, and thus realizing empathetic dialogue and improving the accuracy of dialogue text generation.
[0277] The above is an illustrative scheme of a virtual dialogue device according to this embodiment. It should be noted that the technical solution of this virtual dialogue device and the technical solution of the above-described virtual dialogue method belong to the same concept. For details not described in detail in the technical solution of the virtual dialogue device, please refer to the description of the technical solution of the above-described virtual dialogue method.
[0278] Corresponding to the above-described embodiments of the data processing method for dialogue content, this specification also provides embodiments of the data processing apparatus for dialogue content. Figure 12 A schematic diagram of a data processing apparatus for dialogue content provided in one embodiment of this specification is shown. Figure 12 As shown, this device is applied to cloud-side equipment and includes:
[0279] The second acquisition module 1202 is configured to acquire a first sample set, wherein the first sample set includes multiple training emotional dialogue contents, the training emotional dialogue contents include multiple training emotional words, and the training emotional dialogue contents carry response word tags;
[0280] The third extraction module 1204 is configured to extract features from the training emotional words corresponding to each training emotional dialogue content, and obtain the training emotional features corresponding to each training emotional dialogue content.
[0281] The first input module 1206 is configured to input the training sentiment features and the training sentiment association graph corresponding to each training sentiment dialogue content into the graph attention model to obtain the predicted response words corresponding to each training sentiment dialogue content. The training sentiment association graph is constructed based on the association relationship between multiple training sentiment words.
[0282] The first training module 1208 is configured to train the graph attention model based on the response word tags and the predicted response words, and obtain the model parameters of the trained graph attention model.
[0283] The second sending module 1210 is configured to send the model parameters of the trained graph attention model to the end-side device.
[0284] Applying the scheme of the embodiments of this specification, a first sample set is obtained, wherein the first sample set includes multiple training emotional dialogue contents, each training emotional dialogue content includes multiple training emotional words, and each training emotional dialogue content carries a response word tag; features are extracted from the training emotional words corresponding to each training emotional dialogue content to obtain training emotional features corresponding to each training emotional dialogue content; the training emotional features and the training emotional association graph corresponding to each training emotional dialogue content are input into a graph attention model to obtain the predicted response word corresponding to each training emotional dialogue content, wherein the training emotional association graph is constructed based on the association relationship between multiple training emotional words; the graph attention model is trained according to the response word tag and the predicted response word to obtain the model parameters of the trained graph attention model; the model parameters of the trained graph attention model are sent to the end device. By continuously adjusting the parameters of the graph attention model, the final graph attention model becomes more accurate. The above is an illustrative scheme of a dialogue content data processing device according to this embodiment. It should be noted that the technical solution of this dialogue content data processing device and the technical solution of the above-described dialogue content data processing method belong to the same concept. Details not described in detail in the technical solution of the dialogue content data processing device can be found in the description of the technical solution of the above-described dialogue content data processing method.
[0285] Corresponding to the above-described embodiments of the data processing method for dialogue content, this specification also provides embodiments of the data processing apparatus for dialogue content. Figure 13 A schematic diagram of another data processing apparatus for dialogue content provided in one embodiment of this specification is shown. Figure 13 As shown, this device is applied to cloud-side equipment and includes:
[0286] The third acquisition module 1302 is configured to acquire a second sample set, wherein the second sample set includes multiple training emotional dialogue contents and training response words corresponding to each training emotional dialogue content, and the training emotional dialogue contents carry response content tags;
[0287] The second input module 1304 is configured to input multiple training emotional dialogue contents and the training response words corresponding to each training emotional dialogue content into the encoder of the response generation model to obtain the predicted encoding representation;
[0288] The third input module 1306 is configured to input the predicted response content into the decoder of the response generation model to obtain the predicted response content corresponding to the training emotional dialogue content;
[0289] The second training module 1308 is configured to train the response generation model based on the predicted response content and the response content tags, and obtain the model parameters of the trained response generation model.
[0290] The third sending module 1310 is configured to send the model parameters of the trained response generation model to the end device.
[0291] Applying the scheme of the embodiments in this specification, a second sample set is obtained, wherein the second sample set includes multiple training emotional dialogue contents and training response words corresponding to each training emotional dialogue content, and the training emotional dialogue contents carry response content tags; the multiple training emotional dialogue contents and the training response words corresponding to each training emotional dialogue content are input into the encoder of the response generation model to obtain a predictive encoding representation; the predictive encoding representation is input into the decoder of the response generation model to obtain the predicted response content corresponding to the training emotional dialogue content; the response generation model is trained based on the predicted response content and the response content tags to obtain the model parameters of the trained response generation model; the model parameters of the trained response generation model are sent to the edge device. By continuously adjusting the parameters of the response generation model, the final response generation model becomes more accurate.
[0292] The above is an illustrative scheme of a data processing device for dialogue content according to this embodiment. It should be noted that the technical solution of this data processing device for dialogue content belongs to the same concept as the technical solution of the data processing method for dialogue content described above. For details not described in detail in the technical solution of the data processing device for dialogue content, please refer to the description of the technical solution of the data processing method for dialogue content described above.
[0293] Figure 14 A structural block diagram of a computing device according to one embodiment of this specification is shown. The components of the computing device 1400 include, but are not limited to, a memory 1410 and a processor 1420. The processor 1420 is connected to the memory 1410 via a bus 1430, and a database 1450 is used to store data.
[0294] The computing device 1400 also includes an access device 1440, which enables the computing device 1400 to communicate via one or more networks 1460. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1440 may include one or more of any type of wired or wireless network interface (e.g., Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Networks (WLAN) interface, a Wi-MAX (World Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0295] In one embodiment of this specification, the above-described components of the computing device 1400 and Figure 14 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 14 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0296] The computing device 1400 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1400 can also be a mobile or stationary server.
[0297] The processor 1420 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-described dialogue content generation method, virtual dialogue method, or dialogue content data processing method.
[0298] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the technical solutions of the dialogue content generation method, the virtual dialogue method, and the dialogue content data processing method described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the dialogue content generation method, the virtual dialogue method, or the dialogue content data processing method described above.
[0299] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described dialogue content generation method, virtual dialogue method, or dialogue content data processing method.
[0300] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solutions of the dialogue content generation method, the virtual dialogue method, and the dialogue content data processing method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the dialogue content generation method, the virtual dialogue method, or the dialogue content data processing method described above.
[0301] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described dialogue content generation method, virtual dialogue method, or dialogue content data processing method.
[0302] The above is an illustrative scheme of a computer program according to this embodiment. It should be noted that the technical solution of this computer program belongs to the same concept as the technical solutions of the above-mentioned dialogue content generation method, virtual dialogue method, and dialogue content data processing method. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solutions of the above-mentioned dialogue content generation method, virtual dialogue method, or dialogue content data processing method.
[0303] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0304] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0305] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0306] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0307] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for generating dialogue content, comprising: Obtain emotional conversation content; Extract at least one emotional keyword from the emotional dialogue content, wherein the emotional keyword includes at least emotional expression words and emotional reason words, and the emotional reason words refer to words that evoke the user's emotional response in the dialogue; From a pre-constructed sentiment association graph, a target subgraph containing the at least one sentiment keyword is determined, wherein the sentiment association graph is constructed using sentiment keywords in each sentence of the sample dialogue set as nodes and contextual adjacency relationships as edges; Based on the at least one sentiment keyword and the target subgraph, predict the target keywords in the response content, wherein the target keywords are obtained by filtering from the target subgraph; Based on the emotional dialogue content and the target keywords, generate dialogue response content.
2. The method according to claim 1, further comprising, before determining the target subgraph containing the at least one sentiment keyword from the pre-constructed sentiment association graph: Obtain a sample dialogue set, wherein the sample dialogue set includes multiple sample dialogue contents; Perform sentiment recognition on the multiple sample dialogue contents to obtain sentiment sample fragments in the multiple sample dialogue contents; Keyword recognition is performed on the emotional sample fragments to determine the emotional sample words in the multiple sample dialogue contents; Based on the dialogue relationships of the multiple sample dialogue contents, the association relationships between the sentiment sample words are determined; Using the sentiment sample words as nodes and the association relationships as edges, a sentiment association graph is constructed.
3. The method according to claim 2, wherein performing emotion recognition on the plurality of sample dialogue contents to obtain emotion sample fragments from the plurality of sample dialogue contents includes: The sample dialogue content is input into the text extraction model. After processing by the text extraction model, sentiment sample fragments from the multiple sample dialogue contents are obtained. The text extraction model is trained based on multiple training dialogue contents and the sentiment tags corresponding to each training dialogue content.
4. The method according to claim 1, wherein extracting at least one emotional keyword from the emotional dialogue content includes: The emotional dialogue content is input into the word extraction model, and after processing by the word extraction model, at least one emotional keyword in the emotional dialogue content is obtained.
5. The method according to claim 1 or 2, wherein determining the target subgraph containing the at least one sentiment keyword from a pre-constructed sentiment association graph comprises: In the sentiment association graph, starting from the sentiment keyword, we find the target node that has an edge with the sentiment keyword; The emotional keywords and the target nodes are used as nodes in the target subgraph, and the target subgraph is obtained by segmenting from the emotional association graph.
6. The method according to claim 1, wherein predicting the target keyword in the response content based on the at least one sentiment keyword and the target subgraph comprises: Obtain the contextual dialogue content of the emotional dialogue content; Extract at least one contextual sentiment keyword from the contextual dialogue content; Based on the contextual sentiment keywords and the at least one sentiment keyword, determine the sentiment sequence features; The emotional sequence features and the target subgraph are input into a graph attention model to obtain the target keywords corresponding to the emotional dialogue content.
7. The method according to claim 6, wherein determining the sentiment sequence features based on the contextual sentiment keywords and the at least one sentiment keyword comprises: The contextual sentiment keywords and at least one sentiment keyword are input into a gated loop unit, and the sentiment sequence features are obtained through processing by the gated loop unit.
8. The method according to claim 6, wherein inputting the emotional sequence features and the target subgraph into a graph attention model to obtain the target keywords corresponding to the emotional dialogue content includes: The emotional sequence features and the target subgraph are input into the graph attention model to determine the graph attention features; Based on the graph attention features, determine the subgraph index corresponding to the target subgraph; Based on the subgraph indicators and the target nodes in the target subgraph, the target keywords corresponding to the emotional dialogue content are determined.
9. The method according to claim 1, wherein generating dialogue response content based on the emotional dialogue content and the target keywords includes: The target keywords and the contextual dialogue content of the emotional dialogue content are input into the cross-attention layer of the response generation model to obtain a fused feature sequence; The fused feature sequence is input into the decoder of the response generation model, and the fused feature sequence is processed using a copy mechanism to obtain the dialogue response content corresponding to the emotional dialogue content.
10. A virtual dialogue method, comprising: Receive a dialogue request sent by the front end, wherein the dialogue request carries emotional dialogue text; Extract at least one emotional keyword from the emotional dialogue text, wherein the emotional keyword includes at least emotional expression words and emotional cause words, and the emotional cause words refer to words that evoke the user's emotional response in the dialogue; From a pre-constructed sentiment association graph, a target subgraph containing the at least one sentiment keyword is determined, wherein the sentiment association graph is constructed using sentiment keywords in each sentence of the sample dialogue set as nodes and contextual adjacency relationships as edges; Based on the at least one sentiment keyword and the target subgraph, predict the target keyword in the response text, wherein the target keyword is obtained by filtering from the target subgraph; Generate dialogue response text based on the emotional dialogue text and the target keywords; The dialogue reply text is sent to the front end so that the front end displays the dialogue reply text.
11. A method for processing dialogue content, applied to a cloud-side device, the method comprising: Obtain a first sample set, wherein the first sample set includes multiple training emotional dialogue contents, the training emotional dialogue contents include multiple training emotional words, the training emotional dialogue contents carry response word tags, and the training emotional words include at least emotional expression words and emotional reason words, the emotional reason words refer to words that evoke the user's emotional dialogue. Feature extraction is performed on the training emotional words corresponding to each training emotional dialogue content to obtain the training emotional features corresponding to each training emotional dialogue content. The training sentiment features and the training sentiment association graph corresponding to each training sentiment dialogue content are input into the graph attention model to obtain the predicted response words corresponding to each training sentiment dialogue content. The training sentiment association graph is constructed using sentiment keywords in each sentence of the first sample set as nodes and contextual adjacency relationships as edges. The predicted response words are selected from the training sentiment association graph corresponding to the training sentiment dialogue content. The graph attention model is trained based on the response word tags and the predicted response words to obtain the model parameters of the trained graph attention model; The trained graph attention model parameters are sent to the edge device.
12. A method for processing dialogue content, applied to a cloud-side device, the method comprising: Obtain a second sample set, wherein the second sample set includes multiple training emotional dialogue contents and training response words corresponding to each training emotional dialogue content, the training emotional dialogue contents carry response content tags, and the training emotional dialogue contents include at least emotional expression words and emotional reason words, the emotional reason words being words that evoke the user's emotional dialogue; The multiple training sentiment dialogue contents and the corresponding training response words are input into the encoder of the response generation model to obtain a predicted encoding representation. The training response words are predicted based on the training sentiment association graph and the training sentiment dialogue contents. The training sentiment association graph is constructed with sentiment keywords in each sentence of the second sample set as nodes and contextual adjacency relationships as edges. The predicted response words are selected from the training sentiment association graph. The predicted encoding representation is input into the decoder of the response generation model to obtain the predicted response content corresponding to the trained emotional dialogue content; The response generation model is trained based on the predicted response content and the response content tags to obtain the model parameters of the trained response generation model; The model parameters of the trained response generation model are sent to the edge device.
13. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 9, or claim 10, or claim 11, or claim 12.
14. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method of any one of claims 1 to 9, or claim 10, or claim 11, or claim 12.
15. A computer program product comprising computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 9, or claim 10, or claim 11, or claim 12.
Citation Information
Patent Citations
Emotion guiding method and system based on emotion semantic transfer graph
CN111914556A
Knowledge and emotion integrated end-to-end dialogue method based on variational auto-encoder
CN114610861A
Visual question and answer method and device based on graph attention neural network and visual relation
CN115588193A