Self-adaptive dialogue generation method based on personalized memory graph and related device

By building a personalized memory map on the intelligent outgoing call robot all-in-one machine, using voice emotion and semantic analysis technology, key information in user conversations is extracted and adaptive conversation audio data is generated, which solves the problem that the intelligent outgoing call robot all-in-one machine cannot continue to chat, and achieves a more realistic and continuous chat experience.

CN120412653APending Publication Date: 2025-08-01GUANGDONG BAIYUN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510377786.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing intelligent outgoing robot all-in-one machine cannot realize continuous chat with users, resulting in poor authenticity of chats, easy repetition or inability to continue chat topics.

Method used

By constructing a personalized memory map on the intelligent outgoing call robot integrated machine, using voice sentiment and semantic analysis technology, key information in user conversations is extracted, and adaptive conversation audio data is matched and generated in the personalized memory map.

Benefits of technology

It realizes local lightweight storage of the intelligent outgoing robot all-in-one machine, reducing storage pressure, avoiding repetition of conversations, and improving the authenticity and continuity of chats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412653A_ABST
    Figure CN120412653A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive dialogue generation method based on a personalized memory map and a related device, and the method comprises the steps: enabling a long-term memory database on a far-end server to be associated with an intelligent outbound robot all-in-one machine after logging in the far-end server; loading the personalized memory map to the local of the intelligent outbound robot all-in-one machine for local storage; the method comprises the following steps: when entering a chat conversation mode, collecting conversation audio data of a service object user, and extracting audio emotion information, semantic information and semantic keywords in the conversation audio data; matching knowledge nodes of the semantic keywords in the personalized memory graph; and dialogue audio data corresponding to the dialogue audio data is generated based on the semantic information, the knowledge nodes and the audio emotion information, and dialogue playing processing is carried out. In the embodiment of the invention, the lightweight local storage of the intelligent outbound robot all-in-one machine is realized, the conversation with the user is realized through the personalized memory map, and the sense of reality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to an adaptive dialogue generation method and related device based on a personalized memory graph. Background Art

[0002] In some existing elderly care devices, an intelligent outbound call robot integrated machine generally only integrates some outbound call functions and does not integrate a chatting function with service object users. It cannot implement the chatting function of the outbound call device, or integrates some simple chatting modules on the outbound call device to implement simple conversations or greetings. However, because the continuation of chat memory cannot be achieved, during conversations or greetings, repetition may not be avoided or the corresponding chat topics cannot be continued, resulting in poor authenticity of the chat and weak user experience. Summary of the Invention

[0003] The purpose of the present invention is to overcome the deficiencies of the prior art. The present invention provides an adaptive dialogue generation method and related device based on a personalized memory graph, which realizes the lightweight local storage of the intelligent outbound call robot integrated machine, and realizes conversations with users through the personalized memory graph, and improves the sense of reality.

[0004] To solve the above technical problems, an embodiment of the present invention provides an adaptive dialogue generation method based on a personalized memory graph, which is applied to an intelligent outbound call robot integrated machine. The intelligent outbound call robot integrated machine is communicatively connected to a remote server. The method includes:

[0005] After a service object user logs in to the remote server through the intelligent outbound call robot integrated machine, the long-term memory database on the remote server is associated with the intelligent outbound call robot integrated machine based on the login information of the service object user; meanwhile,

[0006] The remote server loads the personalized memory graph stored in the long-term memory database to the local of the intelligent outbound call robot integrated machine for local storage;

[0007] When the intelligent outbound call robot integrated machine enters the chat dialogue mode, the dialogue audio data of the service object user is collected based on an audio collection device, and the audio emotion information, semantic information, and semantic keywords in the dialogue audio data are extracted;

[0008] The semantic keywords are used for matching processing in the personalized memory graph stored locally, and the knowledge nodes of the semantic keywords in the personalized memory graph are matched;

[0009] Adaptive generation of the dialogue audio data corresponding to the said dialogue audio data based on the said semantic information, the said knowledge nodes and the said audio emotion information, and carrying out dialogue playback processing.

[0010] Optionally, the association of the long-term memory database on the said remote server with the intelligent outbound robot all-in-one machine based on the login information of the said service object user includes:

[0011] The said remote server obtains the user identifier in the said login information of the said service object user, the user identifier is unique and is assigned when the said service object user registers on the said remote server;

[0012] The said remote server indexes the long-term memory database bound to the said user identifier by using the said user identifier. The long-term memory database is a database created by the said remote server for the said service object user by using the said user identifier. The long-term memory database stores the user information of the said service object user and the key information captured when the said service object user has a dialogue with the intelligent outbound robot all-in-one machine. The key information includes keywords and the semantic information corresponding to the keywords;

[0013] The said remote server associates the long-term memory database with the local account of the said service object user on the intelligent outbound robot all-in-one machine by using the said user identifier. The local account is an account created locally on the intelligent outbound robot all-in-one machine by using the user identifier in the login information of the said service object user.

[0014] Optionally, the said personalized memory graph is also stored in the long-term memory database;

[0015] The said personalized memory graph is a knowledge graph constructed by using the user information stored in the said long-term memory database and the key information captured when the service object user has a dialogue with the intelligent outbound robot all-in-one machine. And when the key information stored in the long-term memory database is updated, the personalized memory graph is updated by using the updated key information.

[0016] Optionally, the said remote server loads the personalized memory graph stored in the long-term memory database to the local of the intelligent outbound robot all-in-one machine for local storage, including:

[0017] The said remote server loads the personalized memory graph stored in the long-term memory database to the intelligent outbound robot all-in-one machine through a communication connection;

[0018] The intelligent outbound call robot all-in-one machine indexes to the corresponding local storage space through the association of the long-term memory database with the local account, and updates and stores the personalized memory map in the local storage space. The local storage space is the storage space opened by the intelligent outbound call robot all-in-one machine for the local account in the local storage device.

[0019] Optionally, extracting the audio emotion information, semantic information, and semantic keywords in the conversation audio data includes:

[0020] The intelligent outbound call robot all-in-one machine performs tone and speech rate analysis processing on the conversation audio data to obtain the tone data and speech rate data corresponding to the audio data;

[0021] Based on the emotion analysis model, audio emotion analysis processing is performed using the tone data and speech rate data corresponding to the audio data to obtain the audio emotion information corresponding to the audio data. The emotion analysis model is trained by a deep neural network model using the tone data and speech rate data marked with audio emotion information corresponding to historical audio data;

[0022] Convert the conversation audio data into conversation text data, and input the conversation text data and the audio emotion information into a natural language analysis model for semantic analysis processing to obtain the semantic information of the conversation audio data under the audio emotion information;

[0023] Use a keyword extraction algorithm to perform keyword extraction processing on the semantic information to obtain the semantic keywords corresponding to the semantic information.

[0024] Optionally, using the semantic keywords to perform matching processing in the personalized memory map stored locally and matching to the knowledge nodes of the semantic keywords in the personalized memory map includes:

[0025] The intelligent outbound call robot all-in-one machine retrieves the personalized memory map stored locally and uses the semantic keywords to perform matching processing with the key information in each knowledge node in the personalized memory map to obtain a matching result;

[0026] Obtain the knowledge nodes of the semantic keywords in the personalized memory map through the matching result.

[0027] Optionally, adaptively generating the conversation audio data corresponding to the conversation audio data based on the semantic information, the knowledge nodes, and the audio emotion information includes:

[0028] Obtain a number of relational knowledge nodes that have a direct edge relationship with the knowledge node in the personalized memory graph, and use the first semantic information in the knowledge node and the number of relational knowledge nodes as knowledge semantic information;

[0029] Generate dialogue text data based on the semantic information under the knowledge semantic information, and adaptively generate dialogue audio data from the dialogue text data under the audio emotion information.

[0030] In addition, an embodiment of the present invention further provides an adaptive dialogue generation device based on a personalized memory graph, which is applied to an intelligent outbound robot all-in-one machine. The intelligent outbound robot all-in-one machine is communicatively connected to a remote server. The device includes:

[0031] Association module: After the service object user logs in to the remote server through the intelligent outbound robot all-in-one machine, it is used to associate the long-term memory database on the remote server with the intelligent outbound robot all-in-one machine based on the login information of the service object user; at the same time,

[0032] Loading and storage module: It is used for the remote server to load the personalized memory graph stored in the long-term memory database to the local of the intelligent outbound robot all-in-one machine for local storage;

[0033] Information extraction module: When the intelligent outbound robot all-in-one machine enters the chat dialogue mode, it is used to collect the dialogue audio data of the service object user based on the audio acquisition device, and extract the audio emotion information, semantic information and semantic keywords in the dialogue audio data;

[0034] Matching module: It is used to perform matching processing on the local stored personalized memory graph by using the semantic keywords, and match the knowledge nodes of the semantic keywords in the personalized memory graph;

[0035] Audio generation module: It is used to adaptively generate the dialogue audio data corresponding to the dialogue audio data based on the semantic information, the knowledge node and the audio emotion information, and perform dialogue playback processing.

[0036] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory. The processor runs a computer program or code stored in the memory to implement the adaptive dialogue generation method as described in any one of the above.

[0037] In addition, an embodiment of the present invention further provides a computer-readable storage medium for storing a computer program or code. When the computer program or code is executed by a processor, the adaptive dialogue generation method as described in any one of the above is implemented.

[0038] In an embodiment of the present invention, the long-term memory database on the remote server is associated with the intelligent outbound robot all-in-one machine through the login information of the service object user, and the personalized memory map stored in the long-term memory database is loaded to the local of the intelligent outbound robot all-in-one machine for local storage; realizing the lightweight storage of the local of the intelligent outbound robot all-in-one machine and reducing the storage pressure of the intelligent outbound robot all-in-one machine; at the same time, by constructing a personalized memory map, and in the subsequent chat mode, extracting the audio emotion information, semantic information and semantic keywords in the dialogue audio data; matching the corresponding knowledge nodes in the personalized memory map through the semantic keywords, and then adaptively generating the dialogue audio data corresponding to the dialogue audio data according to the semantic information, knowledge nodes and audio emotion information, and performing dialogue playback processing; thereby realizing chat conversations according to the key information in the service object user's chat, and generating the final dialogue audio data according to the audio emotion information of the service object user for dialogue playback; effectively avoiding the repetition of the dialogue and chatting to death the topic of the dialogue chat, and improving the realism of the chat conversation. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0040] Figure 1 It is a schematic flowchart of an adaptive dialogue generation method based on a personalized memory map in an embodiment of the present invention;

[0041] Figure 2 It is a schematic flowchart of an adaptive dialogue generation method based on a personalized memory map in another embodiment of the present invention;

[0042] Figure 3 It is a schematic structural composition diagram of an adaptive dialogue generation device based on a personalized memory map in an embodiment of the present invention;

[0043] Figure 4 It is a schematic structural composition diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] Example 1. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an adaptive dialogue generation method based on a personalized memory map in an embodiment of the present invention.

[0046] As Figure 1 shown, an adaptive dialogue generation method based on a personalized memory map is applied to an intelligent outbound robot all-in-one machine, and the intelligent outbound robot all-in-one machine is communicatively connected to a remote server. The method includes:

[0047] S101: After the service object user logs in to the remote server through the intelligent outbound robot all-in-one machine, associate the long-term memory database on the remote server with the intelligent outbound robot all-in-one machine based on the login information of the service object user;

[0048] In the specific implementation process of the present invention, the associating the long-term memory database on the remote server with the intelligent outbound robot all-in-one machine based on the login information of the service object user includes: the remote server obtains the user identifier in the login information of the service object user, and the user identifier is unique and is assigned when the service object user registers on the remote server; the remote server uses the user identifier to index the long-term memory database bound to the user identifier, and the long-term memory database is a database created by the remote server for the service object user using the user identifier. The long-term memory database stores the user information of the service object user and the key information captured when the service object user communicates with the intelligent outbound robot all-in-one machine, and the key information includes keywords and semantic information corresponding to the keywords; the remote server uses the user identifier to associate the long-term memory database with the local account of the service object user on the intelligent outbound robot all-in-one machine, and the local account is an account created locally on the intelligent outbound robot all-in-one machine using the user identifier in the login information of the service object user.

[0049] The personalized memory graph is also stored in the long-term memory database; the personalized memory graph is a knowledge graph constructed by using the user information stored in the long-term memory database and the key information captured when the service object user converses with the intelligent outbound robot all-in-one machine. When the key information stored in the long-term memory database is updated, the personalized memory graph is updated by using the updated key information.

[0050] Specifically, first, when the service object user needs to use the intelligent outbound robot all-in-one machine, the service object user first needs to log in to the remote server through the intelligent outbound robot all-in-one machine by using the corresponding information. After binding the long-term memory database of the service object user on the remote server with the intelligent outbound robot all-in-one machine device, the intelligent outbound robot all-in-one machine device can provide targeted dialogue services for the service object user.

[0051] That is, when associating the long-term memory database with the intelligent outbound robot all-in-one machine, the remote server needs to obtain the user identifier in the login information of the service object user. The user identifier is unique and is assigned when the service object user registers on the remote server. Then the remote server will index the long-term memory database bound to the user identifier through the user identifier. The long-term memory database is a database created by the remote server for the service object user by using the user identifier. The long-term memory database stores the user information of the service object user and the key information captured when the service object user converses with the intelligent outbound robot all-in-one machine. The key information includes keywords and the semantic information corresponding to the keywords. That is, in each chat conversation, the extracted key information is uploaded to the remote server, and the remote server updates and stores these key information in the long-term memory database. The update storage is to match whether the key information already exists in the long-term memory database. If it does not exist, it is stored in the long-term memory database. If it already exists, it is matched whether there is a difference between the existing key information and the latest key information. If there is a difference, it is updated to the latest key information, otherwise it does not need to be updated or stored in the long-term memory database.

[0052] The remote server will associate the long-term memory database with the local account of the service object user on the intelligent outbound robot all-in-one machine by using the user identifier. The local account is an account created locally on the intelligent outbound robot all-in-one machine by using the user identifier in the login information of the service object user. By binding the long-term memory database with the local account created by the service object user on the intelligent outbound robot all-in-one machine, it can make the subsequent dialogue more real by accurately using the personalized memory graph stored in the long-term memory database to adaptively generate the dialogue.

[0053] The personalized memory graph is also stored in the long-term memory database; the personalized memory graph is a knowledge graph constructed by using the user information stored in the long-term memory database and the key information captured when the user of the service object communicates with the intelligent outbound robot all-in-one machine. When the key information stored in the long-term memory database is updated, the personalized memory graph is updated by using the updated key information; thus, the timeliness of the update of the personalized memory graph is ensured, and the subsequent dialogue generation is made more realistic.

[0054] S102: The remote server loads the personalized memory graph stored in the long-term memory database to the local of the intelligent outbound robot all-in-one machine for local storage;

[0055] In the specific implementation process of the present invention, the remote server loads the personalized memory graph stored in the long-term memory database to the local of the intelligent outbound robot all-in-one machine for local storage, including: the remote server loads the personalized memory graph stored in the long-term memory database to the intelligent outbound robot all-in-one machine through a communication connection; the intelligent outbound robot all-in-one machine indexes to the corresponding local storage space through the association between the long-term memory database and the local account, and updates and stores the personalized memory graph to the local storage space, and the local storage space is the storage space opened by the intelligent outbound robot all-in-one machine for the local account in the local storage device.

[0056] Specifically, first, a separate local storage space is allocated for each user of the service object on the local storage device of the intelligent outbound robot all-in-one machine, and this storage space is used for subsequent lightweight storage; that is, first, the remote server loads the personalized memory graph stored in the long-term memory database to the intelligent outbound robot all-in-one machine through a communication connection; after the intelligent outbound robot all-in-one machine receives the personalized memory graph, it indexes to the corresponding local storage space through the association between the long-term memory database and the local account, and finally updates and stores the personalized memory graph to the local storage space, thereby realizing the local lightweight storage of the intelligent outbound robot all-in-one machine.

[0057] S103: When the intelligent outbound robot all-in-one machine enters the chat dialogue mode, based on the audio acquisition device, it acquires the dialogue audio data of the user of the service object, and extracts the audio emotion information, semantic information and semantic keywords in the dialogue audio data;

[0058] In the specific implementation process of the present invention, extracting the audio emotion information, semantic information, and semantic keywords in the dialogue audio data includes: the intelligent outbound robot all-in-one machine performs pitch and speech rate analysis processing on the dialogue audio data to obtain the pitch data and speech rate data corresponding to the audio data; based on an emotion analysis model, uses the pitch data and speech rate data corresponding to the audio data to perform audio emotion analysis processing to obtain the audio emotion information corresponding to the audio data, and the emotion analysis model is obtained by training a deep neural network model using the pitch data and speech rate data marked with audio emotion information corresponding to historical audio data; converts the dialogue audio data into dialogue text data, and inputs the dialogue text data and the audio emotion information into a natural language analysis model for semantic analysis processing to obtain the semantic information of the dialogue audio data under the audio emotion information; uses a keyword extraction algorithm to perform keyword extraction processing on the semantic information to obtain the semantic keywords corresponding to the semantic information.

[0059] Specifically, at this time, when the intelligent outbound robot all-in-one machine enters the chat dialogue mode, the audio collection device (MIC device) set on the intelligent outbound robot all-in-one machine will be started, and the dialogue audio of the service object user during the dialogue will be collected through the audio collection device, and then the dialogue audio data can be obtained; after obtaining the dialogue audio data, further processing is required to extract the audio emotion information, semantic information, and semantic keywords in the dialogue audio data.

[0060] That is, first, an audio analysis model is called on the intelligent outbound robot all-in-one machine, and the pitch and speech rate analysis processing is performed on the dialogue audio data through the audio analysis model to obtain the pitch data and speech rate data corresponding to the audio data; then the emotion analysis model is called, and the audio emotion analysis processing is performed on the pitch data and speech rate data corresponding to the audio data through the emotion analysis model to obtain the audio emotion information corresponding to the audio data; among them, the emotion analysis model is obtained by training a deep neural network model using the pitch data and speech rate data marked with audio emotion information corresponding to historical audio data.

[0061] Then, the audio needs to be converted into text through text extraction software to convert the dialogue audio data into dialogue text data. After obtaining the dialogue text data, the dialogue text data and the audio emotion information need to be input into a natural language analysis model for semantic analysis processing to extract the semantic information of the dialogue audio data under the audio emotion information; then, the keyword extraction algorithm is used to perform keyword extraction processing on the semantic information, and finally the semantic keywords corresponding to the semantic information will be obtained; through the above processing method, the audio emotion information, semantic information, and semantic keywords in the dialogue audio data can be extracted, which is convenient for subsequent processing.

[0062] S104: Perform a matching process in the personalized memory graph stored locally using the semantic keywords, and match the knowledge nodes of the semantic keywords in the personalized memory graph;

[0063] In the specific implementation process of the present invention, the performing a matching process in the personalized memory graph stored locally using the semantic keywords and matching the knowledge nodes of the semantic keywords in the personalized memory graph includes: the intelligent outbound robot all-in-one machine retrieves the personalized memory graph stored locally, and uses the semantic keywords to perform a matching process with the key information in each knowledge node in the personalized memory graph to obtain a matching result; the knowledge nodes of the semantic keywords in the personalized memory graph are obtained through the matching result.

[0064] Specifically, at this time, it is necessary for the intelligent outbound robot all-in-one machine to retrieve the personalized memory graph stored locally, and then use the semantic keywords to perform a matching process with the key information in each knowledge node in the personalized memory graph to obtain a matching result; finally, the knowledge nodes of the semantic keywords in the personalized memory graph are confirmed through the matching result; so that in the subsequent generation of the dialogue, the corresponding key information in the knowledge nodes can be referred to for the generation of the dialogue, so that the generated dialogue is more in line with the dialogue content required by the service object user, and thus the generated dialogue will be more realistic.

[0065] S105: Adaptively generate the dialogue audio data corresponding to the dialogue audio data based on the semantic information, the knowledge nodes, and the audio emotion information, and perform dialogue playback processing.

[0066] In the specific implementation process of the present invention, the adaptively generating the dialogue audio data corresponding to the dialogue audio data based on the semantic information, the knowledge nodes, and the audio emotion information includes: obtaining a plurality of relational knowledge nodes that have a direct edge relationship with the knowledge nodes in the personalized memory graph, and using the first semantic information in the knowledge nodes and the plurality of relational knowledge nodes as knowledge semantic information; generating dialogue text data based on the semantic information under the knowledge semantic information, and adaptively generating dialogue audio data from the dialogue text data under the audio emotion information.

[0067] Specifically, in order to make the generated dialogue more formal, more knowledge related to the service object user is needed as support. Therefore, it is necessary to obtain several relational knowledge nodes that have a direct edge relationship with the knowledge nodes in the personalized memory graph, and then extract the key information in these nodes as the corresponding knowledge, that is, take the first semantic information in the knowledge nodes and several relational knowledge nodes as the knowledge semantic information; then generate the dialogue text data corresponding to the user reply dialogue audio data under the knowledge semantic information through this semantic information (this is the semantic information extracted from the dialogue audio data when the service object user is chatting with the intelligent outbound robot integrated machine), and then adaptively generate the dialogue audio data from the dialogue text data under the audio emotion information, so as to perform dialogue playback processing.

[0068] In the embodiment of the present invention, the long-term memory database on the remote server is associated with the intelligent outbound robot integrated machine through the login information of the service object user, and the personalized memory graph stored in the long-term memory database is loaded to the local of the intelligent outbound robot integrated machine for local storage; realizing the lightweight storage of the intelligent outbound robot integrated machine locally and reducing the storage pressure of the intelligent outbound robot integrated machine; at the same time, by constructing a personalized memory graph, and in the subsequent chat mode, extracting the audio emotion information, semantic information and semantic keywords in the dialogue audio data; matching the corresponding knowledge nodes in the personalized memory graph through the semantic keywords, and then adaptively generating the dialogue audio data corresponding to the dialogue audio data through the semantic information, knowledge nodes and audio emotion information, and performing dialogue playback processing; thus realizing chat conversations according to the key information in the chat of the service object user, and generating the final dialogue audio data for dialogue playback according to the audio emotion information of the service object user; effectively avoiding the repetition of the dialogue and getting the chat topic stuck, and enhancing the realism of the chat conversation.

[0069] Embodiment 2, please refer to Figure 2 , Figure 2 is a schematic flowchart of an adaptive dialogue generation method based on a personalized memory graph in another embodiment of the present invention.

[0070] As Figure 2 shown, an adaptive dialogue generation method based on a personalized memory graph is applied to an intelligent outbound robot integrated machine, and the intelligent outbound robot integrated machine is communicatively connected to a remote server. The method includes:

[0071] S201: After the service object user logs in to the remote server through the intelligent outbound robot integrated machine, associate the long-term memory database on the remote server with the intelligent outbound robot integrated machine based on the login information of the service object user; at the same time,

[0072] S202: The remote server loads the personalized memory graph stored in the long-term memory database to the local of the intelligent outbound robot all-in-one machine for local storage;

[0073] S203: When the intelligent outbound robot all-in-one machine enters the chat dialogue mode, based on the audio acquisition device, the dialogue audio data of the service object user is collected, and the intelligent outbound robot all-in-one machine analyzes and processes the pitch and speech rate of the dialogue audio data to obtain the pitch data and speech rate data corresponding to the audio data;

[0074] S204: Based on the emotion analysis model, audio emotion analysis processing is performed using the pitch data and speech rate data corresponding to the audio data to obtain the audio emotion information corresponding to the audio data. The emotion analysis model is obtained by training a deep neural network model using the pitch data and speech rate data marked with audio emotion information corresponding to historical audio data;

[0075] S205: Convert the dialogue audio data into dialogue text data, and input the dialogue text data and the audio emotion information into a natural language analysis model for semantic analysis processing to obtain the semantic information of the dialogue audio data under the audio emotion information;

[0076] S206: Use the keyword extraction algorithm to perform keyword extraction processing on the semantic information to obtain the semantic keywords corresponding to the semantic information;

[0077] S207: Use the semantic keywords to perform matching processing in the personalized memory graph stored locally, and match the knowledge nodes of the semantic keywords in the personalized memory graph;

[0078] S208: Based on the semantic information, the knowledge nodes, and the audio emotion information, adaptively generate the dialogue audio data corresponding to the dialogue audio data, and perform dialogue playback processing.

[0079] For the specific implementation of the second embodiment, reference can be made to the above embodiment, which will not be elaborated here.

[0080] Embodiment 3, please refer to Figure 3 , Figure 3 which is a schematic structural composition diagram of the adaptive dialogue generation device based on the personalized memory graph in the embodiments of the present invention.

[0081] As Figure 3 shown, an adaptive dialogue generation device based on a personalized memory graph is applied to an intelligent outbound robot all-in-one machine. The intelligent outbound robot all-in-one machine is communicatively connected to a remote server. The device includes:

[0082] Association Module 301: After the service object user logs in to the remote server through the intelligent outbound robot all-in-one machine, it associates the long-term memory database on the remote server with the intelligent outbound robot all-in-one machine based on the login information of the service object user;

[0083] In the specific implementation process of the present invention, the associating the long-term memory database on the remote server with the intelligent outbound robot all-in-one machine based on the login information of the service object user includes: the remote server obtains the user identifier in the login information of the service object user, the user identifier is unique and is assigned when the service object user registers on the remote server; the remote server uses the user identifier to index to the long-term memory database bound to the user identifier, the long-term memory database is a database created by the remote server for the service object user using the user identifier, and the long-term memory database stores the user information of the service object user and the key information captured when the service object user communicates with the intelligent outbound robot all-in-one machine, the key information includes keywords and semantic information corresponding to the keywords; the remote server uses the user identifier to associate the long-term memory database with the local account of the service object user on the intelligent outbound robot all-in-one machine, and the local account is an account created locally on the intelligent outbound robot all-in-one machine using the user identifier in the login information of the service object user.

[0084] The long-term memory database also stores the personalized memory graph; the personalized memory graph is a knowledge graph constructed using the user information stored in the long-term memory database and the key information captured when the service object user communicates with the intelligent outbound robot all-in-one machine, and when the key information stored in the long-term memory database is updated, the personalized memory graph is updated using the updated key information.

[0085] Specifically, first, when the service object user needs to use the intelligent outbound robot all-in-one machine, first, the service object user needs to use the corresponding information to log in to the remote server through the intelligent outbound robot all-in-one machine, and then after binding the long-term memory database of the service object user on the remote server with the intelligent outbound robot all-in-one machine device, the intelligent outbound robot all-in-one machine device can provide targeted dialogue services to the service object user.

[0086] That is, when associating the long-term memory database with the intelligent outbound call robot all-in-one machine, the remote server needs to obtain the user identifier in the login information of the service object user, where the user identifier is unique and is assigned when the service object user registers on the remote server. Then, the remote server will index the long-term memory database bound to the user identifier through the user identifier. Among them, the long-term memory database is a database created by the remote server for the service object user using the user identifier. The long-term memory database stores the user information of the service object user and the key information captured when the service object user has a conversation with the intelligent outbound call robot all-in-one machine. The key information includes keywords and the semantic information corresponding to the keywords. That is, in each chat conversation, the extracted key information is uploaded to the remote server, and the remote server updates and stores these key information in the long-term memory database. The update storage is to check whether the key information already exists in the long-term memory database. If it does not exist, it is stored in the long-term memory database. If it already exists, it is checked whether there is a difference between the existing key information and the latest key information. If there is a difference, it is updated to the latest key information, otherwise, there is no need to update or store it in the long-term memory database.

[0087] The remote server will associate the long-term memory database with the local account of the service object user on the intelligent outbound call robot all-in-one machine using the user identifier. The local account is an account created locally on the intelligent outbound call robot all-in-one machine using the user identifier in the login information of the service object user. By binding the long-term memory database to the local account created by the service object user on the intelligent outbound call robot all-in-one machine, it is possible to accurately utilize the personalized memory graph stored in the long-term memory database to adaptively generate conversations in subsequent conversations, making the conversations more realistic.

[0088] The long-term memory database also stores the personalized memory graph. The personalized memory graph is a knowledge graph constructed using the user information stored in the long-term memory database and the key information captured when the service object user has a conversation with the intelligent outbound call robot all-in-one machine. When the key information stored in the long-term memory database is updated, the personalized memory graph is updated using the updated key information. Thus, the timeliness of the update of the personalized memory graph is ensured, and the generation of subsequent conversations is made more realistic.

[0089] Loading and storage module 302: used for the remote server to load the personalized memory graph stored in the long-term memory database to the local of the intelligent outbound call robot all-in-one machine for local storage;

[0090] In the specific implementation process of the present invention, the remote server loads the personalized memory graph stored in the long-term memory database to the local of the intelligent outbound call robot all-in-one machine for local storage, including: the remote server loads the personalized memory graph stored in the long-term memory database to the intelligent outbound call robot all-in-one machine through a communication connection; the intelligent outbound call robot all-in-one machine indexes to the corresponding local storage space through the association between the long-term memory database and the local account, and updates and stores the personalized memory graph to the local storage space, and the local storage space is the storage space opened by the intelligent outbound call robot all-in-one machine for the local account in the local storage device.

[0091] Specifically, first, a separate local storage space is allocated for each service object user on the local storage device of the intelligent outbound call robot all-in-one machine, and this storage space is used for subsequent lightweight storage; that is, first, the remote server loads the personalized memory graph stored in the long-term memory database to the intelligent outbound call robot all-in-one machine through a communication connection; after the intelligent outbound call robot all-in-one machine receives the personalized memory graph, it indexes to the corresponding local storage space through the association between the long-term memory database and the local account, and finally updates and stores the personalized memory graph to the local storage space, thereby realizing the local lightweight storage of the intelligent outbound call robot all-in-one machine.

[0092] Information extraction module 303: It is used to collect the dialogue audio data of the service object user based on the audio collection device when the intelligent outbound call robot all-in-one machine enters the chat dialogue mode, and extract the audio emotion information, semantic information and semantic keywords in the dialogue audio data.

[0093] In the specific implementation process of the present invention, the extraction of the audio emotion information, semantic information and semantic keywords in the dialogue audio data includes: the intelligent outbound call robot all-in-one machine performs tone and speech rate analysis processing on the dialogue audio data to obtain the tone data and speech rate data corresponding to the audio data; based on the emotion analysis model, uses the tone data and speech rate data corresponding to the audio data to perform audio emotion analysis processing to obtain the audio emotion information corresponding to the audio data, and the emotion analysis model is obtained by training the deep neural network model with the tone data and speech rate data marked with audio emotion information corresponding to the historical audio data; converts the dialogue audio data into dialogue text data, and inputs the dialogue text data and the audio emotion information into the natural language analysis model for semantic analysis processing to obtain the semantic information of the dialogue audio data under the audio emotion information; uses the keyword extraction algorithm to perform keyword extraction processing on the semantic information to obtain the semantic keywords corresponding to the semantic information.

[0094] Specifically, at this time, when the intelligent outbound call robot all-in-one machine enters the chat dialogue mode, the audio acquisition device (MIC device) set on the intelligent outbound call robot all-in-one machine will be activated, and the dialogue audio of the service object user during the dialogue will be acquired through the audio acquisition device, and the dialogue audio data can be obtained; after the dialogue audio data is obtained, further processing is required to extract the audio emotion information, semantic information, and semantic keywords in the dialogue audio data.

[0095] That is, first, an audio analysis model is called on the intelligent outbound call robot all-in-one machine, and the tone and speech rate of the dialogue audio data are analyzed and processed through the audio analysis model to obtain the tone data and speech rate data corresponding to the audio data; then, an emotion analysis model is called, and the audio emotion analysis of the tone data and speech rate data corresponding to the audio data is performed through the emotion analysis model to obtain the audio emotion information corresponding to the audio data; among them, the emotion analysis model is obtained by training a deep neural network model using the tone data and speech rate data marked with audio emotion information corresponding to historical audio data.

[0096] Then, it is necessary to convert the audio to text through text extraction software to convert the dialogue audio data into dialogue text data. After the dialogue text data is obtained, the dialogue text data and audio emotion information need to be input into the natural language analysis model for semantic analysis to extract the semantic information of the dialogue audio data under the audio emotion information; then, the keyword extraction algorithm is used to extract keywords from the semantic information, and finally, the semantic keywords corresponding to the semantic information are obtained; through the above processing method, the audio emotion information, semantic information, and semantic keywords in the dialogue audio data can be extracted, which is convenient for subsequent processing.

[0097] Matching module 304: It is used to perform matching processing with the personalized memory graph stored locally using the semantic keywords and match the knowledge nodes of the semantic keywords in the personalized memory graph.

[0098] In the specific implementation process of the present invention, the step of performing matching processing with the personalized memory graph stored locally using the semantic keywords and matching the knowledge nodes of the semantic keywords in the personalized memory graph includes: the intelligent outbound call robot all-in-one machine retrieves the personalized memory graph stored locally and uses the semantic keywords to perform matching processing with the key information in each knowledge node of the personalized memory graph to obtain a matching result; the knowledge nodes of the semantic keywords in the personalized memory graph are obtained through the matching result.

[0099] Specifically, at this time, the intelligent outbound robot all-in-one machine needs to retrieve the personalized memory graph stored locally, and then use the semantic keywords to match the key information in each knowledge node of the personalized memory graph to obtain the matching result; finally, the knowledge node of the semantic keywords in the personalized memory graph is confirmed through the matching result; so that when generating the subsequent conversation, the corresponding key information in the knowledge node can be referred to for generating the conversation, so that the generated conversation is more in line with the conversation content required by the service object user, and thus the generated conversation will be more realistic.

[0100] Audio generation module 305: used to adaptively generate the dialogue audio data corresponding to the dialogue audio data based on the semantic information, the knowledge node, and the audio emotion information, and perform dialogue playback processing.

[0101] In the specific implementation process of the present invention, the step of adaptively generating the dialogue audio data corresponding to the dialogue audio data based on the semantic information, the knowledge node, and the audio emotion information includes: obtaining a plurality of relationship knowledge nodes that have a direct edge relationship with the knowledge node in the personalized memory graph, and using the first semantic information in the knowledge node and the plurality of relationship knowledge nodes as knowledge semantic information; generating dialogue text data based on the semantic information under the knowledge semantic information, and adaptively generating dialogue audio data from the dialogue text data under the audio emotion information.

[0102] Specifically, in order to make the generated conversation more formal, more knowledge related to the service object user is needed as support. Therefore, it is necessary to obtain a plurality of relationship knowledge nodes that have a direct edge relationship with the knowledge node in the personalized memory graph, so as to extract the key information in these nodes as the corresponding knowledge, that is, using the first semantic information in the knowledge node and the plurality of relationship knowledge nodes as knowledge semantic information; then generating the dialogue text data corresponding to the user reply dialogue audio data through this semantic information (this is the semantic information extracted from the dialogue audio data when the service object user is talking to the intelligent outbound robot all-in-one machine), and then adaptively generating dialogue audio data from the dialogue text data under the audio emotion information, so as to perform dialogue playback processing.

[0103] In an embodiment of the present invention, the long-term memory database on the remote server is associated with the intelligent outbound robot all-in-one machine through the login information of the service object user, and the personalized memory map stored in the long-term memory database is loaded to the local of the intelligent outbound robot all-in-one machine for local storage; realizing the lightweight storage of the local of the intelligent outbound robot all-in-one machine and reducing the storage pressure of the intelligent outbound robot all-in-one machine; at the same time, by constructing a personalized memory map and extracting the audio emotion information, semantic information and semantic keywords in the dialogue audio data in the subsequent chat mode; matching the corresponding knowledge nodes in the personalized memory map through the semantic keywords, and adaptively generating the dialogue audio data corresponding to the dialogue audio data according to the semantic information, knowledge nodes and audio emotion information, and performing dialogue playback processing; thereby realizing chatting according to the key information in the chat of the service object user, and generating the final dialogue audio data according to the audio emotion information of the service object user for dialogue playback; effectively avoiding the repetition of the dialogue and chatting to death of the dialogue topic, and enhancing the authenticity of the chat dialogue.

[0104] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored, and when the program is executed by a processor, it implements the adaptive dialogue generation method in any one of the above embodiments. Among them, the computer-readable storage medium includes, but is not limited to, any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards or optical cards. That is, the storage device includes any medium that can store or transmit information in a readable form by a device (such as a computer, a mobile phone), and can be a read-only memory, a magnetic disk or an optical disk, etc.

[0105] An embodiment of the present invention also provides a computer application program that runs on a computer, and the computer application program is used to execute the adaptive dialogue generation method in any one of the above.

[0106] In addition, Figure 4 It is a schematic diagram of the structural composition of the electronic device in the embodiment of the present invention.

[0107] An embodiment of the present invention also provides an electronic device, such as Figure 4As shown in the figure. The electronic device includes devices such as a processor 402, a memory 403, an input unit 404, and a display unit 405. Those skilled in the art can understand that Figure 4 the structural devices of the electronic device shown do not limit all devices, and may include more or fewer components than shown, or combine certain components. The memory 403 can be used to store the application program 401 and each functional module. The processor 402 runs the application program 401 stored in the memory 403, thereby performing various functional applications and data processing of the device. The memory can be an internal memory or an external memory, or include both an internal memory and an external memory. The internal memory can include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, or a random access memory. The external memory can include a hard disk, a floppy disk, a ZIP disk, a USB flash drive, a magnetic tape, etc. The memory disclosed in the present invention includes but is not limited to these types of memories. The memory disclosed in the present invention is only an example and not a limitation.

[0108] The input unit 404 is used to receive the input of signals and the keywords input by the user. The input unit 404 can include a touch panel and other input devices. The touch panel can collect the touch operations of the user on or near it (such as the operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel), and drive the corresponding connection device according to a pre-set program; the other input devices can include but are not limited to one or more of a physical keyboard, function keys (such as play control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc. The display unit 405 can be used to display the information input by the user or the information provided to the user and various menus of the terminal device. The display unit 405 can be in the form of a liquid crystal display, an organic light emitting diode, etc. The processor 402 is the control center of the terminal device, connects various parts of the entire device using various interfaces and lines, and performs various functions and processes data by running or executing the software programs and / or modules stored in the memory 403, and calling the data stored in the memory.

[0109] As an embodiment, the electronic device includes: one or more processors 402, a memory 403, one or more application programs 401, where the one or more application programs 401 are stored in the memory 403 and are configured to be executed by the one or more processors 402, and the one or more application programs 401 are configured to execute the corresponding adaptive dialogue generation method in any one of the above embodiments.

[0110] In the embodiment of the present invention, the long-term memory database on the remote server is associated with the intelligent outbound robot all-in-one through the login information of the service object user, and the personalized memory map stored in the long-term memory database is loaded to the local of the intelligent outbound robot all-in-one for local storage; realizing the lightweight storage of the intelligent outbound robot all-in-one locally and reducing the storage pressure of the intelligent outbound robot all-in-one; at the same time, by constructing a personalized memory map, and in the subsequent chat mode, extracting the audio emotion information, semantic information and semantic keywords in the dialogue audio data; matching the corresponding knowledge nodes in the personalized memory map through the semantic keywords, and adaptively generating the dialogue audio data corresponding to the dialogue audio data through the semantic information, knowledge nodes and audio emotion information, and performing dialogue playback processing; thus realizing chat conversations according to the key information in the service object user's chat, and generating the final dialogue audio data according to the audio emotion information of the service object user for dialogue playback; effectively avoiding the repetition of conversations and chatting to death the topics of the dialogue chat, and improving the authenticity of the chat conversation.

[0111] In addition, the above has introduced in detail a method and related device for adaptive dialogue generation based on a personalized memory map provided by the embodiments of the present invention. In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An adaptive dialogue generation method based on a personalized memory graph, characterized in that, Applied to an intelligent outbound call robot all-in-one machine, the intelligent outbound call robot all-in-one machine is communicatively connected to a remote server, and the method includes: After the service object user logs in to the remote server through the intelligent outbound call robot all-in-one machine, associating the long-term memory database on the remote server with the intelligent outbound call robot all-in-one machine based on the login information of the service object user; meanwhile, The remote server loads the personalized memory graph stored in the long-term memory database to the local of the intelligent outbound call robot all-in-one machine for local storage; When the intelligent outbound call robot all-in-one machine enters the chat dialogue mode, collecting the dialogue audio data of the service object user based on an audio collection device, and extracting the audio emotion information, semantic information, and semantic keywords in the dialogue audio data; Performing a matching process on the local stored personalized memory graph using the semantic keywords, and matching the knowledge nodes of the semantic keywords in the personalized memory graph; Generating corresponding dialogue audio data for the dialogue audio data adaptively based on the semantic information, the knowledge nodes, and the audio emotion information, and performing dialogue playback processing.

2. The adaptive dialogue generation method according to claim 1, wherein The associating the long-term memory database on the remote server with the intelligent outbound call robot all-in-one machine based on the login information of the service object user includes: The remote server obtains the user identifier in the login information of the service object user, the user identifier is unique and is assigned by the service object user when registering on the remote server; The remote server indexes the long-term memory database bound to the user identifier using the user identifier. The long-term memory database is a database created by the remote server for the service object user using the user identifier. The long-term memory database stores the user information of the service object user and the key information captured when the service object user has a dialogue with the intelligent outbound call robot all-in-one machine. The key information includes keywords and the semantic information corresponding to the keywords; The remote server associates the long-term memory database with the local account of the service object user on the intelligent outbound call robot all-in-one machine using the user identifier. The local account is an account created locally on the intelligent outbound call robot all-in-one machine using the user identifier in the login information of the service object user.

3. The adaptive dialogue generation method according to claim 2, wherein The personalized memory graph is also stored in the long-term memory database; The personalized memory graph is a knowledge graph constructed using the user information stored in the long-term memory database and the key information captured when the service object user has a dialogue with the intelligent outbound call robot all-in-one machine. When the key information stored in the long-term memory database is updated, the personalized memory graph is updated using the updated key information.

4. The adaptive dialogue generation method according to claim 1, wherein The remote server loading the personalized memory graph stored in the long-term memory database to the local of the intelligent outbound call robot all-in-one machine for local storage includes: The remote server loads the personalized memory map stored in the long-term memory database to the intelligent outbound robot all-in-one machine through a communication connection; The intelligent outbound robot all-in-one machine indexes to the corresponding local storage space through the association between the long-term memory database and the local account, and updates and stores the personalized memory map to the local storage space, where the local storage space is the storage space opened by the intelligent outbound robot all-in-one machine for the local account in the local storage device.

5. The adaptive dialogue generation method according to claim 1, wherein The extraction of audio emotion information, semantic information, and semantic keywords from the dialogue audio data includes: The intelligent outbound robot all-in-one machine performs pitch and speech rate analysis and processing on the dialogue audio data to obtain the pitch data and speech rate data corresponding to the audio data; Based on an emotion analysis model, audio emotion analysis and processing are performed using the pitch data and speech rate data corresponding to the audio data to obtain the audio emotion information corresponding to the audio data. The emotion analysis model is trained by a deep neural network model using the pitch data and speech rate data marked with audio emotion information corresponding to historical audio data; The dialogue audio data is converted into dialogue text data, and the dialogue text data and the audio emotion information are input into a natural language analysis model for semantic analysis and processing to obtain the semantic information of the dialogue audio data under the audio emotion information; The keyword extraction algorithm is used to perform keyword extraction processing on the semantic information to obtain the semantic keywords corresponding to the semantic information.

6. The adaptive dialogue generation method according to claim 1, wherein The use of the semantic keywords to perform matching processing in the personalized memory map stored locally and matching the knowledge nodes of the semantic keywords in the personalized memory map includes: The intelligent outbound robot all-in-one machine retrieves the personalized memory map stored locally and uses the semantic keywords to perform matching processing with the key information in each knowledge node in the personalized memory map to obtain a matching result; The knowledge node of the semantic keyword in the personalized memory map is obtained through the matching result.

7. The adaptive dialogue generation method according to claim 1, wherein The generation of the dialogue audio data corresponding to the dialogue audio data based on the semantic information, the knowledge node, and the audio emotion information adaptively includes: Obtaining a number of relational knowledge nodes that have a direct edge relationship with the knowledge node in the personalized memory map, and using the first semantic information in the knowledge node and the number of relational knowledge nodes as knowledge semantic information; Generating dialogue text data based on the semantic information under the knowledge semantic information, and adaptively generating the dialogue text data into dialogue audio data under the audio emotion information.

8. An adaptive dialogue generation device based on a personalized memory graph, characterized in that, Applied to an intelligent outbound robot all-in-one machine, the intelligent outbound robot all-in-one machine is communicatively connected to a remote server, and the device includes: An association module: used for, after a service object user logs in to the remote server through the intelligent outbound robot all-in-one machine, associating the long-term memory database on the remote server with the intelligent outbound robot all-in-one machine based on the login information of the service object user; meanwhile, Loading and storing module: used for the remote server to load the personalized memory graph stored in the long-term memory database to the local of the intelligent outbound robot all-in-one machine for local storage; Information extraction module: used for when the intelligent outbound robot all-in-one machine enters the chat dialogue mode, collecting the dialogue audio data of the service object user based on the audio acquisition device, and extracting the audio emotion information, semantic information and semantic keywords in the dialogue audio data; Matching module: used for performing matching processing on the local stored personalized memory graph by using the semantic keywords, and matching the knowledge nodes of the semantic keywords in the personalized memory graph; Audio generation module: used for adaptively generating the dialogue audio data corresponding to the dialogue audio data based on the semantic information, the knowledge nodes and the audio emotion information, and performing dialogue playback processing.

9. An electronic device, comprising a processor and a memory, characterized in that, The processor runs the computer program or code stored in the memory to implement the adaptive dialogue generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium for storing a computer program or code, characterized in that, When the computer program or code is executed by the processor, the adaptive dialogue generation method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Method, device and equipment for processing agent data based on knowledge graph

    CN120975116A