Method and apparatus for dialog interaction, and device and medium

By generating dialogue feature information to improve the dialogue memory ability of digital assistants, the problem of poor user experience in existing technologies is solved, and a more human-like dialogue interaction is achieved.

WO2026090952A1PCT designated stage Publication Date: 2026-05-07BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-10-30
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing digital assistants cannot effectively remember conversations when users are having real-time voice calls with them, resulting in unhuman-like responses and a poor user experience.

Method used

By processing the historical dialogue content between users and digital assistants, dialogue feature information is generated. Based on this feature information, response messages are generated, reducing the amount of data and improving the processing speed of the model, thus enhancing the human-like nature of the dialogue.

Benefits of technology

It improves the digital assistant's conversation memory capabilities, enabling it to respond more in line with user expectations and enhance the user's conversation experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128696_07052026_PF_FP_ABST
    Figure CN2024128696_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method and apparatus for dialog interaction, and a device and a medium. One method comprises: in response to receiving a dialog message sent by a user to at least one digital assistant, on the basis of dialog feature information related to the user and the at least one digital assistant, determining prompt information, wherein the dialog feature information is key information generated by means of processing historical dialog content between the user and the at least one digital assistant; and presenting to the user a response message for the dialog message, wherein the response message is generated on the basis of the dialog message and the prompt information. In this way, a response message can be generated on the basis of dialog memory between a user and a digital assistant, so that the digital assistant is made more human-like, thereby improving the dialog experience of the user with the digital assistant.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus, devices and media for dialogue interaction Technical Field

[0001] Exemplary implementations of this disclosure generally relate to computer technology, and in particular to methods, apparatus, devices, and computer-readable storage media for dialogic interaction. Background Technology

[0002] With the development of information technology, digital assistants such as voice assistants, robots (BOTs), and intelligent agents have seen significant development and application. Various types of digital assistants have brought great convenience to users in many ways. Users can use digital assistants for various leisure and entertainment activities; for example, users can have real-time voice calls and chat with their digital assistants.

[0003] Summary of the Invention

[0004] In a first aspect of this disclosure, a method for dialogue interaction is provided. In this method, in response to receiving a dialogue message from a user to at least one digital assistant, prompt information is determined based on dialogue feature information related to the user and the at least one digital assistant, the dialogue feature information being key information generated by processing historical dialogue content between the user and the at least one digital assistant. A response message to the dialogue message is then presented to the user, the response message being generated based on the dialogue message and the prompt information.

[0005] In a second aspect of this disclosure, an apparatus for dialogue interaction is provided. The apparatus includes: a prompt information determination module configured to determine prompt information based on dialogue feature information associated with the user and the at least one digital assistant in response to receiving a dialogue message from a user to at least one digital assistant; the dialogue feature information being key information generated by processing historical dialogue content between the user and the at least one digital assistant; and a response message generation module configured to present a response message to the user in response to the dialogue message; the response message being generated based on the dialogue message and the prompt information.

[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.

[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.

[0008] In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent in the following detailed description, taken in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0011] Figure 1 shows a schematic diagram of an example environment in which the implementation of this disclosure can be implemented;

[0012] Figure 2 shows a flowchart of an example process for dialogue interaction according to some embodiments of the present disclosure;

[0013] Figure 3 shows a flowchart of an example process for dialogue memory according to some embodiments of the present disclosure;

[0014] Figure 4 illustrates an example of the signaling flow of a dialogue interaction according to some embodiments of the present disclosure;

[0015] Figure 5 shows a flowchart of a method for dialogue interaction according to some embodiments of the present disclosure;

[0016] Figure 6 shows a block diagram of a device for dialogue interaction according to some embodiments of the present disclosure; and

[0017] Figure 7 shows a block diagram of an apparatus capable of implementing several embodiments of the present disclosure. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.

[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0022] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information, thereby enabling the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers or storage media that perform the operation of the technical solution disclosed herein, based on the prompt message.

[0023] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0025] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.

[0026] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.

[0027] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an application service component 112 and a digital assistant 114 are installed on a client device 110. A user 130 can interact with the application service component 112 and the digital assistant 114 via the client device 110 and / or an attached device to the client device 110.

[0028] In some embodiments, application service component 112 and digital assistant 114 can be downloaded and installed on client device 110. In some embodiments, application service component 112 and digital assistant 114 can also be accessed in other ways, such as through a web page. In some embodiments, in environment 100 of FIG1, in response to application service component 112 being started, client device 110 can present interface 140 of application service component 112 and digital assistant 114. Interface 140 can be, for example, an interactive interface of application service component 112 and digital assistant 114.

[0029] Application business components 112 include, but are not limited to, one or more of the following: chat application components (also known as instant messaging application components), document application components, audio and video conferencing application components, email application components, task application components, calendar application components, goal and key results (OKR) application components, etc. It is understood that although a single application business component is shown in Figure 1, multiple application business components may actually be installed on the client device 110. In some embodiments, application business component 112 may include a multi-functional collaboration platform, such as an office collaboration platform (also known as an office suite) that can provide integration of various types of business components to facilitate people's office work, communication, and other activities. In a multi-functional collaboration platform, people can launch different business components as needed to complete corresponding information processing, sharing, communication, etc.

[0030] In some embodiments, the digital assistant 114 may be provided by a separate application business component, or it may be integrated into an application business component 120 capable of providing content entities. The application business component providing the client interface for the digital assistant may correspond to a single-function application business component or a multi-functional collaboration platform, such as an office suite or other collaboration platform capable of integrating multiple components. It is understood that, similar to application business components, although a single digital assistant is shown in Figure 1, there may actually be multiple digital assistants.

[0031] In some embodiments, the digital assistant 114 supports the use of plugins. Each plugin can provide one or more functions of the application. Such plugins include, but are not limited to, one or more of the following: search plugin, contact plugin, messaging plugin, document plugin, form plugin, email plugin, calendar plugin, schedule plugin, task plugin, etc.

[0032] Digital assistant 114 is a user's intelligent assistant, possessing intelligent dialogue and information processing capabilities. In embodiments of this disclosure, digital assistant 114 is used to interact with user 130 to assist user 130 in using terminal devices or applications. In some embodiments, an interaction window for digital assistant 114 may be presented in interface 140. In the interaction window, user 130 can converse with digital assistant 114 by inputting natural language, images, audio files, video files, web page files, etc., to instruct digital assistant to assist in completing various tasks, including operations on content entities.

[0033] In some embodiments, multiple interaction modes can be provided for the user 130 and the digital assistant 114, and the user can flexibly switch between these modes. When a certain interaction mode is triggered, a corresponding interaction area is presented to facilitate interaction between the user 130 and the digital assistant 114. The interaction methods between the user 130 and the digital assistant 114 differ under different interaction modes, thus flexibly adapting to the interaction needs of different application scenarios.

[0034] In some embodiments, information processing services specific to user 130 can be provided based on historical interaction information between user 130 and digital assistant 114 and / or data ranges specific to user 130. In some embodiments, historical interaction information of user 130 interacting with digital assistant 114 in multiple interaction modes can all be associated with user 130 and stored. Thus, in one of the multiple interaction modes (any or a specified interaction mode), digital assistant 114 can provide services to user 130 based on the historical interaction information associated with user 130 stored.

[0035] Digital assistant 114 can be invoked or activated by an appropriate means (e.g., shortcut keys, buttons, or voice) to present an interaction window with user 130. Selecting digital assistant 114 opens an interaction window with it. The interaction window may include interface elements for information interaction, such as input boxes, message lists, message bubbles, etc. In other embodiments, digital assistant 114 can be invoked through entry controls or menus provided in interface 140, or by inputting preset commands.

[0036] The interaction window between the digital assistant 114 and the user 130 may include a session window, such as a session window in an instant messaging application or an instant messaging module of a specific application. In the session window, the interaction between the digital assistant 114 and the user 130 may be presented in the form of session messages. Alternatively or additionally, the interaction window between the digital assistant 114 and the user 130 may also include other types of windows, such as a floating window, in which the user 130 can trigger the digital assistant 114 to perform corresponding operations by entering commands, selecting shortcuts, etc.

[0037] In some embodiments, the digital assistant 114 may support a conversation window interaction mode, also known as conversation mode. In this interaction mode, a conversation window is presented between the user 130 and the digital assistant 114, where the user 130 and the digital assistant 114 interact through conversation messages. In conversation mode, the digital assistant 114 can perform tasks based on the conversation messages in the conversation window. In the interaction window, the user 130 inputs interaction messages, and the digital assistant 114 responds to the user's input by providing a reply message. A conversation window with the digital assistant 114 can be opened by selecting the digital assistant 114. The conversation window may include interface elements for information interaction, such as input boxes, message lists, message bubbles, etc.

[0038] In some embodiments, the digital assistant 114 can operate in a voice call interaction mode, also known as a call mode. In this interaction mode, a call interface is presented between the user 130 and the digital assistant 114, where the user 130 interacts with the digital assistant 114 via real-time voice calls using an attached device (e.g., a speaker and microphone). In call mode, the digital assistant 114 can provide response messages based on the user's voice input.

[0039] In some embodiments, client device 110 communicates with server device 120 to provide services to digital assistant 114 and business component 125. Client device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, client device 110 can also support any type of user-facing interface (such as "wearable" circuitry). Server device 120 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0040] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0041] As mentioned earlier, users can converse with digital assistants. In some scenarios, these conversations can be quite information-rich. For example, users can engage in real-time voice calls to chat with their digital assistants. In this scenario, the conversations are often quite in-depth, typically involving a lot of content about daily life. Therefore, in real-time call scenarios, users expect their digital assistants to feel more like a real "friend" than a chat tool. However, current digital assistants typically do not remember the content of users' voice calls, making it impossible for them to provide more human-like responses based on previous conversations, resulting in a poor user experience. Similar problems may exist in other conversational scenarios.

[0042] In view of this, according to embodiments of the present disclosure, an improved scheme for dialogue interaction is provided. Specifically, in response to receiving a dialogue message from a user to at least one digital assistant, prompt information is determined based on dialogue feature information related to the user and at least one digital assistant. The dialogue feature information is key information generated by processing historical dialogue content between the user and at least one digital assistant. A response message for the dialogue message is presented to the user, the response message being generated based on the dialogue message and the prompt information.

[0043] According to the scheme disclosed herein, the content of the dialogue between the user and the digital assistant is memorized, and response messages are generated based on this memorization. This allows the digital assistant to generate response messages that are more closely aligned with the user's expectations based on the memorized dialogue content, providing the user with a more lifelike dialogue experience. Furthermore, according to the scheme disclosed herein, what is memorized is key information about the dialogue between the user and the digital assistant; in other words, the dialogue content between the user and the digital assistant is abstracted into feature information rather than preserved verbatim. This reduces the amount of data input to the model that generates the response messages, lowers the model's processing difficulty, and speeds up the model's processing. Therefore, the scheme disclosed herein can also improve the problems of insufficient context window and the model's inability to remember user dialogue in models used to generate response messages. In this way, the model is more adept at generating response messages based on memorized information, improving the user experience.

[0044] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0045] Figure 2 illustrates a flowchart of an example process 200 for dialogue interaction according to some embodiments of the present disclosure. For ease of discussion, process 200 will be described with reference to the environment of Figure 1. In environment 100, process 200 may be implemented at client device 110 and / or server device 120. For ease of description, it will be illustrated by assuming that process 200 is implemented at server device 120.

[0046] It should be noted that process 200 can be implemented at the client device 110. Some operations described with reference to client device 110 may require the assistance of server device 120. It should be noted that the operations performed by client device 110 may specifically be performed by relevant applications and / or digital assistants installed on client device 110.

[0047] In some embodiments, server device 120 can provide conversational interaction functions such as voice calls, allowing user 130 to chat with digital assistant 114. For example, user 130 can chat with digital assistant 114 via real-time voice calls. Conversation message 210 can be, for example, a real-time voice call message input by user 130 through client device 110. Of course, in other embodiments, conversation message 210 can also be other types of messages input by user 130 through client device 110, such as non-real-time voice messages or text messages.

[0048] In some embodiments, in response to receiving a dialogue message 210 from user 130 to digital assistant 114, server device 120 retrieves dialogue feature information 230 related to user 130 and digital assistant 114 from dialogue feature information database 220. Dialogue feature information 230 is generated based on historical dialogue content between user 130 and digital assistant 114. Specifically, dialogue feature information 230 is key information obtained by processing historical dialogue content. Dialogue feature information 230 is used to characterize key dialogue content between user 130 and digital assistant 114. In this way, the dialogue content between user 130 and digital assistant 114 is abstracted or summarized into dialogue feature information and stored, without storing the dialogue content as is. Such dialogue feature information can be considered a refined, essential memory. In this way, the amount of information stored is minimized while preserving the dialogue context.

[0049] In some embodiments, the dialogue feature information 230 includes entity feature information and event feature information. Entity feature information describes aspects such as the identity, social relationships, and preferences of user 130 and / or digital assistant 114. For example, entity feature information may describe the names of user 130 and digital assistant 114, their relationship, and user 130's hobbies. Event feature information describes aspects of the events that user 130 and / or digital assistant 114 intend to perform. For example, information about user 130 planning a trip at a certain time.

[0050] In some embodiments, the server device 120 may determine the prompt information 240 based on the acquired dialogue feature information 230. This prompt information 240 is used by the digital assistant 114 for reference when generating the response message 250. In some embodiments, the prompt information 240 includes system prompt information and dialogue message 210. The system prompt information is generated based on the dialogue feature information 230. The system prompt information describes entity feature information and / or event feature information related to the user 130 and the digital assistant 114. For example, as shown in Figure 2, "User A is in elementary school and is good friends with Doraemon," and "User A likes to travel, likes the beach, Doraemon also likes the beach, and Doraemon likes to go to the moon," belong to entity feature information. "[Date: xx / xx / xx / 12:20:00] User plans to go to the moon with Doraemon during the Dragon Boat Festival, and Doraemon suggests using a bamboo dragonfly to fly to the moon," belongs to event feature information.

[0051] In some embodiments, the server device 120 generates a response message 250 for the dialogue message 210 based on the prompt information 240. Since the response message 250 is generated based on the prompt information 240, and the prompt information 240 is related to the historical dialogue between the user 130 and the digital assistant 114, the response message 250 is generated based on the historical dialogue between the user 130 and the digital assistant 114. In other words, in this embodiment of the disclosure, the response message 250 for the current dialogue message (i.e., dialogue message 210) is generated by remembering the dialogue between the user 130 and the digital assistant 114. This enables the digital assistant 114 to have dialogue memory capabilities and to complete real-time dialogue with the user based on the remembered dialogue content. This makes the digital assistant 114 more human-like, allowing the user to obtain a near-real-person dialogue experience.

[0052] In some embodiments, the dialogue feature information 230 includes multiple current dialogue feature information items. Each of these multiple current dialogue feature information items is configured with a corresponding weight. The weight represents the importance and / or chronological order of the corresponding dialogue feature information item. When determining the prompt information 240 based on the dialogue feature information 230, the server device 120 determines the order of the multiple current feature information items in the prompt information 240 based on their respective weights. For example, the server device 120 places dialogue feature information items with higher weights at the beginning of the prompt information 240 and places dialogue feature information items with lower weights at the end of the prompt information 240.

[0053] In some embodiments, the weight of a dialogue feature item can be determined based on the dialogue occurrence time corresponding to the dialogue feature item and the importance of the dialogue feature item. For example, the server device 120 can assign greater weights to information related to user identity, preferences, and relationships. As another example, the server device 120 can assign weights based on the proximity of the dialogue occurrence time corresponding to the dialogue feature item to the current time; the closer the dialogue occurrence time is to the current time, the greater the weight of the dialogue feature item. It should be understood that the server device 120 can use various suitable weight allocation methods to configure weights for dialogue feature items as needed, and this disclosure does not limit this.

[0054] Taking Figure 2 as an example, information regarding the identity and relationship between user A and the robot cat is assigned a higher weight and therefore appears at the beginning of prompt information 240. Information regarding user A and the robot cat's preferences is assigned a moderate weight and therefore appears in the middle of prompt information 240. Information regarding events involving user A and the robot cat is assigned a lower weight and therefore appears at the end of prompt information 240. It should be understood that this weighting method is merely an example and does not constitute a limitation of this disclosure.

[0055] In some embodiments, the server device 120 may provide a prompt message 240, including dialogue message 210, to a machine learning model (e.g., a language model) and determine a response message 250 based on the output of the machine learning model. In some embodiments, the machine learning module may, when processing the prompt message 240, prioritize referring to the information that appears earlier in the prompt message 240. After receiving the response message 250, the client device 110 may present the response message 250 to the user 130, for example, through a speaker.

[0056] It should be understood that although the embodiments of this disclosure are illustrated using a user's conversation with one digital assistant as an example, the solution of this disclosure can be applied to scenarios where a user converses with multiple digital assistants. For example, a scenario where a user engages in group chat with multiple digital assistants. When applied to a scenario where a user converses with multiple digital assistants, the dialogue feature information can be memorized separately for each user's conversation with each digital assistant, or the dialogue feature information can be shared across the conversations between the user and multiple digital assistants. That is, the dialogue feature information can be memorized separately for each digital assistant, or multiple digital assistants can share the same dialogue feature information.

[0057] It should also be understood that the dialogue feature information database 230 can be located at the server device 120 or at the client device 110. In other words, dialogue memory can be completed at the server device 120 or at the client device 110. Furthermore, the entity feature information database 231 and the event information database 232 can be located in different areas of the same memory or on different memories.

[0058] Figure 3 shows a flowchart of an example process 300 for dialogue memory according to some embodiments of the present disclosure. For ease of description, the process 300 is illustrated as being implemented at server device 120.

[0059] It should be noted that process 300 can also be implemented at the client device 110. Some operations described with reference to client device 110 may require the assistance of server device 120. It should be noted that the operations performed by client device 110 may specifically be performed by relevant applications and / or digital assistants installed on client device 110.

[0060] In some embodiments, user 130 can engage in multiple rounds of dialogue with digital assistant 114, and when the number of dialogue rounds reaches a predetermined number, dialogue memory can be triggered to remember the dialogue between user 130 and digital assistant 114. For example, dialogue memory is triggered when the number of dialogue rounds between user 130 and digital assistant 114 reaches 20 or 30. It should be understood that the predetermined number can be set to any number as needed; the 20 or 30 submitted herein is merely an example. In this disclosure, the process of user 130 providing a dialogue message and digital assistant 114 replying with a response message constitutes one round of dialogue.

[0061] In some embodiments, when the response message corresponds to the end of a predetermined number of rounds of dialogue between user 130 and digital assistant 114, server device 120 generates memory information corresponding to the predetermined number of rounds of dialogue by extracting key information from the dialogue content of the predetermined number of rounds of dialogue. For example, if the predetermined number is 20, when the response message is the response message of the 20th round of dialogue between the user and digital assistant, the server device generates memory information based on the dialogue content of the 20 rounds of dialogue between user 130 and digital assistant 114. This memory information retains only the content that affects the dialogue context and removes meaningless content, and can therefore be considered as essential memory. As shown in Figure 3, the dialogue content 310 includes key information such as personal information, preferences, and events, as well as non-key information such as greetings and questions. Memory information 330 is the essential memory obtained by extracting key information from the dialogue content 310, rather than all information. It can be seen that non-key information such as greetings and questions is not retained in memory information 330. This reduces the amount of data in the memory information, facilitating subsequent model processing.

[0062] In some embodiments, when the response message corresponds to the end of a predetermined number of dialogue rounds between user 130 and digital assistant 114, the dialogue content 310 of the predetermined number of dialogue rounds is provided to the first sub-model 320. The first sub-model 320 generates memory information 330 corresponding to the dialogue content 310 of the predetermined number of dialogue rounds by extracting key information from the dialogue content 310 of the predetermined number of dialogue rounds.

[0063] In some embodiments, the memory retrieval module 320 may include or utilize a machine learning model, such as one or more models including a speech processing model, a language model, etc. The memory retrieval module 320 can extract one or more memory information items related to the user 130 and the digital assistant 114 from the dialogue content 310 of the predetermined number of rounds of dialogue. The memory information items are, for example, refined information about the main content included in the dialogue 310.

[0064] In some embodiments, the memory retrieval module 320 can determine the corresponding category of one or more memory information items through semantic analysis of the dialogue content 310 or the memory information items. The corresponding categories of memory information items include events and non-events. Non-event categories include personal information (e.g., identity information, such as ID), preferences, social relationships (including the user's social relationship with the digital assistant), etc. As shown in Figure 3, the memory information 330 includes four memory information items, each labeled with corresponding category information.

[0065] In some embodiments, the memory retrieval module can also determine the importance of one or more memory information items based on the semantics of the dialogue content 310 or the memory information items themselves. For example, the importance of one or more memory information items can be determined based on information such as the user's voice, tone, and volume corresponding to different dialogue messages in the dialogue content 310. For example, when the user's voice, tone, and volume indicate that the dialogue message is a strong expression of the user, the corresponding memory information item can be marked with a higher level of importance. The level of importance reflects the importance of the information to the characteristics representing the user 130. For example, if the user strongly expresses a strong liking for basketball in the dialogue content, then the memory information item related to the user's liking for basketball can have a higher level of importance because this information reflects the user's strong personal preference.

[0066] In some embodiments, the memory retrieval module 320 can also determine the corresponding occurrence time of one or more memory information items based on the time information corresponding to the dialogue content 310. For example, if the time of occurrence of a memory information item in the dialogue content 310 is 12:10, then the timestamp of the memory information item is 12:10.

[0067] In some embodiments, the memory retrieval module 320 may also configure weights for one or more memory information items based on at least one of the importance of one or more memory information items and the corresponding occurrence time, wherein the weights indicate the prompting priority of the one or more memory information items when determining the prompting information.

[0068] After the memory retrieval module 320 generates memory information 330, the feature update module 340 generates one or more feature information items to characterize the predetermined number of rounds of dialogue based on the memory information 330; and updates the dialogue feature information using one or more feature information items.

[0069] In some embodiments, the dialogue feature information includes entity feature information and event feature information. The feature update module can generate one or more entity feature information items based on memory information items categorized as non-events from one or more memory information items. The feature update module 340 can also generate one or more event information items based on memory information items categorized as events from one or more memory information items. Taking Figure 3 as an example, the feature update module 340 can generate three entity feature information items based on the three memory information items (person information, preferences) in the memory information 330. The feature update module 340 can also generate event feature information items based on the event memory information items in the memory information 330.

[0070] In some embodiments, the feature update module 340 can determine the corresponding weight of one or more feature information items based on at least one of the occurrence time of the dialogue corresponding to one or more feature information items and the corresponding importance of one or more feature information items. It should be understood that the corresponding weight of the one or more feature information items can be determined based on the corresponding weight of the memory information items corresponding to the one or more feature information items. Alternatively, it can be determined directly based on at least one of the occurrence time of the dialogue corresponding to one or more feature information items and the corresponding importance of one or more feature information items. That is, in some embodiments, the weight of one or more memory information items is determined by the memory retrieval module 320, and the feature update module 340 can directly call the weight of the one or more memory information items to determine the weight of the generated one or more feature information items. In some embodiments, without determining the weight of one or more memory information items through the memory retrieval module 320, the feature update module 340 directly determines the occurrence time of the dialogue corresponding to one or more feature information items and the corresponding importance of one or more feature information items based on the dialogue content 310, and then determines the corresponding weight of one or more feature information items based on at least one of the occurrence time of the dialogue corresponding to one or more feature information items and the corresponding importance of one or more feature information items.

[0071] When one or more feature information items are generated, the feature update module 340 can use these feature information items to update the dialogue feature information. For example, the second sub-model 340 can use the generated entity feature information items to update the entity feature information in the entity feature information database 231. The feature update module 340 can use the generated event feature information items to update the event feature information in the event feature information database 232. In some embodiments, when the feature update module 340 uses the generated event feature information items to update the event feature information in the event feature information database 232, it adds a timestamp to the event feature information item to identify the event that occurred in the dialogue. This facilitates prioritizing the reference to the most recent event information when generating subsequent response messages.

[0072] In some embodiments, the feature update module 340 may add one or more feature items to the dialogue feature information in response to the availability of the information base (i.e., entity feature information base 231 or event feature information base 232) being greater than or equal to the storage capacity required for one or more feature information items, so as to update the dialogue feature information.

[0073] As an example, the entity feature information database 231 stores the information that user A is in elementary school at time t0. The feature update module 340 generates three entity feature information items at time t1 (a time after t0): user A's name is XX; user A and Doraemon are good friends; user A likes to travel and likes to go to the moon; Doraemon likes to go to the moon. Since the available capacity of the entity feature information database 231 is greater than the storage capacity required for these three entity feature information items, the feature update module 340 can directly add these three entity feature information items to the entity feature information database 231, obtaining the entity feature information at time t1: user A's name is XX, is in elementary school, and is good friends with Doraemon; user A likes to travel and likes to go to the moon; Doraemon also likes to go to the moon. After the update at time t1, the entity feature information database 231 includes two entity feature information items. It should be understood that this update process is only an example and does not constitute a limitation of this disclosure.

[0074] In some embodiments, the feature update module 340 may replace at least a portion of the dialogue feature information with at least a portion of one or more feature information items in response to the available capacity of the information base (i.e., entity feature information base 231 or event feature information base 232) being less than the storage capacity required by one or more feature information items.

[0075] In some embodiments, one or more feature information items (entity feature information items or event feature information items) generated by the feature update module 340 are each configured with a corresponding first weight. The dialogue feature information includes multiple current feature information items (entity feature information items or event feature information items), and each of the multiple current feature information items is configured with a corresponding second weight. The feature update module 340 can select one or more target feature information items from the multiple current feature information items based on the corresponding first weight of one or more feature information items and the corresponding second weight of the multiple current feature information items; and replace one or more target feature information items with one or more feature information items. As an example, at time t1, the second sub-model 340 generates three feature information items, configured with weights A1, A2, and A3 respectively. The dialogue feature information at time t0 (the time before t1) includes five current feature information items, configured with weights B1, B2, B3, B4, and B5 respectively. An information database (e.g., entity feature information database 231 or event feature information database 232) is used. A1, A2, and A3 are all greater than B3, B4, and B5, but the information database can only store five feature information items. Therefore, the information database can be updated by replacing the current feature information items corresponding to B3, B4, and B5 with three feature information items generated by the feature update module 340 at time t1. It should be understood that this replacement process is only an example and does not constitute a limitation of this disclosure.

[0076] Figure 4 illustrates an example of a signaling flow 400 for a dialogue interaction according to some embodiments of the present disclosure. The signaling flow 400 relates to a client device 110, a server device 120, a machine learning model 401, and feature information 402.

[0077] Client device 110 can receive (411) dialogue messages entered by the user in a dialogue interface (e.g., a call interface) and send (412) the dialogue messages to server device 120. Server device 120 can send (413) a request to database 402 to obtain dialogue feature information. Database 402 can return (414) dialogue feature information to server device 120 based on the request.

[0078] Server device 120 may determine (415) prompt information based on dialogue messages and dialogue feature information, and provide (416) the prompt information to machine learning model 401. Machine learning model 401 generates a response message based on the prompt information and sends the response message to client device 110 (417).

[0079] The client device 110 presents (418) a response message to the user, for example, through a speaker.

[0080] In response to the end of a predetermined number of rounds of dialogue, server device 120 generates (419) memory information. Server device 120 generates (420) one or more feature information items based on the memory information, and updates (421) dialogue feature information using one or more feature information items.

[0081] Figure 5 illustrates a schematic diagram of a process 500 for dialogue interaction according to some embodiments of the present disclosure. Process 500 can be implemented at the server device 120 and / or client device 110 of Figure 1.

[0082] In box 510, in response to receiving a dialogue message from a user to at least one digital assistant, the server device 120 determines a prompt message based on dialogue feature information related to the user and at least one digital assistant. The dialogue feature information is key information generated by processing the historical dialogue content between the user and at least one digital assistant.

[0083] In box 520, server device 120 presents a response message to the user in response to the dialogue message, which is generated based on the dialogue message and prompt information.

[0084] In some embodiments, the user engages in multiple rounds of dialogue with at least one digital assistant, and process 500 further includes: in response to a response message corresponding to the end of a predetermined number of rounds of dialogue between the user and at least one digital assistant, generating memory information corresponding to the predetermined number of rounds of dialogue by extracting key information from the dialogue content of the predetermined number of rounds of dialogue; generating one or more feature information items to characterize the predetermined number of rounds of dialogue based on the memory information; and updating the dialogue feature information using the one or more feature information items.

[0085] In some embodiments, the dialogue feature information includes entity feature information and event feature information, and generating memory information includes: extracting one or more memory information items related to the user or at least one digital assistant from the dialogue content; and determining the corresponding category of one or more memory information items through semantic analysis of the one or more memory information items.

[0086] In some embodiments, generating one or more feature information items includes: generating one or more entity feature information items based on memory information items of the category of non-event among one or more memory information items; and generating one or more event information items based on memory information items of the category of event among one or more memory information items.

[0087] In some embodiments, generating one or more feature information items includes: determining the corresponding weight of one or more feature information items based on at least one of the following: the corresponding occurrence time of the dialogue corresponding to one or more feature information items, or the corresponding importance of one or more feature information items.

[0088] In some embodiments, dialogue feature information items are stored in a database, and updating the dialogue feature information includes: adding one or more feature information items to the dialogue feature information to update the dialogue feature information in response to the available capacity of the database being greater than or equal to the storage capacity required by one or more feature information items; and replacing at least a portion of the dialogue feature information with at least a portion of one or more feature information items in response to the available capacity of the database being less than the storage capacity required by one or more feature information items.

[0089] In some embodiments, one or more feature information items are each configured with a corresponding first weight. The dialogue feature information includes a plurality of current feature information items, each of which is configured with a corresponding second weight. Replacing at least a portion of the dialogue feature information with at least a portion of one or more feature information items includes: selecting one or more target feature information items from the plurality of current feature information items based on the corresponding first weights of the one or more feature information items and the corresponding second weights of the plurality of current feature information items; and replacing one or more target feature information items with the one or more feature information items.

[0090] In some embodiments, the dialogue feature information includes multiple current feature information items, each of which is configured with a corresponding weight, and determining the prompt information includes: determining the order of the current feature information items in the information based on their respective weights; and generating prompt information based on the feature information items according to the order.

[0091] In some embodiments, the response message is generated by: providing a prompt to a machine learning model to obtain the output of the machine learning model; and determining the response message based on the output of the machine learning model.

[0092] In some embodiments, conversational messages include messages during a real-time voice call between the user and at least one digital assistant.

[0093] Figure 6 shows a block diagram of a device 600 for dialog interaction according to some embodiments of the present disclosure. The device 600 may be implemented as or included in the client device 110 or server device 120 of Figure 1. The various modules / components in the device 600 may be implemented by hardware, software, firmware, or any combination thereof.

[0094] As shown in the figure, device 600 includes a prompt information determination module 610, configured to determine prompt information in response to receiving a dialogue message from a user to at least one digital assistant, based on dialogue feature information related to the user and at least one digital assistant. The dialogue feature information is key information generated by processing historical dialogue content between the user and at least one digital assistant. Device 600 also includes a response message generation module 620, configured to present a response message to the user in response to the dialogue message, the response message being generated based on the dialogue message and the prompt information.

[0095] In some embodiments, a user engages in multiple rounds of dialogue with at least one digital assistant, and the apparatus 600 further includes a memory module configured to: in response to a response message corresponding to the end of a predetermined number of rounds of dialogue between the user and at least one digital assistant, generate memory information corresponding to the predetermined number of rounds of dialogue by extracting key information from the dialogue content of the predetermined number of rounds of dialogue; based on the memory information, generate one or more feature information items to characterize the predetermined number of rounds of dialogue; and update the dialogue feature information using the one or more feature information items.

[0096] In some embodiments, the dialogue feature information includes entity feature information and event feature information, and the memory module is further configured to: extract one or more memory information items related to at least one of the user or at least one digital assistant from the dialogue content; and determine the corresponding category of one or more memory information items through semantic analysis of the one or more memory information items.

[0097] In some embodiments, the apparatus 600 further includes an update module configured to: generate one or more entity feature information items based on memory information items of the category of non-event among one or more memory information items; and generate one or more event information items based on memory information items of the category of event among one or more memory information items.

[0098] In some embodiments, the update module is further configured to determine the corresponding weight of one or more feature information items based on at least one of the following: the corresponding occurrence time of the dialogue corresponding to one or more feature information items, or the corresponding importance of one or more feature information items.

[0099] In some embodiments, dialogue feature information items are stored in a database, and the update module is further configured to: add one or more feature information items to the dialogue feature information to update the dialogue feature information in response to the available capacity of the database being greater than or equal to the storage capacity required by one or more feature information items. The update module is also configured to: replace at least a portion of the dialogue feature information with at least a portion of one or more feature information items in response to the available capacity of the database being less than the storage capacity required by one or more feature information items.

[0100] In some embodiments, one or more feature information items are each configured with a corresponding first weight. The dialogue feature information includes multiple current feature information items, each of which is configured with a corresponding second weight. The update module is further configured to: select one or more target feature information items from the multiple current feature information items based on the corresponding first weights of the one or more feature information items and the corresponding second weights of the multiple current feature information items; and replace one or more target feature information items with the one or more feature information items.

[0101] In some embodiments, the dialogue feature information includes multiple current feature information items, each of which is configured with a corresponding weight, and the prompt information determination module 610 is further configured to: determine the order of the current feature information items in the information based on the corresponding weights of the current feature information items; and generate prompt information based on the feature information items according to the order.

[0102] In some embodiments, the response message generation module 620 is further configured to: provide prompt information to a machine learning model to obtain the output of the machine learning model; and determine a response message based on the output of the machine learning model.

[0103] In some embodiments, conversational messages include messages during a real-time voice call between the user and at least one digital assistant.

[0104] Figure 7 shows a block diagram of a device 700 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 700 shown in Figure 7 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 700 shown in Figure 7 can be used to implement the methods described above.

[0105] As shown in Figure 7, the computing device 700 is in the form of a general-purpose computing device. Components of the computing device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage devices 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 700.

[0106] Computing device 700 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible within computing device 700.

[0107] The computing device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.

[0108] The communication unit 740 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 700 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 700 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.

[0109] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 700 can also communicate as needed with one or more external devices (not shown) via communication unit 740. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 700, or with any device (e.g., network card, modem, etc.) that enables computing device 700 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interface (not shown).

[0110] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.

[0111] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0112] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0113] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0115] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1.A method for conversational interaction, comprising: in response to receiving a conversational message issued by a user to at least one digital assistant, determining prompt information based on conversational feature information related to the user and the at least one digital assistant, the conversational feature information being key information generated by processing historical conversational content of the user and the at least one digital assistant; and presenting a response message to the user for the conversational message, the response message being generated based on the prompt information. 2.The method of claim 1, wherein the user has a plurality of rounds of conversation with the at least one digital assistant, and the method further comprises: in response to the response message corresponding to an end of a predetermined number of rounds of conversation between the user and the at least one digital assistant, generating memory information corresponding to the predetermined number of rounds of conversation by extracting key information in conversational content of the predetermined number of rounds of conversation; generating one or more feature information items for representing the predetermined number of rounds of conversation based on the memory information; and updating the conversational feature information with the one or more feature information items. 3.The method of claim 2, wherein the conversational feature information comprises entity feature information and event feature information, and generating the memory information comprises: extracting one or more memory information items related to at least one of the user or the at least one digital assistant from the conversational content; and determining respective categories of the one or more memory information items by semantic analysis of the one or more memory information items. 4.The method of claim 3, wherein generating the one or more feature information items comprises: generating one or more entity feature information items based on memory information items of the one or more memory information items having a category of non-event; and generating one or more event information items based on memory information items of the one or more memory information items having a category of event. 5.The method of claim 3, wherein generating the one or more feature information items comprises: determining respective weights of the one or more feature information items based on at least one of: respective occurrence times of conversations to which the one or more feature information items correspond, or respective importance degrees of the one or more feature information items. 6.The method of claim 2, wherein the conversational feature information items are stored in an information base, and updating the conversational feature information comprises: in response to an available capacity of the information base being greater than or equal to a storage capacity required by the one or more feature information items, adding the one or more feature information items to the conversational feature information to update the conversational feature information; and in response to the available capacity of the information base being less than the storage capacity required by the one or more feature information items, replacing at least a portion of the conversational feature information with at least a portion of the one or more feature information items. ​ ​ ​ 7.The method of claim 6, wherein the one or more feature information items are respectively configured with respective first weights, the conversation feature information comprises a plurality of current feature information items, the plurality of current feature information items are respectively configured with respective second weights, and replacing at least part of the conversation feature information with at least part of the one or more feature information items comprises: selecting one or more target feature information items from the plurality of current feature information items based on the respective first weights of the one or more feature information items and the respective second weights of the plurality of current feature information items; and replacing the one or more target feature information items with the one or more feature information items. 8.The method of claim 1, wherein the conversation feature information comprises a plurality of current feature information items, the plurality of current feature information items are respectively configured with respective weights, and determining the prompt information comprises: determining an order of the plurality of current feature information items in the prompt information based on the respective weights of the plurality of current feature information items; and generating the prompt information based on the plurality of feature information items according to the order. 9.The method of claim 1, wherein the response message is generated by: providing the prompt information to a machine learning model to obtain an output of the machine learning model; and determining the response message based on the output of the machine learning model. 10.The method of claim 1, wherein the conversation messages comprise messages during a real-time voice call between the user and the at least one digital assistant. 11.An apparatus for conversation interaction, comprising: a prompt information determining module configured to, in response to receiving a conversation message issued by a user to at least one digital assistant, determine prompt information based on conversation feature information related to the user and the at least one digital assistant, the conversation feature information being generated by processing historical conversation content of the user and the at least one digital assistant; and a response message generating module configured to present a response message to the user for the conversation message, the response message being generated based on the conversation message and the prompt information. 12.An electronic device, comprising: a set of processing units; and a set of memories coupled to the set of processing units and storing instructions for execution by the set of processing units, the instructions, when executed by the set of processing units, cause the electronic device to perform the method of any one of claims 1-10. 13.A computer-readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, causing the processor to implement the method of any one of claims 1-10. 14.A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method of any one of claims 1-10. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​

Citation Information

Patent Citations

  • Man-machine conversation generation method and device, electronic equipment and storage medium

    CN114969250A

  • Human-computer interaction method and device, electronic equipment and storage medium

    CN118098217A

  • Method and device for dialogue interaction

    CN118445380A

  • Contextually inferred talking points for improved communication

    US20190042086A1