Intelligent dialogue processing method, device and equipment

By dividing user dialogue records into short-term and long-term memory, and using large models for reasoning, the dialogue continuity problem of dialogue agents is solved, and the naturalness of dialogue and emotional companionship experience is improved.

CN120336474APending Publication Date: 2025-07-18ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510412559.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing dialogue agents lack dialogue continuity and contextual understanding ability, resulting in unnatural or incoherent communication. Especially in the emotional companionship scenario, the conversation experience is fragmented and cannot provide personalized and emotionally resonant dialogue responses.

Method used

By dividing user's dialogue records into short-term memory and long-term memory, they are processed independently, and using large models to reason, generating more natural and personalized dialogue replies.

Benefits of technology

A more natural and smooth intelligent dialogue process is achieved, the emotional companionship experience is improved, and the anthropomorphic and personalized capabilities of dialogue agents are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336474A_ABST
    Figure CN120336474A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an intelligent dialogue processing method, device and equipment. Comprising the steps of receiving first dialogue content input by a user; obtaining a recent dialogue historical record with the user as a short-term memory; obtaining matched long-term memory from long-term memory recorded for the user in advance according to the first dialogue content; generating second dialogue content according to the first dialogue content, the short-term memory and the matched long-term memory; inputting the second dialogue content into a large model for reasoning to obtain corresponding output content, and displaying the output content to the user as a dialogue reply to the first dialogue content; and according to the output content, carrying out record updating on long-term memory recorded for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of automatic dialogue technology, and in particular, to an intelligent dialogue processing method, apparatus, and device. Background Art

[0002] With the continuous development of large model technology, dialogue agents have been widely used in question-and-answer robot fields such as customer service and virtual assistants.

[0003] Ordinary dialogue agents generate appropriate responses by recognizing and parsing user inputs. However, these dialogue agents often lack the ability of dialogue continuity and context understanding, resulting in unnatural or incoherent communication.

[0004] For example, in the scenario of an emotional companionship agent, if a user has multiple conversations with the emotional companionship agent, the initial background of each conversation is blank and the content between conversations is isolated, which will lead to a very fragmented chat experience, far from the dialogue experience between real people in the real world.

[0005] Based on this, a better intelligent dialogue solution is needed. Summary of the Invention

[0006] One or more embodiments of this specification provide an intelligent dialogue processing method, apparatus, and device to solve the following technical problem: a better intelligent dialogue solution is needed.

[0007] To solve the above technical problem, one or more embodiments of this specification are implemented as follows:

[0008] An intelligent dialogue processing method provided by one or more embodiments of this specification includes:

[0009] Receiving first dialogue content input by a user;

[0010] Obtaining the most recent dialogue history record of the user as short-term memory;

[0011] According to the first dialogue content, obtaining a matching long-term memory in the long-term memory pre-recorded for the user;

[0012] Generating second dialogue content according to the first dialogue content, the short-term memory, and the matching long-term memory;

[0013] Inputting the second dialogue content into a large model for inference to obtain corresponding output content, and presenting it to the user as a dialogue reply to the first dialogue content;

[0014] Updating the record of the long-term memory for the user according to the output content.

[0015] An intelligent dialogue processing device provided by one or more embodiments of this specification includes:

[0016] A dialogue content receiving module, which receives the first dialogue content input by the user;

[0017] A short-term memory acquisition module, which acquires the most recent dialogue history with the user as short-term memory;

[0018] A long-term memory acquisition module, which acquires a matching long-term memory from the long-term memory pre-recorded for the user according to the first dialogue content;

[0019] A dialogue content generation module, which generates the second dialogue content according to the first dialogue content, the short-term memory, and the matching long-term memory;

[0020] A dialogue reply generation module, which inputs the second dialogue content into a large model for reasoning, obtains the corresponding output content, and presents it to the user as a dialogue reply to the first dialogue content;

[0021] A long-term memory update module, which updates the long-term memory recorded for the user according to the output content.

[0022] An intelligent dialogue processing device provided by one or more embodiments of this specification includes:

[0023] At least one processor; and,

[0024] A memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute:

[0026] Receive the first dialogue content input by the user;

[0027] Acquire the most recent dialogue history with the user as short-term memory;

[0028] According to the first dialogue content, acquire a matching long-term memory from the long-term memory pre-recorded for the user;

[0029] Generate the second dialogue content according to the first dialogue content, the short-term memory, and the matching long-term memory;

[0030] Input the second dialogue content into a large model for reasoning, obtain the corresponding output content, and present it to the user as a dialogue reply to the first dialogue content;

[0031] Update the record of the long-term memory recorded for the user according to the output content.

[0032] At least one of the above technical solutions adopted by one or more embodiments of this specification can achieve the following beneficial effects: This solution can be implemented for the dialogue agent or its server. The historical conversation records between the user and the dialogue agent are divided into two types: short-term memory and long-term memory, and the independent processing logics of these two memory types are set. When the user is having a real-time conversation with the dialogue agent, according to the current conversation content input by the user, the corresponding short-term and long-term memories can be obtained, and the conversation content can be reconstructed to replace the actual conversation content input by the user and input into the large model for reasoning. Therefore, during the reasoning process of the large model, not only can the short-term memory of the user be utilized naturally, but also the long-term memory of the user can be utilized, which helps to infer a more sufficient and accurate conversation reply that fits the actual conversation content input by the user, making the intelligent conversation process with the user more natural and fluent, and also helping to bring a more real emotional companionship experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0034] Figure 1 It is a flowchart of an intelligent conversation processing method provided for one or more embodiments of this specification;

[0035] Figure 2 Provided for one or more embodiments of this specification Figure 1 It is a flowchart of a specific implementation scheme of the method;

[0036] Figure 3 It is a flowchart of a user long-term memory processing scheme based on other user assistance provided for one or more embodiments of this specification;

[0037] Figure 4 It is a flowchart of a reconstruction processing scheme for the conversation content input by the user provided for one or more embodiments of this specification;

[0038] Figure 5 It is a structural diagram of an intelligent conversation processing device provided for one or more embodiments of this specification;

[0039] Figure 6A schematic structural diagram of an intelligent dialogue processing device provided for one or more embodiments of this specification. Detailed implementation manners

[0040] Embodiments of this specification provide an intelligent dialogue processing method, device, equipment, and storage medium.

[0041] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0042] An intelligent agent refers to a software entity that can autonomously perceive the environment and affect the environment by performing actions. The intelligent agent makes decisions based on its perception and preset rules to complete tasks.

[0043] Regarding the problems mentioned in the background art, this application particularly focuses on specific application scenarios such as emotional companionship. In this scenario, the dialogue intelligent agent is specifically an emotional companionship intelligent agent (for example, a smart home robot, a pet robot, a nursing robot, etc.). Different from general intelligent customer service, the emotional companionship intelligent agent needs to have stronger emotional resonance and communication capabilities. As an artificial intelligence system mainly used to provide emotional support and companionship services, it should be able to better understand and respond to the user's emotional state, and provide psychological comfort and emotional support to the user through dialogue, interaction, etc.

[0044] With the rise of large models, the question-and-answer intelligent agents constructed using large models usually have the ability to reason based on the current dialogue context. However, to construct an emotional companionship intelligent agent, it is not enough to only reason based on the current dialogue context.

[0045] The applicant initially tried two types of solutions, but found that there were still many problems. The first type of solution is that the intelligent agent responds based on rule configuration. The disadvantages of this type of solution include: poor scalability, the need to continuously update rules to handle new situations, and difficulty in processing fuzzy or undefined inputs; the conversations between different users and the emotional companionship intelligent agent are all based on the same rule configuration, which will cause the emotional companionship intelligent agent to lose its character characteristics and lack personalization. The second type of solution is to fine-tune the large model. The disadvantages of this type of solution include: fine-tuning the large model requires a large amount of corpus and machines, with high costs and low returns, and also lacks flexibility in response to demand changes; the emotional companionship intelligent agent based on the fine-tuned large model has data bias and limitations in background information, and cannot reflect personalization during the interaction with users, failing to meet the expectations of true emotional companionship.

[0046] To solve the above problems and meet the actual needs, this application considers that by summarizing and refining historical conversations, the dialogue agent can have short-term and long-term memory capabilities, making the agent more anthropomorphic. This allows users to experience a more natural and fluent conversation process with the dialogue agent, and makes the conversation experience closer to the real world.

[0047] Based on the above general idea, the solution of this application will be further described below.

[0048] Figure 1 It is a schematic flowchart of an intelligent dialogue processing method provided for one or more embodiments of this specification. It can be applied to a dialogue agent, which is carried on a user terminal, such as a smart phone, a tablet computer, a smart watch, a service robot, a car computer, a game console, etc. It should be noted that some steps in this process can also be executed on the corresponding server, which is more efficient.

[0049] Figure 1 The process in [figure] includes the following steps:

[0050] S102: Receive the first conversation content input by the user.

[0051] The user initiates a conversation with the dialogue agent by means of text input or voice input, etc. The specific content input is called the first conversation content. The first conversation content is generally a question raised by the user. In this way, the information of the first conversation content is more abundant, which is convenient for more targeted reasoning and reply in the next step. Of course, the first conversation content can also be a chat-type active statement content, rather than a clear question, and the user's intention is inferred through reasoning.

[0052] S104: Obtain the most recent conversation history record with the user as short-term memory.

[0053] The most recent conversation history record can be the original chat content or the content summarized after refining the chat content. For example, the most recent N rounds of conversations between the user and the dialogue agent are used as short-term memory, and N can be dynamically configured. At the same time, the short-term memory can be more refined by combining dimensions such as the most recent conversation times and the time range from the current time.

[0054] As the conversation continues, the content of the most recent conversation history record can be updated in a timely manner accordingly, so as to ensure the timeliness of "the most recent".

[0055] S106: Obtain the matching long-term memory in the long-term memory pre-recorded for the user according to the first conversation content.

[0056] In one or more embodiments of this specification, a long-term memory for the user is maintained for the dialogue agent, so that the dialogue agent can process and store data on the dialogue content over a longer period of time. Long-term memory allows the dialogue agent to retain and retrieve important historical dialogue information over a longer time span, and in particular enables the emotional companion agent to have its own character characteristics. This application makes it an important component of the anthropomorphism of the agent.

[0057] Furthermore, long-term memory is maintained through a special database, rather than relying on the large model itself to learn long-term memory. Long-term memory can be the chat content of a longer period of time in the past (for example, months, years or even longer), or it can be the content summarized after refining the chat content. Preferably, the latter method is adopted, for example, all historical conversation records of the user are handed over to the large model for induction and summary, and the user generates long-term memory. Long-term memory more fully and specifically reflects the personalized overall picture of the user. In the process of continuous dialogue between the user and the intelligent agent, the recent conversation content and the long-term memory that has been mastered can be handed over to the large model for processing based on predefined rules, and the latest long-term memory is generated again. This allows long-term memory to have the ability to accumulate and update continuously, just like the memory of people in reality has a process of continuous updating and strengthening. In the process of handing over long-term memory to the large model for processing, the large model can be required to subdivide the long-term memory summary into multiple types, such as: key events, key times, key locations, character preferences, interpersonal relationships, etc., so that the long-term memory can be tuned more finely.

[0058] After receiving the first conversation content input by the user, the long-term memory related to the first conversation content is matched in the long-term memory, and the matching method may adopt, for example, text similarity comparison, vector similarity comparison, etc.

[0059] Take vector similarity comparison as an example, in order to improve the comparison efficiency. In addition to maintaining long-term memory in natural language form, a long-term memory vector library is also maintained, which contains long-term memory vectors obtained by vectorizing long-term memory (directly converted into embedding vectors; or further condensing knowledge, inference to obtain inference result vectors, which helps to reduce the total amount of vector data). In this case, when matching long-term memory, specifically, the first conversation content is converted into a first vector; the first vector is compared for similarity in the long-term memory vector library to obtain a second vector similar to the first vector. The vectors in the long-term memory vector library are obtained by converting the long-term memory recorded in advance for the user; the long-term memory corresponding to the second vector is obtained (according to the previous conversion correspondence during vectorization, the long-term memory used for converting the second vector is found) as the long-term memory matching the first conversation content.

[0060] S108: Generate second conversation content based on the first conversation content, the short-term memory, and the matched long-term memory.

[0061] The short-term memory supplements the context knowledge of this conversation to the first conversation content; the matched long-term memory supplements more persistent background knowledge specific to the current user to the first conversation content and the short-term memory. Based on this, reconstruct the first conversation content based on the short-term memory and the matched long-term memory, and intelligently and automatically complete the conversation content without user intervention as the second conversation content. Assume the second conversation content is the more comprehensive conversation content input by the user for subsequent large model inference processing. It should be noted that the second conversation content may not be shown to the user to avoid user misunderstanding.

[0062] In one or more embodiments of this specification, assemble the first conversation content, the short-term memory, and the matched long-term memory, and use the assembled content as the second conversation content. For example, if the first conversation content is expressed in natural language, and the short-term memory and the matched long-term memory can also be expressed in natural language, then the first conversation content, the short-term memory, and the matched long-term memory can be assembled into a longer text expressed in natural language as the second conversation content.

[0063] Of course, during the assembly process, more refined information can also be further extracted as the second conversation content.

[0064] S110: Input the second conversation content into a large model for inference, obtain the corresponding output content, and display it to the user as a conversation reply to the first conversation content.

[0065] A large model refers to a model trained using a large amount of data in the field of machine learning, especially deep learning. These models usually have millions to billions of parameters. Large models have powerful functions and can show excellent performance in tasks such as natural language processing and computer vision. They are trained through large-scale datasets to achieve more complex and advanced tasks. Existing large models can be used for inference, or the large model can be fine-tuned for the current user and then used for inference.

[0066] Since the second conversation content contains memory knowledge specific to the current user, especially long-term memory knowledge, the large model can perform inference more specifically for the current user, which helps to obtain strongly personalized output content with more emotional resonance specific to the current user. In this way, the output content of the large model may behave as if it is a real old friend who has known the current user for many years and knows the current user very well making a sincere response to the current user's conversation, rather than a formulaic and impersonal mechanical response.

[0067] After seeing the output content, the user can further send new conversation content, and the agent can continue to generate new conversation responses in a similar processing manner to continue the conversation with the user.

[0068] S112: Update the long-term memory recorded for the user according to the output content.

[0069] When the output content is obtained, both the short-term memory and the long-term memory are updated. For the short-term memory, the output content can be updated to the most recent conversation history with the user.

[0070] In one or more embodiments of this specification, as the conversation with the user increases, the long-term memory will accumulate. To prevent unnecessary data accumulation, the original conversation records can be refined and used as the long-term memory. In this way, on the one hand, the long-term memory will not expand violently in terms of data volume, and on the other hand, it can also focus on the high-value content part in the memory that is helpful for more effective communication with the user (for example, the part that can more accurately reflect the user's personality, important energy, or the things the user cares about, and the content part that has a relatively large impact on the user).

[0071] Based on this idea, when updating the long-term memory, specifically: the most recent conversation history including the first conversation content and the output content with the user can be obtained as the updated short-term memory; according to all the long-term memory recorded for the user and the updated short-term memory, generate the full-update input content; input the full-update input content into the large model for inference to obtain the updated long-term memory, and use it to overwrite all the long-term memory recorded for the user. To improve efficiency, for example, the server of the conversation agent is used to update the long-term memory.

[0072] Furthermore, the updated long-term memory is vectorized and stored in the long-term memory vector library for obtaining the matching long-term memory when conversing with the user.

[0073] It should be noted that since the output content can also be updated to the most recent conversation history in a timely manner, the update time of the long-term memory can be relaxed and does not necessarily need to be updated immediately, so as to avoid the inference time-consuming of updating the long-term memory from affecting the response speed to the user. For example, the long-term memory can be updated appropriately after the end of the current conversation with the user (which may include multiple rounds of conversations, for example, considering one question and one answer as one round of conversation).

[0074] Through Figure 1The method can implement this solution for the dialogue agent or its server, classify the historical conversation records between the user and the dialogue agent into two types: short-term memory and long-term memory, and set independent processing logics for these two memory types. When the user is having a real-time conversation with the dialogue agent, according to the current conversation content input by the user, the corresponding short-term and long-term memories can be obtained, and the conversation content can be reconstructed to replace the actual conversation content input by the user, and then input into the large model for reasoning. Therefore, during the reasoning process of the large model, not only can the short-term memory of the user be utilized naturally, but also the long-term memory of the user can be utilized, which helps to infer a more sufficient and accurate conversation reply that fits the actual conversation content input by the user, thus making the intelligent conversation process with the user more natural and fluent, and also helping to bring a more real emotional companionship experience.

[0075] Based on Figure 1 the method, this specification also provides some specific implementation schemes and extended schemes of this method, which will be continued to be described below.

[0076] Intuitively, Figure 2 is a schematic flowchart of a specific implementation scheme of the method provided by one or more embodiments of this specification. In Figure 1 the scenario of, this method is specifically applied to the emotional companionship agent. In addition to this end, the process also involves the user, the large model for reasoning, the embedding model for vectorization, the relational database for storing memories, and the vectorized database for storing memory vectors (including the above-mentioned long-term memory vector library, and if there are temporary conversation content vectors, they can also be temporarily stored here). Figure 2 The solution in

[0077] Figure 2 mainly includes the following four parts of steps:

[0078] The first part of the steps: Each time the emotional companionship agent receives the conversation content input by the user, it will respectively obtain the most recent conversation record with the user from the relational database, obtain the matching long-term memory vector from the vectorized database, and recall the corresponding long-term memory; the emotional companionship agent processes the conversation record obtained from the database according to the preset rules (for example, specifically taking the most recent few rounds, whether to perform data cleaning, etc.) as the short-term memory of this conversation.

[0079] The second part of the steps: After obtaining the short-term memory and the relevant long-term memory, the emotional companionship agent assembles the final large model prompt content (usually called prompt), gives it to the large model, and allows the large model to perform reasoning and generate output content.

[0080] Third part steps: After the emotional companion agent obtains the dialogue output content of the large model, it returns it to the front end for presentation to reply to the user. At the same time, the content of this conversation is stored in a relational database, becoming the latest historical record of the conversation between the user and the emotional companion agent, and also serving as the source material for the short-term and long-term memories of subsequent conversations.

[0081] Fourth part steps: After each conversation is completed, the emotional companion agent can initiate an asynchronous process (the immediate requirement is not high, so it can be processed asynchronously) to update the long-term memory. If the update process of the long-term memory for this time has not been executed, it is started; if it has been executed, it is skipped. All historical full long-term memories and the most recent conversation content are handed over to the large model for reasoning to obtain the latest full long-term memory. The latest full long-term memory is vectorized through an embedding model and stored in a vector database so that relevant long-term memories can be quickly and accurately recalled in subsequent conversations.

[0082] The above four parts of the steps can be automatically triggered and executed along with the communication between the user and the emotional companion agent. In this way, it repeats in a cycle. Each conversation can obtain the latest short-term memory and the long-term memory content that is continuously strengthened and updated. Eventually, it can make the emotional companion agent anthropomorphic, and the user can have a more natural and fluent conversation with it, and the agent will have the ability of personalized emotional companionship.

[0083] In one or more embodiments of this specification, it is considered to use the long-term memories of other users to update and correct the long-term memories of the current user. In this way, it helps to eliminate misunderstandings and prejudices between users, and also helps to generate more interesting and friendly reply content for users. Based on this idea, a flow diagram of a user long-term memory processing solution based on the assistance of other users is provided. See Figure 3 。

[0084] Figure 3 The solution in

[0085] S302: Take the user as the target user and determine one or more users related to the target user.

[0086] Users related to the target user can be automatically mined according to the social relationship data of the target user. More reliably, the target user can proactively specify one or more users in their social user set as the users related to the target user.

[0087] In addition, the user can also be provided with the option of whether to use their own long-term memory to assist other users. Of course, the degree of assistance can also be subdivided. For example, the user can open a general level of assistance for friends, and a higher level of assistance for relatives or partners.

[0088] S304: In the long-term memory recorded for the relevant user, check whether there is a target long-term memory that can assist the target user.

[0089] The step of checking whether there is a target long-term memory that can assist the target user in the long-term memory recorded for the relevant user specifically includes:

[0090] In one or more embodiments of the present specification, check whether there is at least one of the following types of long-term memories in the long-term memory recorded for the relevant user: a target long-term memory related to the target user, such as a long-term memory that directly mentions the target user, a long-term memory about something that the target user also participated in, etc.; a target long-term memory similar to the first conversation content, such as a long-term memory with a similar question to the current user, a long-term memory in a similar scenario to the current user, etc.; a target long-term memory that can reply to the first conversation content, such as a long-term memory that can directly serve as the answer to the first conversation content, a long-term memory that can provide indirectly inspiring information for obtaining the answer, etc.

[0091] S306: If there is such a memory, then perform subsequent steps to update or correct the long-term memory recorded for the target user according to the target long-term memory.

[0092] Generally, the long-term memory is updated by adding the target long-term memory to the existing long-term memory of the target user (of course, actions such as refinement and vector transformation can be performed as needed).

[0093] To further improve reliability, it is also possible to analyze whether the target long-term memory conflicts with the existing long-term memory of the target user. If there is a conflict, then further comparative analysis can be performed to determine whether there is a misunderstanding in the target long-term memory or in the existing long-term memory of the target user. If there is a misunderstanding in the existing long-term memory of the target user, then corresponding corrections are made according to the target long-term memory. On the contrary, the target long-term memory can be abandoned. Additionally, in practical applications, there may not necessarily be a clear distinction between right and wrong in both cases. In such a situation, the target long-term memory can be retained to a certain extent, and then when inferring the dialogue reply, the retained target long-term memory can be introduced or its weight can be increased, so that the output dialogue reply can attempt to weakly guide the user in the direction of the target long-term memory, and then, based on the user's reaction, it is decided whether to retain the target long-term memory permanently.

[0094] If there is no target long-term memory that can assist the target user, then the long-term memory can be updated using the newly generated conversation record of the target user itself.

[0095] S308: Calculate the positive sentiment degree of the target long-term memory for the target user.

[0096] Perform semantic analysis on the target long-term memory to distinguish the content with positive emotional color from other content (if necessary, the content with negative emotional color can be distinguished from it). The higher the proportion of the content with positive emotional color, the higher the positive emotion degree.

[0097] Furthermore, especially according to the situation of the target user, it can be judged whether the part of the content with positive emotional color has a higher or lower bonus for the target user, so as to introduce the corresponding weighting factor to calculate the positive emotion degree of the target long-term memory for the target user.

[0098] S310: Judge whether the positive emotion degree is greater than the set threshold.

[0099] For the target user, it is inclined to introduce the target long-term memory with a sufficiently high positive emotion degree of other users. In this way, it can avoid bringing negative impacts to the target user, and it also helps to protect the interests of other users. Due to this assistance, it may be carried out mutually among different users. Therefore, globally, it helps to improve the positive impact of the long-term memory and also helps to improve the enthusiasm of users for intelligent conversations.

[0100] S312: If so, update or correct the long-term memory recorded for the target user according to the target long-term memory.

[0101] If the positive emotion degree is greater than the set threshold, it is inclined to use the target long-term memory to assist the target user; otherwise, it is not inclined to assist. Of course, in addition to the threshold comparison method, some other decision-making means can also be adopted to judge whether the positive emotion degree is high enough.

[0102] If the positive emotion degree is not greater than the set threshold, the update or correction can be abandoned, or the part with a high positive emotion degree can be segmented from the target long-term memory for update or correction.

[0103] Furthermore, one or more embodiments of this specification also provide a flowchart of a reconstruction processing scheme for the dialogue content input by the user. Refer to Figure 4 .

[0104] Figure 4 The scheme in

[0105] S402: Before updating or correcting the long-term memory recorded for the target user according to the target long-term memory, generate the second dialogue content based on the first dialogue content, the short-term memory, the matched long-term memory, and the target long-term memory, so as to infer the corresponding output content.

[0106] In one or more embodiments of this specification, the long-term memories of other users (i.e., the target long-term memories) are introduced to reconstruct the conversation content of the current user, which helps to improve the objectivity of the reasoning basis and also helps to learn the blind spot knowledge about the current user.

[0107] S404: After obtaining the corresponding output content and presenting it to the user, receive the third conversation content input by the user in response to the output content.

[0108] S406: Calculate the first similarity between the third conversation content and the matched long-term memory, and calculate the second similarity between the third conversation content and the target long-term memory.

[0109] Based on the user's reaction to the output content, inversely verify whether the actual value of the introduced target long-term memory for replying to the current user is high enough.

[0110] If the similarity between the third conversation content and the target long-term memory is high enough, it indicates that the user is more interested in the content part inferred based on the introduced target long-term memory, and the user may be further discussing this content part by inputting the third conversation content. In this case, it can be considered whether the actual value of the target long-term memory is high enough.

[0111] To more objectively measure the actual value of the target long-term memory, it is horizontally compared with the user's own long-term memory (i.e., the matched long-term memory) used for this reasoning. Calculate the similarities between the two and the third conversation content respectively to determine which part of the reasoning content the user tends to focus on.

[0112] S408: Determine whether the second similarity is greater than the first similarity.

[0113] If the second similarity is greater than the first similarity, it indicates that the user may be more inclined to focus on the reply content part inferred based on the target long-term memory. Otherwise, it indicates that the user may be more inclined to focus on the reply content part inferred based on the matched long-term memory.

[0114] S410: If so, determine to update or correct the long-term memory recorded for the target user according to the target long-term memory.

[0115] If so, it indicates that the actual value of the target long-term memory for the user may be high, so it can be used to update or correct the user's existing long-term memory. Otherwise, such update or correction may not be performed.

[0116] By Figure 4The solution in [it] helps to more reliably and carefully utilize the long-term memory of other users to assist the dialogue agent in automatically replying to the target user and to assist in maintaining the long-term memory of the target user himself / herself.

[0117] In addition, based on a similar idea, the target long-term memory can be introduced first to try to simulate the user's reaction. For example, not only can the second dialogue content be inferred, but also an attempt can be made to infer the third dialogue content, without the user having to input the third dialogue content himself / herself for the time being. Only when the inference result is relatively positive will the second dialogue content be officially presented to the user.

[0118] Based on the same idea, one or more embodiments of this specification also provide the corresponding devices and equipment for the above method, as Figure 5 、 Figure 6 shown. The devices and equipment can correspondingly execute the above method and related alternative solutions.

[0119] Figure 5 is a schematic structural diagram of an intelligent dialogue processing device provided by one or more embodiments of this specification. The device includes:

[0120] A dialogue content receiving module 502 that receives the first dialogue content input by the user;

[0121] A short-term memory acquisition module 504 that acquires the most recent dialogue history record of the user as the short-term memory;

[0122] A long-term memory acquisition module 506 that acquires the matching long-term memory in the long-term memory pre-recorded for the user according to the first dialogue content;

[0123] A dialogue content generation module 508 that generates the second dialogue content according to the first dialogue content, the short-term memory, and the matching long-term memory;

[0124] A dialogue reply generation module 510 that inputs the second dialogue content into a large model for inference, obtains the corresponding output content, and presents it to the user as a dialogue reply to the first dialogue content;

[0125] A long-term memory update module 512 that updates the record of the long-term memory recorded for the user according to the output content.

[0126] Optionally, the long-term memory acquisition module 506 converts the first dialogue content into a first vector;

[0127] Performs a similarity comparison of the first vector in the long-term memory vector library to obtain a second vector similar to the first vector. The vectors in the long-term memory vector library are obtained by converting the long-term memory pre-recorded for the user;

[0128] Obtain the long-term memory corresponding to the second vector as the long-term memory matching the first conversation content.

[0129] Optionally, the conversation content generation module 508 assembles the first conversation content, the short-term memory, and the matching long-term memory expressed in natural language to generate text expressed in natural language as the second conversation content.

[0130] Optionally, the long-term memory update module 512 obtains the most recent conversation history record of the user including the first conversation content and the output content as the updated short-term memory;

[0131] Generate a full-update input content according to all the long-term memories recorded for the user and the updated short-term memory;

[0132] Input the full-update input content into the large model for inference to obtain the updated long-term memory, and use it to overwrite all the long-term memories recorded for the user.

[0133] Optionally, after inputting the full-update input content into the large model for inference to obtain the updated long-term memory, the long-term memory acquisition module 506 performs vectorization processing on the updated long-term memory and stores it in the long-term memory vector library for obtaining the matching long-term memory when conversing with the user.

[0134] Optionally, the long-term memory update module 512 determines one or more users related to the target user with the user as the target user;

[0135] Search in the long-term memories recorded for the related users to find whether there is a target long-term memory that can assist the target user;

[0136] If so, update or correct the long-term memory recorded for the target user according to the target long-term memory.

[0137] Optionally, the long-term memory update module 512 calculates the positive sentiment degree of the target long-term memory for the target user;

[0138] Judge whether the positive sentiment degree is greater than a set threshold;

[0139] If so, update or correct the long-term memory recorded for the target user according to the target long-term memory;

[0140] Otherwise, abandon the update or correction, or split out the part with high positive sentiment degree from the target long-term memory for the update or correction.

[0141] Optionally, before updating or correcting the long-term memory recorded for the target user according to the target long-term memory, the conversation content generation module 508 generates second conversation content based on the first conversation content, the short-term memory, the matched long-term memory, and the target long-term memory.

[0142] After obtaining the corresponding output content and presenting it to the user, the conversation content receiving module 502 receives third conversation content input by the user in response to the output content.

[0143] The long-term memory update module 512 calculates a first similarity between the third conversation content and the matched long-term memory, and calculates a second similarity between the third conversation content and the target long-term memory.

[0144] Determine whether the second similarity is greater than the first similarity.

[0145] If so, determine to update or correct the long-term memory recorded for the target user according to the target long-term memory.

[0146] Otherwise, do not perform the update or correction.

[0147] Optionally, the long-term memory update module 512 determines one or more users actively designated by the target user in its social user set as the users related to the target user.

[0148] In the long-term memory recorded for the related users, check whether there is at least one of the following types of long-term memory: the target long-term memory related to the target user; the target long-term memory similar to the first conversation content; the target long-term memory that can reply to the first conversation content.

[0149] Optionally, for an emotional companionship intelligent agent.

[0150] Figure 6 The figure is a schematic structural diagram of an intelligent conversation processing device provided by one or more embodiments of the present specification. The device includes:

[0151] At least one processor; and,

[0152] A memory communicatively connected to the at least one processor; wherein,

[0153] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute:

[0154] Receive the first conversation content input by the user;

[0155] Obtain the most recent conversation history of the user as short-term memory;

[0156] According to the first conversation content, obtain the matching long-term memory in the long-term memory pre-recorded for the user;

[0157] Generate the second conversation content according to the first conversation content, the short-term memory, and the matching long-term memory;

[0158] Input the second conversation content into the large model for reasoning, obtain the corresponding output content, and display it to the user as the conversation reply to the first conversation content;

[0159] Update the record of the long-term memory recorded for the user according to the output content.

[0160] Based on the same idea, one or more embodiments of this specification also provide a non-volatile computer storage medium, and the medium stores computer-executable instructions, and the computer-executable instructions are set as:

[0161] Receive the first conversation content input by the user;

[0162] Obtain the most recent conversation history of the user as short-term memory;

[0163] According to the first conversation content, obtain the matching long-term memory in the long-term memory pre-recorded for the user;

[0164] Generate the second conversation content according to the first conversation content, the short-term memory, and the matching long-term memory;

[0165] Input the second conversation content into the large model for reasoning, obtain the corresponding output content, and display it to the user as the conversation reply to the first conversation content;

[0166] Update the record of the long-term memory recorded for the user according to the output content.

[0167] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, today, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply making a little logical programming of the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0168] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0169] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0170] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0171] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0172] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0173] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0174] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0175] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0176] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0177] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0178] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity or device including the element.

[0179] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0180] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0181] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0182] The foregoing is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and variations can be made to one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.

Claims

1. An intelligent dialogue processing method, comprising: Receiving the first dialogue content input by the user; Obtaining the most recent dialogue history of the user as short-term memory; Obtaining a matching long-term memory in the long-term memory pre-recorded for the user according to the first dialogue content; Generating a second dialogue content according to the first dialogue content, the short-term memory, and the matching long-term memory; Inputting the second dialogue content into a large model for inference to obtain corresponding output content, and presenting it to the user as a dialogue reply to the first dialogue content; Updating the record of the long-term memory for the user according to the output content.

2. The method according to claim 1, wherein the step of obtaining a matching long-term memory in the long-term memory pre-recorded for the user according to the first dialogue content specifically comprises: Converting the first dialogue content into a first vector; Performing a similarity comparison on the first vector in a long-term memory vector library to obtain a second vector similar to the first vector, and the vectors in the long-term memory vector library are obtained by converting the long-term memory pre-recorded for the user; Obtaining the long-term memory corresponding to the second vector as the long-term memory matching the first dialogue content.

3. The method according to claim 1, wherein the step of generating a second dialogue content according to the first dialogue content, the short-term memory, and the matching long-term memory specifically comprises: Assembling the first dialogue content, the short-term memory, and the matching long-term memory expressed in natural language to generate a text expressed in natural language as the second dialogue content.

4. The method according to claim 1 or 2, wherein the step of updating the record of the long-term memory for the user according to the output content specifically comprises: Obtaining the most recent dialogue history of the user including the first dialogue content and the output content as updated short-term memory; Generating full-update input content according to all the long-term memories recorded for the user and the updated short-term memory; Inputting the full-update input content into a large model for inference to obtain updated long-term memory, and using it to overwrite all the long-term memories recorded for the user.

5. The method according to claim 4, after the step of inputting the full-update input content into a large model for inference to obtain updated long-term memory, the method further comprises: Performing vectorization processing on the updated long-term memory and storing it in the long-term memory vector library for obtaining a matching long-term memory when conversing with the user.

6. The method according to claim 1, further comprising: Taking the user as the target user and determining one or more users related to the target user; Searching in the long-term memories recorded for the related users to find out whether there is a target long-term memory that can assist the target user; If so, updating or correcting the long-term memory recorded for the target user according to the target long-term memory.

7. The method according to claim 6, wherein updating or correcting the long-term memory recorded for the target user according to the target long-term memory specifically includes: Calculating the positive sentiment degree of the target long-term memory for the target user; Judging whether the positive sentiment degree is greater than a set threshold; If so, updating or correcting the long-term memory recorded for the target user according to the target long-term memory; Otherwise, abandoning the update or correction, or splitting out the part with a high positive sentiment degree from the target long-term memory for the update or correction.

8. The method according to claim 6, wherein generating the second dialogue content according to the first dialogue content, the short-term memory and the matched long-term memory specifically includes: Before updating or correcting the long-term memory recorded for the target user according to the target long-term memory, generating the second dialogue content according to the first dialogue content, the short-term memory, the matched long-term memory, and the target long-term memory; After presenting the corresponding output content to the user, the method further includes: Receiving the third dialogue content input by the user for the output content; Calculating a first similarity between the third dialogue content and the matched long-term memory, and calculating a second similarity between the third dialogue content and the target long-term memory; Judging whether the second similarity is greater than the first similarity; If so, determining to update or correct the long-term memory recorded for the target user according to the target long-term memory; Otherwise, not performing the update or correction.

9. The method according to claim 7, wherein determining one or more users related to the target user specifically includes: Determining one or more users actively designated by the target user in his social user set as the users related to the target user; The step of searching for whether there is a target long-term memory in the long-term memory recorded for the related user that can assist the target user specifically includes: Searching in the long-term memory recorded for the related user for whether there is at least one of the following types of long-term memories: the target long-term memory related to the target user; The target long-term memory similar to the first dialogue content; the target long-term memory that can reply to the first dialogue content.

10. The method according to claim 1 is used for an emotional companionship intelligent agent.

11. An intelligent dialogue processing device includes: A dialogue content receiving module, which receives the first dialogue content input by the user; A short-term memory acquisition module, which acquires the most recent dialogue history record of the user as the short-term memory; A long-term memory acquisition module, which acquires the matched long-term memory in the long-term memory pre-recorded for the user according to the first dialogue content; A dialogue content generation module, which generates the second dialogue content according to the first dialogue content, the short-term memory and the matched long-term memory; A dialogue reply generation module, which inputs the second dialogue content into a large model for inference to obtain the corresponding output content and presents it to the user as a dialogue reply to the first dialogue content; The long-term memory update module updates the long-term memory recorded for the user according to the output content.

12. The apparatus according to claim 11, wherein the long-term memory acquisition module converts the first conversation content into a first vector; Performs a similarity comparison on the first vector in the long-term memory vector library to obtain a second vector similar to the first vector, and the vectors in the long-term memory vector library are obtained by converting the long-term memory previously recorded for the user; Obtains the long-term memory corresponding to the second vector as the long-term memory matching the first conversation content.

13. The apparatus according to claim 11, wherein the conversation content generation module assembles the first conversation content expressed in natural language, the short-term memory, and the matching long-term memory to generate a text expressed in natural language as the second conversation content.

14. The apparatus according to claim 11 or 12, wherein the long-term memory update module obtains the most recent conversation history record including the first conversation content and the output content of the user as the updated short-term memory; Generates a full-update input content according to all the long-term memories recorded for the user and the updated short-term memory; Inputs the full-update input content into a large model for inference to obtain an updated long-term memory, and uses it to overwrite all the long-term memories recorded for the user.

15. The apparatus according to claim 14, wherein the long-term memory acquisition module, after inputting the full-update input content into a large model for inference to obtain an updated long-term memory, vectorizes the updated long-term memory and stores it in the long-term memory vector library for obtaining a matching long-term memory when conversing with the user.

16. The apparatus according to claim 11, wherein the long-term memory update module takes the user as the target user and determines one or more users related to the target user; Checks whether there is a target long-term memory in the long-term memories recorded for the related users that can assist the target user; If so, updates or corrects the long-term memory recorded for the target user according to the target long-term memory.

17. The apparatus according to claim 16, wherein the long-term memory update module calculates the positive sentiment degree of the target long-term memory for the target user; Determines whether the positive sentiment degree is greater than a set threshold; If so, updates or corrects the long-term memory recorded for the target user according to the target long-term memory; Otherwise, abandons the update or correction, or splits the part with a high positive sentiment degree from the target long-term memory for the update or correction.

18. The apparatus according to claim 16, wherein the conversation content generation module generates a second conversation content according to the first conversation content, the short-term memory, the matching long-term memory, and the target long-term memory before updating or correcting the long-term memory recorded for the target user according to the target long-term memory. After obtaining the corresponding output content and presenting it to the user, the dialogue content receiving module receives the third dialogue content input by the user for the output content. The long-term memory updating module calculates a first similarity between the third dialogue content and the matched long-term memory, and calculates a second similarity between the third dialogue content and the target long-term memory. Determine whether the second similarity is greater than the first similarity. If so, determine to update or correct the long-term memory recorded for the target user according to the target long-term memory. Otherwise, do not perform the update or correction.

19. The apparatus according to claim 17, wherein the long-term memory updating module determines one or more users actively designated by the target user in the target user's social user set as the users related to the target user. In the long-term memory recorded for the related users, check whether there is at least one of the following types of long-term memories: the target long-term memory related to the target user; the target long-term memory similar to the first dialogue content; the target long-term memory that can reply to the first dialogue content.

20. The apparatus according to claim 1, for an emotional companionship intelligent agent.

21. An intelligent dialogue processing device, comprising: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute: Receive the first dialogue content input by the user; Obtain the most recent dialogue history record of the user as the short-term memory; According to the first dialogue content, obtain the matched long-term memory in the long-term memory pre-recorded for the user; Generate the second dialogue content according to the first dialogue content, the short-term memory and the matched long-term memory; Input the second dialogue content into a large model for inference to obtain the corresponding output content, and present it to the user as a dialogue reply to the first dialogue content; Record and update the long-term memory recorded for the user according to the output content.

Citation Information

Cited By

  • Virtual human stylized interactive long-term emotion accompanying dialogue generation method

    CN121168497A