Intelligent dialogue method and electronic equipment

By introducing the division and asynchronous update mechanism of memory management modules into large language models, the problems of incomplete dialogue memory and delayed response in multiple rounds of dialogue are solved, and fast response and high-accurate response content generation are achieved.

CN119988532APending Publication Date: 2025-05-13ZHEJIANG TMALL TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202411803002.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-05-13

Smart Images

  • Figure CN119988532A_ABST
    Figure CN119988532A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an intelligent dialogue method and electronic equipment, and the method comprises the steps: receiving dialogue content inputted by a user in a current dialogue round in a dialogue process of performing dialogue with the user through a large-scale language model (LLM); requesting to obtain the memory content of the current session from a memory management module; wherein the memory management module comprises a first memory module and a second memory module, and the first memory module is used for storing memory abstract content generated according to historical dialogue content; the second memory module is used for storing a newly generated dialogue content original text in the current dialogue; according to the memory abstract content currently stored in the first memory module and the dialogue content original text stored in the second memory module, prompt information used for guiding an LLM model to execute a content generation task is generated. Through the embodiment of the invention, the user request can be quickly responded, and the memory integrity is realized, so that the accuracy of the reply content is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent dialogue technology, and in particular to an intelligent dialogue method and electronic equipment. Background Art

[0002] LLMs (Large Language Models) have been widely used in multi-round dialogue applications such as intelligent customer service, intelligent assistants, and emotional companionship because they can understand and generate more natural and richer text content. In these application scenarios, the memory ability of the dialogue system is particularly important because it directly affects the coherence and contextual understanding of the dialogue. However, human-computer dialogue systems based on large models face several challenges in achieving effective dialogue memory, which are specifically manifested in the following aspects:

[0003] 1) Limitation on the number of input tokens:

[0004] Token is the basic unit of natural language processing by LLM, which converts the original natural language text into a form that the model can understand and operate. By segmenting the text into tokens, LLM can better understand and generate text content. However, most large language models have strict limits on the number of tokens for a single input. This limitation means that as the number of conversation rounds increases, the complete conversation history cannot be passed to the model as input. For example, in a customer support scenario involving complex problem solving, if the conversation has been going on for dozens of rounds, the early conversation content may be truncated because it exceeds the token limit, affecting the model's understanding of the entire conversation context. In this case, the model may not be able to effectively use early conversation information to generate coherent and contextually relevant responses.

[0005] 2) Limitations of long text processing capabilities:

[0006] Although large language models perform well in processing long texts, not all information is equally important. In a multi-round conversation, the most recent rounds of communication are often more critical to the response in the current round. If too much early conversation content is fed into the model together, it will not only consume limited token resources, but may also make it difficult for the model to focus on recent important information. For example, in an emotional support conversation, a user may express an emotion or need in the first few rounds and ask new questions in subsequent conversations. If the model focuses too much on early expressions, it may ignore the user's latest needs, resulting in responses that are not accurate or relevant enough. Therefore, how to effectively filter and retain key information becomes an important issue.

[0007] 3) Impact of response time:

[0008] Large language models usually take longer to process long inputs than short inputs. In real-time dialogue systems, response time is a very important performance metric, and users expect fast feedback. Processing long inputs not only increases the computational burden, but can also cause response delays, which can affect the user experience. For example, in intelligent customer service scenarios, fast and accurate responses are essential, while long processing times can increase user anxiety. Therefore, optimizing the model's processing speed and response time is one of the key factors in improving user experience. Summary of the invention

[0009] The present application provides an intelligent dialogue method and an electronic device that can quickly respond to user requests and achieve memory integrity so that the accuracy of the reply content is guaranteed.

[0010] This application provides the following solutions:

[0011] An intelligent dialogue method, comprising:

[0012] In a conversation process with a user through a large language model LLM, receiving the conversation content input by the user in the current conversation round;

[0013] Requesting the memory management module to obtain the memory content of the current session; wherein the memory management module includes a first memory module and a second memory module, the first memory module is used to store the memory summary content generated according to the historical conversation content generated in the current session, and update the memory summary content asynchronously with the process of generating the reply content; the second memory module is used to store the original text of the conversation content newly generated in the current session and not yet formed into a memory summary in the first memory module;

[0014] Based on the memory summary content currently stored in the first memory module and the original text of the conversation content stored in the second memory module, prompt information is generated to guide the LLM model to perform a content generation task, so that the LLM model can generate reply content for replying to the conversation content input by the user.

[0015] Among them, the second memory module manages the original text of the conversation content by means of a sliding window, and the length of the sliding window changes dynamically with the increase in the number of conversation content rounds newly generated in the current session that have not yet formed a memory summary in the first memory module; so that the second memory module can save the original text of the conversation content within the current sliding window range.

[0016] After the length of the sliding window reaches a first threshold length, a process of regenerating memory summary content based on the conversation content within the first threshold length starting from the first round of conversation content in the second memory module is triggered, wherein the process of regenerating memory summary content is performed asynchronously with the process of generating reply content through a separate thread;

[0017] After the length of the sliding window reaches a second threshold length, the conversation content within the first threshold length starting from the first round of conversation content in the second memory module is deleted, and the regenerated memory summary content is updated to the first memory module.

[0018] Among them, after triggering the process of regenerating the memory summary content, the memory summary content stored in the first memory module remains unchanged, and the length of the sliding window in the second memory module gradually increases with the newly generated dialogue content rounds based on the first threshold length, until the length of the sliding window in the second memory module reaches the second threshold length, and the regenerated memory summary content is updated to the first memory module.

[0019] Among them, when the memory summary content is regenerated, it is generated by the LLM model based on the memory summary content currently stored in the first memory module and the original text of the conversation content within the first threshold length starting from the first round of conversation content in the second memory module.

[0020] The second threshold length is greater than the first threshold length.

[0021] An intelligent dialogue device, comprising:

[0022] A dialogue request receiving unit, used for receiving the dialogue content input by the user in the current dialogue round in the conversation process of conducting a dialogue with the user through the large language model LLM;

[0023] A memory content acquisition unit, used for requesting the memory management module to acquire the memory content of the current session; wherein the memory management module includes a first memory module and a second memory module, wherein the first memory module is used for storing the memory summary content generated according to the historical conversation content generated in the current session, and performing asynchronous memory summary content update with the process of generating the reply content; and the second memory module is used for storing the original text of the conversation content newly generated in the current session and not yet forming a memory summary in the first memory module;

[0024] A reply content generation unit is used to generate prompt information for guiding the LLM model to perform a content generation task based on the memory summary content currently stored in the first memory module and the original text of the conversation content stored in the second memory module, so that the LLM model can generate reply content for replying to the conversation content input by the user.

[0025] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any of the methods described above.

[0026] An electronic device, comprising:

[0027] one or more processors; and

[0028] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of any of the methods described above.

[0029] A computer program product comprises a computer program / computer executable instructions, wherein the computer program / computer executable instructions implement the steps of any of the aforementioned methods when executed by a processor in an electronic device.

[0030] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0031] Through the embodiment of the present application, the memory management module can be divided into a first memory module (which can be called long-term memory) and a second memory module (which can be called short-term memory), wherein the first memory module can be used to save the memory summary content generated according to the historical conversation content generated in the current session, and update the memory summary content asynchronously with the process of generating the reply content; the second memory module is used to save the original text of the conversation content newly generated in the current session that has not yet formed a memory summary in the first memory module. In this way, when it is necessary to generate the reply content according to the conversation content input by the user, the prompt information for guiding the LLM model to perform the content generation task can be generated according to the memory summary content currently saved in the first memory module and the original text of the conversation content saved in the second memory module, so that the LLM model can generate the reply content for replying to the conversation content input by the user. In this way, since the memory summary content is generated for the historical conversation record, the length of the input data can be shortened, and the requirements of the LLM model for the length of the input data can be met. At the same time, there is no need to wait for the update of the memory summary content, so as to avoid blocking the main processing thread, shorten the reply delay, and ensure a quick response to the user request. As for the conversation content that has not yet been updated in the first memory module, it can be supplemented by the original text of the conversation content stored in the second memory module, so as to achieve the integrity of the memory and ensure the accuracy of the generated reply content.

[0032] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0034] Figure 1 It is a schematic diagram of the system architecture provided by the embodiment of the present application;

[0035] Figure 2 is a flow chart of the method provided in the embodiment of the present application;

[0036] Figure 3 is a schematic diagram of a conversation memory structure provided in an embodiment of the present application;

[0037] Figure 4 It is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0039] In order to facilitate understanding of the solution provided by the embodiment of the present application, it is first necessary to explain that, in order to solve the various problems described in the background technology section, a feasible solution is to limit the overall input of the model by using a sliding window with a fixed number of rounds, so as to solve the problem of too long model input content caused by directly inputting all historical conversation content. For example, the sliding window size of the model input can be limited to 30 rounds to retain the memory of the most recent conversation round. However, this method simply truncates the conversation history by the number of rounds, and cannot achieve the memory of longer conversation rounds.

[0040] Another solution is a synchronous update mechanism based on long-term memory, that is, after each round of dialogue (a question and answer between a user and an intelligent customer service is called a round of dialogue), all historical dialogue contents including the previous round of dialogue can be summarized, and the memory summary content can be obtained and updated and saved; after receiving the user's dialogue content in the next round of dialogue, the updated memory summary content and the currently received dialogue content can be used as the input content of the LLM model to generate specific reply content. In this way, since the specific memory summary content is generated after summarizing the historical dialogue content, its length will be much smaller than the original text of all historical dialogue content. Therefore, it can solve the problem of too long model input content. At the same time, it does not involve truncation of the dialogue history, so it can achieve memory of all dialogue rounds. However, the problem with this solution is that the synchronous update mechanism of the memory summary content may cause response delays. This is because, in the above-mentioned synchronous update mechanism, each time the conversation content input by the user is received and the memory summary content is requested, it is necessary to wait for the last round of memory summary content to be generated and updated before the latest generated memory summary content can be obtained. However, since each round of memory summary content generation requires the historical memory summary content and the original text of the previous round of conversation content to be re-understood and processed, this process may also involve the processing of a large model, which takes a long time. Therefore, when receiving the next round of conversation content from the user, the memory summary content of the previous round has not been updated. At this time, it is necessary to wait until the memory summary content is updated, and the LLM model can generate the reply content. Obviously, this will cause response delays, that is, after the user enters a new round of conversation content, it may take a long time to receive the reply content.

[0041] In view of the above situation, in the embodiment of the present application, a processing method combining asynchronously updated long-term memory and short-term memory is provided to solve the above problem. Specifically, the memory management module in the embodiment of the present application can be divided into two parts, which can be respectively referred to as the first memory module and the second memory module for the convenience of description. Among them, the first memory module corresponds to the asynchronously updated long-term memory module, and the second memory module corresponds to the short-term memory module. Regarding the first memory module, the historical conversation content generated in the current session can be generated into a memory summary content and saved. The memory summary content in the first memory module can be asynchronously updated, that is, it does not need to be processed synchronously with the process of generating the reply content, but after receiving the conversation content input by the user, it is not necessary to wait for the currently executed memory summary content to be updated to complete, and the memory summary content currently saved in the first memory module can be directly returned to generate the reply content. Of course, since the memory summary content may not have been updated yet, the memory summary content at this time may be incomplete, that is, the conversation content of the most recent rounds may not have been reflected in the memory summary content. Therefore, in an embodiment of the present application, in order to ensure the integrity of the conversation content, the conversation content that has not yet generated a memory summary can also be completed through a short-term memory mechanism, so the aforementioned second memory module appears. In the second memory module, the original text of the conversation content that has been recently generated in the current session and has not yet formed a memory summary in the first memory module can be saved. In this way, after the user enters the conversation content, the memory summary content currently saved in the first memory module and the newly generated original text of the conversation content that has not yet formed a memory summary in the first memory module can be assembled to jointly generate a prompt information (Prompt) for guiding the LLM model, so that the LLM model can generate corresponding reply content based on the above input information.

[0042] In this way, since the historical conversation content is summarized and summarized by the first memory module, the memory summary content is generated, which is convenient for meeting the LLM model's requirements for the length of the input content. At the same time, since the first memory module can adopt an asynchronous update mechanism, it can avoid waiting for the update to be completed, thereby avoiding response delays. In addition, for the part of the first memory module that has not yet completed the memory summary update, it can be supplemented by the original text of the conversation content stored in the second memory module, so as to ensure the integrity of the memory and avoid content truncation, thereby ensuring the accuracy of the generated reply content.

[0043] Among them, regarding the asynchronous update of the memory summary content by the first memory module, a specific memory summary content generation task can be executed by a thread running in the background, and the memory summary content in the first memory module can be updated, as long as the update of the memory summary content and the provision of the memory summary content query result are performed asynchronously. Of course, there can be a variety of specific implementation methods. For example, in one method, the memory summary content can be generated and updated again after each round of dialogue. Of course, since the embodiment of the present application also involves cooperation with the second memory module, in another way, in order to achieve more efficient cooperation between the first memory module and the second memory module, a sliding window method with a dynamic adaptive length can also be used to manage the short-term memory in the second memory module. At the same time, the generation and update of the memory summary content can be triggered according to the length of the sliding window.

[0044] Specifically, the so-called sliding window with dynamic adaptive length means that the length of the sliding window in the embodiment of the present application is not fixed. Before the memory summary content in the first memory module has not been updated, the length of the sliding window can gradually increase with the increase in the number of newly generated dialogue rounds. When the length of the sliding window reaches a preset first threshold length (for example, 20 rounds), the generation of the memory summary content can be triggered, that is, it is not necessary to regenerate a new memory summary content after each round of dialogue, but the memory summary content can be regenerated once every certain number of rounds. Since the generation process of the memory summary content requires a certain amount of time, after the regeneration of the memory summary content is triggered, the specific generation task can be entered into the background. During this process, the memory summary content of the first memory module will not be updated, but since the dialogue between the user and the model is still continuing, the length of the sliding window in the second memory module continues to increase. Before the sliding window length reaches the second threshold length, the memory summary content is usually generated. When the sliding window length reaches the second threshold length, the memory summary content in the first memory module is updated, and the dialogue content from the first position to the first threshold length in the second memory module is deleted. In this way, the generation and update of the memory summary content can be performed at certain intervals, so it is easier to achieve mutual cooperation between the first memory module and the second memory module.

[0045] From the perspective of system architecture, see Figure 1, the embodiment of the present application relates to the improvement of the intelligent dialogue system, and the memory management module can be divided into a first memory module (also called a long-term memory module) and a second memory module (also called a short-term memory module), which are respectively used to store the memory summary content generated according to the historical dialogue content generated in the current session, and the original dialogue content that is newly generated in the current session and has not yet formed a memory summary in the first memory module. The first memory module can also update the memory summary content by asynchronous update. In this way, after receiving the dialogue content input by the user each time, the dialogue management module can request the memory management module to obtain the memory content. The memory management module does not need to wait for the update processing of the memory summary content, but can directly return the current memory summary content stored in the first memory module and the original dialogue content that is newly generated and has not yet formed a memory summary in the first memory module stored in the second memory module to the dialogue management module. In this way, the dialogue management module can assemble a prompt according to the above content and input it to the LLM model, and the LLM model generates the reply content. Among them, the second memory module can use a sliding window with a dynamic adaptive length to achieve the management of dialogue memory. Correspondingly, the generation and update timing of the memory summary content can also be controlled according to the length of the sliding window. That is, the second memory module can be used to trigger the memory summary content background generation module to generate the memory summary content and update the memory summary content to the first memory module.

[0046] The specific implementation scheme provided in the embodiments of the present application is described in detail below.

[0047] First, the present application embodiment provides an intelligent dialogue method, see Figure 2 , the method may include:

[0048] S201: In a conversation process with a user through a large language model LLM, receiving conversation content input by the user in the current conversation round.

[0049] First of all, it should be noted that a "conversation" can refer to a reception process of a user by an intelligent customer service or voice robot, which usually has a clear "beginning" and "end". For example, a conversation can generally start from the user initiating a conversation and end when the intelligent customer service closes the conversation or the system prompts the conversation to be closed. Usually, a conversation involves multiple rounds of conversations between the two parties, and a question and answer between the two parties can be called a conversation round. Therefore, in a conversation, multiple conversation rounds will be generated, and the user may enter the conversation content in multiple conversation rounds. Accordingly, the intelligent customer service needs to generate the corresponding reply content after each user enters the conversation content.

[0050] S202: Request the memory management module to obtain the memory content of the current conversation; wherein the memory management module includes a first memory module and a second memory module, the first memory module is used to save the memory summary content generated according to the historical conversation content generated in the current conversation, and update the memory summary content asynchronously with the process of generating the reply content; the second memory module is used to save the original text of the conversation content newly generated in the current conversation that has not yet formed a memory summary in the first memory module.

[0051] After receiving the conversation content input by the user in the current conversation round, the conversation management module can request the memory management module to obtain the memory content of the current conversation. As mentioned above, in the embodiment of the present application, the memory management module may include a first memory module and a second memory module, wherein the first memory module is used to save the memory summary content generated according to the historical conversation content generated in the current conversation, and update the memory summary content asynchronously with the process of generating the reply content; the second memory module is used to save the original text of the conversation content newly generated in the current conversation that has not yet formed a memory summary in the first memory module. Therefore, after receiving the request of the conversation management module, the memory management module can directly return the memory summary content currently saved in the first memory module and the original text of the conversation content saved in the second memory module to the conversation management module.

[0052] It should be noted here that, in the embodiment of the present application, "historical conversation content" may specifically refer to conversation content that has formed memory summary content in the first memory module, that is, if the content of a round of conversation has not been updated to the first memory module, it cannot be called historical conversation content for the time being, but is stored in the second memory module in the form of short-term memory. Of course, as the memory summary content in the first memory module is updated, the conversation content in the second memory module will gradually become historical conversation content, and the memory summary content in the first memory module can also reflect these conversation contents.

[0053] S203: Generate prompt information for guiding the LLM model to perform a content generation task based on the memory summary content currently stored in the first memory module and the original text of the conversation content stored in the second memory module, so that the LLM model can generate reply content for replying to the conversation content input by the user.

[0054] After receiving the memory content returned by the memory management module, the dialogue management module can generate a prompt for guiding the LLM model to perform the content generation task based on the memory summary content currently stored in the first memory module and the original dialogue content stored in the second memory module, so that the LLM model can generate a reply content for replying to the dialogue content input by the user. Among them, how the LLM model specifically generates the reply content based on the above input information is not the focus of the protection of the embodiments of this application, so it will not be described in detail here.

[0055] As described above, in the embodiment of the present application, the memory summary content in the first memory module can be updated asynchronously, that is, each time the user's conversation content is received and the memory management module is queried, the memory management module does not need to wait for the update of the memory summary content in the first memory module, but can directly return the current memory summary content in the first memory module. Among them, the memory summary content in the first memory module can be regenerated and updated after each round of conversation ends, or, in another way, since the second memory module can manage the original text of the conversation content in a sliding window manner, and the length of the sliding window can be dynamically changed with the increase in the number of conversation content rounds newly generated in the current session that have not yet formed a memory summary in the first memory module. Therefore, after the length of the sliding window reaches the first threshold length, the process of regenerating the memory summary content according to the conversation content within the first threshold length starting from the first round of conversation content in the second memory module can be triggered, wherein the process of regenerating the memory summary content is performed asynchronously with the main thread that generates the reply content through a separate thread. Among them, after triggering the process of regenerating the memory summary content, the memory summary content stored in the first memory module remains unchanged, and the length of the sliding window in the second memory module gradually increases with the newly generated rounds of dialogue content on the basis of the first threshold length, until the length of the sliding window in the second memory module reaches the second threshold length, the dialogue content within the first threshold length from the first round of dialogue content in the second memory module can be deleted, and the regenerated memory summary content can be updated to the first memory module.

[0056] When the memory summary content is regenerated, it can also be generated by the LLM model, wherein the data specifically outputting the LLM model may include the memory summary content currently stored in the first memory module, and the original text of the conversation content within the first threshold length from the first round of conversation content in the second memory module. In other words, the LLM model can generate new memory summary content based on the above content. When specifically generating the memory summary content, it can automatically extract key information of the conversation, etc., to generate a summary content that is shorter than the original text of the conversation content.

[0057] Among them, the first threshold length is less than the second threshold length, and the second memory module ensures the integrity of the memory by changing the sliding window length between the first threshold and the second threshold, and provides sufficient update time for the memory summary in the first memory module. For example, the first threshold length is 20 rounds, and the second threshold length is 40 rounds. When the sliding window in the second memory module reaches 20 rounds, the background program can be triggered to regenerate new memory summary content according to the conversation content of the first 20 rounds in the second memory module. After that, the length of the sliding window in the second memory module continues to grow. When it reaches 40 rounds, the conversation content of the first 20 rounds in the second memory module can be deleted, and the corresponding new memory summary content generated can be updated to the first memory module. At the same time, since the length of the sliding window in the second memory module becomes 20 again, the regeneration of the memory summary content of the conversation content of the 20 rounds in the second memory module can be triggered, and so on.

[0058] In specific implementation, the interaction process between multiple terminals may include:

[0059] 1. The user enters the conversation content and initiates a conversation request;

[0060] 2. The dialogue management module initiates a request to the memory management module to obtain the memory content;

[0061] 3. The memory management module sends the session ID of the current session to the memory database;

[0062] 4. The memory database returns the current memory summary content in the first memory module and the original conversation content saved in the second memory module, and can also return the length information of the current sliding window;

[0063] 5. The memory management module returns the current memory summary content in the first memory module and the original conversation content stored in the second memory module to the conversation management module;

[0064] 6. The dialogue management module assembles prompts based on the received memory content and generates reply content through the LLM model;

[0065] 7. Return the reply content to the user.

[0066] 8. When the memory management module finds that the sliding window reaches a first threshold length, it can trigger the generation of memory summary content. When the sliding window reaches a second threshold length, it can trigger the update of the memory summary content in the first memory module.

[0067] In order to better understand the specific implementation scheme provided by the embodiment of the present application, the processing process of the embodiment of the present application is described in detail below through a specific example.

[0068] like Figure 3 As shown in the figure, the long rectangle composed of multiple small rectangles represents all memories of the current conversation, where each small rectangle represents a round of conversation. The dark gray small rectangle represents long-term memory, and the memory summary content is generated and saved by the first memory module. The first memory module can exist in the form of a plug-in knowledge base. The light gray small rectangle represents short-term memory, which is specifically expressed by directly inputting the original text of the conversation content as the model input. Among them, Figure 3 The first row shown shows that rounds 0-59 are long-term memory, and rounds 60-79 are short-term memory. The combination of long and short-term memory is used as the input of the model to ensure the integrity of all memories.

[0069] When the short-term memory reaches the first threshold length of the sliding window, the memory management module can be triggered to perform a memory summary on the short-term memory of the first threshold length in the second memory module. The specific method is: current long-term memory + the original text of the conversation content from the first position to the first threshold length in the current second memory module = updated memory summary content; the update mechanism of the memory summary can be performed asynchronously in the background without affecting the response of the current dialogue system.

[0070] After the asynchronous update mechanism of the memory summary content is triggered, the memory summary content will be regenerated first. Since the generation process takes some time, after the above update mechanism is triggered, the long-term memory will remain unchanged temporarily, but the sliding window of the short-term memory will increase in size round by round.

[0071] For example, Figure 3 As shown in the first row of , at a certain moment, the long-term memory stores the memory summary content of the conversation rounds 0 to 59, and the sliding window of the short-term memory is the 60th to 79th rounds, with a length of 20 rounds. Assuming that the first threshold length mentioned above is 20 rounds, this state of the short-term memory will trigger the update mechanism of the memory summary content. However, since the memory summary content is being regenerated at this time, specifically, the new memory summary content is regenerated based on the previously generated memory summary content and the conversation content of rounds 60 to 79. At this time, the long-term memory remains unchanged temporarily. For example, after the memory summary update mechanism is triggered in the state of the first row, as shown in Figure 3 As shown in the second row of , a new round of dialogue content is generated, that is, the 80th round of dialogue content is generated. At this time, the long-term memory still stores the memory summary content of the 0th to 59th round of dialogue content, but the sliding window of the short-term memory becomes the 60th to 80th round, with a length of 21 rounds; when the 81st round of dialogue content is generated, as shown in Figure 3As shown in the third row of , the content stored in the long-term memory remains unchanged, which is still the memory summary of the conversation content from round 0 to 59. However, the sliding window of the short-term memory becomes the 60th to 81st round, with a length of 22 rounds. And so on, until the 99th round of conversation content is newly generated, as shown in Figure 3 As shown in the fourth row in the figure, the content stored in the long-term memory remains unchanged, which is the memory summary content of the conversation content from rounds 0 to 59. The sliding window of the short-term memory changes to rounds 60-99, with a length of 40 rounds. Assuming that the second threshold length is 40 rounds, the update of the memory summary content in the long-term memory module will be triggered under the state shown in the fourth row above, that is, the memory summary content corresponding to rounds 60-79 will be updated to the long-term memory module. At this time, Figure 3 As shown in the fifth row of , the short-term memory module stores the original text of the conversation from round 80 to round 99. At this time, since the sliding window length of the short-term memory becomes 20 rounds, the generation of memory summaries for rounds 80 to 99 can be triggered again, and so on.

[0072] In summary, through the embodiment of the present application, the memory management module can be divided into a first memory module (which can be called long-term memory) and a second memory module (which can be called short-term memory), wherein the first memory module can be used to save the memory summary content generated according to the historical conversation content generated in the current session, and update the memory summary content asynchronously with the process of generating the reply content; the second memory module is used to save the original text of the conversation content newly generated in the current session that has not yet formed a memory summary in the first memory module. In this way, when it is necessary to generate the reply content according to the conversation content input by the user in the current conversation round, the prompt information for guiding the LLM model to perform the content generation task can be generated according to the memory summary content currently saved in the first memory module and the original text of the conversation content saved in the second memory module, so that the LLM model can generate the reply content for replying to the conversation content input by the user. In this way, since the memory summary content is generated for the historical conversation record, the length of the input data can be shortened, and the requirements of the LLM model for the length of the input data can be met. At the same time, there is no need to wait for the update of the memory summary content, so that the reply delay can be shortened. As for the conversation content that has not yet been updated in the first memory module, it can be supplemented by the original text of the conversation content stored in the second memory module, so as to ensure the integrity of the memory and ensure the accuracy of the generated reply content.

[0073] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data can be used in the scheme described herein within the scope permitted by applicable laws and regulations, subject to the requirements of applicable laws and regulations of the country where the user is located (for example, with the user's explicit consent, effective notification to the user, etc.).

[0074] Corresponding to the aforementioned method embodiment, the embodiment of the present application further provides an intelligent dialogue device, which may include:

[0075] A dialogue request receiving unit, used for receiving the dialogue content input by the user in the current dialogue round in the conversation process of conducting a dialogue with the user through the large language model LLM;

[0076] A memory content acquisition unit, used for requesting the memory management module to acquire the memory content of the current session; wherein the memory management module includes a first memory module and a second memory module, wherein the first memory module is used for storing the memory summary content generated according to the historical conversation content generated in the current session, and performing asynchronous memory summary content update with the process of generating the reply content; and the second memory module is used for storing the original text of the conversation content newly generated in the current session and not yet forming a memory summary in the first memory module;

[0077] A reply content generation unit is used to generate prompt information for guiding the LLM model to perform a content generation task based on the memory summary content currently stored in the first memory module and the original text of the conversation content stored in the second memory module, so that the LLM model can generate reply content for replying to the conversation content input by the user.

[0078] In a specific implementation, the second memory module manages the original text of the conversation content by means of a sliding window, and the length of the sliding window changes dynamically as the number of conversation content rounds newly generated in the current session that have not yet formed a memory summary in the first memory module increases; so that the second memory module can save the original text of the conversation content within the current sliding window.

[0079] Wherein, after the length of the sliding window reaches a first threshold length, a process of regenerating memory summary content according to the conversation content within the first threshold length starting from the first round of conversation content in the second memory module may be triggered, wherein the process of regenerating memory summary content is performed asynchronously with the process of generating reply content through a separate thread;

[0080] After the length of the sliding window reaches a second threshold length, the conversation content within the first threshold length starting from the first round of conversation content in the second memory module is deleted, and the regenerated memory summary content is updated to the first memory module.

[0081] Specifically, after triggering the process of regenerating the memory summary content, the memory summary content stored in the first memory module remains unchanged, and the length of the sliding window in the second memory module gradually increases based on the first threshold length as the new rounds of dialogue content are generated, until the length of the sliding window in the second memory module reaches the second threshold length, and then the regenerated memory summary content is updated to the first memory module.

[0082] Among them, when the memory summary content is regenerated, it is generated by the LLM model based on the memory summary content currently stored in the first memory module and the original text of the conversation content within the first threshold length starting from the first round of conversation content in the second memory module.

[0083] The second threshold length is greater than the first threshold length.

[0084] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.

[0085] And an electronic device, comprising:

[0086] one or more processors; and

[0087] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.

[0088] A computer program product includes a computer program / computer executable instructions, which, when executed by a processor in an electronic device, implement the steps of the method described in the aforementioned method embodiment.

[0089] in, Figure 4 The architecture of the electronic device is shown as an example, which may include a processor 410, a video display adapter 411, a disk drive 412, an input / output interface 413, a network interface 414, and a memory 420. The processor 410, the video display adapter 411, the disk drive 412, the input / output interface 413, the network interface 414, and the memory 420 may be communicatively connected via a communication bus 430.

[0090] Among them, the processor 410 can be implemented by a general-purpose CPU (Central Processing Unit, processor), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solution provided in this application.

[0091] The memory 420 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 420 can store an operating system 421 for controlling the operation of the electronic device 400, and a basic input and output system (BIOS) for controlling the low-level operation of the electronic device 400. In addition, a web browser 423, a data storage management system 424, and an intelligent dialogue processing system 425, etc. can also be stored. The above-mentioned intelligent dialogue processing system 425 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 420 and is called and executed by the processor 410.

[0092] The input / output interface 413 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0093] The network interface 414 is used to connect to a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0094] The bus 430 comprises a pathway for transmitting information between the various components of the device (eg, the processor 410, the video display adapter 411, the disk drive 412, the input / output interface 413, the network interface 414, and the memory 420).

[0095] It should be noted that, although the above device only shows a processor 410, a video display adapter 411, a disk drive 412, an input / output interface 413, a network interface 414, a memory 420, a bus 430, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include components necessary for implementing the solution of the present application, and does not necessarily include all the components shown in the figure.

[0096] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.

[0097] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.

[0098] The above is a detailed introduction to the intelligent dialogue method and electronic device provided by the present application. The principles and implementation methods of the present application are described in detail using specific examples. The description of the above embodiments is only used to help understand the method and core ideas of the present application. At the same time, for those skilled in the art, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present application.

Claims

1. An intelligent dialogue method, characterized in that: include: In a conversation process with a user through a large language model LLM, receiving the conversation content input by the user in the current conversation round; Requesting the memory management module to obtain the memory content of the current session; wherein the memory management module includes a first memory module and a second memory module, the first memory module is used to store the memory summary content generated according to the historical conversation content generated in the current session, and update the memory summary content asynchronously with the process of generating the reply content; the second memory module is used to store the original text of the conversation content newly generated in the current session and not yet formed into a memory summary in the first memory module; Based on the memory summary content currently stored in the first memory module and the original text of the conversation content stored in the second memory module, prompt information is generated to guide the LLM model to perform a content generation task, so that the LLM model can generate reply content for replying to the conversation content input by the user.

2. The method according to claim 1, characterized in that: The second memory module manages the original text of the conversation content by means of a sliding window, and the length of the sliding window changes dynamically as the number of conversation content rounds newly generated in the current session that have not yet formed a memory summary in the first memory module increases; so that the second memory module can save the original text of the conversation content within the current sliding window.

3. The method according to claim 2, characterized in that After the length of the sliding window reaches a first threshold length, a process of regenerating memory summary content according to the conversation content within the first threshold length starting from the first round of conversation content in the second memory module is triggered, wherein the process of regenerating memory summary content is performed asynchronously with the process of generating reply content through a separate thread; After the length of the sliding window reaches a second threshold length, the conversation content within the first threshold length starting from the first round of conversation content in the second memory module is deleted, and the regenerated memory summary content is updated to the first memory module.

4. The method according to claim 3, characterized in that After triggering the process of regenerating the memory summary content, the memory summary content stored in the first memory module remains unchanged, and the length of the sliding window in the second memory module gradually increases based on the first threshold length as the new round of dialogue content is generated, until the length of the sliding window in the second memory module reaches the second threshold length, and then the regenerated memory summary content is updated to the first memory module.

5. The method according to claim 3, characterized in that: When the memory summary content is regenerated, it is generated by the LLM model according to the memory summary content currently stored in the first memory module and the original text of the conversation content within the first threshold length starting from the first round of conversation content in the second memory module.

6. The method according to claim 3, characterized in that The second threshold length is greater than the first threshold length.

7. An intelligent dialogue device, characterized in that: include: A dialogue request receiving unit, used for receiving the dialogue content input by the user in the current dialogue round in the conversation process of conducting a dialogue with the user through the large language model LLM; A memory content acquisition unit, used for requesting the memory management module to acquire the memory content of the current session; wherein the memory management module includes a first memory module and a second memory module, wherein the first memory module is used for storing the memory summary content generated according to the historical conversation content generated in the current session, and performing asynchronous memory summary content update with the process of generating the reply content; and the second memory module is used for storing the original text of the conversation content newly generated in the current session and not yet forming a memory summary in the first memory module; A reply content generation unit is used to generate prompt information for guiding the LLM model to perform a content generation task based on the memory summary content currently stored in the first memory module and the original text of the conversation content stored in the second memory module, so that the LLM model can generate reply content for replying to the conversation content input by the user.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.

9. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of claims 1 to 6.

10. A computer program product comprising a computer program / computer executable instructions, characterized in that: When the computer program / computer executable instructions are executed by a processor in an electronic device, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Question and answer management method and device based on large model, storage medium and program product

    CN121233825A

  • Question answering management method and device based on large model, storage medium and program product

    CN121233825B

  • Software file generation method and device, storage medium and electronic device

    CN121858083A

  • Software file generation method and device, storage medium and electronic device

    CN121858083B