A method, device, equipment and medium for processing large language model dialogue memory

By introducing long-term memory boxes and memory management mechanisms into the large language model, the problem that large language model is difficult to achieve cross-long cycle memory in scenarios such as emotional companionship and psychological counseling is solved, and long-term dialogue memory and personalized services of the agent are realized, enhancing the accuracy of anthropomorphism and memory.

CN119003728BActive Publication Date: 2025-06-20BEIJING FACE WALL INTELLIGENT TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411097797.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-12
Publication Date
2025-06-20
Estimated Expiration
2044-08-12

AI Technical Summary

Technical Problem

The large language model is difficult to achieve long-term memory and personalized services in scenarios such as emotional companionship and psychological counseling, resulting in the lack of anthropomorphism and the integrity of dialogue historical memory difficult to guarantee.

Method used

By segmenting and packaging the conversation information between the agent and the user into session information, and combining environmental information from other systems, long-term memories are created or updated in the long-term memory box, thereby building a large language model to generate requested long-term memory prompts.

Benefits of technology

It realizes long-term dialogue memory and personalized services of the agent, enhances the sense of anthropomorphism, ensures the timeliness and accuracy of long-term memory, and avoids inaccurate or incomplete responses caused by information lag.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119003728B_ABST
    Figure CN119003728B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device and medium for processing dialogue memory of a large language model. The method includes: segmenting and packing the dialogue information between the agent and the user to obtain session information; the session information includes a plurality of dialogue information belonging to the same round of dialogue; creating or updating long-term memory in a long-term memory box according to the session information and environmental information from other systems; obtaining target long-term memory from the long-term memory box, and constructing a long-term memory prompt for the large language model generation request according to the target long-term memory. Embodiments of the present invention can improve the long-term dialogue ability of the agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, apparatus, device and medium for processing dialogue memory of large language models. Background Art

[0002] Large language models (LLMs) for chatting are essentially a stateless service mechanism. Unless state information is carried in the requests of the large language model, the large language model has no memory and perception of the state. To make up for this deficiency, generally, the historical dialogue messages of the current session are placed in the context of the request to obtain the multi-round dialogue ability. However, due to the length limit of the request context of the large language model, it is difficult to obtain long-term memory across sessions using this solution. Especially in scenarios such as emotional companionship and psychological counseling, it is difficult for an agent built based on the large language model to show the feeling of growing together with the user in the interaction over a long period.

[0003] Although some research has explored the ultra-long context of large language models and achieved certain results, for scenarios such as emotional companionship, the introduction of ultra-long context still cannot meet the need for memory. Typical problems include: the lack of memory update and forgetting mechanisms, resulting in the lack of anthropomorphic sense; the randomness in the ChitChat scenario, where the mutually related memories are often scattered and interspersed in the dialogue history, making it difficult to ensure the integrity of the retrieval results of long-term memory; the retrieval mechanism introduces too much noise (such as dialogue information about the same event at different times), posing a great challenge to the accuracy of generation. Summary of the Invention

[0004] The present invention provides a method, apparatus, device and medium for processing dialogue memory of large language models to improve the long-term dialogue ability of the agent.

[0005] According to one aspect of the present invention, there is provided a method for processing dialogue memory of large language models, including:

[0006] Segmenting and packaging the dialogue information between the agent and the user to obtain session information; the session information includes multiple dialogue information belonging to the same round of dialogue;

[0007] Creating or updating long-term memory in the long-term memory box according to the session information and environmental information from other systems;

[0008] Obtaining target long-term memory from the long-term memory box, and constructing a long-term memory prompt for the large language model generation request according to the target long-term memory.

[0009] According to another aspect of the present invention, there is provided a processing device for large language model dialogue memory, including:

[0010] A dialogue segmentation module for segmenting and packaging the dialogue information between the agent and the user to obtain session information; the session information includes multiple pieces of dialogue information belonging to the same round of dialogue;

[0011] A memory management module for creating or updating long-term memory in a long-term memory box according to the session information and environmental information from other systems;

[0012] A memory usage module for obtaining target long-term memory from the long-term memory box and constructing a long-term memory prompt for the large language model generation request according to the target long-term memory.

[0013] According to another aspect of the present invention, there is provided a computer program product including a computer program, which when executed by a processor implements the processing method for large language model dialogue memory according to any embodiment of the present invention.

[0014] According to another aspect of the present invention, there is provided an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the processing method for large language model dialogue memory according to any embodiment of the present invention.

[0015] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the processing method for large language model dialogue memory according to any embodiment of the present invention when executed.

[0016] The embodiments of the present invention establish an anthropomorphic long-term memory storage, update and recall mechanism, which is applied to scenarios such as anthropomorphic companionship, enabling the agent to provide a more human-like memory mechanism through the simulation of memory, realizing the personalization of user services; at the same time, capturing the latest dialogue information and updating the long-term memory in real time to ensure the timeliness and accuracy of the long-term memory, and avoiding inaccurate or incomplete agent responses caused by information lag.

[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0019] Figure 1 is a flowchart of a method for processing the dialogue memory of a large language model according to an embodiment of the present invention;

[0020] Figure 2A is a flowchart of a method for processing the dialogue memory of a large language model according to another embodiment of the present invention;

[0021] Figure 2B is a schematic diagram of the interaction process of each component according to another embodiment of the present invention;

[0022] Figure 3 is a schematic structural diagram of a method for processing the dialogue memory of a large language model according to another embodiment of the present invention;

[0023] Figure 4 is a schematic structural diagram of an electronic device for implementing the embodiments of the present invention. Detailed implementation manners

[0024] In order to enable those skilled in the art of the present technology to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0026] Figure 1The flowchart of a method for processing dialogue memory of a large language model provided by an embodiment of the present invention. This embodiment is applicable to the situation of designing an overall framework covering long-term and short-term memory, memory update mechanism, and memory invocation mechanism, designing the processes and interfaces for the interaction between the long-term memory box and the intelligent agent and other systems, and improving the dialogue ability of the intelligent agent accordingly. This method can be executed by a processing device for dialogue memory of a large language model. This device can be implemented in the form of hardware and / or software, and can be configured in an electronic device with corresponding data processing capabilities. As Figure 1 shown, the method includes:

[0027] S110. Segment and package the dialogue information between the intelligent agent and the user to obtain session information;

[0028] S120. Create or update long-term memory in the long-term memory box according to the session information and the environmental information from other systems.

[0029] S130. Obtain the target long-term memory from the long-term memory box, and construct a long-term memory prompt for the large language model generation request according to the target long-term memory.

[0030] Among them, the session message contains multiple chat messages belonging to the same round of dialogue. Other systems are the systems involved in the dialogue content of the dialogue message, and the environment message is the information referred to when generating the dialogue content. For example, if the dialogue content is movie recommendation, other systems are third-party movie review websites, and the environment message is the movie review information, actor information, and director information of the recommended movie. The long-term memory box is a specially established memory processing unit, which is used to dynamically update the memory according to various collected information to ensure the timeliness and accuracy of the memory, and can reflect the latest communication situation.

[0031] Specifically, a large amount of dialogue information will be generated during the process of the user's dialogue with the (general chat) intelligent agent. These dialogue information will be collected in real time and segmented and packaged into one or more session information according to the session to which they belong after the collection is completed. The session information and the environmental information from other systems will be used by the long-term memory box to update its existing long-term memory or create new long-term memory (Memory Reflection).

[0032] When a user has a conversation with an agent, the agent sends a memory retrieval request (MemoryRecall) to the long-term memory box. The long-term memory box retrieves the long-term memory stored in itself according to the retrieval request and returns the retrieval result as the target long-term memory to the agent. The agent splices the target long-term memory as a long-term memory augmented prompt into the context of the large language model generation request, so as to provide more personalized and history-communication-compliant services for the user.

[0033] The embodiment of the present invention establishes an anthropomorphic long-term memory storage, update, and recall mechanism, which is applied to scenarios such as anthropomorphic companionship. It enables the agent to provide a more human-like memory mechanism through the simulation of memory, realizing the personalization of user services. At the same time, it captures the latest conversation information and updates the long-term memory in real time to ensure the timeliness and accuracy of the long-term memory, and avoid inaccurate or incomplete responses of the agent caused by information lag.

[0034] Based on the above embodiment, optionally, the method further includes: regularly organizing the long-term memory in the long-term memory box in terms of the agent.

[0035] Specifically, in order to further optimize the organization and management of memory, set a regular task in terms of the agent to regularly organize the long-term memory in the long-term memory box, remove redundant and invalid information, and improve the quality and usability of the memory.

[0036] Based on the above embodiment, optionally, the long-term memory includes important information memory between the user and the agent, and setting information memory of the agent.

[0037] Specifically, the content that long-term memory can contain is diverse. Generally, it includes: personal portrait memory of the user (name, address, age, occupation, hobbies, social and family relationships, personality traits, emotional state, etc.); important information memory between the user and the agent (such as the relationship and intimacy with the agent, major events); setting information memory of the agent (such as the identity setting and plot line setting of the agent).

[0038] Figure 2A The figure is a flowchart of a method for processing large language model dialogue memory provided by another embodiment of the present invention. This embodiment is optimized and improved based on the above embodiment. As Figure 2A shown, the method includes:

[0039] S210: Real-time obtain and store the dialogue information between the agent and the user through the dialogue center; publish the dialogue information to the session message broker through the dialogue center; and segment and package the dialogue information through the session message broker to obtain session information.

[0040] Specifically, as Figure 2B shown, the conversation information between the user and the agent is updated in real time to the Conversation Center to ensure that the latest communication information can be captured in a timely manner. The Conversation Center is responsible for processing and storing the conversation information in real time. In order for the long-term memory box to obtain the conversation messages in a timely manner to create or update the long-term memory, when the Conversation Center receives the conversation messages, it publishes the conversation messages to the Message Block Broker through the message queue (real-time message queue and / or delayed message queue). The Message Block Broker splits and packages the received conversation message stream to obtain the session information, which helps to effectively manage and process a large amount of conversation information.

[0041] S220: Publish the session information to the Observation Message Queue through the Message Block Broker; publish the environmental information from other systems to the Observation Message Queue; create or update the long-term memory in the long-term memory box through the session information and environmental information in the Observation Message Queue.

[0042] Specifically, continue to refer to Figure 2B , the session information generated by the Message Block Broker and the environmental information generated by other systems will be published to the Observation Message Queue as information related to the agent's memory, realizing the concentration and integration of information. The long-term memory box reads the messages from the Observation Message Queue and uses them to create or update the memory, so that the memory in the box can be continuously improved and accurate. Creating a unified message queue to collect and summarize various types of relevant information can achieve centralized management of information, avoid information dispersion and omission, and provide a basis for creating a comprehensive and accurate memory; at the same time, various specialized processing modules manage and optimize a large number of message streams in an orderly manner, which can improve the efficiency and quality of message processing, making the subsequent memory creation and update more organized.

[0043] Optionally, the number of conversation information included in the session information is not greater than the set number, and the interaction time between any two adjacent conversation information is not greater than the set time.

[0044] Specifically, in data evaluation and training, annotations and training are usually carried out in units of sessions. So far, there is no concept of sessions in the intelligent agent products for individual users. When performing online inference, directly attaching the past preset number (N) of conversations as the context of the large model request will bring some problems. Therefore, the present invention provides the following method for splitting conversations:

[0045] a. After the user resets or deletes the conversation, it is regarded as clearing all sessions and memories, and a new conversation channel and session are generated.

[0046] b. When there is no conversation interaction between the user and the agent for more than the set time, a new session is generated, and the conversation information between the end of the previous session and the most recent conversation history is packaged as a session information.

[0047] c. If the user continuously has conversations with the agent and the number of conversation messages exceeds the set number, but there is no time interval without interaction, then these set number of conversation messages are packaged as a session information. However, when the large language model is performing inference, we still attach the previous set number of conversation messages of this session as context information. (This will cause the problem that the same information exists in both the long-term memory and the conversation history, but this problem can be solved through model training).

[0048] In addition, for other cooperating roles: After splitting the session, increase the user's perception in a relatively natural way, such as adding an active prompt from the agent like "It's been a while since we last talked. Let's talk about something new." The "backtracking" function is used for re-consideration, but do not backtrack across sessions.

[0049] S230. Obtain the target long-term memory from the long-term memory box, and construct a long-term memory prompt for the large language model generation request according to the target long-term memory.

[0050] S240. Combine the long-term memory prompt and other prompts to generate the context for the large language model generation request; input the context and the large language model generation request into the large language model in the agent to obtain the output of the large language model and new conversation information.

[0051] Specifically, continue to refer to Figure 2B , the long-term memory prompt and other prompts, such as System Prompt, Retrieval-Augmented Generation Prompt, and In-session ChatHistory, are concatenated into the context of the large language model generation request, and the context and the large language model generation request are input into the large language model in the agent, allowing the large language model to generate content following the memory and output. Record the large language model generation request and the corresponding output, and send them as new conversation information to the conversation center.

[0052] In addition, essentially, long-term memory is a special type of "knowledge", and prompts similar to retrieval-augmented generation can be used to prompt the large language model to generate responses. Currently, retrieval-augmented generation has not been specifically optimized for long-term memory. We can first use instructions similar to long-term memory to evaluate the generation effect. For model training, it is necessary to fine-tune the model with the <Instructions, RelatedMemories, GenrationResult> triple.

[0053] Exemplarily, an example based on the conversation content is as follows:

[0054] User: Hello.

[0055] Agent: Hello, how can I help you?

[0056] User: Do you have any movie recommendations?

[0057] Agent: Have you watched the <Movie A> recommended to you last time?

[0058] User: Yes, I have.

[0059] Agent: This time I think you should watch <Movie B>.

[0060] User: Okay.

[0061] In this example, the agent retrieves the long-term memory from the long-term memory box and obtains that the user has just watched Movie A and the user's long-term movie-watching preferences. Based on these two, the agent recommends Movie B to the user. By establishing an anthropomorphic long-term memory mechanism, the present invention provides a more considerate, coherent and personalized communication experience for users, meets the needs of users for in-depth interaction and companion growth in these scenarios, and thus attracts more users to choose the company's services. Secondly, the anthropomorphic long-term memory mechanism helps to expand the application scenarios of the company's products. It enables the large language model to be not only applicable to short-term and general communication, but also to perform excellently in complex scenarios that require long-term memory and in-depth understanding, such as educational tutoring, medical consultation, etc., opening up new markets and customer groups for the company. Furthermore, the anthropomorphic long-term memory mechanism improves users' satisfaction and loyalty to the company's products. Better memory processing capabilities can allow users to feel a more intelligent and user-friendly service, enhance the emotional connection between users and the product, reduce user churn, and promote long-term use and word-of-mouth dissemination of the product.

[0062] The embodiment of the present invention creates a unified message queue to collect and summarize various types of relevant information, which can achieve centralized management of information, avoid information dispersion and omission, and provide a basis for comprehensive and accurate memory creation; at the same time, various special processing modules manage and optimize a large number of message flows in an orderly manner, which can improve the efficiency and quality of message processing, making subsequent memory creation and update more organized.

[0063] Figure 3 It is a schematic structural diagram of a processing device for large language model dialogue memory provided by another embodiment of the present invention. As Figure 3 shown, the device includes:

[0064] The dialogue segmentation module 310 is used to segment and package the dialogue information between the agent and the user to obtain session information; the session information includes multiple dialogue information belonging to the same round of dialogue.

[0065] The memory management module 320 is used to create or update long-term memory in the long-term memory box according to the session information and environmental information from other systems.

[0066] The memory usage module 330 is used to obtain the target long-term memory from the long-term memory box and construct a long-term memory prompt for the large language model generation request according to the target long-term memory.

[0067] The processing device for large language model dialogue memory provided by the embodiments of the present invention can execute the processing method for large language model dialogue memory provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0068] Optionally, the dialogue segmentation module 310 includes:

[0069] The dialogue storage unit is used to obtain and store the dialogue information between the agent and the user in real time through the dialogue center.

[0070] The dialogue publishing unit is used to publish the dialogue information to the session message broker through the dialogue center.

[0071] The dialogue segmentation unit is used to segment and package the dialogue information through the session message broker to obtain session information.

[0072] Optionally, the memory management module 320 includes:

[0073] The session publishing unit is used to publish the session information to the observation message queue through the session message broker.

[0074] The information publishing unit is used to publish the environmental information from other systems to the observation message queue.

[0075] The memory management unit is used to create or update long-term memory in the long-term memory box through the session information and environmental information in the observation message queue.

[0076] Optionally, the number of dialogue information included in the session information is not greater than the set number, and the interaction time between any two adjacent dialogue information is not greater than the set time.

[0077] Optionally, the long-term memory includes the important information memory between the user and the agent and the set information memory of the agent.

[0078] Optionally, the device further includes:

[0079] A memory organizing module, which is used to regularly organize the long-term memory in the long-term memory box by agent granularity.

[0080] Optionally, the device further includes:

[0081] A context generation module, which is used to combine the long-term memory prompt and other prompts to generate the context of the large language model generation request;

[0082] A context usage module, which is used to input the context and the large language model generation request into the large language model in the agent to obtain the output of the large language model and new conversation information.

[0083] Furthermore, the processing device for large language model dialogue memory described above can also execute the method for processing large language model dialogue memory provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0084] Figure 4 FIG. shows a schematic structural diagram of an electronic device 40 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0085] As Figure 4 shown, the electronic device 40 includes at least one processor 41, and a memory communicatively connected to at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc. Among them, the memory stores a computer program executable by at least one processor, and the processor 41 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other through a bus 44. The input / output (I / O) interface 45 is also connected to the bus 44.

[0086] Multiple components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a disk, an optical disc, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0087] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the processing method of the large language model dialogue memory.

[0088] In some embodiments, the processing method of the large language model dialogue memory can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the processing method of the large language model dialogue memory described above can be executed. Alternatively, in other embodiments, the processor 41 can be configured to execute the processing method of the large language model dialogue memory by any other suitable means (e.g., by means of firmware).

[0089] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0090] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0091] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0092] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0093] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0094] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0095] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0096] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for processing dialogue memory of a large language model, characterized in that: The method comprises: The conversation information between the agent and the user is divided and packaged to obtain conversation information; the conversation information includes multiple conversation information belonging to the same round of conversation; Creating long-term memory or updating long-term memory in the long-term memory box according to the conversation information and environmental information from other systems; the other systems are systems involved in the conversation content in the conversation information; the environmental information is information referenced when generating the conversation content; Obtaining a target long-term memory from the long-term memory box, and constructing a large language model based on the target long-term memory to generate a requested long-term memory prompt; The process of splitting and packaging is as follows: If the user resets or deletes a conversation, all conversations and memories will be cleared, and a new conversation channel and conversation will be generated; If the user and the agent have no dialogue interaction for more than the set time, a new session is generated, and the dialogue information between the end of the last session and the most recent dialogue history is packaged into one session information; If the user continues to talk to the agent and the number of dialogue messages exceeds the set number and there is no time interval without interaction, the set number of dialogue messages are packaged into one session message. When the large language model performs inference later, the previous set number of dialogue messages of this session are attached as context information; Wherein, creating long-term memory or updating long-term memory in the long-term memory box according to the session information and the environment information from other systems comprises: Publish session information to the observation message queue through the session message broker; Publish environmental information from other systems to the observation message queue; Long-term memory is created or updated in the long-term memory box by observing the session information and the environment information in the message queue.

2. The method according to claim 1, characterized in that The step of dividing and packaging the dialogue information between the agent and the user to obtain the session information includes: Acquire and store the conversation information between the agent and the user in real time through the conversation center; Publishing the conversation information to the conversation message agent through the conversation center; The conversation information is segmented and packaged by the conversation message agent to obtain the conversation information.

3. The method according to any one of claims 1 to 2, characterized in that: The number of dialogue messages included in the session information is not greater than a set number and the interaction time between any two adjacent dialogue messages is not greater than a set time.

4. The method according to claim 1, characterized in that: The long-term memory includes important information memory between the user and the agent, and setting information memory of the agent.

5. The method according to claim 1, characterized in that The method further comprises: The long-term memory in the long-term memory box is regularly sorted out based on the granularity of the agent.

6. The method according to claim 1, characterized in that Target long-term memory builds a large language model to generate the requested long-term memory prompt, which also includes: Combining the long-term memory prompt with other prompts to generate a context for a large language model to generate a request; The context and the large language model generation request are input into the large language model in the agent to obtain the output of the large language model and new dialogue information.

7. A processing device for large language model dialogue memory, characterized in that: The device comprises: A dialogue segmentation module is used to segment and package the dialogue information between the agent and the user to obtain session information; the session information includes multiple dialogue information belonging to the same round of dialogue; A memory management module, used to create long-term memory or update long-term memory in the long-term memory box according to the conversation information and environmental information from other systems; the other systems are systems involved in the conversation content in the conversation information; the environmental information is information referenced when generating the conversation content; A memory using module, used to obtain a target long-term memory from the long-term memory box, and to construct a large language model based on the target long-term memory to generate a requested long-term memory prompt; The process of splitting and packaging is as follows: If the user resets or deletes a conversation, all conversations and memories will be cleared, and a new conversation channel and conversation will be generated; If the user and the agent have no dialogue interaction for more than the set time, a new session is generated, and the dialogue information between the end of the last session and the most recent dialogue history is packaged into one session information; If the user continues to talk to the agent and the number of dialogue messages exceeds the set number and there is no time interval without interaction, the set number of dialogue messages are packaged into one session message. When the large language model performs inference later, the previous set number of dialogue messages of this session are attached as context information; Wherein, the memory management module includes: A session publishing unit, used for publishing session information to an observation message queue through a session message agent; An information publishing unit, used to publish environmental information from other systems to the observation message queue; The memory management unit is used to create or update long-term memory in the long-term memory box by observing the session information and the environment information in the message queue.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for processing large language model dialogue memory according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method for processing large language model dialogue memory according to any one of claims 1 to 6 when executed.

Citation Information

Patent Citations

  • Service quality management system based on community system

    CN110414999A

  • Data processing method for chat and related device

    CN114003702A

  • Human-computer interaction method and device, electronic equipment and storage medium

    CN118098217A