Long-term memory-based agent interaction method, medium, and device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN SCHRÖDINGER TECHNOLOGY CO LTD
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]然而,现有基于向量数据库或关系型数据库的长时记忆构建技术,在应用于数字人、智能玩偶等需要塑造个性化虚拟人格的智能语音交互产品时,存在适配性不足,对个性化的虚拟人格构建的优化缺失的技术缺陷:向量数据库的核心在于文档拆分与向量索引构建,关系型数据库的核心在于信息间的关联建立,二者在问答机器人、客服机器人等通用交互场景中具备适配性,但针对数字人虚拟人格构建这一特定场景,处理方式过于粗暴,未围绕人格塑造进行针对性优化,无法支撑数字人形成独特、连贯的个性化交互特征
[0018]本申请中的基于长时记忆的智能体交互方法、介质和设备,通过同时调取固定化的画像信息、动态化的状态信息、浓缩性的摘要数据和细节性的陈述数据。画像信息决定了响应的核心方向贴合用户长期固定偏好,状态信息让响应适配用户当下的场景与状态,摘要数据把握了用户历史交互的核心规律,陈述数据则为响应提供了具体的历史细节支撑。多类型数据的融合避免了单一数据维度的片面性,使得智能体能够精准捕捉用户需求的核心,生成贴合用户个人特征的个性化响应,而非通用化、无差别的反馈,解决了现有技术中数据结构单一、缺乏人格化设计的问题。
Smart Images

Figure CN122527273A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to a method, medium and device for intelligent agent interaction based on long-term memory. Background Technology
[0002] In the field of intelligent voice interaction products with personalized personalities, such as digital humans and smart dolls, building long-term memory for the AI agent is a core technological step in realizing personalized interaction and shaping unique virtual personalities. Existing technologies, such as those disclosed in patents CN120688545A and CN120688545A, generally rely on vector databases and / or relational databases to build a knowledge base. This knowledge base provides the AI agent with an interactive context, thereby enabling long-term memory-related functions. The technology link based on vector databases is referred to as retrieval enhancement.
[0003] However, existing long-term memory construction technologies based on vector databases or relational databases have technical shortcomings when applied to intelligent voice interaction products such as digital humans and smart dolls that require the creation of personalized virtual personalities. These shortcomings include insufficient adaptability and a lack of optimization for personalized virtual personality construction. The core of vector databases lies in document splitting and vector index construction, while the core of relational databases lies in establishing connections between information. Both are adaptable in general interaction scenarios such as question-answering robots and customer service robots, but for the specific scenario of digital human virtual personality construction, the processing methods are too crude and have not been optimized around personality shaping, thus failing to support the formation of unique and coherent personalized interaction characteristics of digital humans.
[0004] Therefore, there is an urgent need for a novel digital human interaction method based on long-term memory to solve the above-mentioned technical problems. Summary of the Invention
[0005] The purpose of this application is to provide an intelligent agent interaction method, medium, and device based on long-term memory to solve at least one of the above-mentioned technical problems.
[0006] In a first aspect, this application provides an intelligent agent interaction method based on long-term memory, the method comprising: A query command is triggered based on the user's interaction information, and the profile information and status information of the intelligent agent are obtained based on the query command. Retrieve a first number of summary data that have the highest matching degree with the interaction information from the first long-term memory; Extract a second number of statement data associated with the first number of summary data from the second long-term memory, wherein the storage duration of the statement data in the first long-term memory is longer than the storage duration of the summary data in the second long-term memory; Response data to the interactive information is generated based on the extracted statement data, summary data, profile information, and status information.
[0007] Optionally, the method further includes: When the memory data in the second storage duration exceeds the preset second storage duration threshold, the memory data exceeding the second storage duration threshold will undergo initial summary processing to form summary data; The resulting summary data is added to the first duration memory, and memory data that exceeds the second storage duration threshold is deleted from the second duration memory.
[0008] Optionally, the step of performing initial summarization processing on memory data exceeding a second storage duration threshold to form summary data includes: Selective forgetting is performed on memory data that exceeds the second storage duration threshold. The memory data retained after the forgetting is then subjected to initial summarization processing to form summary data.
[0009] Optionally, the selective forgetting screening of memory data exceeding the second storage duration threshold includes: Identify the first degree of matching between memory data exceeding the second storage duration threshold and the agent's profile information, and the second degree of matching with the emotional state; Identify the importance of events reflected in memory data that exceeds a second storage duration threshold; The probability of forgetting is determined based on the first matching degree, the second matching degree, and the importance of the event; The determination of whether to retain the corresponding memory data is based on the forgetting probability, wherein the forgetting probability is negatively correlated with the memory retention probability, and the first matching degree, the second matching degree, and the event importance are negatively correlated with the forgetting probability.
[0010] Optionally, generating response data to the interactive information based on the extracted statement data, summary data, profile information, and status information includes: Based on the interaction information, the extracted first quantity of summary data and the second quantity of statement data are reordered; The third number of summary data and the fourth number of statement data are selected based on the reordering; Response data to the interactive information is generated based on the reordered summary data, statement data, profile information, and status information.
[0011] Optionally, generating response data to the interactive information based on the reordered summary data, statement data, profile information, and status information includes: According to the preset prompt words, the interactive information, reordered summary data, statement data, profile information and status information are converted into input information for the preset large language model. The input information is used as input to the large language model to generate corresponding response data.
[0012] Optionally, before generating response data to the interaction information based on the reordered summary data, statement data, profile information, and status information, the method further includes: directly extracting a fifth number of statement data from the second time-long memory based on the interaction information.
[0013] Optionally, the step of reordering the extracted first number of summary data and the second number of statement data based on the interaction information includes: reordering the extracted first number of summary data, the fifth number of statement data, and the second number of statement data based on the interaction information.
[0014] Optionally, the first time-limited memory includes multiple levels of time-limited memory, and the storage duration of the statement data in different levels of time-limited memory is different. The step of recalling a first number of summary data with the highest matching degree with the interaction information from the first time-limited memory includes: Extract the corresponding number of summary data with the highest matching degree from each level of the first time memory, so that the sum of the summary data extracted from each level of the time memory is the first number.
[0015] Optionally, the method further includes: when the storage duration of the statement data of one of the first-level time memories in the first time memory exceeds a preset time threshold, the statement data exceeding the corresponding time threshold is re-summarized to form the statement data of the corresponding next-level time memory, and stored in the corresponding next-level time memory.
[0016] In a second aspect, this application provides a computer-readable storage medium storing executable instructions that, when executed by a processor, cause the processor to perform the method described in any embodiment of this application.
[0017] A third aspect of this application provides an electronic device, comprising: One or more processors; A memory for storing one or more programs that, when executed by one or more processors, cause the one or more processors to perform the method as described in any embodiment of this application.
[0018] The long-term memory-based intelligent agent interaction method, medium, and device in this application simultaneously retrieve fixed profile information, dynamic state information, condensed summary data, and detailed descriptive data. Profile information determines the core direction of the response, aligning with the user's long-term fixed preferences; state information allows the response to adapt to the user's current scenario and state; summary data captures the core patterns of the user's historical interactions; and descriptive data provides specific historical details to support the response. The fusion of multiple data types avoids the one-sidedness of a single data dimension, enabling the intelligent agent to accurately capture the core of user needs and generate personalized responses tailored to the user's individual characteristics, rather than generic, undifferentiated feedback. This solves the problems of single data structure and lack of personalized design in existing technologies.
[0019] Furthermore, this application first retrieves condensed summary data from the first-time memory with a long storage duration, achieving rapid coarse screening of massive memory data and ensuring response efficiency. This simulates the rapid retrieval of summary memory data by humans, although it lacks memory details. For the rapidly retrieved summary memory data, related descriptive data is then extracted from the second-time memory with a short storage duration, enabling the tracing and expansion of condensed memory into original detailed memory. This allows for the recall of details of distant / important events, providing the response with specific historical interaction details. This linkage mode between high- and low-order memories simulates the human brain's memory retrieval logic of "first recalling core conclusions, then tracing specific details," avoiding the inefficiency caused by directly retrieving massive amounts of original descriptive data and solving the problem of insufficient detail and authenticity in the response caused by using only condensed summary data. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation on the scope of this application.
[0021] Figure 1 This is a flowchart illustrating an intelligent agent interaction method based on long-term memory in one embodiment; Figure 2 This is a schematic diagram of a process for selectively forgetting memory data that exceeds a second storage duration threshold in one embodiment; Figure 3 This is a schematic diagram of a process for generating response data to interactive information based on extracted statement data, summary data, profile information, and status information in one embodiment. Figure 4 This is a flowchart illustrating an electronic device in one embodiment. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0023] All terms used in this application (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0024] For example, the terms "first," "second," etc., used in this application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from another element.
[0025] For example, the terms "comprising" or "including" used in this application indicate the presence of features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0026] This application provides an intelligent agent interaction method based on long-term memory, such as... Figure 1 As shown, the method includes: Step 110: Trigger a query command based on the user's interaction information, and obtain the profile information and status information of the intelligent agent based on the query command.
[0027] The AI Agent in this application refers to an artificial intelligence entity equipped with an innovative long-term memory system. Its core characteristics are human-like abilities to organize, internalize, forget, and recall memories. It can construct a unique personality / behavioral logic through a hierarchical and structured memory system, making it suitable for scenarios requiring long-term interaction with users and the formation of unique cognition, such as virtual personalities, intelligent companions, dedicated customer service, and personalized question-and-answer robots. For example, this AI Agent could be an intelligent doll, a digital human, a virtual boyfriend / girlfriend, a dedicated intelligent customer service representative, or a virtual game character.
[0028] User interaction information represents various interaction data generated between the intelligent agent and the user through online applications and offline devices. It serves as the original input that triggers the intelligent agent's response, including forms such as text dialogue, voice commands, operational behaviors, and scene feedback. After receiving user interaction information, the intelligent agent automatically generates query commands to retrieve its own memory data and initiate response logic, and responds to the interaction information based on the corresponding information retrieved.
[0029] Profile information is structured data used to describe the virtual personality of an intelligent agent, its corresponding interactive objects, and the virtual world it inhabits. It is organized in a tree / graph + key-value pair format, and its attribute characteristics are stable over a long period and do not change frequently with interaction scenarios. For example, it may include data such as the agent's and / or user's persona, personality, and speech characteristics. It may also include the agent's functional positioning, service areas, interaction style, and professional scope. For example, if the intelligent agent is a virtual boyfriend, its profile information would be: {Role: Virtual Boyfriend, Age: 28, Occupation: Designer, Personality: Gentle and considerate, User: XX, User Age: 25, User Occupation: Teacher, User Preferences: Likes watching movies and drinking milk tea}; if the intelligent agent is a dedicated customer service representative, its profile information would be: {Brand: XX Cosmetics, Role: Brand Dedicated Customer Service Representative, User: XX, User Skin Type: Dry Skin, User Membership Level: VIP2, User Purchase History: Face Cream, Lipstick}; if the intelligent agent is a virtual game character, its profile information would be: {Role: Disciple of a Martial Arts Sect, Sect: Wudang, Martial Arts Skill: Tai Chi Sword, World: Jianghu, User: XX, User Role: Junior Sister, User Level: Level 50}. Based on the user identifier and scene information in the query command, the corresponding profile information is retrieved from the graph database of the memory database. For example, the user's profile information is "tree structure: diet preference > sweets > cream; key-value pairs: taste = sweet, dietary restrictions = none, consumption level = mid-range". At the same time, the virtual personality profile of the intelligent agent itself is retrieved: "attributes: food recommendation assistant, style: friendly and down-to-earth, professionalism: familiar with local offline dessert shops".
[0030] Status information is key-value pair data used to describe the real-time dynamic attributes of an agent. It can be frequently updated as the interaction progresses and over time, reflecting the agent's current dynamic characteristics. For example, it may include the agent's current emotional characteristics (such as pleasure, doubt, anxiety, satisfaction, etc.). Similarly, if the agent is a virtual boyfriend, its status information could be {agent's mood: happy, user's mood: depressed, user's hunger: yes, current interaction scenario: user is going home from get off work}; if it is a dedicated customer service agent, its status information could be {user's current need: inquiring about the moisturizing effect of face cream, agent's current reception status: idle, brand's current promotion: buy one get one free face cream}; if it is a virtual game character agent, its status information could be {agent's health: 80%, agent's stamina: 50%, user's health: 90%, current scenario: fighting monsters}.
[0031] The intelligent agent collects user interaction information such as dialogue and operation through online applications and offline devices. After semantic recognition, it triggers standardized query commands. The query command sends a retrieval request to the intelligent agent's memory database and extracts complete profile information and real-time status information that match the current interaction scenario from the dedicated profile data module and status data module, laying the foundation data for subsequent response generation.
[0032] Specifically, intelligent agents can acquire user text / voice interaction information from online chat apps and intelligent customer service interfaces. They can also simultaneously acquire offline interaction information such as user voice commands and body language feedback from offline devices such as smart speakers and smart robots, thus collecting interaction information and storing it as memory data in a memory database. For example, a user sending a text message to an intelligent virtual assistant, "I want to go out for something sweet today, any recommendations?" is a typical example of online interaction information.
[0033] Step 120: Retrieve the first number of summary data that have the highest matching degree with the interaction information from the first time memory.
[0034] Step 130: Retrieve a second number of statement data associated with the first number of summary data from the second temporal memory. The storage duration of the statement data in the first temporal memory is longer than the storage duration of the summary data in the second temporal memory.
[0035] First-term memory, within an agent's memory system, is a collection of memory data formed by summarizing, filtering, and accumulating declarative data one or more times. This memory data is called summary data. Summary data is condensed data formed by summarizing and refining the original interactive declarative data, retaining the core information of the original data and eliminating redundant content. It is the compressed form of information in each level of long-term memory.
[0036] The second-term memory, within the agent's long-term memory system, stores uncondensed, raw interactive statements, or a collection of statements more closely resembling the original interactions. The storage duration of data in the second-term memory is shorter than that in the first-term memory; the storage duration represents the time elapsed from the moment the memory was created to the present. The descriptive data consists of original, detailed data formed by semantically recording the user's original interaction process with the agent and / or the agent's own behavior. It retains information such as the specific context of the interaction, dialogue content, and operational details, serving as the basis for generating summary data. Understandably, the data in the second-term memory is continuously updated based on data generated from user-agent interactions or the agent's own behavior.
[0037] Data in the second-term memory can be periodically or irregularly cleared. The cleared data can be summarized and transferred to the first-term memory. For example, when the clearing time arrives, all data in the second-term memory can be cleared and processed into summary data; or only data whose storage time exceeds the corresponding duration threshold can be cleared and processed into summary data.
[0038] Matching degree represents the degree of association between memory data and user interaction information, calculated through semantic similarity algorithms, keyword matching algorithms, etc. The higher the matching degree value, the stronger the relevance between the data and user needs. The first quantity is a threshold value for the number of recalled summary data that is pre-set according to the intelligent agent's business scenario and response efficiency requirements. It is a fixed value or can be dynamically adjusted according to the interaction scenario. The second quantity is a threshold value for the number of statement data that is extracted from the second duration memory and associated with one or more summary data that is pre-set according to the information capacity of the summary data and the richness of detail of the statement data. It can be the same as or different from the first quantity.
[0039] Targeting the core features of user interaction information, a retrieval algorithm is used to filter out a predetermined number of summary data with the highest relevance to the interaction information in the first-term memory, thereby enabling the retrieval of memory data. Specifically, all summary data in the user interaction information and the first-term memory are vectorized using an embedding model. A semantic similarity retrieval algorithm is used to calculate the similarity between vectors to represent the matching degree. The data is then sorted from high to low according to the matching degree, and finally, the top N (first quantity) summary data are retrieved, completing the filtering of the first-term memory.
[0040] Since the summary data is derived from the statement data, although the statement data after summary processing no longer exists in the second long-term memory, based on the source association between the summary data and the statement data, it is still possible to associate some statement data that is related to and / or similar to the summary data from the existing second long-term memory, so as to realize the tracing and expansion from condensed memory to detailed memory.
[0041] Step 140: Generate response data to the interactive information based on the extracted statement data, summary data, profile information, and status information.
[0042] In this embodiment, response data is feedback data generated by the intelligent agent based on the user's original interaction information and by integrating various types of memory data. This feedback data meets the user's needs and its format matches the user's interaction information, potentially including one or more of the following: text replies, voice recommendations, and image / text links. The extracted statement data, summary data, profile information, and status information can be fused. For example, they can be integrated, filtered, and reorganized according to the needs of the business scenario to form unified contextual data that meets the input requirements of the large language model.
[0043] The extracted detailed statement data, condensed summary data, fixed profile information, and dynamic status information are integrated in multiple dimensions and input into the large language model. The model then combines the core needs of the interaction information to generate accurate response data that fits the user characteristics.
[0044] Specifically, by filtering and reorganizing statement data, summary data, profile information, and state information, redundant information is eliminated, and key data related to the core needs of user interaction information is retained to construct a structured large language model input context. This context, along with the user's original interaction information, is then input into a pre-trained large language model (LLM). The model generates response data to the interaction information according to the virtual personality style, response logic, and mood state of the agent preset in the summary data and profile information. This makes the generated response data more in line with the agent's personalized response, and the response data is sent to the agent so that the agent can respond to the user's interaction information.
[0045] The long-term memory-based agent interaction method in this application simultaneously retrieves fixed profile information, dynamic state information, condensed summary data, and detailed descriptive data. Profile information determines the core direction of the response, aligning with the user's long-term fixed preferences; state information adapts the response to the user's current scenario and state; summary data captures the core patterns of the user's historical interactions; and descriptive data provides specific historical details to support the response. The fusion of multiple data types avoids the one-sidedness of a single data dimension, enabling the agent to accurately capture the core of user needs and generate personalized responses tailored to individual user characteristics, rather than generic, undifferentiated feedback. This solves the problems of single data structure and lack of personalized design in existing technologies.
[0046] Furthermore, this application first retrieves condensed summary data from the first-time memory with a long storage duration, achieving rapid coarse screening of massive memory data and ensuring response efficiency. This simulates the rapid retrieval of summary memory data by humans, although it lacks memory details. For the rapidly retrieved summary memory data, related descriptive data is then extracted from the second-time memory with a short storage duration, enabling the tracing and expansion of condensed memory into original detailed memory. This allows for the recall of details of distant / important events, providing the response with specific historical interaction details. This linkage mode between high- and low-order memories simulates the human brain's memory retrieval logic of "first recalling core conclusions, then tracing specific details," avoiding the inefficiency caused by directly retrieving massive amounts of original descriptive data and solving the problem of insufficient detail and authenticity in the response caused by using only condensed summary data.
[0047] In one embodiment, the method further includes: when the memory data in the second long-term memory exceeds a preset second storage duration threshold, performing initial summary processing on the memory data exceeding the second storage duration threshold to form summary data; adding the formed summary data to the first long-term memory; and deleting the memory data exceeding the second storage duration threshold from the second long-term memory.
[0048] In this embodiment, the second storage duration threshold is a pre-set time limit used to determine whether the memory data in the second long-term memory needs to undergo digest processing. When the duration of the memory data stored in the second long-term memory reaches the preset second storage duration threshold, the digest mechanism is triggered to perform initial digest processing on all expired memory data, generating corresponding digest data. The generated digest data is stored in the first long-term memory for archiving, while the expired original memory data that has completed digest processing is cleared from the storage medium of the second long-term memory, releasing the storage resources of the second long-term memory.
[0049] The initial summarization process involves extracting and compressing information from the target data (descriptive data) to retain key information while reducing data storage volume. This includes extracting the core semantics, key features, and core interactive value from the descriptive data to generate lightweight, high-value-density summary data. The summary data generated after initial summarization is structured data; it is not a simple text string but a composite data structure containing core semantic identifiers, key interactive elements, and value weight labels.
[0050] Specifically, semantic parsing, value screening, core feature extraction, and structured generation can be performed on target data that requires summary processing to form summary data.
[0051] During the semantic parsing process, natural language processing algorithms are used to segment the target data into sentences, words, and parts of speech to identify basic semantic units. Dependency parsing is used to analyze the subject-predicate, verb-object, and modifier relationships between semantic units. Finally, the correlation between semantic units is calculated to filter out semantic fragments related to the core interaction theme, remove meaningless modifiers, redundant modifiers, and other interfering information, and output the basic semantic parsing results.
[0052] For example, the BERT-base model can be used as the basic semantic parsing model to perform word segmentation and vector embedding on the original memory data. For example, the original memory data is "On March 1, 2026 at 10:00, a user consulted about spring health recipes, asking for light recipes suitable for people with weak spleen and stomach. Three recipes were recommended: stir-fried vegetables, yam and pork rib soup, and millet porridge. The user expressed satisfaction." After word segmentation, the basic semantic units are obtained as follows: [March 1, 2026 at 10:00, user, spring health recipes, people with weak spleen and stomach, light recipes, 3 recipes, stir-fried vegetables, yam and pork rib soup, millet porridge, satisfied]. Dependency relationship analysis: Through dependency parsing, the relationships between the core semantic units are analyzed as follows: "user (subject) - consultation (action) - spring health recipes (object)", "people with weak spleen and stomach (modifier) - light recipes (object)", "3 recipes (quantity) - recipes (headword)", "satisfied (emotion) - user (subject)". Interference removal: By setting interference removal rules, modal words (such as "ah" and "ya") and redundant modifiers (such as "very" and "especially") are filtered out, and only semantic units related to the core interaction are retained, and the basic semantic parsing results are output.
[0053] During the value selection process, a multi-value assessment index system, including pre-constructed service relevance, emotional value, and information uniqueness, is used to quantitatively score each basic semantic unit and semantic fragment. A scoring threshold is set, and core value fragments with scores exceeding the threshold are selected to ensure that the retained fragments possess high service value, high emotional value, and uniqueness. Service relevance reflects the degree of matching between the semantic fragment and the agent's profile information; emotional value reflects the supporting value of the user's emotional state (such as satisfaction, doubt, and anxiety) contained in the semantic fragment for subsequent interactions; and information uniqueness reflects whether the content of the semantic fragment is appearing for the first time and is free of duplication or redundancy, avoiding the retention of duplicate information. By quantifying the index values of service relevance, emotional value, and information uniqueness, and then weighting and summing these index values, a comprehensive value score for each semantic fragment is obtained. High-value semantic fragments with scores exceeding the threshold are retained, while worthless semantic fragments with scores below the threshold are removed.
[0054] In the core feature extraction process, deep semantic mining is performed on the retained high-value semantic fragments to extract core features that can be used for subsequent retrieval, matching, and response generation. High-dimensional feature vectors are generated through vector mapping technology to achieve the digital expression of semantic features. Specifically, a multi-level feature extraction logic, including entity features, scene features, and value features, can be adopted to extract core entity and key scene information, mine the core features of the interaction scene, and associate value feature labels. Through a pre-trained model (such as the Sentence-BERT model), the multi-level features are mapped into high-dimensional feature vectors, so that the vectors not only retain the core semantic information but also adapt to the algorithm requirements of subsequent retrieval and matching, thus completing the digital transformation of semantic features.
[0055] Entity feature extraction is used to extract key entities from core value segments, such as time, subject, core object, and core action. Scene feature extraction is used to mine the core attributes of interactive scenes, such as scene type (health consultation), scene needs (recipe recommendation), and scene constraints (weak spleen and stomach). Scene features include health consultation scene, recipe recommendation needs, and constraints related to people with weak spleen and stomach. Value feature extraction is used to correlate with the value assessment results of the previous stage and extract value tags, such as high service relevance, high emotional value, and unique information. Value features include: high service relevance, high emotional value, and unique information.
[0056] During the structured generation process, the extracted core feature vectors and value tags are integrated according to a pre-defined standardized structure to generate structured summary data containing multiple fields. This ensures that the summary data conforms to the storage specifications of the first-time memory and can be directly used for subsequent retrieval, matching, and response generation. For example, the core identifiers obtained from the semantic parsing of the original memory data, the core content features extracted from the core feature vectors, value features, and association features are integrated into structured data. At the same time, metadata such as data volume compression ratio and timestamp generation are added to complete the final output of the initial summary processing.
[0057] Abstraction processing removes interfering information and worthless fragments from the original data, reducing data storage while preserving the core interactive value. Its core feature vectors accurately map the semantic and value information of the original memory data. In subsequent retrieval and matching processes, the agent can directly calculate the similarity of feature vectors to quickly locate the abstract data most relevant to the current interactive information, improving data retrieval efficiency. Simultaneously, the addition of value tags further filters high-value data, preventing low-value data from interfering with the retrieval process and further improving the accuracy of retrieval and matching.
[0058] In one embodiment, performing initial summarization processing on memory data exceeding a second storage duration threshold to form summary data includes: selectively forgetting the memory data exceeding the second storage duration threshold, and performing initial summarization processing on the memory data retained by the forgetting screening to form summary data.
[0059] In this embodiment, selective forgetting screening refers to the process of filtering target data (overdue memory data) according to preset rules before the initial summarization processing, retaining only a portion of the data for summarization. The data removed by this screening is typically of low retention value, avoiding the omission of important data due to random forgetting. This selective forgetting screening process mimics human memory, preserving important and deeply memorable data. The memory data retained by the forgetting screening are overdue memory data that, after selective forgetting screening, are determined to have retention value and require further initial summarization processing.
[0060] Specifically, selective forgetting filtering can include an interaction value dimension and an information uniqueness dimension. The interaction value dimension is used to determine whether the data is related to the core interaction scenario of the agent, while the information uniqueness dimension is used to determine whether the data is duplicate or redundant information.
[0061] For example, there are 3 pieces of data to be deleted in the second long-term memory: Data A "User asks about the function introduction of the agent", Data B "User repeatedly asks about the function introduction of the agent (completely consistent with the content of Data A)", and Data C "User consults about office software usage skills".
[0062] For data A: the interaction value dimension is judged as "high" (belonging to the core interaction scenario of the intelligent agent), and the information uniqueness dimension is judged as "unique," therefore the screening result is "retain." For data B: the interaction value dimension is judged as "high," but the information uniqueness dimension is judged as "repetitive and redundant," therefore the screening result is "forgotten." For data C: the interaction value dimension is judged as "high" (belonging to a commonly used interaction scenario of the intelligent agent), and the information uniqueness dimension is judged as "unique," therefore the screening result is "retain." Through the selective forgetting screening mechanism, only data A and data C are retained as memory data through forgetting screening, while data B is deleted.
[0063] In one embodiment, such as Figure 2 As shown, selective forgetting filtering is performed on memory data that exceeds the second storage duration threshold, including: Step 210: Identify the first degree of matching between memory data exceeding the second storage duration threshold and the agent's profile information, and the second degree of matching with the emotional state.
[0064] The system includes a pre-set matching degree calculation model. It extracts the core content features and auxiliary emotional features from the memory data, and then extracts the core attribute features and core service target features from the agent profile information. The first matching degree is obtained by calculating the feature vector similarity between the core content features and the core attribute features. The second matching degree is obtained by calculating the auxiliary emotional features and the core service target features, and then the matching degree is associated with the corresponding memory data.
[0065] Step 220: Identify the importance of events reflected in memory data that exceeds the second storage duration threshold.
[0066] Event importance reflects the value of user interaction events recorded in overdue memory data to the agent's subsequent provision of personalized services to that user. It is a core indicator for measuring the long-term retention value of memory data; the higher the event importance value, the more important the event. By extracting the core features of the interaction events corresponding to the overdue memory data (such as event type, user participation, and service frequency), and quantifying and scoring each feature according to preset evaluation weights, the quantitative results of event importance are finally obtained.
[0067] Step 230: Determine the forgetting probability based on the first matching degree, the second matching degree, and the importance of the event.
[0068] Specifically, based on the preset forgetting probability calculation formula, the quantitative values of the first matching degree, the second matching degree, and the event importance are substituted into the calculation to obtain the forgetting probability of each overdue memory data. The forgetting probability is negatively correlated with the first matching degree, the second matching degree, and the event importance.
[0069] The formula for calculating the probability of forgetting can be expressed as: Probability of forgetting P = 1 - (w1×p1 + w2×p2 + w3×p3). Here, w1, w2, and w3 are weighting coefficients, and p1, p2, and p3 represent the first matching degree, the second matching degree, and the importance of the event, respectively.
[0070] Step 240: Determine whether the corresponding memory data should be retained based on the forgetting probability.
[0071] The probability of forgetting is negatively correlated with the probability of memory retention. The first match degree, the second match degree, and the importance of the event are also negatively correlated with the probability of forgetting.
[0072] Optionally, a forgetting probability threshold can be set, and the forgetting probability of each overdue memory data can be compared with the threshold. Based on the comparison result, it is determined whether the data should be retained. If the forgetting probability is lower than the threshold, the data is retained; if it is higher than or equal to the threshold, the data is forgotten.
[0073] This application, through the first matching degree between profile information, the second matching degree with emotional state, and the quantitative indicators of event importance, can avoid the interference of human subjective factors and ensure that the screening results are highly matched with the actual value of the memory data.
[0074] In one embodiment, such as Figure 3 As shown, response data to interactive information is generated based on the extracted statement data, summary data, profile information, and status information, including: Step 310: Reorder the extracted first number of summary data and the second number of statement data based on the interaction information.
[0075] The first and second quantities are the pre-set initial quantities of recalled summary data and statement data, respectively, and are the basic data volume for re-ranking. For example, the first quantity is preset to 10 items, and the second quantity to 5 items.
[0076] The core semantic features of the current interaction information can be extracted, and then the core semantic features of the first number of recalled summary data and the second number of recalled statement data can be extracted. The matching degree between each data and the interaction information is obtained by similarity calculation. Then, the summary data and statement data are uniformly reordered according to the matching degree from high to low to form a single ordered data list.
[0077] Step 320: Select the third number of summary data and the fourth number of statement data based on the reordering.
[0078] The third and fourth quantities are pre-defined numbers of summary and statement data selected from the reordered list of data to generate the response. The third quantity must be less than or equal to the first quantity, and the fourth quantity must be less than or equal to the second quantity. The summary and statement data with the highest matching degree are selected from the reordered list of data until the pre-defined third and fourth quantities are reached, forming a dedicated memory dataset for response generation.
[0079] Step 330: Generate response data to the interactive information based on the reordered summary data, statement data, profile information, and status information.
[0080] In one embodiment, before step 330, the method further includes: extracting a fifth number of summary data directly from the second duration memory based on interactive information; step 310 includes: reordering the extracted first number of summary data, the fifth number of summary data, and the second number of statement data based on interactive information.
[0081] In addition to extracting statement data based on summary data, the extraction process further directly extracts statement data based on interaction data, ensuring that the extracted statement data is the sum of the second and fifth quantities. Specifically, the matching degree between the core semantic features of the interaction information and each statement data is calculated, and the fifth quantity of statement data with the highest matching degree is selected. The first quantity of summary data from the first time period is integrated with the second and fifth quantity of statement data from the second time period. Based on the core features of the current interaction information, all integrated data are uniformly reordered to generate a new ordered data list.
[0082] In one embodiment, interactive information, reordered summary data, statement data, profile information, and status information are converted into input information for a preset large language model according to a preset prompt word engineering method; the input information is used as input to the large language model to generate corresponding response data.
[0083] Specifically, according to the preset prompt word engineering template, the current interaction information, the reordered selected summary data, statement data, profile information and status information are structurally integrated to generate standardized input information for the large language model; the input information is input into the preset large language model, and the model outputs the corresponding response data.
[0084] For example, the template for this prompt word project is: { <Example of prompt word> Your close human friend has said something to you. Please refer to your emotions, memories, and task characteristics when generating your response.
[0085] Your friend's words: <chat> [What the user said at the beginning] < / chat> Your character's personality: Arrogant Your catchphrase: I don't like it. Your relationship with your friends: Current mood: Happy Your past memories with friends: <memory> [Memories of Recall] <memory> < / Prompt example> } According to the corresponding prompt engineering template, the extracted memory data and interaction data to be carried are injected into the large language model as context information. Based on its own semantic understanding and generation ability, the model outputs response data.
[0086] In one embodiment, the first duration memory includes multiple levels of duration memories, and the storage durations of the statement data in different levels of duration memories are different. Retrieving the first quantity of summary data with the highest matching degree to the interaction information from the first duration memory includes: extracting the corresponding quantity of summary data with the highest matching degree from each level of duration memory in the first duration memory, so that the sum of the summary data extracted from each level of duration memory is the first quantity.
[0087] In this embodiment, the summary data in the first duration memory is further divided into multiple levels according to the storage duration. For example, it can be divided into three levels, namely the first-level duration memory (7 days < storage duration ≤ 30 days), the second-level duration memory (30 days < storage duration ≤ 180 days), and the third-level duration memory (storage duration > 180 days). The first duration memory is divided into multiple levels of duration memories, and the storage duration thresholds of each level are different; from each level of duration memory, the preset corresponding quantity of summary data with the highest matching degree to the interaction information is extracted respectively, ensuring that the total amount of summary data extracted from all levels is equal to the first quantity.
[0088] For each level of duration memory, the interaction information matching degree calculation and precise extraction operations are independently performed to extract the preset corresponding quantity of summary data; finally, the summary data extracted from each level is aggregated to form a multi-level recall summary data set with a total quantity of the first quantity. The statement data association extraction is performed on the summary data in the formed multi-level recall summary data set.
[0089] In one embodiment, the above method further includes: when the storage duration of the statement data in one level of the first duration memory exceeds the preset duration threshold, the statement data exceeding the corresponding duration threshold is re-summarized to form the statement data of the corresponding next level of duration memory, and it is stored in the corresponding next level of duration memory.
[0090] In this embodiment, the second-level summarization process refers to a secondary information extraction and compression process performed on expired statements at a certain level of the first-level long-term memory. Compared to the initial summarization process, the second-level summarization process achieves a higher degree of compression, with the core objective of retaining the "core of the core" information. The next level of long-term memory is specifically the storage level below the current level. After the second-level summarization process, the expired statements will be archived to this level, realizing the hierarchical sinking of memory data. The process of the second-level summarization process is similar to that of the initial summarization process, and may also include processes such as semantic parsing of the summary data, value screening, core feature extraction, and structure generation.
[0091] By employing a method of re-summarizing expired data and hierarchical decentralization, expired summary data is compressed to a greater extent. While retaining core information, this further reduces data storage volume, achieving a balance between preserving core information from long-term memories and conserving storage resources. This results in an effect similar to internalization and forgetting in life. Simultaneously, a multi-level retrieval mechanism ensures that core long-term memories are accurately retrieved when needed, solving the technical problems of difficulty in preserving long-term memories and their inefficient use after preservation. This allows each user to possess unique data and their own exclusive model context, continuously accumulating knowledge and past experiences.
[0092] Furthermore, embodiments of this application also provide a computer-readable storage medium storing executable instructions thereon, which, when executed by a processor, cause the processor to perform the methods described in the above-described method embodiments.
[0093] Furthermore, this application also provides a computer program product, the computer program product including a non-transitory computer-readable storage medium storing a computer program, which, when connected to a computer device, is executed by one or more processors of the computer device, is capable of performing all or part of the steps of the method described in any embodiment of this application.
[0094] In one embodiment, an electronic device is also provided, including one or more processors; and a memory storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the steps in the above method embodiments.
[0095] In one embodiment, such as Figure 4 The diagram illustrates the structure of an electronic device used to implement an embodiment of this application. The electronic device includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 402 or a program loaded from a storage portion 408 into a random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the electronic device 400. The CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0096] The following components are connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. Drive 410 is also connected to I / O interface 405 as needed. Removable media 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 410 as needed so that computer programs read from them can be installed into storage section 408 as needed.
[0097] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer-readable medium carrying instructions that, in such embodiments, can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the instructions are executed by central processing unit (CPU) 401, the various method steps described in this application are performed.
[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0099] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of this application and form different embodiments. For example, all the embodiments above can be used in any combination. The information disclosed in this background section is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that such information constitutes prior art known to those skilled in the art.< / memory> < / memory>
Claims
1. A method for intelligent agent interaction based on long-term memory, characterized in that, The method includes: A query command is triggered based on the user's interaction information, and the profile information and status information of the intelligent agent are obtained based on the query command. Retrieve a first number of summary data that have the highest matching degree with the interaction information from the first long-term memory; Extract a second number of statement data associated with the first number of summary data from the second long-term memory, wherein the storage duration of the statement data in the first long-term memory is longer than the storage duration of the summary data in the second long-term memory; Response data to the interactive information is generated based on the extracted statement data, summary data, profile information, and status information.
2. The method according to claim 1, characterized in that, The method further includes: When the memory data in the second storage duration exceeds the preset second storage duration threshold, the memory data exceeding the second storage duration threshold will undergo initial summary processing to form summary data; The resulting summary data is added to the first duration memory, and memory data that exceeds the second storage duration threshold is deleted from the second duration memory.
3. The method according to claim 2, characterized in that, The step of performing initial summarization processing on memory data exceeding the second storage duration threshold to form summary data includes: Selective forgetting is performed on memory data that exceeds the second storage duration threshold. The memory data retained after the forgetting is then subjected to initial summarization processing to form summary data.
4. The method according to claim 3, characterized in that, The selective forgetting screening of memory data exceeding the second storage duration threshold includes: Identify the first degree of matching between memory data exceeding the second storage duration threshold and the agent's profile information, and the second degree of matching with the emotional state; Identify the importance of events reflected in memory data that exceeds a second storage duration threshold; The probability of forgetting is determined based on the first matching degree, the second matching degree, and the importance of the event; The determination of whether to retain the corresponding memory data is based on the forgetting probability, wherein the forgetting probability is negatively correlated with the memory retention probability, and the first matching degree, the second matching degree, and the event importance are negatively correlated with the forgetting probability.
5. The method according to claim 1, characterized in that, The generation of response data to the interactive information based on the extracted statement data, summary data, profile information, and status information includes: Based on the interaction information, the extracted first quantity of summary data and the second quantity of statement data are reordered; The third number of summary data and the fourth number of statement data are selected based on the reordering; Response data to the interactive information is generated based on the reordered summary data, statement data, profile information, and status information.
6. The method according to claim 5, characterized in that, The step of generating response data to the interactive information based on the reordered summary data, statement data, profile information, and status information includes: According to the preset prompt words, the interactive information, reordered summary data, statement data, profile information and status information are converted into input information for the preset large language model. The input information is used as input to the large language model to generate corresponding response data.
7. The method according to claim 5, characterized in that, Before generating response data for the interactive information based on the reordered summary data, statement data, profile information, and status information, the method further includes: Based on the interaction information, a fifth number of statement data are directly extracted from the second duration memory; The step of reordering the extracted first number of summary data and the second number of statement data based on the interaction information includes: reordering the extracted first number of summary data, the fifth number of statement data, and the second number of statement data based on the interaction information.
8. The method according to any one of claims 1 to 7, characterized in that, The first temporal memory contains multiple levels of temporal memory, and the storage duration of the statement data in different levels of temporal memory is different. The step of retrieving a first number of summary data points from the first temporal memory that have the highest matching degree with the interaction information includes: Extract the corresponding number of summary data with the highest matching degree from each level of the first time memory, so that the sum of the summary data extracted from each level of the time memory is the first number; The method further includes: when the storage duration of the statement data of one of the first-level time memories in the first time memory exceeds a preset time threshold, the statement data exceeding the corresponding time threshold is re-summarized to form the statement data of the corresponding next-level time memory, and stored in the corresponding next-level time memory.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores executable instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1 to 8.
10. An electronic device, characterized in that, include: One or more processors; A memory for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to perform the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Intelligent agent long-term memory modeling method based on memory network
CN120688545A