Memory retrieval method and related device based on large language model
Through the two-way transmission mechanism of short-term and long-term memory modules and personalized knowledge base processing, the memory capacity and retrieval efficiency of the large language model are solved, efficient and personalized information management and response are achieved, and user experience is improved.
Patent Information
- Application Number
- CN202411984396.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The existing large language models have limited memory capacity, high forgetting rate, low information retrieval efficiency, lack of effective classification mechanisms, and incomplete information updates, which affect user experience and interaction efficiency.
The two-way transmission mechanism of the short-term memory module and the long-term memory module is adopted, and information processing is carried out in combination with a personalized knowledge base. The short-term module temporarily stores real-time interactive information, and the long-term module stores key information that conforms to the dynamic transfer strategy, and responds through multimodal information perception and personalized adaptation mechanism.
It improves the efficiency of information retrieval, realizes the effective classification and organization of personalized information, reduces the information forgetting rate, and improves user experience and interaction quality.
Smart Images

Figure CN119903125B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a memory retrieval method based on a large language model and related devices. Background Art
[0002] With the rapid development of information technology, the amount of information people need to process and remember every day is exploding, which brings many challenges to the application of large language models (LLMs).
[0003] In related technologies, large language models primarily rely on traditional contextual storage methods, which have significant drawbacks. First, this storage method results in limited memory capacity for the model, making it unable to accommodate the large amount of personalized information accumulated by users over time, making it difficult to meet their actual needs. Second, the model has a high forgetfulness rate, making it difficult to persistently store specific user information. Frequent "forgetfulness" issues lead to a lack of coherence in interactions, impacting the user experience. Furthermore, information retrieval efficiency is low, making it difficult to accurately and quickly locate the personalized information required by users within a vast amount of information, significantly impacting the model's response speed and accuracy. Furthermore, related technologies lack effective information classification mechanisms, making it difficult to intelligently classify and organize user personalized information, resulting in disorganized information management and numerous obstacles in the retrieval process. Furthermore, imperfect information update mechanisms prevent the efficient integration and updating of new and old user information, which can easily lead to information redundancy or inconsistencies.
[0004] These issues are particularly evident during the interaction between large language models and users, severely hindering improvements in user experience and information processing efficiency. Therefore, a new technical solution is urgently needed to address at least one of these issues and promote the development and improvement of large models in practical applications. Summary of the Invention
[0005] In response to the technical problems existing in the prior art, the present application provides a memory retrieval method and related devices based on a large language model to solve at least one technical problem existing in the prior art.
[0006] In a first aspect, an embodiment of the present application provides a memory retrieval method based on a large language model, the method comprising:
[0007] In response to an interaction instruction issued by a target user, performing context analysis on the interaction instruction to obtain the real-time interaction intention of the target user;
[0008] Based on the real-time interaction intention, a short-term memory module and a long-term memory module are searched to obtain target search information; wherein, the short-term memory module and the long-term memory module support bidirectional transmission of stored information; the short-term memory module is used to temporarily store the real-time interaction information and real-time search information with the target user within a preset time period; the long-term memory module is used to store key information in the real-time interaction information and the real-time search information that meets the preset dynamic transfer strategy;
[0009] The target search information is personalized in combination with the personalized knowledge base of the target user to obtain response information that matches the target user's usage habits and is used to respond to the interactive instruction.
[0010] In a second aspect, an embodiment of the present application provides a memory retrieval device based on a large language model, the device comprising at least the following units:
[0011] An interaction unit is configured to respond to an interaction instruction of a target user, perform context analysis on the interaction instruction, and obtain a real-time interaction intention of the target user;
[0012] A retrieval unit is configured to perform memory retrieval on a short-term memory module and a long-term memory module based on the real-time interaction intention to obtain target retrieval information; wherein the short-term memory module and the long-term memory module support bidirectional transmission of stored information; the short-term memory module is used to temporarily store real-time interaction information and real-time retrieval information with a target user within a preset time period; and the long-term memory module is used to store key information in the real-time interaction information and the real-time retrieval information that complies with the preset dynamic transfer strategy;
[0013] The generating unit is configured to perform personalized processing on the target search information in combination with the personalized knowledge base of the target user, and obtain response information that matches the target user's usage habits to respond to the interactive instruction.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising:
[0015] at least one processor, memory, and input-output unit;
[0016] The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the memory retrieval method based on a large language model of the first aspect.
[0017] In a fourth aspect, a computer-readable storage medium is provided, comprising instructions, which, when executed on a computer, cause the computer to execute the memory retrieval method based on a large language model of the first aspect.
[0018] The beneficial effects of the present application are: providing a memory retrieval method and related device based on a large language model. In this technical solution, first, in response to an interaction instruction issued by a target user, the interaction instruction is contextually parsed to obtain the target user's real-time interaction intent. Then, based on the real-time interaction intent, a memory search is performed on the short-term memory module and the long-term memory module to obtain target retrieval information. The short-term memory module and the long-term memory module support bidirectional transmission of stored information; the short-term memory module is used to temporarily store real-time interaction information and real-time retrieval information with the target user within a preset time period; and the long-term memory module is used to store key information in the real-time interaction information and the real-time retrieval information that complies with a preset dynamic transfer strategy. Finally, the target retrieval information is personalized based on the target user's personalized knowledge base to obtain response information that matches the target user's usage habits in response to the interaction instruction. In summary, this solution can improve the information retrieval efficiency of the large language model, enhance the performance of the large language model in interactive scenarios, and improve the user experience by accurately understanding user intent, efficiently and comprehensively retrieving memory information, and achieving personalized responses. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flowchart of a memory retrieval method based on a large language model according to an embodiment of the present application;
[0020] Figure 2 Schematic diagram of a memory retrieval device based on a large language model according to an embodiment of the present application;
[0021] Figure 3 This is a schematic structural diagram of an electronic device according to an embodiment of the present application;
[0022] Figure 4 It is a structural diagram of a medium device according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0024] To address at least one technical problem in the related art, embodiments of the present application provide a memory retrieval method and related apparatus based on a large language model. In this technical solution, in response to an interaction instruction issued by a target user, the interaction instruction is contextually parsed to obtain the target user's real-time interaction intent. Furthermore, based on the real-time interaction intent, a memory search is performed on a short-term memory module and a long-term memory module to obtain target retrieval information. The short-term memory module and the long-term memory module support bidirectional transmission of stored information; the short-term memory module is used to temporarily store real-time interaction information and real-time retrieval information with the target user within a preset time period; and the long-term memory module is used to store key information in the real-time interaction information and real-time retrieval information that complies with a preset dynamic transfer strategy. Finally, the target retrieval information is personalized based on the target user's personalized knowledge base to obtain response information that matches the target user's usage habits in response to the interaction instruction. In summary, this solution can improve the information retrieval efficiency of the large language model, enhance the performance of the large language model in interactive scenarios, and enhance the user experience by accurately understanding user intent, efficiently and comprehensively retrieving memory information, and achieving personalized responses.
[0025] Specifically, the embodiment of the present application proposes a memory retrieval solution based on a large language model, which is composed of multiple specially trained large language models. For example, the memory retrieval system based on the large language model includes the following functions: multimodal information perception, information understanding and extraction, short-term and long-term memory, memory integration and updating, forgetting mechanism, information classification and indexing, context understanding, personalized adaptation, memory retrieval and reasoning, output generation, and central coordination modules. These modules simulate the working principles of the human memory system through a sophisticated collaborative mechanism, achieving efficient processing, storage, and retrieval of multimodal information. The workflow of the system begins with the reception and preliminary processing of multimodal information, which is then stored in short-term memory after deep understanding and extraction. At the same time, the embodiment of the present application analyzes the current context, activates relevant long-term memory, and performs information retrieval and reasoning according to needs. The output generation module generates personalized responses based on the retrieved memory and the current context. The embodiment of the present application also continuously optimizes the memory management strategy through modules such as memory integration and updating, forgetting mechanism, and personalized adaptation, thereby improving the long-term operating efficiency and personalization of the system.
[0026] It is understandable that the improvements of the technical solution of this application over the related art are analyzed from multiple perspectives below:
[0027] The short-term memory module focuses on real-time interaction with users and information retrieval within a preset time period, enabling rapid location of recently relevant content. The long-term memory module stores key information consistent with dynamic transfer strategies and can tap into important knowledge accumulated in the past. Both modules support bidirectional transfer of stored information, meaning they complement and coordinate with each other, enabling rapid recall of fresh content from short-term memory based on the current situation and recall of in-depth knowledge from long-term memory. This makes memory retrieval more comprehensive and efficient, encompassing as much information as possible that is valuable in responding to interactive commands and increasing the probability of retrieving the correct target information.
[0028] The target search information is personalized based on the target user's personalized knowledge base, ultimately resulting in a response that matches the user's usage habits. This eliminates the need for a one-size-fits-all answer and instead fully considers individual factors such as the user's unique preferences and past operating habits. For example, for users who prefer concise statements, the response will be as concise and clear as possible; for those who prefer detailed and in-depth explanations, richer and more detailed content can be provided, significantly enhancing the user experience and increasing user satisfaction with the system and their willingness to continue using it.
[0029] The short-term memory module can temporarily store real-time interaction information and real-time retrieval information within a preset time period, which helps to quickly grasp the latest situation within a certain time range. It can also update or clean up these temporary information according to certain time rules to avoid the accumulation of excessive redundant data affecting retrieval efficiency, allowing the system to focus on the most relevant interaction content at the moment to retrieve information.
[0030] The long-term memory module filters and stores key information according to the preset dynamic transfer strategy, ensuring that the data retained in the long-term memory are of important reference value for subsequent interactions. It effectively refines and integrates the information, making it possible to locate the core content more quickly during long-term memory retrieval, thereby improving the resource utilization efficiency and operational performance of the entire memory retrieval system.
[0031] In general, the embodiments of the present application significantly improve the memory capacity of large language models during interaction, reduce the rate of information forgetting, improve information retrieval efficiency, and achieve effective classification and organization of personalized information. At the same time, through improved context understanding capabilities and personalized adaptation mechanisms, the embodiments of the present application can maintain a high degree of coherence and personalized features in long-term interactions, thereby significantly improving the user experience and interaction quality.
[0032] The technical solution of the present application, the memory retrieval scheme based on the large language model provided in the embodiment of the present application can also be executed by an electronic device, which can be a server, a server cluster, or a cloud server. The electronic device can also be a terminal device such as a mobile phone, a computer, a tablet computer, a wearable device, or a special device (such as a special terminal device with a memory retrieval method system based on a large language model). These electronic devices can also be equipped with the chips or other hardware processing units introduced in the above embodiments. Alternatively, these electronic devices can also be installed with a service program for executing a memory retrieval scheme based on a large language model.
[0033] Figure 1 A flowchart of a memory retrieval method based on a large language model provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes the following steps:
[0034] 101, in response to an interaction instruction issued by a target user, performing context analysis on the interaction instruction to obtain a real-time interaction intention of the target user;
[0035] 102. Perform memory retrieval on a short-term memory module and a long-term memory module based on the real-time interaction intention to obtain target retrieval information;
[0036] 103 , combining the personalized knowledge base of the target user, performing personalized processing on the target search information, and obtaining response information that matches the target user's usage habits, so as to respond to the interactive instruction.
[0037] In an embodiment of the present application, the short-term memory module and the long-term memory module support bidirectional transmission of stored information. This bidirectional transmission mechanism enhances the synergy between the two memory modules. For example, when the real-time interaction information temporarily stored in the short-term memory module has been precipitated for a period of time and meets the key information standards set by the dynamic transfer strategy of the long-term memory module, it can be transferred to the long-term memory for long-term storage, so as to facilitate the subsequent more in-depth and comprehensive excavation of the user's past related situations. Conversely, when processing the current interactive instruction for memory retrieval, if there is historical key information related to the current real-time interaction intention in the long-term memory, it can be transmitted to the short-term memory module in a timely manner, and participate in the retrieval process together with the temporary information in the short-term memory, so that the retrieved target retrieval information is more accurate and comprehensive, and fully integrates the valuable data of the recent and past.
[0038] In an embodiment of the present application, the short-term memory module is used to temporarily store real-time interaction information and real-time retrieval information with the target user within a preset time period. The short-term memory module is mainly responsible for temporarily storing real-time interaction information and real-time retrieval information with the target user within a preset time period. The preset time period here is a pre-set time range, such as the past 1 hour, 1 day, etc. (the specific duration is set according to the actual application scenario and needs). Real-time interaction information covers various texts, data, etc. generated by the interaction process, such as the content of the instructions input by the user, the corresponding feedback made by the system, etc.; real-time retrieval information is all kinds of relevant information involved in the retrieval operation for the user interaction instruction, such as some candidate results retrieved in the intermediate process.
[0039] The temporary storage function of the short-term memory module helps focus on the current interaction context. During this time period, the system can quickly obtain the freshest and most directly related information to the current interaction, avoiding the need to sift through massive amounts of historical data from scratch when searching. This improves the timeliness of information retrieval and allows the system to quickly locate recent key situations when responding to recent related interaction instructions, ensuring the timeliness of the response and its relevance to the current scenario. At the same time, by setting a time limit for temporary storage of information, unlimited data accumulation is avoided, which helps maintain the rational use of system storage resources and stable retrieval efficiency.
[0040] In an embodiment of the present application, the long-term memory module is used to store the key information in the real-time interaction information and the real-time retrieval information that complies with the preset dynamic transfer strategy. The long-term memory module is used to store the key information in the real-time interaction information and the real-time retrieval information that complies with the preset dynamic transfer strategy. The dynamic transfer strategy is a set of pre-established rule systems for determining which information is worth transferring from the short-term memory module and storing for a long time. For example, these rules may involve considerations of multiple dimensions such as the importance of information (such as those involving core user needs, frequently occurring key knowledge points, etc.), the strength of relevance (content closely related to the user's long-term areas of concern), and the long-term value in the time dimension (information that may be of reference significance for subsequent interactions over a longer period of time).
[0041] By filtering and storing key information according to a dynamic transfer strategy, the long-term memory module effectively builds a refined knowledge base that is deeply valuable to user interactions. When faced with complex and diverse user interaction commands, especially those that require tracing back history and uncovering deeply connected content, the long-term memory module provides reliable and targeted key data support, ensuring that the system can more accurately understand intent and retrieve appropriate target information based on the complete picture of user interactions over time. This improves the quality and effectiveness of the entire system's responses, while also making the storage and management of user-related information more scientific and organized, focusing on the truly valuable elements.
[0042] As an optional embodiment, in 101, in response to the interaction instruction issued by the target user, the interaction instruction is context-parsed to obtain the real-time interaction intention of the target user.
[0043] As an optional embodiment, before performing memory retrieval on the short-term memory module and the long-term memory module based on the real-time interaction intention in step 102 and obtaining target retrieval information, the following steps may be further implemented:
[0044] 201. Collect multimodal information of target users;
[0045] 202. Extracting information to be stored from the multimodal information;
[0046] 203. Storing the information to be stored in a short-term memory module according to the importance score and the timeliness score;
[0047] 204. Transfer the key information stored in the short-term memory module to the long-term memory module according to the dynamic transfer strategy.
[0048] In the embodiment of the present application, the multimodal information includes at least: image information, audio information, and text information. Multimodal information refers to a collection of information containing multiple different types of data modalities, and in the embodiment of the present application, it covers at least image information, audio information, and text information. Image information can be pictures uploaded by users, visual image content displayed in the system interaction interface, etc., which convey corresponding meanings through elements such as color, shape, and object layout; audio information, such as user voice input, background sounds in the environment, etc., contains voice intonation, timbre, and various sound event characteristics; text information is text instructions input by users, text content fed back by the system, etc., which convey information through the meaning of text. These different modal information reflect the various situations of the target user in the process of interacting with the system from multiple angles, and together constitute a rich and three-dimensional information source, providing a basis for subsequent comprehensive understanding of users and the realization of personalized services.
[0049] From a visual information processing perspective, the Vision Transformer (ViT) serves as the backbone network, combined with pre-trained models such as CLIP to process image information. These models accurately capture key image features and convert them into vector representations that can interact with language models, enabling better integration of image information into subsequent processes. Furthermore, integrated object detection and scene segmentation algorithms capture more granular visual information, such as distinguishing between different objects and scene regions within an image, helping to visually understand user intent. From an audio processing perspective, self-supervised learning models such as Wav2Vec are used for speech recognition, converting speech into an analyzable form. Voiceprint recognition technology is also integrated to identify the speaker's identity, facilitating differentiation between different users. For ambient sounds, a Sound Event Detection (SED) algorithm is used to classify and extract features, comprehensively analyzing audio information and providing critical audio evidence for understanding user interaction intent. From a text processing perspective, pre-trained language models such as multilingual BERT or RoBERTa are employed, combined with Named Entity Recognition (NER) and Relation Extraction (RE) techniques to deeply mine textual information, identify key entities and their relationships, and achieve deep understanding. Then, sentiment analysis algorithms are used to capture the emotional characteristics of the text, because emotions can reflect the user's potential intentions, further improve the understanding of the meaning conveyed by the text, and help clarify the interaction intentions.
[0050] In the embodiment of the present application, with the help of cross-modal attention mechanism and contrastive learning method, different modal information such as visual, audio, and text are mapped to the same semantic space, eliminating modal barriers, facilitating association and integration, laying a solid foundation for subsequent memory storage and retrieval, and also helping to analyze the embodiment and association of user intentions in various modal information. In addition, the fault tolerance mechanism for modal loss ensures that even if some modal information is missing, the system can still process information stably to avoid affecting the analysis of user intentions.
[0051] To further improve the efficiency of multimodal processing, a pipeline parallel processing architecture is adopted in the embodiment of the present application. Each modal information can be processed simultaneously, fully utilizing resources and speeding up processing. Finally, the Feature Fusion Network is used to integrate all information into a unified representation, organically combining multimodal information. This not only significantly enhances the real-time performance of the system, enabling it to quickly respond to interactive commands, but also ensures the integrity and accuracy of information processing, avoids information problems, and ensures accurate analysis of the user's real-time interactive intentions.
[0052] In actual applications, relying on such a multimodal collaborative processing mechanism, the embodiments of the present application can synchronously perceive and understand multiple information such as images, videos, audio, text, etc., provide high-quality input data for personalized memory processing, thereby more comprehensively understanding and remembering user input, and then more accurately providing personalized services to meet the user's real-time interaction intentions.
[0053] In an embodiment of the present application, the information to be stored includes at least: key content information associated with the target user in the multimodal information. The information to be stored is key content information that is further extracted from the collected multimodal information, has certain value and is associated with the target user. It focuses on those parts that are important for subsequent memory storage and serving users, such as object features that may be closely related to the user's current needs in image information, key voice content that can reflect the user's identity or current intentions in audio information, and parts of text information that involve the user's core demands and key entity relationships. After screening and extraction, these contents will serve as the main objects subsequently stored in the memory module. They are the key link in the entire information processing process, which determines which information can be retained for a long time for further analysis and response to user interaction instructions.
[0054] Furthermore, optionally, the information to be stored is extracted using a multi-task learning model based on the information extraction model. The multi-task learning model enables the information extraction model to simultaneously process multiple related tasks (e.g., performing named entity recognition and relationship extraction in text information), fully utilizing the correlation between different tasks to improve model performance, thereby more accurately identifying and extracting key content from complex multimodal information, and ensuring the quality and effectiveness of the information to be stored.
[0055] In step 201, multimodal information of the target user is collected through various channels such as corresponding input interfaces and sensors. For image information, visual image content may be obtained using image acquisition devices such as cameras; for audio information, sound data is recorded with the help of audio acquisition devices such as microphones; and text information is mainly collected from text content entered by users in the interactive interface through keyboard input, voice-to-text conversion, etc. At the same time, it is also necessary to ensure the synchronization and relevance of the collection of these different modal information, such as knowing which image a certain audio segment corresponds to, which text segment is related to it, etc., so as to carry out unified processing and analysis later and comprehensively capture the complete information of the user at the same time or in a specific interactive scenario. This lays the foundation for a deeper understanding of users, analysis of their interactive intentions, and corresponding information storage and processing. Only by comprehensively and accurately collecting multimodal information can valuable content be mined to serve users.
[0056] In step 202, the collected multimodal information from different modalities is first input into the information extraction model. This model operates based on a multi-task learning model. For example, when processing the text modality, it simultaneously performs tasks such as named entity recognition (identifying key entities such as names of people, places, and organizations in the text), relationship extraction (analyzing the relationships between different entities), and sentiment analysis (determining the emotional tendencies implied by the text). For the image modality, computer vision-related technologies are used to extract features of key objects in the image and identify specific scenes. For the audio modality, key acoustic features of speech, speaker voiceprint characteristics, and the categories of ambient sounds are extracted. Then, the results of the analysis of each modality are combined to select key content information that is closely related to the target user and of great value to subsequent services, and output as the information to be stored. In this way, the massive and complex multimodal information is refined, avoiding the storage of excessive redundant data in the memory module, improving the efficiency of subsequent storage and retrieval, and the effectiveness of information utilization, and ensuring that the stored information is key content that serves to accurately understand users and respond to interactive instructions.
[0057] In step 203, the system calculates the importance and timeliness scores of the extracted information to be stored using a preset evaluation mechanism. Factors determining the importance score may include the degree of relevance of the information to the user's current interaction intent and its value to understanding the user's long-term behavior patterns. For example, information directly related to the core needs clearly stated by the user will have a higher importance score. The timeliness score considers the time sensitivity of the information. For example, information that appears frequently recently and is closely related to the current interaction scenario will have a higher timeliness score, while information that is outdated or from a long time ago and not strongly relevant will have a lower timeliness score. Then, based on these two scores, the information to be stored is stored in the short-term memory module according to certain rules. For example, information with both high importance and timeliness scores is prioritized. Information with lower scores can be selectively stored or temporarily not stored based on the capacity of the short-term memory module. This ensures that the short-term memory module only stores currently critical and time-sensitive information, facilitating rapid response to subsequent interactive retrieval needs. In this way, the information storage of the short-term memory module is managed reasonably, allowing the short-term memory module to focus on the information that is most helpful in understanding user interactions and retrieving related content in the near future, so that valuable content can be quickly located when memory retrieval is based on real-time interaction intentions. At the same time, the storage situation is dynamically adjusted through the scoring mechanism to maintain the effectiveness and practicality of the short-term memory module information.
[0058] In step 204, the dynamic transfer strategy is a set of rules that comprehensively considers multiple factors and is used to determine which information in the short-term memory module should be transferred to the long-term memory module to become key information for long-term storage. These factors can include the frequency of information (such as information that frequently appears and is always relevant to the user's key needs), representativeness of the user's overall behavior pattern (information that can reflect the user's long-term preferences and habits), and versatility in different interaction scenarios (information applicable to a variety of common interaction scenarios). According to this dynamic transfer strategy, the system regularly screens and evaluates the information stored in the short-term memory module or when specific trigger conditions are met (such as when the accumulation of a certain type of information in the short-term memory module reaches a certain threshold), extracts key information that meets the conditions, and transfers it to the long-term memory module, achieving a reasonable transition and effective retention of information from the short-term to the long-term, and providing a deep and valuable data foundation for long-term memory retrieval. In this way, through dynamic transfer, the long-term memory module can continuously accumulate key information that has long-term value to users, and build a more complete user interaction knowledge base, which is convenient for providing reliable and rich historical data support when it is necessary to deeply explore the user's past situation, fully understand the user's long-term behavioral characteristics, and respond to complex interaction instructions in the future, thereby improving the quality and accuracy of the system's response to users based on memory retrieval.
[0059] In summary, steps 201 to 204 above, the collection of multimodal information and the subsequent extraction and storage of stored information work closely together to serve the memory retrieval based on real-time interactive intentions and the realization of personalized services for users by the entire system.
[0060] For example, the following is a detailed introduction to the above multimodal information processing solution:
[0061] Text information processing: Use the improved BERT model for semantic understanding and leverage the multi-head self-attention mechanism to capture contextual relationships in the text, thereby accurately grasping the meaning of the text and providing basic text semantic analysis for subsequent processing.
[0062] Non-text information processing: For image information, specific modal encoders (such as ResNet, ViT, etc.) are used to extract features and convert them into a form that can be integrated with other modal information so that they can interact and process in a unified semantic space, achieving effective feature extraction and preliminary understanding of different types of non-text information.
[0063] Multi-task learning framework: This framework integrates named entity recognition (NER), relation extraction (RE), and event extraction (EE). Within this framework, conditional random fields (CRFs) and bidirectional LSTM networks work together to identify key entities, graph neural networks (GNNs) are used to model relationships between entities, and a hierarchical attention mechanism is employed to capture event structure. This advanced combination of technologies allows for comprehensive and accurate extraction of key information elements from user interaction information. Furthermore, these extractors undergo domain-specific adaptive training to better meet the information extraction needs of diverse fields, improving both accuracy and effectiveness.
[0064] For example, pre-built knowledge graphs and external knowledge bases are used to inject knowledge into language models, enabling them to better understand professional terminology and common sense, addressing the knowledge deficiencies of simple language models and improving their ability to understand complex semantics. A contrastive learning mechanism is also introduced, enabling the model to discern subtle semantic differences by constructing positive and negative sample pairs for comparative analysis. This further optimizes the accuracy and depth of semantic understanding, allowing for a more accurate grasp of the underlying meaning and potential intent of user interaction information.
[0065] This embodiment of the application further implements an adaptive quality control mechanism, assigning reliability scores to extraction results through a confidence assessment module and combining it with a multi-round verification mechanism to strictly control the accuracy of the extracted information. For information with low confidence or uncertainty, an interactive clarification mechanism is activated, asking users precise questions to obtain supplementary information. This ensures that the information that ultimately enters the subsequent processing flow is of high quality and reliability, reducing the impact of erroneous and ambiguous information on the system.
[0066] Here, the extracted information will be organized into a structured knowledge representation form, using a hierarchical semantic framework to encode entity, relationship and event information into vector form, and attach metadata such as timestamps and sources. This makes the information not only have a clear semantic structure, but also contains rich context and source information, providing a solid, orderly and easy-to-operate foundation for subsequent memory storage and retrieval, making it easier for the system to efficiently manage and utilize this information, and better serve personalized memory processing and subsequent user interaction responses.
[0067] Through the close coordination and organic combination of the technologies in the above-mentioned links, it is possible to accurately understand and extract key information from multimodal user interaction information, and continuously optimize and improve this process. As time goes by, continuous model optimization and knowledge accumulation will further enhance the system's understanding ability, provide users with higher quality and more personalized service experience, meet users' diverse needs in different scenarios, play a key role in the interaction process between users and the system, and improve the overall interaction effect and user satisfaction.
[0068] In the embodiment of the present application, the dynamic transfer strategy is dynamically adjusted using a reinforcement learning algorithm. Reinforcement learning is a machine learning method, the core of which is that the intelligent agent (in this context, it can be understood as the relevant module or mechanism for executing the dynamic transfer strategy) interacts with the environment and learns what actions to take based on the reward signal fed back by the environment to maximize the long-term accumulated rewards. In the embodiment of the present application, the dynamic transfer strategy is dynamically adjusted using a reinforcement learning algorithm, which means that the system regards the process of transferring information from the short-term memory module to the long-term memory module as the behavioral decision-making process of the intelligent agent. The environment is a comprehensive situation such as the entire user interaction scenario and the existing information storage status of the system. For example, after a piece of information in the short-term memory is transferred to the long-term memory module, if the subsequent response to the user interaction instruction based on this information achieves a good effect (such as improving the accuracy of the response, user satisfaction, etc.), then the system will give corresponding "reward" feedback, otherwise if the effect is not good, a lower "reward" or even "penalty" will be given.
[0069] In this way, the dynamic transfer strategy can continuously optimize and adjust itself based on actual interaction feedback. It does not use fixed rules to determine which information is transferred from short-term memory to long-term memory. Instead, as the system and user interaction practices increase, it can adaptively discover which types of information are more valuable for long-term memory accumulation, better meeting the diverse and dynamically changing interaction needs of users. For example, at the beginning, a certain type of information may be considered unimportant and does not need to be stored long-term. However, as the reinforcement learning process discovers that it frequently plays a key role in different interaction scenarios, the strategy will be adjusted to incorporate it into the long-term memory module, so that the key information stored in the long-term memory module always remains in the most valuable state for user interaction, improving the flexibility and effectiveness of the entire memory retrieval system.
[0070] In an embodiment of the present application, the storage mechanism in the short-term memory module is a temporary storage that depends on the scoring priority sorting method. It is understandable that the storage mechanism adopted by the short-term memory module is to temporarily store information based on the scoring priority. As mentioned above, the importance score and timeliness score will be calculated for the stored information. Here, the priority of each information is determined based on these two scores or a combination of other relevant scoring indicators. For example, information with high importance and strong timeliness will be prioritized, and then the information will be stored in the short-term memory module in order from high to low priority. And this storage is temporary and will be subject to the restrictions of a preset time period. Within a certain time range, this information is retained to facilitate the system to quickly retrieve recent content that is closely related to the current interaction. Once this time period is exceeded or due to capacity limitations caused by the continuous influx of new information, low-priority information may be gradually replaced or deleted.
[0071] A temporary storage mechanism based on scoring and prioritization enables the short-term memory module to focus on the most valuable information at the moment, ensuring its timeliness and relevance to the current user interaction. This efficiently manages the short-term memory module's storage space, preventing unrestricted data accumulation. This allows the system to quickly locate the most critical recent content when performing memory retrieval in response to real-time interaction intent, improving retrieval efficiency and timely response, and better serving user needs in the current interaction scenario.
[0072] In an embodiment of the present application, the storage mechanism in the long-term memory module is a classified compressed storage based on the knowledge graph structure. The long-term memory module uses a method based on the knowledge graph structure to store information. The knowledge graph is a structured representation that constructs various types of knowledge into a network through nodes and edges. In this application, the system will classify the key information transferred from the short-term memory module according to its inherent semantic relationships, category attributes, etc. For example, information related to user interests and hobbies will be classified into one category, and information related to the user's work field will be classified into another category. Then, the nodes in the knowledge graph are used to represent these different categories of key information, and edges are used to describe the associations between them, such as causal relationships, affiliations, etc. At the same time, in order to store and manage this information more efficiently, a certain degree of compression processing will be performed to remove some redundant data representations or use a more compact data structure to store the information in the knowledge graph, saving storage space while ensuring the integrity of the information.
[0073] Storage based on the knowledge graph structure helps to sort out the logical relationships between complex user interaction information, and can clearly show how different information is related to each other, making it easier for the system to grasp the user's overall behavior patterns, long-term interests and preferences from a macro perspective. In addition, classified storage can make memory retrieval more targeted. When searching based on real-time interaction intentions, relevant information can be quickly located in the corresponding category. Compressed storage solves the storage resource pressure problem that may be caused by the continuous growth of long-term memory data, ensuring that the long-term memory module can continuously and stably accumulate and save key information that has long-term value to users, and provide strong knowledge support for the subsequent comprehensive and in-depth understanding of users and response to interaction instructions.
[0074] For example, the long-term memory (LTM) subsystem utilizes a knowledge graph structure based on a graph neural network (GNN), combined with a vector database for efficient information storage and retrieval. The system uses the BERT model to encode input information, compresses information dimensions through knowledge distillation techniques, and establishes relationships between information using a graph attention network (GAT). To process large amounts of information, the system employs a distributed storage architecture, combined with the LSH (locally sensitive hashing) algorithm for fast approximate nearest neighbor search.
[0075] Overall, the reinforcement learning adjustment mechanism of the dynamic transfer strategy and the respective storage mechanisms of the short-term and long-term memory modules ensure the rational flow, efficient storage and effective utilization of information in the system from different angles, and jointly help achieve accurate memory retrieval and high-quality user interaction services.
[0076] As an optional embodiment, in 203, storing the information to be stored in the short-term memory module according to the importance score and the timeliness score can be implemented as follows:
[0077] A short-term memory model is used to determine the information importance and timeliness of the information to be stored to obtain a real-time importance score of the information to be stored in the current interactive session; the information to be stored is prioritized based on the real-time importance score to obtain a real-time sorting result; based on the real-time sorting result, a dynamic storage space corresponding to the information to be stored in the short-term memory module is adaptively configured; and the information to be stored is stored in the corresponding dynamic storage space in the short-term memory module as short-term storage information.
[0078] Specifically, importance assessment involves the short-term memory model analyzing the relevance of the information to be stored to the user's current interaction intent. For example, if the information to be stored directly relates to the user's current and explicit question, need, or task, its importance score will be increased accordingly. Furthermore, the information's potential value to the user's long-term behavior patterns is considered. For example, recurring information related to the user's core business or interests, even if it's not central to the current interaction, will be assigned a higher importance score due to its long-term value.
[0079] Timeliness can be assessed primarily based on the time the information was generated and its timeliness within the current interaction context. Information generated recently and closely related to the current interaction receives a higher timeliness score, while information older and less relevant to the current interaction receives a lower timeliness score. For example, in a real-time conversation, a key term or event-related information that the user just mentioned has high timeliness; conversely, insignificant interaction details from a few days ago have lower timeliness. By combining these two dimensions, a real-time importance score is calculated for each piece of information to be stored within the current interaction session. This score quantitatively reflects the current value of the information.
[0080] Then, based on the calculated real-time importance scores, all information to be stored is sorted from highest to lowest, forming a priority sequence. Information with higher scores is placed first, indicating that it has a higher priority in the current interaction and needs to be processed and stored first so that it can be quickly retrieved and used later. Information with lower scores is placed later, indicating that it is relatively less urgent and less important in the current interaction. This ranking guides subsequent storage space allocation and information storage, ensuring that the short-term memory module focuses on the most important information first.
[0081] Next, the storage space of the short-term memory module is not fixed, but is dynamically adjusted according to the priority of the information. For the information to be stored at the front and with high priority, a relatively large and more accessible storage space will be allocated to ensure that this critical information can be quickly stored and retrieved. For example, it may be stored near the front of the storage area or stored using a more efficient storage structure. As the priority decreases, the storage space allocated to the information to be stored will be reduced accordingly, and the storage location may also be relatively far back or a more compact storage method may be used. This adaptive storage space configuration mechanism can make full use of limited short-term memory resources, ensuring that at any time, the short-term memory module stores the information most needed for the current interaction, while also preventing low-priority information from taking up too much valuable space, thereby improving space utilization and the rationality of information storage.
[0082] Finally, according to the previously determined dynamic storage space allocation scheme, the information to be stored is stored one by one in the corresponding location in the short-term memory module, completing the short-term storage process of the information. At this point, this information officially becomes short-term storage information, available for rapid retrieval and use by the system in the current interactive session and subsequent possible related interactions, providing timely and targeted data support for memory retrieval based on real-time interactive intent. Furthermore, as new information to be stored continues to arrive and the interaction process progresses, the short-term memory module will continue to update, replace, and reorder the information according to the above steps, always maintaining the timeliness and importance of its stored information to adapt to the ever-changing user interaction needs.
[0083] Through this series of steps, the short-term memory module can efficiently manage the information to be stored, maximize its support for the current interaction under limited space and time conditions, and provide strong guarantees for the efficient operation of the entire system and accurate response to user interaction instructions.
[0084] For example, to improve processing efficiency, the above steps can employ a hierarchical attention network (HAN) to perform preliminary information screening and importance scoring, and an LSTM network to capture temporal dependencies. The capacity of the short-term memory module is dynamically adjusted through an adaptive mechanism within the short-term memory model, automatically optimizing storage space based on the importance and timeliness of the information.
[0085] Optionally, the corresponding dynamic storage space in the short-term memory module is a dynamic buffer. This means that the size of the buffer and the content of the stored information are not fixed, but are dynamically adjusted based on the system's operating state and the relevant characteristics of the information. It acts as a flexible temporary storage area for information extracted from multimodal information, evaluated for importance and timeliness, and requiring short-term retention. This facilitates subsequent rapid memory retrieval based on real-time interaction intent.
[0086] The short-term memory module is provided with an interactive window mechanism for maintaining and managing the dynamic buffer. This mechanism is used to determine which information in the dynamic buffer needs to be focused on, retained, and when to update or replace the information. For example, it can be set to focus only on the interactive information in the most recent period of time. As new information continues to enter the dynamic buffer, the window will slide accordingly, moving earlier, relatively less important or outdated information out of the window range (that is, removing it from the dynamic buffer), and at the same time incorporating new recent interactive information into the window range (that is, storing it in the dynamic buffer), thereby achieving effective management of the dynamic buffer information and always keeping the information in the buffer in the state with the strongest relevance to the current interaction and the best timeliness.
[0087] For example, the short-term memory (STM) subsystem utilizes a modified Transformer architecture and a multi-head attention mechanism to process real-time information. This subsystem maintains a dynamic buffer and manages recent interactive information using a sliding window mechanism. Specifically, the STM subsystem utilizes a modified Transformer architecture, which inherently offers advantages such as strong parallel computing capabilities and the ability to effectively capture long-range dependencies. This architecture has been improved and applied to the STM subsystem, making it more adaptable to the STM module's requirements for rapid information processing and dynamic management. For example, it can more efficiently process real-time information, rapidly encoding and integrating new information entering the dynamic buffer, allowing it to be quickly converted into a form suitable for subsequent retrieval and analysis, providing a powerful computational architecture to enable the system to promptly respond to user interactions. Within this modified Transformer architecture, a multi-head attention mechanism is employed to process real-time information. This multi-head attention mechanism can simultaneously focus on different parts of information and their relationships from multiple perspectives, making it particularly suitable for real-time information, which is often complex and requires rapid extraction of key content. It can mine the correlations between different elements in instant messages, distinguishing between important and relevant content and less important content. This helps the system better understand the key intent of the user's interaction. It then stores the processed information in a dynamic buffer, facilitating flexible retrieval and utilization under the interaction window mechanism, further improving the short-term memory subsystem's processing and management of instant messages. Specifically, a sliding window mechanism manages recent interactions. Similar to the interaction window mechanism mentioned above, it uses a sliding window-like mechanism to determine which recent interactions fall within the window (i.e., to be retained in the dynamic buffer). For example, the window size can be set to the past 10 minutes or 20 interactions (depending on the application scenario and system performance). As time progresses and new interactions are generated, the window slides accordingly, continuously updating the dynamic buffer to ensure that it always stores the most recent and most relevant key information. This allows the short-term memory subsystem to quickly locate currently useful data during memory retrieval and other operations, enhancing the timeliness and accuracy of the system's response to user interactions.
[0088] In summary, the dynamic buffer in the short-term memory module and the corresponding interactive window (sliding window) mechanism, combined with the improved Transformer architecture and multi-head attention mechanism, together build an efficient, flexible short-term memory management system that can dynamically adapt to changes in user interactions. This plays a vital role in enabling the entire system to accurately understand user intent, quickly retrieve memory information, and provide appropriate responses.
[0089] Specifically, the interaction window mechanism first needs to be initialized. This includes determining the size and initial position of the window. The window size can be set based on factors such as system performance, the frequency and characteristics of user interactions, and other factors. For example, in a high-frequency interaction scenario, such as a real-time customer service chat system, the window size might be set to cover the most recent 10-20 interactions to ensure that sufficient recent context is captured. The initial position typically starts when the system begins processing user interaction information, at which point the window covers the initial interaction information.
[0090] When new interactive information is generated, the system will determine whether this information should be included in the interactive window. The basis for judgment is mainly the timeliness and importance of the information. In terms of timeliness, newly generated information meets the "recent" standard in time and is usually considered for inclusion in the window. Importance is measured by the information importance score mentioned earlier. For example, if the new information is related to the core needs of the current user (such as the product details that the user is asking about), its importance score is higher and it will be included in the window first. This information will be stored in the dynamic buffer of the short-term memory module. The dynamic buffer serves as the actual storage area of the interactive window mechanism and provides storage space for the information included in the window. For example, after processing, the new text information is stored in the corresponding position of the dynamic buffer in the form of a vector. The key features extracted from the image or audio information will also be stored here.
[0091] As new information is continuously included in the window, the window will slide. The sliding method is similar to an observation box moving on the timeline. For example, if the window size is set to the most recent 10 interactive messages, when the 11th message is generated and included in the window, the oldest message will be moved out of the window. This process is dynamic and automatic, ensuring that the window is always focused on the most recent interactive information. During the sliding process, the system will re-evaluate the priority of the information in the window based on the importance and timeliness of the information. For example, a piece of information that was originally at the edge of the window and had low importance may increase its priority as new information is included and time passes because of its increased relevance to the new information (such as the new information mentions related content), making its storage location in the dynamic buffer more conducive to retrieval.
[0092] When the system needs to retrieve memories based on real-time interaction intent, the information in the dynamic buffer managed by the interaction window mechanism comes into play. The system prioritizes information within the window because it is recent, filtered, and more likely relevant to the current interaction intent. For example, when answering a user's follow-up question about a product, previous discussions about the product stored in the window are quickly retrieved as a reference for answering the question.
[0093] If the information within the window doesn't meet the retrieval requirements, the system might expand the search scope based on a dynamic transfer strategy, potentially seeking relevant information from long-term memory. However, the interactive window mechanism ensures that the system can, in most cases, first respond quickly using the most timely and relevant short-term information.
[0094] The interactive window mechanism can also adjust the window size based on the system's operating state and user interaction. For example, in scenarios where user interaction is relatively simple and information relevance is not strong, the window can be appropriately reduced to reduce unnecessary information storage and improve system efficiency. Conversely, in complex interactive scenarios, such as in-depth discussions involving multiple topics, the window may be expanded to ensure that sufficient contextual information is retained, providing more data support for the system to fully understand user intent. This ability to dynamically adjust the window size allows the interactive window mechanism to better adapt to different user interaction patterns.
[0095] As an optional embodiment, in 204, transferring the key information stored in the short-term memory module to the long-term memory module according to the dynamic transfer strategy can be implemented as follows:
[0096] A dynamic priority queue is established based on the real-time sorting results through a memory integration model; short-term storage information in the dynamic priority queue whose priority meets the transfer conditions indicated in the dynamic transfer strategy is obtained according to the memory update cycle or update event indicated in the dynamic transfer strategy; and the short-term storage information that meets the transfer conditions is transferred from the short-term memory module to the long-term memory module.
[0097] For example, the embodiment of the present application introduces a transformer-based attention mechanism to achieve bidirectional information flow between the above two memory systems, ensuring the coherence and integrity of memory.
[0098] For example, to prevent catastrophic forgetting, the system also integrates the Elastic Weight Consolidation (EWC) algorithm and Progressive Neural Network (PNN) technology to maintain model plasticity while preserving existing knowledge. Through this complex combination of technologies, the memory storage module achieves efficient information management and long-term knowledge accumulation.
[0099] Specifically, the memory integration and update module is one of the core components of the entire personalized memory system. It mainly adopts a multi-level neural network architecture and innovative algorithm mechanism to achieve dynamic memory management.
[0100] Specifically, first, an attention mechanism based on the Transformer architecture is used to evaluate the relevance and importance between new and old memories. Through the multi-head self-attention mechanism, the system can automatically identify the semantic associations between different memory fragments and establish a dynamic association weight matrix. This mechanism enables the system to accurately capture the intrinsic connections between memories, providing a basis for subsequent integration. Secondly, an improved memory compression algorithm is used in the memory integration process. This algorithm merges semantically similar memory information through a learnable compression matrix while retaining key information features. This compression mechanism can not only effectively reduce storage space usage, but also improve the efficiency of memory retrieval. The system adopts a structure based on variational autoencoder (VAE) to achieve efficient encoding and decoding of memory information.
[0101] Regarding update strategies, the module implements a priority-based memory update mechanism that prioritizes memories based on multiple dimensions, including frequency of use, timeliness, and importance. For example, frequently retrieved memories are more critical to the current interaction scenario and therefore receive a higher frequency score. Memories that are closely related to the current user interaction and are relatively recent receive a higher timeliness score. Furthermore, memories that address core user needs and key topics receive an advantage in importance. Based on these evaluations, the system constructs a dynamic priority queue. The order of memories in this queue changes in real time, adjusting dynamically based on changes in the memory's performance across various dimensions, constantly reflecting the importance of the memory to the current system operation. This mechanism prioritizes the retention and updating of higher-priority memories, as they are more crucial for the system to accurately understand user intent and respond to interaction commands. Lower-priority memories, on the other hand, face the possibility of being consolidated or eliminated to free up space and resources for more valuable memories. To better track the temporal characteristics of memory, the system uses an improved LSTM network. LSTM networks excel at processing data with time series properties and can capture the chronological relationships and changes in memory at different time points, providing support for accurately determining characteristics such as the timeliness of memory. At the same time, reinforcement learning algorithms are combined to optimize update strategies. Reinforcement learning uses continuous interactive feedback between the system and the environment (which can be understood here as user interaction scenarios, etc.) to adjust update strategies based on the actual effects of memory updates (such as whether retrieval efficiency has been improved or whether it is more conducive to responding to users). This makes priority judgments and memory retention, integration, and elimination operations more in line with actual application needs, continuously optimizing the rationality and effectiveness of memory updates.
[0102] To ensure the consistency and integrity of memory, the module also implements a memory conflict detection and resolution mechanism (Memory Conflict Resolution). When new memory information conflicts with existing memory, the system will initiate the corresponding processing flow. First, semantic similarity calculation is used to measure the semantic similarity between the new memory and the existing memory to determine whether there are contradictions or inconsistencies in their meaning. At the same time, combined with contextual analysis, the specific interaction scenarios in which these memory information are located, the previous and subsequent related content, etc., are considered to more comprehensively determine whether there is a conflict. In this process, a BERT-based semantic understanding model is used. BERT's powerful semantic representation ability can accurately analyze the semantics of memory content and accurately identify possible semantic conflicts. In addition, in conjunction with knowledge graph technology, the existing knowledge structure and entity relationship information in the knowledge graph are used to assist in judging the rationality of memory information in the entire knowledge system, ensuring the coherence and accuracy of the memory system, and avoiding deviations in system understanding and response due to conflicting memory information.
[0103] In this way, the memory integration and update module enables dynamic memory management and optimization, ensuring efficient system operation while maintaining memory quality and usability. This multi-layered technical architecture enables the system to flexibly process and organize memory information, similar to the human brain, providing reliable memory management support for the entire personalized memory system.
[0104] As an optional embodiment, in the above steps, transferring the short-term stored information that meets the transfer condition from the short-term memory module to the long-term memory module can be implemented as follows:
[0105] receiving short-term storage information satisfying the transfer condition from the short-term memory module;
[0106] Through the long-term memory model, the received short-term storage information is embedded into the long-term memory map related to the target user, and based on the embedding result, the received short-term storage information is stored in the corresponding long-term storage space in the vector database as long-term storage information.
[0107] In the above steps, first, it is necessary to clarify that short-term storage information is information that has been screened, evaluated, and stored according to certain rules in the short-term memory module. This information has been considered in the short-term memory module based on dimensions such as importance and timeliness. When they meet specific transfer conditions (for example, according to the dynamic transfer strategy, the frequency of occurrence of the information reaches a certain threshold, it is representative of the user's long-term behavior patterns, etc.), they will be selected and transferred from the short-term memory module to the long-term memory module. This receiving process is to obtain this short-term storage information that has met the transfer requirements and prepare it for subsequent integration into the long-term memory system.
[0108] Next, the long-term memory model is used to process the received short-term stored information. The long-term memory model plays a key role here, embedding the short-term stored information into the long-term memory graph related to the target user. The long-term memory graph is a structured knowledge representation that can display the complex relationships between different pieces of information. Like a network, it clearly displays the various knowledge nodes surrounding the target user and their connections. During the embedding process, specific algorithms and techniques are used to "place" the short-term stored information in this graph structure in an appropriate form, making it part of the graph and establishing an intrinsic connection with other existing relevant information.
[0109] Then, based on the embedding results, the received short-term information is stored in the corresponding long-term storage space within the vector database, officially transforming it into long-term storage information. A vector database is a database suitable for storing vector-formatted data, enabling efficient management and retrieval of stored data. Storing information in the form of vectors in the corresponding long-term storage space facilitates rapid retrieval based on vector features while ensuring the orderly storage of long-term memory information. This makes the storage structure of the entire long-term memory module more rational, providing a good foundation for subsequent operations such as memory retrieval.
[0110] For example, the long-term memory (LTM) subsystem uses a knowledge graph structure based on a graph neural network (GNN), combined with a vector database to achieve efficient information storage and retrieval. The system uses the BERT model to encode input information, compresses information dimensions through knowledge distillation technology, and then establishes associations between information through a graph attention network (GAT).
[0111] Specifically, the long-term memory (LTM) subsystem uses a knowledge graph structure based on a graph neural network (GNN), which means it leverages the characteristics of GNNs to process and analyze information in the knowledge graph. Through the information transfer mechanism between nodes and edges, GNNs can automatically learn the complex relationships between nodes in the knowledge graph (representing different information elements) and explore deep-level association patterns. For example, for different types of nodes such as a user's long-term interests and hobbies, past interaction history, etc., GNNs can analyze which interests and hobbies have mutually reinforcing or related relationships, and which interaction histories have a significant impact on the current user's behavior patterns. As a result, the information in the knowledge graph is no longer isolated, but rather interconnected and logically integrated, providing strong structural support for the system's comprehensive and in-depth understanding of users and accurate retrieval of memory information.
[0112] For example, the BERT model is first used to encode the input information. With its powerful language understanding capabilities, the BERT model can transform information in various textual forms into vector representations with rich semantic features, laying the foundation for subsequent processing. Then, knowledge distillation is used to compress the information dimension. Knowledge distillation can refine the high-dimensional vector information encoded by BERT, removing some redundant information dimensions. While preserving key semantic features as much as possible, it reduces the amount of data, improves the efficiency of information processing and storage, and avoids the storage and computational burdens caused by excessive data dimensionality.
[0113] To process large amounts of information, the system adopts a distributed storage architecture, combined with a locality-sensitive hashing (LSH) algorithm to achieve fast approximate nearest neighbor searches. In practice, as user interactions increase, the amount of information required to store in the long-term memory module becomes enormous. Using traditional centralized storage methods can easily lead to problems such as insufficient storage capacity and slow retrieval speeds. A distributed storage architecture, on the other hand, stores large amounts of information across multiple storage nodes. Through a rational distributed management strategy, this enables efficient storage and management of large amounts of information. Each storage node can work in parallel, sharing the burden of storage and retrieval. This is like dividing a large library into multiple smaller branches, each responsible for managing a portion of the books. This allows users to search across multiple branches simultaneously when searching for books (or information). This improves overall processing efficiency and ensures the system's stable operation despite the massive amount of long-term memory information.
[0114] Incorporating the Locality Sensitive Hashing (LSH) algorithm, its main purpose is to achieve fast approximate nearest neighbor search. In a long-term memory vector database, when searching for other information most similar to a given piece of information (i.e., nearest neighbor information), traditional precise search methods are computationally intensive and inefficient in large-scale data situations. The LSH algorithm cleverly hashes the vector space, so that similar vectors are assigned to the same or similar hash buckets with a high probability. This allows searches to quickly find approximate nearest neighbor vectors by searching only a few relevant hash buckets, greatly improving retrieval speed and providing an efficient search method for the system to quickly locate relevant information in the long-term memory module, further enhancing the operational efficiency and practicality of the entire long-term memory subsystem.
[0115] As an optional embodiment, determining the timing of transferring short-term stored information is a complex process that involves comprehensive consideration of multiple factors and strategies. If certain short-term stored information appears frequently within a certain time frame, this indicates that this information may be of high importance to the user or is closely related to the user's core needs. For example, in a shopping recommendation system, if a user mentions his interest in a certain brand of electronic products, such as mobile phone models, computer configurations, and other related information many times in a short period of time, the system will consider transferring this frequently appearing information from the short-term memory module to the long-term memory module. Because this high-frequency information is likely to be content that users pay attention to for a long time, it is of great value for understanding users' long-term preferences and intentions.
[0116] To quantify this frequency of occurrence, the system can set a frequency threshold. When the number of occurrences of short-term stored information reaches or exceeds this threshold, a transfer operation is triggered. The threshold setting can be determined based on factors such as the system's specific application scenario, user behavior patterns, and storage resources. For example, in a simple question-and-answer system, if information related to a certain topic appears six times out of ten interactions, the transfer threshold is reached and the information is transferred to the long-term memory module so that it can be better utilized in subsequent questions and answers to provide more accurate answers.
[0117] Short-term stored information has a time-sensitive nature, but over time, some information, while less time-sensitive, may demonstrate long-term value. For example, a user mentions a new technical requirement during a product consultation. While the timeliness of this requirement may gradually decrease after the consultation concludes, if the system determines that this technical requirement may still play a key role in the user's future product selection or usage, then it possesses long-term value. The system determines the timing of transfer based on dynamic monitoring of the information's timeliness and its estimated long-term value.
[0118] Several factors can be considered when determining the long-term value of information. These include whether the information is relevant to the user's long-term goals. For example, a grammatical point repeatedly mentioned by a user while learning a foreign language is relevant to this long-term goal. Does the information relate to the user's basic preferences and habits, such as frequently mentioned reading preferences or exercise habits? Or, does the information relate to the system's core service functions? For a financial services system, information about a financial management method that users frequently inquire about has long-term value. When the system determines that short-term stored information has sufficient long-term value, it transfers it to the long-term memory module at the appropriate time.
[0119] The timing of this shift can be determined by observing changes in user interaction patterns. For example, when a user's interaction shifts from the initial information-gathering phase to the in-depth comparison and selection phase, key information previously acquired during the information-gathering phase, such as the basic product categories and the user's primary functional parameters, can be transferred from short-term memory to long-term memory. This information will continue to play an important supporting role in subsequent interactions, and this shift in interaction patterns suggests the long-term usefulness of this information.
[0120] The timing of this transfer is determined based on the system's own tasks and functional requirements. If the system's primary task is to provide users with personalized services, such as personalized recommendations and personalized learning path planning, then the transfer should occur when short-term stored information is important for the long-term planning of completing these tasks. Taking a personalized learning system as an example, when the short-term stored user learning progress information and knowledge mastery reaches a certain level and can provide a reference for formulating long-term learning plans, this information will be transferred to the long-term memory module, allowing the system to better track the user's long-term learning trajectory and provide more effective learning suggestions.
[0121] The reward-feedback mechanism uses reinforcement learning algorithms to dynamically adjust transfer strategies. The system can provide rewards or penalties based on the actual effect of information transfer. For example, if a piece of short-term stored information is successfully used to provide more accurate services or answers in subsequent user interactions after transferring it to the long-term memory module, such as improving the accuracy of recommendations or increasing satisfaction with answers, the system will provide a positive reward signal. Conversely, if the transferred information does not have a positive effect or causes system performance to degrade, such as increasing retrieval time without providing any real value, negative feedback will be provided.
[0122] By continuously accumulating these feedback signals, reinforcement learning algorithms can optimize dynamic transfer strategies and adjust the timing of transferring short-term stored information. For example, transfer timing might initially be set based on simple frequency rules. However, as reinforcement learning discovers that certain high-frequency but ineffective information is consuming excessive resources after transfer, the strategy can be adjusted to take into account long-term value and actual results, more rationally determining transfer timing so that the information transferred to the long-term memory module truly contributes to the long-term operation of the system and improves service quality.
[0123] As an optional embodiment, the storage information of the short-term memory module and the long-term memory module carries a memory code; the memory code is used to identify the information content type of the corresponding storage information.
[0124] Memory coding plays a key role in identifying the type of information content corresponding to the stored information.
[0125] For the short-term memory module, since it temporarily stores real-time interactions and retrieval information with users within a preset period of time, these information sources are diverse and complex. Memory encoding can quickly distinguish whether the information is interactive discourse in text form, key voice content in audio, or key feature descriptions related to images, etc. For example, "T" represents text, "A" represents audio, and "I" represents image. In this way, when performing retrieval based on real-time interactive intent, the system can quickly locate the corresponding type of information based on the encoding, thereby improving retrieval efficiency.
[0126] The same is true for the key information stored in the long-term memory module. It covers important content from many past interactions, and memory encoding facilitates the classification and management of key information of different content types. For example, when building a knowledge graph based on user interests, encoding can easily filter out information nodes that belong to the same interest and hobby content type, and better sort out the relationship between information. Moreover, when information needs to be integrated or updated, memory encoding can help the system accurately determine which types of information can be merged and which require special processing, ensuring the orderliness of the memory module information and the convenience of subsequent use, helping the system achieve more accurate and efficient memory management and retrieval utilization.
[0127] As an optional embodiment, in the above steps, the received short-term storage information is embedded into the long-term memory map related to the target user through the long-term memory model, and based on the embedding result, the received short-term storage information is stored in the corresponding long-term storage space in the vector database as long-term storage information, which can be implemented as follows:
[0128] The received short-term storage information is encoded to obtain an initial information code; the information dimension of the initial information code is compressed through knowledge distillation processing to obtain an optimized information dimension corresponding to the initial information code; according to the optimized information dimension, the correlation between the initial information codes is established; according to the correlation and the optimized information dimension, the corresponding graph attributes of the initial information code in the long-term memory graph are determined; wherein the graph attributes include at least: the dimension to which the information belongs and the information position; based on the dimension to which the information belongs and the information position, the initial information code is stored in the corresponding long-term storage space in the vector database.
[0129] It's worth noting that after receiving short-term storage information from the short-term memory module, it must first be encoded. This process aims to convert this information of varying forms and content into a specific code form that the computer can better process and recognize, known as the initial information encoding. For example, for short-term storage of textual information, techniques such as word embedding are used to map the text into a vector-like encoding. Image information is then encoded into a corresponding vector after extracting key features. Similarly, audio information is converted into a code representing its features. This creates the initial information encoding, laying the foundation for further processing.
[0130] The initial information encoding may be too high in dimensionality or contain redundant information. This is where knowledge distillation comes in. Based on existing knowledge systems and rules, it refines and compresses the initial information encoding, removing dimensions that are less critical to expressing the core meaning of the information while retaining the most representative and important parts, thereby optimizing the information dimension. This is like extracting the key points from a lengthy article. This method preserves the core content while reducing the data size, facilitating efficient subsequent storage and processing while reducing the computational and storage burden.
[0131] Based on the optimized information dimensions after compression, the inherent connections between the initial information encodings can be analyzed. For example, similarity metrics can be used to determine the degree of similarity between the information represented by different encodings in terms of semantics, features, and so on, thereby establishing connections. If two information encodings are semantically similar, or address different aspects of the same topic, then corresponding connections can be established, eliminating the isolation of information and paving the way for the construction of a complete long-term memory map.
[0132] Based on the previously established relevance and optimized information dimensions, we further clarify the graph attributes of each initial piece of information encoded in the long-term memory graph, such as the dimension to which the information belongs (user interest dimension, behavior dimension, etc.) and the location of the information (specific node position in the graph structure, etc.). This is equivalent to accurately "positioning" the information in the graph so that it can be integrated into the logical structure of the graph.
[0133] Finally, according to the determined information dimension and information location, the initial information encoding is accurately stored in the corresponding long-term storage space of the vector database, making it a long-term storage information, which is convenient for subsequent efficient retrieval and utilization based on the graph structure and vector features, and improves the information storage and management mechanism of the long-term memory module.
[0134] Further optionally, in the above steps, after the information to be stored is stored in the short-term memory module according to the importance score and the timeliness score, or after the key information stored in the short-term memory module is transferred to the long-term memory module according to the dynamic transfer strategy, the following steps may be further implemented:
[0135] 301. Assign corresponding time weights to the short-term storage information stored in the short-term memory module and the long-term storage information stored in the long-term memory module respectively;
[0136] 302. Dynamically adjust the time weight using an exponential decay function; wherein the exponential decay function is constructed based on a dynamic forgetting algorithm of time decay, and the dynamic forgetting algorithm is set according to the Ebbinghaus forgetting curve;
[0137] 303. Monitor the access frequency, recent usage time, and relevance of short-term and long-term stored information to other information;
[0138] 304. Based on the monitoring results, the forgetting assessment model is used to rank the importance of short-term storage information and long-term storage information;
[0139] 305. Perform context relevance analysis on the short-term stored information and the long-term stored information to obtain interactive topic relevance between the short-term stored information and the long-term stored information;
[0140] 306. Determine the short-term stored information and / or long-term stored information to be deleted based on the importance ranking result, the interaction topic relevance, and the time weight.
[0141] It is understandable that, first, the above steps adopt a dynamic forgetting algorithm based on time decay (Time-Based Decay Algorithm). This algorithm draws on the principle of the Ebbinghaus forgetting curve, assigns a time weight to each piece of memory information, and dynamically adjusts the weight value through an exponential decay function. As time goes by, the system will automatically reduce the weight of older information, but retain key knowledge points. Drawing on the principle of the Ebbinghaus forgetting curve, the curve reveals that people forget knowledge in a pattern of first fast and then slow over time. In the system, corresponding time weights are configured for the short-term storage information stored in the short-term memory module and the long-term storage information stored in the long-term memory module, which means that the newly stored information will be given a relatively high weight because it is closer to the user interaction context at the moment and has stronger timeliness. For example, the product details that the user has just inquired about have a larger weight when initially stored.
[0142] These time weights are dynamically adjusted through an exponential decay function. As time goes by, just as people gradually forget the knowledge they learned in the past, the system automatically reduces the weight of older information. However, for information that is a key knowledge point, even if time passes, the reduction in its weight will be relatively small or it will still maintain a certain importance, ensuring that it will not be easily forgotten. For example, information related to core interests and hobbies that users have been paying attention to for a long time, even if it has not been accessed for a long time, due to its critical nature, the weight can still be maintained at a level that is not likely to be deleted. Therefore, in the overall memory management, it not only reflects the impact of time on memory, but also ensures the retention of key content.
[0143] Second, the module integrates a frequency-based importance assessment mechanism (Frequency-Based Importance Assessment). The system builds a multi-dimensional importance scoring model by tracking the access frequency, recent usage time, and correlation with other information for each piece of information. A multi-dimensional importance scoring model is built by tracking the access frequency, recent usage time, and correlation with other information for each piece of information. Information with high access frequency means that it is often called during the user interaction process and is often closely related to the user's current needs, so its importance is naturally higher; the more recent the usage time, the more valuable it is in the current interaction scenario; and the correlation with other information reflects the position of the information in the entire knowledge system. The more correlations there are, the greater the role it may play in understanding the overall user interaction intention. For example, in a knowledge question-and-answer system, a knowledge point that is frequently queried and related to multiple related topics will perform well in these dimensions and its importance score will be higher.
[0144] An improved TD-IDF algorithm, combined with an attention mechanism, ranks information by importance, providing a basis for forgetting decisions. The TD-IDF algorithm excels at measuring the importance of words in a document and, after improvements, is even better suited for assessing the importance of memorized information. The attention mechanism focuses on key components of information and the importance of relationships between different pieces of information. By combining these two, it accurately determines the order of importance of each piece of information, providing a reliable basis for subsequent forgetting decisions. This allows the system to more scientifically determine which information is worth retaining and which can be considered for forgetting.
[0145] Third, a contextual relevance analysis engine (Contextual Relevance Engine) was introduced. This engine uses an improved Transformer architecture to analyze the semantic relevance between information through a self-attention mechanism. The system will retain information that is highly relevant to the user's recent interaction topics, while gradually fading or cleaning up low-relevance content. The system uses a reinforcement learning framework to optimize the forgetting strategy (Reinforcement Learning for Forgetting Strategy). By designing a suitable reward function, the system can continuously learn and adjust the forgetting strategy from interaction feedback, achieving a balance between maintaining the efficiency of the memory system and retaining important information. The Deep Q-Network (DQN) model is used for forgetting decisions, which make decisions based on the temporal characteristics, importance score, and contextual relevance of the information.
[0146] During the interaction between the user and the system, different interaction topics will be formed. Information that is highly relevant to the user's recent interaction topics is critical for understanding the current user's intentions and providing accurate responses. The system will retain this content. For example, in a conversation about travel planning, recently mentioned information such as attractions and transportation will be retained because of their strong relevance to the interaction topic, while content that is irrelevant to the topic or has very low relevance will be gradually faded or cleared to ensure that the memory system focuses on currently useful information. The system uses a reinforcement learning framework to optimize the forgetting strategy. By designing a suitable reward function, it continuously learns and adjusts the forgetting strategy based on actual interaction feedback after information is forgotten or retained (such as whether it improves retrieval efficiency, whether it is more conducive to accurately responding to users, etc.), and finds a balance between maintaining the efficiency of the memory system and retaining important information. The forgetting decision adopts the Deep Q-Network (DQN) model, which comprehensively considers the temporal characteristics of information (i.e., the temporal weight), importance score (from the previous importance assessment mechanism), and contextual relevance. After comprehensive consideration, it makes decisions on which short-term stored information and / or long-term stored information should be deleted, making the entire forgetting process more intelligent and reasonable, in line with the system's operating requirements and user interaction characteristics.
[0147] To sum up, these three aspects work together to build a relatively complete memory information management and forgetting decision-making mechanism, so that the information in the short-term memory module and the long-term memory module can be dynamically adjusted according to its actual value, time factors, contextual relevance and other factors, to ensure the efficiency and effectiveness of the memory system.
[0148] As an optional embodiment, in 306, the short-term stored information and / or long-term stored information to be deleted is determined based on the importance ranking result, the interaction topic relevance, and the time weight, including at least one of the following:
[0149] Setting the short-term stored information and / or long-term stored information whose relevance to the interactive topic within the set time period is lower than the set relevance threshold as the short-term stored information and / or long-term stored information to be deleted;
[0150] Setting the short-term storage information and / or long-term storage information whose importance ranking is lower than the set ranking threshold as the short-term storage information and / or long-term storage information to be deleted;
[0151] The short-term storage information and / or long-term storage information whose time weight is lower than the set time weight threshold is set as the short-term storage information and / or long-term storage information to be deleted.
[0152] In 102, based on the real-time interaction intention, the short-term memory module and the long-term memory module are memory-retrieved to obtain the target retrieval information. Specifically, in the information classification and indexing module, a multi-level technical architecture is adopted to realize an efficient information retrieval system. First, an improved vector database technology is adopted at the bottom layer, and a high-performance vector indexing framework such as FAISS or Milvus is used to convert all information content into high-dimensional vector representations. These vectors are generated by pre-trained Transformer encoders (such as BERT or RoBERTa) to ensure the accurate capture of semantic information. At the classification level, the system organically combines the hierarchical topic model (HTM) and the dynamic clustering algorithm. HTM can automatically mine the hierarchical structure contained in the text topic, sort out a clear context from large topics to subdivided small topics, and make the information category hierarchy clear. At the same time, the dynamic clustering algorithm can adjust the category system in real time based on the newly added information, so that the classification system can always adapt to the changing information content. In specific implementation, an adaptive multi-level classification system is constructed by combining the improved LDA (latent Dirichlet allocation) model with the HDBSCAN clustering algorithm to more accurately classify various types of information, making it easier to quickly locate the corresponding category and find the target retrieval information during retrieval.
[0153] For example, suppose there is a knowledge base of an online learning platform, which stores a massive amount of learning materials, including various types of information such as various subject knowledge, learning methods, and examination information. At the classification level, a hierarchical topic model (HTM) is used. For example, for the subject knowledge section, it can first sort out major themes such as "natural sciences" and "humanities and social sciences", and then subdivide "physics", "chemistry", "biology" and other small themes under "natural sciences", clearly presenting a hierarchical structure so that the knowledge categories are clear at a glance. At the same time, the dynamic clustering algorithm comes into play. When new learning materials are uploaded, such as content related to the emerging interdisciplinary subject "biochemistry", it will adjust the category system in real time based on this new information, and reasonably classify "biochemistry" into relevant levels, so that it can be integrated into the existing classification framework. Specifically, by combining an improved Latent Dirichlet Allocation (LDA) model with the HDBSCAN clustering algorithm, taking learning methods materials as an example, the LDA model can analyze the method themes focused on by different materials, such as "memory methods" and "problem-solving techniques." The HDBSCAN clustering algorithm then accurately clusters the materials based on these themes, building an adaptive multi-level classification system. This allows users to quickly locate the corresponding category when searching for materials related to "problem-solving techniques for physics," efficiently finding the target information.
[0154] To improve retrieval efficiency, a hybrid index structure was designed, combining inverted and vector indices. The inverted index facilitates fast keyword matching, while the vector index handles semantic similarity searches. Furthermore, an attention-enhanced relevance ranking algorithm was introduced, dynamically adjusting retrieval weights based on query context. The system also employs locality-sensitive hashing (LSH) technology to improve retrieval speed for large-scale data through dimensionality reduction.
[0155] To improve search efficiency, a hybrid index structure was designed, integrating an inverted index and a vector index. The inverted index enables fast keyword matching. When specific keywords are included in real-time interaction intent, it quickly filters out information containing these keywords. The vector index focuses on semantic similarity search, using previously generated high-dimensional vector representations to find information with semantically similar query intent. The two complement each other to comprehensively cover diverse search needs. Firstly, an attention-enhanced relevance ranking algorithm is introduced, which dynamically adjusts search weights based on the query context. For example, during a search, if the user's real-time interaction intent focuses on a certain aspect, the relevant information will be weighted higher, ranking it higher in the search results, making the target information more relevant to the user's current needs. Secondly, the use of locality-sensitive hashing (LSH) technology reduces the dimensionality of high-dimensional vectors, effectively accelerating retrieval on large-scale data, preventing inefficiencies caused by large data volumes, and ensuring that the target information can be quickly filtered from a large amount of information.
[0156] For example, in a news information retrieval system, a user's real-time interaction intent is "the latest application results of artificial intelligence in science and technology." The hybrid index structure operates. The inverted index, based on keywords such as "science and technology," "artificial intelligence," and "application results," quickly selects content containing these terms from the vast amount of news information. For example, relevant articles with titles containing these terms are first selected. The vector index, leveraging the high-dimensional vector representations generated from the news text, searches for semantically similar information. For example, articles describing new applications of artificial intelligence in different scenarios can be retrieved even if they don't use identical wording. The two work together to meet multifaceted search needs. Next, an attention-enhanced relevance ranking algorithm takes effect. Because the user focuses on "latest application results," the system prioritizes content related to recent innovative application cases, placing them higher in the search results. Finally, in the face of a large volume of news data (in large-scale scenarios), locality-sensitive hashing (LSH) technology reduces the dimensionality of the news vectors, accelerating retrieval and helping to quickly select the news and information that best meets the user's needs as target information, effectively improving overall search efficiency.
[0157] In terms of index updates, we adopt an incremental update strategy, combined with Bloom filters for rapid duplication checking, and use the Skip-gram model to continuously optimize the word vector space. This allows updates to be made only for newly added or changed information, avoiding the need to rebuild the entire index on a large scale, saving resources and time. Furthermore, combined with Bloom filters for rapid duplication checking, it is possible to quickly determine whether information is already in the index, improving update efficiency. At the same time, the Skip-gram model is used to continuously optimize the word vector space, allowing word vectors to better reflect the semantic relationship between words, further improving retrieval accuracy, and ensuring that the index is in good condition to serve memory retrieval, making the acquisition of target retrieval information more accurate and efficient.
[0158] For example, consider an e-commerce product information indexing system. As new products are constantly added and product information is updated, the index needs to be maintained. Using an incremental update strategy, for example, when a new smartwatch is launched, the system doesn't rebuild the entire product index. Instead, it simply adds the brand, model, features, price, and other information to the corresponding index location, saving significant time and resources. During this process, a Bloom filter is incorporated for rapid duplicate checking. For example, if a merchant mistakenly uploads the same smartwatch information twice, a Bloom filter can quickly determine that the information already exists in the index, avoiding duplicate insertions and improving update efficiency. Simultaneously, the Skip-gram model is used to continuously optimize the word embedding space. For product description terms like "high-definition screen" and "long battery life," the Skip-gram model can analyze large amounts of product text data to learn the semantic relationships between these terms and other related terms (such as "clear display" and "durable battery"), enabling the word embedding to more accurately represent the meaning of the terms. In this way, when a user searches for "watch with clear screen", even if there is no exact matching word in the index, the smart watch can be found through the semantic association of the word vector, further improving the retrieval accuracy and ensuring that the target retrieval information can be obtained more accurately and efficiently, allowing users to quickly find their favorite products.
[0159] To handle multimodal information, the system integrates a cross-modal attention network, enabling unified indexing and retrieval of diverse data types, such as text and images. This means that regardless of the modality of information involved in a user's real-time interaction, the system can search within a unified framework, eliminating modal differences and more comprehensively locating target information relevant to the user's intent, enhancing retrieval completeness and accuracy.
[0160] For example, on a travel guide sharing platform, a user's real-time interactive intent is to find information about the "beautiful spring scenery of West Lake in Hangzhou." The platform contains both travelogues shared by tourists describing the spring scenery of West Lake, such as "West Lake's weeping willows sway, peach blossoms are in full bloom, and the lake's surface shimmers," and photos of the beautiful spring scenery of West Lake taken by everyone, showing images of flowers, willows, and mountains. The system's integrated cross-modal attention network comes into play. When a user initiates a search, it can associate key elements mentioned in the text with corresponding features presented in the image, indexing and searching within a unified framework. Even if the user simply enters the text "West Lake peach blossoms," not only can the travelogue content containing these words be retrieved, but also images of West Lake that focus on the blooming peach blossoms. Similarly, if a user uploads a new photo of West Lake in spring, the system can also analyze the photo content through a cross-modal network, combine it with the existing text description, accurately determine the theme of the photo, and index it reasonably. When a user searches for relevant information in the future, it can fully present various modal information related to it, eliminating the modal differences between text and images, enhancing the completeness and accuracy of the retrieval, and better meeting user needs.
[0161] To fully optimize retrieval performance, the system implements an adaptive caching mechanism and uses the LRU-K algorithm to manage hot data. This effectively caches frequently accessed hot data, allowing it to be retrieved directly from the cache the next time it is retrieved, speeding up retrieval. A distributed storage architecture is used to improve system scalability, enabling it to cope with growing amounts of data and easily handle large-scale memory information storage and retrieval tasks. In addition, a reinforcement learning-based index optimizer is introduced, which can automatically adjust index strategies based on changing query patterns, continuously optimize the index structure to better adapt to actual retrieval needs, further improve retrieval efficiency, and help more efficiently obtain target retrieval information that meets real-time interactive intentions.
[0162] As an optional embodiment, in 102, performing memory retrieval on the short-term memory module and the long-term memory module based on the real-time interaction intention to obtain target retrieval information can be implemented as follows:
[0163] 401. Performing deep semantic analysis on the real-time interaction intention to obtain a real-time search target corresponding to the target user;
[0164] 402. Retrieve first candidate information associated with the real-time retrieval target from the short-term memory module;
[0165] 403. Retrieve second candidate information associated with the real-time retrieval target and / or the first candidate information from the long-term memory module;
[0166] 404. Map the real-time retrieval target, the first candidate context information, and the second candidate information into the same vector space to obtain respective corresponding projection vectors;
[0167] 405. Calculate a first similarity between a first projection vector corresponding to the first candidate information and a reference projection vector corresponding to the real-time retrieval target;
[0168] 406. Calculate a second similarity between a second projection vector corresponding to the second candidate information and the reference projection vector;
[0169] 407. Generate the target retrieval information based on the first similarity and the second similarity; the target retrieval information is obtained based on inference using a multi-hop Bayesian inference network.
[0170] In an embodiment of the present application, the corresponding proportions of the first candidate information and the second candidate information in the target retrieval information are dynamically configured using an attention mechanism; the attention mechanism is associated with a retrieval strategy pre-configured in the retrieval scenario.
[0171] Consider an intelligent learning assistant system. A user enters the real-time interactive intent "Learn about real-life applications of the buoyancy principle in physics." First, the system performs a deep semantic analysis of the intent, identifying the real-time search target as real-life applications of the buoyancy principle. Next, the system retrieves recent notes on buoyancy from the user's physics studies as first-tier candidates, such as notes on their understanding of the buoyancy calculation formula. Then, the system retrieves previously accumulated knowledge points about real-life applications of the buoyancy principle, such as how boats float and how to use swimming rings, from the long-term memory module as second-tier candidates. The real-time search target, first-tier candidates, and second-tier candidates are then mapped into the same vector space to obtain their respective projection vectors. The first and second-tier similarities are then calculated. For example, if someone has recently learned about buoyancy calculations and is exploring its application, the first-tier similarity for the first-tier candidate is high, while the second-tier similarity for commonly used application cases accumulated previously is also high. Finally, based on these similarities, the target search information is inferred using a multi-hop Bayesian inference network. The attention mechanism dynamically configures the ratio of the two. For example, when discussing the principle application scenario at this moment, the system uses the attention mechanism to increase the proportion of the second candidate information (life application cases) in the target search information based on the pre-configured retrieval strategy, and gives priority to displaying specific application content such as boats and swimming rings, so that the presented search results are more in line with the current user's needs to deeply understand the application cases, helping to efficiently obtain useful information.
[0172] In the specific implementation of the above step 102, in terms of retrieval technology, a hybrid retrieval strategy combining vector retrieval (Vector Retrieval) and semantic retrieval (Semantic Search) is adopted. The query content and stored memory information are mapped to the same vector space through the dense passage retrieval (DPR) model to achieve fast similarity matching. At the same time, a BERT-based semantic understanding model is introduced to perform deep semantic analysis on the retrieval content to improve retrieval accuracy. The system also integrates an attention mechanism to dynamically adjust the weights of different retrieval strategies.
[0173] To improve retrieval efficiency, the system implements a hierarchical index structure and employs locality-sensitive hashing (LSH) technology to build a fast retrieval index. When processing large amounts of memorized data, a distributed retrieval architecture is used to increase retrieval speed through asynchronous parallel processing. Furthermore, the system integrates an adaptive caching mechanism that dynamically adjusts the storage level of hot data based on access frequency.
[0174] For example, within an enterprise's internal knowledge management system, an employee might want to search for content related to "recommendations for selecting online marketing channels for new product promotion plans." Using a hybrid search strategy combining vectorized and semantic search, the system first uses a dense paragraph search (DPR) model to map the employee's query and the system's stored product promotion knowledge documents, among other information, into the same vector space. This quickly identifies documents mentioning keywords like "new product promotion" and "online marketing channels," completing a preliminary similarity match.
[0175] At the same time, the BERT-based semantic understanding model deeply analyzes the semantics of the query content, accurately grasps the key semantic requirement of "selecting suggestions", and further filters out documents that truly meet the deep meaning of employees' desire to obtain suggestions, thereby improving retrieval accuracy.
[0176] The attention mechanism dynamically adjusts the weights. If the system determines that semantic retrieval is more in line with the needs at this moment, it will increase its weight and focus on presenting results with high semantic matching.
[0177] A hierarchical index structure, combined with locality-sensitive hashing (LSH) technology, creates a fast search index that can quickly locate relevant documents within large-scale knowledge databases. A distributed search architecture enables asynchronous parallel processing, accelerating overall search speed.
[0178] The adaptive caching mechanism comes into play. If the knowledge related to "online marketing channels" has been frequently accessed recently, it will be upgraded to a higher storage level as hot data. The next time there is a similar search, it can respond faster. Employees can then efficiently obtain accurate and useful content related to online marketing channel selection recommendations for new product promotion plans.
[0179] For example, in an intelligent customer service system, when a user engages in multiple rounds of conversations with customer service regarding product usage issues, the Transformer-based multi-head attention mechanism comes into play. For example, if a user first describes some symptoms of a product failure and then, a few sentences later, mentions the previous operation steps, the system can adaptively focus on this historical information by encoding it into a query, key, and value matrix. This allows the system to connect these historical information even if they are far apart, understand the conversation context, and grasp the overall situation. In the hierarchical semantic understanding framework, the underlying BERT-like pre-trained model encodes the basic semantics of each user sentence. The middle-level graph neural network (GNN) conversation topic tracking mechanism can focus on the topic of discussing product failures, and the top-level attention pooling completes global semantic integration. For example, if a user describes the symptoms of a problem and then raises operational questions, the system can simultaneously capture semantics at different levels to accurately understand the user's intention.
[0180] In long-term conversations with many turns, a dynamic context management mechanism based on a sliding window adjusts the window size and, based on the importance score calculated by the improved REINFORCE algorithm, adaptively retains key content related to the product failure and removes irrelevant chatter. A conversation state tracking mechanism based on a hierarchical recurrent neural network (HRED) maintains coherence. For example, if a user shifts from discussing the problem to inquiring about repair methods, it accurately captures the topic transition and resets the context when necessary. An integrated conditional random field (CRF) model better models the sequence of state transitions. Finally, an ensemble learning strategy weights the outputs of sub-models such as intent recognition and sentiment analysis. If the intent recognition sub-model is more critical for determining user needs at that moment, the adaptive online learning algorithm increases its weight, enabling the customer service system to more reliably understand the context and accurately respond to customers.
[0181] Furthermore, a Transformer-based multi-head attention mechanism is optionally employed to capture long-range dependencies in conversations. By encoding the conversation history into three matrices: query, key, and value, the system can adaptively focus on relevant historical information, achieving dynamic understanding of context. Furthermore, the introduction of positional encoding ensures the preservation of conversation order. The module integrates a hierarchical semantic understanding framework. At the bottom layer, a pre-trained BERT-like model is used for basic semantic encoding. A graph neural network (GNN)-based conversation topic tracking mechanism is employed in the middle layer. Finally, a global semantic integration based on attention pooling is implemented at the top layer. This hierarchical design simultaneously captures semantic information at the word, sentence, and conversation levels. To handle long-term conversation scenarios, the module also implements a dynamic context management mechanism based on a sliding window. By setting an adjustable context window size and combining it with an importance scoring algorithm, the system can adaptively retain key contextual information while avoiding the accumulation of irrelevant information. Importance scoring utilizes an improved REINFORCE algorithm, which optimizes the context selection strategy through reinforcement learning. When handling multi-turn conversations, the module introduces a dialogue state tracking mechanism based on a hierarchical recurrent neural network (HRED). This mechanism maintains conversational coherence, accurately captures topic transitions, and triggers context resets when necessary. Furthermore, by integrating a conditional random field (CRF) model, the system can better model the sequence of state transitions within a conversation.
[0182] Optionally, to improve system robustness, an ensemble learning strategy is employed, combining the weighted outputs of multiple specialized sub-models (such as intent recognition, sentiment analysis, and coreference resolution) to achieve more reliable contextual understanding. Weights are updated using an adaptive online learning algorithm, dynamically adjusting the importance of each sub-model based on actual conversation performance.
[0183] As an optional embodiment, in step 103, the target search information is personalized in combination with the personalized knowledge base of the target user to obtain response information that matches the target user's usage habits, including:
[0184] Perform multi-head attention calculation on the target retrieval information; based on the calculation results, evaluate the consistency between the target retrieval information and the historical interaction information stored in the short-term memory module and the long-term memory module; if the consistency evaluation is passed, generate initial response information based on the target retrieval information; use control codes to perform conditional text generation processing on the initial response information to optimize the personalized expression of the initial response information; wherein the personalized expression includes at least: tone, style, and idioms; generate dynamic image adjustment parameters based on the image information in the historical interaction information, and optimize the image information in the initial response information through the dynamic image adjustment parameters to obtain response information that matches the target user's usage habits.
[0185] Specifically, in this step, a multi-head attention mechanism is used to process the target retrieval information. This multi-head attention mechanism acts like multiple "perspectives," focusing on the key content of the target retrieval information and the relationships between its internal elements from different perspectives. For example, for target retrieval information in various forms, such as text and images, it can focus on the connections between different semantic blocks in the text, as well as the correspondence between image elements and text descriptions. Through such calculations, richer and more detailed information features are mined, providing a foundation for subsequent operations such as consistency evaluation.
[0186] The results of multi-head attention calculations are used to measure the consistency between the target retrieval information and the historical interaction information stored in the short-term memory module and the long-term memory module. The historical interaction information records every detail of previous interactions with the target user, reflecting the user's past concerns, preferences, and so on. Through comparative analysis, for example, we can check whether the topics and key content mentioned in the target retrieval information are consistent with the content frequently discussed by previous users, and whether the expression style of the text is close to the previous communication style. If a high degree of consistency is shown in many aspects, it means that the target retrieval information is consistent with the user's past interaction habits, and the consistency assessment will pass. This means that the information is in line with the user's consistent interaction context and can proceed to the next step of generating response information.
[0187] Once the consistency assessment is passed, the initial response information is generated based on the target search information. This process comprehensively considers various elements in the target search information and transforms it into a prototype of content that can be replied to the user. For example, if the target search information focuses on a product feature that the user inquired about, the initial response information will integrate relevant feature introductions, usage suggestions, and other content, presenting it in a relatively complete and coherent manner, forming a basic text content that can respond to the user's needs. Of course, it may also be supplemented with corresponding elements such as images to enrich the content.
[0188] Next, control codes are used to conditionally generate text for the initial response, aiming to optimize its personalized expression. Control codes serve as a guide, as different target users have different desired tones, styles, and idiomatic expressions. For example, for young and lively users, control codes might guide the generation of a lighthearted, humorous, and colloquial response; whereas for more professional and rigorous users, they might encourage the generation of a formal and professional expression. In this way, the initial response's expression is more aligned with the user's preferences, enhancing their sense of familiarity and identification with the response content.
[0189] Finally, considering that the response message may contain image elements, image information is extracted from historical interaction information. The user's past image preferences, such as preferred color tones and composition styles, are analyzed to generate dynamic image adjustment parameters. These parameters are then used to optimize the image information in the initial response message. For example, the image's brightness and contrast can be adjusted to better suit the user's visual experience, or the composition can be altered to achieve a more user-pleasing presentation. This series of processing ultimately results in response information that matches the target user's usage habits, fully meeting their personalized needs and enhancing their interactive experience with the system.
[0190] As an optional embodiment, the personalized adaptation module has built a comprehensive and advanced user personalized adaptation system, covering a variety of cutting-edge technologies and innovative strategies, aiming to accurately grasp user needs and provide highly tailored personalized services. For example, the underlying vector embedding technology is used to convert various types of data such as user interaction behaviors and preference characteristics into high-dimensional vector forms, so that they can be effectively mathematically represented in high-dimensional vector space for subsequent calculations and analysis. In this process, combined with the attention mechanism, dynamic weights are flexibly assigned to different user features based on the user's real-time interaction situation and interest change trends. For example, when a user frequently browses a certain type of specific product information recently, the weight of the features related to the product will increase accordingly, thereby accurately capturing the dynamic fluctuations of user interests and ensuring that the user portrait is always close to the user's current interest status.
[0191] For example, we can further leverage graph neural networks (GNNs) to construct user knowledge graphs, integrating various types of scattered user information and clearly presenting the inherent relationships between them, presenting a comprehensive and structured picture of the user. For example, we can connect information nodes such as user purchase records, browsing history, search keywords, and behavior in different scenarios through graphs to form an organic knowledge network, deeply exploring the potential connections and patterns behind user behavior.
[0192] For example, through incremental learning algorithms, the system can acquire and integrate new user interaction data in real time, promptly updating and optimizing existing user models and seamlessly adapting to the continuous evolution of user interests. Each new interaction is viewed as a learning opportunity, enabling the model to keep pace with the times, accurately reflecting the latest changes in users' interests and needs, and ensuring the timeliness and accuracy of personalized services.
[0193] The introduction of a meta-learning framework, specifically the Model-Agnostic Meta-Learning (MAML) algorithm, enables the system to rapidly learn and adapt to new users. This algorithm can rapidly adjust model parameters with a small sample size to find a personalized configuration suitable for new users, effectively addressing the cold start problem for new users, providing them with an efficient and accurate personalized experience and accelerating the adaptation process between the system and new users.
[0194] In terms of adaptive strategies, the module uses the policy gradient-based REINFORCE algorithm from reinforcement learning. Through continuous interaction with users, it continuously collects feedback signals and optimizes personalized strategies based on these signals. For example, when the system's recommendations receive positive feedback (such as clicks, favorites, and positive reviews), the system strengthens the recommendation strategy. Conversely, when the system's recommendations receive negative feedback, the system adjusts and improves the strategy, gradually improving the quality and effectiveness of personalized recommendations and making the interaction between the system and users smoother and more satisfying.
[0195] To find the optimal balance between exploring new personalized strategies and leveraging existing effective ones, the system uses a multi-armed bandit algorithm. When faced with multiple possible personalized strategies, this algorithm can try new strategies to discover potentially better solutions (exploration) while also avoiding excessive abandonment of proven strategies (utilization). This ensures the stability of overall service quality and improves efficiency while continuously optimizing personalized services.
[0196] To enhance the robustness of the system, the module integrates contrastive learning technology. By performing self-supervised learning on user behavior data, it can automatically discover the inherent structure and similarities in the data, thereby generating more stable and reliable user representations. For example, when processing user browsing data, contrastive learning can identify similar patterns and different features between different browsing behaviors. Even in the presence of certain noise or incomplete data, it can accurately extract key user information, providing a solid foundation for subsequent personalized services.
[0197] At the same time, an adversarial training mechanism is introduced, allowing the model to continuously learn and optimize while confronting potential noisy data, thereby improving its resistance to noisy data. For example, when faced with abnormal user behavior data (which may be caused by misoperation, system failure, or malicious attack), the model can effectively filter out this noise interference through the feature recognition capabilities learned through adversarial training, maintain an accurate judgment of the user's true intentions and interests, and ensure the stability and reliability of personalized services.
[0198] To make personalized service results explainable and enable users and developers to clearly understand the basis and process of personalized decision-making, the module integrates technologies such as attention visualization and decision trees. This technology intuitively demonstrates the system's focus on different features when processing user information. For example, when generating personalized recommendations, it's clear which user features have a key impact on the results.
[0199] Furthermore, using SHAP (SHapley Additive exPlanations) value analysis, the system can precisely quantify the contribution of different features to personalized decisions, providing an in-depth analysis of the role each feature plays in the final decision. This not only helps users understand why they receive specific personalized services, but also provides developers with a powerful basis for optimizing personalization models. By adjusting and optimizing the weights and processing of key features, the quality and effectiveness of personalized services can be further improved, making the entire personalized adaptation process more transparent, controllable, and efficient.
[0200] In terms of retrieval technology, the memory retrieval and reasoning module uses a comprehensive set of strategies at the retrieval technology level, aiming to achieve efficient and accurate information retrieval. A combination of vector retrieval and semantic search is adopted. With the help of the Dense Passage Retrieval (DPR) model, the query content entered by the user and the memory information stored in the system can be mapped to the same vector space. Based on this space, the similarity matching operation can be completed quickly, and the information related to the query can be quickly filtered out, thereby improving the speed of retrieval. At the same time, a BERT-based semantic understanding model is introduced. This model can conduct in-depth semantic analysis of the search content, explore its inherent meaning, and further accurately grasp the query intent, thereby effectively improving the accuracy of the retrieval, avoiding the retrieval errors that may be caused by relying solely on surface keyword matching, and ensuring that the retrieval results are more in line with actual needs.
[0201] In addition, the system also integrates the Attention Mechanism, which plays an important role in the entire retrieval process. It can dynamically adjust the weights of two different retrieval strategies, vectorized retrieval and semantic retrieval, based on different retrieval scenarios and actual needs, so that the two can better cooperate and optimize the retrieval effect.
[0202] In terms of reasoning mechanisms, a multi-hop reasoning network (MRN) is employed. This network is capable of multi-step logical deduction, enabling the gradual reasoning of conclusions along complex logical chains. In its implementation, a graph neural network (GNN) is employed to construct a knowledge graph. Leveraging the message passing mechanism within the knowledge graph, nodes exchange information and collaborate on computations, completing various complex reasoning tasks. This provides strong structural support and information exchange pathways for the reasoning process.
[0203] Furthermore, considering the uncertainty inherent in actual reasoning, the module introduces an uncertain reasoning mechanism, employing Bayesian networks to handle probabilistic reasoning. Bayesian networks can scientifically calculate posterior probabilities based on prior probabilities and newly acquired evidence, enabling reasonable inferences about uncertain situations. This effectively improves the reliability of reasoning results and makes them more adaptable to complex and changing real-world scenarios.
[0204] To comprehensively improve retrieval efficiency, especially in response to the pressure of retrieval when memorizing large amounts of data, the system deploys a number of targeted technical measures. This enables a hierarchical index structure and employs locality-sensitive hashing (LSH) technology to construct a fast retrieval index. The hierarchical index structure organizes information more systematically, facilitating rapid location and search. Locality-sensitive hashing, through clever algorithmic design, enables fast approximate nearest neighbor searches in high-dimensional vector spaces, further accelerating retrieval and enabling the system to respond quickly even to massive amounts of data.
[0205] When processing large-scale memory data, a distributed retrieval architecture is used. By assigning retrieval tasks to multiple nodes for asynchronous parallel processing, each node works simultaneously, making full use of computing resources, greatly improving the overall retrieval speed and ensuring that the system can efficiently handle large-scale data retrieval needs.
[0206] The system also integrates an adaptive caching mechanism that dynamically adjusts the storage tier of hot data based on data access frequency. Frequently accessed hot data is stored in a more accessible tier, allowing for quick and direct access the next time it's retrieved, avoiding repeated searches and further improving retrieval efficiency.
[0207] To ensure the interpretability of reasoning results, the module uses corresponding mechanisms to make the reasoning process clearly visible and the results reliable. It integrates an attention-based explanation mechanism, which can track the reasoning path during the reasoning process, clearly showing how each step of the reasoning is carried out. It can also generate visual explanations, presenting the abstract reasoning process in an intuitive form, making it easier for users and relevant personnel to understand the logic and basis of the reasoning. At the same time, it is equipped with a confidence assessment mechanism, which will score the reliability of the reasoning results, measuring the credibility of the reasoning results from a quantitative perspective, providing an important reference for subsequent decision-making based on the reasoning results, helping decision-makers better judge the effectiveness and risk level of the results, and make more scientific and reasonable decisions.
[0208] Through the above-mentioned comprehensive design and application in multiple aspects such as retrieval technology, reasoning mechanism, retrieval efficiency improvement and interpretability of reasoning results, the memory retrieval and reasoning module can provide strong support for the operation of the entire system more efficiently, accurately and reliably, and meet diverse application needs.
[0209] The output generation module plays a crucial role in the personalized memory system. As the final key link, it undertakes the core task of organically integrating the retrieved memory information with the current context, thereby generating a coherent and personalized response. The output generation module uses a decoder network based on the Transformer architecture as its main framework, integrating an attention mechanism and a context-aware encoder on this basis. In the specific processing flow, when faced with information obtained from the short-term memory and long-term memory modules, the system uses a multi-head attention calculation method to process it. Through the multi-head attention mechanism, the system is able to accurately capture the inherent connections between information and their connection with historical interaction content, as if examining this information from multiple perspectives. This ensures that the final output content is highly consistent with past historical interactions, allowing the generated response to fit the user's previous interaction context and give people a coherent and natural feeling.
[0210] This module also integrates conditional text generation technology (Conditional Text Generation), which uses control codes to flexibly adjust the tone, style, and personalized features of the output content. The system will fully rely on user portrait information, because the user portrait comprehensively depicts the user's expression habits, preferences and other key factors, and then dynamically adjust the corresponding generation parameters. For example, for users who prefer formal and rigorous styles, the system will use control codes to guide the generation of content that conforms to this style; and for users who are accustomed to relaxed and lively expressions, the system will generate corresponding output content to ensure that the output content can meet the user's personalized needs to the greatest extent possible, thereby enhancing the user's recognition and acceptance of the output content.
[0211] To further improve the quality of output, the module introduces an output optimization mechanism based on reinforcement learning. By carefully designing a specific reward function, the system can learn and explore a better output strategy. This reward function is multi-dimensional, covering multiple aspects such as information accuracy, language fluency, context relevance, and personalized matching. For example, when the output content is accurate in terms of facts, the language expression is fluent and natural, it is closely related to the context, and it highly matches the user's personalized characteristics, the system can obtain higher reward feedback, and then continuously adjust and optimize the output strategy based on such feedback. While ensuring the accuracy and fluency of the output content, it can also fully demonstrate personalized colors, achieve balanced development in all aspects, and provide users with high-quality response content.
[0212] During the actual operation phase, the module adopts a hierarchical decoding strategy. First, the overall framework of the output content is planned, just like constructing the main structure when building a building, clarifying the overall idea and framework layout; then, the specific content details are gradually refined, so that the generated content has clear structural integrity and rigorous logical coherence from macro to micro, avoiding content confusion and illogical situations. At the same time, by integrating a memory-based sampling mechanism, the system can fully explore those high-quality expressions in historical interactions and apply them reasonably to the current output content, so that the generated response can not only reflect the actual needs of the current situation, but also incorporate successful and popular expression elements in the past, thereby improving the overall expression quality.
[0213] The module is equipped with a dedicated output quality control unit, which, through the establishment of multiple constraints and post-processing rules, comprehensively ensures that the generated content meets predetermined quality standards. This includes factual verification of the output content to avoid inaccurate information; grammatical verification to ensure that the language expression conforms to grammatical rules and is smooth and accurate; and an assessment of the degree of compatibility with the user's personalized characteristics to reconfirm that the output content's style and tone truly meet the user's preferences. Through such rigorous quality control, the output content presented to users is ultimately of a high quality level, effectively improving the user experience.
[0214] In summary, through this series of comprehensive and mutually coordinated technical solutions, the output generation module is committed to generating high-quality, personalized, and coherent response content for the personalized memory system to meet users' interaction needs in different scenarios.
[0215] As an optional embodiment, in step 103, before personalizing the target search information in combination with the personalized knowledge base of the target user and obtaining response information that matches the target user's usage habits, the following may also be implemented:
[0216] Obtain personal preference data input by a target user; the personal preference data includes at least: the target user's business field, personal habits, and usage preferences; extract the personal preference data through an information extraction module, and construct the extracted user usage habit data into structured data corresponding to the target user to obtain the personalized knowledge base; based on the target user's historical access data, dynamically set the update strategy of the personalized knowledge base.
[0217] For example, suppose a target user enters personal preference data when starting to use an online reading platform. For example, they may indicate that they work in the education industry, prefer to read before bed at night, and prefer humanities and social science books with a simple layout and moderate font size. The platform's information extraction module extracts this personal preference data, identifying key user usage habits such as "education industry," "reading before bed," "humanities and social science books," and "simple layout and moderate font size." This data is then used to construct structured data specific to the target user, forming a personalized knowledge base. Based on the target user's historical visit data, if, for example, the user frequently views history and philosophy books on the platform and their reading time is generally fixed at night, the platform dynamically sets an update strategy for the personalized knowledge base. For example, based on new reading preferences, the weight of more specific history and philosophy categories in the humanities and social sciences will be increased. If the user stops and views a new typesetting style multiple times during the reading process, the corresponding page presentation preferences will also be updated, so that the personalized knowledge base can always adapt to the user's changing usage habits. When the target retrieval information (such as recommended books and other related content) is subsequently processed, matching response information can be generated, such as recommending new books that suit their preferences.
[0218] It can be understood that based on the above principles and steps, in a core workflow in an embodiment of the present application, the system uses multimodal fusion technology in the information input stage, and uses multiple specialized encoders to cope with different types of input information, thereby building a foundation for multimodal information processing.
[0219] For text information, pre-trained models such as BERT or RoBERTa are selected for encoding, converting the text content into a vector representation suitable for subsequent processing. For image information, visual models such as VisionTransformer or ResNet are used to complete feature extraction and accurately capture the key features of the image. When processing voice information, speech recognition models such as Wav2Vec or Whisper are used to perform corresponding operations. On this basis, the cross-modal attention mechanism is used to align and fuse the features extracted from different modalities into a unified feature representation, eliminating modal differences and providing coherent and integrated input data for the subsequent deep understanding stage.
[0220] Entering the deep understanding stage, the system builds a hierarchical semantic understanding architecture to deeply explore the connotation of information. First, a large-scale language model based on Transformer (such as the GPT series) is used to perform preliminary semantic analysis of the input information, extract key entities, events and the relationship between them, and sort out the basic context of the information. Subsequently, with the help of the reasoning module enhanced by the knowledge graph, the extracted information is analyzed in association with the existing knowledge of the system. This process relies on the graph neural network (GNN) technology, and through its message passing mechanism, the relationship between entities is carefully modeled to achieve a deeper level of semantic understanding and to explore the logical connections hidden behind the information. In addition, the attention mechanism is also used in this process to highlight important information and avoid key content being ignored in complex information processing, ensuring that the core content can be accurately grasped, laying a solid foundation for the further use of information.
[0221] During the short-term memory storage phase, the system introduces an improved neural graph storage network to manage information storage. This utilizes a graph-based storage architecture, where each node corresponds to an information unit and edges between nodes represent the relationships between information, creating a logically structured information storage network. During storage operations, the concept of differential neural computers (DNCs) is incorporated, leveraging differentiable read and write operations to flexibly manage memory content and achieve orderly storage and retrieval of information. Furthermore, an attention-based indexing mechanism is introduced, leveraging vector similarity calculations to enable rapid information retrieval and improve information search efficiency. To avoid information redundancy and forgetting, the system integrates a memory update strategy based on importance scoring. Using reinforcement learning algorithms, the storage strategy is dynamically adjusted based on actual conditions, ensuring that the information stored in short-term memory remains valuable and relevant.
[0222] These three stages are closely connected through an end-to-end deep learning framework, forming a complete information processing pipeline. The stages collaborate with each other and progress sequentially to complete the information processing process.
[0223] In terms of overall training, the system uses optimization algorithms such as gradient descent for comprehensive training, and adopts a multi-task learning approach to simultaneously optimize multiple objectives. Factors such as information comprehension accuracy, storage efficiency, and retrieval speed are all within the scope of optimization considerations, thereby comprehensively improving system performance. Furthermore, by introducing contrastive learning and self-supervised learning techniques, the model's generalization and robustness are further improved, enabling it to better handle information processing tasks in different scenarios and ensuring the stable and efficient operation of the entire information processing workflow. These three stages are tightly connected through an end-to-end deep learning framework to form a complete information processing pipeline. The system uses optimization algorithms such as gradient descent for overall training, and uses multi-task learning to simultaneously optimize multiple objectives, including information comprehension accuracy, storage efficiency, and retrieval speed. At the same time, by introducing contrastive learning and self-supervised learning techniques, the model's generalization and robustness are improved.
[0224] Based on the above principles and steps, in another core workflow in the embodiment of the present application, this module is at the starting link of the entire workflow, mainly relying on the attention mechanism based on the Transformer architecture and combining the BERT-type pre-training model to carry out context encoding. First, the input dialogue history will be segmented and vectorized to convert the text-based dialogue content into a vector form that the computer can better process and analyze. Then, the multi-head self-attention mechanism is used to capture key information and semantic dependencies in the context from multiple angles, just like examining the dialogue content from different perspectives to dig out the key points and internal logical connections. At the same time, in order to accurately grasp the temporal characteristics of the dialogue, a temporal perception mechanism is also introduced to effectively maintain the dialogue sequence information through position encoding to avoid misunderstandings caused by disordered sequences, thereby ensuring that the temporal context of the current dialogue can be accurately understood and providing clear and accurate basic information for the operation of subsequent modules.
[0225] After completing the context analysis, the long-term memory activation module is entered. This module uses a structure similar to a neural graph network (GNN) to present the stored long-term memory in the form of a knowledge graph, building a knowledge network with rich associations.
[0226] Leveraging a specially designed similarity calculation algorithm, the system can rapidly locate and activate relevant memory nodes within a large-scale long-term memory repository based on the current context information output by the context analysis module. This process utilizes an improved attention mechanism and sparse activation strategy. The improved attention mechanism more precisely focuses on key memory nodes, while the sparse activation strategy ensures that, even with massive amounts of memory, only the truly relevant parts are selectively activated, effectively avoiding interference from irrelevant information. This enables efficient and accurate long-term memory activation, preparing the necessary and appropriate knowledge foundation for subsequent information retrieval and reasoning.
[0227] The information retrieval and reasoning module integrates a hybrid architecture of vector retrieval and symbolic reasoning to complete complex and accurate information processing tasks.
[0228] At the vector space level, an improved nearest neighbor search algorithm (such as HNSW) is used to perform efficient similarity matching, quickly finding content similar to the required vector representation among a large amount of information, thereby ensuring retrieval efficiency. At the symbolic level, a rule-based reasoning engine is relied upon to process structured knowledge, performing rigorous reasoning operations according to established rules and logic to ensure the accuracy of reasoning. These two levels do not operate independently, but rather cooperate and work together through a carefully designed fusion mechanism, allowing the entire retrieval and reasoning process to not only quickly obtain relevant information but also perform reasonable reasoning based on rules to produce reliable results. In addition, to further enhance the credibility of the results, a confidence assessment mechanism is introduced to score the reliability of the results obtained from retrieval and reasoning, screen out more trustworthy output results, and provide strong support for the system's final response.
[0229] These three modules work closely together through a specially designed collaborative mechanism, forming an organic whole. The output of the context analysis module directly determines the scope and intensity of long-term memory activation, thereby providing a clear direction and limiting the scope of long-term memory activation. The activated long-term memory, like a "treasure trove of knowledge," provides essential knowledge support for subsequent information retrieval and reasoning, enabling more accurate retrieval and reasoning tasks.
[0230] The entire process thus forms a closed-loop information processing chain, with each link interconnected and linked, ensuring the consistency and accuracy of the system's responses, and ensuring that the output information conforms to the dialogue logic and is trustworthy. It is also worth mentioning that each module is designed with the concept of explainability in mind. Whether it is the capture of key information in context analysis, the selection of nodes when activating long-term memory, or the operations and results judgment during information retrieval and reasoning, they are all easy to track and deeply understand, facilitating subsequent optimization of the system's decision-making process, enabling the entire system to continuously improve performance and better serve information processing needs.
[0231] Based on the above principles and steps, in another core workflow in the embodiment of the present application, the memory integration and update module plays a key basic role in the entire information processing process. It uses a hybrid architecture of the hierarchical attention network (Hierarchical Attention Network) and the dynamic memory network (Dynamic Memory Network) to carry out its work.
[0232] First, the module uses a multi-head self-attention mechanism to calculate the weights of new and old information. This is like examining this information from multiple perspectives, assigning corresponding weights to different pieces of information based on their respective characteristics and importance. This importance scoring method clarifies the priority that different pieces of information should have in memory. On this basis, a graph neural network (GNN) is further used to construct an association network between information. Each piece of information is like a node in the network. Through the message passing mechanism, the nodes communicate and transmit information with each other, thereby updating the node representation and allowing the association between information to be dynamically reflected. Ultimately, dynamic memory integration is achieved, so that the memory content can be continuously optimized and adjusted according to the importance and mutual relationship of the information, forming a more logical and valuable knowledge system.
[0233] After completing the memory integration and update, the forgetting mechanism optimization module begins to play a role. Its core is to adopt an improved version of the LSTM (Long Short-Term Memory) network and introduce a forgetting gate control mechanism based on time decay.
[0234] The system comprehensively considers characteristics across multiple dimensions, such as frequency of use, time intervals, and information importance, to construct an adaptive forgetting strategy. Through reinforcement learning algorithms, the forgetting threshold is dynamically adjusted based on actual information usage and various characteristic performance factors to determine which information should be forgotten and which is worth retaining. At the same time, comparative learning methods are used to differentiate between information of different values. For high-value information, emphasis is placed on maintaining its representational quality so that it can exist stably and accurately in memory. For low-value information, appropriate compression or direct deletion is performed to prevent useless information from excessively occupying memory resources, thereby optimizing the entire memory system and ensuring that the memory content always maintains a high level of effectiveness and value.
[0235] After the memory is processed by the previous two modules, the personalized adaptation module uses the MAML (Model-Agnostic Meta-Learning) algorithm based on the meta-learning framework to achieve rapid adaptation to meet the personalized needs of different users.
[0236] The personalized adaptation module first collects a large amount of user interaction data, extracts the key elements that can reflect the user's characteristics and preferences, and then establishes a personalized user portrait vector. This vector is like the user's "personality label" in the system, which comprehensively and accurately depicts the user's uniqueness. Subsequently, the Transformer encoder is used to conduct a more in-depth modeling of user preferences, converting abstract preferences into a form that computers can better process and analyze. In addition, the system also integrates transfer learning technology. With the help of feature transfer and model fine-tuning, the general features and knowledge learned from existing data are flexibly applied to personalized services for different users. Adaptation and adjustment are made according to the specific situation of each user, truly providing different users with services that meet their personalized needs and improving user experience.
[0237] These three modules work in a close and orderly collaborative relationship, forming a complete and coherent workflow. The memory integration and update module lays a solid memory foundation for subsequent operations. The processed memory content is more organized and valuable, providing a clear target for the forgetting mechanism optimization module. The forgetting mechanism optimization module screens and optimizes the integrated memory according to established strategies to ensure the efficiency of the memory system, while preparing high-quality, redundant memory information for the personalized adaptation module. The personalized adaptation module, based on the optimized memory and combined with the user's personalized needs, provides precise personalized services, further enhancing the applicability and effectiveness of the entire system in practical applications.
[0238] In summary, this workflow, through the unique technical application of each module and the coordination between each other, has formed a complete information processing system from memory integration, forgetting optimization to personalized adaptation, effectively improving the system's capabilities and level in memory management and personalized services.
[0239] Furthermore, the system optionally uses a retrieval-augmented generation (RAG) architecture, combined with an attention mechanism and a context-aware sequence generation model to achieve the task of response generation.
[0240] During the memory retrieval phase, the system mainly relies on a dual encoder architecture. The query encoder is responsible for processing the information currently input by the user and the related contextual content, converting it into a vector form that is convenient for subsequent calculation and comparison; while the document encoder encodes the stored memory content. Afterwards, the degree of association between the two is measured through vector similarity calculation (such as the commonly used cosine similarity), so as to quickly locate the memory fragments that are most relevant to the current needs. In this process, a sparse-dense hybrid retrieval strategy is adopted. The advantage of this strategy is that it not only ensures the efficiency of retrieval, allowing the system to quickly screen out potentially relevant parts from a large number of memory contents, but also fully considers semantic relevance to ensure that the retrieved memory fragments are indeed consistent with the user input and context at the semantic level, laying the foundation for the subsequent generation of high-quality responses.
[0241] After completing the memory retrieval, the stage of deep integration of the retrieved memory information with the current context begins. This link uses the cross-attention mechanism (Cross-Attention) to achieve this goal, and adopts a multi-head attention (Multi-Head Attention) structure. The multi-head attention structure is like multiple "observation perspectives", which allows the model to focus on information features at different levels at the same time, explore the intrinsic connections between information from multiple angles, and then achieve more fine-grained information integration, so that the memory information and the current context can be more closely and comprehensively integrated. At the same time, in order to better dynamically control the importance of memory information and context according to actual conditions, a context-aware gating mechanism (Context-Aware Gating) is also introduced. This mechanism will flexibly adjust the weights of the two according to the specific interaction scenarios and information characteristics, highlight key information, make the fusion process more scientific and reasonable, and further improve the fusion effect.
[0242] Finally, the response generation stage utilizes an improved Transformer decoder architecture and integrates a memory-enhanced self-attention layer. The system generates responses word by word through auto-regressive decoding, generating each word sequentially in a specific order to gradually construct a complete response. During the generation process, a decoding strategy based on nucleus sampling is employed. This strategy cleverly strikes a balance between response diversity and coherence, avoiding overly monotonous and stereotyped content while ensuring smooth, logically coherent sentences that conform to normal language expression. Furthermore, to make the generated content more personalized, a post-processing mechanism is introduced to align the style of the generated responses with the user's historical interaction style, enhancing the user's familiarity and identification with the response content, and comprehensively improving the response quality and user experience.
[0243] In summary, through the orderly collaboration of these key links, the system generates high-quality response content based on retrieval of memory and integration of the current context, providing users with responses that are accurate, coherent, and consistent with their interaction style.
[0244] Summarizing the above steps, an overall example of an embodiment of the present application is as follows: When a user inputs a meeting record via voice, the system begins operation. During the information input and multimodal processing phase, the multimodal information perception module captures the voice information. The voice processing unit first converts the audio signal into a digital signal. After noise reduction and audio enhancement, the speech recognition engine converts it into text data. The multimodal fusion unit integrates the text and the original audio, including features such as intonation and voice, to form a unified information representation. Entering the information understanding and extraction phase, the information understanding module conducts in-depth semantic analysis of the fused meeting record, extracting key elements such as the meeting time, location, participants, main topics, and resolutions. The sentiment analysis module identifies the speaker's tone and attitude to grasp the focus and urgency. The information extractor organizes the structured information according to a preset template to facilitate subsequent operations. This is followed by the memory storage processing phase. Information that needs to be processed immediately is stored in the short-term memory module, while information of long-term value, such as detailed records and resolutions, is stored in the long-term memory module. Timestamps, importance tags, and association tags are added during storage to construct an information association network. Next, in the context analysis and memory retrieval phase, the system analyzes the context of meeting minutes, activates relevant historical memories, and retrieves similar or related historical meeting minutes. It also proactively searches to-do items and schedules to avoid conflicts between new meeting schedules and existing plans. In the memory update and maintenance phase, the memory integration module links new meeting information with historical records and updates the information network. The forgetting mechanism adjusts storage strategies based on timeliness, importance, and frequency of use, handling completed or expired information accordingly. The information classification module updates the index structure to ensure efficient retrieval. In the control and personalized adaptation phase, the central coordination module monitors all processes to ensure coordinated operation. The personalized adaptation module learns user meeting management habits and updates user profiles, covering personalized features such as time management and task prioritization, to optimize subsequent information processing. Finally, in the output generation phase, the system generates personalized responses, such as meeting minutes summary, to-do reminders, and relevant suggestions. The presentation format is selected based on user preferences, and an interactive confirmation mechanism is provided to facilitate modifications and additions.
[0245] Furthermore, in practical applications, the system can handle a variety of special situations. It proactively prompts for additional information when input is incomplete; activates a priority processing mechanism for urgent matters; provides recommended solutions to information conflicts; and establishes specific pattern recognition and processing procedures for recurring meetings. The entire system utilizes a "hierarchical memory management architecture" that dynamically adjusts storage locations and priorities within short- and long-term memory levels based on information importance, timeliness, and relevance. Furthermore, through a "context-aware memory fusion mechanism," it analyzes scenarios in real time, activates associated historical memories, and intelligently integrates them. This enhances contextual understanding and interactive intelligence, simulating the human memory association process to more efficiently serve user meeting management needs.
[0246] The technical solution of this application improves the memory capacity of the large language model during the interaction process, reduces the information forgetting rate, improves information retrieval efficiency, realizes the effective classification and organization of personalized information, and enhances the user experience. At the same time, the embodiments of this application, through improved context understanding capabilities and personalized adaptation mechanisms, can maintain a high degree of coherence and personalized features in long-term interactions, thereby significantly improving the user experience and interaction quality.
[0247] In another embodiment of the present application, a memory retrieval device based on a large language model is also provided. Figure 2 Said device comprises the following units:
[0248] An interaction unit is configured to respond to an interaction instruction of a target user, perform context analysis on the interaction instruction, and obtain a real-time interaction intention of the target user;
[0249] A retrieval unit is configured to perform memory retrieval on a short-term memory module and a long-term memory module based on the real-time interaction intention to obtain target retrieval information; wherein the short-term memory module and the long-term memory module support bidirectional transmission of stored information; the short-term memory module is used to temporarily store real-time interaction information and real-time retrieval information with a target user within a preset time period; and the long-term memory module is used to store key information in the real-time interaction information and the real-time retrieval information that complies with the preset dynamic transfer strategy;
[0250] The generating unit is configured to perform personalized processing on the target search information in combination with the personalized knowledge base of the target user, and obtain response information that matches the target user's usage habits to respond to the interactive instruction.
[0251] Further optionally, the storage unit, before the retrieval unit performs memory retrieval on the short-term memory module and the long-term memory module based on the real-time interaction intention to obtain target retrieval information, is configured to: collect multimodal information of the target user; the multimodal information includes at least image information, audio information, and text information; extract information to be stored from the multimodal information; the information to be stored includes at least key content information associated with the target user in the multimodal information; the information to be stored is extracted based on an information extraction model using a multi-task learning mode; and store the information to be stored in the short-term memory module according to importance scores and timeliness scores;
[0252] The key information stored in the short-term memory module is transferred to the long-term memory module according to the dynamic transfer strategy; the dynamic transfer strategy is dynamically adjusted using a reinforcement learning algorithm; wherein, the storage mechanism in the short-term memory module is temporary storage that depends on the scoring priority sorting method; the storage mechanism in the long-term memory module is classified compressed storage based on the knowledge graph structure.
[0253] Further optionally, the storage unit stores the information to be stored in the short-term memory module according to the importance score and timeliness score, and is configured to: use a short-term memory model to determine the information importance and timeliness of the information to be stored to obtain a real-time importance score of the information to be stored in the current interactive session; prioritize the information to be stored based on the real-time importance score to obtain a real-time sorting result; adaptively configure the dynamic storage space corresponding to the information to be stored in the short-term memory module based on the real-time sorting result; and store the information to be stored in the corresponding dynamic storage space in the short-term memory module as short-term storage information.
[0254] Further optionally, the corresponding dynamic storage space in the short-term memory module belongs to a dynamic buffer zone; and the short-term memory module is provided with an interactive window mechanism for maintaining and managing the dynamic buffer zone.
[0255] Further optionally, the storage unit transfers the key information stored in the short-term memory module to the long-term memory module according to the dynamic transfer strategy, and is configured to: establish a dynamic priority queue based on the real-time sorting result through a memory integration model; obtain short-term storage information in the dynamic priority queue whose priority meets the transfer conditions indicated in the dynamic transfer strategy according to the memory update cycle or update event indicated in the dynamic transfer strategy; and transfer the short-term storage information that meets the transfer conditions from the short-term memory module to the long-term memory module.
[0256] Further optionally, the storage unit transfers the short-term storage information that meets the transfer condition from the short-term memory module to the long-term memory module, and is configured to: receive the short-term storage information that meets the transfer condition from the short-term memory module; embed the received short-term storage information into the long-term memory map related to the target user through the long-term memory model, and store the received short-term storage information in the corresponding long-term storage space in the vector database based on the embedding result as long-term storage information.
[0257] Further optionally, the storage unit embeds the received short-term storage information into the long-term memory graph related to the target user through the long-term memory model, and stores the received short-term storage information in the corresponding long-term storage space in the vector database based on the embedding result as long-term storage information, and is configured to: encode the received short-term storage information to obtain the initial information code; compress the information dimension of the initial information code through knowledge distillation processing to obtain the optimized information dimension corresponding to the initial information code; establish the correlation between the initial information codes according to the optimized information dimension; determine the graph attributes corresponding to the initial information code in the long-term memory graph according to the correlation and the optimized information dimension; wherein the graph attributes include at least: the dimension to which the information belongs and the information position; based on the dimension to which the information belongs and the information position, store the initial information code in the corresponding long-term storage space in the vector database.
[0258] Further optionally, the storage unit, after storing the information to be stored in the short-term memory module according to the importance score and timeliness score, or after transferring the key information stored in the short-term memory module to the long-term memory module according to the dynamic transfer strategy, is also configured to: configure corresponding time weights for the short-term storage information stored in the short-term memory module and the long-term storage information stored in the long-term memory module; dynamically adjust the time weight through an exponential decay function; wherein the exponential decay function is constructed based on a dynamic forgetting algorithm of time decay, and the dynamic forgetting algorithm is obtained according to the setting of the Ebbinghaus forgetting curve; monitor the access frequency, recent usage time and correlation with other information of the short-term storage information and the long-term storage information; use a forgetting assessment model to sort the importance of the short-term storage information and the long-term storage information based on the monitoring results; perform contextual correlation analysis on the short-term storage information and the long-term storage information to obtain the interactive topic relevance of the short-term storage information and the long-term storage information; determine the short-term storage information and / or long-term storage information to be deleted based on the importance sorting result, the interactive topic relevance and the time weight.
[0259] Further optionally, the storage unit determines the short-term storage information and / or long-term storage information to be deleted based on the importance ranking result, the interaction topic relevance and the time weight, and is also configured to: set the short-term storage information and / or long-term storage information whose relevance to the interaction topic within a set time period is lower than the set relevance threshold as the short-term storage information and / or long-term storage information to be deleted; or, set the short-term storage information and / or long-term storage information whose importance ranking is lower than the set ranking threshold as the short-term storage information and / or long-term storage information to be deleted; or, set the short-term storage information and / or long-term storage information whose time weight is lower than the set time weight threshold as the short-term storage information and / or long-term storage information to be deleted.
[0260] Further optionally, the retrieval unit performs memory retrieval on the short-term memory module and the long-term memory module based on the real-time interaction intention to obtain target retrieval information, and is configured to: perform deep semantic analysis on the real-time interaction intention to obtain the real-time retrieval target corresponding to the target user; retrieve the first candidate information associated with the real-time retrieval target from the short-term memory module; retrieve the second candidate information associated with the real-time retrieval target and / or the first candidate information from the long-term memory module; map the real-time retrieval target, the first candidate context information, and the second candidate information into the same vector space to obtain their respective corresponding projection vectors; calculate the first similarity between the first projection vector corresponding to the first candidate information and the reference projection vector corresponding to the real-time retrieval target; calculate the second similarity between the second projection vector corresponding to the second candidate information and the reference projection vector; generate the target retrieval information based on the first similarity and the second similarity; the target retrieval information is obtained based on inference of a multi-hop Bayesian inference network; wherein the corresponding proportions of the first candidate information and the second candidate information in the target retrieval information are dynamically configured using an attention mechanism; the attention mechanism is associated with a retrieval strategy pre-configured in the retrieval scenario.
[0261] Further optionally, the personalization unit, which combines the personalized knowledge base of the target user to personalize the target retrieval information to obtain response information that matches the target user's usage habits, is configured to: perform multi-head attention calculation on the target retrieval information; based on the calculation results, evaluate the consistency between the target retrieval information and the historical interaction information stored in the short-term memory module and the long-term memory module; if the consistency evaluation is passed, generate initial response information based on the target retrieval information; use a control code to perform conditional text generation processing on the initial response information to optimize the personalized expression of the initial response information; wherein the personalized expression includes at least: tone, style, and idioms; generate dynamic image adjustment parameters based on the image information in the historical interaction information, and optimize the image information in the initial response information through the dynamic image adjustment parameters to obtain response information that matches the target user's usage habits.
[0262] Further optionally, the personalization unit, in combination with the personalized knowledge base of the target user, personalizes the target retrieval information to obtain response information that matches the target user's usage habits, and is also configured to: obtain personal preference data input by the target user; the personal preference data at least includes: the target user's business field, personal habits, and usage preferences; extract the personal preference data through the information extraction module, and construct the extracted user usage habit data into structured data corresponding to the target user to obtain the personalized knowledge base; based on the target user's historical access data, dynamically set the update strategy of the personalized knowledge base.
[0263] Further optionally, the storage information of the short-term memory module and the long-term memory module carries a memory code; the memory code is used to identify the information content type of the corresponding storage information.
[0264] The system can implement various steps in the above method embodiments, which will not be expanded here.
[0265] In the embodiment of the present application, a memory retrieval device based on a large language model is adopted.
[0266] See also Figure 3 , Figure 3 This is a schematic diagram of an embodiment of an electronic device provided in an embodiment of the present application. Figure 3 As shown, an embodiment of the present application provides an electronic device 500, including a memory 510, a processor 520, and a computer program 511 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 511, a memory retrieval method based on a large language model is implemented.
[0267] See also Figure 4 , Figure 4 This is a schematic diagram of an embodiment of a computer-readable storage medium provided in an embodiment of the present application. Figure 4 As shown, this embodiment provides a computer-readable storage medium 600 on which a computer program 611 is stored. When the computer program 611 is executed by a processor, the memory retrieval method based on a large language model is implemented.
[0268] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0269] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0270] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0271] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if such changes and modifications of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include such changes and modifications.
Claims
1. A memory retrieval method based on a large language model, characterized in that: The method comprises: In response to an interaction instruction issued by a target user, performing context analysis on the interaction instruction to obtain the real-time interaction intention of the target user; Based on the real-time interaction intention, a short-term memory module and a long-term memory module are searched to obtain target search information; wherein, the short-term memory module and the long-term memory module support bidirectional transmission of stored information; the short-term memory module is used to temporarily store the real-time interaction information and real-time search information with the target user within a preset time period; the long-term memory module is used to store key information in the real-time interaction information and the real-time search information that conforms to a preset dynamic transfer strategy; The target search information is personalized in combination with the personalized knowledge base of the target user to obtain response information that matches the target user's usage habits to respond to the interactive instruction; Wherein, the memory retrieval of the short-term memory module and the long-term memory module based on the real-time interaction intention, before obtaining the target retrieval information, also includes: collecting multimodal information of the target user; the multimodal information includes at least: image information, audio information, and text information; extracting information to be stored from the multimodal information; the information to be stored includes at least: key content information associated with the target user in the multimodal information; the information to be stored is extracted based on the information extraction model using a multi-task learning mode; the information to be stored is stored in the short-term memory module according to the importance score and timeliness score; the key information stored in the short-term memory module is transferred to the long-term memory module according to the dynamic transfer strategy; the dynamic transfer strategy is dynamically adjusted using a reinforcement learning algorithm; wherein the storage mechanism in the short-term memory module is a temporary storage that depends on the scoring priority sorting method; the storage mechanism in the long-term memory module is a classified compression storage based on the knowledge graph structure; Among them, after the information to be stored is stored in the short-term memory module according to the importance score and the timeliness score, or after the key information stored in the short-term memory module is transferred to the long-term memory module according to the dynamic transfer strategy, it also includes: configuring corresponding time weights for the short-term storage information stored in the short-term memory module and the long-term storage information stored in the long-term memory module; dynamically adjusting the time weights through an exponential decay function; wherein the exponential decay function is constructed based on a dynamic forgetting algorithm of time decay, and the dynamic forgetting algorithm is obtained according to the setting of the Ebbinghaus forgetting curve; monitoring the access frequency, recent usage time and correlation between the short-term storage information and the long-term storage information; based on the monitoring results, using a forgetting assessment model to sort the importance of the short-term storage information and the long-term storage information; performing context correlation analysis on the short-term storage information and the long-term storage information to obtain short-term storage information. The short-term storage information and / or long-term storage information are respectively correlated with the interactive topics of the short-term storage information and the long-term storage information; the short-term storage information and / or long-term storage information to be deleted are determined according to the importance sorting result, the interactive topic relevance and the time weight; wherein, the short-term storage information and / or long-term storage information to be deleted according to the importance sorting result, the interactive topic relevance and the time weight includes: setting the short-term storage information and / or long-term storage information whose relevance to the interactive topic within the set time period is lower than the set relevance threshold as the short-term storage information and / or long-term storage information to be deleted; or, setting the short-term storage information and / or long-term storage information whose importance sorting rank is lower than the set rank threshold as the short-term storage information and / or long-term storage information to be deleted; or, setting the short-term storage information and / or long-term storage information whose time weight is lower than the set time weight threshold as the short-term storage information and / or long-term storage information to be deleted.
2. The memory retrieval method based on a large language model according to claim 1, characterized in that: Storing the information to be stored in the short-term memory module according to the importance score and the timeliness score includes: Using a short-term memory model to determine the information importance and timeliness of the information to be stored, so as to obtain a real-time importance score of the information to be stored in the current interactive session; Prioritizing the information to be stored based on the real-time importance score to obtain a real-time ranking result; Adaptively configuring the dynamic storage space corresponding to the information to be stored in the short-term memory module based on the real-time sorting result; Storing the information to be stored in the corresponding dynamic storage space in the short-term memory module as short-term storage information; Wherein, the corresponding dynamic storage space in the short-term memory module belongs to the dynamic buffer zone; The short-term memory module is provided with an interactive window mechanism for maintaining and managing the dynamic buffer zone.
3. The memory retrieval method based on a large language model according to claim 2, characterized in that: The transferring of the key information stored in the short-term memory module to the long-term memory module according to the dynamic transfer strategy includes: Establishing a dynamic priority queue based on the real-time sorting results through a memory integration model; According to the memory update cycle or update event indicated in the dynamic transfer strategy, obtaining short-term storage information of the dynamic priority queue whose priority meets the transfer condition indicated in the dynamic transfer strategy; The short-term storage information that meets the transfer condition is transferred from the short-term memory module to the long-term memory module.
4. The memory retrieval method based on a large language model according to claim 3, characterized in that: The transferring of the short-term storage information that meets the transfer condition from the short-term memory module to the long-term memory module comprises: receiving short-term storage information satisfying the transfer condition from the short-term memory module; Through the long-term memory model, the received short-term storage information is embedded into the long-term memory map related to the target user, and based on the embedding result, the received short-term storage information is stored in the corresponding long-term storage space in the vector database as long-term storage information.
5. The memory retrieval method based on a large language model according to claim 4, characterized in that: The method embeds the received short-term storage information into the long-term memory map related to the target user through the long-term memory model, and stores the received short-term storage information into the corresponding long-term storage space in the vector database based on the embedding result as long-term storage information, including: Encoding the received short-term storage information to obtain an initial information code; Compressing the information dimension of the initial information encoding through knowledge distillation to obtain an optimized information dimension corresponding to the initial information encoding; Establishing correlations between the initial information codes according to the optimized information dimension; Determining the corresponding atlas attributes of the initial information encoding in the long-term memory atlas according to the relevance and optimized information dimension; wherein the atlas attributes include at least: the dimension to which the information belongs and the position of the information; Based on the dimension to which the information belongs and the information location, the initial information codes are stored in the corresponding long-term storage space in the vector database.
6. The memory retrieval method based on a large language model according to claim 1, characterized in that: The memory retrieval of the short-term memory module and the long-term memory module based on the real-time interaction intention to obtain target retrieval information includes: Performing deep semantic analysis on the real-time interaction intention to obtain the real-time retrieval target corresponding to the target user; Retrieving first candidate information associated with the real-time retrieval target from the short-term memory module; Retrieving second candidate information associated with the real-time retrieval target and / or the first candidate information from the long-term memory module; Mapping the real-time retrieval target, the first candidate context information, and the second candidate information into the same vector space to obtain respective corresponding projection vectors; Calculating a first similarity between a first projection vector corresponding to the first candidate information and a reference projection vector corresponding to the real-time retrieval target; Calculating a second similarity between a second projection vector corresponding to the second candidate information and the reference projection vector; generating the target retrieval information based on the first similarity and the second similarity; the target retrieval information is obtained based on multi-hop Bayesian inference network reasoning; Among them, the corresponding proportions of the first candidate information and the second candidate information in the target retrieval information are dynamically configured using an attention mechanism; the attention mechanism is associated with a retrieval strategy pre-configured in the retrieval scenario.
7. The memory retrieval method based on a large language model according to claim 1, characterized in that: The target search information is personalized in combination with the personalized knowledge base of the target user to obtain response information that matches the target user's usage habits, including: Performing multi-head attention calculation on the target retrieval information; Based on the calculation results, evaluating the consistency between the target retrieval information and the historical interaction information stored in the short-term memory module and the long-term memory module; If the consistency assessment passes, generating initial response information based on the target search information; Using the control code to perform conditional text generation processing on the initial response information to optimize the personalized expression of the initial response information; wherein the personalized expression includes at least: tone, style, and idiomatic expressions; Dynamic image adjustment parameters are generated based on the image information in the historical interaction information, and the image information in the initial response information is optimized using the dynamic image adjustment parameters to obtain response information that matches the target user's usage habits.
8. A memory retrieval device based on a large language model, characterized in that: The device is used to implement the memory retrieval method based on a large language model according to any one of claims 1 to 7, and the device includes the following units, wherein: An interaction unit is configured to respond to an interaction instruction of a target user, perform context analysis on the interaction instruction, and obtain a real-time interaction intention of the target user; a retrieval unit configured to perform memory retrieval on a short-term memory module and a long-term memory module based on the real-time interaction intention to obtain target retrieval information; wherein the short-term memory module and the long-term memory module support bidirectional transmission of stored information; the short-term memory module is used to temporarily store real-time interaction information and real-time retrieval information with a target user within a preset time period; and the long-term memory module is used to store key information in the real-time interaction information and the real-time retrieval information that complies with a preset dynamic transfer strategy; a generating unit configured to perform personalized processing on the target search information in combination with the personalized knowledge base of the target user, and obtain response information that matches the target user's usage habits, so as to respond to the interactive instruction; Wherein, the retrieval unit, before performing memory retrieval on the short-term memory module and the long-term memory module based on the real-time interaction intention to obtain the target retrieval information, is further configured to: collect multimodal information of the target user; the multimodal information includes at least: image information, audio information, and text information; extract information to be stored from the multimodal information; the information to be stored includes at least: key content information associated with the target user in the multimodal information; the information to be stored is extracted based on the information extraction model using a multi-task learning mode; the information to be stored is stored in the short-term memory module according to the importance score and the timeliness score; the key information stored in the short-term memory module is transferred to the long-term memory module according to the dynamic transfer strategy; the dynamic transfer strategy is dynamically adjusted using a reinforcement learning algorithm; wherein the storage mechanism in the short-term memory module is a temporary storage that depends on the scoring priority sorting method; the storage mechanism in the long-term memory module is a classified compression storage based on the knowledge graph structure; Wherein, after the retrieval unit stores the information to be stored in the short-term memory module according to the importance score and the timeliness score, or after transferring the key information stored in the short-term memory module to the long-term memory module according to the dynamic transfer strategy, it is further configured to: respectively configure corresponding time weights for the short-term storage information stored in the short-term memory module and the long-term storage information stored in the long-term memory module; dynamically adjust the time weights through an exponential decay function; wherein the exponential decay function is constructed based on a dynamic forgetting algorithm of time decay, and the dynamic forgetting algorithm is obtained according to the setting of the Ebbinghaus forgetting curve; monitor the access frequency, recent usage time and correlation between the short-term storage information and the long-term storage information; based on the monitoring results, use the forgetting assessment model to sort the importance of the short-term storage information and the long-term storage information; perform context correlation analysis on the short-term storage information and the long-term storage information, To obtain the interactive topic relevance of each short-term storage information and long-term storage information; determine the short-term storage information and / or long-term storage information to be deleted according to the importance ranking result, the interactive topic relevance and the time weight; wherein, the short-term storage information and / or long-term storage information to be deleted according to the importance ranking result, the interactive topic relevance and the time weight includes: setting the short-term storage information and / or long-term storage information whose relevance to the interactive topic within a set time period is lower than the set relevance threshold as the short-term storage information and / or long-term storage information to be deleted; or, setting the short-term storage information and / or long-term storage information whose importance ranking is lower than the set ranking threshold as the short-term storage information and / or long-term storage information to be deleted; or, setting the short-term storage information and / or long-term storage information whose time weight is lower than the set time weight threshold as the short-term storage information and / or long-term storage information to be deleted.
Citation Information
Patent Citations
Text classification method and device
CN109918506A
Systems and methods for providing adaptive ai-driven conversational agents
US20240289863A1