Memory system-based interaction method, device, equipment, storage medium and vehicle
By acquiring users' long-term and short-term memories and combining them with query statements to generate prompts, the problem of large language models failing to fully consider users' long-term interests is solved, thus improving the accuracy and personalization of responses.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING CO WHEELS TECH CO LTD
- Filing Date
- 2024-12-09
- Publication Date
- 2026-05-29
AI Technical Summary
Existing large-scale language models fail to adequately consider users' long-term interests and background information during user interactions, resulting in low accuracy of response statements.
By acquiring the user's long-term memory, short-term memory, and memory summaries, and combining them with the query statement to generate prompt words, which are then input into a trained language model to generate a response statement.
It improves the accuracy of language models in generating responses, reduces misunderstandings and context breaks, and generates responses that better meet individual needs.
Smart Images

Figure CN119293191B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of language model technology, and in particular relates to an interaction method, device, equipment, storage medium and vehicle based on a memory system. Background Technology
[0002] During the dialogue between the user and the language model, after the user inputs a query, the electronic device typically concatenates the query with the user's previous dialogue responses to obtain prompts. These prompts are then used as input to the language model. This concatenation method maintains the continuity of the dialogue and the consistency of the context.
[0003] However, this approach cannot incorporate users' historical behavior into the prompts. Therefore, the language model cannot fully consider users' long-term interests and background information, easily overlooking key information and affecting the accuracy and depth of responses. Consequently, the responses generated by the language model based on prompts have low accuracy. Summary of the Invention
[0004] This application provides an interactive method, apparatus, device, storage medium, and vehicle based on a memory system, which can solve the problem of low accuracy of response statements in existing large language models.
[0005] In a first aspect, embodiments of this application provide an interaction method based on a memory system, the method comprising:
[0006] Get the input query statement;
[0007] The target memory associated with the query statement is determined by the language big model, wherein the target memory includes at least one of long-term memory, short-term memory, and memory summary, and the memory summary is the content obtained by the language big model from the long-term memory and the short-term memory;
[0008] The target memory and the query statement are merged to obtain the prompt words;
[0009] The prompt words are input into the trained language model to obtain the response statement output by the language model.
[0010] In some embodiments, prior to determining the target memory associated with the query statement through a large language model, the method further includes:
[0011] Acquire transient memory, which includes at least one of the following: device information of the interactive device, user characteristic information, and interaction information between the user and the interactive device;
[0012] The transient memories are organized according to a preset standard to obtain short-term memories;
[0013] The short-term memory is incrementally processed to obtain behavioral sequence data, and the long-term memory is generated based on the behavioral sequence data.
[0014] In some embodiments, organizing the transient memory according to a preset specification includes at least one of the following:
[0015] Clean up noisy or invalid data in the transient memory;
[0016] The instantaneous memory is normalized.
[0017] The repeated data in the instantaneous memory are aggregated.
[0018] In some embodiments, when the interactive device is a vehicle, the instantaneous memory includes at least one of the vehicle's vehicle control signals, sensor signals, and image information, voice information, and text information received by the vehicle.
[0019] In some embodiments, the long-term memory includes a summary, which includes an overall summary and a topic summary. After obtaining the input query statement, the method further includes:
[0020] Determine the topic tags corresponding to the query statement;
[0021] Query the first topic summary corresponding to the topic tag;
[0022] The overall summary and the first topic summary are updated according to the query statement.
[0023] In some embodiments, the long-term memory includes preference data, which includes primary topic preferences and secondary topic preferences. After obtaining the input query statement, the method further includes:
[0024] Determine the primary and secondary topic tags corresponding to the query statement;
[0025] Obtain the first-level preference data of the first-level topic tags and the second-level preference data corresponding to the second-level topic tags;
[0026] Update the primary preference data and the secondary preference data according to the query statement.
[0027] In some embodiments, the secondary preference data includes memory strength, and updating the secondary preference data according to the query statement includes:
[0028] Determine the relevance between the query statement and the secondary topic tags;
[0029] Get the total number of interactions associated with the secondary topic tags, and a first time interval, wherein the first time interval is the time interval between the time when the query statement is obtained and the time when the interaction information associated with the secondary topic tags is received last time;
[0030] The memory strength is updated based on the correlation, the total number of interactions, and the first time interval.
[0031] In some embodiments, the method for generating the memory summary includes:
[0032] N memory fragments that match the query statement are obtained through a target matching method, wherein the target matching method includes at least one of keyword matching and vector matching, and N is a positive integer;
[0033] The N memory segments are sorted according to their relevance to the query statement, and the P memory segments with the highest relevance are selected from the sorted N memory segments, where P is a positive integer less than or equal to N.
[0034] The P memory fragments, the query statement, and the short-term memory associated with the query statement are input into the language big model to obtain K memory fragments obtained by the language big model based on thought chain reasoning, wherein the P memory fragments include the K memory fragments, and K is a positive integer less than P;
[0035] The memory summary is obtained by summarizing the short-term memories associated with the K memory fragments and the query statement using the large language model.
[0036] Secondly, embodiments of this application provide an interactive device based on a memory system, the device comprising:
[0037] The first acquisition module is used to acquire the input query statement;
[0038] The first determining module is used to determine the target memory associated with the query statement through a language big model, wherein the target memory includes at least one of long-term memory, short-term memory, and memory summary, and the memory summary is the content obtained by the language big model from the long-term memory and the short-term memory;
[0039] The fusion module is used to fuse the target memory and the query statement to obtain prompt words;
[0040] The response module is used to input the prompt words into the trained language model and obtain the response statement output by the language model.
[0041] Thirdly, embodiments of this application provide an interactive device based on a memory system, the device including: a processor and a memory storing computer program instructions;
[0042] The processor implements the above memory-based interactive method when executing computer program instructions.
[0043] Fourthly, embodiments of this application provide a computer storage medium storing computer program instructions, which, when executed by a processor, implement the above-described memory-based interaction method.
[0044] Fifthly, embodiments of this application provide a vehicle that includes computer program instructions, which, when executed by a processor, implement the above-described memory-based interaction method.
[0045] In this application, since target memory includes the user's long-term memory, short-term memory, and memory summaries, it can fully reflect the user's historical behavior, preferences, and interaction records, as well as the current context information. Therefore, by combining target memory with query statements to obtain prompt words, and inputting the prompt words into the language big model to obtain response statements, the model can generate responses that not only rely on the current input but also combine the user's short-term context information and long-term interests and background. As a result, the language big model is more accurate in understanding user needs, can reduce misunderstandings or context breaks, and can generate responses that are more in line with individual needs based on the user's historical behavior, thereby significantly improving the accuracy of responses. Attached Figure Description
[0046] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 This is a flowchart illustrating an embodiment of the interaction method based on a memory system provided in this application;
[0048] Figure 2 This is a schematic diagram of the structure of an interactive device based on a memory system provided in an embodiment of this application;
[0049] Figure 3 This is a schematic diagram of the hardware structure of an interactive device based on a memory system provided in an embodiment of this application. Detailed Implementation
[0050] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples of this application.
[0051] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0052] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The embodiments will now be described in detail with reference to the accompanying drawings.
[0053] Specifically, in order to address the problems of the prior art, embodiments of this application provide an interaction method, apparatus, device, storage medium, and vehicle based on a memory system. The interaction method based on a memory system provided in the embodiments of this application will be described first below.
[0054] Figure 1 A flowchart illustrating an embodiment of the memory-based interaction method provided in this application is shown. This memory-based interaction method can be used in vehicles, smartphones, or other intelligent devices with interactive functions. The method includes the following steps:
[0055] S110, retrieve the input query statement.
[0056] The method in this embodiment can be applied to vehicles. The query statement is a natural language command entered by the user into the vehicle's interactive device to ask a question or express a need, used to obtain information or perform a specific task. For example, the query statement could be "Navigate to the place I went to last night" or "Play the song I played before".
[0057] S120, the target memory associated with the query statement is determined through the language big model, wherein the target memory includes at least one of long-term memory, short-term memory, and memory summary, and the memory summary is the content obtained by the language big model from the long-term memory and the short-term memory.
[0058] In this embodiment, the target memory may include long-term memory, short-term memory, and memory summaries. The vehicle may include a memory module that stores long-term memory. Long-term memory refers to historical memory information stored in the device, which can be memory information obtained by summarizing long-term accumulated data such as the user's historical behavior, interests, preferences, historical device information, and past conversation records.
[0059] After the vehicle receives the user's query, the system can retrieve relevant long-term memory data associated with the query from the memory module. The long-term memory retrieval process can utilize methods such as keyword matching and vector matching to ensure that the retrieved long-term memory data is highly relevant to the query.
[0060] For example, if a user's query is "Please retrieve my past food order information", the system can find specific content related to "food order" from the user's long-term memory through keyword matching or vector matching.
[0061] Short-term memory can be memory information that is close in time to the time of the query retrieval, or the context information of the query. For example, short-term memory can be memory information within 30 minutes before the retrieval time, or it can be the ten rounds of dialogue between the user and the interactive device before the query.
[0062] Memory summaries are the content generated by the language big model after semantic fusion and generalization of long-term and short-term memories.
[0063] Specifically, as an optional embodiment, the method for generating the memory summary includes:
[0064] N memory fragments that match the query statement are obtained through a target matching method, wherein the target matching method includes at least one of keyword matching and vector matching, and N is a positive integer;
[0065] The N memory segments are sorted according to their relevance to the query statement, and the P memory segments with the highest relevance are selected from the sorted N memory segments, where P is a positive integer less than or equal to N.
[0066] The P memory fragments, the query statement, and the short-term memory associated with the query statement are input into the language big model to obtain K memory fragments obtained by the language big model based on thought chain reasoning, wherein the P memory fragments include the K memory fragments, and K is a positive integer less than P;
[0067] The memory summary is obtained by summarizing the short-term memories associated with the K memory fragments and the query statement using the large language model.
[0068] In this embodiment, the system can first retrieve long-term memories associated with the user-input query from the memory bank using target matching. Specifically, it can retrieve N memory segments associated with the query. Target matching methods can include keyword matching, user ID matching, vector matching, time period matching, and memory dimension matching.
[0069] Keyword matching identifies keywords in the query and directly matches them with the content in the memory bank to find memory segments containing the same keywords. Vector matching calculates the semantic vectors of each memory segment in the memory bank and the semantic vector of the query. It then compares the similarity between the semantic vectors of the query and each memory segment to find the memory segments with the highest similarity. User ID matching retrieves memory segments related to a user based on their unique identifier, such as an account ID or device ID. Time period matching matches memory segments within a specified time range. Memory dimension matching matches the most relevant dimension information to the query based on the dimension to which the memory segment belongs. Dimensions of memory segments can include topics, interests, or behavioral types.
[0070] After retrieving N memory fragments, these fragments can be input into the fine-ranking model, which will sort them based on their relevance to the query. This relevance can be measured by keyword matching or vector similarity. After sorting, the fine-ranking model can select the top P most relevant memory fragments from the sorted results.
[0071] After obtaining P memory fragments, these fragments can be input into the language model. The language model, the query statement, and the short-term memory associated with the query statement are all input into the language model. The language model can autonomously analyze the query statement and the content of the P memory fragments through CoT (Chain of Thought) to gradually determine which memory fragments are most valuable to the current query statement, thereby refining the K optimal memory fragments.
[0072] Subsequently, the language big data model semantically fuses the K memory fragments with short-term memory, summarizing and extracting key information to generate a concise and accurate memory summary. This ensures that multiple memory fragments related to the query, as well as short-term memory, can be integrated into a unified memory summary.
[0073] For example, N can be 10, P can be 5, and K can be 3. If the user's query is "recent fuel consumption data", the system can obtain 10 potentially relevant memory fragments through at least one of the above keyword matching and vector matching. Then, the fine-ranking model sorts the fragments according to their relevance to the query and finally selects the top 5 most relevant memory fragments. The language big data model uses CoT to perform inference analysis on the top 5 memory fragments to obtain 3 optimal memory fragments. These 3 memory fragments are then fused with short-term memory to obtain a memory summary related to the query.
[0074] For example, the DSSM (Deep Structured Semantic Model) dual-tower model can be used to accurately sort N memory fragments and filter out P memory fragments. Then, the CoT (Chain of Thought) method can be used to select K memory fragments from the P memory fragments.
[0075] Through the methods described above, precise matching, sorting, and fusion ensure that the system can effectively extract the most relevant long-term memories from a large amount of data, and summarize these long-term and short-term memories to obtain a memory summary. This multi-stage processing approach improves the accuracy of memory retrieval through fine-grained ranking and large-scale model inference, and generates a memory summary that better meets user needs by combining context and semantic logic.
[0076] S130, the target memory and the query statement are merged to obtain prompt words.
[0077] In this embodiment, after retrieving the long-term memory associated with the query statement, the query statement and the retrieved long-term memory can be combined to generate a complete prompt. The prompt is the input to the language model. Specifically, the fusion process of long-term memory and query statement can be as follows: integrating the user's long-term memory with the current query statement to form a complete prompt.
[0078] For example, if a user enters the query "What was the name of the book I ordered last time?", the system retrieves from its memory platform that the user previously purchased a book titled "Introduction to Artificial Intelligence." The resulting suggestion could be: "The user is asking for the name of the book they purchased last time. According to their history, they last ordered 'Introduction to Artificial Intelligence.' Generate a relevant answer."
[0079] S140, input the prompt word into the trained language model to obtain the response statement output by the language model.
[0080] In this embodiment, after fusing long-term memory and query statements to obtain prompt words, the system can input the generated prompt words into a pre-trained language model. The language model can then generate response statements based on the information in the prompt words and the context. Here, the language model refers to a generative model based on natural language processing, capable of generating coherent dialogue responses from the input text.
[0081] Before applying the large-scale language model, it needs to be trained. Specifically, this can begin by collecting a large amount of text data, which may include books, encyclopedias, news articles, articles, and user history conversations. This text data can then be cleaned, labeled, and formatted to obtain the training set for the large-scale language model.
[0082] Then, the untrained language model is initialized by setting its weight parameters to random values. Next, the training data from the training set is input into the language model in batches, yielding the prediction results for each batch of training data output by the language model.
[0083] For each batch of training data, the actual answers are obtained, and the actual answers are compared with the predicted results. The loss between the actual answers and the predicted results is calculated, and then the gradient of the model parameters is calculated based on the loss and propagated to each layer of the neural network. Finally, the model parameters of the large language model are adjusted through an optimization algorithm to minimize the loss.
[0084] The training data of all batches in the training set can be iteratively trained using the above method. When the loss value drops to within the allowable error range during the training process, the language model can be considered to have been trained and the training of the language model can be ended.
[0085] In this embodiment, since the target memory includes the user's long-term memory, short-term memory, and memory summary, the target memory can fully reflect the user's historical behavior, preferences, and interaction records, as well as the current context information. Therefore, by combining the target memory with the query statement to obtain prompt words, and inputting the prompt words into the language big model to obtain the response statement, the model can not only rely on the current input when generating the response, but also combine the user's short-term context information and long-term interests and background. As a result, the language big model is more accurate in understanding user needs, can reduce misunderstandings or context breaks, and can generate responses that are more in line with individual needs based on the user's historical behavior, thereby significantly improving the accuracy of the response.
[0086] As an optional embodiment, before determining the target memory associated with the query statement through a large language model, the method may further include:
[0087] Acquire transient memory, which includes at least one of the following: device information of the interactive device, user characteristic information, and interaction information between the user and the interactive device;
[0088] The transient memories are organized according to a preset standard to obtain short-term memories;
[0089] The short-term memory is incrementally processed to obtain behavioral sequence data, and the long-term memory is generated based on the behavioral sequence data.
[0090] In this embodiment, the system first collects various transient memories from the vehicle's various sensors and devices. Transient memories refer to real-time information related to user-device interaction collected by the devices within a short period of time. Transient memories can include device information of the interactive device, user characteristic information, and interaction information between the user and the interactive device. Device information refers to the hardware or software attributes of the interactive device, such as the device model, system version, or sensor status. Interaction information refers to the operation or input / output behavior between the user and the interactive device, such as the user clicking on the interactive device or sending voice commands to the interactive device. Characteristic information refers to the user's attributes or behavioral characteristics, such as the user's name, voiceprint features, preferences, or historical behavior records.
[0091] When the interactive device is a vehicle, the instantaneous memory includes at least one of the vehicle's vehicle control signals, sensor signals, and image information, voice information, and text information received by the vehicle.
[0092] For example, when the interactive device is a vehicle's infotainment system, the interactive information can include image information, voice information, and text information received by the vehicle; the device information can include vehicle control signals and sensor signals. Specifically, image information can include images captured by the vehicle's onboard camera; voice information can be voice commands used by the user to interact with the vehicle; text information can be text input by the user in a system such as an in-vehicle navigation system; vehicle control signals can be signals such as steering, acceleration, and braking; and sensor signals can be signals collected by sensors in the vehicle, such as temperature, pressure, and GPS positioning signals.
[0093] Whenever an interactive device acquires short-term memory, it can perform standardization processing in real time according to preset specifications. Standardization includes cleaning, normalizing, and aligning the signals to ensure data consistency and usability. For example, images captured by in-vehicle cameras are processed into a uniform format, and output data from in-vehicle sensors are calibrated or aligned to the timeline. This processed data constitutes the short-term memory of the interactive device.
[0094] Next, the interactive device incrementally processes the short-term memory, continuously updating and integrating existing information in the system based on newly acquired short-term memories. The incrementally processed short-term memories are then categorized according to different dimensions, such as time, event type, or user behavior habits. Through incremental processing, short-term memories can be gradually transformed into long-term memories.
[0095] Specifically, after acquiring short-term memory, the system can incrementally process the user's recent behaviors in short-term memory to form user behavior sequence data. The behavior sequence data records the user's operations and habits in the short term, and the system can then extract long-term memory from this behavior sequence data.
[0096] In the above scheme, the instantaneous memory of the interactive device is acquired, standardized into short-term memory, and then incrementally processed and classified to form the long-term memory of the interactive device. In this way, by combining the long-term memory and the user's query statements, the interactive device can provide a more accurate and intelligent interactive experience based on the user's historical behavior and the interactive device's historical records.
[0097] As an optional embodiment, the standardization process of the transient memory to obtain short-term memory may include at least one of the following:
[0098] Clean up noisy or invalid data in the transient memory;
[0099] The instantaneous memory is normalized.
[0100] The repeated data in the instantaneous memory are aggregated.
[0101] In this embodiment, the process of standardizing the transient memory may include data cleaning, normalization, or aggregation.
[0102] Data cleaning, in particular, involves first removing noisy or invalid data from the system's short-term memory. Short-term memory can be affected by environmental noise, sensor malfunctions, or other external factors, resulting in inaccurate or irrelevant information. The cleaning process identifies and removes this interfering data, making the retained short-term memory more accurate. For example, images captured by in-vehicle cameras may contain overexposed areas, and vehicle control signals may generate abnormal readings due to faulty sensors; these will be filtered out through the cleaning process.
[0103] Normalization is a process that standardizes signals with similar dimensions or correlations in transient memory, enabling comparisons and calculations under the same benchmark. For example, it can unify the units of speed data from different speed sensors to km / h, and the units of temperature data from different temperature sensors to degrees Celsius, making signals of the same dimension from different sources comparable. Transient memory may originate from different sensors, and its units, time frequencies, or numerical ranges may differ; normalization ensures that this information is processed under a unified standard.
[0104] Aggregation processing merges duplicate information to avoid data redundancy and improve storage and processing efficiency.
[0105] In this way, by standardizing the transient memory, including cleaning noisy data, normalizing, and aggregating duplicate data, the system can generate high-quality short-term memory.
[0106] As an optional embodiment, after obtaining the input query statement, the method further includes:
[0107] Update the long-term memory according to the query statement.
[0108] In this embodiment, after the user enters a query, the system not only retrieves the target memory to generate a response, but also updates the long-term memory in the memory platform based on the query. Specifically, the system records the user's latest query and adjusts its understanding of user preferences or behavioral patterns based on the content of the query.
[0109] For example, the system can save a user's query history to long-term memory as new information about the user's behavior. If the system discovers new interests or behavioral patterns in the user's query history, it can incorporate this newly discovered information into long-term memory to adjust its understanding of user preferences. Furthermore, the system can update or overwrite existing long-term memory based on the content of the query.
[0110] In this way, the system can continuously learn new user behaviors and interests and incorporate them into long-term memory, thereby improving the real-time nature and accuracy of long-term memory.
[0111] As an optional embodiment, the long-term memory includes a summary, which includes an overall summary and a topic summary. After obtaining the input query statement, the method further includes:
[0112] Determine the topic tags corresponding to the query statement;
[0113] Query the first topic summary corresponding to the topic tag;
[0114] The overall summary and the first topic summary are updated according to the query statement.
[0115] In this embodiment, long-term memory includes summaries, which in turn include overall summaries and topic summaries. The topic summaries are a summary and condensation of information related to a specific topic, encompassing the user's long-term behavior and preferences. The overall summary is a global summary generated by integrating the content of all topic summaries, covering information from all topics.
[0116] Specifically, after a user enters a query, the system analyzes the topic tags of the query using a topic classification model. Then, it retrieves the historical first topic summaries under those topic tags from long-term memory. After determining the topic tags and first topic summaries, the system integrates the user's query with the historical first topic summaries from long-term memory to form a new first topic summary. This new first topic summary then overwrites the historical first topic summaries.
[0117] Specifically, the integration process can involve inputting the query statement and the first topic summary into the language big model, whereby the language big model summarizes the content of the query statement and the first topic summary to generate a new first topic summary.
[0118] Furthermore, the overall summary can be updated based on the query statement. The query statement and all topic summaries can be input into the language big model, which will then summarize the content of the query statement and all topic summaries to generate a new overall summary.
[0119] For example, if a user asks "What were the results of recent basketball games?", the system will identify the topic as basketball. The system will retrieve historical topic summaries and overall summaries related to basketball. A historical topic summary for basketball might be "The Basketball World Cup is underway, and Team A won their last game." A historical overall summary might be "Users are interested in current affairs and entertainment." Then, the system can retrieve the query results for "recent basketball games," which might be "Team B won their last game." The system can then input this query result along with the topic summaries from long-term memory into the language model. The language model will output a new topic summary for "basketball": "The Basketball World Cup is underway, and both Team A and Team B won their last games." Finally, all topic summaries and the query statement will be input into the language model, which will output a new overall summary: "Users are interested in sports, current affairs, and entertainment."
[0120] The above methods, through the identification of topic tags and the querying and fusion of topic summaries, enable dynamic updating of summaries in long-term memory, thus ensuring the real-time nature and accuracy of the summaries.
[0121] As an optional embodiment, the long-term memory includes preference data, which includes primary topic preferences and secondary topic preferences. After obtaining the input query statement, the method further includes:
[0122] Determine the primary and secondary topic tags corresponding to the query statement;
[0123] Obtain the first-level preference data of the first-level topic tags and the second-level preference data corresponding to the second-level topic tags;
[0124] Update the primary preference data and the secondary preference data according to the query statement.
[0125] In this embodiment, long-term memory includes preference data. Preference data refers to specific data dynamically generated by the system for each user based on their behavior, interests, historical interaction records, and other information for personalized services. This preference data helps the system understand the user's habits, concerns, interests, etc., thereby providing more accurate information recommendations during the interaction process. Preference data can include two main categories: topic preferences and entity word preferences.
[0126] Topic preference refers to a user's interest in certain broad topics or areas. The system determines a user's depth of interest in a topic based on their historical queries, number of interactions, and the time of their last interaction. Topics can include sports, technology, and entertainment. Entity preference refers to a user's interest in specific things, such as names of people, places, brands, and products. Topics can be divided into primary and secondary topics. Primary topics are broad classifications of content areas, representing the overall direction of the topic, such as "sports"; secondary topics are specific sub-categories under primary topics, representing specific subcategories of the topic, such as "basketball".
[0127] In this embodiment, after the user enters a query, the system first uses a topic classification model to determine which topic tag the query belongs to. After determining the topic tag, it retrieves preference data related to that topic tag from long-term memory. Once the preference data is obtained, the system updates the preference data based on the user's query.
[0128] For example, a user might input the query "What were the results of recent basketball games?" with the input time being 2024-03-05. The system can use a topic classification model to find that the primary topic tag for this query is "sports," and the secondary topic tag is "basketball." Then, the system can obtain primary preference data for the primary topic tag and secondary preference data for the secondary topic tag. The preference data can include the total number of queries under the corresponding topic tag, the most recent query time, and the memory strength of that topic tag.
[0129] For example, the original primary preference data for the primary topic tag "sports" could include: total number of queries (25), last query time (March 2, 2024), and memory strength of 0.3. The original secondary preference data for the secondary topic tag "basketball" could include: total number of queries (15), last query time (March 1, 2024), and memory strength of 0.7.
[0130] Therefore, updating the primary and secondary preference data based on the query statement can specifically include updating the total number of queries, the most recent query time, and the primary memory strength. The total number of queries in both the primary and secondary preference data can be incremented by 1. For example, the query count for "sports" is updated from 25 to 26, and the query count for "basketball" is updated from 15 to 16. The most recent query time in both the primary and secondary preference data is updated to 2024-03-05. Finally, the memory strength in the primary and secondary preference data is recalculated using the Ebbinghaus forgetting curve.
[0131] In this way, the solution can dynamically parse query statements and update the preference data of primary and secondary topics in real time, making long-term memory more accurately reflect the user's latest interests and behavioral changes.
[0132] Specifically, in some embodiments, the secondary preference data includes memory strength, and updating the secondary preference data according to the query statement includes:
[0133] Determine the relevance between the query statement and the secondary topic tags;
[0134] Get the total number of interactions associated with the secondary topic tags, and a first time interval, wherein the first time interval is the time interval between the time when the query statement is obtained and the time when the interaction information associated with the secondary topic tags is received last time;
[0135] The memory strength is updated based on the correlation, the total number of interactions, and the first time interval.
[0136] In this embodiment, the Ebbinghaus forgetting curve can be used to calculate the user's memory strength for each topic. By calculating the Ebbinghaus forgetting curve, the system can dynamically adjust the user's preference data. The Ebbinghaus forgetting curve describes the gradual decline of human memory over time, indicating that memory decreases rapidly in the initial period, and then the rate of forgetting gradually slows down.
[0137] For example, memory strength can be calculated using the following formula:
[0138] R=e -t / s
[0139] Where t is the first time interval between the time the query statement is retrieved and the time the interaction information associated with the second-level topic tag was received last time; s is the total number of times the user has interacted with the second-level topic tag before; the more interactions, the stronger the memory and the slower the decay. R is used to characterize the user's memory strength of the second-level topic tag.
[0140] In the above method, the system determines the corresponding secondary topic tags through the user's query statement, and analyzes the relevance between the query statement and the secondary topic tags, the total number of user interactions and the time interval, and dynamically updates the memory strength, which can accurately capture changes in the user's interest in specific topics.
[0141] Based on the memory-based interaction method provided in the above embodiments, this application also provides specific implementations of a memory-based interaction device. Please refer to the following embodiments.
[0142] First see Figure 2The memory-based interactive device 200 provided in this application includes the following modules:
[0143] The first acquisition module 201 is used to acquire the input query statement;
[0144] The first determining module 202 is used to determine the target memory associated with the query statement through a language big model, wherein the target memory includes at least one of long-term memory, short-term memory, and memory summary, and the memory summary is the content obtained by the language big model from the long-term memory and the short-term memory;
[0145] The fusion module 203 is used to fuse the target memory and the query statement to obtain prompt words;
[0146] The response module 204 is used to input the prompt words into the trained language model and obtain the response statement output by the language model.
[0147] In this embodiment, since the target memory includes the user's long-term memory, short-term memory, and memory summary, it can fully reflect the user's historical behavior, preferences, and interaction records, as well as the current context information. Therefore, by combining the target memory with the query statement to obtain prompt words, and inputting the prompt words into the language big model to obtain the response statement, the model can not only rely on the current input when generating the response, but also combine the user's short-term context information and long-term interests and background. As a result, the language big model is more accurate in understanding user needs, can reduce misunderstandings or context breaks, and can generate responses that are more in line with individual needs based on the user's historical behavior, thereby significantly improving the accuracy of the response.
[0148] As one implementation of this application, the interactive device 200 of the memory system may further include:
[0149] The second acquisition module is used to acquire transient memory, which includes at least one of the following: device information of the interactive device, user characteristic information, and interaction information between the user and the interactive device.
[0150] The organizing module is used to organize the transient memory according to a preset specification to obtain short-term memory;
[0151] The processing module is used to incrementally process the short-term memory to obtain behavioral sequence data, and generate the long-term memory based on the behavioral sequence data.
[0152] As one implementation of this application, the above-mentioned sorting module is specifically used for:
[0153] Clean up noisy or invalid data in the transient memory;
[0154] The instantaneous memory is normalized.
[0155] The repeated data in the instantaneous memory are aggregated.
[0156] As one implementation of this application, the interactive device 200 of the memory system may further include:
[0157] The second determining module is used to determine the topic tags corresponding to the query statement;
[0158] The query module is used to query the first topic summary corresponding to the topic tag;
[0159] The first update module is used to update the overall summary and the first topic summary according to the query statement.
[0160] As one implementation of this application, the interactive device 200 of the memory system may further include:
[0161] The third determining module is used to determine the first-level topic tags and second-level topic tags corresponding to the query statement;
[0162] The third acquisition module is used to acquire the first-level preference data of the first-level topic tags and the second-level preference data corresponding to the second-level topic tags;
[0163] The second update module is used to update the primary preference data and the secondary preference data according to the query statement.
[0164] As one implementation of this application, the second update module described above is specifically used for:
[0165] Determine the relevance between the query statement and the secondary topic tags;
[0166] Get the total number of interactions associated with the secondary topic tags, and a first time interval, wherein the first time interval is the time interval between the time when the query statement is obtained and the time when the interaction information associated with the secondary topic tags is received last time;
[0167] The memory strength is updated based on the correlation, the total number of interactions, and the first time interval.
[0168] As one implementation of this application, the interactive device 200 of the memory system may further include:
[0169] The matching module is used to obtain N memory fragments that match the query statement through a target matching method, wherein the target matching method includes at least one of keyword matching and vector matching, and N is a positive integer;
[0170] The sorting module is used to sort the N memory segments according to their relevance to the query statement, and select the P memory segments with the highest relevance from the sorted N memory segments, where P is a positive integer less than or equal to N.
[0171] The reasoning module is used to input the P memory fragments, the query statement, and the short-term memory associated with the query statement into the language big model to obtain K memory fragments obtained by the language big model based on the thought chain reasoning, wherein the P memory fragments include the K memory fragments, and K is a positive integer less than P;
[0172] The summary module is used to summarize the short-term memory associated with the K memory fragments and the query statement through the language big model to obtain the memory summary.
[0173] The interactive device based on a memory system provided in this embodiment of the invention can implement the steps in the above method embodiments, and will not be repeated here to avoid repetition.
[0174] Figure 3 A schematic diagram of the hardware structure of the interactive device based on the memory system provided in an embodiment of this application is shown.
[0175] The interactive device based on the memory system may include a processor 1001 and a memory 1002 storing computer program instructions.
[0176] Specifically, the processor 1001 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0177] Memory 1002 may include mass storage for data or instructions. For example, and not limitingly, memory 1002 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 1002 may include removable or non-removable (or fixed) media. Where appropriate, memory 1002 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 1002 is non-volatile solid-state memory.
[0178] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.
[0179] The processor 1001 reads and executes computer program instructions stored in the memory 1002 to implement any of the memory-based interaction methods in the above embodiments.
[0180] In one example, the memory-based interactive device may further include a communication interface 1003 and a bus 1010. For example, Figure 3 As shown, the processor 1001, memory 1002, and communication interface 1003 are connected through bus 1010 and complete communication with each other.
[0181] The communication interface 1003 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0182] Bus 1010 includes hardware, software, or both, that couples components of a memory-based interactive device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 1010 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0183] The memory-based interactive device can be based on the above embodiments to realize the above-described memory-based interactive method and apparatus.
[0184] Furthermore, in conjunction with the memory-based interaction method in the above embodiments, this application embodiment can provide a computer storage medium for implementation. This computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the memory-based interaction methods in the above embodiments and achieve the same technical effect. To avoid repetition, further details are omitted here. The aforementioned computer-readable storage medium may include non-transitory computer-readable storage media, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc., and is not limited thereto.
[0185] In addition, this application also provides a vehicle including computer program instructions, which, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0186] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0187] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0188] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0189] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and vehicles according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0190] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. An interaction method based on a memory system, characterized in that, The method includes: Get the input query statement; The target memory associated with the query statement is determined by the language big model, wherein the target memory includes at least one of long-term memory, short-term memory, and memory summary, and the memory summary is the content obtained by the language big model from the long-term memory and the short-term memory; The target memory and the query statement are merged to obtain the prompt words; The prompt words are input into the trained language model to obtain the response statement output by the language model. The long-term memory includes summary and preference data. The summary includes an overall summary and a topic summary. The preference data includes first-level topic preferences and second-level topic preferences. After obtaining the input query statement, the method further includes: Determine the primary and secondary topic tags corresponding to the query statement; Obtain the first-level preference data of the first-level topic tags and the second-level preference data corresponding to the second-level topic tags; Update the primary preference data and the secondary preference data according to the query statement; Wherein, the secondary preference data includes memory strength, and updating the secondary preference data according to the query statement includes: Determine the relevance between the query statement and the secondary topic tags; Get the total number of interactions associated with the secondary topic tags, and a first time interval, wherein the first time interval is the time interval between the time when the query statement is obtained and the time when the interaction information associated with the secondary topic tags is received last time; The memory strength is updated using the Ebbinghaus forgetting curve based on the correlation, the total number of interactions, and the first time interval.
2. The interaction method based on a memory system according to claim 1, characterized in that, Before determining the target memory associated with the query statement through a large language model, the method further includes: Acquire transient memory, which includes at least one of the following: device information of the interactive device, user characteristic information, and interaction information between the user and the interactive device; The transient memories are organized according to a preset standard to obtain short-term memories; The short-term memory is incrementally processed to obtain behavioral sequence data, and the long-term memory is generated based on the behavioral sequence data.
3. The interaction method based on a memory system according to claim 2, characterized in that, The process of organizing the instantaneous memory according to a preset standard includes at least one of the following: Clean up noisy or invalid data in the transient memory; The instantaneous memory is normalized. The repeated data in the instantaneous memory are aggregated.
4. The interaction method based on a memory system according to claim 2, characterized in that, When the interactive device is a vehicle, the instantaneous memory includes at least one of the vehicle's vehicle control signals, sensor signals, and image information, voice information, and text information received by the vehicle.
5. The interaction method based on a memory system according to claim 1, characterized in that, After obtaining the input query statement, the method further includes: Determine the topic tags corresponding to the query statement; Query the first topic summary corresponding to the topic tag; The overall summary and the first topic summary are updated according to the query statement.
6. The interaction method based on a memory system according to claim 1, characterized in that, The method for generating the memory summary includes: N memory fragments that match the query statement are obtained through a target matching method, wherein the target matching method includes at least one of keyword matching and vector matching, and N is a positive integer; The N memory segments are sorted according to their relevance to the query statement, and the P memory segments with the highest relevance are selected from the sorted N memory segments, where P is a positive integer less than or equal to N. The P memory fragments, the query statement, and the short-term memory associated with the query statement are input into the language big model to obtain K memory fragments obtained by the language big model based on thought chain reasoning, wherein the P memory fragments include the K memory fragments, and K is a positive integer less than P; The memory summary is obtained by summarizing the short-term memories associated with the K memory fragments and the query statement using the large language model.
7. An interactive device based on a memory system, characterized in that, The device includes: The first acquisition module is used to acquire the input query statement; The first determining module is used to determine the target memory associated with the query statement through a language big model, wherein the target memory includes at least one of long-term memory, short-term memory, and memory summary, and the memory summary is the content obtained by the language big model from the long-term memory and the short-term memory; The fusion module is used to fuse the target memory and the query statement to obtain prompt words; The response module is used to input the prompt words into the trained language model and obtain the response statement output by the language model. The third determining module is used to determine the first-level topic tags and second-level topic tags corresponding to the query statement; The third acquisition module is used to acquire the first-level preference data of the first-level topic tags and the second-level preference data corresponding to the second-level topic tags; The second update module is used to update the primary preference data and the secondary preference data according to the query statement; The second update module is specifically used to: determine the relevance between the query statement and the secondary topic tags; Get the total number of interactions associated with the secondary topic tags, and a first time interval, wherein the first time interval is the time interval between the time when the query statement is obtained and the time when the interaction information associated with the secondary topic tags is received last time; The memory strength is updated using the Ebbinghaus forgetting curve based on the correlation, the total number of interactions, and the first time interval.
8. An interactive device based on a memory system, characterized in that, The memory-based interactive device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the memory-based interactive method as described in any one of claims 1-6.
9. A computer storage medium, characterized in that, The computer storage medium stores computer program instructions, which, when executed by a processor, implement the memory-based interactive method as described in any one of claims 1-6.
10. A vehicle, characterized in that, The vehicle includes at least one of the device as described in claim 7, the equipment as described in claim 8, and the medium as described in claim 9.