Active memory method and device, electronic equipment and storage medium
By processing multimodal data in real time on the device side to generate contextual and long-term memory data, the contradiction between privacy and timeliness, as well as the obstacles to collaboration brought about by cloud deployment, are resolved, and an efficient, intelligent, and personalized vehicle interaction experience is achieved.
Patent Information
- Application Number
- CN202511425122.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-02-03
AI Technical Summary
Existing long short-term memory systems suffer from privacy and timeliness conflicts, lack of active memory capabilities, and obstacles to cross-agent collaboration due to cloud deployment, making it difficult to achieve efficient, intelligent, and personalized services.
Multimodal data is acquired in real time at the device side, semantic and feature analysis is performed, and contextual memory data and long-term memory data are generated for vehicle interaction response.
It avoids the privacy leakage risks associated with cloud transmission, improves the timeliness of memory generation, enhances the vehicle's ability to understand user preferences, scenarios, and interaction history, and achieves a more coherent, intelligent, and personalized interactive experience.
Smart Images

Figure CN121457516A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of edge-side large model technology, and in particular to an active memory method, device, electronic device and storage medium. Background Technology
[0002] The long short-term memory system of the large car model is an independent and personalized module that can be used to store, update and retrieve user preferences, interaction history, scene data and other long short-term memories, enabling the large model to achieve a more coherent, intelligent and personalized interactive experience.
[0003] In existing technologies, long short-term memory systems mainly rely on cloud deployment. On the one hand, multimodal data from the device needs to be transmitted to the cloud, which poses a risk of privacy leakage. Furthermore, the device cannot actively generate long-term memories based on local data, nor can it identify scenes and people in real time. On the other hand, the memories of each agent (intelligent agent) on the device are independent, which leads to problems such as information not being synchronized or data being stored repeatedly, which can easily cause fragmentation of the service experience.
[0004] Therefore, existing long short-term memory systems suffer from privacy and timeliness conflicts, lack of active memory capabilities, obstacles to cross-agent collaboration, and insufficient utilization of edge resources due to cloud deployment, making it difficult to achieve efficient, intelligent, and personalized services. Summary of the Invention
[0005] This application provides an active memory system, method, apparatus, electronic device, and storage medium to at least solve the problem in related technologies where long short-term memory systems, due to cloud deployment, struggle to achieve efficient, intelligent, and personalized services. The technical solution of this application is as follows:
[0006] According to a first aspect of the embodiments of this application, an active memory method is provided, comprising:
[0007] Real-time acquisition of multimodal data, which is associated with the in-vehicle user and the vehicle's current situation;
[0008] Semantic analysis is performed on the multimodal data to obtain contextual memory data, which carries a corresponding timestamp.
[0009] Feature analysis is performed on the contextual memory data to obtain long-term memory data, enabling the vehicle to reason and respond to received interactive commands based on the long-term memory data.
[0010] According to a second aspect of the embodiments of this application, an active memory device is provided, comprising:
[0011] The acquisition module is used to acquire multimodal data in real time, and the multimodal data is associated with the in-vehicle user and the vehicle's current situation;
[0012] The scenario analysis module is used to perform semantic analysis on the multimodal data to obtain scenario memory data, which carries a corresponding timestamp.
[0013] The feature analysis module is used to perform feature analysis on the contextual memory data to obtain long-term memory data, so that the vehicle can reason and respond to received interactive instructions based on the long-term memory data.
[0014] According to a third aspect of the embodiments of this application, an active memory electronic device is provided, comprising:
[0015] processor;
[0016] Memory used to store the processor's executable instructions;
[0017] The processor is configured to execute the instructions to implement the active memory method described in any one of the claims.
[0018] According to a fourth aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein when the instructions in the computer-readable storage medium are executed by a processor of an active memory electronic device, the active memory electronic device is enabled to perform the active memory method described in any one of the claims.
[0019] According to a fifth aspect of the embodiments of this application, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the active memory method described in any one of the claims.
[0020] The technical solutions provided by the embodiments of this application bring at least the following beneficial effects:
[0021] By processing multimodal data locally in real time on the device side, the risk of privacy leakage caused by cloud transmission is avoided, and the timeliness of memory generation is improved. At the same time, long-term memory is built based on context and feature analysis, which enhances the vehicle's ability to understand user preferences, scenarios and interaction history, and helps to achieve a more coherent, intelligent and personalized interactive experience. This effectively alleviates the problems of privacy and timeliness contradictions, lack of active memory ability and cross-agent collaboration obstacles caused by cloud deployment in existing technologies.
[0022] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application, and do not constitute an undue limitation of this application.
[0024] Figure 1 This is a flowchart illustrating an active memory method according to an exemplary embodiment.
[0025] Figure 2 This is a schematic diagram illustrating a memory hierarchical architecture and data flow relationship according to an exemplary embodiment.
[0026] Figure 3 This is a flowchart illustrating a memory processing system in a vehicle-mounted scenario according to an exemplary embodiment.
[0027] Figure 4 This is a schematic diagram illustrating a system architecture and data interaction process according to an exemplary embodiment.
[0028] Figure 5 This is a block diagram illustrating an active memory device according to an exemplary embodiment.
[0029] Figure 6 This is a block diagram illustrating an electronic device for active memory according to an exemplary embodiment.
[0030] Figure 7 This is a block diagram illustrating an active memory device according to an exemplary embodiment. Detailed Implementation
[0031] To enable those skilled in the art to better understand the technical solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0032] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0033] Figure 1 This is a flowchart illustrating an active memory method according to an exemplary embodiment, such as... Figure 1 As shown, this active memory method includes:
[0034] In step S11, multimodal data is acquired in real time, and the multimodal data is associated with the in-vehicle user and the vehicle's current situation.
[0035] In this step, the vehicle can continuously collect multimodal data related to user behavior and the vehicle environment through hardware interfaces such as on-board sensor arrays, microphone arrays, cameras, and the vehicle CAN bus.
[0036] For example, multimodal data can include, but is not limited to, the following:
[0037] Voice data is captured through the microphone, showing the natural language communication of users inside the vehicle, including dialogue content, tone of voice, speech rate, and emotional state.
[0038] Visual data is obtained through cameras, including users' facial expressions, body movements, gaze direction, and identity characteristics, while also recording the distribution of people inside the vehicle and seat occupancy.
[0039] Environmental data, combined with information from temperature and humidity sensors, GPS positioning, window status, vehicle speed, and light intensity, allows for a comprehensive perception of the vehicle's external environment and internal atmosphere.
[0040] Interactive behavior data, including touch operations, voice command triggers, and in-vehicle system usage habits.
[0041] By synchronously collecting these multimodal data that are strongly correlated with user behavior and scene dynamics, we can provide comprehensive and relevant original materials for subsequent memory generation, avoiding the disconnect between memory and actual situation due to data loss.
[0042] In step S12, semantic analysis is performed on the multimodal data to obtain contextual memory data, which carries a corresponding timestamp.
[0043] In this step, multimodal fusion and natural language understanding technologies can be used to perform semantic parsing and structured processing on multimodal data, generating contextual memory data with timestamps, thereby leveraging semantic understanding technology to uncover the contextual relationships behind the data.
[0044] For example, user command intent and dialogue topics can be extracted from voice data, user actions and the distribution of people in the vehicle can be identified from image data, and user behavior logic can be sorted out from operation data. At the same time, corresponding timestamps can be added to the processed data to transform scattered multimodal information into contextual memory data with time dimension and clear semantics, and so on.
[0045] These timestamped contextual memory data can clearly record the interaction process between users and vehicles at different time points and the contextual state of the vehicles, forming traceable contextual memory fragments and laying a structured foundation for the generation of subsequent long-term memories.
[0046] In step S13, feature analysis is performed on the contextual memory data to obtain long-term memory data, so that the vehicle can reason and respond to the received interactive instructions based on the long-term memory data.
[0047] In this step, further feature mining can be performed on the structured contextual memory data to extract core features that are stable and regular from multiple sets of contextual memory data. For example, by analyzing contextual memory data from multiple time stamps, we can identify users' recurring behavioral preferences (such as frequently used air conditioning temperatures, preferred music genres), fixed interaction habits (such as specific command expressions), and demand patterns strongly associated with the context (such as navigation preferences during commuting hours), and integrate these feature information with long-term reference value into long-term memory data.
[0048] In this way, when the vehicle receives the user's interaction command, it can directly call on this locally stored long-term memory data, combine it with the current command context to reason, accurately judge the user's potential needs, and then generate a response that fits the user's personalized habits, ensuring the consistency and intelligence of the interaction.
[0049] In one implementation, in step S12, semantic analysis is performed on the multimodal data to obtain contextual memory data, including:
[0050] According to the preset time interval, the multimodal data is divided into time slices to obtain multiple groups. The multimodal data of each group includes: the dialogue information, location information, human characteristic information of the people in the vehicle, as well as the environmental status information of the people and vehicles and the current application operation information.
[0051] Data fusion is performed on the multimodal data of each group to obtain the contextual memory data and corresponding timestamps for each group, wherein the contextual memory data includes: associated personal characteristics and contextual memory information.
[0052] In this implementation, the semantic analysis process of multimodal data first involves dividing the real-time collected continuous multimodal data into multiple independent groups by using a preset time interval (such as 30 seconds, 60 seconds, etc., which are fixed durations that conform to the interaction rhythm of in-vehicle scenarios).
[0053] This time-slicing processing is not a simple mechanical division, but rather an interval setting that combines the natural cycle of user interaction and scenario changes in the in-vehicle scenario. This ensures that the multimodal data in each group can fully reflect the user behavior and vehicle scenario within a short period of time. For example, a group can fully cover the continuous process of "user issuing voice command - vehicle response - user supplementary operation", avoiding the fragmentation of scenario information due to excessive time span or the data redundancy due to excessively short interval.
[0054] In this example, the multimodal data for each group may include: conversation information, location information, personal characteristics information, vehicle and human environment status information, and current application operation information.
[0055] After time segmentation, refined semantic analysis is conducted for each group. Specifically, entity recognition technology can be used to extract core information entities from the multimodal data within the group. For example, from voice data, it can identify the names of people, locations, and key commands mentioned by the user (such as "Teacher Zhou," "shopping mall," "play music"); from image data, it can identify the identity features and postures of people in the vehicle (such as "facial features of the passenger" and "user adjusting the air conditioning knob"); and from vehicle status data, it can identify key parameters (such as "current vehicle speed" and "air conditioning temperature").
[0056] Simultaneously, correlation analysis can be performed to connect the extracted independent entities according to logical relationships. For example, "facial features of the co-pilot" can be associated with "Teacher Zhou", and "user adjusting the air conditioning knob" can be associated with "current air conditioning temperature" to clarify the contextual relationships between different entities.
[0057] In this way, each group will generate a corresponding structured information containing core entities and entity relationships, namely contextual memory data, and automatically attach the timestamp corresponding to the group (such as "2024-10-05 09:15-09:16"), so that each contextual memory can be accurately anchored to a specific time node, which not only ensures the integrity and logic of contextual information, but also provides clear and traceable basic materials for feature extraction of subsequent long-term memory data.
[0058] In one implementation, in step S13, feature analysis is performed on the episodic memory data to obtain long-term memory data, including:
[0059] Multi-dimensional and multi-objective feature analysis was performed on the episodic memory data to obtain the feature data corresponding to each episodic memory data.
[0060] Feature fusion is performed on the feature data to obtain long-term memory data.
[0061] In this implementation, step S13 can perform multi-dimensional and multi-objective feature analysis on each piece of contextual memory data carrying a timestamp.
[0062] Among them, multi-dimensional analysis refers to mining key information at different levels from contextual memory data. For example, from contextual memories related to "user-vehicle interaction", user behavior dimension features (such as "using navigation at 8 o'clock every day" and "preferring voice control of air conditioning"), preference dimension features (such as "often playing Jay Chou songs" and "air conditioning temperature is fixed at 24℃"), as well as person association dimension features (such as "the passenger in the front seat is often Mr. Zhou" and "when interacting with Mr. Zhou, polite words such as 'trouble' are often used") and scene dependence dimension features (such as "the seat heating is turned on when riding in the car in the rain" and "the windows are tended to be closed in high-speed scenarios").
[0063] Multi-objective analysis selectively filters features with long-term reference value. For example, it distinguishes between "occasional temporary actions of users" (such as adjusting the music style when a friend is traveling) and "stable and repetitive habitual behaviors" (such as choosing the same route when commuting for three consecutive months). The latter is given priority as effective feature data, ensuring that the feature data corresponding to each contextual memory data is both comprehensive and practical.
[0064] After acquiring the feature data of each context memory data, feature fusion processing can be further carried out. Based on logical association and priority ranking, the scattered feature data can be integrated into structured long-term memory data.
[0065] For example, features under the same dimension can be categorized and merged. For instance, features such as "prefer to play music by artist AA" and "prefer to play music by artist AA" can be merged into "user's core music preference: works by artist AA".
[0066] Alternatively, a mapping between features of different dimensions can be established. For example, the "person-related feature (Teacher Zhou takes a car)" can be bound to the "scene preference feature (Teacher Zhou turns on the ambient light when taking a car)" to form the scenario-behavior related feature of "Teacher Zhou takes a car → turns on the ambient light".
[0067] In this way, the obtained long-term memory data not only covers stable features such as users' core preferences, fixed habits, and interpersonal relationships, but also includes logical connections between features. It can comprehensively and accurately reflect users' personalized needs and behavioral patterns, providing reliable long-term memory support for vehicles to make inference responses based on interactive commands.
[0068] In one implementation, after performing feature analysis on the episodic memory data in step S13 to obtain long-term memory data, the method further includes:
[0069] Long-term memory data is stored in a database, and the secondary cache is updated based on the long-term memory data;
[0070] Upon receiving a command from the large model to access long-term memory data, the long-term memory data is transferred to the large model so that the large model can reason and respond to the received interactive command based on the long-term memory data.
[0071] In this implementation, after obtaining long-term memory data through feature analysis in step S13, this data can first be stored in a preset database. This ensures that the cache prioritizes core features frequently relied upon by the user. Simultaneously, the secondary cache can be updated based on the long-term memory data. The secondary cache simplifies the context length of memory, ensuring that the memory context contains rich information without exceeding the model context limit while meeting the constraints of the large model context on the client side. The secondary cache is divided into four cache areas: 1. Global topic summary; 2. M-turn current dialogue; 3. Current user feature preferences; 4. Information cache from the last recall; it includes both short-term and long-term memory.
[0072] At the same time, data structuring can be used to organize long-term memory data into a format that is easy to retrieve quickly (such as storing it according to dimensions such as "user preferences", "personal relationships" and "scene adaptation"). This not only avoids the call delay caused by disordered accumulation of data in the cache, but also eliminates the need to transfer data to the cloud through the local storage characteristics of the cache, further protecting user data privacy and security, and laying the foundation for rapid response to subsequent call requests.
[0073] When a long-term memory data retrieval command is received from a large model, data transmission is completed through a pre-defined standardized interface. This interface predefines a unified data interaction format and transmission protocol, enabling efficient data extraction and transmission and avoiding data transmission delays or information loss due to interface compatibility issues.
[0074] After acquiring this long-term memory data, the large model can combine the currently received user interaction instructions (such as "help me plan my travel route today") to quickly associate the user's historical preferences and behavior patterns. For example, based on the feature in long-term memory that "users prefer to take expressways to avoid congestion on weekdays", it can infer a navigation plan that conforms to the user's habits, thereby generating personalized and accurate response results. This ensures that the entire interaction process is both efficient and meets the user's needs, effectively making up for the problems of edge memory retrieval delay and difficulty in ensuring data privacy in existing technologies.
[0075] In one implementation, after storing long-term memory data in the cache, the method further includes:
[0076] Update the shared state in the cache based on long-term memory data;
[0077] Transferring long-term memory data to a large model includes:
[0078] Transfer the shared state to the large model.
[0079] In this implementation, after storing long-term memory data in the cache, the existing shared state in the cache can be dynamically updated based on this newly generated long-term memory data. The shared state is the core component providing a unified memory view for all modules on the client side (including large models). Its update process involves supplementing or optimizing the corresponding dimensions in the shared state by combining the feature types and priorities of the new long-term memory data.
[0080] For example, if new long-term memory data adds the preference feature of "users frequently choosing new energy charging stations as waypoints recently", the system will integrate this feature into the "user travel preference" dimension in the shared state, while retaining the original stable features such as "commuting route preference". If new data corrects old features (such as "user's air conditioning temperature preference is adjusted from 24℃ to 23℃"), the corresponding parameters in the shared state will be updated synchronously to ensure that the shared state always reflects the user's latest personalized needs and behavior patterns, and maintains the consistency and integrity of the data, avoiding service deviations caused by some modules using old data.
[0081] When transferring long-term memory data to a large model, the updated shared state can be directly transmitted as a data carrier through a pre-defined standardized interface. This pre-defined standardized interface has a pre-defined data interaction format adapted to the large model, which can completely and efficiently synchronize various types of long-term memory data covered in the shared state to the large model.
[0082] This transmission method eliminates the need for large models to retrieve scattered long-term memory data from the cache one by one. Instead, it obtains a structured and standardized complete memory view all at once through shared state. This reduces the frequency and latency of data transmission and avoids inference bias in large models caused by data fragmentation.
[0083] For example, when the large model needs to process the user's interactive command "prepare the in-car environment for me", the shared state obtained through the interface will simultaneously contain multi-dimensional long-term memory data such as "user's seat position preference", "frequently used air conditioning temperature" and "ambient lighting settings associated with the current passenger". The large model can directly and quickly infer based on this integrated information to generate an in-car environment adjustment scheme that conforms to the user's overall habits, further improving the efficiency and accuracy of the interactive response.
[0084] In one implementation, the method further includes:
[0085] Store the most recent multimodal data within a preset time period and a preset number of scene memory data in the cache;
[0086] Upon receiving a call instruction from the large model for multimodal data and / or contextual memory data, the cached multimodal data and / or contextual memory data are transferred to the large model, enabling the large model to reason and respond to the received interactive instructions based on the multimodal data and / or contextual memory data.
[0087] In this implementation, during the process of storing long-term memory data in the cache, the system will simultaneously perform targeted caching processing of multimodal data and contextual memory data.
[0088] Specifically, it can filter multimodal data within the most recent preset time period (such as raw data such as in-vehicle voice, images, and vehicle status collected within the last hour) and a preset number of contextual memory data (such as 10 recently generated contextual memory fragments with timestamps) and store them together in the cache.
[0089] In this way, the timeliness requirements of data in the vehicle scenario are taken into account (data from the most recent period is often more closely related to the current user interaction, such as a recently ended conversation or recent operating habits), and the preset duration and preset quantity limits are used to avoid excessive redundant data accumulation in the cache, ensuring that cache resources can efficiently serve high-frequency call requirements.
[0090] For example, if a user adjusted the seat angle just 5 minutes ago (multimodal data), and the system has generated corresponding contextual memory data ("User adjusted the driver's seat back to 110° at 14:30"), these two types of data will be stored in the cache first, providing the original basis and contextual support for quick retrieval of possible related interactions (such as "Help me restore the seat position just now").
[0091] At the same time, this method also supports the direct calling of multimodal data and contextual memory data by large models. When the system receives the calling instructions sent by the large model for these two types of data, it will complete the data transmission through the preset standardized interface.
[0092] For example, when a user gives the interactive instruction "Where did I just say I wanted to go?", the big model will trigger the call to recent contextual memory data. The system transmits the "10 most recent contextual memory data" in the cache to the big model through an interface. The big model can quickly locate the contextual segment containing "the user mentioned the destination" from it, and then accurately infer the user's needs and respond accordingly. If the big model needs to verify the details of a certain interaction (such as "Did the user say they wanted to turn on the air conditioner?"), it can call the multimodal data (such as the original voice recording) of the corresponding time period through an interface to ensure the accuracy of the inference basis.
[0093] In this way, large models can not only provide personalized services by relying on long-term memory data, but also restore specific interaction scenarios and verify detailed information by calling recent multimodal data and contextual memory data, thereby further improving the accuracy of responses and scenario fit.
[0094] like Figure 2 The diagram shown is a schematic representation of the memory hierarchical architecture and data flow relationship of this application.
[0095] The left side shows a three-layer memory architecture, including the MemLay0 raw data layer (raw data cache), which stores multimodal data in the first-level cache; the MemLay1 contextual memory layer (time-series memory cards), which stores contextual memory data; and the MemLay2 feature memory layer (feature memory / preference memory), which stores long-term memory data. The right side shows the shared state (Statement) and the edge cache (second-level cache).
[0096] MemLay0 transmits data to the edge cache in time series; MemLay1 extracts contextual memory data to the edge cache; MemLay2, based on contextual memory mining features, sends long-term memory data to the edge cache. Simultaneously, the edge cache can provide feedback to MemLay1 and MemLay2 to update contextual memories and mining features.
[0097] Figure 3 This is a flowchart illustrating a memory processing system in a vehicle-mounted scenario, based on an exemplary embodiment. The messages, environment data (including vehicle data, camera data, and navi data), and human profile data (covering biometrics, clothing features, and preference features) on the left side are first aggregated onto a memory bus, synchronizing the raw memory data to the "Memory Retrieval and Caching" module. In this module, a first-level cache stores data cached from all raw data for a specific time period and is connected to a log system; time-slice scene memory sequences are retrieved from long-term memory, storing scene memory cards in a vector database; and person feature preferences are extracted from person features, storing person feature cards in MongoDB. The processed data, via the memory bus, flows again to a second-level cache in the form of "updating long-term memory / recalling memory." The second-level cache contains the final scene memory summary, the most recent time-slice dialogue cache, the cache of current in-vehicle personnel features, and recalled data. Ultimately, this data is used for statement structure encapsulation to support memory applications for intelligent interaction in vehicle-mounted scenarios.
[0098] Figure 4 This is a schematic diagram illustrating a system architecture and data interaction process according to an exemplary embodiment.
[0099] The left side represents the AI Core Runtime, which includes the System Agent, Multimodal Question Answering Agent, and Intelligent Scene Agent. It interacts with Action-Service, PerceptionService, MemoryService, ModelCoreService, and LongTermMemory Provider through the AI kit's perception, execution, large model, and memory interfaces. PerceptionService covers sensor perception (human / vehicle / environmental state perception), sound perception (ASR, ambient sound perception), and visual perception (video frame perception and distribution, target feature filtering). MemoryService includes the Memory Manager (data acquisition and synchronization, latent memory graph understanding, environmental and human feature recognition, memory retrieval data caching) and the memory data bus (pub / sub). On the right is the AI Box / Cloud, which manages memory and caches data. Memory management is handled by the long-term memory agent, involving data collection, data retrieval, extraction and storage of contextual memory information, management and analysis of user preference tags, and analysis and extraction of user experience knowledge data. It is connected to the end-side multimodal analysis model, data bus protocol (MQTT), and embedding model. The data cache includes short-term memory cache, person / vehicle / environment status cache, biometric feature cache, operation log cache, long-term contextual memory cards, preference tags and preference data, experience knowledge base, expert knowledge base, and user basic information. It connects to the log system, SQLite database, MongoDB, and vector database. The left and right sides are connected by arrows to realize data interaction, building a complete AI system link from data perception, processing to memory management and application.
[0100] As can be seen from the above, the technical solution provided by the embodiments of this application avoids the risk of privacy leakage caused by cloud transmission by processing multimodal data locally in real time on the edge, and improves the timeliness of memory generation. At the same time, the construction of long-term memory based on scenario and feature analysis enhances the vehicle's ability to understand user preferences, scenarios and interaction history, which helps to achieve a more coherent, intelligent and personalized interactive experience, and effectively alleviates the problems of privacy and timeliness contradiction, lack of active memory ability and cross-Agent collaboration obstacles caused by cloud deployment in the prior art.
[0101] Figure 5 This is a block diagram of an active memory device according to an exemplary embodiment, comprising:
[0102] The acquisition module 201 is used to acquire multimodal data in real time, and the multimodal data is associated with the in-vehicle user and the vehicle's current situation;
[0103] Context analysis module 202 is used to perform semantic analysis on the multimodal data to obtain context memory data, wherein the context memory data carries a corresponding timestamp;
[0104] The feature analysis module 203 is used to perform feature analysis on the context memory data to obtain long-term memory data, so that the vehicle can reason and respond to the received interactive instructions based on the long-term memory data.
[0105] Optionally, the scenario analysis module includes:
[0106] The grouping unit is used to divide the multimodal data into time slices according to a preset time interval to obtain multiple groups. The multimodal data of each group includes: dialogue information, location information, human characteristic information, human and vehicle environmental status information, and current application operation information of the people in the vehicle.
[0107] The fusion unit is used to perform data fusion on the multimodal data of each group to obtain the contextual memory data and corresponding timestamp of each group, wherein the contextual memory data includes: associated personal characteristics and contextual memory information.
[0108] Optionally, the feature analysis module includes:
[0109] The feature data acquisition unit is used to perform multi-dimensional and multi-target feature analysis on the context memory data to obtain the feature data corresponding to each context memory data.
[0110] The memory data acquisition unit is used to perform feature fusion on the feature data to obtain long-term memory data.
[0111] Optionally, the device further includes:
[0112] A memory data storage module is used to store the long-term memory data in a database and update the secondary cache based on the long-term memory data;
[0113] The data transmission module is used to receive a call instruction from the large model for the long-term memory data, and then transmit the long-term memory data to the large model so that the large model can perform reasoning and response to the received interactive instructions based on the long-term memory data.
[0114] Optionally, the device further includes:
[0115] The state update module is used to update the shared state in the cache based on the long-term memory data;
[0116] The data transmission module includes:
[0117] A state transmission unit is used to transmit the shared state to the large model.
[0118] Optionally, the data storage module includes:
[0119] A memory data storage unit is used to store the most recent multimodal data within a preset time period and a preset number of scenario memory data into the cache;
[0120] The context data transmission unit is used to transmit the cached multimodal data and / or context memory data to the large model after receiving a call instruction from the large model for the multimodal data and / or the context memory data, so that the large model can perform reasoning and response to the received interaction instruction based on the multimodal data and / or the context memory data.
[0121] As can be seen from the above, the technical solution provided by the embodiments of this application avoids the risk of privacy leakage caused by cloud transmission by processing multimodal data locally in real time on the edge, and improves the timeliness of memory generation. At the same time, the construction of long-term memory based on scenario and feature analysis enhances the vehicle's ability to understand user preferences, scenarios and interaction history, which helps to achieve a more coherent, intelligent and personalized interactive experience, and effectively alleviates the problems of privacy and timeliness contradiction, lack of active memory ability and cross-Agent collaboration obstacles caused by cloud deployment in the prior art.
[0122] Figure 6 This is a block diagram illustrating an electronic device for active memory according to an exemplary embodiment.
[0123] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions that can be executed by a processor of an electronic device to perform the method. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0124] In an exemplary embodiment, a computer program product is also provided that, when run on a computer, enables the computer to implement the method of active memory.
[0125] As can be seen from the above, the technical solution provided by the embodiments of this application avoids the risk of privacy leakage caused by cloud transmission by processing multimodal data locally in real time on the edge, and improves the timeliness of memory generation. At the same time, the construction of long-term memory based on scenario and feature analysis enhances the vehicle's ability to understand user preferences, scenarios and interaction history, which helps to achieve a more coherent, intelligent and personalized interactive experience, and effectively alleviates the problems of privacy and timeliness contradiction, lack of active memory ability and cross-Agent collaboration obstacles caused by cloud deployment in the prior art.
[0126] Figure 7 This is a block diagram illustrating an active memory device 800 according to an exemplary embodiment.
[0127] For example, device 800 can be a mobile phone, computer, digital broadcasting electronic device, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0128] Reference Figure 7 The device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0129] Processing component 802 typically controls the overall operation of device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps described. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0130] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of this data include instructions for any application or method operating on device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0131] Power supply component 807 provides power to various components of device 800. Power supply component 807 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 800.
[0132] Multimedia component 808 includes a screen that provides an output interface between the device 800 and the account. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the account. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0133] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0134] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, which may be a keyboard, click wheel, buttons, etc. These buttons may include, but are not limited to, a home button, volume buttons, a power button, and a lock button.
[0135] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of device 800, changes in position of device 800 or a component of device 800, the presence or absence of contact between an account and device 800, orientation or acceleration / deceleration of device 800, and temperature changes of device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0136] Communication component 816 is configured to facilitate wired or wireless communication between device 800 and other devices. Device 800 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0137] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described in the first and second aspects.
[0138] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions that can be executed by a processor 820 of the device 800 to perform the method. Optionally, for example, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.
[0139] In an exemplary embodiment, a computer program product including instructions is also provided, which, when run on a computer, causes the computer to perform any of the active memory methods described in the embodiments.
[0140] As can be seen from the above, the technical solution provided by the embodiments of this application avoids the risk of privacy leakage caused by cloud transmission by processing multimodal data locally in real time on the edge, and improves the timeliness of memory generation. At the same time, the construction of long-term memory based on scenario and feature analysis enhances the vehicle's ability to understand user preferences, scenarios and interaction history, which helps to achieve a more coherent, intelligent and personalized interactive experience, and effectively alleviates the problems of privacy and timeliness contradiction, lack of active memory ability and cross-Agent collaboration obstacles caused by cloud deployment in the prior art.
[0141] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0142] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. An active memorization method, characterized in that, include: Real-time acquisition of multimodal data, which is associated with the in-vehicle user and the vehicle's current situation; Semantic analysis is performed on the multimodal data to obtain contextual memory data, which carries a corresponding timestamp. Feature analysis is performed on the contextual memory data to obtain long-term memory data, enabling the vehicle to reason and respond to received interactive commands based on the long-term memory data.
2. The active memory method according to claim 1, characterized in that, The semantic analysis of the multimodal data to obtain contextual memory data includes: The multimodal data is divided into time slices according to a preset time interval to obtain multiple groups. The multimodal data of each group includes: the dialogue information, location information, and human characteristic information of the people in the vehicle, as well as the environmental status information of the people and vehicles and the current application operation information. Data fusion is performed on the multimodal data of each group to obtain the contextual memory data and corresponding timestamps for each group, wherein the contextual memory data includes: associated personal characteristics and contextual memory information.
3. The active memory method according to claim 1, characterized in that, The step of performing feature analysis on the episodic memory data to obtain long-term memory data includes: Multi-dimensional and multi-target feature analysis was performed on the aforementioned contextual memory data to obtain the feature data corresponding to each contextual memory data. Feature fusion is performed on the feature data to obtain long-term memory data.
4. The active memory method according to any one of claims 1-3, characterized in that, After performing feature analysis on the episodic memory data to obtain long-term memory data, the method further includes: The long-term memory data is stored in the database, and the secondary cache is updated based on the long-term memory data; Upon receiving a call instruction from the large model for the long-term memory data, the long-term memory data is transmitted to the large model, enabling the large model to perform reasoning and response to the received interactive instructions based on the long-term memory data.
5. The active memory method according to claim 4, characterized in that, After storing the long-term memory data in the cache, the method further includes: Update the shared state in the cache based on the long-term memory data; The step of transferring the long-term memory data to the large model includes: The shared state is transmitted to the large model.
6. The active memory method according to claim 4, characterized in that, The multimodal data within the most recent preset time period and the preset number of scenario memory data are stored in the cache; Upon receiving a call instruction from the large model for the multimodal data and / or the contextual memory data, the cached multimodal data and / or the contextual memory data are transferred to the large model, so that the large model can perform reasoning and response to the received interactive instructions based on the multimodal data and / or the contextual memory data.
7. An active memory device, characterized in that, include: The acquisition module is used to acquire multimodal data in real time, and the multimodal data is associated with the in-vehicle user and the vehicle's current situation; The scenario analysis module is used to perform semantic analysis on the multimodal data to obtain scenario memory data, which carries a corresponding timestamp. The feature analysis module is used to perform feature analysis on the contextual memory data to obtain long-term memory data, so that the vehicle can reason and respond to received interactive instructions based on the long-term memory data.
8. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the active memory method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the active memory electronic device, the active memory electronic device is enabled to perform the active memory method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the active memory method according to any one of claims 1 to 6.