Data storage method and electronic device

CN122777066APending Publication Date: 2026-09-18LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611164002.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-31
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

目前难以在长期记忆存储中兼顾存储开销的问题

Benefits of technology

[0012] According to a second aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the following operations: determining reference event information for target event information, the reference event information being determined based on at least one type of associated information related to the target event information; obtaining first difference information between the target event information and the reference event information; determining a storage strategy corresponding to the target event information based on the first difference information; and, according to the storage strategy, associating and storing second difference information between the target event information and the reference event information with an event identifier corresponding to the target event information, the second difference information being the same as or associated with the first difference information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122777066A_ABST
    Figure CN122777066A_ABST
Patent Text Reader

Abstract

The present disclosure provides a data storage method and an electronic device, and relates to the technical field of artificial intelligence. The data storage method comprises the following steps: determining reference event information of target event information, the reference event information being determined based on at least one kind of association information related to the target event information; obtaining first difference information between the target event information and the reference event information; determining a storage strategy corresponding to the target event information based on the first difference information; and storing second difference information between the target event information and the reference event information and an event identifier corresponding to the target event information according to the storage strategy, the second difference information being the same as or associated with the first difference information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more specifically to a data storage method and electronic device. Background Technology

[0002] In long-term, continuous operation scenarios for multimodal agents, the agent can continuously receive multimodal data such as text, images, video, audio, and interaction logs. The agent can summarize, vectorize, or graph-store the important content of the multimodal data for later use. Currently, it is difficult to balance storage overhead with long-term memory storage. Summary of the Invention

[0003] According to a first aspect of this disclosure, a data storage method is provided, comprising: determining reference event information for target event information, the reference event information being determined based on at least one type of associated information related to the target event information; obtaining first difference information between the target event information and the reference event information; determining a storage strategy corresponding to the target event information based on the first difference information; and, according to the storage strategy, associating and storing second difference information between the target event information and the reference event information with an event identifier corresponding to the target event information, wherein the second difference information is the same as or associated with the first difference information.

[0004] According to embodiments of this disclosure, at least one type of associated information includes at least one of general knowledge, domain knowledge, user information, current task context, and historical event information; the event identifier corresponding to the target event information has a mapping relationship with at least one type of associated information; the event identifier is generated based on at least one type of associated information, or generated based on the target event information.

[0005] According to embodiments of this disclosure, the reference event information includes predicted event information, which is generated by predicting a target event based on at least one related piece of information; the first difference information includes surprise information of the target event information relative to the predicted event information, which characterizes the expected deviation between the target event information and the predicted event information.

[0006] According to embodiments of this disclosure, obtaining first difference information between target event information and reference event information includes: obtaining one of semantic error, visual error, and behavioral error of the user involved in the target event information between the target event information and the predicted event information; determining surprise level information based on one of the semantic error, visual error, and behavioral error of the user involved in the target event information; or, obtaining at least two of the semantic error, visual error, and behavioral error of the user involved in the target event information between the target event information and the predicted event information; assigning weight information to at least two of the semantic error, visual error, and behavioral error according to the current task context related to the target event information; performing a weighted operation on at least two of the semantic error, visual error, and behavioral error based on the weight information to obtain a weighted result; and determining surprise level information based on the weighted result.

[0007] According to embodiments of this disclosure, determining a storage strategy corresponding to target event information based on first difference information includes: in response to the surprise level information not meeting the target requirements, determining the storage strategy corresponding to the target event information as a first storage strategy, the first storage strategy indicating the storage of the essential representation contained in the second difference information, the second difference information being associated with the first difference information; or, in response to the surprise level information meeting the target requirements, determining the storage strategy corresponding to the target event information as a second storage strategy, the second storage strategy indicating the storage of the full amount of information contained in the second difference information, the second difference information being associated with the first difference information.

[0008] According to embodiments of this disclosure, in accordance with a storage strategy, the second difference information between the target event information and the reference event information is associated and stored with the event identifier corresponding to the target event information. This includes: when the target event information has associated historical event information, determining the historical event identifier of the historical event information as the event identifier corresponding to the target event information; in accordance with a first storage strategy, associating and storing the essential representation contained in the second difference information with the historical event identifier; or, when the target event information does not have associated historical event information, generating a new event identifier based on at least one associated information or the target event information, and using it as the event identifier corresponding to the target event information; in accordance with the first storage strategy, associating and storing the essential representation contained in the second difference information with the historical event identifier. The storage method involves: 1) storing the new event identifier in association with the target event information; 2) storing the new event identifier in association with the target event information, provided that the target event information has associated historical event information; 3) storing the new event identifier in association with the historical event identifier of the historical event information, provided that the new event identifier is generated based on at least one associated information or the target event information; 4) storing the new event identifier in association with the historical event identifier of the historical event information, provided that the new event identifier is generated based on at least one associated information or the target event information; 5) storing the new event identifier in association with the target event information, provided that the target event information does not have associated historical event information; 6) storing the new event identifier in association with the target event information, provided that the target event information does not have associated historical event information; 7) storing the new event identifier in association with the target event information, provided that the target event information does not have associated historical event information; 8) storing the new event identifier in association with the target event information, provided that the target event information does not have associated historical event information; and 9) storing the new event identifier in association with the target event information, provided that the target event information does not have associated historical event information, provided that the target event information has associated historical event information; and 10) storing the new event identifier in association with the target event information, provided that the target event information has associated historical event information.

[0009] According to embodiments of this disclosure, the reference event information includes at least one of event summary information, aggregation information, statistical information, and common representation information, and the reference event information is generated based on intent understanding of at least one related information; the first difference information includes at least one of new information, abnormal information, and uncovered information of the target event information relative to the reference event summary information.

[0010] According to embodiments of this disclosure, determining a storage strategy corresponding to target event information based on first difference information includes: in response to the first difference information containing less than a data volume threshold, determining the storage strategy corresponding to the target event information as a first storage strategy, wherein the first storage strategy indicates storing the essential representation contained in the second difference information; or, in response to the first difference information containing more than or equal to a data volume threshold, determining the storage strategy corresponding to the target event information as a second storage strategy, wherein the second storage strategy indicates storing the full amount of information contained in the second difference information.

[0011] According to embodiments of this disclosure, in accordance with a storage strategy, associating and storing the second difference information between the target event information and the reference event information with the event identifier corresponding to the target event information includes: determining the historical event identifier of the historical event information as the event identifier corresponding to the target event information; associating and storing the essential representation contained in the second difference information with the historical event identifier in accordance with a first storage strategy; or, generating a new event identifier based on at least one associated information or the target event information as the event identifier corresponding to the target event information, and associating the new event identifier with the historical event identifier of the historical event information; and associating and storing the full information contained in the second difference information with the new event identifier in accordance with a second storage strategy.

[0012] According to a second aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the following operations: determining reference event information for target event information, the reference event information being determined based on at least one type of associated information related to the target event information; obtaining first difference information between the target event information and the reference event information; determining a storage strategy corresponding to the target event information based on the first difference information; and, according to the storage strategy, associating and storing second difference information between the target event information and the reference event information with an event identifier corresponding to the target event information, the second difference information being the same as or associated with the first difference information. Attached Figure Description

[0013] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0014] Figure 1 This diagram illustrates an application scenario of a data storage method according to an embodiment of the present disclosure.

[0015] Figure 2 A flowchart illustrating a data storage method according to an embodiment of the present disclosure is shown schematically.

[0016] Figure 3 The illustration schematically shows a first difference information determination process according to an embodiment of the present disclosure;

[0017] Figure 4 A schematic diagram illustrating the storage strategy determination process according to an embodiment of the present disclosure is shown.

[0018] Figure 5 This illustration schematically shows a storage path diagram of the second difference information according to an embodiment of the present disclosure;

[0019] Figure 6A schematic block diagram of a data storage device according to an embodiment of the present disclosure is shown.

[0020] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a data storage method according to an embodiment of the present disclosure. Detailed Implementation

[0021] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0022] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0024] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0025] In the technical solution disclosed herein, the user information (including but not limited to user personal information, user image information, user device information, such as location information) and data (including but not limited to data used for analysis, stored data, and displayed data) involved are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation entry points are provided for users to choose to authorize or refuse.

[0026] In related technologies, agents can recursively compress and store received data. This method primarily focuses on context expansion optimization for long sequences of plain text, aiming to maintain model performance within limited storage space. However, this method does not address long-term experience storage of multimodal data (such as images, videos, and audio) and lacks a dynamic compression mechanism based on user-personalized prediction errors.

[0027] The agent can also dynamically extract and retrieve long-term dialogue memory, reducing latency and cost. This method dynamically extracts salient information from the dialogue, stores it in a structured manner, and supports efficient retrieval. The compression logic, based on importance filtering, stores the extracted complete semantic fragments.

[0028] The intelligent agent can also achieve effective understanding and backtracking of long video content by structurally decomposing the video content and establishing a memory index based on timelines and key objects. This method primarily serves a general understanding of video content and does not establish differentiated predictive compression for different users.

[0029] Intelligent agents can also typically combine techniques such as summary extraction, knowledge graph construction, model distillation, and vector retrieval to build complex multimodal memory systems capable of processing multimodal data such as text and images. The compression core of this method remains traditional information condensation techniques such as summarization, graphing, distillation, and retrieval.

[0030] The above methods all lack a memory compression mechanism that separates routine (predictable) information explained by existing knowledge from novel but unpredictable information, resulting in excessive retention of redundant routine information, high long-term costs, and weak expression of individual differences.

[0031] In view of this, embodiments of the present disclosure provide a data storage method to at least partially solve the above-mentioned technical problems. A detailed description is provided below with reference to specific embodiments.

[0032] Figure 1 The diagram illustrates an application scenario of the data storage method according to an embodiment of the present disclosure.

[0033] like Figure 1 As shown, the application scenario 100 according to the embodiments of this application may include a terminal device 110, a network 120, and a server 130.

[0034] Network 120 serves as a medium for providing a communication link between terminal device 110 and server 130. Network 120 may include various connection types, such as wired, wireless communication links, or fiber optic cables. For example, user 140 can use terminal device 110 to interact with server 130 through network 120 to receive or send information, etc.

[0035] In some embodiments, application scenario 100 includes a terminal device 110, a network 120, and a server 130, which work together to form a distributed memory storage and inference architecture for the long-term operation of a multimodal intelligent agent. The terminal device 110 serves as the local interaction entry point for the user 140, continuously collecting multi-source heterogeneous data streams such as text, images, video, audio, and interaction logs. The server 130 carries the core generative model, prediction engine, event repository, and splitting and compression module, responsible for performing surprise level determination, reference event matching, difference extraction, and hierarchical storage scheduling on the multimodal data uploaded by the terminal. The network 120 serves as the communication link medium between the two, supporting various connection types such as wired, wireless communication links, or fiber optic cables, ensuring low-latency transmission of high-throughput video streams, real-time audio frames, and interaction logs.

[0036] Terminal, fixed terminal or portable terminal, including mobile phones, tablet computers, laptop computers, notebook computers, netbook computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), smart speakers, smart home devices, in-vehicle devices, wearable devices or any combination thereof, including accessories and peripherals of these devices or any combination thereof.

[0037] In some embodiments, terminal device 110 uses hardware modules such as cameras, microphones, touchscreens, and sensors to capture user 140's voice commands, photographs, videos, text inputs, and operation interaction logs in real time. Terminal device 110 can also perform lightweight modal alignment, timestamp annotation, and basic semantic tagging before data upload, reducing the transmission load on network 120.

[0038] Server 130 can be a server providing various services, such as a backend server supporting requests input by a user using terminal device 110. The backend server can retrieve reference event information matching the target event information from the event repository based on association information, using this as a benchmark anchor for difference comparison. It calculates multiple difference information between the target event information and the reference event information, generates or reuses event identifiers based on association information, matches the corresponding storage strategy according to the difference information, and performs differentiated storage of the target event according to the storage strategy.

[0039] User 140 initiates a multimodal interaction request through terminal device 110. After server 130 completes prediction, compression, and storage, it sends the event identifier, summary, or high-fidelity residual fragment back to terminal device 110. User 140 can view the event backtracking results in real time.

[0040] It should be noted that the data storage method provided in this application embodiment can generally be executed by server 130 and / or terminal device 110. Accordingly, the data storage device provided in this application embodiment can generally be located in server 130 and / or terminal device 110.

[0041] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0042] Figure 2 A flowchart illustrating a data storage method according to an embodiment of the present disclosure is shown schematically.

[0043] like Figure 2 As shown, the data storage method of this disclosure embodiment may include S210~S240.

[0044] In S210, the reference event information for the target event information is determined.

[0045] According to embodiments of this disclosure, event information can refer to event fragment information with complete semantic boundaries, which is a structured independent memory unit. Event information can be bound to full-dimensional association features, supporting unified indexing and backtracking of content from different modalities. Full-dimensional association features can include basic spatiotemporal elements, original features, event semantic elements, and event identifiers. Basic spatiotemporal elements can clearly mark the three basic dimensions of the event's subject, time, and location, corresponding to "who, when, and where" in the independent memory unit. Original features can cover all modalities of original input, such as collected text, images, videos, audio, interaction logs, and sensor data, corresponding to "what was seen" and "what was heard" in the independent memory unit. Event semantic elements can completely record the core behavior, causal relationships, object state changes, task-related goals, and other semantic logic of the event, corresponding to "what happened" in the independent memory unit. Event identifiers can serve as identity identifiers for event information.

[0046] Event information can be independent memory units formed after time alignment and event segmentation. For example, the multimodal data stream can be divided into multiple segments based on the semantic boundaries in the multimodal data stream; for each segment, the multimodal data contained in the segment is time-aligned to obtain the time-aligned segment; and event identifiers are assigned to the aligned segment to obtain the event information.

[0047] According to embodiments of this disclosure, the target event information can refer to the event information corresponding to the latest multimodal input currently being processed, that is, the object that needs to be compressed and stored in this round, and can cover any combination of input modes such as text, image, video, audio, and interactive logs.

[0048] According to embodiments of this disclosure, reference event information can be determined based on at least one piece of associated information related to target event information, and used as a reference event information for comparison with the target event information. The associated information can be event information or non-event information, serving as auxiliary information for determining the reference event. The associated information is the retrieval basis for determining whether an event is predictable and whether it matches the reference event.

[0049] In S220, the first difference information between the target event information and the reference event information is obtained.

[0050] According to embodiments of this disclosure, the first difference information may be an abstract description of the difference features extracted after comparing the target event information and the reference event information. It is not the original data itself, but rather an extraction of the differences between the target event information and the reference event information. The first difference information can be the core decision-making basis for subsequent storage strategies.

[0051] In S230, based on the first difference information, the storage strategy corresponding to the target event information is determined.

[0052] According to embodiments of this disclosure, the storage strategy can be a differentiated storage rule determined based on first difference information. The storage strategy may include at least one of the following: storage content, storage precision, storage location, storage level, storage format, update object, retention duration, and retrieval priority.

[0053] Storage content can refer to the specific range of information that needs to be stored as specified by the storage strategy. Storage precision can refer to the fidelity and compression level of the specific information range. Storage location can refer to the physical storage medium for storing the specific information range, such as disks, distributed storage, local cache, etc. Storage hierarchy can refer to the level within the hierarchical memory system to which the specific information range belongs. Storage format can refer to the structured form in which the specific information range is ultimately encapsulated. Updated object can refer to the target unit that will be written to or modified during storage operations. Retention duration can refer to the lifecycle and automatic retention period of the specific information range. Retrieval priority can refer to the weight of different contents being hit and returned during subsequent memory retrieval operations.

[0054] In other words, different storage strategies can be used to store the differences between the target event information and the reference event information.

[0055] In S240, according to the storage strategy, the second difference information between the target event information and the reference event information is associated with and stored with the event identifier corresponding to the target event information.

[0056] According to embodiments of this disclosure, the event identifier can serve as a unique identifier for the target event information and as an anchor point for association and binding. The second difference information is stored under the event identifier of the target event information, forming an index relationship of "event identifier → difference content". In subsequent queries, the corresponding difference content can be queried through the event identifier.

[0057] The second difference information is the same as or related to the first difference information. The first difference information can be an abstract difference feature, or it can refer to the entity data stored on the blockchain. The second difference information can refer to the entity data stored on the blockchain. That is, if the second difference information is the same as the first difference information, it means that the first difference information itself is already lightweight, structured content that can be directly stored. It doesn't require additional transformation or expansion and can be directly bound to the event identifier as the second difference information for storage. If the second difference information is related to the first difference information, it means that the first difference information is only an abstract difference feature identifier / description, which needs to be further matched and associated with the corresponding original entity data to generate the corresponding concrete storage content, which is the final stored second difference information. There is a clear pointing relationship between the two.

[0058] For example, taking an intelligent vehicle terminal as an example, the multimodal input of the intelligent vehicle terminal includes driver voice interaction, in-vehicle camera images, and driving trajectory logs. The intelligent vehicle terminal collects these three types of data in real time, forming a continuous multimodal data stream. The system identifies "semantic boundary points" in the data (such as sudden changes in behavior patterns, scene transitions, etc.) through semantic detection, and segments the continuous data stream into multiple segments.

[0059] The process of segmenting target event information: Assuming a driver is commuting during the morning rush hour, from 07:30 to 08:10, the intelligent vehicle terminal system identifies semantic boundary judgment criteria that may include "vehicle starts, leaves the garage," "stable driving on the regular route A," "detection of temporary traffic control, voice prompt 'road closure ahead, trajectory deviates from the regular route'," and "stable driving resumes after detour." The corresponding segmented semantics are departure commuting, regular cruising, sudden traffic control, detour resumption, etc. The segmentation criteria include the appearance of keywords such as "road closure," trajectory deviation from the target route, and camera capturing roadblocks. These multimodal signals collectively constitute a semantic abrupt change boundary, based on which the system segments the data stream into an independent segment. The multimodal data within the segment is time-aligned to obtain the aligned event segment. For example, using 07:55:05 as a unified base timestamp, the three types of modal data are calibrated to the same time axis to form a time-aligned event segment.

[0060] The obtained reference event information can be: "vehicle starts, leaves the garage", "stable driving on the regular route A", or "trajectory deviates from the regular route".

[0061] By comparing the target event information and the reference event information, it was found that "temporary traffic control" was a difference that was never recorded in the reference itinerary. The "temporary traffic control" event fragment can be bound to the event identifier of the target event for associated storage, and other event fragments can be aligned and no longer stored.

[0062] The data storage method of this disclosure obtains reference event information associated with the target event information as a benchmark for comparison with the target event information to obtain difference information. The storage strategy is determined based on this difference information; different difference information leads to different storage strategies. Therefore, data can be stored in a distributed manner based on the difference information, thereby achieving a memory compression mechanism based on "predictable / unpredictable" distribution. This avoids the repeated storage of content in the target event information that has already been covered by the reference event information or is interpretable, allowing the storage object to focus more on the difference content, thus reducing the accumulation of redundant data and lowering storage overhead.

[0063] Furthermore, storing data streams based on event information essentially upgrades "data storage" to "event memory with semantics, valuable judgments, and structured indexes," enabling the storage itself to have retrieval, tracing, and decision-making capabilities.

[0064] In some embodiments, at least one type of associated information includes at least one of general knowledge, domain knowledge, user information, current task context, and historical event information.

[0065] According to embodiments of this disclosure, general knowledge can be pre-built, global common-sense knowledge that does not depend on specific users or specific domains. Examples include the laws governing the physical world (e.g., "roads are slippery when it rains"), social common sense (e.g., "traffic lights indicate go / stop"), universal semantics of language (e.g., "congestion means slower traffic"), and basic logical reasoning rules. General knowledge can be used to construct the underlying expectation framework for reference event information. For example, if the system knows that "traffic jams are common during weekday morning rush hour," it can generate a reference event benchmark of "regular morning rush hour commuting" and compare it with the target event.

[0066] According to embodiments of this disclosure, domain knowledge can be specialized rules, models, and prior knowledge specific to a particular application domain. Examples include traffic regulations and driving behavior models in the field of autonomous driving, diagnostic guidelines and disease knowledge bases in the medical field, transaction rules and risk control models in the financial field, and equipment operating parameters and fault mode libraries in the industrial manufacturing field. Domain knowledge can provide accurate predictive benchmarks for specific scenarios, making reference event information more closely aligned with the "normal state" of the specialized scenario, and enabling more targeted difference comparisons.

[0067] According to embodiments of this disclosure, user information can refer to personalized data directly related to a specific individual user. Examples include historical behavioral habits (frequently used routes, frequently used functions, daily routines), personal interest tags, device usage preferences, and interaction patterns (such as a user's preference for voice commands over text input). User information allows reference event information and storage strategies to be perfectly matched to the user's personal pattern, resulting in different judgments and storage results for the same event for different users.

[0068] According to embodiments of this disclosure, the current task context can be the task objective, task stage, and task constraints that the system is currently executing. For example, the current task type (e.g., navigation, meeting recording, security monitoring), task objective (e.g., "safely arrive at the destination," "complete meeting minutes"), task stage (start / in progress / ending), task priority, and constraints (e.g., "arrive within 30 minutes," "all meeting decisions must be recorded"), etc.

[0069] According to embodiments of this disclosure, historical event information can refer to stored historical memory data related to the target event information. Examples include records of similar events that have occurred in the past for the user / scenario, statistical summaries of events (e.g., "This road segment has been subject to 5 traffic controls in the past 30 days"), aggregated commonalities of events (e.g., "This user had 3 detour records last week"), stored reference events, and difference information. Historical event information can provide a direct historical baseline for generating reference event information.

[0070] In some embodiments, the event identifier corresponding to the target event information has a mapping relationship with at least one type of associated information. The event identifier is generated based on at least one type of associated information, or based on the target event information.

[0071] According to embodiments of this disclosure, the event identifier is not an isolated, randomly generated number, but can have a clear correspondence with associated information. The event identifier can be found by reverse deduction through the associated information, and the associated information can also be traced back through the event identifier.

[0072] By using associated information (such as "Zhang San + Today + Commuting"), the corresponding event identifier can be directly located without traversing the entire database, significantly improving retrieval efficiency. Event identifiers serve as anchor points, unifying secondary differentiation information from different modalities such as text, images, audio, and sensor data under a single identifier, enabling integrated cross-modal backtracking. Event identifiers can map to related information such as user information; similar events from different users can have different identifiers without interference, supporting independent memory management in multi-user scenarios. When associated information changes (e.g., a user updates their frequently used routes), the mapping relationship can dynamically adjust the event identifier generation rules, ensuring the identifier system remains consistent with the latest associated information.

[0073] According to embodiments of this disclosure, generating an event identifier based on at least one piece of associated information can refer to: extracting key elements (user ID, time, location, task type, etc.) from the associated information and combining them into an encoding to generate an event identifier. Generating an event identifier based on target event information can refer to extracting and generating an event identifier from the characteristic elements (semantic content, modality type, spatiotemporal range, etc.) of the target event information itself.

[0074] Through the data storage method of this disclosure, since the associated information is a multi-dimensional clue for "understanding the context, determining the reference benchmark, and judging whether it is important", and the event identifier is a unified anchor point for "hanging up the memory content, finding it, and being traceable", the two are bidirectionally bound through the mapping relationship, jointly supporting a differentiated, personalized, and searchable efficient memory storage system, which better meets the needs of personalized off-line storage.

[0075] In some embodiments, the reference event information includes predicted event information, which is generated by predicting the target event based on at least one related piece of information.

[0076] The first difference information includes the surprise information of the target event information relative to the predicted event information, which characterizes the expected deviation between the target event information and the predicted event information.

[0077] According to embodiments of this disclosure, the predicted event information can be an event predicted based on at least one type of related information before the target event occurs. The related information can be used as input to an existing prediction model or generative model for inference. The expected event is output as the predicted event information.

[0078] For example, general knowledge can provide the basic laws governing the world's operation as the underlying logic for prediction. For instance, "roads will be slippery on rainy days" predicts "the current driving speed should be reduced." Domain knowledge can provide rule models specific to a particular domain as the basis for prediction. For instance, a traffic domain model predicts "congestion should occur on this road segment at 7:30 am on weekdays." User information can provide users' personal habits as a personalized prediction baseline. For instance, "Zhang San usually takes Route A at 7:30 am" predicts "Zhang San will be on Route A at 7:30 am today." The current task context can provide the current task objectives and constraints as prediction conditions. For instance, if the task is "emergency navigation," it predicts "priority should be given to abnormal road conditions." Historical event information can provide historical statistical patterns as prediction priors. For instance, "this road segment has been subject to traffic control 5 times in the past 30 days" predicts "there is a risk of traffic control on this road segment today."

[0079] Surprise level information quantifies the degree of deviation between target event information (what actually happened) and predicted event information (what the system expects to happen). Surprise level information can accurately quantify the target event information and predicted event information, precisely distinguish between low surprise level (predictable) and high surprise level (unpredictable) content, and match different storage strategies.

[0080] It's important to note that using predicted event information as reference event information is applicable to boundless real-time multimodal sensing streams. For example, cameras and sensors generate continuous data streams 24 / 7. These streams contain many valueless repetitive images and routine values, lacking clear event boundaries. Determining storage strategies based on surprise information allows for real-time comparison of the current frame with historical baselines using target rules or prediction models. This enables the filtering of highly surprising and valid events that "exceed expectations," directly skipping massive amounts of invalid and redundant data and controlling storage costs at the source. Furthermore, surprise information-based calculations are lightweight real-time comparison operations, quickly completing event identification without waiting for the full dataset to be stored, meeting the high latency requirements of real-time sensing streams.

[0081] Figure 3 The illustration shows a first difference information determination process according to an embodiment of the present disclosure.

[0082] like Figure 3 As shown, in some embodiments, obtaining first difference information between target event information and reference event information may include: obtaining one of semantic error, visual error, and user behavior error related to the target event information between the target event information and the predicted event information; and determining surprise information based on one of semantic error, visual error, and user behavior error related to the target event information.

[0083] According to embodiments of this disclosure, semantic error can refer to the difference between target event information and predicted event information at the language / text / logical content level. Visual error can refer to the difference between target event information and predicted event information at the image / video / perceptual level. Behavioral error can refer to the difference between the user behavior involved in the target event information and the expected user behavior in the predicted event information. One of the semantic error, visual error, and behavioral error can be normalized to obtain surprise level information.

[0084] For example, the target event information includes Zhang San encountering traffic control on Road A at 7:30. Multimodal data includes video footage, traffic sign images, Zhang San's GPS detour trajectory, and ambient audio. Reference event information includes Zhang San traveling normally on Road A at 7:30, with an estimated arrival time of 25 minutes. Semantic errors include inconsistencies in the textual description of the event; the predicted event information indicates "normal passage," while the actual event is "traffic control." Visual errors include the predicted event information indicating "sunny, open road," while the actual event is "cloudy with construction barriers." Behavioral errors include the user's actual route differing from the prediction; the predicted event information indicates "taking Road A," while the actual route is "detour via Road B."

[0085] By processing one of the aforementioned errors as surprise information, instead of directly treating the difference as surprise information, continuous quantization values ​​can be output, which can accurately distinguish between low and high surprise content and match different storage strategies.

[0086] like Figure 3 As shown, in some other embodiments, obtaining the first difference information between the target event information and the reference event information may further include: obtaining at least two of the semantic error, visual error, and user behavior error involved in the target event information between the target event information and the predicted event information; assigning weight information to at least two of the semantic error, visual error, and behavioral error based on the current task context related to the target event information; performing a weighted operation on at least two of the semantic error, visual error, and behavioral error based on the weight information to obtain a weighted result; and determining the surprise level information based on the weighted result.

[0087] According to embodiments of this disclosure, different users / tasks corresponding to different current task contexts have different importance. Task understanding can be performed on the current task context to identify user goals, and then weights can be assigned to different errors based on user goals.

[0088] For example, in Zhang San's daily commute navigation, the user's goal is to reach the destination quickly. Therefore, a semantic weight of 0.3 can be assigned to semantic errors, a visual weight of 0.2 to visual errors, and a behavioral weight of 0.5 (behavioral path is the most critical). In Zhang San's emergency rescue scenario, the user's goal is safety first. Therefore, a semantic weight of 0.4 can be assigned to semantic errors (event type is the most critical), a visual weight of 0.3 can be assigned to visual errors, and a behavioral weight of 0.3 can be assigned to behavioral errors. In a traffic management system, the user's goal is global traffic monitoring. Therefore, a semantic weight of 0.3 can be assigned to semantic errors, a visual weight of 0.5 can be assigned to visual errors (visual scene changes are the most critical), and a behavioral weight of 0.1 can be assigned to behavioral errors.

[0089] By assigning weights to the aforementioned errors based on the current task context to calculate surprise information, differentiated storage strategies can be determined for different tasks or different users, thereby achieving personalized storage.

[0090] Figure 4 A schematic flowchart illustrating the storage strategy determination process according to an embodiment of the present disclosure is shown.

[0091] like Figure 4 As shown, in some embodiments, determining the storage strategy corresponding to the target event information based on the first difference information may include: in response to the surprise information not meeting the target requirements, determining the storage strategy corresponding to the target event information as the first storage strategy.

[0092] According to embodiments of this disclosure, the first anomaly information is surprise level information, that is, the first difference information is an abstract difference feature, not entity data that needs to be stored. The second difference information to be stored is then associated with the first difference information. The second difference information may include at least one of semantic error, visual error, and behavioral error.

[0093] According to embodiments of this disclosure, a first storage strategy instructs the storage of the essential representation contained in the second difference information. The essential representation may include at least one of semantic summaries, templates or habit updates, statistical features, background models, or sparse keyframes. Semantic summaries may use natural language or structured text to condense the core content of the second difference information into a short description of one or a few sentences, discarding minor details. Templates may refer to extracting a reusable event "skeleton / framework" from the current event, which can be directly applied when encountering similar events again. Habit updates may refer to not adding new storage if the current event conforms to the user's / system's regular patterns, but instead "merging" the information into an existing habitual pattern and updating the pattern's parameters (such as frequency, time range, etc.). Statistical features may refer to numerical statistics extracted from the second difference information, such as mean, variance, frequency, distribution, count, ratio, etc. Background models may refer to continuously updating a mathematical model of a "normal state" (such as a Gaussian mixture model, neural network representation), which the target event information can be used to fine-tune. Sparse keyframes extract a very small number of representative frames at regular intervals (or only at points of change) as evidence that the "event existed." It does not store continuous video or data streams.

[0094] According to embodiments of this disclosure, the surprise level information not meeting the target requirement can mean that the value of the surprise level information is less than a threshold. The surprise level information can be a continuous value between 0 and 1. For example, if the threshold is 0.3 and the calculated surprise level information is 0.2, since 0.2 is less than 0.3, it indicates that the target event information at this time is a low surprise level event, that is, it is highly consistent with the predicted event information.

[0095] like Figure 4As shown, in some other embodiments, determining the storage strategy corresponding to the target event information based on the first difference information may further include: in response to the surprise information meeting the target requirements, determining the storage strategy corresponding to the target event information as a second storage strategy.

[0096] The second storage strategy instructs the storage of the full amount of information contained in the second difference information, which is associated with the first difference information.

[0097] It should be understood that since the commonalities between the target event information and the reference event information can be predicted based on the correlation information, these commonalities do not need to be stored; only the full information contained in the second difference information between the target event information and the reference event information should be stored. In practical applications, the commonalities can be based on the already stored correlation information. Alternatively, the entire high-surprise-level target event information can be stored.

[0098] For example, the full amount of information included in the second difference information may include at least one of the following: episodic memory, residual memory, key fragments, changes in object state, causal chains, and consequences.

[0099] Episodic memory refers to recording the second difference between the target event information and the predicted event information as a "complete story," including narrative elements such as time, place, characters, process, and result. Residual memory records the second difference between the target event information and the predicted event information. Key fragments can be the most crucial moments selected from massive amounts of target event information and saved in the highest quality (high definition, complete). Object state changes refer to "from what state to what state" each key object changes to, including the time of change and the triggering cause. Causal chains and consequence effects record "why it happened" (causal chain) and "what it led to" (effect chain).

[0100] For example, for target event information with low surprise level, the stored essential representation can be:

[0101] Semantic summary: "2024-06-12 07:32 Zhang San A Road normal traffic, sunny, 60km / h, travel time 24min".

[0102] Template / Habit Update:

[0103] Template ID: T_ZhangSan_Commute

[0104] Update: count=121, avg_time=23.9min, var=2.9min.

[0105] Statistical characteristics:

[0106] Number of normal traffic trips this week: 6; average speed: 58.2 km / h.

[0107] Background model fine-tuning:

[0108] Average speed: 58.1, with slight adjustments to visual weights.

[0109] Sparse keyframes:

[0110] No new additions will be made (07:30 frame already covers this).

[0111] For target event information with high surprise level, the full information stored can be:

[0112] Complete semantic information:

[0113] A full text description of the event type, location, time, cause, and scope of impact;

[0114] "2024-06-12 07:30 Construction barriers erected on the eastern section of Zhangsan A Road;"

[0115] "Both directions were closed, and traffic police directed detours via Heping Road, causing a 20-minute delay."

[0116] Complete visual information:

[0117] 7:30-7:45 Full-time video stream (or complete set of keyframes);

[0118] Images of construction site barriers, traffic police hand signals, and detour intersections.

[0119] Complete behavioral information:

[0120] GPS complete track: 7:30 emergency braking → 7:31 reversing → 7:32 changing lanes → 7:35 detouring;

[0121] CAN data: Complete timing sequence of throttle / brake / steering;

[0122] Speed ​​curve: 60→0→15→45→60 km / h (Complete record).

[0123] Full context information:

[0124] Weather, lighting, surrounding traffic density, and timestamps accurate to milliseconds.

[0125] The data processing method of this disclosure, when the target event information and the predicted event information are highly consistent, indicates that the information can be explained by existing knowledge. Only the essential representation of the information can be retained, converting repetitive or routine information into a regular or low-dimensional representation, minimizing storage overhead. When the target event information and the predicted event information are highly inconsistent, it indicates that an anomaly or new information has occurred. The system executes a high-fidelity retention strategy, focusing on retaining information of high value for future decision-making.

[0126] Figure 5 The diagram illustrates the storage path of the second difference information according to an embodiment of the present disclosure.

[0127] like Figure 5 As shown, in some embodiments, according to the storage strategy, associating and storing the second difference information between the target event information and the reference event information with the event identifier corresponding to the target event information may include: when there is associated historical event information with the target event information, determining the historical event identifier of the historical event information as the event identifier corresponding to the target event information; and associating and storing the essential representation contained in the second difference information with the historical event identifier according to the first storage strategy.

[0128] In other embodiments, according to the storage strategy, associating and storing the second difference information between the target event information and the reference event information with the event identifier corresponding to the target event information may further include: generating a new event identifier based on at least one association information or the target event information when there is no associated historical event information for the target event information, and using it as the event identifier corresponding to the target event information; and associating and storing the essential representation contained in the second difference information with the new event identifier according to the first storage strategy.

[0129] In other embodiments, according to the storage strategy, associating and storing the second difference information between the target event information and the reference event information with the event identifier corresponding to the target event information may further include: if the target event information has associated historical event information, generating a new event identifier based on at least one associated information or the target event information as the event identifier corresponding to the target event information, and associating the new event identifier with the historical event identifier of the historical event information; and associating and storing the full information contained in the second difference information with the new event identifier according to the second storage strategy.

[0130] In other embodiments, according to the storage strategy, the second difference information between the target event information and the reference event information is associated with and stored with the event identifier corresponding to the target event information. This may further include: if there is no associated historical event information for the target event information, generating a new event identifier based on at least one associated information or the target event information, and using it as the event identifier corresponding to the target event information; and according to the second storage strategy, associating and storing all the information contained in the second difference information with the new event identifier.

[0131] According to embodiments of this disclosure, when there is related historical event information for the target event information, the historical event identifier of the historical event information can be directly used as the event identifier of the target event information, so that similar event information does not need to be repeatedly archived, saving storage space; when there is no related historical event information for the target event information, a new event identifier can be generated based on at least one related information or the target event information, extracting elements from the target event itself and / or surrounding related information to form a brand new event identifier for archiving, ensuring that each new scenario has an independent identity, which facilitates subsequent retrieval.

[0132] In some embodiments, the reference event information may include at least one of event summary information, aggregation information, statistical information, and commonality representation information, and the reference event information is generated based on intent understanding of at least one related piece of information. The first difference information may include at least one of new information, anomaly information, and uncovered information in the target event information relative to the reference event summary information.

[0133] According to embodiments of this disclosure, reference event information can be known, historically relevant historical event information. Event summary information can be a brief description summarizing historical event information. Aggregation information can be information obtained by merging and summarizing multiple similar historical event information. Statistical information can be numbers / indicators obtained by quantifying and statistically analyzing historical event information. Commonality representation information can be a general "template" or "feature representation" that abstracts the common features of multiple historical event information.

[0134] According to embodiments of this disclosure, intent-based generation is not a simple retrieval, but rather the agent first understands what the user wants to refer to, and then generates the corresponding information.

[0135] For example, scenario: Zhang San takes route A again. The agent's intent understanding process: The user / system wants to know: What is the relationship between Zhang San taking route A this time and the previous construction event? Therefore, it retrieves:

[0136] Event summary information: "The previous construction detour";

[0137] Statistics: "There is a 35% probability of construction on Jianguo Road during the morning rush hour."

[0138] Common information: "Road construction event template"

[0139] Reference event information is generated based on the above information combination and used as a comparison benchmark.

[0140] According to embodiments of this disclosure, newly added information can refer to content that is not present in historical event information but newly appears in the target event information. Abnormal information can refer to abnormal content appearing in the target event information. Uncovered information can refer to content that is present in historical event information but not in the target event information.

[0141] It should be noted that the method of generating reference event information based on intent understanding using at least one related piece of information is more suitable for offline events with well-defined storage adaptation boundaries and clear business rules (such as travel itineraries and budget management). Offline events with well-defined boundaries and clear business rules come with clear timestamps, event identifiers, and structured attribute fields. No additional filtering or identification is required; incremental writing and snapshot binding can be directly completed according to the target rules without data omissions or logical conflicts. For example, budget events automatically validate the validity of the amount field, and travel itinerary events automatically associate the flow status of previous and subsequent nodes. Without the need for complex real-time calculations, strong data consistency can be guaranteed, adapting to the storage requirements of offline structured events.

[0142] In some embodiments, determining the storage strategy corresponding to the target event information may include: in response to the fact that the amount of data contained in the first difference information is less than a data amount threshold, determining the storage strategy corresponding to the target event information as a first storage strategy, wherein the first storage strategy indicates that the essential representation contained in the second difference information is stored.

[0143] In other embodiments, determining the storage strategy corresponding to the target event information further includes: in response to the first difference information containing a data amount greater than or equal to a data amount threshold, determining the storage strategy corresponding to the target event information as a second storage strategy, wherein the second storage strategy indicates storing the full amount of information contained in the second difference information.

[0144] According to embodiments of this disclosure, in a scenario where reference event information is generated based on intent understanding using at least one related piece of information, the first difference information can be obtained by directly comparing the data of each modality contained in the target event information with the data of each modality contained in the reference event information based on a predetermined rule dimension. In this case, the first difference information can be entity data that needs to be stored, i.e., the second difference information.

[0145] In some embodiments, the training data sources for the prediction model that obtains prediction event information include general multimodal corpora, role / domain data, user history data, and online feedback data.

[0146] For example, during the training data acquisition phase, general data, domain data, application data, and feedback data can be collected as the basic inputs for training and evolution. These data cover multiple layers of information, from general knowledge to information used for behavior.

[0147] In the pre-training phase, a base model is trained on large-scale multimodal data to learn general semantics, temporal relationships, and cross-modal prediction capabilities. This phase forms the basic prediction model.

[0148] During the role adaptation phase, the basic predictive model is fine-tuned by incorporating domain-specific or professional data to enable it to grasp the cognitive patterns of a specific role. The model learns "how a certain type of user understands and predicts the environment."

[0149] In the individualized modeling phase, historical user data is used for individual adaptation, learning user preferences, habits, and common behavioral patterns. This forms a prediction distribution and attentional tendencies tailored to each individual, resulting in the final prediction model.

[0150] In some embodiments, according to a storage strategy, the second difference information between the target event information and the reference event information is associated with and stored with the event identifier corresponding to the target event information, including: determining the historical event identifier of the historical event information as the event identifier corresponding to the target event information; and according to a first storage strategy, associating and storing the essential representation contained in the second difference information with the historical event identifier.

[0151] In other embodiments, according to the storage strategy, the second difference information between the target event information and the reference event information is associated and stored with the event identifier corresponding to the target event information. The method further includes: generating a new event identifier based on at least one associated information or the target event information as the event identifier corresponding to the target event information, and associating the new event identifier with the historical event identifier of the historical event information; and associating and storing the full information contained in the second difference information with the new event identifier according to the second storage strategy.

[0152] It should be understood that the storage strategy in this scenario is the same as the aforementioned surprise-driven storage strategy, and the resulting technical effects are the same, so it will not be elaborated here.

[0153] In some embodiments, the data storage method may further include storing second difference information according to different time scales and / or different levels of abstraction.

[0154] For example, the compression results can be uniformly written into a hierarchical memory system, which may include:

[0155] Working memory: stores short-term context;

[0156] Episodic memory: Stores key events;

[0157] Semantic memory: storing abstract rules and knowledge;

[0158] Personal Model: Stores prediction priors and parameters.

[0159] This hierarchical memory mechanism supports information management at different time scales and abstraction levels, enabling agents to respond quickly and accumulate knowledge over a long period.

[0160] In some embodiments, the data storage method may further include: continuously optimizing the memory structure through a periodic "consolidation and learning" module.

[0161] For example, the main processes include: merging and deduplication, abstraction and generalization, forgetting mechanisms, and model updates. Merging and deduplication eliminates duplicate information and improves storage efficiency. Abstraction and generalization extract rules or patterns from multiple similar events and write them into semantic memory. Forgetting mechanisms delete low-value or long-unused information. Predictive model updates feed new patterns back into the predictive model, continuously improving its predictive capabilities.

[0162] In addition, the agent can use user feedback as training signals, such as retrieval behavior, error correction operations, task success / failure, and usage frequency, to dynamically adjust the surprise threshold and compression strategy.

[0163] In some embodiments, the agent can be trained with a compression control strategy, enabling it to make storage decisions based on initial difference information and correlations. The agent is deployed in an online system to generate predictions in real time for multimodal inputs and participate in memory compression. The results are simultaneously written to a hierarchical memory structure. User behavior feedback is collected, including information on retrieval, error correction, task results, and usage frequency. These signals are used to evaluate the prediction and compression effectiveness. The agent is incrementally updated using the feedback data to optimize its predictive capabilities and compression strategy. The agent gradually transforms high-frequency anomalies into predictable patterns, achieving long-term evolution.

[0164] In other words, the intelligent agent can realize a closed loop from "data → model → policy → online operation → feedback → continuous learning", which enables the predictive model to continuously improve and drives more efficient memory compression.

[0165] Based on the above data storage method, embodiments of this disclosure also provide a data storage device. The following will be combined with... Figure 6 The device is described in detail.

[0166] Figure 6 A schematic block diagram of a data storage device according to an embodiment of the present disclosure is shown.

[0167] likeFigure 6 As shown, the data storage device 600 of this embodiment includes a first determining module 610, an obtaining module 620, a second determining module 630, and a storage module 640.

[0168] The first determining module 610 is used to determine reference event information for the target event information. The reference event information is determined based on at least one piece of associated information related to the target event information. In one embodiment, the first determining module 610 can be used to execute step S210 described above, which will not be repeated here.

[0169] The obtaining module 620 is used to obtain first difference information between the target event information and the reference event information. In one embodiment, the obtaining module 620 can be used to perform step S220 described above, which will not be repeated here.

[0170] The second determining module 630 is used to determine the storage strategy corresponding to the target event information based on the first difference information. In one embodiment, the second determining module 630 can be used to execute the step S230 described above, which will not be repeated here.

[0171] The storage module 640 is used to associate and store, according to a storage strategy, the second difference information between the target event information and the reference event information with the event identifier corresponding to the target event information, wherein the second difference information is the same as or related to the first difference information. In one embodiment, the storage module 640 can be used to execute step S240 described above, which will not be repeated here.

[0172] According to embodiments of this disclosure, any plurality of modules among the first determining module 610, obtaining module 620, second determining module 630, and storage module 640 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the first determining module 610, obtaining module 620, second determining module 630, and storage module 640 can be at least partially implemented as hardware circuitry, such as field-programmable gate arrays, programmable logic arrays, systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits, or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any appropriate combination of any of these three implementation methods. Alternatively, at least one of the first determining module 610, obtaining module 620, second determining module 630, and storage module 640 can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.

[0173] Figure 7 A block diagram schematically illustrates an electronic device suitable for implementing a data storage method according to embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0174] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded into random access memory (RAM) 703 from storage unit 708. The RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0175] Multiple components in device 700 are connected to input / output (I / O) interface 705, including: input unit 706, such as a keyboard, mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as a disk, optical disk, etc.; and communication unit 709, such as a network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0176] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as model training and text processing methods. For example, in some embodiments, the model training and text processing methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the model training and text processing methods described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform model training and text processing methods by any other suitable means (e.g., by means of firmware).

[0177] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0178] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0179] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0180] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0181] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0182] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0183] Embodiments of this disclosure also provide a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0184] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In embodiments of this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include the read-only memory, and / or random access memory, and / or one or more memories other than read-only memory and random access memory described above.

[0185] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this disclosure.

[0186] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices or magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed as signals over a network medium, and downloaded and installed via a communication component, and / or installed from a removable medium. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0187] In embodiments of this disclosure, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a processor, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0188] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The program code can execute entirely on a user computing device, partially on a user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0189] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0190] Those skilled in the art will understand that the features described in the various embodiments of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments of this disclosure can be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.

Claims

1. A data storage method, comprising: Reference event information for determining target event information, wherein the reference event information is determined based on at least one type of associated information related to the target event information; Obtain the first difference information between the target event information and the reference event information; Based on the first difference information, determine the storage strategy corresponding to the target event information; According to the storage strategy, the second difference information between the target event information and the reference event information is associated with the event identifier corresponding to the target event information and stored together. The second difference information is the same as or related to the first difference information.

2. The method according to claim 1, wherein the at least one piece of associated information includes: At least one of the following: general knowledge, domain knowledge, user information, current task context, and historical event information; The event identifier corresponding to the target event information has a mapping relationship with the at least one associated information; the event identifier is generated based on the at least one associated information, or generated based on the target event information.

3. The method according to claim 1 or 2, wherein the reference event information includes predicted event information, and the predicted event information is generated by predicting the target event based on the at least one related information; The first difference information includes surprise information of the target event information relative to the predicted event information, wherein the surprise information characterizes the expected deviation between the target event information and the predicted event information.

4. The method according to claim 3, wherein obtaining the first difference information between the target event information and the reference event information comprises: The objective event information is obtained from one of the following: semantic error, visual error, and user behavior error related to the target event information; The surprise level information is determined based on one of the semantic error, visual error, and user behavior error related to the target event information. or, Obtain at least two of the following: semantic error, visual error, and user behavior error related to the target event information between the target event information and the predicted event information; Based on the current task context related to the target event information, assign weight information to at least two of the semantic error, visual error, and behavioral error; Based on the weight information, at least two of the semantic error, visual error, and behavioral error are weighted to obtain a weighted result. The surprise level information is determined based on the weighted result.

5. The method according to claim 3, wherein determining the storage strategy corresponding to the target event information based on the first difference information includes: In response to the surprise level information not meeting the target requirements, the storage strategy corresponding to the target event information is determined to be a first storage strategy. The first storage strategy indicates that the essential representation contained in the second difference information is stored, and the second difference information is associated with the first difference information. Alternatively, in response to the surprise level information meeting the target requirements, the storage strategy corresponding to the target event information is determined to be a second storage strategy, the second storage strategy indicating the storage of all information contained in the second difference information, the second difference information being associated with the first difference information.

6. The method according to claim 5, wherein the step of associating and storing the second difference information between the target event information and the reference event information with the event identifier corresponding to the target event information according to the storage strategy includes: If the target event information has associated historical event information, the historical event identifier of the historical event information is determined as the event identifier corresponding to the target event information; according to the first storage strategy, the essential representation contained in the second difference information is associated and stored with the historical event identifier; Alternatively, if the target event information does not have associated historical event information, a new event identifier is generated based on the at least one associated information or the target event information, and used as the event identifier corresponding to the target event information; according to the first storage strategy, the essential representation contained in the second difference information is associated and stored with the new event identifier; Alternatively, if the target event information has associated historical event information, a new event identifier is generated based on the at least one associated information or the target event information, and used as the event identifier corresponding to the target event information, and the new event identifier is associated with the historical event identifier of the historical event information. According to the second storage strategy, the full information contained in the second difference information is associated with and stored with the new event identifier; Alternatively, if the target event information does not have associated historical event information, a new event identifier is generated based on the at least one associated information or the target event information, and used as the event identifier corresponding to the target event information. According to the second storage strategy, the full information contained in the second difference information is associated with and stored with the new event identifier.

7. The method according to claim 1 or 2, wherein the reference event information includes at least one of event summary information, aggregation information, statistical information, and common representation information, and the reference event information is generated based on intent understanding of the at least one related information; The first difference information includes at least one of the following: new information, abnormal information, and uncovered information of the target event information relative to the event summary information.

8. The method according to claim 7, wherein determining the storage strategy corresponding to the target event information based on the first difference information includes: In response to the fact that the amount of data contained in the first difference information is less than the data amount threshold, the storage strategy corresponding to the target event information is determined to be the first storage strategy, and the first storage strategy indicates that the essential representation contained in the second difference information is stored. Alternatively, in response to the fact that the amount of data contained in the first difference information is greater than or equal to the data amount threshold, the storage strategy corresponding to the target event information is determined to be the second storage strategy, and the second storage strategy indicates that the full amount of information contained in the second difference information is stored.

9. The method according to claim 8, wherein the step of associating and storing the second difference information between the target event information and the reference event information with the event identifier corresponding to the target event information according to the storage strategy comprises: The historical event identifier of the historical event information is determined as the event identifier corresponding to the target event information; According to the first storage strategy, the essential representation contained in the second difference information is associated with and stored with the historical event identifier; Alternatively, a new event identifier may be generated based on the at least one associated information or the target event information, and used as the event identifier corresponding to the target event information, and the new event identifier may be associated with the historical event identifier of the historical event information; According to the second storage strategy, the full information contained in the second difference information is associated with and stored with the new event identifier.

10. An electronic device, comprising: At least one processor; as well as The memory connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform the following operations: Reference event information for determining target event information, wherein the reference event information is determined based on at least one type of associated information related to the target event information; Obtain the first difference information between the target event information and the reference event information; Based on the first difference information, determine the storage strategy corresponding to the target event information; According to the storage strategy, the second difference information between the target event information and the reference event information is associated with the event identifier corresponding to the target event information and stored together. The second difference information is the same as or related to the first difference information.