Data management method, electronic device, and computer program product

By calculating multiple metrics and effective time of data records, and combining them with business categories to calculate retention index, the problem of inaccurate data value assessment in existing technologies has been solved, achieving efficient data management, reducing storage costs, and improving data retrieval efficiency and service quality.

CN122284918APending Publication Date: 2026-06-26KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610395508.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-27
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing data management methods such as LRU and LFU cannot effectively distinguish the intrinsic value of data, leading to the accidental deletion of important data or the retention of invalid data, resulting in increased storage costs and reduced retrieval efficiency, which affects the quality of artificial intelligence services.

Method used

By acquiring multiple metrics, effective time, and business category of the data records, a first value, a second value, and a third value are calculated. These values ​​are then combined to determine the retention index, and the retention index is used to decide whether to retain or delete the data.

Benefits of technology

It enables accurate assessment of data value, avoids accidental deletion of important data, reduces storage costs, and improves data retrieval efficiency and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122284918A_ABST
    Figure CN122284918A_ABST
Patent Text Reader

Abstract

This disclosure provides a data management method, electronic device, and computer program product. The method first scans data records to obtain multiple indicators, effective time, and corresponding business categories of the data records. The multiple indicators are used to represent the usage of the data records. Then, a first value representing the importance of the data records is determined through the multiple indicators. A second value representing the memory strength of the data records is determined through the effective time. A third value of the data records is determined through the number of data records contained in the business category corresponding to the data records. The second value is negatively correlated with the cumulative duration since the effective time, and the third value is negatively correlated with the number of data records. Then, the retention index of the data records is determined by combining the first, second, and third values. If the retention index is less than the index threshold, the data records are deleted from their storage locations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure specifically relates to data management methods, electronic devices, and computer program products. Background Technology

[0002] In the context of the widespread application of artificial intelligence, data memory has become a key technology for realizing personalized, continuous, and contextualized intelligent services. By storing user data through data memory, it is possible to provide users with "long-term companionship" services.

[0003] When storage space is limited, it is necessary to remove some non-critical data, that is, to "forget" some data, in order to make room for new data.

[0004] Existing strategies largely rely on mechanical rules such as "Least Recently Used (LRU)" or "Least Frequently Used (LFU)" to determine which data to discard. However, these methods have inherent flaws: they cannot distinguish the intrinsic value of the data being stored. For example, data that should be retained as important may be considered low-value data and discarded due to its low usage frequency; while expired data that could be discarded may continue to be considered high-value data and retained because of its high past usage frequency. This leads to increased storage costs, decreased retrieval efficiency of stored data, and consequently, deviations or errors in the services provided by artificial intelligence. Summary of the Invention

[0005] This disclosure provides data management methods, electronic devices, readable storage media, and computer program products.

[0006] This disclosure firstly proposes a data management method, comprising: scanning data records to obtain multiple indicators, effective time, and corresponding business categories of the data records, wherein the multiple indicators are used to represent the usage of the data records; determining a first value representing the importance of the data records based on the multiple indicators, determining a second value representing the memory strength of the data records based on the effective time, and determining a third value of the data records based on the number of data records included in the business category corresponding to the data records, wherein the second value is negatively correlated with the cumulative duration since the effective time, and the third value is negatively correlated with the number of data records; determining a retention index of the data records by combining the first value, the second value, and the third value; and deleting the data records from their storage locations if the retention index is less than an index threshold.

[0007] According to some embodiments of this disclosure, the step of scanning the data records is performed in response to the fulfillment of a triggering condition. The triggering condition is fulfilled when any of the following is met: the start time of the data scanning task is reached, or the amount of stored data reaches a storage threshold.

[0008] According to some embodiments of this disclosure, the multiple indicators include one or more of the following: popularity, number of user feedbacks, number of updates, recent usage time, and effective status.

[0009] According to some embodiments of this disclosure, the first value is a weighted average of the multiple indicators.

[0010] According to some embodiments of this disclosure, the second value decreases non-linearly with the increase of the accumulated duration, and the rate of decrease decreases with the increase of the accumulated duration.

[0011] According to some embodiments of this disclosure, different business categories differ in at least one of business scenarios, business tags, and business matters.

[0012] According to some embodiments of this disclosure, one or more of the following are satisfied: different business scenarios correspond to different artificial intelligence services, the business tags are used to represent business characteristics, and the business items are used to represent user intent.

[0013] According to some embodiments of this disclosure, the retention index is obtained by weighted fusion calculation of the first value, the second value, and the third value.

[0014] A second aspect of this disclosure provides an electronic device, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, causing the processor to perform the method described in any of the above embodiments.

[0015] A third aspect of this disclosure provides a readable storage medium storing a computer program that, when executed by a processor, is used to implement the method described in any of the above embodiments.

[0016] The fourth aspect of this disclosure provides a computer program product comprising a computer program that, when executed by a processor, is used to implement the method described in any of the above embodiments. Attached Figure Description

[0017] The accompanying drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.

[0018] Figure 1 A schematic diagram of the overall flow of a data management method M100 according to some embodiments of the present disclosure is shown.

[0019] Figure 2 A schematic diagram of the overall flow of a data management method M100 according to other embodiments of the present disclosure is shown.

[0020] Figure 3 A flowchart of a data management method according to some embodiments of the present disclosure is shown.

[0021] Figure 4 This is a schematic block diagram of the structure of a data management device according to one embodiment of the present disclosure.

[0022] Figure 5 This is a schematic block diagram of an electronic device 1000 according to one embodiment of the present disclosure. Detailed Implementation

[0023] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present disclosure are shown in the accompanying drawings.

[0024] It should be noted that, where there is no conflict, the embodiments and features described in this disclosure can be combined with each other. The technical solutions of this disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0025] Unless otherwise stated, the exemplary implementations / embodiments shown are to be understood as providing exemplary features of various details that provide ways in which the technical concepts of this disclosure can be implemented in practice. Therefore, unless otherwise stated, the features of various implementations / embodiments may be additionally combined, separated, interchanged and / or rearranged without departing from the technical concepts of this disclosure.

[0026] The terminology used herein is for the purpose of describing particular embodiments and is not restrictive. As used herein, unless the context clearly indicates otherwise, the singular forms “a” and “the” are intended to include the plural forms as well. Furthermore, when the terms “comprising” and / or “including” and variations thereof are used in this specification, it indicates the presence of the stated features, integrals, steps, operations, parts, components, and / or groups thereof, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, parts, components, and / or groups thereof. It should also be noted that, as used herein, the terms “substantially,” “about,” and other similar terms are used as approximate terms rather than as terms of degree, thus explaining the inherent biases in measurements, calculated values, and / or provided values ​​that would be recognized by one of ordinary skill in the art.

[0027] With the rapid development of artificial intelligence technology, especially the deep integration of large language models (LLMs) and intelligent agents, building efficient and sustainable long-term memory systems has become crucial for achieving higher levels of intelligence. As the core carrier of personalized and contextualized interactions by intelligent agents, the management mechanism of long-term memory directly determines the system's performance ceiling and application value.

[0028] Currently, a common practice is to use mechanical linear rules to evict and forget already memorized data, such as the LRU (Least Recently Used) rule or the LFU rule. The LRU rule evicts data that has not been used in the last few weeks. The LFU rule evicts data that has been used the least frequently.

[0029] However, the aforementioned data eviction rules have inherent flaws: they fail to distinguish the intrinsic value of memories, easily discarding low-frequency but crucial data while retaining previously frequently used but now outdated "noise data." For example, user allergy history data is low-frequency but important. If a data point indicates a user is allergic to pollen, discarding that data might lead to recommending properties near parks or gardens when the user has rental needs, resulting in ineffective property viewings. Similarly, outdated user preference data is high-frequency but invalid. If a data point indicates a user was interested in large homes in the past, but their latest interest is only in small homes, retaining that data might lead to recommending properties the user is not interested in.

[0030] The above situations can lead to low data retrieval efficiency, high data storage costs, and a significant increase in the risk of the model generating illusions. When providing services, artificial intelligence may give incorrect answers, or there may be situations such as disjointed dialogue, failure of personalized services, or even accidental deletion of security data such as permissions, authentication status, and sensitive operation records, creating security risks.

[0031] Therefore, this disclosure proposes a data management method.

[0032] The data management method provided in this disclosure can be automatically executed by a data management system, which can be deployed on a server or on a terminal device, and manages data in a database on the server or cached data on the terminal device, such as removing stored data. In this disclosure, both the data management system and the server can each include at least one processor and at least one memory.

[0033] In this disclosure, "terminal device" can be any type of electronic device, such as a tablet computer, a laptop computer, or a desktop computer. Additionally, the server can be a physical server or a cloud server; this disclosure does not limit the type of server.

[0034] Figure 1 A schematic diagram of the overall flow of a data management method M100 according to some embodiments of this disclosure is shown. For example... Figure 1 The method shown includes steps S110, S120, S130, and S140. This method can be executed by electronic devices such as computers and servers.

[0035] S110: Scan the data records to obtain multiple metrics, effective time, and corresponding business category of the data records. These metrics indicate the usage status of the data records.

[0036] Data records can be stored in a server's database. Each data record is equivalent to a data entry, and the business can set up one or more dedicated databases as a memory to store the stored data records. These stored data records can be generated when users use services provided by an agent. For example, data records can be user behavior data generated when users use an app (application), specifically including text information entered by the user on a page, the user's search content and browsing history, and dialogue content entered by the user when interacting with the agent provided by the app, etc. In this disclosure, all these data records are data that has been acquired and stored (remembered) in a database or other types of data sources.

[0037] A single data record can correspond to multiple metrics, which act as the record's metadata. These metrics correspond to various attribute dimensions, describing characteristics such as the usage of the data record. For example, among these metrics might be a popularity metric (a1), which describes the frequency with which the data record is actively accessed or used by the system. Another example is a recent usage time metric (a2), which describes the time when the data record was last accessed or used. These metrics, such as popularity metric (a1) and recent usage time metric (a2), can reflect the importance and value of the data record to a certain extent. It's understandable that the metrics corresponding to a data record can also include other metrics that reflect its importance and value; the specific metrics included can be set based on the business type and scenario of the business entity.

[0038] Data records can also correspond to an effective time and a business category, which are essentially the metadata of the data record. The effective time refers to the moment when the data record becomes valid. This could be the moment the data record was created, or the moment when changes in relevant regulations or systems make the data record valid again. The business category describes the contextual information such as the scenario and theme in which the data record was created. For example, if the data record is user demand information recorded when a user fills in their purchase needs in the app's recommendation module, then the business category of the data record could include "recommendation module" corresponding to the scenario context and "determine demand" corresponding to the theme context.

[0039] S120, a first value representing the importance of a data record is determined using the aforementioned multiple indicators; a second value representing the memory strength of a data record is determined using the effective date; and a third value representing the business category corresponding to the data record is determined using the number of data records contained in the business category corresponding to the data record. The second value is negatively correlated with the cumulative duration since the effective date, and the third value is negatively correlated with the number of data records.

[0040] The various metrics, effective times, and business categories obtained by scanning data records are each used to perform relevant numerical calculations.

[0041] The obtained metrics are used to calculate a first value H, which reflects the importance of the data record, that is, the memory value of the data record. For example, the higher the popularity of a data record and the smaller the time difference between its most recent use and the implementation time of the method disclosed herein, the higher the importance of the data record represented by the calculated first value H, the higher its memory value, and the higher the demand for the data record to be retained.

[0042] The first value H can be calculated by fitting the above-mentioned multiple indicators to a machine learning model. For example, a neural network model can be trained using supervised learning to predict the first value H based on the above-mentioned multiple indicators.

[0043] The obtained effective time can be used to calculate the second value F. The cumulative duration refers to the time difference between the effective time and the implementation time of the method disclosed herein. The memory strength is negatively correlated with the cumulative duration. The higher the value of this time difference, the longer the effective duration of the data record, and the lower the memory strength. The memory strength is positively correlated with the second value F. Therefore, the smaller the second value F, the lower the need for the data record to be retained.

[0044] The specific calculation method for the second value F can be obtained through a linear proportional function, such as using the function F=-kx, where F is the second value and x is the cumulative duration.

[0045] The obtained business category can be used to calculate the third value S. The third value S represents the scarcity of the data record. The fewer the total number of data records in the business category L to which the data record belongs, the higher the third value S corresponding to business category L, indicating a higher scarcity of data records in business category L and a higher demand for retaining the data record. The third value of each data record in the same business category is the same as the third value of the business category to which it belongs; therefore, the third value of the data record is the third value S.

[0046] The specific calculation method for the third value S can be obtained through an inverse proportional function, such as S=1 / [ln(1+N)], where S is the third value and N is the total number of data records contained in the business category to which this data record belongs.

[0047] It is understandable that the calculation order of the first value H, the second value F, and the third value S can be arbitrary; for example, they can be calculated sequentially or in parallel. S130, combine the first, second and third values ​​of the data record to determine the retention index of the data record.

[0048] The first value H, the second value F, and the third value S are the three factors used to calculate the retention index D. The retention index is calculated by comprehensively considering the importance of the data record (represented by the first value H), the memory strength (timeliness) of the data record (represented by the second value F), and the scarcity of the data record (represented by the third value S). A higher retention index means a higher probability that the data record will be retained; a lower retention index means a higher probability that the data record will be deleted from memory.

[0049] S140, if the retention index is less than the index threshold, the data record is deleted from its storage location.

[0050] After assessing the retention index of all scanned data records, if the retention index is less than the threshold, it indicates that the importance, timeliness, and scarcity of the data record are insufficient to warrant retention, thus failing to meet the memory requirements and allowing for deletion. Conversely, if the retention index is not less than the threshold, it indicates that the importance, timeliness, and scarcity of the data record support its retention, thus meeting the memory requirements and requiring no deletion. This allows the memory bank to retain critical data while removing lower-value memories to maintain its size, preventing it from continuously expanding with the addition of new data, thereby reducing storage costs and ensuring efficient retrieval during memory searches.

[0051] It's important to note that the combination of the first value (H), the second value (F), and the third value (S) is highly effective in determining the value of remembered data. However, relying solely on H and F to judge the value of data records could lead to the deletion of low-intensity, long-standing data with low recall strength. This data might be crucial; for example, in real estate transactions, if user U mentioned a pollen allergy in a past consultation, relying solely on H and F might mistakenly delete this data, resulting in the business recommending properties near gardens and parks, which the user is unlikely to be interested in, thus lowering service quality. The introduction of the third value (S) mitigates this risk. It protects this type of data, which is scarce but crucial in a specific context (such as budget). This information, representing the user's true needs, may have a much higher value than ordinary data such as a user's property browsing history.

[0052] A / B testing (a testing method used to compare different versions by conducting control experiments) in real business scenarios revealed that if a third value, S, is used in addition to the first value H and the second value F when calculating the retention index, the false deletion rate of key customer information (such as user needs) will be significantly reduced, which is beneficial for follow-up on business opportunities (user needs) and improvement of service quality.

[0053] If the value of a data record is judged solely by the second value F and the third value S, then long-standing and less scarce memory data will be deleted. While it may retain relatively scarce key data, it will delete some highly important data and retain a lot of low-value data with low importance, resulting in a lot of "noisy memories" in the stored data.

[0054] If the value of a data record is judged solely by the first value H and the third value S, then memorized data that is less important and less scarce will be deleted. This may include some data with high memory intensity but not high importance or scarcity, leading to the erroneous deletion of critical data.

[0055] According to the data management method proposed in the embodiments of this disclosure, a complete biomimetic memory management framework is constructed. By simulating the inherent laws of human brain memory, the dynamic optimization problem of long-term memory of intelligent agents is solved. A first value H, a second value F, and a third value S are used to judge the memory value of memorized data records. The first value H, the second value F, and the third value S each provide their own judgment rules, providing stable judgment ability and business practicality of memory management in a collaborative manner. It can remove some non-critical data, thereby reserving storage space for new memory data. Furthermore, based on the salience memory enhancement effect, the third value S can protect unique knowledge, accurately distinguish the memory value of data records, avoid accidentally deleting some low-frequency but crucial memory data, and ensure the quality of business services.

[0056] For example, the step of scanning data records may be performed in response to a triggering condition being met. Figure 2 A schematic diagram of the overall flow of a data management method M100 according to other embodiments of this disclosure is shown. (See also...) Figure 2 Step S110 can specifically be as follows: In response to the fulfillment of the triggering condition, scan the data records to obtain multiple indicators, effective time, and corresponding business categories of the data records. The triggering condition is fulfilled when any of the following conditions are met: the start time of the data scanning task is reached, or the amount of stored data reaches the storage threshold.

[0057] The execution of step S110 (scanning data records to obtain multiple indicators, effective time and corresponding business category of data records) can be triggered by triggering conditions.

[0058] The triggering condition can include only a time condition, which can be a user-preset start time. When the actual time arrives at the start time, the server automatically starts a data scanning task on its stored data records. The time condition can include only a timed trigger, only a delayed trigger, or both timed and delayed triggers.

[0059] In the timed triggering method, the time node is an absolute time. Users can set up regular data scanning tasks in the system (e.g., 3 AM every calendar day). The system will then allocate computing resources to implement the data management method disclosed herein during off-peak server hours, automatically initiating the memory optimization process. The timed triggering method can be implemented through a scheduled task scheduler.

[0060] In the delayed triggering method, the time node is a relative time. Users can set a delay after completing a specific data processing task, after which the system automatically starts the memory optimization process. For example, after receiving and memorizing a large amount of data, memory optimization begins to ensure sufficient free storage space in the system's memory bank. The delayed triggering method can be implemented using scripts.

[0061] The triggering condition can also include only a threshold condition, which corresponds to the threshold triggering method. The threshold condition can be a user-preset threshold for the remaining capacity of the memory bank. The system can monitor the remaining capacity of the memory bank in real time. When the remaining capacity falls below the threshold, it indicates that the memory bank has limited free space, and therefore the memory optimization process can be automatically initiated immediately to remove some unimportant memory data. The threshold triggering method can be implemented using a capacity counter.

[0062] It is understandable that the triggering conditions can also include time conditions and threshold conditions at the same time, and the memory optimization process can be automatically triggered when either condition is met.

[0063] Therefore, by triggering the system to automatically start the memory optimization process through timed, delayed, or capacity thresholds, the system can proactively clean up memory data when the system load is controllable, avoiding runtime performance degradation or storage overflow.

[0064] The scanned data records may include some or all of the data records in the data source. When scanning the data records in the memory (step S110), all the data in the memory can be scanned, allowing for a comprehensive check of all data records and maximizing the removal of unnecessary data. Alternatively, only a portion of the data in the memory can be scanned, excluding critical data that is clearly deemed undeletable from the removal scope. For example, this ensures that high-privilege, security-related data is not accidentally deleted.

[0065] The multiple metrics obtained by scanning data records (step S110) may include one or more of the following: popularity, number of user feedbacks, number of updates, recent usage time, and active status.

[0066] Heat, also known as popularity, represents the frequency with which a data record is actively accessed or associated with by the system. It reflects the activity level of the remembered data during actual interactions, emphasizing proactive access from the system side rather than passive browsing by the user. Each time an agent references or activates a data record in a dialogue, recommendation, or decision-making process, the heat value of that data record increases by 1. For example, if a user views properties multiple times, and the agent recommends properties based on the remembered data record of "preferring three-bedroom apartments," each recommendation can be considered an associated use.

[0067] The number of user feedback responses indicates the total number of times a user provides feedback on a data record. This is used to determine whether the remembered data is accepted by the user or triggers effective interaction; the more responses, the more "useful" the memory is considered. User feedback can be explicit or implicit. Explicit feedback refers to users clicking buttons to express their feedback, giving a "like," or correcting the information. Buttons can indicate whether the information was helpful or inaccurate. Correcting information could involve a user typing, "I meant a four-bedroom apartment, not a three-bedroom." Implicit feedback occurs when a user focuses their conversation on the data record during dialogue with the agent. For example, mentioning the memory data and continuing the conversation to a conclusion (such as paying a deposit) can be considered positive feedback by the system.

[0068] Understandably, the aforementioned metrics can also include user feedback scores, which refer to the average score given by users to a particular data record. For example, if a user repeatedly inputs phrases like "budget information is accurate" to indicate the accuracy of a data record during a conversation with the agent or in other scenarios, the feedback frequency is high, and the feedback scores are consistently high. Combining the number of feedback instances and the feedback scores can be used to distinguish between "high-frequency, low-quality" and "low-frequency, high-quality" memory data. For instance, a data record that receives many feedback instances but has a low average score may be less valuable than a data record that receives only one feedback instance but has a high feedback score.

[0069] The update count indicates the number of times a data record has been modified. It is used to characterize the dynamic accuracy of data memory. Frequently updated data records may indicate that the data record is on a critical decision-making path and has high business sensitivity. For example, if the data record is initially "budget X yuan", and later the user enters "actually I can give Y yuan" when talking to the agent, the system will update the data record about the budget, and the update count will be incremented by 1.

[0070] The most recent time refers to the moment when the data record was last actively accessed or associated with by the system.

[0071] It's understandable that the aforementioned metrics can also include internal similarity counts, which represent the number of other data records in the current memory that are semantically or content-wise highly similar to the given data record. Specifically, the semantic similarity between two data records can be determined by calculating the cosine similarity. If the semantic similarity is higher than a similarity threshold, the internal similarity count of each data record is incremented by 1. For example, "I want to buy a house near a school" and "My child needs to go to school and needs to be near a good school district" can be considered similar data records. A higher internal similarity count indicates that the data record is more "redundant" or "common." A lower internal similarity count indicates that the data record is more unique, and its scarcity can be reflected by a third value S, thus preventing the data record from being discarded as low-value data.

[0072] The "effectiveness status" indicates whether a data record is valid and can be used in conjunction with its "effectiveness time." For example, if a user inputs "I am allergic to a certain material" during a conversation with an agent, this information is permanent health information, and its effective status will usually remain "valid." Conversely, if a user inputs "I want to buy a used house" on day d1 and then inputs "Used houses are too expensive, I want to rent" on day d2, the previously input "I want to buy a used house" will change from an effective status to an invalid status. This is understandable because the two user inputs belong to different business categories; therefore, the content input on day d1 will not be updated (overwritten) by the content input on day d2. Instead, a new data record will be created to store the content input on day d2.

[0073] Therefore, by combining indicators such as popularity, update frequency, and feedback frequency, a first value H can be quantitatively calculated to represent the importance of information in a data record. This simulates the physiological process in which the human brain strengthens memory synapses through repeated activation and emotional intensity. The higher the first value H, the more important the data record.

[0074] The first value can be a weighted average of multiple indicators. For example, different weights can be assigned to different indicators, and the first value H can be obtained by weighting the weights with the corresponding indicator values. Since the currently selected indicators (popularity, feedback, etc.) are efficient and stable proxy indicators that can directly reflect the value of memory, a weighted summation method is used to determine the first value. In actual industrial scenarios, this method can better balance effectiveness and computational overhead, avoiding increased system complexity and latency, which would otherwise be detrimental to large-scale applications with high-frequency triggering.

[0075] The second value can decay non-linearly with increasing cumulative duration, and the decay rate can decrease as the cumulative duration increases. For example, the second value can be calculated based on the following formula: F = 1 - k × M, where M is ΔT raised to the power of c, F is the second value, k and c are constants greater than 0, and ΔT is the cumulative duration (the time difference between the effective date and the implementation date of the method disclosed herein). This formula can calculate the degree of non-linear decay of memory over time, achieving a precise digital simulation of the non-linear decay law of memory. The larger the calculated second value F, the lower the priority of memory forgetting, and the more necessary it is to retain this data record.

[0076] It should be noted that the calculation formula for the second value F is one of the core engines driving the entire automated forgetting decision-making system. This disclosure uses this calculation formula as a pre-prediction control instruction, and its main function is to evaluate the value of stored long-term memories and intelligently eliminate them, rather than to select which short-term memories can be stored.

[0077] Compared to managing memory data through post-event analysis, the pre-event prediction method can predict the current value (F value) of the data before it is used and proactively decide whether to discard the data. This changes the memory management approach from "discovering the problem and discarding the data after providing low-quality service to users due to expired data" to "pre-determining and discarding expired data in advance through batch calculations to achieve preventive optimization" and thus provide high-quality service.

[0078] Different business categories can differ in at least one of the following: business scenario, business tag, and business item. In other words, if two data records differ in any one of the following three aspects, then these two data records belong to two different business categories.

[0079] For the data management method provided in this disclosure, one or more of the following can be satisfied: different business scenarios correspond to different artificial intelligence services, business tags can be used to represent business characteristics, and business matters can be used to represent user intent.

[0080] A business scenario can correspond to an AI service. For example, if a business provides multiple different AI services in an app, such as a customer acquisition assistant and an office assistant, and a user uses multiple different AI services, the data records stored will include the corresponding multiple business scenarios.

[0081] Business characteristics can be keywords or key semantic features in data records, while user intent is equivalent to the theme of the data record. For example, if the data record is "I want to buy a new house, budget X yuan" entered by the user, then the business tag could include "purchase amount" and the user intent could be "new house purchase". Similarly, if the data record is "I want to buy a three-bedroom second-hand house" entered by the user, then the business tag could include "housing preferences" and the user intent could be "second-hand house purchase". For other data records, the user intent could also be themes such as "writing weekly reports", "drawing flowcharts", or "data analysis". The specific content that business scenarios, business tags, and business matters can include can be set by the business party according to its own business and working conditions; this disclosure does not impose any restrictions on this.

[0082] When determining the retention index of a data record by combining the first, second, and third values ​​(step S130), the retention index can be obtained by weighted fusion calculation of the first, second, and third values. For example, the first value H, the second value F, and the third value S are used as three core biomimetic factors and input into the decision function D=f(H, F, S) for weighted fusion to obtain the retention index D.

[0083] Weighted fusion can be achieved through either weighted linear summation or weighted geometric mean. Assuming w1, w2, and w3 are the weights of the first value H, the second value F, and the third value S, respectively, then when calculating using weighted linear summation, D = w1 × H + w2 × F + w3 × S. Alternatively, when calculating using weighted linear summation, w1, w2, and w3 can be normalized so that their sum equals 1, and then a weighted geometric mean can be applied, where D = H. w1 +F w2 +S w3 Both of these calculation methods have low computational overhead and are suitable for high-frequency batch processing.

[0084] After weighted fusion, a retention index D is obtained. The retention index D is compared with a predefined index threshold θ. If D < θ, the memory is considered to have low value, and a forgetting operation is performed, that is, it is hard deleted from the active memory bank to release storage space. If D ≥ θ, the memory is considered to have retention value and is retained, continuing to remain in the fast-retrieval memory bank.

[0085] Figure 3 A flowchart of a data management method according to some embodiments of the present disclosure is shown. Figure 3 The detailed description provided is for the purpose of better understanding the technical solution of this disclosure and should not be considered as a limitation on the scope of protection of this disclosure. In the process of implementing the technical concept of this disclosure, one or more steps may be omitted, or other alternative methods may be adopted.

[0086] See Figure 3 The data management method includes a data acquisition phase, a calculation phase, and a judgment phase. In the scanning phase, all memory data (i.e., data records) in the memory bank are scanned and metadata is extracted to obtain a multi-dimensional attribute vector for each memory data entry. This vector includes metrics such as popularity, number of feedbacks, feedback score, number of internal similarities, recent usage time, number of updates, effective status, effective time, and the total number of memories under the business category. Specifically, the multi-dimensional attribute vector for one memory data entry Q is: {"id": "M789", "content": "My actual budget is X yuan", "Business scenario": "Agent1", "Business tag": "Purchase amount", "Business theme": "New house purchase", "Effective time": "y1 year m1 month d1 day", "Recent usage time": "y1 year m1 month d2 day", "User feedback score": 5, "Number of user feedbacks": 1, "Popularity": 2, "Number of updates": 0, "Number of internal similarities": "2", "Effective status": "Valid"}.

[0087] During the calculation phase, the first value H, the second value F, and the third value S are calculated in parallel. Specifically, the first value H is calculated using a weighted linear summation method, quantifying the degree to which the memory is valued by the user or verified by the system, simulating the "re-activation + emotional reinforcement" mechanism; the third value S is calculated using an inverse proportional function, simulating the natural decay of information that has not been used for a long time in the human brain; and the third value S is calculated based on the total number of memories, protecting key information unique in a specific context and preventing accidental deletion due to low frequency. Then, the retention index D = w1 × H + w2 × F + w3 × S is calculated using a weighted linear summation method.

[0088] Finally, in the judgment phase, the retention index D is compared with the index threshold θ. When D < θ, the memory data Q is deleted from the memory bank; when D ≥ θ, no deletion is performed. Through the organic integration of multi-dimensional biomimetic models, the intelligent agent memory system can simulate the inherent laws of biological memory, gaining autonomous "metabolism" capabilities. This enables precise evaluation and adaptive optimization of memory value, ensuring the retention of high-value, scarce, and not-completely-forgotten information to support subsequent accurate services. It significantly improves the information quality and retrieval efficiency of the memory bank, effectively reduces storage load and computational overhead, and enhances the decision-making reliability and interactive intelligence level of the intelligent agent. This provides key technical support for building efficient and robust next-generation artificial intelligence systems.

[0089] Based on any of the above embodiments, this disclosure also provides a data management device. Figure 4 This is a schematic block diagram of the structure of a data management device according to one embodiment of the present disclosure. Figure 4 As shown, the data management device includes: a data scanning module 110, a numerical calculation module 120, an index determination module 130, and a data deletion module 140.

[0090] The data scanning module 110 is used to scan data records and obtain multiple indicators, effective time, and corresponding business categories of the data records. Among them, the aforementioned multiple indicators are used to indicate the usage status of the data records.

[0091] The numerical calculation module 120 is used to determine a first value representing the importance of a data record based on the aforementioned multiple indicators, a second value based on the effective date, and a third value based on the number of data records contained in the corresponding business category. The second value is negatively correlated with the cumulative duration since the effective date, and the third value is negatively correlated with the number of data records.

[0092] The index determination module 130 is used to determine the retention index of the data record by combining the first, second and third values ​​of the data record.

[0093] The data deletion module 140 is used to delete data records from their storage location when the retention index is less than the index threshold.

[0094] The aforementioned data management device can be in the form of computer software, and each module of the data management device can be implemented through computer software modules. The specific implementation process of the functions and roles of each module in the aforementioned device is detailed in the corresponding steps of the above method, and will not be repeated here.

[0095] The data management method implemented in the specific embodiments of this disclosure can be executed by electronic devices such as computers and servers.

[0096] Therefore, based on any of the above embodiments, this disclosure also provides an electronic device that can execute the data management method of any of the embodiments described above.

[0097] Figure 5 This is a schematic block diagram of an electronic device 1000 according to one embodiment of the present disclosure.

[0098] The hardware architecture of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnect buses and bridges, depending on the specific application of the hardware and overall design constraints. Bus 1100 connects various circuits, including one or more processors 1200, memory 1300, and / or hardware modules. Bus 1100 can also connect various other circuits 1400, such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0099] Bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one connection line is used in this diagram, but this does not imply that there is only one bus or only one type of bus.

[0100] The processor 1200 can be a central processing unit (CPU). The processor 1200 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0101] The memory 1300 can serve as a non-transitory computer-readable storage medium, used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions of a computer program in the embodiments of this disclosure. The processor 1200 implements the data management method by running the non-transitory software programs, instructions, and modules stored in the memory 1300.

[0102] The memory 1300 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function. The data storage area may store data created by the processor 1200, such as a first value H, a second value F, a third value S, a retention index D, etc. Furthermore, the memory 1300 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 1300 may optionally include memory remotely located relative to the processor 1200, and these remote memories may be connected to the processor 1200 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0103] This disclosure also provides a readable storage medium storing a computer program that, when executed by a processor, is used to implement the methods described above. A "readable storage medium" can be any means capable of containing, storing, communicating, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples of a readable storage medium include: an electrical connection with one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable read-only memory (CDROM), etc.

[0104] This disclosure also provides a computer program product, the methods of which can be implemented wholly or partially through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented wholly or partially as a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of this disclosure are performed wholly or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, network equipment, user equipment, core network equipment, OAM, or other programmable device.

[0105] Computer programs or instructions can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The readable storage medium can be any available medium capable of access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a digital video optical disc; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or it can include both volatile and non-volatile types of storage media.

[0106] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0107] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, systems, and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0108] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0109] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0110] In the description of this specification, the references to terms such as "one embodiment / mode," "some embodiments / modes," "example," "specific example," or "some examples," etc., refer to specific features, structures, or characteristics described in connection with that embodiment / mode or example, which are included in at least one embodiment / mode or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment / mode or example. Moreover, the specific features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments / modes or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments / modes or examples described in this specification, as well as the features of different embodiments / modes or examples.

[0111] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0112] Those skilled in the art should understand that the above embodiments are merely for illustrating the present disclosure and are not intended to limit the scope of the disclosure. Those skilled in the art can make other changes or modifications based on the above disclosure, and these changes or modifications still fall within the scope of the present disclosure.

[0113] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0114] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0115] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0116] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0117] At the same time, it is understood that the data involved in this disclosed technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

Claims

1. A data management method characterized by, The method comprises: scanning data records to obtain a plurality of indicators of the data records, a validity time of the data records, and a corresponding business category of the data records, the plurality of indicators being used to represent usage of the data records; determining a first value representing importance of the data records by the plurality of indicators, determining a second value representing memory strength of the data records by the validity time, and determining a third value of the data records by a number of data records contained in the corresponding business category of the data records, the second value being negatively correlated with an accumulated time length since the validity time, and the third value being negatively correlated with the number of data records; determining a retention index of the data records in combination with the first value, the second value, and the third value of the data records; and if the retention index is less than an index threshold, deleting the data records from a storage location where the data records are located. The step of scanning the data records is performed in response to a trigger condition being met, the trigger condition being met when any of the following conditions is met: a start time of a data scanning task is reached, and a data amount of stored data reaches a storage threshold.

2. The data management method according to claim 1, characterized by, The plurality of indicators include one or more of a hotness, a number of user feedbacks, a number of updates, a latest use time, and a validity state.

3. The data management method of claim 1, wherein, The first value is a weighted average of the plurality of indicators.

4. The data management method according to claim 1 or 3, characterized by, The second value is non-linearly attenuated with an increase of the accumulated time length, and an attenuation speed decreases with an increase of the accumulated time length.

5. The data management method of claim 1, wherein, Different business categories differ in at least one of a business scenario, a business label, and a business matter.

6. The data management method of claim 1, wherein, Any one or more of the following conditions is met: different business scenarios correspond to different artificial intelligence services, the business label is used to represent a business feature, and the business matter is used to represent a user intent.

7. The data management method of claim 6, wherein, The retention index is calculated by weighted fusion of the first value, the second value, and the third value.

8. The data management method of claim 1, wherein, The method comprises:

9. An electronic device, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, so that the processor performs the method of any one of claims 1 to 8. The computer program product comprises a computer program, which, when executed by a processor, is used to implement the method of any one of claims 1 to 8.

10. A computer program product, characterised in that, The computer program product comprises a computer program, which, when executed by a processor, is used to implement the method of any one of claims 1 to 8.