Data caching method, system and device and storage medium
By determining the storage popularity of target files in instant messaging tools and prioritizing the retention of important files, the problem of unintelligent cache strategies in the existing technology is solved, efficient and intelligent file management is achieved, and communication efficiency and storage resource utilization are improved.
Patent Information
- Application Number
- CN202511022966.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-07-23
AI Technical Summary
Existing instant messaging tools cannot distinguish the importance of content and business relevance in file cache management, resulting in important files being deleted, while irrelevant file retention cannot meet the needs of refined management of file life cycles.
By obtaining the target files in the target conversation of the instant messaging software, determining the target objects and tasks, measuring the importance of the file based on the storage popularity, and then determining the storage strategy, including storage locations and algorithms, and giving priority to retaining files with high storage popularity.
It realizes intelligent file cache management, improves communication efficiency, reduces storage and bandwidth costs, ensures accessibility of key files, reasonably allocates equipment resources, and reduces the impact of redundant data on equipment performance.
Smart Images

Figure CN120583063A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of data management, and in particular to a data caching method, system, device, and storage medium. Background Art
[0002] Instant messaging (IM) tools have become a crucial platform for work collaboration. Documents, images, videos, audio, and other data are often exchanged and shared through IM. For example, in the construction sector (such as architecture, municipal engineering, and road and bridge construction), workers frequently share and discuss construction drawings through IM.
[0003] Currently, mainstream instant messaging software primarily uses fixed-time expiration policies or fixed-capacity limit policies for media file cache management. For example, received images and files are automatically cleared after being stored for a specific number of days (e.g., 30 days). Alternatively, when the cache reaches a preset size, a simple FIFO (first-in, first-out) or timestamp sorting method is used to purge the oldest content. While simple and easy to implement, this strategy fails to distinguish between content importance and business relevance. Consequently, important files and data may be deleted due to expiration, while less important files and data may be retained due to recent receipt. This fails to meet the need for refined file lifecycle management.
[0004] Therefore, it is necessary to provide a data caching method, system, device and storage medium to accurately and intelligently cache drawings. Summary of the Invention
[0005] One or more embodiments of the present specification provide a data caching method, which includes: obtaining a target file in a target conversation of an instant messaging software; determining a target object of the target conversation; determining a target task based on the target object and the target file; determining the storage heat of the target file based on the target object, the target file and the target task; wherein the storage heat is used to measure the importance of the target file to the target object and the target task; and determining a storage strategy for the target file based on the storage heat.
[0006] One or more embodiments of the present specification also provide a data caching system, which includes: an acquisition module configured to acquire a target file in a target conversation of an instant messaging software; an object determination module configured to determine a target object of the target conversation; a task determination module configured to determine a target task based on the target object and the target file; a storage heat determination module configured to determine the storage heat of the target file based on the target object, the target file and the target task; wherein the storage heat is used to measure the importance of the target file to the target object and the target task; and a storage strategy determination module configured to determine a storage strategy for the target file based on the storage heat.
[0007] One or more embodiments of the present specification also provide a data caching device, which includes at least one processor and at least one memory; the at least one memory is used to store computer instructions; and the at least one processor is used to execute at least part of the computer instructions to implement the data caching method.
[0008] One or more embodiments of this specification further provide a computer-readable storage medium, wherein the storage medium stores computer instructions, and when the computer instructions are executed by a processor, the data caching method is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein: Figure 1 is a schematic diagram of an application scenario of a data cache system according to some embodiments of this specification; Figure 2 is an exemplary flow chart of a data caching method according to some embodiments of this specification; Figure 3 is an exemplary flow chart of determining the storage popularity of a target file according to some embodiments of this specification; Figure 4 is an exemplary flow chart of determining a storage state according to some embodiments of this specification; Figure 5 This is an exemplary module diagram of a data cache system according to some embodiments of this specification. DETAILED DESCRIPTION
[0010] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.
[0011] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0012] As used in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not refer to the singular but also include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0013] Flowcharts are used throughout this specification to illustrate the operations performed by systems according to embodiments of this specification. It should be understood that preceding or following operations do not necessarily need to be performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0014] Instant messaging tools face numerous challenges in data caching. For example, mobile device storage burden: When transferring large files, they quickly consume limited storage space, leading to performance degradation or even storage exhaustion. Bandwidth waste: In situations with limited network conditions (such as construction sites), repeatedly transferring large files consumes valuable bandwidth, prolongs download times, and impacts communication efficiency. Unintelligent caching strategies: Current IM tools often employ simple "cache everything" or "manual management" strategies, lacking intelligent judgment of file importance and timeliness. This results in irrelevant files occupying storage space while critical files may be accidentally deleted. Lack of file lifecycle management: Traditional IM tools fail to consider the relationship between files and task cycles and are unable to identify which files are outdated or no longer relevant, resulting in a large accumulation of expired content. Group sharing is chaotic: When multiple tasks are discussed in different chat groups, the same files may be sent and stored repeatedly, lacking an intelligent cross-group file sharing mechanism. Lack of context: Files shared in IM tools lack clear associations with the corresponding tasks, making it difficult for relevant personnel to quickly understand the file's application scenario and importance.
[0015] Therefore, some embodiments of the present specification provide a data caching method, system, device and storage medium, which can achieve at least the following technical effects by obtaining the target file in the target conversation of the instant messaging software; determining the target object of the target conversation; determining the target task based on the target object and the target file; determining the storage heat of the target file based on the target object, the target file and the target task; and determining the storage strategy of the target file based on the storage heat: solving the storage management problem of files in the instant messaging environment, improving communication efficiency, reducing storage and bandwidth costs, and ensuring the accessibility of key files; ensuring the access speed of high-priority target files, reasonably allocating the storage resources of the target device, and reducing the impact of redundant data on the performance of the target device, thereby achieving efficient and intelligent cache management.
[0016] Figure 1 This is a schematic diagram of an application scenario of a data cache system according to some embodiments of this specification.
[0017] In some embodiments, as Figure 1 As shown, the application scenario 100 of the data cache system includes a target file 110, a target device 120, a network 130, a processor 140, and a storage device 150.
[0018] The target file 110 refers to data that needs to be cached. For example, the target file 110 can be cached in the target device 120. Taking the construction field as an example, the target file 110 can be a construction drawing.
[0019] It should be noted that for ease of description, this specification uses the example of a construction drawing as the target file 110. However, this is merely illustrative and not restrictive. The target file 110 may also be other types of data (e.g., documents, images, videos, audio, etc.) that individuals, organizations, or enterprises need to cache within instant messaging software. The data caching method and system described in this specification are also applicable to other scenarios, such as caching personal files, corporate documents, or university data.
[0020] The target device 120 is a device that needs to store the target file. The target device can be a terminal device used by a user. By way of example only, the target device 120 can be a mobile device, a tablet computer, a laptop computer, a desktop computer, or any other device with input and / or output functions, or any combination thereof. In some embodiments, the target file 110 can be stored and read by the target device 120.
[0021] The network 130 may connect various components in the application scenario 100 of the data cache system and / or connect other components outside the application scenario 100. In some embodiments, one or more components of the application scenario 100 of the data cache system (e.g., the target device 120, the processor 140, and the storage device 150) may be connected and / or communicate with each other via the network 130.
[0022] The processor 140 can process information and / or data related to the data cache system to perform one or more functions described in this specification. In some embodiments, the processor 140 can obtain a target file in a target conversation of an instant messaging software; determine a target object of the target conversation; determine a target task based on the target object and the target file; determine the storage popularity of the target file based on the target object, the target file, and the target task; and determine a storage strategy for the target file based on the storage popularity. Detailed descriptions of the relevant content can be found later (e.g., Figure 2 、 Figure 3 etc.) related descriptions.
[0023] In some embodiments, processor 140 may include a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU), a computer, a user console, or the like, or any combination thereof. In some embodiments, processor 140 may include a single server or a server group. The server group may be centralized or distributed. In some embodiments, processor 140 may be local or remote. In some embodiments, processor 140 may be implemented on a cloud platform. By way of example only, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an on-premises cloud, a multi-layer cloud, or the like, or any combination thereof.
[0024] Storage device 150 can store data, instructions, and / or any other information. In some embodiments, storage device 150 can store target files, etc. In some embodiments, storage device 150 can include mass storage, removable storage, volatile read-write memory, read-only memory (ROM), etc., or any combination thereof. In some embodiments, storage device 150 can be executed on a cloud platform. In some embodiments, processor 140 and storage device 150 can be part of target device 120.
[0025] It is worth noting that the application scenario 100 of the data caching system is provided for illustrative purposes only and is not intended to limit the scope of this specification. For those skilled in the art, various changes and modifications can be made based on the description of this specification. For example, the application scenario 100 of the data caching system can also include a database, an information source, etc. For another example, the application scenario 100 of the data caching system can be implemented on other devices to achieve similar or different functions. However, these changes and modifications will not deviate from the scope of this specification.
[0026] Figure 2 is an exemplary flow chart of a data caching method according to some embodiments of this specification. In some embodiments, process 200 may be executed by a processing device (eg, processor 140). Figure 2 As shown, the process 200 includes the following steps.
[0027] Step 210: Acquire a target file in a target conversation of the instant messaging software.
[0028] Instant messaging software (also known as instant messaging tools) is a digital communication platform that allows users to send and receive text, images, audio, video, and other information / data in real time over the internet or a local area network. Its core features include low-latency interaction, allowing for both individual and group chats, and often integrating features such as file transfer and voice and video calls. It is a crucial tool for both personal social interaction and business communication.
[0029] In some embodiments, the instant messaging software may include internal messaging software within an organization, third-party messaging software, etc. Taking the construction field as an example, the instant messaging software may be messaging software in a construction system used by a user.
[0030] In some embodiments, instant messaging software is installed in the target device.
[0031] The target conversation refers to the conversation that contains the target file. Target conversations include one-on-one conversations and group chat conversations (i.e., multi-party conversations). For more information about target files, see Figure 1 and its related parts.
[0032] In some embodiments, the target conversation can be determined by user selection.For example, the user can select one or more conversations in the instant messaging software as the target conversation.
[0033] In some embodiments, the target conversation may also be determined by the processor 140. For example, the processor 140 may determine a conversation involving file sending and receiving in an instant messaging software as the target conversation.
[0034] In some embodiments, processor 140 may obtain a file in a target conversation and use the obtained file as the target file. In some embodiments, the target file may also be selected by the user. For example, processor 140 may obtain a file selected by the user as the target file. In some embodiments, there may be multiple target conversations, and processor 140 may obtain the target file corresponding to each of the multiple target conversations.
[0035] Step 220: Determine the target object of the target conversation.
[0036] The target object refers to a user related to the target file. For example, the target object can be the recipient of the target file.
[0037] In some embodiments, when the target conversation is a one-to-one conversation, the processor 140 may determine the two chat objects in the one-to-one conversation as target objects.
[0038] In some embodiments, when the target conversation is a multi-party conversation, processor 140 may extract attribute information of the target file and determine the target object. Attribute information refers to various data and information describing the target file. For example, attribute information includes the target file's name, number, type, and sending time.
[0039] For example, if the target file is a construction drawing, processor 140 can extract the drawing's attribute information, including the file name, drawing number, type, size, sending time, associated annotations (text format), and drawing metadata (including title, content, etc.). It can also parse information such as the project number and area identifier that may be included in the file name and metadata. For files whose attributes cannot be directly extracted (for example, PDF and image files), processor 140 can use OCR technology to identify text information in the drawings to improve the accuracy of association analysis. In some embodiments, for video and audio files, processor 140 can determine the target file's attribute information through methods such as image recognition and speech recognition.
[0040] In some embodiments, the processor 140 may determine the target object based on the attribute information of the target file. For example, if the target file is named "XX Project Construction Design Drawing," the processor may determine personnel related to the project construction (e.g., construction personnel, project managers, safety supervisors, etc.) as the target object.
[0041] Step 230: Determine the target task based on the target object and the target file.
[0042] A target task refers to the task that needs to be completed for the target file. For example, if the target file is a construction drawing, the target task can be the construction task that the target object needs to complete as shown in the construction drawing. For example, destroying a wall shown in the construction drawing.
[0043] In some embodiments, the processor 140 can obtain the target object and the target task corresponding to the target file from a preset database. The preset database stores the mapping relationship between the target object and the target file and the target task. Exemplarily, the processor 140 can directly obtain the associated mapping relationship between the drawing, the target object and the construction task from a preset database (such as an enterprise construction task management system), including the drawing number (determined from the file name or drawing content in the drawing's attribute information), the correspondence table between the target object's name and work number and the construction task ID, the matching relationship between the project area to which the drawing belongs and the construction task, etc., and determine the target task based on the mapping relationship, the drawing, and the target object.
[0044] In some embodiments, for a target file not recorded in the preset database, the processor 140 may analyze the context of the target conversation to identify the discussion content related to the target task, thereby associating the target file with the target task. In some embodiments, the processor 140 may also determine a mapping relationship between the newly associated target file and the target task and store it in the preset database.
[0045] In some embodiments, the processor 140 can also use the API of the enterprise construction task management system to obtain ongoing construction task information, such as the construction task leader, construction period, area, related drawings, etc., so as to determine the target task corresponding to the drawing information based on the drawing information and the matching content in the construction task information (for example, the construction task leader, construction period, and area in the drawing content information are the same as the corresponding content in the construction task information).
[0046] In some embodiments, the processor 140 may also associate the target file with the target task based on text similarity and keyword matching, and extract key terms for semantic matching.
[0047] Step 240 : Determine the storage popularity of the target file based on the target object, the target file, and the target task.
[0048] Storage popularity is an indicator used to measure the importance of a target file to the target object and target task. The greater the storage popularity, the more important the target file.
[0049] In some embodiments, storage popularity may be represented by a score or a grade.
[0050] In some embodiments, processor 140 may determine the storage popularity of a target file based on preset rules. For example, if the target file is a construction drawing corresponding to a construction task, the storage popularity is high. For another example, if the target file is a physical examination form for a construction worker involved in the construction task, the storage popularity is low.
[0051] In some embodiments, the processor 140 may also determine the timeliness score of the target file, the task relevance score between the target file and the target task, and the user relevance score between the target object and the target file, and determine the storage popularity of the target file based on one or more of the timeliness score, task relevance score, and user relevance score. For more information on related content, please refer to Figure 3 .
[0052] Step 250: Determine a storage strategy for the target file based on the storage popularity.
[0053] The storage strategy refers to a method for storing a target file. In some embodiments, the storage strategy includes at least a storage location and a storage algorithm.
[0054] Storage locations include local storage, cloud storage, and edge storage.
[0055] Storage algorithms include long-term cache strategy - LRU-2 cache algorithm, conditional cache strategy - SLRU cache algorithm, temporary cache strategy, etc.
[0056] In some embodiments, there is a preset correspondence between storage heat and storage strategy, and the processor 140 can determine the storage strategy based on the preset correspondence. For example, when the storage heat is high, the corresponding storage strategy can be local storage (for easy reading) using a long-term cache strategy - LRU-2 cache algorithm.
[0057] In some embodiments, the processor 140 may determine the storage priority of the target file based on the storage heat and the heat threshold; and determine the storage algorithm of the target file based on the storage priority.
[0058] The heat threshold is used to determine the storage heat level of the target file. In some embodiments, the heat threshold includes multiple heat thresholds. For example, the heat threshold includes a high heat threshold and a low heat threshold.
[0059] In some embodiments, the heat threshold is dynamically adjusted based on the storage heat of all cached files on the target device and / or the free storage space on the target device. The free storage space on the target device refers to the remaining storage space on the target device. The free storage space can be expressed in absolute terms (e.g., 20 GB of free storage space remaining) or in relative terms (e.g., the percentage of free storage space to total storage space).
[0060] In some embodiments, the processor 140 can adjust the heat threshold according to the storage heat corresponding to all cached files in the target device. For example, the high / low heat threshold can be determined directly based on the average heat and standard deviation of all cached files in the target device. Exemplarily, the processor 140 can use the average heat plus a certain multiple of the standard deviation as the high heat threshold. For example, the average heat is M and the standard deviation is S, and the high heat threshold can be set to M+2S. The processor 140 can also use the average heat minus a certain multiple of the standard deviation to determine the low heat threshold. For example, the low heat threshold is set to M-2S.
[0061] For another example, the processor 140 may divide all cached files in the target device into high / medium / low priority files in a ratio of 2:5:3, and divide the heat thresholds based on the aforementioned ratios and the storage heat of all cached files in the target device. For example, the average storage heat of high-priority files may be used as the high heat threshold, and the average storage heat of low-priority files may be used as the low heat threshold.
[0062] In some embodiments, the processor 140 may determine the heat threshold based on the free storage space. For example, the less free storage space there is, the higher the high / low heat threshold.
[0063] In some embodiments, the processor 140 can determine a heat threshold based on the storage heat corresponding to all cached files in the target device and the free storage space of the target device. For example, when there is sufficient free storage space (e.g., free storage space is greater than 30%), the original high / low heat thresholds are maintained to ensure that the long-term caching strategy for high-priority drawings is fully implemented; when a storage warning occurs (e.g., free storage space is 10% to 30%), the low heat threshold is increased, some medium-priority drawings are downgraded to low priority, and some low-priority drawings are temporarily not cached, but the high heat threshold is maintained unchanged to ensure stable storage of high-priority drawings; when storage is urgent (e.g., free storage space is less than 10%), all thresholds are adjusted, the high heat threshold is also increased, and only the most critical drawings are retained locally, while triggering an emergency cleanup mechanism (e.g., cleaning up some low-storage heat files that have not been accessed for a long time).
[0064] In some embodiments of the present specification, the heat threshold is adaptively adjusted according to the current storage space of the target device and the stored cache files, so that the heat threshold is more consistent with the current situation of the target device.
[0065] Storage priority refers to a parameter used to evaluate the priority of target file storage.
[0066] In some embodiments, the processor 140 may determine the storage priority based on a heat threshold. For example, when the storage heat of the target file is greater than a high heat threshold, the storage priority of the target file is high; when the storage heat of the target file is greater than a low heat threshold but less than a high heat threshold, the storage priority of the target file is medium; and when the storage heat of the target file is less than a high heat threshold, the storage priority of the target file is low.
[0067] In some embodiments, when the storage priority of the target file is high, the processor 140 can set a long-term cache strategy for the target file and apply the LRU-2 cache algorithm to ensure a high hit rate and fast response for high-priority target files; when the storage priority of the target file is medium, the processor 140 can set a conditional cache strategy for the target file and adopt the SLRU cache algorithm to divide the cache into hot areas and cold areas, and adjust the migration of the target file between the two intervals according to the access frequency, taking into account cache efficiency and memory occupancy; when the storage priority of the target file is low, the processor 140 can set a temporary cache strategy for the target file and use a combined cache elimination mechanism based on survival time and access frequency to avoid cache pollution.
[0068] In some embodiments of the present specification, target files with different storage hotness correspond to different storage algorithms, which ensures the access speed of high-priority target files, reasonably allocates the storage resources of the target device, and reduces the impact of redundant data on the performance of the target device, thereby realizing efficient and intelligent cache management.
[0069] In some embodiments, the storage policy also includes storage status. The processor 140 may also obtain access information of the target file and task information of the target task; determine the current status score of the target file based on the storage popularity, access information, and task information; and determine the storage status of the target file based on the current status score. For more information on related content, please refer to Figure 4 .
[0070] The data caching method described in some embodiments of this specification solves the storage management problem of target files in an instant messaging environment, improves communication efficiency, reduces storage and bandwidth costs, and ensures the accessibility of important files.
[0071] Figure 3is an exemplary flow chart of determining the storage popularity of a target file according to some embodiments of this specification. In some embodiments, process 300 may be executed by a processing device (eg, processor 140). Figure 3 As shown, the process 300 includes the following steps.
[0072] Step 310: Obtain chat information of the target object in the instant messaging software.
[0073] Chat information refers to conversation data related to the target file between the target party and the target file in the instant messaging software. For example, conversation data mentioning the target file, conversation data related to the content of the target file, etc.
[0074] In some embodiments, the chat information may be conversation data related to the target file in the target conversation, or may be conversation data related to the target file in other conversations.
[0075] In some embodiments, with permission from the target object, the processor 140 may identify chat records in each conversation in the instant messaging software and obtain chat information of the target object in the instant messaging software.
[0076] Step 320: Obtain user information of the target object.
[0077] User information refers to information / data related to the target object. For example, user information may include the target object's name, position, employee number, etc.
[0078] In some embodiments, the instant messaging software may store user information, and the processor 140 may obtain the user information of the target object through the instant messaging software. In some embodiments, the processor 140 may also obtain the user information of the target object through a preset database.
[0079] Step 330: Obtain access information and file information of the target file.
[0080] The access information refers to information related to the operation of the target object accessing the target file. In some embodiments, the access information includes access frequency, last access time, etc.
[0081] File information refers to data / information that characterizes the content of a target file. For example, file information includes the name, number, and specific content of the target file.
[0082] In some embodiments, the processor 140 may record each operation of the target object accessing the target file to obtain access information.
[0083] In some embodiments, the file information can be obtained by parsing the target file. For example, the processor 140 can parse the target file through methods such as metadata extraction, content recognition (image / text / audio recognition), structured data parsing, and file feature analysis to obtain the file information.
[0084] Step 340: Obtain task information of the target task.
[0085] Task information refers to data / information related to the target task. For example, task information includes task content, task progress, etc.
[0086] In some embodiments, task information can be determined based on chat information and the target file. For example, processor 140 can determine the progress of the target task (including not yet started, in progress, completed, etc.) based on the target object's conversation in the instant messaging software. For another example, processor 140 can determine the content of the target task by parsing the target file.
[0087] Step 350: Determine the timeliness score of the target file based on the chat information and the access information.
[0088] The timeliness score represents the validity and time sensitivity of the target file. The more frequently the target file is mentioned in chat messages, the more frequently it is accessed, and the more recently it was last accessed, the higher the validity and time sensitivity of the target file, and the higher the corresponding timeliness score.
[0089] In some embodiments, the processor 140 may calculate the timeliness score based on the chat information and the access information using a time decay model. The time decay model may be a preset comparison table, mapping rules, vector database, etc. The time decay model includes a correspondence between the chat information and the access information and the timeliness score.
[0090] Step 360 : Determine a task relevance score between the target file and the target task based on the file information and the task information.
[0091] The task relevance score is a score that represents the degree of relevance between the target file and the target task in terms of content.
[0092] In some embodiments, the processor 140 may compare the semantic similarity between the target file and the target task description through a text vector space to determine a task relevance score between the target file and the target task.
[0093] Step 370: Determine a user relevance score between the target object and the target file based on the user information and the task information.
[0094] The user relevance score represents the degree of association between the target file and the target object. The degree of association between a target file and different target objects varies. For example, if the target file is a construction drawing, different target objects such as project managers, team leaders, and construction personnel have different requirements for construction drawings, resulting in different degrees of association.
[0095] In some embodiments, the processor 140 can determine a user relevance score based on the degree of match between the user information, task information, and the target file. For example, the processor 140 can identify the target subject's role (e.g., project manager, team leader) and department / project based on the user information, and determine whether their pre-set responsibilities are directly related to the target file (e.g., construction drawings) and task information. The processor 140 can also analyze the target subject's recent operational behavior, counting the number of recent (e.g., past 30 days) operations on the target file (e.g., viewing, modifying, commenting on, and associating the file with a task). Users with highly relevant responsibilities (e.g., project managers) are assigned a higher base score, and additional points are awarded based on the target subject's recent frequency of operations (the more operations, the more points are added) to determine the user relevance score.
[0096] Step 380 : Determine the storage popularity of the target file based on one or more of the timeliness score, the task relevance score, and the user relevance score.
[0097] In some embodiments, the processor 140 may select any one of a timeliness score, a task relevance score, and a user relevance score as the storage popularity of the target file.
[0098] In some embodiments, the processor 140 may perform a weighted summation of the timeliness score, the task relevance score, and the user relevance score, and determine the result of the weighted summation as the storage popularity of the target file.
[0099] In some embodiments, the processor 140 can also determine the first storage heat of the target file based on the target object, target file and target task; obtain the associated tasks of the target task; determine the second storage heat of the target file based on the associated tasks; and determine the storage heat based on the first storage heat and the second storage heat.
[0100] The first storage heat is an indicator that measures the storage importance of the target file to the target device, target object, and target task.
[0101] In some embodiments, the processor 140 may determine the first storage heat using a first model, wherein the first model is a machine learning model. For example, the first model may be a neural network model, a deep neural network model, etc.
[0102] In some embodiments, the processor 140 may input access information and file information of the target file, free storage space of the target device, user information of the target object, and task information of the target task into the first model, and the first model outputs a first storage heat of the target file.
[0103] The first model can be trained using a first sample and a first label. The first sample includes historical target file access information and file information, historical target device free storage space, historical target object user information, and historical target task information. The first label includes the historical first storage popularity. The first label and first sample can be collected and annotated based on the historical data caching process.
[0104] Associated tasks refer to tasks associated with the target task. Associated tasks can include tasks related to the target task's task area. For example, if the target task requires installing water pipes within a certain area, an associated task could be installing electrical wiring within the same area. In some embodiments, associated tasks can also include tasks related to the target task's task content. For example, if the target task requires pouring a floor slab, an associated task could be pre-buried water and electricity pipelines.
[0105] The second storage heat is an indicator that measures the storage importance of the target file to the associated task.
[0106] In some embodiments, the processor 140 may determine the second storage heat using a second model, wherein the second model is a machine learning model. For example, the second model may be a neural network model, a deep neural network model, etc.
[0107] In some embodiments, the processor 140 may input file information of the target file, task information of the target task, and task information of the associated tasks into the second model, and the second model outputs a second storage heat of the target file.
[0108] The second model can be trained using the second sample and the second label. The second sample includes the file information of the historical target file, the task information of the historical target task, and the task information of the historical associated task. The second label includes the historical second storage heat. The second label and the second sample can be collected and annotated based on the historical data caching process.
[0109] In some embodiments, the processor 140 may determine the storage heat based on the first storage heat, the first weight, the second storage heat, and the second weight. For example, the storage heat may be determined based on a weighted sum of the first storage heat, the first weight, the second storage heat, and the second weight.
[0110] The first weight is the weight of the first storage heat, and the second weight is the weight of the second storage heat. The first weight and the second weight are related to the task progress of the target task. In some embodiments, based on the task progress of the target task, the task completion degree of the target task can be determined. The lower the task completion degree of the target task, the greater the first weight and the smaller the second weight. That is, if the target task has not been completed yet, attention should be paid to the storage importance of the target file to the target device, target user and the target task itself. If the target task has been completed, more attention should be paid to the storage importance of the target file to other related tasks, so that the storage heat can be determined more accurately.
[0111] In some embodiments, the second weight is also related to the task progress of the associated task. In some embodiments, based on the task progress of the associated task, the task completion degree of the associated task can be determined, and the lower the task completion degree of the associated task, the greater the second weight.
[0112] In some embodiments of the present specification, each task may not exist independently, and there may be associations between different tasks. By determining the first storage heat of the target task and the second storage heat of the associated task, and further determining the storage heat of the target file, the coordinated storage of target files corresponding to different tasks can be achieved, thereby obtaining better data storage and management effects.
[0113] Figure 4 is an exemplary flow chart of determining a storage state according to some embodiments of this specification. In some embodiments, process 400 may be executed by a processing device (eg, processor 140). Figure 4 As shown, process 400 includes the following steps.
[0114] Step 410: Obtain access information of the target file and task information of the target task.
[0115] The processor 140 may obtain the access information of the target file and the task information of the target task based on the same methods as steps 330 and 340 , respectively.
[0116] In some embodiments, the processor 140 may obtain chat information of the target object in the instant messaging software; analyze the chat information to determine associated chat information related to the target task; and determine task information of the target task based on the associated chat information.
[0117] Related chat information refers to conversation data / information associated with the target task. In some embodiments, processor 140 may analyze the chat information to determine related chat information related to the target task. For example, processor 140 may determine chat information that mentions the target task and / or target file as related chat information.
[0118] In some embodiments, processor 140 may determine the task information of the target task based on the associated chat information. For example, the target party may have mentioned in a chat message (e.g., a chat message in the target conversation or a chat message in another conversation) that the target task has been completed, thereby adjusting the task progress information in the task information. For another example, the target party may have mentioned in a chat message that the target task has changed, thereby adjusting the task content in the task information.
[0119] In some embodiments, the target task includes an associated file. When the associated file meets a preset adjustment condition, the processor 140 may adjust the task information of the target task.
[0120] Associated files are files related to the target task. For example, when the target task is a construction task, associated files may include bidding documents, rectification documents, acceptance documents, etc.
[0121] The preset adjustment condition includes whether the associated file is related to the adjustment of the target task. For example, the associated file is a rectification file, an acceptance file, etc. of the target task. In some embodiments, the processor 140 can parse the associated file to determine whether the associated file is related to the adjustment of the target task. The parsing method of the associated file is similar to the parsing method of the target file. For details, please refer to the relevant part above, such as Figure 2 wait.
[0122] In some embodiments, processor 140 can adjust the task information of the target task based on the associated file. For example, the task information of the target task can be adjusted based on the rectification information in the rectification file (e.g., adjusting the position of a certain line in a routing task). For another example, if the acceptance file contains unacceptable content, the corresponding task in the target task can be adjusted.
[0123] Step 420 : Determine the current status score of the target file based on the storage popularity, access information, and task information.
[0124] The current status score is an indicator that evaluates the activity level of the target file. The higher the current status score, the more active the target file is.
[0125] In some embodiments, the processor 140 may determine the current status score based on the storage heat, the timeliness score of the target file, the access frequency score, and the task progress score. For example, the current status score of the target file It can be determined based on the following formula (1): (1) in, Indicates storage heat; Indicates timeliness rating; Indicates the access frequency score, which can be determined based on the access frequency in the access information. For example, the access frequency can be standardized to serve as the access frequency score; Indicates the task progress score, which can be determined based on the task progress in the task information. For example, the percentage of task completion can be normalized and used as the task progress score; It is a dynamically adjusted weight coefficient that can adapt to the current status of the target file and user needs.
[0126] Step 430: Determine the storage status of the target file based on the current status score.
[0127] The storage state refers to the activity state of the target file when it is stored. For example, in descending order of activity, the storage state of the target file includes active state, reference state, archive state, and pending cleanup state.
[0128] In some embodiments, the processor 140 may determine the storage state of the target file directly based on the current state score. For example, a mapping relationship between the current state score and the storage state of the target file may be preset, and the processor 140 may determine the storage state of the target file based on the current state score and the mapping relationship.
[0129] The storage state of a target file often changes over time. For example, the active state may become the reference state, and the reference state may become the active state. Exemplarily, when the current state score of the target file exceeds the active threshold or is marked as a critical file, the processor 140 determines that the storage state of the target file is the active state. When the current state score of the target file drops below the reference threshold, the processor 140 determines that the storage state of the target file is the archive state. When the target file is referenced again or the access frequency suddenly increases, and the state score exceeds the reference threshold, the processor 140 determines that the storage state of the target file is converted from the archive state to the reference state. When the target file state score continues to be below the cleanup threshold and the related project has been completed for more than a preset period of time (e.g., 3 months), the processor 140 determines that the storage state of the target file is converted from the archive state to the reference state. When the target file is accidentally accessed or manually restored, and the storage state score again exceeds the cleanup threshold, the processor 140 determines that the storage state of the target file is converted from the pending cleanup state to the archive state.
[0130] In some embodiments, the processor 140 may further obtain a historical status score of the target file; and determine the storage status of the target file based on the historical status score, the current status score, and a status score threshold.
[0131] The historical status score is an indicator for evaluating the historical activity of the target file. In some embodiments, the processor 140 may determine the historical status score using a method similar to that used to determine the current status score. For example, the processor 140 may determine the historical status score of the target file based on the historical storage popularity, the historical timeliness score of the target file, the historical access frequency score, and the historical task progress score.
[0132] The status score threshold is used to determine the storage status of the target file. In some embodiments, the status score threshold may include multiple thresholds. For example, the status score threshold may include an active threshold, a reference threshold, a cleanup threshold, and the like. When the status score of the target file (such as the current status score, the historical status score) is greater than the active threshold, the storage status of the target file is active; when the status score of the target file (such as the current status score, the historical status score) is greater than the reference threshold and less than the active threshold, the storage status of the target file is reference; when the status score of the target file (such as the current status score, the historical status score) is greater than the cleanup threshold and less than the reference threshold, the storage status of the target file is archived; when the status score of the target file (such as the current status score, the historical status score) is less than the cleanup threshold, the storage status of the target file is to be cleaned.
[0133] In some embodiments, the status score threshold may be set based on experience.
[0134] In some embodiments, the state score threshold may also be adjusted based on a reinforcement learning algorithm.
[0135] For example, the processor 140 may design a reward function to balance storage efficiency and access performance, for example, the reward function R = α × storage saving + β × access speed − γ × state switching frequency.
[0136] Storage savings measures how much high-speed storage space (such as local storage) is freed up by a policy. For example, moving infrequently accessed data out of the cache saves cache space. Metrics can be expressed as the number of bytes saved or as a percentage of the total cache space.
[0137] The access speed refers to the speed of accessing the stored target file. For example, the access speed can be determined based on the number of successful accesses per unit time or the average access delay.
[0138] The state switching frequency refers to the switching frequency of the storage state of the target file. Exemplarily, the state switching frequency can be determined based on the number of times the storage state is switched within a preset time period.
[0139] α, β, and γ represent the weights for storage savings, access speed, and state switching frequency, respectively. These weights can be adjusted dynamically based on specific user needs and application scenarios. For example, when there is less free storage space, the value of α can be appropriately increased; when higher response speed is required, the value of β can be increased.
[0140] In some embodiments, the processor 140 can continuously optimize the conversion strategy through the Q-learning algorithm and dynamically adjust the state scoring threshold. In some embodiments, the processor 140 can also regularly evaluate system performance indicators, including cache hit rate, storage utilization, and access latency, as feedback for reinforcement learning.
[0141] For example, the state S of the Q-learning algorithm t Indicates the current operating status of the system, which may include: cache status, such as cache hit rate (i.e., the proportion of target files found in the cache when obtaining the target file, which can be determined based on the number of hits / total number of requests), load distribution (i.e., the access frequency distribution of each storage unit in the cache (such as cache lines, partitions)), hot data pattern (i.e., the data characteristics of active target files, including temporal locality, spatial locality, data type, etc.); storage status, such as storage space utilization, tiered storage space allocation, data access pattern, etc.; performance indicators, such as average access latency, throughput, state switching history, etc.; current policy parameters, such as the currently used state score threshold itself or its historical change trend.
[0142] Action A t It represents the optimization operations that the processor 140 can perform, including the adjustment of the status score threshold.
[0143] Processor 140 performs action A t , enter the new state S t+1 Afterwards, the reward function R can be updated to obtain a new reward function R1 to quantify the action A t In state S t The immediate benefits (including positive benefits and negative benefits / penalties) brought by the new state S. t+1 The new reward function R1 is calculated based on the storage savings, access speed and state switching frequency.
[0144] In some embodiments, whenever the processor 140 is in state S t Execute action A t , receives reward R1 and moves to new state S t+1 After that, the Q value can also be updated. For example, based on the following formula (2): (2) Where η is the learning rate, which controls the extent to which new information overwrites old information (0 < η ≤ 1). Smaller η results in slower but more stable learning, while larger η results in faster learning but potentially oscillatory. Represents the new state S t+1 The maximum Q value among all possible actions under St +1 An estimate of the best future cumulative reward that can be obtained; the λ discount factor is used to measure the importance of future rewards relative to immediate rewards (0≤λ≤1), where λ=0 means that the system only cares about immediate rewards, and λ=1 means that the system regards all future rewards as equally important as immediate rewards.
[0145] In some embodiments of this specification, the storage efficiency and access performance of the target file can be balanced by adjusting the status score threshold through a reinforcement learning algorithm. In addition, indicators such as cache hit rate, storage utilization, and access latency are regularly evaluated as feedback for reinforcement learning to improve the accuracy of storage status judgment.
[0146] In some embodiments, the processor 140 may determine or update the storage state based on the historical state score. For example, if the state score (e.g., the previous two historical state scores and the current state score) is between the reference threshold and the active threshold for three consecutive times, the storage state of the target file is determined to be the reference state.
[0147] In some embodiments, the processor 140 may also determine whether the storage state of the target file needs to be adjusted based on the historical storage state and the state score of the target file. For example, when the number of times the state score of the target file (including the historical state score and the current state score) does not match the historical storage state exceeds a matching threshold, the storage state of the target file is adjusted. Exemplarily, the historical storage state of the target file is an archive state. When the state score (for example, the previous two historical state scores and the current state score) is higher than the active threshold for three consecutive times, it is determined that the storage state of the target file needs to be upgraded, that is, changed from the archive state to the reference state.
[0148] In some embodiments of the present specification, combined with the historical storage status and status score of the target file, when the number of times the status score of the target file does not match the historical storage status is higher than the matching threshold, the storage status of the target file is adjusted to ensure that the storage status can only change step by step, ensuring orderly management of the file life cycle, reducing the risk of data loss and maintaining the compliance of the storage system.
[0149] In some embodiments, the storage status indicates whether the target file needs to continue to be stored and the storage location of the target file.
[0150] For example, the active state indicates that the target file is currently being used frequently or is highly relevant to the task being performed. A target file in the active state can be stored on the local device and edge nodes to maintain high-speed access there, supporting real-time rendering and interaction. The reference state indicates that the target file is currently used less frequently but still has reference value. This may be due to a completed phased task that still requires occasional reference, or an associated task in preparation. A target file in the reference state can retain an abbreviated or low-resolution version locally, while the full version remains accessible on the edge node. The archived state indicates that the target task associated with the target file has been substantially completed but needs to be retained for recordkeeping and possible subsequent review. A target file in the archived state is removed from the local device and retained only in cloud storage, but its metadata index is fully maintained to ensure rapid recovery when needed. The pending cleanup state indicates that the target file has become outdated, the associated project or task has concluded, and it has not been accessed for a long time. A target file in the pending cleanup state is ready to be cleaned from the target device. However, before complete deletion, the processor 140 creates a lightweight metadata backup and index, and may trigger user confirmation under certain conditions.
[0151] In some embodiments of the present specification, the activity level of the target file is evaluated, and the storage status is dynamically adjusted according to the activity level of the target file, thereby reducing the storage overhead of inactive data, and adjusting the storage location according to the storage status of the target file to achieve an optimal balance between performance, cost and storage resource occupancy, thereby reducing the complexity of manual maintenance.
[0152] In some embodiments, the processor 140 can also manage target files. For example, it can clean target files and adjust their storage locations. For example, for active target files, by default, they are not cleaned, but intelligent compression and selective caching are implemented. For large active target files, lossless or lossless compression techniques are applied to reduce storage usage. For extremely large target files (e.g., larger than 50MB), only commonly used areas and low-resolution full images can be cached. A preload boundary is established to dynamically load the full content only when needed. For reference target files, partial cleanup and quality degradation strategies are implemented: the original high-definition target file is removed from the local device, retaining only a thumbnail or optimized version. The original target file is removed from the local device, retaining only a thumbnail or optimized version. A LRU cache eviction queue is established to prioritize the removal of reference target files when space is insufficient. For archived target files, they are completely removed from local devices and edge nodes, and a hot-cold split storage approach is adopted: the target files are transferred to cloud storage, using a cold storage architecture to reduce storage costs. Only lightweight metadata and thumbnail indexes are retained locally to support content previews. An on-demand lazy loading mechanism is implemented, with content retrieved from the cloud upon user request. For target files to be cleaned, a multi-stage cleaning process is implemented. The first stage: completely remove all content on local and edge nodes, retaining only cloud copies; the second stage: downgrade cloud storage from hot storage to archive storage to further reduce costs; the third stage: perform final cleaning, but retain complete metadata records and minimize thumbnails.
[0153] In some embodiments, the processor 140 can also protect the target file to prevent the target file from being mistakenly cleared, improve the accuracy of the target file storage status, and obtain better target file storage effects. For example, for active target files, no special protection is required, and the system ensures the highest access performance and data integrity. For reference target files, notification and recovery opportunities are provided before downgrading (i.e., lowering the storage status level, for example, from reference state to archive state): 7 days before the planned downgrade, a notification is sent to users who have recently accessed the target file; a one-click upgrade to active state operation option is provided to prevent important target files from being mistakenly downgraded; downgrade history and decision basis are recorded to support subsequent analysis and improvement. For archive target files, preventive protection is implemented before transferring to cold storage: a complete index record is created, including all related messages and discussion content; structured metadata is generated to ensure that even if archived, it can be quickly located through search; and mandatory retention marks are set for specific types of critical target files (such as those related to project security). For target files to be cleaned, a strict multi-layer protection mechanism is implemented: before performing the final cleanup, the system automatically generates a risk level assessment report; high-risk target files (such as those with high historical importance or possible legal requirements) require manual confirmation by the administrator; digital fingerprints and complete metadata backups are created to ensure that they can be retrieved based on records in the future; cleanup logs and operation audit trails are maintained to support responsibility tracing and recovery after accidental cleanup.
[0154] In some embodiments, processor 140 can also associate chat messages from instant messaging software with target files. In some embodiments, processor 140 can construct a bidirectional index between chat messages and target files, including: when a target file is shared in an IM, creating a unique reference identifier and applying a hash algorithm to ensure unique identification of the target file content; establishing a bidirectional link relationship between the target file and related chat messages, constructing a target file-chat message directed acyclic graph data structure; maintaining reference records of the target file in different sessions to achieve efficient and fast search; and using a sparse matrix to represent the target file-chat message association relationship to optimize storage efficiency.
[0155] In some embodiments, the processor 140 can also implement context-based target file retrieval, including: analyzing message content to identify implicit and explicit references to target files; developing a conversation content search function to support searching for relevant target files by keywords; and implementing reverse query of target files to quickly locate relevant discussions from the target files.
[0156] In some embodiments, the processor 140 can also synchronize the update status of the target file, including: automatically pushing notifications in related sessions when a new version of the target file is updated; displaying the version and update status of the target file in the message; and providing a new and old version comparison function to facilitate users to understand the changes.
[0157] In some embodiments, the processor 140 can also optimize target file sharing across multiple conversations within the instant messaging software. In some embodiments, the processor 140 can also implement a cross-conversation file sharing mechanism, including: identifying the same target file shared across multiple conversations; establishing a global target file resource pool to avoid duplicate storage; and providing unified access rights management for different conversations.
[0158] In some embodiments, the processor 140 can also optimize the display of target files within the conversation, including: automatically classifying and organizing target files according to the task topic of the conversation; providing a target file timeline view to display related target files in the order of discussion; and highlighting the latest versions and important changes of target files.
[0159] In some embodiments, the processor 140 can also implement role-based target file access control, including: dynamically adjusting target file access permissions based on the user's role in the project; supporting fine-grained control of permissions for specific parts of target files; maintaining access logs to record target file viewing and operation history.
[0160] In some embodiments, the processor 140 can also implement a multi-level cache architecture, including: designing a three-level architecture of target device local cache, edge node cache and cloud storage, and using a consistent hash distribution algorithm to ensure the scalability and load balancing of cache nodes; deciding which level of cache to store in based on the target file association and access pattern, and applying a dynamic programming algorithm to optimize cache allocation; achieving data synchronization and consistency maintenance between caches at all levels, and using conflict-free replication data type technology to ensure eventual consistency.
[0161] In some embodiments, the processor 140 can also optimize the local cache strategy, including: monitoring the target device storage space and usage; dynamically adjusting the local cache content according to the target file priority and device status; and implementing cache preheating and prefetching to load target files that may be needed in advance.
[0162] In some embodiments, the processor 140 may also manage edge node caches, including: deploying edge cache servers at the target task execution site; optimizing cache content of edge nodes based on regional target tasks; and enabling fast data exchange between local devices and edge nodes.
[0163] In some embodiments, the processor 140 can also implement intelligent transmission and rendering of target files. For example, the processor 140 can implement an adaptive file transfer strategy, including: detecting the user's network environment and device performance; automatically adjusting the transmission quality and speed according to the network bandwidth; supporting breakpoint resumption and background transmission to optimize transmission efficiency. In some embodiments, for special types of target files, such as target files such as pictures and drawings, the processor 140 can also implement on-demand rendering and loading: support block loading and progressive rendering of target files; prioritize loading content within the viewport range based on the user's viewing area; implement multi-resolution rendering and smoothly switch between overview and detail modes. In some embodiments, the processor 140 can also optimize the offline access experience, including: based on target task planning and predictive analysis, caching target files that may need to be accessed offline in advance; supporting offline marking and annotation, and synchronizing changes after the network is restored; supporting offline marking and annotation, and synchronizing changes after the network is restored.
[0164] In some embodiments, processor 140 can also perform user behavior analysis. For example, processor 140 can collect user interaction data, including: recording the frequency, duration, and patterns of users' access to target files, implementing event stream sampling and compressed storage; tracking users' mentions of specific target files in IM discussions, applying natural language processing techniques to identify implicit references; analyzing users' actions on target files (such as zooming and annotating) to construct a feature vector for the interaction behavior sequence; implementing a privacy-preserving data collection mechanism, and using differential privacy techniques to ensure user data security. Processor 140 can also build user behavior models, including: establishing a model of users' target file usage habits based on historical data; identifying typical behavior patterns of users with different user roles; and discovering correlations between target file usage and target task progress. Processor 140 can also generate personalized predictions, including: predicting target files that users may need to access based on the behavior model; calculating a confidence score for the prediction; and adjusting caching and recommendation strategies based on the prediction results.
[0165] In some embodiments, the processor 140 can also sense the progress of the target task. Taking the construction task as an example, the processor 140 can obtain construction progress data, including: connecting to the project management system to obtain construction task progress information in real time; analyzing the construction status reflected in IM communication; collecting on-site feedback and progress reports. The processor 140 can also evaluate the changes in the relevance of target files, including: evaluating the current relevance of different target files based on the progress of the target task; predicting the changing trend of the relevance of target files as the target task progresses; identifying the changes in the target file requirements at key time points. The processor 140 can also dynamically adjust the life cycle of the target file, including: automatically adjusting the life cycle status of the target file based on the task progress of the target task; identifying the target files that will be needed soon and preparing the cache in advance; identifying the target files that may no longer be needed and arranging for gradual cleanup.
[0166] In some embodiments, the processor 140 may further determine an adaptive lifetime of the target file, and when the adaptive lifetime of the target file expires, the target file is cleaned up. The adaptive lifetime is the longest lifetime of the target file calculated by the system. For example, the adaptive lifetime of the target file may be determined based on the following formula (3): (3) TTL stands for adaptive lifetime; T represents the base lifetime of the target file, which can be preset based on the storage state of the target file, for example, the active state is the longest, the reference state is the second longest, and the archive state is the shortest; ω represents the adjustment factor, which can be determined based on the current state score of the target file. For example, the higher the current state score, the larger the adjustment factor, to ensure that highly active target files have a longer lifetime. It represents the task progress factor. The greater the completion degree of the target task, the greater the task progress factor. It means that as the target task approaches completion, the survival time of the target file gradually decreases. G represents the usage frequency coefficient, which can be determined based on the number of accesses to the target file. The more the number of accesses, the higher the usage frequency coefficient, which avoids premature failure of the hotspot drawing.
[0167] It should be noted that the above descriptions of processes 200, 300, and 400 are for illustration and purpose only and do not limit the scope of this specification. Those skilled in the art may, under the guidance of this specification, make various modifications and alterations to processes 200, 300, and 400. However, such modifications and alterations remain within the scope of this specification.
[0168] Figure 5 is an exemplary module diagram of a data cache system according to some embodiments of this specification. Figure 5 As shown, the data cache system 500 includes an acquisition module 510 , an object determination module 520 , a task determination module 530 , a storage heat determination module 540 and a storage strategy determination module 550 .
[0169] The acquisition module 510 is configured to acquire a target file in a target conversation of an instant messaging software installed in a target device.
[0170] The object determination module 520 is configured to determine a target object of the target conversation.
[0171] The task determination module 530 is configured to determine a target task based on the target object and the target file.
[0172] The storage heat determination module 540 is configured to determine the storage heat of the target file based on the target object, the target file and the target task; wherein the storage heat is used to measure the importance of the target file to the target object and the target task.
[0173] The storage policy determination module 550 is configured to determine a storage policy for the target file based on the storage heat.
[0174] It should be noted that the above description of the data cache system 500 and its modules is for convenience only and does not limit this specification to the scope of the embodiments. Figure 5The acquisition module 510, object determination module 520, task determination module 530, storage heat determination module 540, and storage policy determination module 550 disclosed in the specification may be different modules in a system, or a single module may implement the functions of two or more of the aforementioned modules. For example, each module may share a storage module, or each module may have its own storage module. Such variations are within the scope of protection of this specification.
[0175] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only an example and does not constitute a limitation of this specification. At the same time, this specification uses specific words to describe the embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" refer to a certain feature, structure or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more in different places in this specification does not necessarily refer to the same embodiment.
[0176] It should be noted that, in order to simplify the description of this specification and facilitate understanding of one or more embodiments, the foregoing description of the embodiments of this specification sometimes combines multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not mean that the subject matter of this specification requires more features than those recited in the claims. In fact, an embodiment may have fewer features than all of the features of a single embodiment disclosed above.
Claims
1. A data caching method, characterized in that: The method comprises: Obtain the target file in the target conversation of the instant messaging software; determining a target object of the target conversation; Determining a target task based on the target object and the target file; Determining the storage popularity of the target file based on the target object, the target file, and the target task; wherein the storage popularity is used to measure the importance of the target file to the target object and the target task; Based on the storage popularity, a storage strategy for the target file is determined.
2. The method according to claim 1, characterized in that The storage strategy includes a storage algorithm, and determining the storage strategy of the target file based on the storage heat includes: Determining the storage priority of the target file based on the storage heat and the heat threshold; Based on the storage priority, a storage algorithm for the target file is determined.
3. The method according to claim 2, characterized in that The instant messaging software is installed in the target device, and the heat threshold is dynamically adjusted according to the storage heat corresponding to all cache files in the target device and / or the free storage space of the target device.
4. The method according to claim 1, wherein The storage policy includes a storage state, and the method further includes: Acquire access information of the target file and task information of the target task; Determining a current status score of the target file based on the storage popularity, the access information, and the task information; Based on the current status score, the storage status of the target file is determined; wherein the current status score is used to evaluate the activity level of the target file, and the storage status indicates whether the target file needs to continue to be stored and the storage location of the target file.
5. The method according to claim 4, characterized in that Determining the storage status of the target file based on the current status score includes: Obtaining a historical status score of the target file; The storage status of the target file is determined based on the historical status score, the current status score, and a status score threshold.
6. The method according to claim 4, characterized in that The state scoring threshold is adjusted based on a reinforcement learning algorithm.
7. The method according to claim 1, characterized in that The determining, based on the target object, the target file, and the target task, the storage popularity of the target file includes: Obtaining chat information of the target object in the instant messaging software; Obtaining user information of the target object; Obtaining access information and file information of the target file; Obtaining task information of the target task; Determining a timeliness score of the target file based on the chat information and the access information; Determining a task relevance score between the target file and the target task based on the file information and the task information; Determining a user relevance score between the target object and the target file based on the user information and the task information; The storage popularity of the target file is determined based on one or more of the timeliness score, the task relevance score, and the user relevance score.
8. A data cache system, characterized in that: The system comprises: an acquisition module configured to acquire a target file in a target conversation of an instant messaging software; an object determination module, configured to determine a target object of the target conversation; A task determination module is configured to determine a target task based on the target object and the target file; a storage popularity determination module configured to determine the storage popularity of the target file based on the target object, the target file, and the target task; wherein the storage popularity is used to measure the importance of the target file to the target object and the target task; The storage policy determination module is configured to determine the storage policy of the target file based on the storage heat.
9. A data cache device, characterized in that: The apparatus comprises at least one processor and at least one memory; The at least one memory is for storing computer instructions; The at least one processor is configured to execute at least part of the computer instructions to implement the data caching method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The storage medium stores computer instructions, and when the computer instructions are executed by a processor, the data caching method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
File storage method and related equipment
CN112463727A
Data migration method and device, equipment and storage medium
CN117762898A
File storage method and device, equipment and medium
CN119620952A
File sending method and device of multi-level cache architecture, equipment and medium
CN119697176A
Storage resource management method and system and storage medium
CN120353397A
Cited By
Data processing method, electronic equipment, system, storage medium and program product
CN121166043A
Knowledge base document data caching method and related equipment
CN121681485A
Data caching method for knowledge base document and related device
CN121681485B