Data caching processing method and device

By building a data knowledge graph and generating a cache strategy, the problem of low data cache processing efficiency and accuracy in cloud disk services is solved, and accurate prediction of user file access needs and optimized use of storage resources is achieved.

CN119938715APending Publication Date: 2025-05-06CHINA MOBILE INTERNET CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510057644.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-06

Smart Images

  • Figure CN119938715A_ABST
    Figure CN119938715A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data caching processing method and device, and the method comprises the steps: firstly constructing a data knowledge graph based on the resource data of a cloud disk and the access data of a user to the cloud disk, and then extracting the node features of each node of the data knowledge graph, the node features are input into an access prediction module for access data prediction to obtain target access data, a caching strategy is generated according to the resource data, and caching processing of the target access data is performed according to the caching strategy, so that accurate prediction of future file access requirements of the user is realized, the waiting time of the user is shortened, and the user experience is improved. Optimal use of storage resources is achieved, and unnecessary resource waste is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of data processing, and in particular to a data cache processing method and device. Background Art

[0002] With the development of information technology, people have higher and higher requirements for cloud disk services. In cloud disk services, the main purpose of optimizing the cache of cloud disk resource data is to increase data access speed and improve user experience. The current mainstream technical solutions mainly include full disk cache and cache strategies based on user operation behavior; However, when using cloud disk services, users often face pain points including data access delays, waste of resources caused by unnecessary data synchronization, and storage and battery life limitations on mobile devices. Especially for users who store large amounts of data, efficient and intelligent file caching mechanisms become the key to improving the service experience. Therefore, how to optimize cloud disk services has become a focus of future attention. Summary of the invention

[0003] An object of an embodiment of the present specification is to provide a data cache processing method and device to solve the problems of low efficiency and low accuracy of data cache processing.

[0004] To solve the above technical problems, an embodiment of this specification is implemented as follows: In a first aspect, an embodiment of the present specification provides a data cache processing method, the method comprising: Extracting node features of each node of a data knowledge graph; the data knowledge graph is constructed based on resource data of a cloud disk and user access behavior data to the cloud disk; Inputting the node features into an access prediction module to predict access data and obtain target access data; A cache strategy is generated according to the resource data, and cache processing of the target access data is performed according to the cache strategy.

[0005] In a second aspect, an embodiment of the present specification provides a data cache processing device, the device comprising: A feature extraction module is configured to extract node features of each node of a data knowledge graph; the data knowledge graph is constructed based on resource data of a cloud disk and user access behavior data to the cloud disk; An access data prediction module is configured to input the node features into the access prediction module to perform access data prediction and obtain target access data; The cache strategy generation module is configured to generate a cache strategy according to the resource data, and perform cache processing on the target access data according to the cache strategy.

[0006] In a third aspect, an embodiment of the present specification provides a data cache processing device, comprising: a memory, a processor, and computer executable instructions stored in the memory and executable on the processor, wherein the computer executable instructions, when executed by the processor, implement the steps of the data cache processing method described in the first aspect above.

[0007] In a fourth aspect, an embodiment of the present specification provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the data cache processing method described in the first aspect above are implemented.

[0008] In a fifth aspect, an embodiment of the present specification provides a computer program product, wherein the computer program product includes a data cache processing program, and the data cache processing program is executed by a processor to implement the steps of the data cache processing method described in the first aspect above.

[0009] The data cache processing method provided in this embodiment first constructs a data knowledge graph based on the resource data of the cloud disk and the user's access data to the cloud disk, then extracts the node features of each node of the data knowledge graph, and again inputs the node features into the access prediction module to predict the access data to obtain the target access data, and finally generates a cache strategy based on the resource data, and caches the target access data according to the cache strategy, so as to achieve accurate prediction of the user's future file access needs, reduce user waiting time, achieve optimal use of storage resources, and avoid unnecessary waste of resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in one or more embodiments of the present specification, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0011] Figure 1 A processing flow chart of a data cache processing method provided by one embodiment of this specification; Figure 2 A processing flow chart of a cache processing method applied to a cache processing scenario of cloud disk resource data provided by an embodiment of this specification; Figure 3 A schematic diagram of a data cache processing device provided by an embodiment of this specification; Figure 4 A schematic diagram of the structure of a data cache processing device provided in one embodiment of this specification. DETAILED DESCRIPTION

[0012] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will be combined with the drawings in one or more embodiments of this specification to clearly and completely describe the technical solutions in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this document.

[0013] One or more embodiments of a data cache processing method provided in this specification are as follows: Reference Figure 1 , which shows a processing flow chart of a data cache processing method provided in this embodiment, specifically including the following steps S102 to S106.

[0014] Step S102, extracting node features of each node in the data knowledge graph.

[0015] In this embodiment, the data knowledge graph can be used to reflect the interest preferences of users using the cloud disk, the association between the resource data of the cloud disk, etc. Optionally, the data knowledge graph is constructed based on the resource data of the cloud disk and the user's access behavior data to the cloud disk. In addition, the data knowledge graph can also be constructed based on the user's user data, the resource data of the cloud disk, and the user's access behavior data to the cloud disk; for example, the user data can be the user's mobile phone number, the login ID (Identification) of the cloud disk application, the user ID, and other data that can uniquely identify the user's identity. The user's access behavior data can be the user's access record, upload record, modification record, sharing record, etc. of the resource data in the cloud disk. The resource data of the cloud disk can be the files stored in the cloud disk. In addition, the resource data of the cloud disk can also include the cache hit rate, average response time, and data utilization rate of the cloud disk; optionally, the cloud disk can be a cloud disk application. It should be noted that the node feature can be represented by the embedding of the node.

[0016] During the specific execution process, a heterogeneous knowledge graph (i.e., a data knowledge graph) may be constructed in advance based on the data related to the cloud disk. In an optional implementation provided in this embodiment, the data knowledge graph is constructed in the following manner: Determine the nodes of the data knowledge graph according to the resource data and the user data of the user, and determine the edges of the data knowledge graph according to the access behavior data; The data knowledge graph is constructed based on the nodes and the edges.

[0017] For example, mapping various types of data into a heterogeneous knowledge graph .in, Represents a collection of nodes (such as users, files, tags, etc.), represents a set of edges (such as user-file access relations, file-label annotation relations, etc.), Represents a collection of relationship types.

[0018] It should be noted that after obtaining the resource data, user data and access behavior data, the analysis module can also be used to pre-clean, process and statistically analyze the resource data, user data and access behavior data to obtain more effective target data, and build a data knowledge graph based on the target data to improve the accuracy of the data knowledge graph construction.

[0019] During the specific execution process, after the data knowledge graph is constructed, the node features of each node of the data knowledge graph can be extracted. In an optional implementation provided in this embodiment, the node features of each node of the data knowledge graph are extracted in the following manner: Inputting the data knowledge graph into a feature extraction module to extract node features and obtain a first node feature; The first node feature is input into a feature fusion module to perform node feature fusion to obtain the node feature.

[0020] Specifically, the node features of each node in the data knowledge graph can be extracted through the knowledge graph embedding algorithm of heterogeneous information fusion (HIN-Embed), where the HIN-Embed algorithm includes a heterogeneous graph convolution (HIN-Conv) layer and a graph encoder, and node feature extraction is performed through HIN-Conv (i.e., feature extraction module), and node feature fusion is performed through the graph encoder (i.e., feature fusion module).

[0021] Furthermore, the feature extraction module may include two submodules: a node feature conversion module and a relationship feature aggregation module. Therefore, in the process of node feature extraction, node feature extraction may be performed based on the cooperation of the node feature conversion module and the relationship feature aggregation module. In an optional implementation provided in this embodiment, feature extraction is performed in the following manner: Extracting initial node features of each node of the data knowledge graph based on a node feature conversion matrix; Performing feature aggregation processing on edge features of each edge of the data knowledge graph based on the attention weight matrix to obtain aggregated edge features; The initial node feature and the aggregated edge feature are subjected to feature splicing processing to obtain the first node feature.

[0022] For example, first perform node feature transformation: for each type of node , construct a specific feature transformation matrix (in is the dimension of the node feature), the converted node feature It is expressed as:

[0023] in, For Node In the The feature representation of the layer, is the activation function (such as ReLU).

[0024] Next, perform relationship feature aggregation: for nodes Each relationship type , define a relation-specific attention weight matrix ,node In relationship The set of neighbor nodes on .

[0025] Aggregated relationship features for:

[0026] Among them, the attention weight Measured in relationship Next, neighbor nodes For Node Importance:

[0027] Finally, The nodes of the layer are represented as the concatenation of the transformed features and the aggregated relational features:

[0028] in, Representation Node In the The final feature vector of the layer, It is the relationship feature after aggregation.

[0029] In specific implementation, in the process of node feature fusion, node feature fusion can be performed based on attention weight. In an optional implementation provided in this embodiment, node feature fusion is performed in the following manner: Calculate the attention weight of the first node feature based on the query matrix; Perform weighted sum processing on the attention weight and the first node feature to obtain the node feature.

[0030] Specifically, the feature fusion module can be a graph encoder that contains multiple HIN-Conv layers and introduces a layer-specific attention mechanism to adaptively fuse node representations at different levels.

[0031] For example, suppose the graph encoder has Layer , then The output node of the layer represents for:

[0032] The final node embedding representation Obtained by weighted summation of hierarchical attention :

[0033] Among them, the level attention weight Reflects the Layer nodes represent their contribution to the final embedding:

[0034] in, is a learnable attention query vector.

[0035] In addition, in order to further improve the quality of embedded representation, pre-training and fine-tuning strategies can be designed in HIN-Embed. Specifically, by using large-scale cloud disk system log data, the parameters of the heterogeneous graph convolutional layer and encoder are pre-trained in an unsupervised manner (such as random walks, neighbor predictions, etc.), which helps to capture common structural and semantic information. In downstream tasks (such as cache strategy generation), task-specific supervisory signals (such as cache hit rate) can be used to fine-tune the pre-trained model, which enables the embedded representation to better adapt to specific application scenarios.

[0036] In practical applications, in order to provide users with intuitive, hierarchical, and personalized knowledge graph visualization services, it is also possible to adaptively generate multi-granular, multi-perspective knowledge graph views based on user interests and interactive behaviors, and provide rich interactive and exploration functions. In an optional implementation provided in this embodiment, the data knowledge graph is displayed in the following manner: Building a user data model based on the access behavior data; A knowledge graph subgraph is generated according to the user data model and the data knowledge graph, and the knowledge graph subgraph is displayed based on a display device.

[0037] Specifically, a user-oriented hierarchical knowledge graph visualization algorithm (HKG-Vis) can be used to provide cloud disk users with intuitive, hierarchical, personalized knowledge graph visualization services, that is, the above-mentioned process of building a user data model based on the data knowledge graph and generating a knowledge graph subgraph according to the user data model and the data knowledge graph can be implemented by the HKG-Vis algorithm; The implementation process of the HKG-Vis algorithm is as follows: First, based on the user's access behavior data (such as file access, retrieval, etc.), a user interest model (i.e., user data model) is constructed. Specifically, the topic model and attention mechanism are used to extract implicit interest topics and weights from the user behavior sequence: 1. Build a topic model , discover implicit interest topics from user behavior data:

[0038] in, Indicates words (or entities), Indicates Themes, Indicates Theme generation The probability of a word, Indicates user For The preference probability of a topic; 2. Use attention mechanism , dynamically adjust the topic weight according to the user behavior sequence:

[0039] in, Indicates user For The attention weight of each action, Indicates user For The attention score of each behavior, represents the parameter vector in the attention mechanism, represents the parameter matrix in the attention mechanism, Indicates user The hidden representation vector of Indicates The latent representation vector of each action, Indicates user The set of behaviors; 3. Generate user interest vector representation , as the basis for subsequent knowledge graph visualization:

[0040] in, Indicates The vector representation of the topic, Indicates user The interest vector of Secondly, based on the user interest model and data knowledge graph, multi-granularity subgraphs are dynamically generated as the basis for visualization. Specifically, this is achieved by designing a random walk-based graph sampling algorithm and a clustering-based graph abstraction algorithm: 1. Interest-driven random walk , by introducing the user interest vector, guiding the random walk process, and generating a local subgraph related to the user's interest:

[0041] in, , represents a node in the knowledge graph, Representation Node and The edge between Representation Node The set of neighbor nodes; 2. Evaluate the importance of nodes , evaluate the importance of knowledge graph nodes:

[0042] in, Representation Node The importance score of Representation Node The out-degree, represents the restart probability in the algorithm, Representation Node The initial importance score of 3. Perform node / edge filtering: Extract key nodes and edges from the original knowledge graph based on the importance thresholds of nodes and edges to simplify visualization complexity:

[0043] in, , represents the importance threshold of nodes and edges, , Represents the filtered set of nodes and edges; 4. Perform community discovery and graph abstraction: Use spectral clustering algorithms to discover semantic communities in knowledge graphs and achieve multi-level abstraction of graph structures:

[0044] in, represents the community division result, Indicates communities, Representation Node The characteristic vector of Indicates Based on the generation of multi-granularity knowledge graphs, we have designed rich interactive and visualization functions to provide users with a multi-perspective and in-depth knowledge exploration experience.

[0045] Step S104: input the node features into an access prediction module to perform access data prediction to obtain target access data.

[0046] After constructing the data knowledge graph based on the resource data of the cloud disk and the user's access behavior data to the cloud disk, and extracting the node features of each node of the data knowledge graph, in this step, the node features are input into the access prediction module to predict the access data and obtain the target access data.

[0047] Specifically, the access prediction module predicts access data by calling the dynamic attention enhanced behavior sequence modeling algorithm (DAB-Seq). The core idea of ​​the DAB-Seq algorithm is to construct a dynamic attention mechanism and interest drift detection to adaptively capture the temporal dependency of user behavior and the trend of interest changes.

[0048] In the specific implementation process, in an optional implementation mode provided by this embodiment, access data prediction is performed in the following manner: Inputting the node features into at least one attention module to perform attention calculations to obtain respective attention calculation results; Fusing the attention calculation results to obtain a fused attention matrix; Access data prediction is performed based on the fused attention matrix to obtain the target access data.

[0049] Optionally, at least one attention module includes a short-term self-attention module and a long-term self-attention module, so that the node features are input into at least one attention module for attention calculation to obtain each attention calculation result, including: inputting the node features into the short-term self-attention module for self-attention calculation to obtain a short-term self-attention vector, and inputting the node features into the long-term self-attention module for self-attention calculation to obtain a long-term self-attention vector.

[0050] Specifically, in DAB-Seq, a novel dynamic attention mechanism is designed to capture the long-term and short-term dependencies in user behavior sequences. Two key modules are introduced: the short-term attention module and the long-term attention module. The short-term attention module is designed based on the self-attention mechanism of Transformer to capture the short-term dependencies of user behaviors, and the long-term attention module is designed based on the long-term attention mechanism of GRU to capture the long-term dependencies and interest evolution of user behaviors.

[0051] For example, the short-term attention module in DAB-Seq:

[0052] in, They represent the query, key, and value vectors in the short-term attention module, respectively, and are used to calculate the correlation between the current behavior and the short-term historical behavior.

[0053] Long-term attention module:

[0054] in, They represent the query, key, and value vectors in the long-term attention module, respectively, and are used to calculate the correlation between the current behavior and the long-term historical behavior.

[0055] Attention function: Use the dot product attention function to calculate the correlation weight between the current behavior and the historical behavior.

[0056]

[0057] in, To calculate the inner product of the query vector q and each row in the key matrix K (i.e. each historical behavior), a relevance score vector is obtained. Represents the dimension of the scaling factor in the attention function, usually set to the square root of the embedding dimension, used to adjust the magnitude of the dot product result; Dynamic Fusion : Adaptively fuse short-term and long-term behavior representations via a learnable fusion matrix.

[0058]

[0059] in, The fusion matrices representing the outputs of the short-term and long-term attention modules are used to adaptively fuse short-term and long-term behavior representations.

[0060] In addition, the DAB-Seq algorithm also includes an interest drift detection module, which determines whether the user's interest has changed significantly by comparing the similarity between the current behavior representation and the historical interest vector.

[0061] For example, interest drift detection : Introducing interest drift metrics to quantify the difference between current behavior and historical interests .

[0062]

[0063] in, is the fusion matrix, indicating that at the time step The generated dynamic representation of user behavior combines the information extracted by the short-term attention module and the long-term attention module, and can fully capture the temporal dependencies of user behavior. Adaptive Updates : Through the adaptive interest update strategy, the interest vector is adjusted in time when significant drift occurs:

[0064] in, Respectively represent the time step and The historical interest vector is used to describe the dynamic evolution of user interests; Update weights :Design learnable update weights to dynamically adjust the update strength based on the relevance of current behavior and historical interests:

[0065] in, Represents the adaptive update weight of the interest vector, dynamically adjusting the update strength based on the correlation between current behavior and historical interests. Represent the learnable parameter vector and bias term of the updated weights, respectively, which are used to calculate the updated weights based on the current behavior and historical interests.

[0066] Furthermore, in the process of predicting access data based on the fused attention matrix, at least one access parameter of the user can be predicted based on the fused attention matrix, and access data prediction is performed according to the at least one access parameter to obtain target access data: Predicting at least one access parameter of the user to the resource data based on the fused attention matrix; the access parameter includes access probability, access time and access type; Access data prediction is performed according to the at least one access parameter to obtain the target access data.

[0067] Optionally, the access parameters include access probability, access time and / or access type.

[0068] For example, in the process of predicting the user's future behavior through prediction algorithms, auxiliary tasks (such as behavior type prediction, time interval prediction, etc.) can be constructed to enhance the representation ability of the model: Main task prediction: Based on dynamic behavior representation, predict the user's future behavior probability distribution . ,in, They respectively represent the output layer weight matrix and bias term of the main task (future behavior prediction), which are used to map the behavior representation to the probability distribution of the behavior category. Indicates that at time step The user's real future behavior. Indicates that at time step The generated dynamic representation of user behavior integrates the information extracted by the short-term attention module and the long-term attention module, and can fully capture the temporal dependencies of user behavior; Auxiliary Task 1: Constructing Behavior Type Prediction Task ,enhancing the model’s understanding of behavioral semantics. ,in, They respectively represent the output layer weight matrix and bias term of auxiliary task 1 (behavior type prediction), which are used to map the behavior representation to the probability distribution of the behavior type. Indicates that at time step The actual type of user behavior; Auxiliary Task 2: Constructing time interval prediction tasks to enhance the model’s ability to model behavior temporal dependencies. ,in, They represent the output layer weight matrix and bias term of auxiliary task 2 (time interval prediction), which are used to map the behavior representation to the predicted value of the time interval. Indicates that at time step The actual time interval between a user's action and the previous action; Joint Optimization: Design multi-task joint optimization objectives to fully utilize auxiliary information while predicting future behaviors. ,in, They respectively represent the weight coefficients of auxiliary task 1 and auxiliary task 2 in the joint optimization objective, which are used to balance the contributions of different tasks. They represent the loss functions of the main task (future behavior prediction), auxiliary task 1 (behavior type prediction), and auxiliary task 2 (time interval prediction), respectively.

[0069] In addition, in the process of predicting access data, in order to improve the prediction accuracy, the node features can be fused with the user's access behavior data, and access data prediction can be performed based on the fused access behavior feature data to obtain the target access data.

[0070] For example: , where behavior_features is the access behavior data, and graph_features is the node features, including: the embedding vector of the file node (generated by the HIN-Embed algorithm), the semantic association information of related entities, and the graph structure features (such as node centrality, clustering coefficient, etc.). It should be noted that the access behavior feature data can also be called the user behavior sequence.

[0071] Step S106: Generate a cache strategy according to the resource data, and perform cache processing on the target access data according to the cache strategy.

[0072] After the node features are input into the access prediction module to predict the access data and the target access data is obtained, in this step, a cache strategy is generated according to the resource data, and the target access data is cached according to the cache strategy. Specifically, a cache strategy can be generated based on the multi-objective resource scheduling algorithm for cloud disk cache optimization (MORC-Sched), and the target access data is cached according to the cache strategy.

[0073] In the specific implementation process, in an optional implementation mode provided by this embodiment, the cache strategy is generated in the following manner: Performing a weighted combination on the cache hit rate, average response time and data utilization rate contained in the resource data to obtain cache parameters; A target cache policy is determined in a cache policy space based on the cache parameters, and the target cache policy is used as the cache policy.

[0074] For example, the implementation process of the MORC-Sched algorithm is as follows: First, the system state is represented, specifically, the current state of the cloud disk is represented as a multidimensional vector , which includes key performance indicators such as cache hit rate, average response time, resource utilization, etc. In addition, it also considers the dynamic characteristics of the system, such as user request patterns, hot file distribution, etc.; Then design a multi-objective reward function: In order to balance multiple performance objectives, a weighted multi-objective reward function is designed. Specifically, for each performance indicator, a normalized reward component is defined and linearly combined through learnable weight parameters, including: Cache Hit Ratio: , that is, building cache hit rate as a key performance indicator to encourage algorithms to improve cache utilization efficiency; Average response time: , that is, building a response time reward to encourage the algorithm to reduce the waiting delay of user requests; Resource Utilization: , that is, building resource utilization rewards to encourage algorithms to improve the efficiency of cache resource usage; Weighted combination: , that is, through weighted combination, the trade-offs between different performance objectives are balanced to achieve flexible multi-objective optimization.

[0075] Then design the action space: In MORC-Sched, cache resource allocation and replacement strategies are abstracted into a unified action space. Specifically, for each file request, the algorithm needs to decide whether to cache it and choose the replacement file when the cache space is insufficient, including: Caching Decision: , that is, building a binary cache decision variable to control whether each file request is cached; Replacement decision: , that is, construct discrete replacement decision variables and select the cache file to be replaced; Joint Action Space: , that is, designing a joint action space, optimizing cache and replacement decisions at the same time, and improving the flexibility of the strategy.

[0076] Finally, multi-objective reinforcement learning optimization is performed: Based on the above defined states, actions, and rewards, a multi-objective reinforcement learning algorithm is used to optimize the cache scheduling strategy. Specifically, a Q-learning algorithm variant based on Pareto optimality is selected, including: Q value update: , that is, using the Q-learning algorithm to update the action value function through temporal difference learning, where, In Status Take action The action value function represents the expected value of future cumulative rewards. Respectively represent the current state and the next state, that is, the cloud disk system state represented by a multidimensional vector, They represent the action taken in the current state and the action taken in the next state, which is a combination of caching and replacement decisions. In Status Take action The immediate reward obtained after is the output of the multi-objective reward function. Discount factor, which indicates the importance of future rewards relative to current rewards, has a value range of ; Comparison of the advantages and disadvantages of objects: , that is, to build a comparison relationship between the advantages and disadvantages of objects and judge the advantages and disadvantages of different actions in the multi-objective space, where No. The action value function corresponding to the target is expressed in state Take action At that time, The expected value of future cumulative rewards for each performance indicator; Action selection: , that is, according to the learned Q value, select the optimal cache scheduling action, where, In Status The optimal action selection strategy under , that is, selecting the optimal action based on the current Q value; Experience Replay: , that is, building an experience replay mechanism, sampling from historical trajectories, improving sample utilization efficiency and learning stability, where, Experience replay buffer, used to store historical state transitions and reward samples .

[0077] To summarize, the data cache processing method provided in this embodiment first constructs a data knowledge graph based on the resource data of the cloud disk and the user's access data to the cloud disk, then extracts the node features of each node of the data knowledge graph, and again inputs the node features into the access prediction module to predict the access data to obtain the target access data, and finally generates a cache strategy based on the resource data, and caches the target access data according to the cache strategy, so as to achieve accurate prediction of the user's future file access needs, reduce user waiting time, achieve optimal use of storage resources, and avoid unnecessary waste of resources.

[0078] The following combination Figure 2 , taking the application of the data cache processing method provided by this embodiment in the cache processing scenario of cloud disk resource data as an example, the data cache processing method provided by this embodiment is further explained. Figure 2 The data cache processing method applied to the cache processing scenario of cloud disk resource data specifically includes steps S202 to S218.

[0079] Step S202, constructing a data knowledge graph based on the resource data of the cloud disk and the user's access behavior data to the cloud disk.

[0080] Step S204: input the data knowledge graph into the feature extraction module to extract node features and obtain the first node feature.

[0081] Step S206: input the first node feature into a feature fusion module to perform node feature fusion to obtain a node feature.

[0082] Step S208: input the node features into at least one attention module to perform attention calculation to obtain various attention calculation results.

[0083] Step S210, fuse the attention calculation results to obtain a fused attention matrix.

[0084] Step S212, predict access data based on the fused attention matrix to obtain target access data.

[0085] Step S214: Perform a weighted combination on the cache hit rate, average response time and data utilization rate included in the resource data to obtain cache parameters.

[0086] Step S216: determine a target cache strategy in the cache strategy space based on the cache parameters, and use the target cache strategy as the cache strategy.

[0087] Step S218: Cache the target access data according to the target cache strategy.

[0088] This specification provides a data cache processing method embodiment: Figure 3 A schematic diagram of a data cache processing device provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown, the device comprises: A feature extraction module 302 is configured to extract node features of each node of a data knowledge graph; the data knowledge graph is constructed based on resource data of the cloud disk and user access behavior data to the cloud disk; The access data prediction module 304 is configured to input the node features into the access prediction module to perform access data prediction and obtain target access data; The cache strategy generation module 306 is configured to generate a cache strategy according to the resource data, and perform cache processing on the target access data according to the cache strategy.

[0089] The data cache processing device provided in this embodiment first runs the feature extraction module 302 to extract the node features of each node of the data knowledge graph; the data knowledge graph is constructed based on the resource data of the cloud disk and the user's access behavior data to the cloud disk, and then runs the access data prediction module 304 to input the node features into the access prediction module to predict the access data and obtain the target access data. Finally, the cache strategy generation module 306 generates a cache strategy according to the resource data, and caches the target access data according to the cache strategy, so as to achieve accurate prediction of the user's future file access needs, reduce user waiting time, achieve optimal use of storage resources, and avoid unnecessary waste of resources.

[0090] The data cache processing device provided in an embodiment of this specification can implement each process in the aforementioned method embodiment and achieve the same functions and effects, which will not be repeated here.

[0091] Furthermore, an embodiment of the present specification also provides a data cache processing device, Figure 4 A schematic diagram of a data cache processing device provided in an embodiment of this specification is shown in FIG. Figure 4 As shown, the device includes: a memory 401, a processor 402, a bus 403 and a communication interface 404. The memory 401, the processor 402 and the communication interface 404 communicate through the bus 403, and the communication interface 404 may include an input and output interface, which includes but is not limited to a keyboard, a mouse, a display, a microphone, a loudspeaker, etc.

[0092] Figure 4 In the embodiment, the memory 401 stores computer executable instructions that can be run on the processor 402. When the computer executable instructions are executed by the processor 402, the following process is implemented: Extracting node features of each node of a data knowledge graph; the data knowledge graph is constructed based on resource data of a cloud disk and user access behavior data to the cloud disk; Inputting the node features into an access prediction module to predict access data and obtain target access data; A cache strategy is generated according to the resource data, and cache processing of the target access data is performed according to the cache strategy.

[0093] The data cache processing device provided in this embodiment, through the cooperation of the memory 401, the processor 402, the bus 403 and the communication interface 404, first builds a data knowledge graph based on the resource data of the cloud disk and the user's access data to the cloud disk, and then extracts the node features of each node of the data knowledge graph, and again inputs the node features into the access prediction module to predict the access data to obtain the target access data, and finally generates a cache strategy according to the resource data, and caches the target access data according to the cache strategy, so as to achieve accurate prediction of the user's future file access needs, reduce user waiting time, achieve optimal use of storage resources, and avoid unnecessary waste of resources.

[0094] The data cache processing device provided in an embodiment of this specification can implement each process in the aforementioned method embodiment and achieve the same functions and effects, which will not be repeated here.

[0095] Furthermore, an embodiment of the present specification also provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store computer-executable instructions, and when the computer-executable instructions are executed by a processor, the following process is implemented: Extracting node features of each node of a data knowledge graph; the data knowledge graph is constructed based on resource data of a cloud disk and user access behavior data to the cloud disk; Inputting the node features into an access prediction module to predict access data and obtain target access data; A cache strategy is generated according to the resource data, and cache processing of the target access data is performed according to the cache strategy.

[0096] The computer-readable storage medium provided in this embodiment first constructs a data knowledge graph based on the resource data of the cloud disk and the user's access data to the cloud disk, then extracts the node features of each node of the data knowledge graph, and again inputs the node features into the access prediction module to predict the access data to obtain the target access data, and finally generates a cache strategy based on the resource data, and caches the target access data according to the cache strategy, so as to achieve accurate prediction of the user's future file access needs, reduce user waiting time, achieve optimal use of storage resources, and avoid unnecessary waste of resources.

[0097] The computer-readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0098] The computer-readable storage medium provided in an embodiment of this specification can implement each process in the aforementioned method embodiment and achieve the same functions and effects, which will not be repeated here.

[0099] Furthermore, an embodiment of the present specification also provides a computer program product, which implements the following process when executed by a processor: Extracting node features of each node of a data knowledge graph; the data knowledge graph is constructed based on resource data of a cloud disk and user access behavior data to the cloud disk; Inputting the node features into an access prediction module to predict access data and obtain target access data; A cache strategy is generated according to the resource data, and cache processing of the target access data is performed according to the cache strategy.

[0100] The computer program product provided in this embodiment first constructs a data knowledge graph based on the resource data of the cloud disk and the user's access data to the cloud disk, then extracts the node features of each node of the data knowledge graph, and again inputs the node features into the access prediction module to predict the access data to obtain the target access data, and finally generates a cache strategy based on the resource data, and caches the target access data according to the cache strategy, so as to achieve accurate prediction of the user's future file access needs, reduce user waiting time, achieve optimal use of storage resources, and avoid unnecessary waste of resources.

[0101] The computer program product provided in one embodiment of this specification can implement each process in the aforementioned method embodiment and achieve the same functions and effects, which will not be repeated here.

[0102] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0103] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0104] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0106] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0107] Memory may include non-permanent storage in a computer-readable storage medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable storage medium.

[0108] Computer-readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable storage media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0109] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0110] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0111] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A data cache processing method, characterized in that: The method comprises: Extracting node features of each node of a data knowledge graph; the data knowledge graph is constructed based on resource data of a cloud disk and user access behavior data to the cloud disk; Inputting the node features into an access prediction module to predict access data and obtain target access data; A cache strategy is generated according to the resource data, and cache processing of the target access data is performed according to the cache strategy.

2. The data cache processing method according to claim 1, characterized in that: The data knowledge graph is constructed in the following way: Determine the nodes of the data knowledge graph according to the resource data and the user data of the user, and determine the edges of the data knowledge graph according to the access behavior data; The data knowledge graph is constructed based on the nodes and the edges.

3. The data cache processing method according to claim 1, characterized in that: The node features of each node of the extracted data knowledge graph include: Extracting initial node features of each node of the data knowledge graph based on a node feature conversion matrix; Performing feature aggregation processing on edge features of each edge of the data knowledge graph based on the attention weight matrix to obtain aggregated edge features; Performing feature concatenation processing on the initial node feature and the aggregated edge feature to obtain the first node feature, and calculating the attention weight of the first node feature based on the query matrix; Perform weighted sum processing on the attention weight and the first node feature to obtain the node feature.

4. The data cache processing method according to claim 1, characterized in that: The step of inputting the node feature into an access data prediction module to perform access data prediction to obtain target access data includes: Inputting the node features into at least one attention module to perform attention calculations to obtain respective attention calculation results; Fusing the attention calculation results to obtain a fused attention matrix; Access data prediction is performed based on the fused attention matrix to obtain the target access data.

5. The data cache processing method according to claim 4, characterized in that: The predicting of access data based on the fused attention matrix to obtain the target access data includes: Predicting at least one access parameter of the user to the resource data based on the fused attention matrix; the access parameter includes access probability, access time and access type; Access data prediction is performed according to the at least one access parameter to obtain the target access data.

6. The data cache processing method according to claim 1, characterized in that: The cache strategy is generated in the following way: Performing a weighted combination on the cache hit rate, average response time and data utilization rate contained in the resource data to obtain cache parameters; A target cache policy is determined in a cache policy space based on the cache parameters, and the target cache policy is used as the cache policy.

7. A data cache processing device, characterized in that: The device comprises: A feature extraction module is configured to extract node features of each node of a data knowledge graph; the data knowledge graph is constructed based on resource data of a cloud disk and user access behavior data to the cloud disk; An access data prediction module is configured to input the node features into the access prediction module to perform access data prediction and obtain target access data; The cache strategy generation module is configured to generate a cache strategy according to the resource data, and perform cache processing on the target access data according to the cache strategy.

8. A data cache processing device, characterized in that: The device includes a memory and a processor, wherein the memory stores computer executable instructions, and when the computer executable instructions are executed on the processor, the steps of the method described in any one of claims 1 to 6 can be implemented.

9. A computer-readable storage medium storing computer-executable instructions, characterized in that: When the computer executable instructions are executed by a processor, the steps of the method described in any one of claims 1 to 6 can be implemented.

10. A computer program product, characterized in that The computer program product comprises a data cache processing program, and when the data cache processing program is executed by a processor, the steps of the method as described in any one of claims 1 to 6 can be implemented.

Citation Information

Cited By

  • Resource loading method and system based on user behavior prediction

    CN121412457A

  • Data preloading method and device, electronic equipment and storage medium

    CN121455554A