Data manager, data management method, electronic equipment and storage medium

Through the synergy between the processor and the controller, data classification is stored in different memory or partitions according to the data access probability, which solves the contradiction between storage capacity and access delay and improves data management efficiency.

CN120335731AActive Publication Date: 2025-07-18INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510822147.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The prior art cannot meet the requirements of storage capacity and access latency at the same time, resulting in low data management efficiency.

Method used

The data type is determined by the processor according to the preset algorithm, and the controller stores different types of data in different memory or different partitions of the same memory, including storing high-access probability data in high-performance memory, storing medium-access probability data in the cache area, and storing low-access probability data for compression.

Benefits of technology

While meeting storage capacity, it reduces data access latency and improves data management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335731A_ABST
    Figure CN120335731A_ABST
Patent Text Reader

Abstract

The invention discloses a data manager, a data management method, an electronic device and a storage medium, and relates to the technical field of data processing.The data type of multiple pieces of data is determined through a processor according to a preset algorithm, and the data is managed through a controller according to the data type; according to the method and the device, the data of different data types are respectively stored in different memories or different partitions of the same memory, so that the problem of relatively low efficiency of data management can be solved, the access delay of the data is reduced while the storage capacity is met, and the efficiency of data management is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical field of data processing of the present application, in particular, relates to a data manager, a data management method, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of data analysis, the requirements for system storage capacity and access latency are getting higher and higher. Different types of memories have different storage characteristics. For example, solid-state drives can provide high-capacity storage, but have a large access latency and a low write lifespan; Dynamic Random-Access Memory (DRAM) has low latency, but its physical density and cost limit its application in large-scale data processing.

[0003] In the related art, the requirements for storage capacity and access latency cannot be satisfied simultaneously, resulting in low efficiency of data management. Summary of the Invention

[0004] The present application provides a data manager, a data management method, an electronic device, and a storage medium to at least solve the problem of low efficiency of data management in the related art.

[0005] The present application provides a data manager, which includes a processor and a controller. The controller is respectively connected to a first memory and a second memory;

[0006] The processor is configured to determine multiple first data, multiple second data, and multiple third data from multiple data according to an access log within a preset period through a preset algorithm. The predicted access probability corresponding to the second data is less than the predicted access probability corresponding to the first data and greater than the predicted access probability corresponding to the third data. The multiple data are the data stored in the first memory and the second memory;

[0007] The controller is configured to store the multiple first data in the first memory, store the multiple second data in the buffer area of the second memory, perform data compression processing on the third data, store the compressed multiple third data in the compressed storage area of the second memory, and update the address conversion relationship.

[0008] The present application further provides a data management method, including:

[0009] Obtaining an access log within a preset period;

[0010] Performing feature extraction processing on the access log to obtain multiple-dimensional features;

[0011] Determining multiple first data from multiple data according to the multiple-dimensional features and an access prediction model;

[0012] Determine multiple second data and multiple third data according to multiple dimensional features and multiple first data, wherein the predicted access probability corresponding to the second data is less than the predicted access probability corresponding to the first data and greater than the predicted access probability corresponding to the third data.

[0013] This application also provides a data management device, including:

[0014] An acquisition module, configured to acquire access logs within a preset time period;

[0015] An extraction module, configured to perform feature extraction processing on the access logs to obtain multiple dimensional features;

[0016] A first determination module, configured to determine multiple first data from multiple data according to multiple dimensional features and an access prediction model;

[0017] A second determination module, configured to determine multiple second data and multiple third data according to multiple dimensional features and multiple first data, wherein the predicted access probability corresponding to the second data is less than the predicted access probability corresponding to the first data and greater than the predicted access probability corresponding to the third data.

[0018] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above data management methods when executing the computer program.

[0019] This application also provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program implements the steps of any of the above data management methods when executed by a processor.

[0020] This application also provides a computer program product, including a computer program, which implements the steps of any of the above data management methods when executed by a processor.

[0021] Through this application, the processor determines the data types of multiple data according to a preset algorithm, and the controller stores data of different data types in different memories or different partitions of the same memory according to the data types. Therefore, the problem of low data management efficiency can be solved, the access latency of data can be reduced while meeting the storage capacity, and the data management efficiency is improved. Description of the Drawings

[0022] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 A schematic structural diagram of a data manager provided by an embodiment of the present application;

[0024] Figure 2 A schematic flowchart of a data management method provided by an embodiment of the present application;

[0025] Figure 3 A schematic diagram of a second threshold adjustment curve provided by an embodiment of the present application;

[0026] Figure 4 A schematic flowchart of a model incremental training method provided by an embodiment of the present application;

[0027] Figure 5 A schematic structural diagram of a data management device provided by an embodiment of the present application;

[0028] Figure 6 A schematic structural diagram of an electronic device provided by the present application. Detailed implementation manners

[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0030] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0031] With the rapid development of data analysis, the requirements for system storage capacity and access latency are getting higher and higher. Different types of memories have different storage characteristics. For example, solid-state drives can provide high-capacity storage, but have a large access latency and a low write life; Dynamic Random-Access Memory (DRAM) has low latency, but its physical density and cost limit its application in large-scale data processing.

[0032] In the related art, the requirements for storage capacity and access latency cannot be satisfied simultaneously, resulting in low data management efficiency.

[0033] To solve the above technical problems, an embodiment of the present application provides a data manager, which includes a processor and a controller. The controller is respectively connected to a first memory and a second memory. The processor determines the data types of multiple data according to a preset algorithm, and the controller stores data of different data types into different memories or different partitions of the same memory respectively according to the data types. In this way, through the intelligent data classification, storage and management by the controller, while meeting the storage capacity, the data access latency is reduced, and the efficiency of data management is improved.

[0034] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] Embodiment 1

[0036] Figure 1 It is a schematic structural diagram of a data manager provided by an embodiment of the present application. Please refer to Figure 1 , Figure 1 at least including a data manager.

[0037] The data manager includes a processor and a controller. The controller is respectively connected to a first memory and a second memory;

[0038] The processor is configured to determine multiple first data, multiple second data, and multiple third data from multiple data according to an access log within a preset period through a preset algorithm. The predicted access probability corresponding to the second data is less than the predicted access probability corresponding to the first data and greater than the predicted access probability corresponding to the third data. The multiple data are the data stored in the first memory and the second memory;

[0039] The controller is configured to store the multiple first data into the first memory, store the multiple second data into the buffer area of the second memory, perform data compression processing on the third data, store the compressed multiple third data into the compressed storage area of the second memory, and update the address conversion relationship.

[0040] Among them, the data manager can be used to manage and optimize data storage and access.

[0041] The processor can be used to classify and analyze multiple data stored in the first memory and the second memory to determine the storage strategy of the multiple data.

[0042] The processor can be a central processing unit or other types of processors, which are not limited herein.

[0043] The controller can be used to manage the specific operations of the first memory and the second memory. The specific operations can include data reading, writing, storage location allocation, etc.

[0044] The controller can be a field programmable gate array or other types of controllers, which are not limited herein.

[0045] The controller can include an address translation engine and a compression engine.

[0046] The address translation engine can be used for the conversion between virtual addresses and physical addresses.

[0047] In the operating system, virtual addresses can be used to manage data, and in hardware devices, the physical addresses corresponding to the virtual addresses can be used to access the actual physical locations.

[0048] The address translation engine can include a Translation Lookaside Buffer (TLB).

[0049] The TLB can be used to store address translation relationships to quickly find the mapping relationship from virtual addresses to physical addresses.

[0050] The TLB can store 1024 mapping relationships from virtual addresses to physical addresses simultaneously, greatly improving the hit rate of address translation and reducing the number of accesses to the page table.

[0051] The TLB can be set with a multi-stage pipeline architecture, and the multi-stage pipeline architecture can process different tasks in parallel to improve the speed of address translation and reduce latency.

[0052] The TLB can be used to support the direct conversion from virtual addresses to physical addresses, avoiding the overhead of page table traversal and improving the processing speed and efficiency of address translation.

[0053] The speed of address translation is significantly improved through the TLB and the pipeline architecture, and the latency is reduced. The mapping latency can reach 95 nanoseconds.

[0054] The compression engine can be used for data compression and decompression.

[0055] Data compression processing can significantly reduce the storage space occupied by data to improve storage efficiency and transmission speed.

[0056] Data compression processing can be achieved through a preset compression algorithm.

[0057] The preset compression algorithm can be implemented through a dedicated hardware circuit to realize the hardware implementation of the preset compression algorithm, and utilize the parallel processing ability of the hardware to improve compression efficiency.

[0058] The compression engine can be set with a streaming processing architecture to support real-time data compression and decompression.

[0059] The first memory can be a high-performance and low-latency storage device, which is not limited here.

[0060] The first memory can be used to store data with a high access probability. This data requires a fast access speed to meet the high-performance requirements of the system.

[0061] The second memory can be a storage device with a relatively large capacity and a slower read / write speed compared to the first memory, which is not limited here.

[0062] The second memory can include a buffer area and a compressed storage area.

[0063] The second memory can be used to store data with a medium access probability. This data still requires a relatively fast access speed.

[0064] The second memory can also be used to store data that has been compressed. This data has a low access probability, and the compression process can save storage space.

[0065] The second memory can reduce the actual write volume through data compression processing, improving the daily full-disk write count.

[0066] The second memory can monitor the wear state of the storage units in real time and dynamically adjust the data write position, evenly distributing the write operations to all available units to avoid local overwear.

[0067] The access log can be used to record the access information of each data within a preset time period.

[0068] The preset time period can be a preset time range. For example, the preset time period is the past 60 seconds.

[0069] The preset algorithm can be used to analyze the access log and classify the data according to specific logic.

[0070] The first data can be data with a relatively high access probability and can be stored in the high-performance first memory to ensure fast access.

[0071] The second data can be data with a medium access probability, between the first data and the third data, and can be stored in a storage device with slightly lower performance but lower cost to ensure relatively fast access.

[0072] The third data can be data with a low access probability and can be compressed for storage to save space.

[0073] Since the access log is dynamically changing and the data type is also dynamically adjusted, data management can adapt to different application scenarios and access patterns, improving the efficiency of data management.

[0074] For example, multiple data can be illustrated by Table 1:

[0075] Table 1

[0076]

[0077] In a possible implementation, the preset algorithm includes an access prediction model, and the processor is specifically configured to:

[0078] Perform feature extraction processing on the access log to obtain multi-dimensional features, where the multi-dimensional features include the number of accesses and the survival duration;

[0079] Determine multiple first data according to the number of accesses and the access prediction model;

[0080] Determine multiple third data according to the survival duration;

[0081] Determine multiple second data according to the multiple first data and the multiple third data.

[0082] Among them, the access prediction model can be a pre-trained learning model, which is not limited here.

[0083] The access log may include the following information: virtual address and physical address, access type (read / write), timestamp, thread identifier, process identifier, processor identifier, etc.

[0084] The multi-dimensional features may include temporal locality, spatial correlation, thread affinity, access intensity, number of accesses, survival duration, access time interval, access duration, data size, data type, data access pattern, data access heat change rate, processor utilization rate, load rate, etc., which are not limited here.

[0085] The multiple first data can be determined according to the number of accesses of the data, or can be determined according to the output result of the access prediction model.

[0086] The multiple third data can be determined according to the survival duration of the data.

[0087] The multiple second data can be data other than the multiple first data and the multiple third data.

[0088] In a possible implementation, the processor is further configured to send an operation instruction to the controller, and the operation instruction includes a virtual address;

[0089] The controller is further configured to determine the physical address corresponding to the virtual address according to the address conversion relationship, and transmit the data corresponding to the physical address to the processor through a preset transmission mode.

[0090] Among them, the operation instruction can be a write operation, a read operation, etc., which are not limited here.

[0091] The preset transfer mode can be a transfer mode that allows the hardware subsystem to directly access the system.

[0092] The controller can receive the virtual address sent by the processor, query whether the virtual address exists in the address translation relationship. If it exists, the physical address is directly determined. If it does not exist, the page table is accessed, and the address translation relationship is updated. Through the preset transfer mode, the data corresponding to the physical address is transmitted to the processor.

[0093] Through the address translation relationship, the access times to the page table can be significantly reduced, the latency is reduced, and through the preset transfer mode, the burden on the CPU is reduced, improving the overall performance of the system.

[0094] In a possible implementation manner, the controller is specifically configured to:

[0095] Split and process multiple third data to obtain multiple data blocks;

[0096] Through the target compression algorithm, perform dictionary encoding processing and compression processing on the multiple data blocks to obtain multiple compressed data blocks;

[0097] Determine the multiple compressed data blocks as the compressed multiple third data.

[0098] Among them, the multiple third data can be split and processed in the order of the timestamps of the multiple third data.

[0099] The dictionary encoding processing can reduce the data volume by finding duplicate strings or patterns and replacing them with shorter symbols or indexes.

[0100] The compression processing can be a process of further reducing data redundancy on the basis of dictionary encoding.

[0101] In this way, not only the storage space is saved, but also efficient decompression and data recovery are supported, which is suitable for data storage scenarios with low access frequencies.

[0102] Embodiment 2

[0103] Figure 2 It is a schematic flow chart of a data management method provided by an embodiment of the present application. The execution subject of the embodiment of the present application can be a processor. The processor can be implemented by software or by a combination of software and hardware. Please refer to Figure 2 and the method includes:

[0104] S201. Obtain the access logs within a preset time period.

[0105] The preset time period and the target object can be determined, and through the log collection tool, the access logs of the target object within the preset time period are obtained.

[0106] Among them, the target object can be any one or more of a server, a database, an operating system, a network device, and an application program.

[0107] Optionally, the processor can set a kernel module to implement monitoring of access logs and control of the controller, so that the controller can perform direct memory access.

[0108] S202. Perform feature extraction processing on the access log to obtain multi-dimensional features.

[0109] Multiple files in the access log can be parsed and converted into log data in a preset data format, and feature extraction processing is performed on the log data through a feature extraction algorithm to obtain multi-dimensional features.

[0110] Among them, the feature extraction algorithm can be a preset algorithm, which is not limited here.

[0111] S203. Determine multiple first data from multiple data according to the multi-dimensional features and the access prediction model.

[0112] Optionally, the multi-dimensional features can be input into the access prediction model to determine multiple first data from multiple data.

[0113] Optionally, the multi-dimensional features include access frequency and processor utilization rate. The multi-dimensional features can be input into the access prediction model to obtain multiple prediction data; the multiple prediction data are determined as multiple first data, and the multiple data other than the multiple prediction data are determined as multiple first processing data; for any one of the first processing data, it is judged whether the access frequency corresponding to the first processing data is greater than a second threshold. If so, the first processing data is determined as the first data.

[0114] Among them, the prediction data are data with an access probability greater than a first threshold within a preset time period.

[0115] The second threshold can be determined according to the processor utilization rate.

[0116] Optionally, the second threshold can be determined by the following formula:

[0117] HOT_THRESH = 1000 × (1 + CPU_util / 100)

[0118] Among them, HOT_THRESH can represent the second threshold, and CPU_util can represent the processor utilization rate.

[0119] Through this formula, the second threshold can be dynamically adjusted. Whenever the processor utilization rate increases by 1%, the threshold increases by 1%.

[0120] Next, in combination withFigure 3 , an example of the second threshold is given.

[0121] Figure 3 This is a schematic diagram of a second threshold adjustment curve provided by an embodiment of the present application. Please refer to Figure 3 , Figure 3 It includes a relationship diagram between processor utilization and the second threshold.

[0122] Among them, the X-axis of the relationship diagram can represent processor utilization, and the Y-axis can represent the second threshold.

[0123] In this way, more hot data can be retained under low load, and high-frequency access data can be preferentially guaranteed under high load.

[0124] Optionally, when determining multiple first data, set the survival duration corresponding to each first data to 0.

[0125] It should be noted that multiple first data can be determined according to any feasible implementation manner, and the embodiments of the present application do not limit this.

[0126] S204. Determine multiple second data and multiple third data according to multiple dimensional features and multiple first data.

[0127] The predicted access probability corresponding to the second data is less than the predicted access probability corresponding to the first data and greater than the predicted access probability corresponding to the third data.

[0128] Optionally, data other than multiple first data can be determined as multiple second processing data, input the multiple dimensional features corresponding to the multiple second processing data into the cold data prediction model to determine multiple third data, and determine multiple second data according to the multiple third data and multiple first data.

[0129] Among them, the cold data prediction model can be a pre-trained model, which is not limited here.

[0130] Optionally, the multiple dimensional features further include survival duration. Among multiple data, data other than multiple first data is determined as multiple second processing data; for any one of the second processing data, determine whether the survival duration corresponding to the second processing data is greater than the third threshold. If so, determine the second processing data as the third data; determine data other than multiple first data and multiple third data as multiple second data.

[0131] Among them, the third threshold is determined according to the load rate of the second memory.

[0132] Optionally, the third threshold can be determined by the following formula:

[0133] AGE_MAX = 5000 / (1 + IOPS / 10000)

[0134] Among them, AGE_MAX can represent the third threshold, and IOPS can represent the load rate.

[0135] In this way, when the memory load rate is higher, the third data retention time is shorter, so as to improve the data management efficiency.

[0136] For the implementation content of each step in the embodiments of the present application, reference may be made to the description of Embodiment 1 above, and repeated content will not be elaborated.

[0137] A data management method provided in this embodiment includes: obtaining access logs within a preset time period; performing feature extraction processing on the access logs to obtain multi-dimensional features; determining multiple first data from multiple data according to the multi-dimensional features and an access prediction model; determining multiple second data and multiple third data according to the multi-dimensional features and the multiple first data, where the predicted access probability corresponding to the second data is less than the predicted access probability corresponding to the first data and greater than the predicted access probability corresponding to the third data. In this way, through intelligent data classification, storage and management by the controller, while meeting the storage capacity, the access latency of data is reduced, and the data management efficiency is improved.

[0138] In a possible implementation manner, perform at least one incremental training step on the access prediction model, and determine whether the access prediction model converges; if so, determine the converged access prediction model as the optimized access prediction model; if not, repeat the incremental training step until the access prediction model converges.

[0139] Among them, at least one incremental training step can be performed on the access prediction model at an interval of the first time period.

[0140] The convergence condition can be a pre-set condition.

[0141] For example, the convergence condition is that the accuracy rate of the access prediction model reaches a threshold.

[0142] Next, in combination with Figure 4 , the process of the incremental training step will be explained.

[0143] Figure 4 It is a schematic flow chart of a model incremental training method provided in an embodiment of the present application. On the basis of the above embodiment, refer to Figure 4 , this method includes:

[0144] S401. Determine multiple training data according to the access logs.

[0145] The training data includes multiple-dimensional features and data labels within a preset time period.

[0146] Data labels can represent the actual access situations of each training data.

[0147] Among them, multiple training data can be divided into a training set, a validation set, and a test set according to a ratio of 8:1:1.

[0148] Feature extraction processing can be performed on the access logs to determine multiple training data.

[0149] S402: Input multiple training data into the access prediction model, perform forward propagation, and obtain multiple model output values.

[0150] Among them, the multiple model output values are multiple predicted data output by the access prediction model.

[0151] The network structure of the access prediction model can include:

[0152] ① Input layer: Input multiple-dimensional features and a preset duration.

[0153] ② First long short-term memory layer: Consisting of 128 units, with the activation function being the hyperbolic tangent function, and returning the complete sequence.

[0154] ③ Regularization layer: The dropout rate is 0.2 to prevent overfitting.

[0155] ④ Second long short-term memory layer: Consisting of 128 units, with the activation function being the hyperbolic tangent function, and returning the final output.

[0156] ⑤ Fully connected layer: The output dimension is N, where N is the N data that may be accessed within the preset future duration for prediction.

[0157] S403: Determine the loss value of the loss function based on the multiple model output values and multiple data labels, and determine the gradient value of the access prediction model based on the loss value.

[0158] Among them, the loss function can be a classification task function and a regression task function.

[0159] For example, the classification task function can be the cross-entropy loss function, and the regression task function can be the mean squared error.

[0160] The training data can be input into the model to obtain the output value of the model. According to the selected loss function, substitute the model output value and the data label into the loss function to calculate the loss value. Based on the loss value, calculate the gradient value of the loss function with respect to each parameter through the backpropagation algorithm.

[0161] S404: Determine the updated parameters of the access prediction model through the optimizer based on the gradient value, and update the access prediction model according to the updated parameters.

[0162] Among them, the preset learning rate of the optimizer can be 0.001, and the batch size can be 256, which is not limited here.

[0163] The optimizer can calculate the update value of each parameter according to the gradient value and the preset learning rate, and update the access prediction model according to the updated parameters.

[0164] The access prediction model can adopt Elastic Weight Consolidation (EWC) to prevent catastrophic forgetting.

[0165] For the implementation content of each step in the embodiments of the present application, reference may be made to the description of the corresponding steps or operations in the above method embodiments, and repeated content will not be elaborated.

[0166] A model incremental training method provided in this embodiment determines whether the access prediction model converges by performing at least one incremental training step on the access prediction model; if so, determines the converged access prediction model as the optimized access prediction model; if not, repeats the incremental training step until the access prediction model converges; wherein, the incremental training step includes: determining a plurality of training data according to the access log, and the training data includes a plurality of dimensional features and data labels within a preset time period; inputting the plurality of training data into the access prediction model for forward propagation to obtain a plurality of model output values; determining the loss value of the loss function according to the plurality of model output values and the plurality of data labels, and determining the gradient value of the access prediction model according to the loss value; determining the update parameters of the access prediction model through the optimizer according to the gradient value, and updating the access prediction model according to the update parameters. In this way, the access prediction model can continuously optimize its own parameters according to the latest access log data, thereby improving the prediction accuracy.

[0167] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.

[0168] Figure 5 It is a schematic structural diagram of a data management device provided in an embodiment of the present application. Please refer to Figure 5 The data management device 500 includes an acquisition module 501, an extraction module 502, a first determination module 503, and a second determination module 504, where

[0169] The acquisition module 501 is configured to acquire the access log within a preset time period;

[0170] The extraction module 502 is configured to perform feature extraction processing on the access log to obtain a plurality of dimensional features;

[0171] The first determination module 503 is configured to determine a plurality of first data from a plurality of data according to a plurality of dimensional features and an access prediction model;

[0172] The second determination module 504 is configured to determine a plurality of second data and a plurality of third data according to a plurality of dimensional features and a plurality of first data, wherein the predicted access probability corresponding to the second data is less than the predicted access probability corresponding to the first data and greater than the predicted access probability corresponding to the third data.

[0173] In a possible implementation manner, the plurality of dimensional features include access frequency and processor utilization rate, and the first determination module 503 is specifically configured to:

[0174] Input the plurality of dimensional features into the access prediction model to obtain a plurality of prediction data, where the prediction data are data with an access probability greater than a first threshold within a preset time period;

[0175] Determine the plurality of prediction data as the plurality of first data, and determine the plurality of data other than the plurality of prediction data as the plurality of first processing data;

[0176] For any one of the first processing data, determine whether the access frequency corresponding to the first processing data is greater than a second threshold. If so, determine the first processing data as the first data, and the second threshold is determined according to the processor utilization rate.

[0177] In a possible implementation manner, the apparatus further includes a training module 505, and the training module 505 is configured to:

[0178] Perform at least one incremental training step on the access prediction model, and determine whether the access prediction model converges;

[0179] If so, determine the converged access prediction model as the optimized access prediction model;

[0180] If not, repeat the incremental training step until the access prediction model converges;

[0181] Wherein, the incremental training step includes:

[0182] Determine a plurality of training data according to the access log, where the training data includes a plurality of dimensional features and data labels within a preset time period;

[0183] Input the plurality of training data into the access prediction model for forward propagation to obtain a plurality of model output values;

[0184] Determine the loss value of the loss function according to the plurality of model output values and the plurality of data labels, and determine the gradient value of the access prediction model according to the loss value;

[0185] Based on the gradient value, the optimizer determines the update parameters for accessing the prediction model, and updates the access prediction model according to the update parameters.

[0186] In a possible implementation manner, the multiple dimensional features further include the survival duration, and the second determination module 504 is specifically configured to:

[0187] Among the multiple data, the data other than the multiple first data is determined as the multiple second processed data;

[0188] For any one of the second processed data, it is judged whether the survival duration corresponding to the second processed data is greater than a third threshold. If so, the second processed data is determined as the third data, and the third threshold is determined according to the load rate of the second memory;

[0189] The data other than the multiple first data and the multiple third data is determined as the multiple second data.

[0190] For the description of the features in the embodiments corresponding to the data management device, reference can be made to the relevant descriptions in the embodiments corresponding to the data management method, which will not be elaborated here one by one.

[0191] Figure 6 This is a schematic structural diagram of the electronic device provided by this application. As Figure 6 shown, the electronic device 60 provided in this embodiment includes: at least one processor 601 and a memory 602. Optionally, the electronic device 60 further includes a communication component 603. Among them, the processor 601, the memory 602, and the communication component 603 are connected through a bus.

[0192] In a specific implementation process, at least one processor 601 executes the computer execution instructions stored in the memory 602, so that at least one processor 601 executes the above-mentioned data management method embodiment.

[0193] For the specific implementation process of the processor 601, reference can be made to the above-mentioned method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0194] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0195] The memory may include random access memory (RAM), or may also include non-volatile memory (NVM), such as at least one disk memory.

[0196] The bus may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the buses in the drawings of the present application are not limited to only one bus or one type of bus.

[0197] The embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any of the above embodiments of the data management method when running.

[0198] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs, etc., which are various media that can store computer programs.

[0199] The embodiments of the present application also provide a computer program product, the above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the data management method are implemented.

[0200] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program which, when executed by a processor, implements the steps in any of the above-described embodiments of the data management method.

[0201] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0202] The above has introduced in detail a data management method provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A data manager, characterized in that, The data manager includes a processor and a controller, and the controller is respectively connected to a first memory and a second memory; The processor is configured to determine, from multiple data, multiple first data, multiple second data, and multiple third data according to an access log within a preset time period through a preset algorithm, where the predicted access probability corresponding to the second data is less than the predicted access probability corresponding to the first data and greater than the predicted access probability corresponding to the third data, and the multiple data are the data stored in the first memory and the second memory; The controller is configured to store the multiple first data in the first memory, store the multiple second data in the buffer area of the second memory, perform data compression processing on the third data, store the compressed multiple third data in the compressed storage area of the second memory, and update the address translation relationship.

2. The data manager according to claim 1, wherein, The preset algorithm includes an access prediction model, and the processor is specifically configured to: Perform feature extraction processing on the access log to obtain multiple dimensional features, where the multiple dimensional features include the number of accesses and the survival duration; Determine multiple first data according to the number of accesses and the access prediction model; Determine the multiple third data according to the survival duration; Determine the multiple second data according to the multiple first data and the multiple third data.

3. The data manager according to claim 1 or 2, characterized in that, The processor is further configured to send an operation instruction to the controller, and the operation instruction includes a virtual address; The controller is further configured to determine a physical address corresponding to the virtual address according to the address translation relationship, and transmit the data corresponding to the physical address to the processor through a preset transmission mode.

4. The data manager according to claim 1 or 2, characterized in that The controller is specifically configured to: Perform segmentation processing on the multiple third data to obtain multiple data blocks; Perform dictionary encoding processing and compression processing on the multiple data blocks through a target compression algorithm to obtain multiple compressed data blocks; Determine the multiple compressed data blocks as the compressed multiple third data.

5. A data management method, characterized in that, Including: Obtain an access log within a preset time period; Perform feature extraction processing on the access log to obtain multiple dimensional features; Determine multiple first data from multiple data according to the multiple dimensional features and the access prediction model; Determine multiple second data and multiple third data according to the multiple dimensional features and the multiple first data, where the predicted access probability corresponding to the second data is less than the predicted access probability corresponding to the first data and greater than the predicted access probability corresponding to the third data.

6. The method according to claim 5, wherein The multiple dimensional features include the access frequency and the processor utilization rate. Determining multiple first data from multiple data according to the multiple dimensional features and the access prediction model includes: Input the multiple dimensional features into the access prediction model to obtain multiple prediction data, where the prediction data are the data with an access probability greater than a first threshold within a preset time duration; Determine the multiple prediction data as multiple first data, and determine multiple data other than the multiple prediction data as multiple first processing data; For any first processed data, determine whether the access frequency corresponding to the first processed data is greater than a second threshold. If so, determine the first processed data as first data, where the second threshold is determined according to the processor utilization rate.

7. The method according to claim 6, wherein The method further includes: Performing at least one incremental training step on the access prediction model, and determining whether the access prediction model converges; If so, determine the converged access prediction model as the optimized access prediction model; If not, repeat the incremental training step until the access prediction model converges; Wherein, the incremental training step includes: Determining a plurality of training data according to the access log, where the training data includes a plurality of dimensional features and data labels within a preset time period; Inputting the plurality of training data into the access prediction model for forward propagation to obtain a plurality of model output values; Determining the loss value of the loss function according to the plurality of model output values and the plurality of data labels, and determining the gradient value of the access prediction model according to the loss value; Determining the update parameters of the access prediction model according to the gradient value through an optimizer, and updating the access prediction model according to the update parameters.

8. The method according to any one of claims 5 to 7, characterized in that, The plurality of dimensional features further includes the survival duration. Determining a plurality of second data and a plurality of third data according to the plurality of dimensional features and the plurality of first data includes: Among the plurality of data, determining the data other than the plurality of first data as a plurality of second processed data; For any one of the second processed data, determining whether the survival duration corresponding to the second processed data is greater than a third threshold. If so, determine the second processed data as third data, where the third threshold is determined according to the load rate of the second memory; Determining the data other than the plurality of first data and the plurality of third data as a plurality of second data.

9. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for implementing the steps of the data management method according to any one of claims 5 to 8 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the data management method according to any one of claims 5 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Report access duration prediction method and device, processor and electronic equipment

    CN117033726A

  • Information recommendation method and device based on portrait system, electronic equipment and medium

    CN117372094A

  • Log data storage method and device, equipment and medium

    CN119311226A

  • Cache management method, electronic device, storage medium and program product

    CN119620961A

  • Data access processing method and storage control device

    JP2008293111A