Data placement method, system and equipment and computer readable storage medium
By generating hot and cold attribute scores through machine learning, the data placement of solid-state drives is optimized, solving the problems of performance degradation and shortened lifespan caused by write amplification, and achieving performance improvement and life extension.
Patent Information
- Application Number
- CN202510898626.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-19
AI Technical Summary
In the prior art, the write amplification factor of solid-state drives leads to performance degradation and shortened lifespan, and flexible data placement methods are complex and prone to improper operation, resulting in performance and lifespan degradation.
Through machine learning, a hot and cold data classifier is built to generate target hot and cold attribute scores based on data size, usage frequency, recent usage time, and usage cycle, and the data is placed in the corresponding file system directory to reduce unnecessary operations and lower the write amplification effect.
Improves the access performance of solid-state drives, extends their lifespan, simplifies user operations, and reduces write amplification effects.
Smart Images

Figure CN120669925A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of storage technology, and more specifically, to a data placement method, system, device, and computer-readable storage medium. Background Art
[0002] With the development of cloud computing, big data, and artificial intelligence, data centers and high-performance computing (HPC) are increasingly demanding storage performance. However, these scenarios are often accompanied by high write loads, which places higher demands on the durability and performance of SSDs (Solid State Disks). Write amplification is also a major challenge for SSDs. Due to the characteristics of NAND flash memory, SSDs require additional internal operations such as garbage collection (GC) and data migration when writing data, resulting in the actual amount of data written being far greater than the amount requested by the user.
[0003] To improve SSD performance and lifespan, Flexible Data Placement (FDP) is an interface in the NVM Express (NVMe) storage standard designed to reduce the write amplification factor (WAF) of SSDs through user-controlled data placement. Embedded Software (ES) operators can use the FDP interface to explicitly control data placement on FDP-enabled SSDs to minimize garbage collection (GC) overhead, improve performance, and extend SSD lifespan. However, FDP requires operators to understand and perform complex and tedious system programming, which can lead to issues such as decreased SSD access performance and lifespan due to improper operation.
[0004] In summary, how to improve the access performance and extend the life of solid-state drives is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of this application is to provide a data placement method that can, to a certain extent, solve the technical problem of how to improve the access performance and extend the life of solid-state drives. This application also provides a data placement system, an electronic device, and a computer-readable storage medium.
[0006] In order to achieve the above objectives, this application provides the following technical solutions:
[0007] A data placement method, comprising:
[0008] Get the target data to be placed;
[0009] Determine the data size, usage frequency, most recent usage time, and usage period of the target data;
[0010] Obtain a pre-trained hot and cold data classifier, wherein the hot and cold data classifier is built based on machine learning;
[0011] Processing the data size, the usage frequency, the most recent usage time, and the usage period by the hot and cold data classifier to generate a target hot and cold attribute score for the target data;
[0012] Obtaining a file system directory set in the solid state drive, wherein the file system directory is generated based on flexible data placement;
[0013] The target data is placed in a file system directory corresponding to the target hot / cold attribute score.
[0014] In an exemplary embodiment, the processing of the data size, the usage frequency, the most recent usage time, and the usage period by the hot and cold data classifier to generate a target hot and cold attribute score of the target data includes:
[0015] According to the generation condition that the hot and cold attribute scores are positively correlated with the data size, the usage frequency, and the most recent usage time and negatively correlated with the usage cycle, the data size, the usage frequency, the most recent usage time, and the usage cycle are processed by the hot and cold data classifier to generate the target hot and cold attribute scores of the target data.
[0016] In an exemplary embodiment, the processing of the data size, the usage frequency, the most recent usage time, and the usage period by the hot and cold data classifier to generate a target hot and cold attribute score of the target data includes:
[0017] inputting the data size, the usage frequency, the most recent usage time, and the usage period into each hot and cold decision tree in the hot and cold data classifier;
[0018] receiving an initial hot and cold attribute score output by the hot and cold decision tree;
[0019] The initial cold and hot attribute scores are processed to generate target cold and hot attribute scores of the target data.
[0020] In an exemplary embodiment, receiving the initial hot and cold attribute scores output by the hot and cold decision tree includes:
[0021] receiving an initial hot and cold attribute score output by the hot and cold decision tree according to a hot and cold attribute score generation formula;
[0022] The formula for generating the cold and hot attribute scores includes:
[0023] ;
[0024] f(S)=log(S+1); f(F)=(FF min ) / (F max -F min );
[0025] f(P)=1 / P; f(R)=(RR min ) / (R max -R min );
[0026] Wherein, Y represents the initial hot and cold attribute score; w1, w2, w3, and w4 represent the set weights; f(S) represents the quantization result of the data size S; f(F) represents the quantization result of the usage frequency F; f(P) represents the quantization result of the usage period P; f(R) represents the quantization result of the most recent usage time R; log represents logarithmic operation; F max Indicates the set maximum frequency of use; F min Indicates the set minimum frequency of use; R max Indicates the earliest usage time set; R min Indicates the latest usage time set.
[0027] In an exemplary embodiment, processing the initial hot and cold attribute scores to generate target hot and cold attribute scores of the target data includes:
[0028] Get the set number of clusters;
[0029] Clustering the initial hot and cold attribute scores according to the number of clusters to obtain a clustering result;
[0030] Counting the cluster center of each cluster in the clustering result;
[0031] Counting the number of members in each cluster in the clustering result;
[0032] Generate a ratio of the number of members of a cluster to the total number of the initial hot and cold attribute scores, and use the ratio as a weight value of the cluster;
[0033] The cluster centers are weighted averaged according to the weight values of the clusters to generate the target hot and cold attribute scores of the target data.
[0034] In an exemplary embodiment, placing the target data into a file system directory corresponding to the target hot / cold attribute score includes:
[0035] Count the frequency of users' operations on file system directories;
[0036] determining hot and cold labels of the file system directories based on the operation frequency;
[0037] The target data is placed in a file system directory whose hot / cold label corresponds to the target hot / cold attribute score.
[0038] In an exemplary embodiment, placing the target data into a file system directory corresponding to the target hot / cold attribute score includes:
[0039] Determine a file system directory corresponding to the target hot / cold attribute score as a candidate directory;
[0040] Prompting the candidate directory;
[0041] Check whether a forced placement instruction is received;
[0042] In response to not receiving a forced placement instruction, placing the target data into the candidate directory.
[0043] A data placement system comprising:
[0044] A data acquisition module, used to acquire target data to be placed;
[0045] A parameter determination module, configured to determine the data size, usage frequency, most recent usage time, and usage period of the target data;
[0046] A classifier acquisition module is used to obtain a pre-trained hot and cold data classifier, wherein the hot and cold data classifier is built based on machine learning;
[0047] A score generating module, configured to process the data size, the usage frequency, the most recent usage time, and the usage period by the hot and cold data classifier to generate a target hot and cold attribute score for the target data;
[0048] A directory acquisition module, configured to acquire a file system directory set in the solid-state drive, wherein the file system directory is generated based on flexible data placement;
[0049] A data placement module is configured to place the target data into a file system directory corresponding to the target hot / cold attribute score.
[0050] An electronic device, comprising:
[0051] memory for storing computer programs;
[0052] A processor is configured to implement the steps of any of the above-mentioned data placement methods when executing the computer program.
[0053] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-mentioned data placement methods.
[0054] The present application provides a data placement method, which includes obtaining target data to be placed; determining the data size, usage frequency, recent usage time, and usage cycle of the target data; obtaining a pre-trained hot and cold data classifier, where the hot and cold data classifier is built based on machine learning; processing the data size, usage frequency, recent usage time, and usage cycle through the hot and cold data classifier to generate target hot and cold attribute scores for the target data; obtaining a file system directory set in a solid-state drive, where the file system directory is generated based on flexible data placement; and placing the target data in the file system directory corresponding to the target hot and cold attribute scores. In this application, a hot and cold data classifier built based on machine learning is used to comprehensively generate target hot and cold attribute scores for target data based on data size, usage frequency, recent usage time, and usage cycle. On the one hand, this enhances the reference factors for generating target hot and cold attribute scores. On the other hand, since the hot and cold data classifier built based on machine learning has strong generalization capabilities and can effectively reduce overfitting, the generation accuracy of the target hot and cold attribute scores is higher. In this way, if the target data is subsequently placed in a file system directory corresponding to the target hot and cold attribute scores, the hot and cold attributes can be used to more accurately place the target data in the adapted file system directory, reducing access latency to improve the access performance of the solid-state drive, and then reducing unnecessary operations on the solid-state drive, reducing the write amplification effect to extend its life. The data placement system, electronic device, and computer-readable storage medium provided in this application also solve corresponding technical problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0056] Figure 1 A flowchart of a data placement method provided in an embodiment of the present application;
[0057] Figure 2 This is a diagram of the training of hot and cold decision trees;
[0058] Figure 3 This is a schematic diagram of the FDP SSD interface;
[0059] Figure 4 This is a schematic diagram of the structure of the nvme general interface io_uring;
[0060] Figure 5 This is a structural diagram of the FUSE architecture;
[0061] Figure 6 This is a schematic diagram of the structure of the FDP file system built based on this application solution;
[0062] Figure 7 This is an application diagram of the FDP file system built based on this application solution.
[0063] Figure 8 A schematic diagram of the structure of a data placement system provided in an embodiment of the present application;
[0064] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0065] Figure 10 Another structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0066] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0067] See also Figure 1 , Figure 1 A flowchart of a data placement method provided in an embodiment of the present application.
[0068] A data placement method provided in an embodiment of the present application may include the following steps:
[0069] Step S101: Acquire target data to be placed.
[0070] In actual applications, the target data to be placed can be obtained first. The type and amount of the target data to be placed can be flexibly determined according to the application scenario. For example, the target data can be a server log file, fault information, etc.
[0071] Step S102: Determine the data size, usage frequency, most recent usage time, and usage period of the target data.
[0072] In practical applications, the data size of the target data can determine whether it is hot or cold data to a certain extent; the usage frequency indicates the number of accesses to the same data within a period of time, which can indicate the hot and cold characteristics of requests in the past period of time. For example, a small number of accesses to the data means that the data is cold; otherwise, the data is hot data. Hot requests with high access frequency indicate that the data may be accessed multiple times. The recent usage time can also determine whether the data is hot or cold data to a certain extent. Therefore, the data size, usage frequency, and recent usage time of the target data can be used to evaluate whether the target data is hot or cold data. However, frequency cannot well reflect temporal locality. Recency only provides information about recent accesses. The smaller the period, the smaller the average interval between each access request, while the larger the frequency and the longer the period, the larger the average interval between each access request. The opposite is true for the smaller the frequency. Therefore, the data size, usage frequency, recent usage time, and usage period of the target data can be determined, so that the target data can be subsequently evaluated as hot or cold data based on the data size, usage frequency, recent usage time, and usage period of the target data.
[0073] Step S103: Obtain a pre-trained hot and cold data classifier, which is built based on machine learning.
[0074] Step S104: Process the data size, usage frequency, most recent usage time, and usage cycle through the hot and cold data classifier to generate a target hot and cold attribute score for the target data.
[0075] In actual applications, this application pre-builds a hot and cold data classifier based on machine learning for hot and cold evaluation of data. Therefore, after determining the data size, usage frequency, recent usage time and usage cycle of the target data, the hot and cold data classifier can be used to process the data size, usage frequency, recent usage time and usage cycle to generate a target hot and cold attribute score for the target data, so as to reflect the hot and cold degree of the target data with the help of the target hot and cold attribute score. For example, when the target hot and cold attribute score is greater than or equal to the set value, the target data can be determined to be hot data, and when the target hot and cold attribute score is less than the set value, the target data can be determined to be cold data, etc.
[0076] In an exemplary embodiment, considering that the hot and cold attributes are positively correlated with data size, usage frequency and recent usage time, and negatively correlated with usage cycle, in the process of processing data size, usage frequency, recent usage time and usage cycle by a hot and cold data classifier to generate target hot and cold attribute scores of target data, the target hot and cold attribute scores of target data can be generated by processing data size, usage frequency, recent usage time and usage cycle by a hot and cold data classifier according to the generation condition that the hot and cold attribute scores are positively correlated with data size, usage frequency and recent usage time, and negatively correlated with usage cycle.
[0077] In an exemplary embodiment, the type of machine learning can be flexibly determined according to the application scenario. For example, in the process of processing the data size, usage frequency, recent usage time and usage cycle through the hot and cold data classifier to generate the target hot and cold attribute scores of the target data, the data size, usage frequency, recent usage time and usage cycle can be input into each hot and cold decision tree in the hot and cold data classifier; the initial hot and cold attribute scores output by the hot and cold decision tree are received; the initial hot and cold attribute scores are processed to generate the target hot and cold attribute scores of the target data. In other words, the target hot and cold attribute scores can be generated with the help of a decision tree. It should be noted that the training process of the hot and cold decision tree can be carried out in combination with the training principle through data size, usage frequency, recent usage time and usage cycle. For example, the training process can be as follows. Figure 2 shown.
[0078] In an exemplary embodiment, in the process of receiving the initial hot and cold attribute scores output by the hot and cold decision tree, the initial hot and cold attribute scores output by the hot and cold decision tree according to the hot and cold attribute score generation formula may be received; the hot and cold attribute score generation formula includes:
[0079] ;
[0080] f(S)=log(S+1); f(F)=(FF min ) / (F max -F min );
[0081] f(P)=1 / P; f(R)=(RR min ) / (R max -R min );
[0082] Where Y represents the initial hot and cold attribute score; w1, w2, w3, and w4 represent the set weights; f(S) represents the quantization result of the data size S; f(F) represents the quantization result of the usage frequency F; f(P) represents the quantization result of the usage period P; f(R) represents the quantization result of the most recent usage time R; log represents the logarithmic operation; F max Indicates the set maximum frequency of use; F min Indicates the set minimum frequency of use; R max Indicates the earliest usage time set; R min Indicates the latest usage time. The earliest usage time and the latest usage time can be obtained by counting the usage time within the set time period.
[0083] In an exemplary embodiment, the initial hot and cold attribute scores are processed to generate the target hot and cold attribute scores of the target data. A set number of clusters can be obtained, for example, the number of clusters can be 3, 5, etc.; the initial hot and cold attribute scores are clustered according to the number of clusters to obtain a clustering result. If the number of clusters is 3, a clustering result with 3 clusters is obtained. If the number of clusters is 5, a clustering result with 5 clusters is obtained, and the cluster center and cluster members of each cluster are different; the cluster center of each cluster in the clustering result is counted; the number of members of each cluster in the clustering result is counted; the ratio of the number of cluster members to the total number of initial hot and cold attribute scores is generated, and the ratio is used as the weight value of the cluster. Assuming that there are 10 members in a cluster and the total number of initial hot and cold attribute scores is 50, the weight value of the cluster is 0.2; the cluster centers are weighted averaged according to the cluster weight values to generate the target hot and cold attribute scores of the target data.
[0084] In this way, each hot and cold decision tree can be used to process the data size, usage frequency, recent usage time and usage cycle to obtain the initial hot and cold attribute scores, so as to fully analyze the initial hot and cold attribute scores of the target data through multiple decision trees; clustering can be used to analyze the distribution of the initial hot and cold attribute scores, so that the fluctuation of the hot and cold attributes can be analyzed according to the distribution, and the fluctuation of the hot and cold attributes can be quantified by the cluster center and cluster weight. Finally, the cluster center can be weighted averaged according to the cluster weight value to generate the target hot and cold attribute score of the target data, that is, the fluctuation of the hot and cold attributes can be comprehensively analyzed to determine the target hot and cold attribute score, which expands the considerations for generating the target hot and cold attribute score, improves the generation accuracy of the target hot and cold attributes, and thus can better place the target data.
[0085] Step S105: Acquire a file system directory set in the solid state drive, where the file system directory is generated based on flexible data placement.
[0086] Step S106: placing the target data into a file system directory corresponding to the target hot / cold attribute score.
[0087] In actual applications, this application abstracts the solid-state drive into different file system directories based on flexible data placement so that users can flexibly operate the solid-state drive. In this process, the user's operation on the file system directory will cause the file system directory to carry hot and cold attributes, so the file system directory set in the solid-state drive can be obtained, and the target data can be placed in the file system directory corresponding to the target hot and cold attribute score.
[0088] In an exemplary embodiment, in the process of placing target data into a file system directory corresponding to a target hot or cold attribute score, the frequency of user operations on the file system directory can be counted; based on the operation frequency, the hot and cold labels of the file system directory are determined, for example, if the frequency of user operations on the file system directory is greater than the average operation frequency, the hot and cold labels of the file system directory are determined to be hot labels; for another example, if the frequency of user operations on the file system directory is less than or equal to the average operation frequency, the hot and cold labels of the file system directory are determined to be cold labels, etc.; the target data is placed into a file system directory corresponding to the hot and cold labels and the target hot and cold attribute scores, for example, if the target hot and cold attribute scores indicate that the target data is hot data, the target data is placed into a file system directory with a hot label; for another example, if the target hot and cold attribute scores indicate that the target data is cold data, the target data is placed into a file system directory with a cold label.
[0089] In an exemplary embodiment, during the process of placing target data into a file system directory corresponding to a target hot or cold attribute score, a user may need to force the target data into a certain file system directory due to their own operations. To meet this requirement of the user, the file system directory corresponding to the target hot or cold attribute score may be determined as a candidate directory; the candidate directory may be prompted; whether a forced placement instruction has been received may be detected; in response to not receiving the forced placement instruction, the target data may be placed into the candidate directory; correspondingly, in response to receiving the forced placement instruction, the target data may be placed into the file system directory corresponding to the forced placement instruction.
[0090] It should be noted that flexible data placement uses data placement instructions to specify the location where data is placed, thereby optimizing data storage. Figure 3 It is an FDP SSD interface, where each EG (Endurance Group) represents the FDP configuration, multiple RUs (reclaim units) form an RG (reclaim Group), multiple RGs form a RUH (reclaim unit handle), and multiple RUHs form an EG. Users can write data to different RUHs by controlling the data placement instructions, so that different data can be allocated to different blocks, thereby reducing write amplification and improving the overall SSD performance and lifespan. Currently, systems using the FDP interface, such as Figure 4 As shown in the figure, the specific nvme general interface io_uring is used to communicate data between the application and the kernel through asynchronous operations. User space file system (FUSE) allows developers to write custom file systems as user programs, such as Figure 5As shown in the FUSE core architecture, when a program interacts with a mounted FUSE (file system in user space), the operating system's request is first passed to the FUSE kernel module. This module encapsulates the request into a structure and places it in a queue for processing. During this time, the program may pause execution and enter a wait state. Meanwhile, the FUSE daemon running in user space is awakened. The daemon retrieves the request from the kernel queue by reading the special device file / dev / fuse and processes it according to the file system's logic. During processing, the request may be passed to the underlying file system or other kernel subsystems. When the FUSE daemon completes request processing, it writes the result back to the kernel via / dev / fuse. After receiving the response, the kernel module returns the result to the requesting program, completing the entire interaction process.
[0091] The present application provides a data placement method, which includes obtaining target data to be placed; determining the data size, usage frequency, recent usage time, and usage cycle of the target data; obtaining a pre-trained hot and cold data classifier, where the hot and cold data classifier is built based on machine learning; processing the data size, usage frequency, recent usage time, and usage cycle through the hot and cold data classifier to generate target hot and cold attribute scores for the target data; obtaining a file system directory set in a solid-state drive, where the file system directory is generated based on flexible data placement; and placing the target data in the file system directory corresponding to the target hot and cold attribute scores. In this application, a hot and cold data classifier built based on machine learning is used to comprehensively generate target hot and cold attribute scores for target data based on data size, usage frequency, recent usage time and usage cycle. On the one hand, the reference factors for generating target hot and cold attribute scores are enhanced. On the other hand, since the hot and cold data classifier built based on machine learning has strong generalization ability and can effectively reduce overfitting, the generation accuracy of the target hot and cold attribute scores is higher. In this way, if the target data is subsequently placed in a file system directory corresponding to the target hot and cold attribute score, the hot and cold attributes can be used to more accurately place the target data in the adapted file system directory, thereby reducing access delay to improve the access performance of the solid-state drive, and then reducing unnecessary operations on the solid-state drive, reducing the write amplification effect and extending its life.
[0092] In order to facilitate the understanding of the data placement solution provided by this application, the present application solution is now described in conjunction with flexible data placement. The FDP file system built based on this application solution can be as follows Figure 6 As shown, the application can be Figure 7As shown in the figure, when mounting a file system on a specific folder, the file system first obtains placement identifier (PID) information from an FDP-enabled device. These PIDs indicate the optimal storage location for data. After obtaining the device information, FDPFS maps each PID to a separate directory and exposes it to user space. Users can directly manage data through these directories, leveraging FDP technology to optimize data storage placement and improve file system performance and efficiency. For each mounted directory, a thread is started to listen to the created I / O message queue. Each thread sends nvme_uring_cmd commands to the device. Therefore, when the file system receives read or write requests based on the directory, it transmits I / O requests for the corresponding PID to the FDP SSD. Each PID corresponds to a separate directory, with directory names starting at 0 and incrementing sequentially until the total number of PIDs is reached. This design allows users to intuitively select storage locations based on data characteristics, fully leveraging the advantages of FDP technology. By abstracting PIDs into file system directories, users can achieve explicit data placement on FDP-enabled SSDs without having to deal with complex low-level operations such as pipeline ring buffer management. This abstraction not only simplifies user operations, but also significantly improves the flexibility and efficiency of data management.
[0093] On this basis, the data placement process of the FDP file system can be as follows:
[0094] Obtain placement identifier information from the SSD and map each placement identifier information to an independent file system directory;
[0095] Start a thread to listen to the created IO message queue to obtain the target data to be placed;
[0096] Determine the data size, usage frequency, most recent usage time, and usage period of the target data;
[0097] Input data size, usage frequency, recent usage time, and usage period into each hot and cold decision tree in the hot and cold data classifier;
[0098] receiving the initial hot and cold attribute scores output by the hot and cold decision tree according to the hot and cold attribute score generation formula;
[0099] Get the set number of clusters;
[0100] Cluster the initial hot and cold attribute scores according to the number of clusters to obtain the clustering results;
[0101] Calculate the cluster center of each cluster in the clustering results;
[0102] Count the number of members in each cluster in the clustering results;
[0103] Generate the ratio of the number of cluster members to the total number of initial hot and cold attribute scores, and use the ratio as the cluster weight value;
[0104] Perform weighted averaging of cluster centers according to cluster weight values to generate target hot and cold attribute scores of target data;
[0105] Count the frequency of users' operations on file system directories;
[0106] Determine hot and cold labels of file system directories based on operation frequency;
[0107] Determine the file system directory corresponding to the hot / cold label and the target hot / cold attribute score as the candidate directory;
[0108] Prompt candidate directories;
[0109] Check whether a forced placement instruction is received;
[0110] In response to not receiving a forced placement instruction, the target data is placed in a candidate directory.
[0111] See also Figure 8 , Figure 8 A schematic diagram of the structure of a data placement system provided in an embodiment of the present application.
[0112] An embodiment of the present application provides a data placement system, which may include:
[0113] The data acquisition module 101 is used to acquire target data to be placed;
[0114] Parameter determination module 102, for determining the data size, usage frequency, most recent usage time and usage period of target data;
[0115] The classifier acquisition module 103 is used to obtain a pre-trained hot and cold data classifier, which is built based on machine learning;
[0116] The score generating module 104 is used to process the data size, usage frequency, recent usage time and usage cycle through the hot and cold data classifier to generate the target hot and cold attribute score of the target data;
[0117] A directory acquisition module 105 is used to acquire a file system directory set in the solid state drive, where the file system directory is generated based on flexible data placement;
[0118] The data placement module 106 is configured to place the target data into a file system directory corresponding to the target hot / cold attribute score.
[0119] In an embodiment of the present application, a data placement system is provided, wherein a score generation module may include:
[0120] The score generation unit is used to process the data size, usage frequency, recent usage time and usage cycle through the hot and cold data classifier according to the generation condition that the hot and cold attribute scores are positively correlated with the data size, usage frequency and recent usage time and negatively correlated with the usage cycle, so as to generate the target hot and cold attribute scores of the target data.
[0121] In a data placement system provided by an embodiment of the present application, a score generation unit can be used to: input data size, usage frequency, recent usage time, and usage cycle into each hot and cold decision tree in a hot and cold data classifier; receive initial hot and cold attribute scores output by the hot and cold decision trees; and process the initial hot and cold attribute scores to generate target hot and cold attribute scores for target data.
[0122] In a data placement system provided in an embodiment of the present application, a score generation unit may be used to:
[0123] receiving the initial hot and cold attribute scores output by the hot and cold decision tree according to the hot and cold attribute score generation formula;
[0124] The formula for generating the hot and cold attribute scores includes:
[0125] ;
[0126] f(S)=log(S+1); f(F)=(F-Fmin) / (Fmax-Fmin);
[0127] f(P)=1 / P; f(R)=(R-Rmin) / (Rmax-Rmin);
[0128] Among them, Y represents the initial hot and cold attribute score; w1, w2, w3, and w4 represent the set weights; f(S) represents the quantization result of the data size S; f(F) represents the quantization result of the usage frequency F; f(P) represents the quantization result of the usage period P; f(R) represents the quantization result of the most recent usage time R; log represents logarithmic operation; Fmax represents the set maximum usage frequency; Fmin represents the set minimum usage frequency; Rmax represents the set earliest usage time; and Rmin represents the set latest usage time.
[0129] An embodiment of the present application provides a data placement system, in which a score generation unit can be used to: obtain a set number of clusters; cluster the initial hot and cold attribute scores according to the number of clusters to obtain a clustering result; count the cluster center of each cluster in the clustering result; count the number of members in each cluster in the clustering result; generate a ratio of the number of cluster members to the total number of initial hot and cold attribute scores, and use the ratio as the weight value of the cluster; and perform weighted averaging of the cluster centers according to the cluster weight values to generate a target hot and cold attribute score for the target data.
[0130] An embodiment of the present application provides a data placement system, wherein the data placement module may include:
[0131] A statistics unit is used to count the frequency of users' operations on file system directories;
[0132] a label determination unit, configured to determine a hot or cold label of a file system directory based on an operation frequency;
[0133] The first data placement unit is configured to place target data into a file system directory corresponding to the hot / cold label and the target hot / cold attribute score.
[0134] An embodiment of the present application provides a data placement system, wherein the data placement module may include:
[0135] a candidate determination unit, configured to determine a file system directory corresponding to a target hot / cold attribute score as a candidate directory;
[0136] A prompting unit, used for prompting candidate directories;
[0137] The second data placement unit is configured to detect whether a forced placement instruction is received; in response to not receiving the forced placement instruction, place the target data into the candidate directory.
[0138] This application also provides an electronic device and a computer-readable storage medium, both of which have the corresponding effects of the data placement method provided in the embodiment of this application. Figure 9 , Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0139] An electronic device provided in an embodiment of the present application includes a memory 201 and a processor 202. The memory 201 stores a computer program, and when the processor 202 executes the computer program, the steps of the data placement method described in any of the above embodiments are implemented.
[0140] See also Figure 10Another electronic device provided in an embodiment of the present application may further include: an input port 203 connected to the processor 202 for transmitting commands inputted from the outside to the processor 202; a display unit 204 connected to the processor 202 for displaying the processing results of the processor 202 to the outside world; and a communication module 205 connected to the processor 202 for enabling communication between the electronic device and the outside world. The display unit 204 may be a display panel, a laser scanning display, etc. The communication method adopted by the communication module 205 includes but is not limited to Mobile High-Definition Link (MHL), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), wireless connection: Wireless Fidelity (WiFi), Bluetooth communication technology, Bluetooth low energy communication technology, and communication technology based on IEEE802.11s.
[0141] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the data placement method described in any of the above embodiments are implemented.
[0142] The computer-readable storage medium involved in this application includes random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs (Compact Disc Read-Only Memory), or any other form of storage medium known in the technical field.
[0143] A computer program product provided in an embodiment of the present application includes a computer program / instruction, which, when executed by a processor, implements the steps of the data placement method described in any of the above embodiments.
[0144] For descriptions of the relevant portions of the data placement system, electronic device, computer program product, and computer-readable storage medium provided in the embodiments of this application, please refer to the detailed description of the corresponding portions in the data placement method provided in the embodiments of this application, and will not be repeated here. Furthermore, portions of the above-mentioned technical solutions provided in the embodiments of this application that are consistent with the implementation principles of corresponding technical solutions in the prior art are not described in detail to avoid redundant description.
[0145] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0146] The above description of the disclosed embodiments will enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data placement method, characterized in that: include: Get the target data to be placed; Determine the data size, usage frequency, most recent usage time, and usage period of the target data; Obtain a pre-trained hot and cold data classifier, wherein the hot and cold data classifier is built based on machine learning; Processing the data size, the usage frequency, the most recent usage time, and the usage period by the hot and cold data classifier to generate a target hot and cold attribute score for the target data; Obtaining a file system directory set in the solid state drive, wherein the file system directory is generated based on flexible data placement; The target data is placed in a file system directory corresponding to the target hot / cold attribute score.
2. The data placement method according to claim 1, characterized in that: The processing of the data size, the usage frequency, the most recent usage time, and the usage period by the hot and cold data classifier to generate a target hot and cold attribute score of the target data includes: According to the generation condition that the hot and cold attribute scores are positively correlated with the data size, the usage frequency, and the most recent usage time and negatively correlated with the usage cycle, the data size, the usage frequency, the most recent usage time, and the usage cycle are processed by the hot and cold data classifier to generate the target hot and cold attribute scores of the target data.
3. The data placement method according to claim 2, wherein: The processing of the data size, the usage frequency, the most recent usage time, and the usage period by the hot and cold data classifier to generate a target hot and cold attribute score of the target data includes: inputting the data size, the usage frequency, the most recent usage time, and the usage period into each hot and cold decision tree in the hot and cold data classifier; receiving an initial hot and cold attribute score output by the hot and cold decision tree; The initial cold and hot attribute scores are processed to generate target cold and hot attribute scores of the target data.
4. The data placement method according to claim 3, characterized in that: The receiving the initial hot and cold attribute scores output by the hot and cold decision tree includes: receiving an initial hot and cold attribute score output by the hot and cold decision tree according to a hot and cold attribute score generation formula; The formula for generating the cold and hot attribute scores includes: ; f(S)=log(S+1);f(F)=(F-F min ) / (F max -F min ); f(P)=1 / P;f(R)=(R-R min ) / (R max -R min ); Wherein, Y represents the initial hot and cold attribute score; w1, w2, w3, and w4 represent the set weights; f(S) represents the quantization result of the data size S; f(F) represents the quantization result of the usage frequency F; f(P) represents the quantization result of the usage period P; f(R) represents the quantization result of the most recent usage time R; log represents logarithmic operation; F max Indicates the set maximum frequency of use; F min Indicates the set minimum frequency of use; R max Indicates the earliest usage time set; R min Indicates the latest usage time set.
5. The data placement method according to claim 2, wherein: The processing of the initial hot and cold attribute scores to generate target hot and cold attribute scores of the target data includes: Get the set number of clusters; Clustering the initial hot and cold attribute scores according to the number of clusters to obtain a clustering result; Counting the cluster center of each cluster in the clustering result; Counting the number of members in each cluster in the clustering result; Generate a ratio of the number of members of a cluster to the total number of the initial hot and cold attribute scores, and use the ratio as a weight value of the cluster; The cluster centers are weighted averaged according to the weight values of the clusters to generate the target hot and cold attribute scores of the target data.
6. The data placement method according to claim 1, wherein: Placing the target data into a file system directory corresponding to the target hot / cold attribute score includes: Count the frequency of users' operations on file system directories; determining hot and cold labels of the file system directories based on the operation frequency; The target data is placed in a file system directory whose hot / cold label corresponds to the target hot / cold attribute score.
7. The data placement method according to claim 1, wherein: Placing the target data into a file system directory corresponding to the target hot / cold attribute score includes: Determine a file system directory corresponding to the target hot / cold attribute score as a candidate directory; Prompting the candidate directory; Check whether a forced placement instruction is received; In response to not receiving a forced placement instruction, placing the target data into the candidate directory.
8. A data placement system, characterized in that: include: A data acquisition module, used to acquire target data to be placed; A parameter determination module, configured to determine the data size, usage frequency, most recent usage time, and usage period of the target data; A classifier acquisition module is used to obtain a pre-trained hot and cold data classifier, wherein the hot and cold data classifier is built based on machine learning; A score generating module, configured to process the data size, the usage frequency, the most recent usage time, and the usage period by the hot and cold data classifier to generate a target hot and cold attribute score for the target data; A directory acquisition module, configured to acquire a file system directory set in the solid-state drive, wherein the file system directory is generated based on flexible data placement; A data placement module is configured to place the target data into a file system directory corresponding to the target hot / cold attribute score.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the data placement method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data placement method according to any one of claims 1 to 7 are implemented.