Data Distribution Method, Device, Electronic Device and Medium Based on Heterogeneous Storage
By adopting a data distribution method based on heterogeneous storage in a distributed file system, the migration probability and cost are calculated based on the device load rate and file characteristics, the problem of unreasonable traditional data distribution is solved, and more efficient data layout and system stability are achieved.
Patent Information
- Application Number
- CN202111192178.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-13
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-10-13
AI Technical Summary
In traditional distributed file systems, due to the failure to reasonably consider the characteristics of the storage device and the file itself, the data distribution is unreasonable, which affects the data access performance.
Through a data distribution method based on heterogeneous storage, the migration probability of file blocks is determined based on the device load rate of the multi-layer storage device, the affinity and migration cost of the target storage device layer are calculated, and the migration operation is triggered when the forward migration benefits are met.
It improves the rationality of data layout, reduces the impact on the system, ensures the normal service of data blocks through decentralized migration operations, and reduces the impact on upper-level applications.
Smart Images

Figure CN114281756B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of heterogeneous storage technologies, and in particular, to a data distribution method, apparatus, electronic device, and storage medium based on heterogeneous storage. Background Art
[0002] Due to the limitations of device bandwidth and storage capacity of traditional distributed file systems that use a single storage medium, they have gradually become a performance bottleneck for upper-layer computing frameworks. Therefore, a distributed heterogeneous storage technology based on multiple storage media has emerged to solve the problems of device bandwidth and storage capacity in traditional distributed file systems.
[0003] Data layout is an important factor affecting data access performance in distributed file systems. However, due to only considering file popularity, that is, the access frequency of files, traditional data locality methods place files with high access popularity on high-layer storage devices with fast transmission speed and low latency. However, due to not considering the characteristics of storage devices and files themselves, the data layout is not reasonable enough.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] An object of the present disclosure is to provide a data distribution method, apparatus, storage medium, and electronic device based on heterogeneous storage, which at least to some extent overcome the problem of unreasonable data distribution in heterogeneous storage devices in related technologies.
[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.
[0007] According to one aspect of the present disclosure, a data distribution method based on heterogeneous storage is provided, including: determining the migration probability of file blocks in each layer of the storage device based on the device load rate of heterogeneous multi-layer storage devices, where the migration probability includes a promotion probability of promoting upward and / or an elimination probability of eliminating downward; calculating the affinity between the target storage device layer and the migrated file blocks and the migration cost of the migrated file blocks respectively based on the migration probability; when a positive migration benefit is determined based on the affinity and the migration cost, triggering the migration operation of the migrated file blocks based on the migration probability.
[0008] In one embodiment, calculating the affinity between the migrated file block and the target storage device layer and the migration cost of the migrated file block respectively based on the migration probability specifically includes: determining the bandwidth of the target storage device layer occupied based on the size of the migrated file block and the read / write operations of the target storage device layer on the migrated file block; predicting the popularity value of the migrated file block; determining the affinity based on the migration probability, the bandwidth, the popularity value, and the identification parameter of the target storage device layer; and determining the migration cost based on the migration probability, the size of the migrated file block, the bandwidth, and the popularity value.
[0009] In one embodiment, determining the affinity based on the migration probability, the bandwidth, the popularity value, and the identification parameter of the target storage device layer specifically includes: calculating the affinity based on a first formula, and the first formula is:
[0010] F = P · BandWidth(size, r / w) · [Popularity() + α · tier(id)],
[0011] where P is the migration probability, BandWidth(size, r / w) is the bandwidth, Popularity() is the popularity value, tier(id) is the identification parameter of the target storage device layer, and α is a first adjustment parameter.
[0012] In one embodiment, determining the migration cost based on the migration probability, the size of the migrated file block, the bandwidth, and the popularity value specifically includes: calculating the migration cost based on a second formula, and the second formula is:
[0013]
[0014] where P is the migration probability, size is the size of the migrated file block, BandWidth(size, r / w) is the bandwidth, Popularity() is the popularity value, β is a second adjustment parameter, and γ is a third adjustment parameter.
[0015] In one embodiment, predicting the popularity value of the migrated file block specifically includes: collecting the access timestamp information of the migrated file block within a preset sliding window; and predicting the popularity value of the migrated file block based on the number of the access timestamp information and the length of the sliding window.
[0016] In one embodiment, determining the migration probability of file blocks in each layer of the storage device based on the device load rate of the heterogeneous multi-layer storage device specifically includes: detecting the IO load rate and storage load of each layer of the storage device, and determining one of the IO load rate and storage load as the device load rate of each layer of the storage device; when designating the storage device as the target storage device layer, calculating the promotion probability of the migration file blocks to be promoted to the target storage device layer based on the device load rate of the designated storage device; calculating the elimination probability of the migration file blocks to be eliminated from the target storage device layer based on the promotion probability.
[0017] In one embodiment, calculating the promotion probability of the migration file blocks to be promoted to the target storage device layer based on the device load rate of the designated storage device specifically includes: calculating the promotion probability based on the third formula, and the third formula is:
[0018]
[0019] where u is the device load rate and K is a custom parameter.
[0020] In one embodiment, when determining that there is a positive migration benefit based on the affinity and the migration cost, triggering the migration operation of the migration file blocks based on the migration probability specifically includes: when detecting that the difference between the affinity and the migration cost is greater than 0, determining that there is the positive migration benefit; randomly generating a migration parameter greater than or equal to 0 and less than or equal to 1, and when the migration parameter is less than the migration probability, performing the migration operation.
[0021] In one embodiment, the migration file blocks include file blocks to be promoted and file blocks to be eliminated. Triggering the migration operation of the migration file blocks based on the migration probability specifically further includes: when triggering the migration operation of the file blocks to be promoted based on the promotion probability, performing a locking operation on the file blocks to be promoted, copying the file blocks to be promoted to the upper target storage device layer, unlocking and deleting the original file blocks to be promoted; when triggering the migration operation of the file blocks to be eliminated based on the elimination probability, performing a locking operation on the file blocks to be eliminated, copying the file blocks to be eliminated to the lower target storage device layer, unlocking and deleting the original file blocks to be eliminated.
[0022] In one embodiment, the heterogeneous multi-layer storage device includes a memory disk, an NVMe hard disk, a solid-state drive, and a mechanical hard disk in sequence from top to bottom.
[0023] According to another aspect of the present disclosure, there is provided a data distribution apparatus based on heterogeneous storage, including: a determination module configured to determine the migration probability of file blocks in each layer of the storage device based on the device load rate of heterogeneous multi-layer storage devices, where the migration probability includes the promotion probability of promoting to the upper layer and / or the elimination probability of eliminating to the lower layer; a calculation module configured to calculate the affinity between the migrated file blocks and the target storage device layer respectively based on the migration probability, and the migration cost of the migrated file blocks; and a migration module configured to trigger the migration operation of the migrated file blocks based on the migration probability when it is determined that there is a positive migration benefit based on the affinity and the migration cost.
[0024] According to still another aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory configured to store executable instructions of the processor; the processor is configured to execute the above-mentioned data distribution method based on heterogeneous storage by executing the executable instructions.
[0025] According to yet another aspect of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the above-mentioned data distribution method based on heterogeneous storage is implemented.
[0026] The data distribution method based on heterogeneous storage provided by the embodiments of the present disclosure determines the probability of promoting the migrated file blocks to the upper layer or eliminating them to the lower layer, that is, the migration probability, by calculating the device load rate of each layer of load devices, that is, the device layer. On the basis of considering the characteristics of the storage device and the characteristics of the file itself, the affinity between the migrated file blocks and the target storage device layer and the migration cost are obtained through the migration probability. Further, when it is determined that the migration operation has a positive migration benefit based on the affinity and the migration cost, that is, the migration operation of the migrated file blocks is triggered based on the migration probability. On the one hand, the rationality of the data layout can be improved. On the other hand, the way of determining whether to migrate based on the migration probability realizes a decentralized migration operation. The decentralized migration operation can reduce the impact on the system. By further detecting whether to perform the migration operation in units of the migrated file blocks, a decentralized migration operation of the data file is realized. The decentralized migration operation is suitable for ensuring the normal service of the data blocks, and thus can reduce the impact on the upper-layer applications.
[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings
[0028] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0029] Figure 1 Schematic diagram of a four-layer structure of a heterogeneous storage device in an embodiment of the present disclosure;
[0030] Figure 2 Flowchart of a data distribution method based on heterogeneous storage in an embodiment of the present disclosure;
[0031] Figure 3 Flowchart of another data distribution method based on heterogeneous storage in an embodiment of the present disclosure;
[0032] Figure 4 Flowchart of yet another data distribution method based on heterogeneous storage in an embodiment of the present disclosure;
[0033] Figure 5 Schematic diagram of a system framework of yet another data distribution scheme based on heterogeneous storage in an embodiment of the present disclosure;
[0034] Figure 6 Flowchart of yet another data distribution method based on heterogeneous storage in an embodiment of the present disclosure;
[0035] Figure 7 Schematic diagram of a data distribution device based on heterogeneous storage in an embodiment of the present disclosure;
[0036] Figure 8 Schematic block diagram of a computer device in an embodiment of the present disclosure. Detailed implementation manners
[0037] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments.
[0038] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0039] The solution provided by the present application, on the one hand, can improve the rationality of data layout. On the other hand, the method of determining whether to migrate based on the migration probability realizes a decentralized migration operation. The decentralized migration operation can reduce the impact on the system. By further detecting whether to perform the migration operation in units of migrated file blocks, a decentralized migration operation of data files is realized. The decentralized migration operation is suitable for ensuring the normal service of data blocks, and thus can reduce the impact on upper-layer applications.
[0040] For the sake of easy understanding, the terms (abbreviations) involved in the present application will be explained first below.
[0041] In order to understand heterogeneous storage policies and Hadoop (distributed storage system) archival storage, the performance characteristics of different storages are compared.
[0042] Hard disk drive (HDD) is the standard disk storage device for Hadoop, which can provide a relatively high throughput and is inexpensive. High throughput is ideal for batch data processing. However, the disk may fail at any time.
[0043] Solid state drive (SSD) provides a measured throughput and I / O per second, but it is several times more expensive than disk storage devices. Like disks, SSD devices have a medium failure rate and may fail at any time.
[0044] Currently, there are roughly two main types of SSD interfaces, namely M.2 and SATA. The SSD with SATA interface executes the AHCI protocol standard, which is a relatively mature and common SSD interface at present. The M.2 interface is divided into the NVMe protocol and the AHCI protocol. According to different protocols, the SSD with M.2 interface will also have some differences in performance. The highest theoretical speed of the NVMe protocol is 32 Gbps.
[0045] RAM-based storage provides extremely high performance for all types of work, but it is very expensive. RAM does not provide persistent storage because everything is stored in memory.
[0046] Next, each step of the data distribution method based on heterogeneous storage in this exemplary embodiment will be described in more detail with reference to the accompanying drawings and embodiments.
[0047] As Figure 1 shown, the heterogeneous multi-layer storage device includes a memory disk 102, an NVMe hard disk 104, a solid-state drive 106, and a mechanical hard disk 108 in sequence from top to bottom.
[0048] Among them, in the four-layer architecture, there are at least the following 6 possible paths for data migration, namely from the memory disk 102 to the NVM 104, from the NVM 104 to the memory 102, from the NVM 104 to the SSD 106, from the SSD 106 to the NVM 104, from the SSD 106 to the mechanical hard disk 108, and from the mechanical hard disk 108 to the SSD 106. There are some differences in migrating data between different storage layers.
[0049] Specifically, when data A is accessed, first check its source and determine whether it is worth promoting. If data A comes from the memory disk 102, there is no need to promote it. If data A comes from the NVM hard disk 104, the promotion process is 2, and with a certain probability, it triggers process 1 to eliminate the least valuable file block B in the upper-layer device. The present disclosure restricts that the elimination process can only be directly triggered by the promotion process.
[0050] Figure 2 Shows a flowchart of a data distribution method based on heterogeneous storage in an embodiment of the present disclosure.
[0051] As Figure 2 shown, the data distribution method based on heterogeneous storage according to an embodiment of the present disclosure includes the following steps:
[0052] Step S202, based on the device load rate of the heterogeneous multi-layer storage device, determine the migration probability of file blocks in each layer of the storage device.
[0053] Among them, as Figure 1 shown, the migration probability includes the promotion probability of promoting to the upper layer and the elimination probability of eliminating to the lower layer. The upper layer includes the adjacent upper layer and the separated upper layer, and the lower layer includes the adjacent lower layer and the separated lower layer.
[0054] Step S204, based on the migration probability, calculate the affinity between the migrated file block and the target storage device layer, and the migration cost of the migrated file block.
[0055] In the related art, when performing data layout among storage devices of each layer, the file popularity, that is, the access frequency of the file, is mainly considered, and the files with high popularity are placed in the storage devices of the higher layer, and the files with lower popularity are placed in the storage devices of the lower layer.
[0056] The present disclosure takes into account the characteristics of the storage device and the characteristics of the file itself on the basis of considering the predicted value of file popularity, and analyzes its impact on data transmission from the perspectives of data files and storage devices respectively. On this basis, the affinity and migration cost between the migrated file block and the target storage device layer are obtained. The higher the affinity, the better the adaptability between the migrated file block and the device layer to which it is to be migrated, and the lower the migration cost, the higher the feasibility of the migration operation.
[0057] Among them, the migration operation may include only the promotion operation to the upper layer, only the elimination operation to the lower layer, or both the promotion operation to the upper layer and the elimination operation to the lower layer.
[0058] Therefore, the promotion operation corresponds to the affinity between the file block to be promoted and the target promotion device layer, as well as the promotion cost, and the elimination operation corresponds to the affinity between the file block to be eliminated and the target eliminated device layer, as well as the elimination cost.
[0059] Step S206, when determining that there is a positive migration gain based on the affinity and migration cost, trigger the migration operation of the migrated file block based on the migration probability.
[0060] In this embodiment, by calculating the device load rate of each layer of load devices, that is, the device layer, the probability of promoting the migrated file block to the upper layer or eliminating it to the lower layer, that is, the migration probability, is determined. On the basis of considering the characteristics of the storage device and the characteristics of the file itself, the affinity and migration cost between the migrated file block and the target storage device layer are obtained through the migration probability. Further, when it is determined that the migration operation has a positive migration gain based on the affinity and migration cost, that is, when triggering the migration operation of the migrated file block based on the migration probability, on the one hand, the rationality of the data layout can be improved, and on the other hand, the way of determining whether to migrate based on the migration probability realizes a decentralized migration operation. The decentralized migration operation can reduce the impact on the system. By further detecting whether to perform the migration operation in units of migrated file blocks, a decentralized migration operation of data files is realized. The decentralized migration operation is suitable for ensuring the normal service of data blocks, and thus can reduce the impact on upper-layer applications.
[0061] Such as Figure 3 shown, in one embodiment, in step S204, a specific implementation manner of calculating the affinity between the migrated file block and the target storage device layer and the migration cost of the migrated file block respectively based on the migration probability includes:
[0062] Step S302, based on the size of the migrated file block and the read and write operations of the target storage device layer on the migrated file block, determine the bandwidth of the target storage device layer occupied.
[0063] Specifically, from the perspective of files, since the sequential read and write speeds of almost all storage devices are much higher than the random read and write speeds, and when reading and writing small files, it is necessary to update the file directory and file control block more frequently, which increases the processing burden of the device controller. The file size has the greatest impact on the read and write speed of the storage device. The smaller the file size, the slower the read and write speed, and the larger the file size, the faster the read and write speed. For the same storage device and the same total file volume, the read and write speed of a small number of large files is generally higher than that of a large number of small files.
[0064] Based on the above factors, by introducing the relationship between the size of the migrated file block, the read and write operations of the migrated file block on the target storage device layer, and the bandwidth of the target storage device layer into the calculation of affinity, the reliability and rationality of the data migration operation are improved.
[0065] Step S304, predict the popularity value of the migrated file block.
[0066] Specifically, for different files in the same layer of storage devices, their access characteristics are also different. This different access characteristic is called file popularity imbalance. The imbalance of file popularity is reflected in two aspects: spatial imbalance and temporal imbalance. Spatial imbalance is the difference in access frequency between different files, which is relatively fixed. Temporal distribution imbalance is the change in the access frequency of the same file or some related files after being written, and its distribution often has no fixed form. Obviously, the temporal distribution of file popularity is very important for the data layout of heterogeneous storage devices. Therefore, by taking into account the above factors to predict the popularity value of the migrated file block, it is possible to predict in advance to some extent which files will have an increase in popularity, and then these files can be transferred to the upper-layer storage device in advance.
[0067] Step S306, determine the affinity based on the migration probability, bandwidth, popularity value, and identification parameters of the target storage device layer.
[0068] Step S308, determine the migration cost based on the migration probability, the size of the migrated file block, bandwidth, and popularity value.
[0069] In this embodiment, when considering the affinity and migration cost between files and devices, the characteristics of data files and storage devices are comprehensively considered, including the size, popularity of files, and the bandwidth and load of storage devices, and their impacts on data transmission are analyzed respectively to select the data blocks that need to be promoted during the promotion process and the data blocks that need to be migrated during the elimination process. Further, by calculating the affinity between the data blocks that need to be promoted and the storage devices of the target layer, as well as the promotion cost, and the affinity between the data blocks that need to be eliminated and the storage devices of the target layer, as well as the elimination cost, the data layout of heterogeneous storage is guided based on the above calculation results, realizing the value optimization of the migration operation.
[0070] In one embodiment, in step S306, based on the migration probability, bandwidth, popularity value, and identification parameter of the target storage device layer, the affinity is determined, specifically including:
[0071] Based on the first formula, the affinity is calculated. The first formula is:
[0072] F = P·BandWidth(size, r / w)·[Popularity() + α·tier(id)] (1)
[0073] Where P is the migration probability, size is the size of the file block, r / w is the read / write operation of the target storage device layer on the migrated file block, BandWidth(size, r / w) is the bandwidth, Popularity() is the popularity value, tier(id) is the identification parameter of the target storage device layer, and α is the first adjustment parameter.
[0074] Specifically, the above-mentioned bandwidth BandWidth(size, r / w) can also be understood as a relational function of the bandwidth with the file size and read / write operation. Popularity() is the file access frequency, that is, the popularity value. id represents the file block, and the tier() function represents the storage level where the file block is located. When it is at the bottom layer, it is recorded as 0, and it increases by 1 for each additional layer.
[0075] Specifically, when the migration probability is the promotion probability P p the first formula is:
[0076] F p = P p ·BandWidth(size, r / w)·[Popularity() + α·tier(id)] (2)
[0077] Specifically, when the migration probability is the promotion probability P e the first formula is:
[0078] F e = P e ·BandWidth(size, r / w)·[Popularity() + α·tier(id)] (3)
[0079] In one embodiment, based on the migration probability, the size of the migrated file block, bandwidth, and popularity value, the migration cost is determined, specifically including: Based on the second formula, the migration cost is calculated. The second formula is:
[0080]
[0081] Among them, P is the migration probability, size is the size of the migrated file block, BandWidth(size, r / w) is the bandwidth, Popularity() is the popularity value, β is the second adjustment parameter, and γ is the third adjustment parameter.
[0082] In the heterogeneous storage system, if a file block is transmitted, a part of the IO bandwidth of the storage device will be occupied, resulting in a certain degree of impact on the transmission of other file blocks. Since the occupation of the bandwidth by the file block is related to its volume, the occupation of the IO bandwidth is part of the migration cost.
[0083] For the file block itself, since it cannot provide services externally during transmission, file blocks with higher popularity may encounter multiple requests, and file blocks with lower popularity may not have access requests even after migration is completed. Therefore, the popularity value of the file is also a major factor in the migration cost.
[0084] In addition, the duration of data migration, that is, the ratio of the file block volume to the transmission rate, is also a factor in the migration cost.
[0085] Since data migration is divided into two processes: promotion and elimination, and the two correspond to different occurrence probabilities, they need to be calculated separately when calculating the migration cost. Therefore, the specific formula for the migration cost is shown in the following formulas (5) and (6), where formula (5) is the promotion cost and formula (6) is the elimination cost.
[0086]
[0087]
[0088] Among them, β is the second parameter and γ is the third parameter, both of which are scaling factors. β is used to adjust the migration cost to a value that can be compared with the file device affinity, and γ is used to adjust the file block popularity to a value that can be compared with the bandwidth.
[0089] In one embodiment, predicting the popularity value of a migrated file block specifically includes: collecting the access timestamp information of the migrated file block within a preset sliding window; predicting the popularity value of the migrated file block based on the number of access timestamp information and the length of the sliding window.
[0090] Specifically, in the popularity value prediction scheme of the present disclosure, when performing access frequency statistics, by using the sliding window mechanism, the window stores the recent access timestamp records of the file. Dividing the number of records in the window by the time difference between the access records at both ends of the window is used as the predicted access frequency. The specific formula is:
[0091]
[0092] Among them, n is the number of access timestamp information, Tmax is the end point of the sliding window, T min is the starting point of the sliding window.
[0093] As Figure 4 shown, in one embodiment, in step S202, based on the device load rate of the heterogeneous multi-layer storage device, determining the migration probability of file blocks in each layer of the storage device specifically includes:
[0094] Step S402, detecting the IO load rate and storage load of each layer of the storage device, and determining one of the IO load rate and the storage load as the device load rate of each layer of the storage device.
[0095] The load level of the storage device is a very important factor affecting its performance. When the device load level reaches a certain degree, it may cause a sharp drop in bandwidth or even inability to work properly. Generally speaking, for a storage device, the load includes two aspects: storage load and IO load.
[0096] The storage load is relatively simple, that is, the ratio of the used capacity to the total capacity of the device, as shown in formula (8):
[0097]
[0098] For the IO load, the present disclosure adopts a standardized processing index, that is, the ratio of the actual IO operation time of the storage device to the total time, as shown in formula (9):
[0099]
[0100] This index represents the utilization rate of the IO device. The utilization rate is not the higher the better. Excessive utilization rate will lead to excessive latency. When the IO load is high, the IO queue will accumulate a large number of IO operations, resulting in a significant increase in the average waiting time of IO operations, thereby affecting the data throughput or response time of the upper-layer application.
[0101] Because the quantization forms of both the storage load and the IO load are in the form of percentages and have the same value range, the present disclosure uses the larger of the two to represent the load level of the device. Combining the method of probability migration can maintain the storage load and the IO load at a reasonable level.
[0102] Determining one of the IO load rate and the storage load as the device load rate of any layer of the storage device specifically is:
[0103] u = max(StorageLoad, IOLoad) (10)
[0104] That is, the device load rate is represented by the maximum value of the storage load and the IO load.
[0105] Step S404, when the specified storage device serves as the target storage device layer, calculate the promotion probability of the migration file blocks to be promoted to the target storage device layer based on the device load rate of the specified storage device.
[0106] In one embodiment, calculating the promotion probability of the migration file blocks to be promoted to any layer of storage devices based on the device load rate specifically includes:
[0107] Calculate the promotion probability based on the third formula, and the third formula is:
[0108]
[0109] where u is the device load rate and K is a custom parameter.
[0110] Step S406, calculate the elimination probability of the migration file blocks to be eliminated from the target storage device layer based on the promotion probability.
[0111] Specifically, the calculation method is as shown in formula (12).
[0112] P p =1 - P e (12)
[0113] First, formula (11) indicates that the device load rate is represented by the maximum value of the storage load and the IO load. Among them, the storage load is the occupancy rate of the storage capacity of the device, and the IO load is the IO occupancy rate. Secondly, formula (12) is the functional form of the elimination probability, where K is a parameter defined by the user. Finally, formula (12) is the functional form of the promotion probability, that is, 1 minus the elimination probability. By adjusting the value of K, the device load rate level at load balancing can be controlled.
[0114] In this embodiment, each data migration consumes a certain amount of time and system resources. At the same time, the external service of the data will be affected during this period, and the more concentrated the migration time of each data block, the more serious the impact. To solve this problem, a data migration method of probability migration is proposed in the present disclosure, that is, whether to migrate the file blocks is determined according to a specific probability. In this way, the actual amount of data migration can be effectively suppressed. Compared with the traditional high and low water level method, the time of data migration is also more dispersed. The promotion probability and the elimination probability must meet the following conditions:
[0115] (1) When the device load rate is 0, the promotion probability must be 1, and the elimination probability must be 0.
[0116] (2) When the device load rate is 1, the promotion probability must be 0, and the elimination probability must be 1.
[0117] (3) There is a device load rate value. When this value is reached, the promotion probability and the elimination probability are equal.
[0118] (4) This value can be set by the user according to needs.
[0119] In addition, those skilled in the art can also understand that for the file blocks in the bottom - layer storage device, the elimination probability is 0, and for the file blocks in the top - layer storage device, the promotion probability is 0.
[0120] In one embodiment, when determining that there is a positive migration benefit based on affinity and migration cost, triggering the migration operation of the migrated file block based on the migration probability, which specifically includes: when detecting that the difference between the affinity and the migration cost is greater than 0, determining that there is a positive migration benefit; randomly generating a migration parameter that is greater than or equal to 0 and less than or equal to 1, and when the migration parameter is less than the migration probability, performing the migration operation.
[0121] In this embodiment, if the affinity is greater than the migration cost, it indicates a beneficial migration. At this time, by randomly generating a migration parameter, if the migration parameter is less than the migration probability, triggering the execution of the migration operation to ensure the reliability of the migration operation.
[0122] In one embodiment, the migrated file block includes the file block to be promoted and the file block to be eliminated. Triggering the migration operation of the migrated file block based on the migration probability specifically further includes: when triggering the migration operation of the file block to be promoted based on the promotion probability, performing a locking operation on the file block to be promoted, copying the file block to be promoted to the target storage device layer in the upper layer, unlocking and deleting the original file block to be promoted; when triggering the migration operation of the file block to be eliminated based on the elimination probability, performing a locking operation on the file block to be eliminated, copying the file block to be eliminated to the target storage device layer in the lower layer, unlocking and deleting the original file block to be eliminated.
[0123] As Figure 5 shown, the data distribution scheme based on heterogeneous storage in the present disclosure can be executed based on the following four modules: file access frequency prediction module 502, device load monitoring module 504, affinity calculation module 506, and data migration module 508.
[0124] The present disclosure can be developed based on the distributed file system Alluxio. The internal storage nodes use a heterogeneous storage system, which is divided into a memory layer, namely memory disk 102, NVM hard disk layer 104, solid - state hard disk layer 106, and mechanical hard disk layer 108.
[0125] By monitoring the changes in characteristics such as the load status of the storage node and the access frequency of the stored data, adjusting the layout of the data among the storage levels according to the algorithm designed by the invention.
[0126] The file access frequency prediction module 502 is used to predict the access frequency of each file in the next period of time and use it as a factor in calculating the file-device affinity.
[0127] Specifically, by using the sliding window mechanism to update data, for a certain file, whenever it is accessed, the current timestamp is recorded. Using the timestamps at the head and end of the window, as well as the number of records within the window, a predicted value of the file access frequency is calculated. It is recalculated every once in a while. And to ensure thread safety, the file access frequency prediction results are stored in a ConcunrrentHashMap container.
[0128] The device load monitoring module 504 is used to consider the dynamic load level of the storage device when calculating the file-device affinity, and monitor the load levels of each storage device during the operation of the system, mainly including separately monitoring the storage load and the IO load.
[0129] First, for the monitoring of the storage load, in the present disclosure, the "du" command can be used to obtain the current size of the directory, and then compared with the pre-configured size of this storage layer to obtain the storage load rate.
[0130] Second, for the monitoring of the IO load, in the present disclosure, the "iostat" command is used to obtain the monitoring data of each storage device, and then the utilization value is extracted from it. Finally, the maximum value of the two loads is used to represent the load of the device.
[0131] The above-mentioned file access frequency prediction module and device load monitoring module are used to respectively obtain the dynamic characteristics of data files and storage devices. Other static characteristics such as file size and device bandwidth do not need to be updated at any time, so they can be directly obtained during calculation.
[0132] The affinity calculation module 506 is implemented as the file block eviction policy of the Alluxio system. The affinity calculation module 506 is used to: when a certain file block in the storage node is accessed, it will be screened through the file block eviction policy, thus triggering the affinity calculation.
[0133] The affinity calculation module obtains data from each data source prepared in advance, calculates the affinity of the file block for the current device and the adjacent storage layer devices according to the previously defined affinity function, and stores it in the ConcunrrentHashMap container for waiting to be used.
[0134] The data migration module 508 is used to: after obtaining the affinity values of each file block and device, it can be used as a basis for file block migration judgment and execution.
[0135] First, the data migration module selects the file block pairs to be migrated: the first is the file block to be promoted to a higher level, i.e., the file block that has just been accessed, and the second is the file block with the lowest affinity in the upper-layer storage device.
[0136] Then, according to the migration cost and migration benefit defined above, the migration benefit is calculated. When the migration benefit is greater than 0, the file block migration is triggered according to the migration probability. During the migration process, to ensure consistency, the file block needs to be locked to block access first, then the file block is copied to the target storage layer, and then the original file block is deleted and the lock is released. The modification of the metadata is the responsibility of the distributed file system Alluxio.
[0137] The specific implementation process is as Figure 6 shown.
[0138] As Figure 6 shown, a data distribution method based on heterogeneous storage according to an embodiment of the present disclosure includes:
[0139] Step S602: In response to a file access request, determine the file block A corresponding to the access request.
[0140] Step S604: Obtain the predicted popularity value of the file block A, the device load rate of the target storage device layer, and the static features of the file block A and the target storage device layer, respectively.
[0141] Step S606: Calculate the affinity between the file block A and the device based on the above data.
[0142] Step S608: Obtain the evicted file block B in the target storage device layer.
[0143] And calculate the affinity between the file block B and the device based on the corresponding data.
[0144] Step S610: Calculate the migration costs of the file block A and the file block B, respectively.
[0145] Step S612: When it is determined that the migration benefit of the file block A is greater than 0 based on the affinity and the migration cost, determine whether to promote the file block A based on the promotion probability. If "yes", go to step S614; if "no", end the process.
[0146] Step S614: Promote the file block A.
[0147] Step S616: When it is determined that the migration benefit of the file block B is greater than 0 based on the affinity and the migration cost, determine whether to evict the file block B based on the eviction probability. If "yes", go to step S618; if "no", end the process.
[0148] Step S618: Evict the file block B.
[0149] Step S620: Lock the file block and copy it to the target storage device layer.
[0150] Step S622: Delete the original file block and unlock the locked file block.
[0151] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0152] Next, refer to Figure 7 to describe the heterogeneous storage-based data distribution device 700 according to the embodiments of the present disclosure. Figure 7 The shown heterogeneous storage-based data distribution device 700 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0153] The heterogeneous storage-based data distribution device 700 is presented in the form of a hardware module. The components of the heterogeneous storage-based data distribution device 700 may include, but are not limited to: a determination module 702, configured to determine the migration probability of file blocks in each layer of storage devices based on the device load rate of heterogeneous multi-layer storage devices, where the migration probability includes the promotion probability of promoting to the upper layer and / or the elimination probability of eliminating to the lower layer; a calculation module 704, configured to calculate the affinity between the migrated file blocks and the target storage device layer respectively based on the migration probability, and the migration cost of the migrated file blocks; a migration module 706, configured to trigger the migration operation of the migrated file blocks based on the migration probability when it is determined that there is a positive migration gain based on the affinity and the migration cost.
[0154] Next, refer to Figure 8 to describe the electronic device 800 according to this embodiment of the present disclosure. The electronic device 800 shown in FIG. 8 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0155] As Figure 8 shown, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: the above-mentioned at least one processing unit 810, the above-mentioned at least one storage unit 820, and a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810).
[0156] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 810 can execute asFigure 2 The solution described in steps S202 to S206 shown in
[0157] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202, and may further include a read-only storage unit (ROM) 8203.
[0158] The storage unit 820 may also include a program / utilities 8204 having a set (at least one) of program modules 8205. Such program modules 8205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0159] The bus 830 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0160] The electronic device 800 may also communicate with one or more external devices 870 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 800, and / or may communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through an input / output (I / O) interface 850. Also, the electronic device 800 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 860. As shown, the network adapter 860 communicates with other modules of the electronic device 800 through the bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0161] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a terminal device, or an electronic device, etc.) to execute the method according to the embodiments of the present disclosure.
[0162] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium having stored thereon a program product capable of implementing the above-described method of this specification. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code that, when the program product runs on a terminal device, is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0163] The program product for implementing the above method according to the embodiments of the present disclosure may adopt a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0164] The program product may adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0165] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0166] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0167] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0168] It should be noted that although several modules or units of the devices for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by a plurality of modules or units.
[0169] In addition, although the various steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be performed in that specific order, or that all of the shown steps must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0170] Those skilled in the art can easily understand from the description of the above embodiments that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or an electronic device, etc.) to execute the method according to the embodiments of the present disclosure.
[0171] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
Claims
1. A data distribution method based on heterogeneous storage, characterized in that, Including: Determining the migration probability of file blocks in each layer of the storage device based on the device load rate of heterogeneous multi-layer storage devices, where the migration probability includes the promotion probability of promoting to the upper layer and / or the elimination probability of eliminating to the lower layer; Calculating the affinity between the migrated file block and the target storage device layer and the migration cost of the migrated file block respectively based on the migration probability, including: determining the affinity based on the migration probability, the bandwidth of the target storage device layer, the popularity value of the migrated file block, and the identification parameter of the target storage device layer; and determining the migration cost based on the migration probability, the size of the migrated file block, the bandwidth, and the popularity value, where the popularity value is predicted for the migrated file block; When it is determined that there is a positive migration benefit based on the affinity and the migration cost, triggering the migration operation of the migrated file block based on the migration probability.
2. The data distribution method based on heterogeneous storage according to claim 1, wherein Also including: Determining the bandwidth of the target storage device layer occupied based on the size of the migrated file block and the read / write operations of the target storage device layer on the migrated file block.
3. The data distribution method based on heterogeneous storage according to claim 2, wherein The determining the affinity based on the migration probability, the bandwidth, the popularity value, and the identification parameter of the target storage device layer specifically includes: Calculating the affinity based on the first formula, where the first formula is: , where P is the migration probability, BandWidth(size, r / w) is the bandwidth, Popularity() is the popularity value, tier(id) is the identification parameter of the target storage device layer, α is the first adjustment parameter, and r / w is the read operation or write operation of the target storage device layer on the migrated file block.
4. The data distribution method based on heterogeneous storage according to claim 2, wherein The determining the migration cost based on the migration probability, the size of the migrated file block, the bandwidth, and the popularity value specifically includes: Calculating the migration cost based on the second formula, where the second formula is: , where P is the migration probability, size is the size of the migrated file block, BandWidth(size, r / w) is the bandwidth, Popularity() is the popularity value, β is the second adjustment parameter, and γ is the third adjustment parameter.
5. The data distribution method based on heterogeneous storage according to claim 2, wherein The predicting the popularity value of the migrated file block specifically includes: Collecting the access timestamp information of the migrated file block within a preset sliding window; Predicting the popularity value of the migrated file block based on the number of the access timestamp information and the length of the sliding window.
6. The data distribution method based on heterogeneous storage according to claim 1, wherein The determining the migration probability of file blocks in each layer of the storage device based on the device load rate of heterogeneous multi-layer storage devices specifically includes: Detecting the IO load rate and storage load of each layer of the storage device, and determining one of the IO load rate and the storage load as the device load rate of each layer of the storage device; When designating the storage device as the target storage device layer, calculating the promotion probability of the migrated file block to be promoted to the target storage device layer based on the device load rate of the designated storage device; Calculating the elimination probability of the migrated file block to be eliminated from the target storage device layer based on the promotion probability.
7. The data distribution method based on heterogeneous storage according to claim 6, wherein Calculating the promotion probability of the migration file block to be promoted to the target storage device layer based on the device load rate of the specified storage device, specifically including: Calculating the promotion probability based on the third formula, and the third formula is: , Where u is the device load rate and K is a custom parameter.
8. The data distribution method based on heterogeneous storage according to claim 1, wherein When determining that there is a positive migration benefit based on the affinity and the migration cost, triggering the migration operation of the migration file block based on the migration probability, specifically including: When detecting that the difference between the affinity and the migration cost is greater than 0, it is determined that there is the positive migration benefit; Randomly generate a migration parameter greater than or equal to 0 and less than or equal to 1. When the migration parameter is less than the migration probability, execute the migration operation.
9. The data distribution method based on heterogeneous storage according to any one of claims 1 to 8, characterized in that The migration file block includes a file block to be promoted and a file block to be eliminated. Triggering the migration operation of the migration file block based on the migration probability specifically further includes: When triggering the migration operation of the file block to be promoted based on the promotion probability, perform a locking operation on the file block to be promoted, copy the file block to be promoted to the upper-layer target storage device layer, unlock and delete the original file block to be promoted; When triggering the migration operation of the file block to be eliminated based on the elimination probability, perform a locking operation on the file block to be eliminated, copy the file block to be eliminated to the lower-layer target storage device layer, unlock and delete the original file block to be eliminated.
10. The heterogeneous storage-based data distribution method according to any one of claims 1 to 8, characterized in that The heterogeneous multi-layer storage device includes a memory disk, an NVMe hard disk, a solid-state drive, and a mechanical hard disk in sequence from top to bottom.
11. A data distribution device based on heterogeneous storage, characterized in that, Including: A determination module, configured to determine the migration probability of the file block in each layer of the storage device based on the device load rate of the heterogeneous multi-layer storage device, where the migration probability includes a promotion probability of being promoted to the upper layer and / or an elimination probability of being eliminated to the lower layer; A calculation module, configured to calculate the affinity between the migration file block and the target storage device layer and the migration cost of the migration file block respectively based on the migration probability, including: determining the affinity based on the migration probability, the bandwidth of the target storage device layer, the popularity value of the migration file block, and the identification parameter of the target storage device layer; and determining the migration cost based on the migration probability, the size of the migration file block, the bandwidth, and the popularity value, where the popularity value is predicted for the migration file block; A migration module, configured to trigger the migration operation of the migration file block based on the migration probability when determining that there is a positive migration benefit based on the affinity and the migration cost.
12. An electronic device, characterized in that, Including: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the heterogeneous storage-based data distribution method according to any one of claims 1 to 10 by executing the executable instructions.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the heterogeneous storage-based data distribution method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Content inquiry method and device based on multiple stages of cache modules
CN104217019A
Hierarchical storage data migration method and system
CN111367469A