Data storage method based on distributed storage and hierarchical storage mechanism

By adopting a cache phase-out algorithm based on Gaussian distribution morphological characteristics in Ceph storage clusters, the problem of limited cache hit rate performance in the existing technology is solved, and more efficient cache management and data access performance is achieved.

CN120045134APending Publication Date: 2025-05-27709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510112416.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, when Ceph storage cluster uses cheap PCs and mechanical hard disks, the IO performance is limited, and the cache hit rate performance is limited, which cannot effectively improve the I/O performance of the back-end storage layer.

Method used

The cache elimination algorithm based on the Gaussian distribution morphological characteristics is adopted, and the data objects in the cache pool are managed through the Gaussian elimination center and the Gaussian elimination factor in the way that the data object farthest from the Gaussian elimination center is first eliminated.

Benefits of technology

It effectively improves the cache hit rate, prioritizes the data that may be frequently accessed in burst access, and reduces the impact of data migration on system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045134A_ABST
    Figure CN120045134A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computers, and particularly discloses a distributed storage and hierarchical storage mechanism-based data storage method, which comprises the following steps of: receiving a data access request for requesting a target data object; under the condition that the space of the cache pool is full and the target data object is not stored in the cache pool, based on a Gaussian elimination center and a Gaussian elimination factor of the target data object, managing the data objects in the cache pool according to a mode that the data object, farthest from the Gaussian elimination center, of the Gaussian elimination factor is firstly eliminated; wherein in the statistical analysis of the data access, the Gaussian elimination factor of the data object conforms to Gaussian distribution, the Gaussian elimination factor is used for representing the feature of a specified dimension of the data object, and the Gaussian elimination center is the mean value of the Gaussian distribution conforming to the Gaussian elimination factor of the data object. Through the cache elimination algorithm based on the Gaussian distribution morphological characteristics provided by the invention, the cache hit rate can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and more specifically, relates to a data storage method based on a distributed storage and hierarchical storage mechanism. Background Art

[0002] With the continuous development of cloud computing and big data technologies, the architecture of storage systems is also evolving. In this evolution process, storage systems are gradually developing from a simple host-based architecture towards networking and virtualization. To achieve unified management of storage systems and solve the compatibility problems of storage devices and heterogeneous hardware problems, storage virtualization technology has emerged. Storage virtualization technology is a technology that separates logical resource mapping and physical storage. Its core idea is to eliminate the heterogeneity of physical devices through the abstraction, hiding, and isolation of system or service functions, and to achieve independent management of data services and physical hardware. A typical application of this technology is a distributed storage system. A distributed storage system stores data dispersedly on multiple independent devices and provides a unified storage service. Compared with traditional centralized storage systems, distributed storage systems have scalability and extensibility, and can add more storage devices as needed to meet the growing data requirements. Ceph is an open-source software-defined storage system that uses network protocols to form a storage cluster, provides multiple access interfaces to meet the distributed storage requirements of various usage scenarios, eliminates the dependence on a single central node of the system, and realizes a truly centerless structure.

[0003] If a Ceph storage cluster is built using inexpensive PCs and traditional mechanical hard disks, the access speed of the disks is limited to a certain extent and cannot reach the ideal IOPS (Input / Output Operations Per Second) performance level. To optimize the IO performance of the system, fast storage devices can be added as caches to reduce data access latency. Among them, the Cache Tier hierarchical storage mechanism is a common solution. Figure 1 is a schematic diagram of the Ceph hierarchical architecture provided by the prior art. As Figure 1 shown, it is widely used in the Ceph server cache and can effectively improve the I / O performance of the backend storage layer. Figure 1 In [Figure], Objecter represents the data object being accessed. The Cache Tier needs to create a storage pool composed of high-speed and expensive storage devices (such as SSDs) as the cache layer, and a backend storage pool composed of relatively inexpensive devices as the economic storage layer.

[0004] Although the Cache Tier works at the storage pool level, it still adopts an elimination rule based on frequency-estimated probability similar to the LRU (Least Recently Used) algorithm. Due to the error in frequency-estimated probability, when designing the elimination algorithm in this way, the cache hit rate performance is limited. Summary of the Invention

[0005] Aiming at the defects of the prior art, the purpose of this application is to improve the cache hit rate based on the cache elimination algorithm of Gaussian distribution morphological characteristics, aiming to solve the problem of limited cache hit rate performance of the existing elimination algorithm.

[0006] To achieve the above object, in the first aspect, this application provides a data storage method based on a distributed storage and hierarchical storage mechanism, including:

[0007] Receiving a data access request for requesting a target data object;

[0008] In the case where the cache pool space is full and the target data object is not stored in the cache pool, based on the Gaussian elimination center and the Gaussian elimination factor of the target data object, in the manner that the data object with the Gaussian elimination factor farthest from the Gaussian elimination center is eliminated first, managing the data objects in the cache pool;

[0009] Among them, in the statistical analysis of data access, the Gaussian elimination factor of the data object conforms to a Gaussian distribution, the Gaussian elimination factor is used to characterize the characteristics of a specified dimension of the data object, and the Gaussian elimination center is the mean value of the Gaussian distribution that the Gaussian elimination factor of the data object conforms to.

[0010] In a possible implementation manner, the Gaussian elimination center is determined by the following formula:

[0011]

[0012] Among them, Gauss i represents the Gaussian elimination factor of any accessed data object within a certain time period, total represents the total number of accessed data objects within a certain time period, represents the Gaussian elimination center.

[0013] In a possible implementation manner, the distance between the Gaussian elimination factor and the Gaussian elimination center is determined by the following formula:

[0014]

[0015] Among them, dis represents the distance between the Gaussian elimination factor and the Gaussian elimination center.

[0016] In a possible implementation, based on the Gaussian elimination center and the Gaussian elimination factor of the target data object, manage the data objects in the cache pool in the way that the data object with the Gaussian elimination factor farthest from the Gaussian elimination center is eliminated first, including:

[0017] Obtain an AVL tree. A node in the AVL tree represents a data object in the cache pool. The position of the data object in the AVL tree is determined based on the Gaussian elimination factor of the data object, and the root node of the AVL tree serves as the Gaussian elimination center;

[0018] Based on the AVL tree and the Gaussian elimination factor of the target data object, manage the data objects in the cache pool in the way that the leaf node farthest from the root node is eliminated first.

[0019] In a possible implementation, based on the AVL tree and the Gaussian elimination factor of the target data object, manage the data objects in the cache pool in the way that the leaf node farthest from the root node is eliminated first, including:

[0020] When gl ≤ g ≤ gr, based on the Gaussian elimination factor of the target data object, delete the leaf node farthest from the root node, remove the data object corresponding to the deleted leaf node from the cache pool, insert the target data object as a node into the AVL tree, and load the target data object into the cache pool;

[0021] where gl represents the Gaussian elimination factor corresponding to the leftmost leaf node in the AVL tree, gr represents the Gaussian elimination factor corresponding to the rightmost leaf node in the AVL tree, and g represents the Gaussian elimination factor of the target data object.

[0022] In a possible implementation, it further includes:

[0023] For the data objects in the cache pool, if Size > Size H , add the migration task of the data object to the data migration task list; if the time interval between the creation time of the data object and the current time is greater than the preset time interval and Size L ≤ Size ≤ Size H and the occupied ratio of the used cache space is greater than or equal to the first ratio, add the migration task of the data object to the data migration task list; if the time interval between the creation time of the data object and the current time is greater than the preset time interval and Size < Size L and the occupied ratio of the used cache space is greater than or equal to the second ratio, add the migration task of the data object to the data migration task list;

[0024] Execute the data migration tasks in the data migration task list;

[0025] Among them, Size represents the file size of the data object in the cache pool, Size H represents the first file size threshold, Size L represents the second file size threshold, Size L <Size H , and the first proportion is less than the second proportion.

[0026] In a possible implementation, before executing the data migration tasks in the data migration task list, it further includes:

[0027] If the occupied ratio of the used cache space is greater than or equal to the second proportion, it is determined to execute the data migration tasks in the data migration task list;

[0028] Or, if the network traffic is less than the traffic threshold and the occupied ratio of the used cache space is less than the first proportion, it is determined to execute the data migration tasks in the data migration task list.

[0029] In a second aspect, the present application provides a data storage device based on a distributed storage and hierarchical storage mechanism, including:

[0030] An access request receiving module, configured to receive a data access request for requesting a target data object;

[0031] A cache management module, configured to, when the cache pool space is full and the target data object is not stored in the cache pool, manage the data objects in the cache pool based on the Gaussian elimination center and the Gaussian elimination factor of the target data object, in such a way that the data object with the farthest distance from the Gaussian elimination center according to the Gaussian elimination factor is eliminated first;

[0032] Among them, in the statistical analysis of data access, the Gaussian elimination factor of the data object conforms to a Gaussian distribution, the Gaussian elimination factor is used to characterize the characteristics of a specified dimension of the data object, and the Gaussian elimination center is the mean of the Gaussian distribution that the Gaussian elimination factor of the data object conforms to.

[0033] In a third aspect, the present application provides an electronic device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory, and when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0034] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program runs on a processor, it causes the processor to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0035] Generally speaking, compared with the prior art, the above technical solutions conceived by this application have the following beneficial effects:

[0036] (1) When the user's access behavior exhibits Gaussian distribution characteristics, compared with the LRU algorithm in the existing Ceph CacheTier mechanism, the cache eviction algorithm based on Gaussian distribution morphological characteristics provided by this application can effectively improve the cache hit rate.

[0037] (2) In the face of sudden access, the LRU algorithm may wrongly evict important data due to the recent access pattern. However, the cache eviction algorithm based on Gaussian distribution morphological characteristics provided by this application can more flexibly adjust the cache content, preferentially retain the data that may be frequently accessed during sudden access, and ensure the cache hit rate in data access.

[0038] (3) When the system cache space is insufficient, by distinguishing different file sizes and the remaining cache space sizes, designing a small file filtering and migration strategy can reduce the performance impact of data migration on the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a schematic diagram of the Ceph hierarchical architecture provided by the prior art;

[0040] Figure 2 is a schematic flow chart of the data storage method based on distributed storage and hierarchical storage mechanism provided by an embodiment of this application;

[0041] Figure 3 is a flow chart of the cache eviction algorithm based on Gaussian distribution morphological characteristics provided by an embodiment of this application;

[0042] Figure 4 is a structural diagram of the data migration service function module provided by an embodiment of this application;

[0043] Figure 5 is a flow chart of the data migration judgment provided by an embodiment of this application;

[0044] Figure 6 is a flow chart of the data migration task execution provided by an embodiment of this application;

[0045] Figure 7 is a structural diagram of the data storage device based on distributed storage and hierarchical storage mechanism provided by an embodiment of this application;

[0046] Figure 8 is a structural diagram of the electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To facilitate a clearer understanding of the embodiments of the present application, some relevant background knowledge is introduced as follows.

[0048] In the field of natural science, there are two typical forms of quantity distribution. One is the bell-shaped Gaussian distribution, and the other is the L-shaped power-law distribution. If a large number of tiny and independent random factors together determine the result of a certain random variable, and the individual effect of each factor is relatively uniformly small, and no single factor can play a dominant role, then this random variable generally approximates the Gaussian distribution. That is, if the information entropy between things is very large or even independent of each other, then it is usually easy to externally exhibit the Gaussian distribution. The Gaussian distribution often exists in scenarios with a wide variety of things and a complex user group. In the cloud computing scenario, users access upper-layer application software, and Ceph storage does not perceive the upper-layer service type, but only provides data storage services for the application software working on cloud computing services. On the one hand, there are many upper-layer application software served by the backend storage, and there are significant differences in the application scenarios and served objects of each application software. Therefore, the data correlation between different application softwares is weak. On the other hand, the application software has its own system architecture, and the functions of each module within the system are often independent. Except for the information associated between modules, the correlation between the information of each module is also weak. The combined effect of these two aspects results in a large information entropy between the data units of the backend storage. The information entropy characteristics of the data objects stored in the Ceph storage backend in this scenario meet the preconditions. Therefore, theoretically, the access behavior of users to these data units is more likely to exhibit the characteristics of the Gaussian distribution. The existing LRU algorithm believes that the recently accessed block is more likely to be accessed in the future, and when performing data block eviction, it preferentially evicts the data block that has not been accessed for the longest time in the cache space. Due to this mechanism, the LRU algorithm has poor resistance to sudden access to cold data (data with a low access frequency), and in extreme cases, the cache hit rate may be 0.

[0049] To overcome the above defects, the present application provides a data storage method based on a distributed storage and hierarchical storage mechanism. Through a cache eviction algorithm based on the morphological characteristics of the Gaussian distribution, the cache hit rate can be effectively improved.

[0050] To make the purpose, technical solution, and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0051] The terms "first" and "second" in the specification and claims of this article are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first proportion and the second proportion are used to distinguish different proportions, rather than to describe a specific order of the proportions.

[0052] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0053] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units, etc.; a plurality of elements refers to two or more elements, etc.

[0054] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.

[0055] Figure 2 is a schematic flowchart of a data storage method based on a distributed storage and hierarchical storage mechanism provided by the embodiments of the present application. As Figure 2 shown, the method includes the following steps S101 and step S102.

[0056] Step S101, receiving a data access request for requesting a target data object;

[0057] Step S102, in the case where the cache pool space is full and the target data object is not stored in the cache pool, based on the Gaussian elimination center and the Gaussian elimination factor of the target data object, manage the data objects in the cache pool in such a way that the data object with the Gaussian elimination factor farthest from the Gaussian elimination center is eliminated first.

[0058] Among them, in the statistical analysis of data access, the Gaussian elimination factor of a data object conforms to a Gaussian distribution. The Gaussian elimination factor is used to characterize the characteristics of a specified dimension of the data object, and the Gaussian elimination center is the mean of the Gaussian distribution to which the Gaussian elimination factor of the data object conforms.

[0059] It is understandable that there is a large information entropy among the data units stored in the backend. The information entropy characteristics of the data objects stored in the Ceph storage backend in this scenario meet the preconditions, and the access behaviors of users to these data units are more likely to exhibit the characteristics of a Gaussian distribution. If it is known through statistical analysis of the data objects that the access requests of users to the data objects conform to a Gaussian distribution in a specified dimension (such as the cumulative number of accesses, the cumulative access duration, the file type of the data object, the file name of the data object, the file size of the data object, the number of updates of the data object, etc.), then this dimension information can be recorded in the metadata of the data object, and this dimension information is called the Gaussian elimination factor of the data object. That is, in the statistical analysis of data access, the Gaussian elimination factor of the data object conforms to a Gaussian distribution, and the Gaussian elimination factor can characterize the characteristics of a specified dimension of the data object.

[0060] Furthermore, when a data access request is received, it can be determined whether the cache pool space is full and whether the target data object is stored in the cache pool. If the cache pool space is full and the target data object is not stored in the cache pool, then based on the Gaussian elimination center and the Gaussian elimination factor of the target data object, in the way that the data object with the Gaussian elimination factor farthest from the Gaussian elimination center is eliminated first, the data objects in the cache pool can be managed. The Gaussian elimination center is the mean of the Gaussian distribution that the Gaussian elimination factor of the data object conforms to. According to the morphological characteristics of the Gaussian distribution, the closer the Gaussian elimination factor of the data object accessed by the user is to the Gaussian elimination center, the higher the probability of being accessed. When the cache space is insufficient, the algorithm will preferentially eliminate the data object with the Gaussian elimination factor farthest from the Gaussian elimination center from the cache pool to improve the cache hit rate in subsequent data accesses.

[0061] Therefore, in the case where the access behaviors of users exhibit the characteristics of a Gaussian distribution, compared with the LRU algorithm in the existing CephCache Tier mechanism, the cache elimination algorithm based on the morphological characteristics of the Gaussian distribution provided by this application can effectively improve the cache hit rate.

[0062] In addition, in the face of bursty accesses, the LRU algorithm may wrongly eliminate important data due to the recent access pattern. However, the cache elimination algorithm based on the morphological characteristics of the Gaussian distribution provided by this application can more flexibly adjust the cache content, preferentially retain the data that may be frequently accessed in bursty accesses, and ensure the cache hit rate in data accesses.

[0063] In a possible implementation, the Gaussian elimination center is determined by the following formula:

[0064]

[0065] where Gauss irepresents the Gaussian elimination factor of any accessed data object within a certain time period, and total represents the total number of accessed data objects within a certain time period. represents the Gaussian elimination center.

[0066] In a possible implementation, the distance between the Gaussian elimination factor and the Gaussian elimination center is determined by the following formula:

[0067]

[0068] where dis represents the distance between the Gaussian elimination factor and the Gaussian elimination center.

[0069] In a possible implementation, based on the Gaussian elimination center and the Gaussian elimination factor of the target data object, in the way that the data object with the Gaussian elimination factor farthest from the Gaussian elimination center is eliminated first, manage the data objects in the cache pool, including:

[0070] Obtain an AVL tree. A node in the AVL tree represents a data object in the cache pool. The position of the data object in the AVL tree is determined based on the Gaussian elimination factor of the data object, and the root node of the AVL tree serves as the Gaussian elimination center;

[0071] Based on the AVL tree and the Gaussian elimination factor of the target data object, manage the data objects in the cache pool in the way that the leaf node farthest from the root node is eliminated first.

[0072] An AVL tree (Balanced Binary Tree) is a special binary tree, and its feature is that the height difference (balance factor) between the left and right subtrees of each node does not exceed a certain value (usually 1).

[0073] The root node of the AVL tree is the topmost node of the tree, and the AVL tree has only one root node. The left and right subtrees of the root node are the trees of the left and right child nodes of the root node respectively. The leftmost leaf node is the leaf node at the bottommost layer reached by starting from the root node and going straight down along the left subtree of each node. The rightmost leaf node is the leaf node at the bottommost layer reached by starting from the root node and going straight down along the right subtree of each node. The main operations of the AVL tree include: searching, inserting, deleting, and rotating.

[0074] In the case where a node is deleted from the AVL tree, remove the data object corresponding to the deleted node from the cache pool. In the case where a node is inserted into the AVL tree, load the data object corresponding to the inserted node into the cache pool.

[0075] In a possible implementation, when gl ≤ g ≤ gr, based on the Gaussian elimination factor of the target data object, the leaf node farthest from the root node is deleted, the data object corresponding to the deleted leaf node is removed from the cache pool, and the target data object is inserted into the balanced binary tree as a node, and the target data object is loaded into the cache pool;

[0076] where gl represents the Gaussian elimination factor corresponding to the leftmost leaf node in the balanced binary tree, gr represents the Gaussian elimination factor corresponding to the rightmost leaf node in the balanced binary tree, and g represents the Gaussian elimination factor of the target data object.

[0077] When g < gl or g > gr, directly write through the target data object bypassing the cache.

[0078] In a possible implementation, it further includes:

[0079] For the data object in the cache pool, if Size > Size H , the migration task of the data object is added to the data migration task list; if the time interval between the creation time of the data object and the current time is greater than the preset time interval and Size L ≤ Size ≤ Size H and the occupied cache space ratio is greater than or equal to the first ratio (for example, 60%), the migration task of the data object is added to the data migration task list; if the time interval between the creation time of the data object and the current time is greater than the preset time interval and Size < Size L and the occupied cache space ratio is greater than or equal to the second ratio (for example, 80%), the migration task of the data object is added to the data migration task list;

[0080] Execute the data migration tasks in the data migration task list;

[0081] where Size represents the file size of the data object in the cache pool, Size H represents the first file size threshold (for example, 4MB), Size L represents the second file size threshold (for example, 128KB), Size L < Size H , and the first ratio is less than the second ratio.

[0082] In a possible implementation, before executing the data migration tasks in the data migration task list, it further includes:

[0083] If the occupied cache space ratio is greater than or equal to the second ratio, determine to execute the data migration tasks in the data migration task list;

[0084] Alternatively, if the network traffic is less than the traffic threshold and the occupied cache space is less than the first ratio, determine to execute the data migration tasks in the data migration task list.

[0085] The data storage method provided by the present application based on the distributed storage and hierarchical storage mechanism will be described below with several examples.

[0086] A cache eviction algorithm based on the morphological characteristics of the Gaussian distribution was designed for the limited cache hit rate performance of the eviction rules designed based on the LRU algorithm in the Ceph Cache Tier mechanism.

[0087] Assume that through the statistical analysis of data objects, it is known that the access requests of users for data objects conform to the Gaussian distribution in a specified dimension (such as the cumulative access times, cumulative access duration, file type of the data object, file name of the data object, file size of the data object, update times of the data object, etc.). Then this dimension information can be recorded in the metadata of the data object, and this dimension information is called the Gaussian eviction factor of the data object. The mean value of the Gaussian eviction factors of the data objects accessed within a certain time period is called the Gaussian eviction center, and the Gaussian eviction center is calculated according to formula (1). Gauss i represents the Gaussian eviction factor of the data object accessed for the i-th time, and total represents the total number of data objects accessed by the user within this period of time.

[0088]

[0089] According to the morphological characteristics of the Gaussian distribution, the closer the Gaussian eviction factor of the data object accessed by the user is to the Gaussian eviction center, the higher the probability of being accessed. When the cache space is insufficient, the algorithm will preferentially evict the data object with the Gaussian eviction factor farthest from the Gaussian eviction center from the cache pool, and the distance between the Gaussian eviction factor and the Gaussian eviction center is calculated according to formula (2). Combining the good symmetry of the Gaussian distribution, an AVL tree is used to manage the cache space. Using this management structure can, on the one hand, according to the principle of estimating the mean value with the median, replace the mean value calculation process of the Gaussian center with the root node of the AVL tree, reducing the calculation amount; on the other hand, it can also transform the calculation and comparison process of the absolute value into the process of finding the leftmost or rightmost leaf node of the binary tree. Therefore, using this structure can reduce the time overhead and space overhead of the algorithm.

[0090]

[0091] Furthermore, the interest hotspots accessed by users have timeliness. When the user accesses conforms to the Gaussian distribution, this timeliness theoretically manifests as the changes in the mean and variance of the Gaussian distribution that the Gaussian elimination factor conforms to. According to the "3σ principle" of the Gaussian distribution, the change in variance will affect the theoretical limit of the algorithm hit rate under a fixed-size cache space, and will not cause the algorithm to fail, but the change in the mean will cause the original Gaussian elimination center to fail, and then lead to the deterioration of the cache space hit rate. Therefore, a steady-state monitoring mechanism is introduced into the algorithm to adapt to the possible migration of the Gaussian elimination center over time. Specifically, it is monitored whether the degree of change of the Gaussian elimination center (the absolute value of the difference between the original Gaussian elimination center and the currently calculated Gaussian elimination center) is within the preset range. If so, it is determined that the Gaussian elimination center is in a steady state and the original Gaussian elimination has not failed; if not, it is determined that the Gaussian elimination center is in a non-steady state and the original Gaussian elimination has failed, and the original Gaussian elimination center is updated to the currently calculated Gaussian elimination center.

[0092] The overall process of the algorithm is as Figure 3 shown in the figure. In the figure, g represents the Gaussian elimination factor of the data object requested by the user, gl represents the Gaussian elimination factor of the leftmost leaf node in the balanced binary tree maintained by the algorithm, and gr represents the Gaussian elimination factor of the rightmost leaf node in the balanced binary tree maintained by the algorithm. "Steady-state monitoring" in the process is used to determine whether the current Gaussian elimination center is in a steady state and generate a corresponding status flag, and this status flag will act on the "cache replacement" process in the process. The replacement algorithm will adopt different strategies to replace the data objects in the cache space for the steady state and the non-steady state.

[0093] To reduce the impact of data migration on the foreground IO, a migration plan is arranged by detecting the real-time throughput of the cluster. The functional modules of the data migration service are as Figure 4 shown. The data migration service includes a migration scan production task module, a task list, a task executor module, a cluster status information collection module, and a data migration module. The migration scan production task module obtains the list of metadata from the metadata center, traverses the metadata list, and judges whether each file meets the migration conditions. If it meets the conditions, a task is inserted into the task list. The task executor scans the task list and judges whether the task can be executed through the cluster status information. If it can be executed, it is handed over to the data migration module for data migration. The cluster status information is regularly reported to the migration service by each object storage device (Object Storage Device, OSD) through heartbeats.

[0094] In view of the impact of reduced system performance caused by data migration when the system cache space is insufficient, a small file filtering and migration strategy is designed: files are classified according to their file size (file dimension), and different migration strategies are adopted for different file sizes. For files with a size of 0 - 128 KB, they are classified as small files. When the used cache space is less than 80%, no migration is performed; when the used cache space is greater than or equal to 80%, the files are merged and migrated. For files with a size of 128 KB - 4 MB, they are classified as medium and small files. When the used cache space is less than 60%, no migration is performed; when the used cache space is greater than or equal to 60%, migration is carried out. For files with a size of more than 4 MB, they are classified as large files and are directly migrated.

[0095] It can be understood that in the case of not using the "small file filtering and migration strategy", when the cache space is insufficient, data will be migrated from the cache space to a slower storage device (disk device). At this time, the IO differences during migration caused by file size are not considered. Using the same cache space, the number of large files that can be stored is much less than the number of small files. At this time, if you want to clear some cache space, the IO operations brought by large file migration are less, while small file migration will bring a large number of IO operations, thus reducing the system performance. If the "small file filtering and migration strategy" is adopted, according to the size of the used cache space, the migration priorities are large files, medium and small files, and small files in turn. Large files will no longer be stored in the cache space, and small files will only be migrated when the used cache space is greater than 80%. The advantage is that the same cache space can store more files (small files), improving the cache utilization rate. Small files are migrated when the system resources (network, disk) are abundant, reducing the impact of system performance degradation caused by IO operations.

[0096] The flow chart of the migration scanning production task module is as Figure 5 shown. Traverse the file list in the cache space, read file metadata, obtain file attributes such as file size (size) and creation time, read the used cache space. If the file size is more than 4 MB, directly insert the production task into the task list. If the file size is less than 4 MB, then determine the creation time. If the creation time is more than N days ago, consider file migration. Then determine whether the file size is between 128 KB - 4 MB. If so, by judging the size of the used cache space, if the used cache space is greater than or equal to 60%, consider migrating this part of the data. The data size is less than 128 KB, and this part of the data is called small files. If the used cache space is greater than or equal to 80%, consider migrating this part of the data.

[0097] When the system cache space is insufficient, by distinguishing different file sizes and the size of the remaining cache space, designing a small file filtering and migration strategy can reduce the impact of data migration on system performance.

[0098] The flow chart of the data migration task execution is asFigure 6 As shown, the task is taken out from the task list, and then it is determined whether the current cache space is insufficient. If the cache space is insufficient, the migration operation is directly performed. Otherwise, it is necessary to determine the network traffic and the target disk usage rate. Only when the network traffic is lower than the threshold and the target disk usage rate is lower than 60%, the task starts to be executed or is forced to be executed. These thresholds can be dynamically configured and take effect in real time. When executing the task, data copy migration can be performed to delete the data. Finally, the migration task is deleted. In the research, the degraded service of the entire cluster is considered. When there is a risk that the remaining space in the cache space is 0 and large object data is received, the data is directly stored in the storage pool.

[0099] The data storage device based on the distributed storage and hierarchical storage mechanism provided by the present application is described below. The data storage device based on the distributed storage and hierarchical storage mechanism described below can be mutually corresponding and referred to the data storage method based on the distributed storage and hierarchical storage mechanism described above.

[0100] Figure 7 is a schematic structural diagram of the data storage device based on the distributed storage and hierarchical storage mechanism provided by the embodiment of the present application. As Figure 7 shown, the device includes: an access request receiving module 10 and a cache management module 20. Among them:

[0101] The access request receiving module 10 is used to receive a data access request for requesting a target data object;

[0102] The cache management module 20 is used to manage the data objects in the cache pool based on the Gaussian elimination center and the Gaussian elimination factor of the target data object in the case where the cache pool space is full and the target data object is not stored in the cache pool, in the manner that the data object with the farthest Gaussian elimination factor from the Gaussian elimination center is eliminated first.

[0103] Among them, in the statistical analysis of data access, the Gaussian elimination factor of the data object conforms to the Gaussian distribution. The Gaussian elimination factor is used to characterize the characteristics of a specified dimension of the data object, and the Gaussian elimination center is the mean value of the Gaussian distribution to which the Gaussian elimination factor of the data object conforms.

[0104] It can be understood that the detailed function implementation of the above-mentioned each unit / module can be referred to the introduction in the foregoing method embodiment, and will not be elaborated here.

[0105] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. For the corresponding program module in the device, its implementation principle and technical effect are similar to the description in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method, and will not be elaborated here.

[0106] Based on the method in the above embodiments, an embodiment of the present application provides an electronic device. Figure 8 It is a schematic structural diagram of the electronic device provided by the embodiment of the present application. As Figure 8 shown, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 may call the logical instructions in the memory 830 to execute the method in the above embodiments.

[0107] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application.

[0108] Based on the method in the above embodiments, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a processor, the processor is caused to execute the method in the above embodiments.

[0109] Based on the method in the above embodiments, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor is caused to execute the method in the above embodiments.

[0110] It can be understood that the processor in the embodiment of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0111] The method steps in the embodiments of this application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0112] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0113] It can be understood that the various numerical numbers involved in the embodiments of this application are only for the convenience of description and are not used to limit the scope of the embodiments of this application.

[0114] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A data storage method based on distributed storage and hierarchical storage mechanism, characterized in that: include: receiving a data access request for a target data object; When the cache pool space is full and the target data object is not stored in the cache pool, the data objects in the cache pool are managed based on the Gaussian elimination center and the Gaussian elimination factor of the target data object, in a manner that the data object with the Gaussian elimination factor farthest from the Gaussian elimination center is eliminated first; Among them, in the statistical analysis of data access, the Gaussian elimination factor of the data object conforms to the Gaussian distribution, the Gaussian elimination factor is used to characterize the characteristics of a specified dimension of the data object, and the Gaussian elimination center is the mean of the Gaussian distribution conformed to by the Gaussian elimination factor of the data object.

2. The data storage method based on distributed storage and hierarchical storage mechanism according to claim 1 is characterized in that: The Gaussian elimination center is determined by the following formula: Among them, Gauss i represents the Gaussian elimination factor of any accessed data object within a certain period of time, and total represents the total number of accessed data objects within a certain period of time. represents Gaussian elimination center.

3. The data storage method based on distributed storage and hierarchical storage mechanism according to claim 2 is characterized in that: The distance between the Gaussian elimination factor and the Gaussian elimination center is determined by the following formula: Where dis represents the distance between the Gaussian elimination factor and the Gaussian elimination center.

4. The data storage method based on distributed storage and hierarchical storage mechanism according to claim 1 is characterized in that: The method of managing data objects in the cache pool based on the Gaussian elimination center and the Gaussian elimination factor of the target data object in a manner that the data object whose Gaussian elimination factor is farthest from the Gaussian elimination center is eliminated first includes: Obtain a balanced binary tree, where a node in the balanced binary tree represents a data object in the cache pool. The position of the data object in the balanced binary tree is determined based on the Gaussian elimination factor of the data object, and the root node of the balanced binary tree serves as the Gaussian elimination center. Based on the balanced binary tree and the Gaussian elimination factor of the target data object, the data objects in the cache pool are managed in such a way that the leaf nodes farthest from the root node are eliminated first.

5. The data storage method based on distributed storage and hierarchical storage mechanism according to claim 4 is characterized in that: The method of managing data objects in the cache pool based on the Gaussian elimination factor of the balanced binary tree and the target data object in a manner that the leaf nodes farthest from the root node are eliminated first includes: In the case of gl≤g≤gr, based on the Gaussian elimination factor of the target data object, the leaf node farthest from the root node is deleted, the data object corresponding to the deleted leaf node is removed from the cache pool, and the target data object is inserted into the balanced binary tree as a node, and the target data object is loaded into the cache pool; Among them, gl represents the Gaussian elimination factor corresponding to the leftmost leaf node in the balanced binary tree, gr represents the Gaussian elimination factor corresponding to the rightmost leaf node in the balanced binary tree, and gg represents the Gaussian elimination factor of the target data object.

6. The data storage method based on distributed storage and hierarchical storage mechanism according to any one of claims 1 to 5, characterized in that: Also includes: For data objects in the cache pool, if Size>Size H , then add the migration task of the data object to the data migration task list; if the time interval between the creation time of the data object and the current time is greater than the preset time interval and Size L ≤Size≤Size H If the used cache space ratio is greater than or equal to the first ratio, the migration task of the data object is added to the data migration task list; if the time interval between the creation time of the data object and the current time is greater than the preset time interval and Size <Size L If the proportion of used cache space is greater than or equal to the second proportion, the migration task of the data object is added to the data migration task list; Execute the data migration tasks in the data migration task list; Among them, Size represents the file size of the data object in the cache pool. H Indicates the first file size threshold, Size L Indicates the second file size threshold, Size L <Size H , the first proportion is smaller than the second proportion.

7. The data storage method based on distributed storage and hierarchical storage mechanism according to claim 6 is characterized in that: Before executing the data migration tasks in the data migration task list, you also need to: If the proportion of used cache space is greater than or equal to the second proportion, determining to execute the data migration task in the data migration task list; Or, if the network traffic is less than the traffic threshold and the used cache space accounts for less than the first proportion, it is determined to execute the data migration task in the data migration task list.

8. A data storage device based on distributed storage and hierarchical storage mechanism, characterized in that: include: An access request receiving module, used for receiving a data access request for requesting a target data object; A cache management module, for managing data objects in the cache pool, based on a Gaussian elimination center and a Gaussian elimination factor of the target data object, in a case where the cache pool space is full and the target data object is not stored in the cache pool, in a manner that the data object with the Gaussian elimination factor farthest from the Gaussian elimination center is eliminated first; Among them, in the statistical analysis of data access, the Gaussian elimination factor of the data object conforms to the Gaussian distribution, the Gaussian elimination factor is used to characterize the characteristics of a specified dimension of the data object, and the Gaussian elimination center is the mean of the Gaussian distribution conformed to by the Gaussian elimination factor of the data object.

9. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor is caused to execute the method according to any one of claims 1 to 7.