Data Management Method, Device, Equipment, Medium and Product in a Storage System

By dividing metadata into hot metadata and cold metadata according to access frequency in a distributed storage system and saving them to different types of storage pools, the difficulty and cost of metadata management are solved, and more efficient storage management and performance improvements are achieved.

CN119690357BActive Publication Date: 2025-06-13JINAN INSPUR DATA TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510200151.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

In distributed storage systems, as the number of objects and the scale of metadata increases, the difficulty and cost of metadata management increases, resulting in excessive management of storage systems.

Method used

By dividing the metadata into hot metadata and cold metadata according to the preset division ratio and access frequency of the target metadata, and saving them to high-speed storage pools and low-speed storage pools respectively, the reasonable management of data and metadata is achieved.

Benefits of technology

It effectively reduces the management cost and management difficulty of the storage system, and improves the overall performance of the storage system and the utilization rate of the storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119690357B_ABST
    Figure CN119690357B_ABST
Patent Text Reader

Abstract

The present invention discloses a data management method, device, equipment, medium and product in a storage system, which is applied to the storage technology field, can manage the stored data and its metadata in the storage system more reasonably and effectively, and further reduces the management cost and management difficulty of the storage system. The method includes: determining data to be stored; wherein, the data to be stored includes target data and target metadata corresponding to the target data; dividing each piece of the target metadata into hot metadata and cold metadata according to a preset division ratio and the access frequency of each piece of the target metadata; saving the hot metadata to a high-speed storage pool, and saving the target data corresponding to the hot metadata to a low-speed storage pool; saving the cold metadata and the target data corresponding to the cold metadata to the low-speed storage pool.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage technologies, and in particular, to a data management method in a storage system, and also relates to a data management device, an electronic device, a non-volatile storage medium, and a computer program product in a storage system. Background Art

[0002] In a distributed storage system, a storage form in which object data and metadata are separated is usually adopted. The object data is stored on an HDD (Hard Disk Drive, mechanical hard disk), while the metadata is stored on an SSD (Solid State Drive, solid-state hard disk). This design makes a trade-off between performance and cost. The low cost of HDDs: suitable for large-scale storage requirements, especially in scenarios where data is not frequently accessed; the high performance of SSDs: can improve the overall performance of the system through fast access, especially in the case of high-concurrency access. However, with the rapid development of related technologies, the scale of the number of objects continues to expand, and its metadata also grows accordingly. Since the size of the metadata is small, the huge metadata will make it difficult to manage object storage, resulting in a huge overhead for the distributed storage system, that is, continuously increasing the cost and management difficulty of the distributed storage system.

[0003] Therefore, how to more reasonably and effectively manage the stored data and its metadata in the storage system, and further reduce the management cost and management difficulty of the storage system is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0004] The object of the present invention is to provide a data management method in a storage system. This data management method in the storage system can more reasonably and effectively manage the stored data and its metadata in the storage system, and further reduce the management cost and management difficulty of the storage system; another object of the present invention is to provide a data management device, an electronic device, a non-volatile storage medium, and a computer program product in a storage system, all of which have the above beneficial effects.

[0005] In a first aspect, the present invention provides a data management method in a storage system, including:

[0006] Determine the data to be stored; wherein, the data to be stored includes target data and target metadata corresponding to the target data;

[0007] According to a preset division ratio and the access frequency of each piece of the target metadata, divide each piece of the target metadata into hot metadata and cold metadata;

[0008] Save the hot metadata to a high-speed storage pool, and save the target data corresponding to the hot metadata to a low-speed storage pool;

[0009] Save the cold metadata and the target data corresponding to the cold metadata to the low-speed storage pool.

[0010] Among them, saving the hot metadata to the high-speed storage pool includes:

[0011] Save the hot metadata to a preset cache;

[0012] When the data storage amount in the preset cache reaches a preset threshold, flush the hot metadata from the preset cache to the high-speed storage pool.

[0013] Among them, flushing the hot metadata from the preset cache to the high-speed storage pool includes:

[0014] In the preset cache, classify each piece of the hot metadata according to the internal record information of the hot metadata to obtain each hot metadata category;

[0015] For each hot metadata category, linearly merge the index information of each piece of the hot metadata under the hot metadata category to construct a first index corresponding to the hot metadata category;

[0016] Flush each piece of the hot metadata under each hot metadata category from the preset cache to the high-speed storage pool according to each first index.

[0017] Among them, classifying each piece of the hot metadata according to the internal record information of the hot metadata to obtain each hot metadata category includes:

[0018] Extract features from the internal record information of the hot metadata to obtain target features; the target features include locality features and / or request combination features;

[0019] Classify each piece of the hot metadata according to each target feature to obtain each hot metadata category.

[0020] Among them, flushing each piece of the hot metadata under each hot metadata category from the preset cache to the high-speed storage pool according to each first index includes:

[0021] For each piece of the hot metadata, generate a key-value pair using the metadata name of the hot metadata and the hot metadata;

[0022] Flush each key-value pair under each hot metadata category from the preset cache to the high-speed storage pool according to each first index.

[0023] Among them, the data management method in the storage system further includes:

[0024] Determine a first query index according to a first data query instruction;

[0025] Determine the target key-value pair corresponding to the first query index in the high-speed storage pool;

[0026] Determine the target hot metadata according to the target key-value pair;

[0027] Query in the low-speed storage pool to obtain the target query data corresponding to the target hot metadata.

[0028] Wherein, after flushing the hot metadata from the preset cache to the high-speed storage pool, it further includes:

[0029] Delete all the hot metadata in the preset cache.

[0030] Wherein, saving the cold metadata and the target data corresponding to the cold metadata to the low-speed storage pool includes:

[0031] Generate a cold data group corresponding to the cold metadata and the target data corresponding to the cold metadata;

[0032] Save each cold data group to the low-speed storage pool.

[0033] Wherein, saving each cold data group to the low-speed storage pool includes:

[0034] For each cold data group, linearly merge the index information of the cold metadata in the cold data group and the index information of the target data corresponding to the cold metadata to construct a second index corresponding to the cold data group;

[0035] Save each cold data group to the low-speed storage pool according to each second index.

[0036] Wherein, the data management method in the storage system further includes:

[0037] Determine a second query index according to a second data query instruction;

[0038] Determine the target cold data group corresponding to the second query index in the low-speed storage pool;

[0039] Determine the target cold metadata and the target query data corresponding to the target cold metadata according to the target cold data group.

[0040] Wherein, the data management method in the storage system further includes:

[0041] In the high-speed storage pool, when the access frequency of the hot metadata is lower than a first threshold, migrate the hot metadata to the low-speed storage pool;

[0042] In the low-speed storage pool, when the access frequency of the cold metadata exceeds the second threshold, migrate the cold metadata to the high-speed storage pool.

[0043] Among them, migrating the hot metadata to the low-speed storage pool includes:

[0044] Lock the hot metadata to obtain locked metadata;

[0045] Migrate the locked metadata to the low-speed storage pool;

[0046] In the low-speed storage pool, unlock the locked metadata to obtain unlocked metadata.

[0047] Among them, locking the hot metadata to obtain locked metadata includes:

[0048] Apply for a target lock from the lock service node, and use the target lock to lock the hot metadata to obtain the locked metadata;

[0049] Correspondingly, after unlocking the locked metadata to obtain unlocked metadata in the low-speed storage pool, it further includes:

[0050] Return the target lock to the lock service node.

[0051] Among them, after unlocking the locked metadata to obtain unlocked metadata in the low-speed storage pool, it further includes:

[0052] Perform consistency verification on the unlocked metadata;

[0053] When the consistency verification fails, output an alarm prompt.

[0054] Among them, the data management method in the storage system further includes:

[0055] When the consistency verification fails, roll back the data of the unlocked metadata to migrate the unlocked metadata back to the high-speed storage pool.

[0056] Among them, the high-speed storage pool is a solid-state drive storage pool; the low-speed storage pool is a mechanical hard disk storage pool.

[0057] In a second aspect, the present invention also discloses a data management device in a storage system, including:

[0058] A determination module for determining data to be stored; among them, the data to be stored includes target data and target metadata corresponding to the target data;

[0059] A partitioning module, configured to partition each of the target metadata into hot metadata and cold metadata according to a preset partitioning ratio and the access frequency of each of the target metadata;

[0060] A first storage module, configured to store the hot metadata in a high-speed storage pool and store the target data corresponding to the hot metadata in a low-speed storage pool;

[0061] A second storage module, configured to store the cold metadata and the target data corresponding to the cold metadata in the low-speed storage pool.

[0062] In a third aspect, the present invention also discloses an electronic device, including:

[0063] A memory, configured to store a computer program;

[0064] A processor, configured to implement the steps of the data management method in any one of the above storage systems when executing the computer program.

[0065] In a fourth aspect, the present invention also discloses a non-volatile storage medium, on which a computer program is stored, and the computer program implements the steps of the data management method in any one of the above storage systems when executed by a processor.

[0066] In a fifth aspect, the present invention also discloses a computer program product, including computer programs / instructions, and the computer programs / instructions implement the steps of the data management method in any one of the above storage systems when executed by a processor.

[0067] The present invention provides a data management method in a storage system, including: determining data to be stored; wherein, the data to be stored includes target data and target metadata corresponding to the target data; partitioning each of the target metadata into hot metadata and cold metadata according to a preset partitioning ratio and the access frequency of each of the target metadata; storing the hot metadata in a high-speed storage pool and storing the target data corresponding to the hot metadata in a low-speed storage pool; storing the cold metadata and the target data corresponding to the cold metadata in the low-speed storage pool.

[0068] Applying the technical solution provided by the present invention, for the target data to be stored currently and the target metadata corresponding to the target data, all the target metadata to be stored can be divided into hot metadata and cold metadata with reference to the preset hot and cold data division ratio and the access frequency of each target metadata within a specific time period. Since the hot metadata has a high access frequency while the cold metadata has a low access frequency, the hot metadata can be saved to the high-speed storage pool in the storage system, and the cold metadata can be saved to the low-speed storage pool in the storage system. At the same time, the target data corresponding to the hot metadata and the target data corresponding to the cold metadata are also saved to the low-speed storage pool together, thus realizing the reasonable utilization of the internal storage space of the storage system, that is, more reasonable and effective management of the stored data and its metadata in the storage system, and further reducing the management cost and management difficulty of the storage system. On this basis, the hot and cold division of metadata effectively reduces the amount of data stored in the high-speed storage pool, thereby effectively reducing the R & D cost of the storage system.

[0069] The data management device, electronic device, non-volatile storage medium, and computer program product in the storage system provided by the present invention also have the above technical effects, which will not be elaborated herein again. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the technical solutions in the prior art and the embodiments of the present invention, the drawings required for the description of the prior art and the embodiments of the present invention will be briefly introduced below. Of course, the drawings described below for the embodiments of the present invention are only a part of the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings, and the other obtained drawings also fall within the protection scope of the present invention.

[0071] Figure 1 It is a schematic flowchart of a data management method in a storage system provided by an embodiment of the present invention;

[0072] Figure 2 It is an overall architecture diagram of a distributed storage system provided by an embodiment of the present invention;

[0073] Figure 3 It is a schematic diagram of the principle of a data classification and merging method provided by an embodiment of the present invention;

[0074] Figure 4 It is a schematic structural diagram of a data management device in a storage system provided by an embodiment of the present invention;

[0075] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0076] The core of the present invention is to provide a data management method in a storage system. The data management method in the storage system can more reasonably and effectively manage the stored data and its metadata in the storage system, further reducing the management cost and management difficulty of the storage system. Another core of the present invention is to provide a data management device, an electronic device, a non-volatile storage medium, and a computer program product in a storage system, all of which have the above beneficial effects.

[0077] In order to describe the technical solutions in the embodiments of the present invention more clearly and completely, the following will introduce the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0078] The embodiments of the present invention provide a data management method in a storage system.

[0079] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a data management method in a storage system provided by an embodiment of the present invention. The data management method in the storage system may include the following S101 to S104.

[0080] S101: Determine the data to be stored; wherein, the data to be stored includes target data and target metadata corresponding to the target data.

[0081] This step aims to realize the determination of the data to be stored, which is various data that need to be saved to the storage system, including target data and the target metadata corresponding to the target data. For example, in the scenario where a client sends an IO (Input / Output) request (specifically an input request) to the storage system, the data to be stored input by the client can be received through a network management service. Here, the storage system may specifically be a distributed object storage system.

[0082] S102: Divide each target metadata into hot metadata and cold metadata according to a preset division ratio and the access frequency of each target metadata.

[0083] This step aims to realize the hot and cold classification of the target metadata. It can be understood that cold metadata is metadata with a lower access frequency, and hot metadata is metadata with a lower access frequency. Therefore, this process can be realized with reference to the access frequency of each target metadata and the preset hot and cold division ratio.

[0084] In the specific implementation process, first, the access frequency of each target metadata within a preset time period is counted, and each target metadata is sorted in descending / ascending order according to the access frequency to obtain a metadata sequence; then, all target metadata in the metadata sequence is divided into cold metadata and hot metadata according to a preset division ratio.

[0085] It should be noted that the specific values of the preset division ratio and the preset time period do not affect the implementation of this technical solution, and can be set by technicians according to actual needs, and the present invention does not limit this. For example, in a possible implementation manner, the preset division ratio can be specifically 7:3, that is, 7 / 10 of the target metadata is determined as hot metadata, and 3 / 10 of the target metadata is determined as cold metadata; the preset time period can be specifically 60 seconds.

[0086] S103: Save the hot metadata to the high-speed storage pool, and save the target data corresponding to the hot metadata to the low-speed storage pool.

[0087] This step aims to achieve the saving of hot metadata and its target data, that is, save the hot metadata to the high-speed storage pool, and save the target data (hot data) corresponding to the hot metadata to the low-speed storage pool. In a possible implementation manner, the high-speed storage pool can specifically be a solid-state drive storage pool (such as an NVme SSD storage pool); the low-speed storage pool can specifically be a mechanical hard disk storage pool (such as an HDD storage pool).

[0088] In an embodiment of the present invention, saving the hot metadata to the high-speed storage pool may include:

[0089] Save the hot metadata to a preset cache;

[0090] When the data storage amount in the preset cache reaches a preset threshold, flush the hot metadata from the preset cache to the high-speed storage pool.

[0091] The embodiment of the present invention provides a method for implementing saving the hot metadata to the high-speed storage pool, that is, it can be implemented based on a preset cache. Specifically, the hot metadata can be saved to the preset cache first, and after the data storage amount in the preset cache reaches a certain threshold condition, the hot metadata in the preset cache is flushed to the high-speed storage pool in batches, so as to effectively avoid the problem of system instability caused by frequent data flushing operations. Similarly, the specific value of the preset threshold can be set by technicians according to the actual situation, and the present application does not limit this.

[0092] Further, after flushing the hot metadata from the preset cache to the high-speed storage pool, it may further include: deleting all the hot metadata in the preset cache. Thereby, the storage space of the preset cache can be effectively saved, unnecessary storage space occupation can be avoided, and sufficient storage space can be provided for the hot metadata in the subsequent new batch of data to be stored.

[0093] Among them, flushing the hot metadata from the preset cache to the high-speed storage pool may include: within the preset cache, classifying each piece of hot metadata according to the internal record information of the hot metadata to obtain each hot metadata category; for each hot metadata category, linearly combining the index information of each piece of hot metadata under the hot metadata category to construct a first index corresponding to the hot metadata category; and flushing each piece of hot metadata under each hot metadata category from the preset cache to the high-speed storage pool according to each first index.

[0094] The embodiment of the present invention provides a method for implementing flushing hot metadata from a preset cache to a high-speed storage pool, which can effectively improve the storage space utilization rate of the high-speed storage pool and is helpful for realizing faster and more efficient data query. Specifically, the hot metadata can be merged with reference to the hot metadata category, and at the same time, a new hot metadata index (i.e., the above-mentioned first index) can be constructed for it to realize the flushing of the hot metadata based on the new hot metadata index.

[0095] Among them, the classification of the hot metadata can be realized with reference to the internal record information of the hot metadata. For example, a certain type of metadata is used to record the index information of the bucket where the corresponding object is located; a certain type of metadata is used to record the user and bucket information of the corresponding object, such as object encryption, object tags, access object audit logs, and user and bucket policy permissions, etc.; a certain type of metadata is used to record the extended permissions of the corresponding object, such as system extended attributes and custom extended attributes, etc.

[0096] In a possible implementation manner, classifying each piece of hot metadata according to the internal record information of the hot metadata to obtain each hot metadata category may include: extracting features from the internal record information of the hot metadata to obtain target features; the target features include locality features and / or request combination features; and classifying each piece of hot metadata according to each target feature to obtain each hot metadata category. That is to say, the classification of the hot metadata can be realized by extracting the relevant features of the internal record information of the hot metadata. Here, the features extracted from the internal record information of the hot metadata may include:

[0097] (1)Locality characteristics of metadata: Locality generally refers to the relationship between the physical locations of objects and their metadata and the access patterns in a distributed storage system. In many application scenarios, the access to objects is often local. For example, in an Internet video stream processing application, multiple segments of a specific video (object) are usually read or written continuously. By optimizing the storage locations of the metadata of these segments, performance can be improved. Based on this, the locality of object metadata includes metadata indexing, object permission policy control, object multi-version control, and custom extended attributes, that is, the three types of metadata mentioned in the above example.

[0098] (2)Request combination characteristics of metadata: Among the operation requests Get, Put, List, Delete, and Head in object storage, the Get operation is usually combined with the Head operation for a request, that is, when downloading an object, a Head request is first made to the storage to query whether the object exists, and only when it exists, a Get request is made; the Put operation is combined with the List operation for a request, that is, when uploading an object, a List request is first made to query and determine whether the object exists, and if it does not exist, the upload operation is performed.

[0099] In a possible implementation manner, according to each first index, each hot metadata in each hot metadata category can be flushed from a preset cache to a high-speed storage pool, which may include: for each hot metadata, generating a key-value pair by using the metadata name and the hot metadata of the hot metadata; according to each first index, flushing each key-value pair in each hot metadata category from the preset cache to the high-speed storage pool. That is to say, the hot metadata can be saved to the high-speed storage pool in the form of a key-value pair, where the name of the hot metadata is used as the key value and the hot metadata is used as the value, forming a key-value pair key-value to implement the saving of the hot metadata in the high-speed storage pool.

[0100] Based on the above embodiments, the data management method in the storage system provided by the embodiments of the present invention may further include:

[0101] Determining a first query index according to a first data query instruction;

[0102] Determining a target key-value pair corresponding to the first query index in the high-speed storage pool;

[0103] Determining target hot metadata according to the target key-value pair;

[0104] Querying to obtain target query data corresponding to the target hot metadata in the low-speed storage pool.

[0105] An embodiment of the present invention provides a data query method based on a first index. Obviously, according to the query instruction, the index information to be queried can be determined. According to the index to be queried, the key-value pair to be queried can be determined. According to the key-value pair to be queried, the metadata to be queried can be determined. According to the metadata to be queried, the data to be queried can be determined.

[0106] S104: Save the cold metadata and the target data corresponding to the cold metadata to the low-speed storage pool.

[0107] This step aims to save the cold metadata and its target data (cold data), that is, both are saved to the low-speed storage pool.

[0108] In an embodiment of the present invention, saving the cold metadata and the target data corresponding to the cold metadata to the low-speed storage pool may include:

[0109] Generate cold data groups corresponding to the cold metadata and the target data corresponding to the cold metadata;

[0110] Save each cold data group to the low-speed storage pool.

[0111] An embodiment of the present invention provides a method for implementing saving the cold metadata and its target data (cold data) to the high-speed storage pool, that is, they can be saved to the low-speed storage pool together in the form of a "cold metadata - cold data" combination (the cold data corresponds to the cold metadata), which is more convenient for obtaining the metadata and its data object simultaneously during subsequent data query.

[0112] Among them, saving each cold data group to the low-speed storage pool may include: for each cold data group, linearly merge the index information of the cold metadata in the cold data group and the index information of the target data corresponding to the cold metadata to construct a second index corresponding to the cold data group; save each cold data group to the low-speed storage pool according to each second index.

[0113] An embodiment of the present invention provides a method for implementing saving the cold data group to the low-speed storage pool, which can effectively improve the storage space utilization rate of the low-speed storage pool and help achieve faster and more efficient data query. Specifically, for each cold data group, the index information of the cold data and its cold metadata can be directly linearly merged to construct a new index (cold data group index, that is, the above-mentioned second index), and the cold data group can be saved based on the new cold data group index.

[0114] Based on the above embodiments, the data management method in the storage system provided by the embodiments of the present invention may further include:

[0115] Determine a second query index according to the second data query instruction;

[0116] Determine the target cold data group corresponding to the second query index in the low-speed storage pool;

[0117] Determine the target cold metadata and the target query data corresponding to the target cold metadata according to the target cold data group.

[0118] An embodiment of the present invention provides a data query method based on a second index. Similar to the above-mentioned data query method based on the first index, obviously, the to-be-query index information can be determined according to the query instruction, the to-be-query data group can be determined according to the to-be-query index, and the to-be-query metadata and its to-be-query data can be determined according to the to-be-query data.

[0119] It can be seen that for the target data to be stored currently and the target metadata corresponding to the target data in the data management method of the storage system provided by the embodiment of the present invention, all the to-be-stored target metadata can be divided into hot metadata and cold metadata with reference to the preset hot and cold data division ratio and the access frequency of each target metadata within a specific time period. Since the hot metadata has a high access frequency while the cold metadata has a low access frequency, the hot metadata can be saved to the high-speed storage pool in the storage system, and the cold metadata can be saved to the low-speed storage pool in the storage system. At the same time, the target data corresponding to the hot metadata and the target data corresponding to the cold metadata are also saved to the low-speed storage pool together, thereby realizing the reasonable utilization of the internal storage space of the storage system, that is, more reasonable and effective management of the stored data and its metadata in the storage system, and further reducing the management cost and management difficulty of the storage system. On this basis, the hot and cold division of metadata effectively reduces the amount of data stored in the high-speed storage pool, thereby effectively reducing the R & D cost of the storage system.

[0120] Based on the above embodiments:

[0121] In an embodiment of the present invention, the data management method in the storage system may further include:

[0122] In the high-speed storage pool, when the access frequency of the hot metadata is lower than the first threshold, migrate the hot metadata to the low-speed storage pool;

[0123] In the low-speed storage pool, when the access frequency of the cold metadata exceeds the second threshold, migrate the cold metadata to the high-speed storage pool.

[0124] An embodiment of the present invention provides a dynamic migration mechanism for metadata in a storage system. It can be understood that the access frequency of metadata in a storage system generally changes in real time, and the change in access frequency will inevitably lead to a change in the hot and cold types of metadata. Based on this, for each hot metadata in the high-speed storage pool, its access frequency can be monitored regularly / real-time. Once its access frequency decreases and it is converted into cold metadata, it can be migrated to the low-speed storage pool; for each cold metadata in the low-speed storage pool, its access frequency can be monitored regularly / real-time. Once its access frequency increases and it is converted into hot metadata, it can be migrated to the high-speed storage pool. It should be noted that the migration of hot and cold metadata will not cause a change in the storage location of its corresponding target data. Therefore, during the migration of hot and cold metadata, the storage location of the corresponding target data can be updated synchronously. In addition, the specific values of the above first threshold and second threshold do not affect the implementation of this technical solution and can be set by those skilled in the art according to the actual situation. This application does not make any limitations in this regard.

[0125] In a possible implementation manner, migrating the hot metadata to the low-speed storage pool may include: locking the hot metadata to obtain the locked metadata; migrating the locked metadata to the low-speed storage pool; and unlocking the locked metadata in the low-speed storage pool to obtain the unlocked metadata.

[0126] An embodiment of the present application provides a method for implementing the migration of hot metadata to the low-speed storage pool, that is, a method for migrating hot metadata based on a locking mechanism, to effectively ensure data consistency during the migration of hot metadata. During the implementation process, the hot metadata to be migrated can be locked first, and then the locked hot metadata (locked metadata) can be migrated to the low-speed storage pool, and finally it can be unlocked. Of course, the locking mechanism is also applicable to the scenario of migrating cold metadata to the high-speed storage pool.

[0127] Among them, locking the hot metadata to obtain the locked metadata may include: applying for a target lock from the lock service node and using the target lock to lock the hot metadata to obtain the locked metadata; correspondingly, after unlocking the locked metadata in the low-speed storage pool to obtain the unlocked metadata, it may further include: returning the target lock to the lock service node. That is to say, a lock service node can be deployed in the storage system, and the lock service node can implement the distribution and recycling of the target lock. It should be noted that during the distribution of the target lock, the lock service node needs to ensure that there is no locking conflict problem.

[0128] In addition, after unlocking the locked metadata in the low-speed storage pool to obtain the unlocked metadata, the process may further include: performing a consistency check on the unlocked metadata; and outputting an alarm prompt when the consistency check fails. Furthermore, when the consistency check fails, data rollback is performed on the unlocked metadata to migrate the unlocked metadata back to the high-speed storage pool.

[0129] That is to say, in order to further ensure the data consistency before and after data migration, a metadata consistency check mechanism and a data rollback mechanism when the check fails can be added after the metadata is unlocked. Specifically, when the consistency check of the unlocked metadata passes, it can be determined that the metadata migration is successful and no other operations are required; when the consistency check of the unlocked metadata fails, it can be determined that the metadata migration has failed, and at this time, an alarm operation and a metadata rollback operation can be performed.

[0130] The embodiment of the present invention takes a distributed storage system as an example and provides another data management method in a storage system.

[0131] First, please refer to Figure 2 , Figure 2 The overall architecture diagram of a distributed storage system provided by an embodiment of the present invention is based on scalability and fault tolerance. The design concept of the distributed storage system is to eliminate performance bottlenecks such as single-point metadata, adopt a distributed and decentralized design architecture, slice data and distribute them in the form of objects on multiple nodes, and ensure that the load and capacity of each node are balanced. Data access and management are performed through network connections, and a unified storage interface is provided. Since the distributed storage system adopts a decentralized design architecture, it is highly scalable, and the failure of a single storage node will not affect the unavailability of the entire distributed storage cluster; it supports dynamic expansion to add or delete storage nodes, and can automatically relocate data objects in the cluster to balance the storage space utilization between storage nodes.

[0132] Based on this distributed storage system, the distributed storage client initiates an object I / O request to the object storage gateway, and the gateway distributes the request to a unified, self-controlled, and scalable distributed storage consistency management system through an extensible hash calculation algorithm. The management system stores the object's index information, object metadata, and data in the storage pool. In the storage pool, the client's data is stored by slicing, which is fixed to 4MB objects by default. The objects are grouped, and the group ID is the value of the object Hash and the number of groups. Finally, it is stored in the underlying device NVmeSSD or HDD.

[0133] like Figure 2As shown, the distributed storage system architecture includes multiple components, such as object gateway services, a unified, self - controlled, and scalable distributed storage consistency management system, a storage pool, a controllable, scalable, and distributed data balancing placement algorithm, a monitoring service cluster, etc. The functions of each component are introduced as follows:

[0134] 1. Object gateway services: Provide a unified storage interface to offer access interface services for distributed storage clients.

[0135] 2. Unified, self - controlled, and scalable distributed storage consistency management system: The core component of the distributed storage system, which provides storage services in the form of objects. In this component, data is stored as objects, each object has a unique identifier and related data, and it is responsible for storing objects distributively on each node of the storage system, and provides functions such as data replication, recovery, and load balancing to ensure data reliability and high - performance access.

[0136] 3. Storage pool: The distributed storage system consists of multiple storage nodes, and each storage node can contain multiple hard disks or storage devices. The storage pool divides storage resources to form different pools to meet different storage requirements. The data storage pool (low - speed storage pool) is composed of HDDs considering cost; the metadata storage pool (high - speed storage pool) is composed of NVme SSDs to ensure efficient metadata access operations. Both data and metadata are stored in the form of objects on OSDs (object storage devices).

[0137] 4. Controllable, scalable, and distributed data balancing placement algorithm: Evenly distributes data objects across each storage node of the storage cluster to avoid data hotspots and improve system performance.

[0138] 5. Monitoring service cluster: Responsible for monitoring the system status and configuration information to ensure system consistency and availability.

[0139] Secondly, this embodiment proposes a predefined value θ for distinguishing the heat of metadata to distinguish the proportion of hot and cold metadata. By default, θ = 3 / 7, that is, 30% of the metadata is classified as hot metadata (metadata with a higher access frequency), and 70% of the metadata is classified as cold metadata (metadata with a lower access frequency). The predefined value θ can be flexibly adjusted according to the size of the NVmeSSD space in the distributed storage system. Among them, the statistical analysis of the access frequency of metadata during the operation of the storage IO service can be realized by designing a time window, with a default of 60 seconds. On this basis, this embodiment provides different merging methods and writes to different storage media for hot and cold metadata. After merging, the hot metadata establishes a new index structure and is directly stored in the NVmeSSD storage pool; the cold metadata is merged with the data to establish a new index structure and is directly stored in the HDD storage pool. The data merging scheme is as follows:

[0140] 1. Merging method for hot metadata: It should be noted that in a distributed object storage system, to effectively reduce the space occupied by object metadata in NvmeSSD, distributed key-value storage key-value can be used to store object metadata, and RocksDB is used as the distributed key-value storage database. On this basis, since the object metadata with high-frequency access is scattered and persisted in different NvmeSSD storage locations in the RocksDB key-Value storage, that is, the hot metadata is non-linearly stored, this embodiment proposes the following merging scheme: First, after the heat statistics and classification of metadata, before writing to NVmeSSD, the storage locations of the metadata with high-frequency access, that is, the index information, are linearly merged to form a new metadata index structure; then it is stored in NVmeSSD. Among them, the metadata with high-frequency access can be merged by constructing a new key-value pair Key-Value pair, and then the Key-Value pair is linearly sorted. Among them, the object name in the object storage represents the key, that is, Key, and the associated metadata with high access heat is stored in Value. Thus, when the client accesses the hot object, first find the main index of the hot metadata, then find the object metadata Key to be accessed in the main index information, and locate the Value of the metadata through the key, so as to obtain the object data to be accessed.

[0141] 2. Merging method for cold metadata and object data: First, the metadata with low-frequency access is merged with the object data, and an index structure is established; then the merged result is stored in the HDD. Thus, when the client makes an IO request for the metadata with low-frequency access, the metadata and data of the object can be obtained simultaneously.

[0142] It is understandable that the purpose of data merging is to reduce the occupancy of object metadata on the NvmeSSD space. At the same time, during the data balancing process in the failure scenario, the number of metadata reconstructions is reduced to ensure efficient access to metadata within the distributed storage system. Establishing a new index structure is to meet the requirement that the merged metadata objects can achieve faster and more efficient data querying.

[0143] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the principle of a data classification and merging method provided by an embodiment of the present invention. When a client initiates an object storage IO request, Object1, Object2,..., ObjectN are received through the object gateway service and processed in the distributed storage system. Among them:

[0144] 1. Object metadata statistical classification component: It makes statistics according to the characteristic information (locality characteristics and request combination characteristics) of the client object storage IO request, and identifies the data type of the object metadata according to the statistical results for the next stage of processing.

[0145] Among them, the extraction of characteristic information can be realized based on the internal record data of the object metadata. In this embodiment, the data type of the object metadata can be divided into the following three categories:

[0146] (1) Metadata A records the index information of the bucket where the object is located, that is, buckets.index contains the information of the bucket where the object is located.

[0147] (2) Metadata B records the user and bucket information of the object, such as object encryption, object tags, access object audit logs, and user and bucket policy permissions, etc.

[0148] (3) Metadata C records the extended permissions of the object, such as system extended attributes, custom extended attributes, etc. Among them, custom extended attributes allow users to add specific information, with "x-amz-meta-" as the prefix. For example, x-amz-meta-author represents the name of the object creator, and x-amz-meta-description identifies the description information of the object; system extended attributes can include the data type Content-Type of the object, such as image / jpeg, application / json, etc., and can also include the version ID of the object, such as x-amz-version-id.

[0149] In addition, according to the type of object metadata, the metadata can also be sorted in the order from high priority to low priority as metadata A, metadata B, and metadata C.

[0150] It should be noted that the statistics and classification process of the object metadata statistical classification component can be carried out dynamically. When a failure occurs, or capacity expansion is needed, or a large number of metadata are written, the distributed storage system will update the metadata according to the dynamic service requests. Therefore, the classification of the metadata will change accordingly, which will trigger the dynamic migration of the metadata between different storage media, namely the high-speed storage pool and the low-speed storage pool.

[0151] 2. High-speed pool component: It uses NvmeSSD as the persistent storage device and stores the frequently accessed object metadata according to the processing results of the object metadata statistical classification component. For example, it merges the hotness statistics of Obj1.idx, Obj1.meta to Obj4.idx, Obj4.meta to form a new index metadata information of NObj1.idx and NObj1.meta.

[0152] 3. Low-speed pool component: It uses HDD as the underlying storage device and stores the infrequently accessed object metadata according to the processing results of the object metadata statistical classification component. For example, it merges Obj1.idx, Obj1.meta, Obj1.data to Obj4.idx, Obj4.meta, Obj4.data to store the metadata and data of the object together in the HDD.

[0153] Furthermore, based on Figure 2 and Figure 3 the architecture diagram shown, the present embodiment proposes the following storage IO process for the data and metadata of distributed objects:

[0154] 1. When the client initiates an object storage write request, it is first managed in memory, including processes such as the hotness metadata buffer, merging to generate a new object index table, and persistent storage. Among them, the memory space allocation method is:

[0155] ;

[0156] Among them, represents the number of written objects, is the total memory occupancy, with a default of 8GB, which can be flexibly adjusted according to the scale of the distributed storage system, is the memory occupancy of the th object.

[0157] 2. For the written data objects, they are cached in the memory in the form of logs and flushed periodically. Among them, the cache flushing is due to the limited memory space.

[0158] 3. Store the heat metadata in the buffer cache structure and generate corresponding index information.

[0159] 4. Merge the objects in the buffer cache structure to generate new indexes and hot metadata, and record them in the new object index table.

[0160] 5. Merge the metadata and data of low-frequency access objects into large object blocks and write them into the distributed storage HDD storage pool.

[0161] 6. Store the merged heat metadata and object indexes in the NVmeSSD storage pool in the form of Key-Value.

[0162] Thus, the logic of the storage I / O process for distributed object data and metadata is as follows:

[0163] "For the object data log written by the client do;

[0164] If the cache is flushed down then;

[0165] Store the heat metadata in the buffer cache structure;

[0166] Merge the objects in the buffer cache structure to generate new indexes and hot metadata, and record them in the new object index table;

[0167] End if;

[0168] If the capacity of the buffer cache structure exceeds the threshold then;

[0169] While the buffer is not empty do;

[0170] Merge the metadata and data of low-frequency access objects into large object blocks and write them into the distributed storage HDD storage pool;

[0171] Store the merged heat metadata and object indexes in the NVmeSSD storage pool in the form of Key-Value;

[0172] Delete the cached objects in the buffer;

[0173] End while;

[0174] End if;

[0175] End for".

[0176] Finally, when a distributed storage system performs object IO storage, there are usually business fluctuations or changes. In this regard, this embodiment proposes a method for dynamically migrating object metadata. In the memory, the access frequency of the object metadata is obtained based on statistics. When the access frequency of the cold object metadata exceeds the predefined threshold, the object metadata is determined to be hot metadata and is migrated back from the HDD storage pool to the NVmeSSD storage pool. Conversely, hot metadata can also be dynamically migrated to the HDD storage pool based on the access frequency obtained from statistics. This process can be implemented based on the object metadata lock mechanism to ensure the consistency of the object metadata during migration. Based on this, the implementation process is as follows:

[0177] 1. Set up the object metadata lock mechanism in the distributed storage system: The implementation principle is to use cluster locks to protect access to metadata. The role of the lock is to ensure that during the migration process, only one process can read and write specific metadata. Specifically, the cluster lock method includes: setting the first node in the distributed storage system as a lock service node, and all requests that need to obtain locks need to apply to this lock service node; when a request is made for a lock, the lock service node will check whether there are other nodes currently holding the lock. If not, the lock will be assigned to the requesting node; if so, the request will be rejected. When the node holding the lock completes the operation, it will return the lock to the lock service node, and other nodes can apply for the lock again.

[0178] 2. Preparation before migration: Get the list of metadata to be migrated and lock the metadata to prevent other operations from modifying it.

[0179] 3. Perform migration: First, obtain the lock, then read the metadata from the NVmeSSD storage pool, lock it and write it to the HDD storage pool. After the migration is complete, update the storage location of the metadata.

[0180] 4. Unlock and consistency check: Unlock to allow other operations to access the migrated metadata. At the same time, perform consistency check on the metadata to ensure that the metadata can be accessed and used normally in the new storage pool.

[0181] 5. Error handling: During the migration process, if an error occurs, the migrated metadata can be rolled back to ensure metadata consistency and ensure that the failed migration does not affect the correctness of other concurrent operations.

[0182] Among them, during the dynamic migration process, the overhead of metadata migration, that is, the total migration time as follows:

[0183] ;

[0184] in, is the migration time, is the locking time (i.e., the time required for all operations on metadata during migration), is the time for consistency check and verification, is the time for other operations (if there is no migration, the time required for other operations on metadata by the application), is the maximum number of metadata in the system, is the number of metadata being migrated.

[0185] It can be seen that for the target data to be stored currently and the target metadata corresponding to the target data, the data management method in the storage system provided by the embodiments of the present invention can divide all the target metadata to be stored into hot metadata and cold metadata with reference to the preset hot and cold data division ratio and the access frequency of each target metadata within a specific time period. Since the hot metadata has a high access frequency while the cold metadata has a low access frequency, the hot metadata can be saved to the high-speed storage pool in the storage system, the cold metadata can be saved to the low-speed storage pool in the storage system, and at the same time, the target data corresponding to the hot metadata and the target data corresponding to the cold metadata are also saved to the low-speed storage pool, thus realizing the reasonable utilization of the internal storage space of the storage system, that is, more reasonable and effective management of the stored data and its metadata in the storage system, and further reducing the management cost and management difficulty of the storage system. On this basis, the hot and cold division of metadata effectively reduces the amount of data stored in the high-speed storage pool, thereby effectively reducing the R & D cost of the storage system.

[0186] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a data management device in a storage system provided by the present invention. The data management device in the storage system may include:

[0187] Determination module 1, configured to determine the data to be stored; wherein, the data to be stored includes target data and target metadata corresponding to the target data;

[0188] Division module 2, configured to divide each target metadata into hot metadata and cold metadata according to a preset division ratio and the access frequency of each target metadata;

[0189] First saving module 3, configured to save the hot metadata to the high-speed storage pool and save the target data corresponding to the hot metadata to the low-speed storage pool;

[0190] Second saving module 4, configured to save the cold metadata and the target data corresponding to the cold metadata to the low-speed storage pool.

[0191] It can be seen that for the target data to be stored currently and the target metadata corresponding to the target data in the data management device in the storage system provided by the embodiment of the present invention, all the target metadata to be stored can be divided into hot metadata and cold metadata with reference to the preset hot and cold data division ratio and the access frequency of each target metadata within a specific time period. Since the hot metadata has a high access frequency while the cold metadata has a low access frequency, the hot metadata can be saved to the high-speed storage pool in the storage system, the cold metadata can be saved to the low-speed storage pool in the storage system, and at the same time, the target data corresponding to the hot metadata and the target data corresponding to the cold metadata are also saved to the low-speed storage pool together, thereby realizing the reasonable utilization of the internal storage space of the storage system, that is, more reasonable and effective management of the stored data and its metadata in the storage system, and further reducing the management cost and management difficulty of the storage system. On this basis, the hot and cold division of metadata effectively reduces the amount of data stored in the high-speed storage pool, thereby effectively reducing the R & D cost of the storage system.

[0192] In an embodiment of the present invention, the above first saving module 3 may include:

[0193] A cache unit for saving the hot metadata to a preset cache;

[0194] A downbrushing unit for, when the data storage amount in the preset cache reaches a preset threshold, downbrushing the hot metadata from the preset cache to the high-speed storage pool.

[0195] In an embodiment of the present invention, the above downbrushing unit may include:

[0196] A classification subunit for classifying each hot metadata according to the internal record information of the hot metadata in the preset cache to obtain each hot metadata category;

[0197] A merging subunit for, for each hot metadata category, linearly merging the index information of the hot metadata under the hot metadata category to construct a first index corresponding to the hot metadata category;

[0198] A downbrushing subunit for downbrushing each hot metadata under each hot metadata category from the preset cache to the high-speed storage pool according to each first index.

[0199] In an embodiment of the present invention, the above classification subunit may specifically be used for extracting features from the internal record information of the hot metadata to obtain target features; the target features include locality features and / or request combination features; and classifying each hot metadata according to each target feature to obtain each hot metadata category.

[0200] In an embodiment of the present invention, the above-mentioned lower brush unit may be specifically configured to, for each piece of thermal metadata, generate a key-value pair by using the metadata name of the thermal metadata and the thermal metadata; and flush each key-value pair under each thermal metadata category from a preset cache to a high-speed storage pool according to each first index.

[0201] In an embodiment of the present invention, the data management in the storage system may further include a first query module, configured to determine a first query index according to a first data query instruction; determine a target key-value pair corresponding to the first query index in the high-speed storage pool; determine target thermal metadata according to the target key-value pair; and query target query data corresponding to the target thermal metadata in the low-speed storage pool.

[0202] In an embodiment of the present invention, the above-mentioned first saving module 3 may further include a deletion unit, configured to delete all the thermal metadata in the preset cache after flushing the thermal metadata from the preset cache to the high-speed storage pool.

[0203] In an embodiment of the present invention, the above-mentioned second saving module 4 may include:

[0204] A combining unit, configured to correspondingly generate a cold data group by combining cold metadata and target data corresponding to the cold metadata;

[0205] A saving unit, configured to save each cold data group to the low-speed storage pool.

[0206] In an embodiment of the present invention, the above-mentioned saving unit may be specifically configured to, for each cold data group, linearly merge the index information of the cold metadata in the cold data group and the index information of the target data corresponding to the cold metadata to construct a second index corresponding to the cold data group; and save each cold data group to the low-speed storage pool according to each second index.

[0207] In an embodiment of the present invention, the data management in the storage system may further include a second query module, configured to determine a second query index according to a second data query instruction; determine a target cold data group corresponding to the second query index in the low-speed storage pool; and determine target cold metadata and target query data corresponding to the target cold metadata according to the target cold data group.

[0208] In an embodiment of the present invention, the data management in the storage system may further include a migration module, configured to, in the high-speed storage pool, when the access frequency of the thermal metadata is lower than a first threshold, migrate the thermal metadata to the low-speed storage pool; and in the low-speed storage pool, when the access frequency of the cold metadata exceeds a second threshold, migrate the cold metadata to the high-speed storage pool.

[0209] In an embodiment of the present invention, the above-mentioned migration module may include:

[0210] A locking unit, configured to lock the hot metadata to obtain locked metadata;

[0211] A migration unit, configured to migrate the locked metadata to a low-speed storage pool;

[0212] An unlocking unit, configured to unlock the locked metadata in the low-speed storage pool to obtain unlocked metadata.

[0213] In an embodiment of the present invention, the above-mentioned locking unit may be specifically configured to apply for a target lock from a lock service node, and use the target lock to lock the hot metadata to obtain locked metadata;

[0214] Correspondingly, the above-mentioned migration module may further include a lock-returning unit, configured to return the target lock to the lock service node after unlocking the locked metadata in the low-speed storage pool to obtain unlocked metadata.

[0215] In an embodiment of the present invention, the above-mentioned migration module may further include a verification unit, configured to perform consistency verification on the unlocked metadata after unlocking the locked metadata in the low-speed storage pool to obtain unlocked metadata; when the consistency verification fails, an alarm prompt is output.

[0216] In an embodiment of the present invention, the above-mentioned migration module may further include a rollback unit, configured to perform data rollback on the unlocked metadata when the consistency verification fails, so as to migrate the unlocked metadata back to the high-speed storage pool.

[0217] In an embodiment of the present invention, the high-speed storage pool may be a solid-state drive storage pool; the low-speed storage pool may be a mechanical hard drive storage pool.

[0218] For the introduction of the device provided in the embodiments of the present invention, please refer to the above method embodiments, and the present invention will not be elaborated herein.

[0219] Embodiments of the present invention provide an electronic device.

[0220] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided by the present invention. The electronic device may include:

[0221] A memory 11, configured to store a computer program;

[0222] A processor 10, configured to implement the steps of the data management method in any of the above storage systems when executing the computer program.

[0223] As Figure 5As shown, it is a schematic diagram of the composition structure of an electronic device. The electronic device may include: a processor 10, a memory 11, a communication interface 12, and a communication bus 13. The processor 10, the memory 11, and the communication interface 12 all complete mutual communication through the communication bus 13.

[0224] In an embodiment of the present invention, the processor 10 may be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices, etc.

[0225] The processor 10 may call the program stored in the memory 11. Specifically, the processor 10 may execute the operations in the embodiments of the data management method in the storage system.

[0226] The memory 11 is used to store one or more programs. The program may include program codes, and the program codes include computer operation instructions. In an embodiment of the present invention, the memory 11 stores at least a program for implementing the following functions:

[0227] Determine the data to be stored; wherein, the data to be stored includes target data and target metadata corresponding to the target data;

[0228] According to a preset division ratio and the access frequency of each target metadata, divide each target metadata into hot metadata and cold metadata;

[0229] Save the hot metadata to the high-speed storage pool, and save the target data corresponding to the hot metadata to the low-speed storage pool;

[0230] Save the cold metadata and the target data corresponding to the cold metadata to the low-speed storage pool.

[0231] In a possible implementation manner, the memory 11 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function, etc.; the data storage area may store the data created during use.

[0232] In addition, the memory 11 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device or other volatile solid-state storage devices.

[0233] The communication interface 12 may be an interface of a communication module, and is used to connect to other devices or systems.

[0234] Of course, it should be noted that Figure 5 the structure shown does not constitute a limitation to the electronic device in the embodiment of the present invention. In actual applications, the electronic device may include more Figure 5More or fewer components as shown, or combine certain components.

[0235] An embodiment of the present invention provides a non - volatile storage medium.

[0236] The computer program stored on the non - volatile storage medium provided by the embodiment of the present invention, when executed by a processor, can implement the steps of the data management method in any of the above - mentioned storage systems.

[0237] Among them, the non - volatile storage medium can be any available medium that a computer can store or a data storage device such as a server or a data center that integrates one or more available media. For example, it can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape, etc.), an optical medium (such as a DVD), or a semiconductor medium (such as a solid - state drive), etc., which are various media that can store computer program code.

[0238] For the introduction of the non - volatile storage medium provided by the embodiment of the present invention, please refer to the above - mentioned method embodiment, and the present invention will not elaborate here.

[0239] An embodiment of the present invention provides a computer program product.

[0240] The computer program product provided by the embodiment of the present invention includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, they can implement the steps of the data management method in any of the above - mentioned storage systems.

[0241] Specifically, in the above - mentioned embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0242] Among them, the computer program product can include one or more computer programs / instructions. When the computer programs / instructions are loaded and executed on a computer, they can generate, in whole or in part, the processes or functions described in the embodiments of the present invention. The computer can be a general - purpose computer, a special - purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a non - volatile storage medium or transmitted from one non - volatile storage medium to another non - volatile storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center by wire (such as coaxial cable, optical fiber, digital subscriber line, etc.) or wirelessly (such as infrared, wireless, microwave, etc.).

[0243] For the introduction of the computer program product provided by the embodiment of the present invention, please refer to the above - mentioned method embodiment, and the present invention will not elaborate here.

[0244] The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0245] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0246] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0247] The technical solutions provided by the present invention have been introduced in detail above. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A data management method in a storage system, characterized in that: include: Determine the data to be stored; wherein the data to be stored includes target data and target metadata corresponding to the target data; According to a preset division ratio and an access frequency of each target metadata, each target metadata is divided into hot metadata and cold metadata; The hot metadata is saved to a preset cache; when the amount of data storage in the preset cache reaches a preset threshold, each of the hot metadata is classified in the preset cache according to the internal record information of the hot metadata to obtain each hot metadata category; for each hot metadata category, the index information of each of the hot metadata under the hot metadata category is linearly merged to construct a first index corresponding to the hot metadata category; for each hot metadata, a key-value pair is generated using the metadata name of the hot metadata and the hot metadata, the metadata name of the hot metadata is used as the key value of the key-value pair, and the hot metadata is used as the value of the key-value pair; according to each first index, each of the key-value pairs under each of the hot metadata categories is flushed from the preset cache to the high-speed storage pool; and the target data corresponding to the hot metadata is saved to the low-speed storage pool; Generate a cold data group by corresponding the cold metadata and the target data corresponding to the cold metadata; for each cold data group, linearly merge the index information of the cold metadata in the cold data group and the index information of the target data corresponding to the cold metadata to construct a second index corresponding to the cold data group; and save each of the cold data groups to the low-speed storage pool according to each of the second indexes.

2. The data management method in a storage system according to claim 1, characterized in that: Classifying each of the hot metadata according to the internal record information of the hot metadata to obtain each hot metadata category includes: Extracting features from the internal record information of the hot metadata to obtain target features; the target features include local features and / or request combination features; The hot metadata are classified according to the target features to obtain categories of the hot metadata.

3. The data management method in a storage system according to claim 1, characterized in that: Also includes: Determine a first query index according to the first data query instruction; Determining a target key-value pair corresponding to the first query index in the high-speed storage pool; Determine target hot metadata according to the target key-value pair; The target query data corresponding to the target hot metadata is obtained by querying in the low-speed storage pool.

4. The data management method in a storage system according to claim 1, characterized in that: After flushing the hot metadata from the preset cache to the high-speed storage pool, the method further includes: All the hot metadata in the preset cache are deleted.

5. The data management method in a storage system according to claim 1, characterized in that: Also includes: Determine a second query index according to the second data query instruction; Determining a target cold data group corresponding to the second query index in the low-speed storage pool; Target cold metadata and target query data corresponding to the target cold metadata are determined according to the target cold data group.

6. The data management method in a storage system according to any one of claims 1 to 5, characterized in that: Also includes: In the high-speed storage pool, when the access frequency of the hot metadata is lower than a first threshold, migrating the hot metadata to the low-speed storage pool; In the low-speed storage pool, when the access frequency of the cold metadata exceeds a second threshold, the cold metadata is migrated to the high-speed storage pool.

7. The data management method in a storage system according to claim 6, characterized in that: Migrating the hot metadata to the low-speed storage pool includes: Locking the hot metadata to obtain locked metadata; Migrating the locking metadata to the low-speed storage pool; In the low-speed storage pool, the locked metadata is unlocked to obtain unlocked metadata.

8. The data management method in a storage system according to claim 7, characterized in that: The hot metadata is locked to obtain the locked metadata, including: Applying for a target lock from a lock service node, and using the target lock to lock the hot metadata to obtain the locked metadata; Correspondingly, in the low-speed storage pool, after the locked metadata is unlocked to obtain the unlocked metadata, the method further includes: Return the target lock to the lock service node.

9. The data management method in a storage system according to claim 7, characterized in that: In the low-speed storage pool, after unlocking the locked metadata to obtain unlocked metadata, the method further includes: Performing consistency check on the unlocking metadata; When the consistency check fails, an alarm is output.

10. The data management method in a storage system according to claim 9, characterized in that: Also includes: When the consistency check fails, data rollback is performed on the unlocking metadata to migrate the unlocking metadata back to the high-speed storage pool.

11. The data management method in a storage system according to claim 1, characterized in that: The high-speed storage pool is a solid-state hard disk storage pool; the low-speed storage pool is a mechanical hard disk storage pool.

12. A data management device in a storage system, characterized in that: include: A determination module, used to determine data to be stored; wherein the data to be stored includes target data and target metadata corresponding to the target data; A division module, used for dividing each target metadata into hot metadata and cold metadata according to a preset division ratio and an access frequency of each target metadata; The first saving module is used to save the hot metadata to a preset cache; when the data storage volume in the preset cache reaches a preset threshold, in the preset cache, each hot metadata is classified according to the internal record information of the hot metadata to obtain each hot metadata category; for each hot metadata category, the index information of each hot metadata under the hot metadata category is linearly merged to construct a first index corresponding to the hot metadata category; for each hot metadata, a key-value pair is generated using the metadata name of the hot metadata and the hot metadata, the metadata name of the hot metadata is used as the key value of the key-value pair, and the hot metadata is used as the value of the key-value pair; according to each first index, each key-value pair under each hot metadata category is flushed from the preset cache to the high-speed storage pool; and the target data corresponding to the hot metadata is saved to the low-speed storage pool; The second saving module is used to generate a cold data group by corresponding the cold metadata and the target data corresponding to the cold metadata; for each cold data group, linearly merge the index information of the cold metadata in the cold data group and the index information of the target data corresponding to the cold metadata to construct a second index corresponding to the cold data group; and save each of the cold data groups to the low-speed storage pool according to each of the second indexes.

13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data management method in the storage system according to any one of claims 1 to 11 when executing the computer program.

14. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data management method in the storage system according to any one of claims 1 to 11 are implemented.

15. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the data management method in the storage system according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Data migration method and device

    CN106855871A

  • Data migration method and device, equipment and medium

    CN115061630A

  • Column storage indexing method and device based on nonvolatile memory

    CN116257523A

  • Data migration method and device, computer equipment and storage medium

    CN116821102A

  • Metadata management method and device, electronic equipment, medium and chip

    CN118170718A