A data storage method, apparatus, system, device, medium, and program product

Through the customized persistent key-value storage component directly interacts with the underlying storage device, the problem of poor compatibility of interfaces between different storage devices is solved and data storage performance is improved.

CN120233954BActive Publication Date: 2025-08-05JINAN INSPUR DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510705336.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-05
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In the existing distributed storage architecture, due to the different interfaces of different storage devices, the RocksDB interface is poor, resulting in low data storage performance.

Method used

Design custom persistent key-value storage components to interact directly with the underlying storage devices, control hardware resources through asynchronous input and output algorithms, direct input and output algorithms, and data mapping algorithms, determine storage locations and convert data formats, and directly pass data through kernel caches.

Benefits of technology

Improve data storage performance, reduce intermediate layer performance losses, and achieve efficient interaction with storage devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233954B_ABST
    Figure CN120233954B_ABST
Patent Text Reader

Abstract

The present invention discloses a data storage method, apparatus, system, device, medium and program product, which are applied to the field of distributed storage technology, including: obtaining a customized persistent key-value storage component; wherein the customized persistent key-value storage component is a component including an interface for directly interacting with an underlying storage device; determining the storage location corresponding to the data to be stored based on the customized persistent key-value storage component, and converting the data format of the data to be stored into target write data; storing the target write data to the storage device corresponding to the storage location based on the customized persistent key-value storage component. Compared with the current need to create a storage layer abstraction before interacting with the storage device, the present invention improves the performance of data storage by designing a customized persistent key-value storage component that can directly interact with the storage device, so that the persistent key-value storage component can directly control the underlying storage device and directly interact with the underlying storage device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed storage technology, and in particular to a data storage method, device, system, equipment, medium and program product. Background Art

[0002] In a distributed storage architecture, storage devices are the components responsible for storing actual data and managing the storage and retrieval of data objects. RocksDB is an embeddable, persistent key-value store (KV-store) for fast data storage. RocksDB provides high-performance key-value storage for OSDs (Object Storage Devices), enabling storage devices to efficiently handle read and write operations, ensuring data reliability and performance. However, the current proliferation of storage devices, coupled with the varying interfaces of different storage devices, results in poor RocksDB interface compatibility. This necessitates multiple abstraction layers for data conversion, resulting in lower data storage performance.

[0003] It can be seen that how to improve the performance of data storage is a technical problem that those skilled in the art need to solve urgently. Summary of the Invention

[0004] In view of this, an object of the present invention is to provide a data storage method, apparatus, system, device, medium and program product, which solve the technical problem of low data storage performance in the prior art.

[0005] To solve the above technical problems, the present invention provides a data storage method, comprising:

[0006] Obtaining a custom persistent key-value storage component; wherein the custom persistent key-value storage component is a component including an interface for directly interacting with an underlying storage device;

[0007] Determine a storage location corresponding to the data to be stored based on the customized persistent key-value storage component, and convert the data format of the data to be stored into target write data;

[0008] The target is written into the storage device corresponding to the storage location based on the customized persistent key-value storage component.

[0009] On the one hand, before obtaining a custom persistent key-value storage component, it also includes:

[0010] Customize the target storage interface; wherein the target storage interface is used to directly control the hardware resources of each storage device;

[0011] The target storage interface is integrated with the persistent key-value storage component to obtain the customized persistent key-value storage component.

[0012] On the one hand, the implementation process of the custom target storage interface includes:

[0013] Determine an algorithm for controlling hardware resources; wherein the algorithm for controlling hardware resources is an algorithm for directly manipulating resources of underlying hardware devices through software logic;

[0014] The algorithm for controlling hardware resources is integrated into the interface to obtain the target storage interface.

[0015] In one aspect, the algorithm for controlling hardware resources includes at least one of an asynchronous input-output algorithm, a direct input-output algorithm, and a data mapping algorithm;

[0016] The asynchronous input and output algorithm is an algorithm that serializes key-value pairs and directly submits asynchronous requests to the storage device;

[0017] The direct input and output algorithm is an algorithm that bypasses the kernel cache and directly controls the storage device;

[0018] The data mapping algorithm is an algorithm for mapping the physical address space of the storage device to the virtual address space of the customized persistent key-value storage component and directly reading and writing data through a pointer.

[0019] On the one hand, determining a storage location corresponding to the data to be stored based on the customized persistent key-value storage component and converting the data format of the data to be stored into target write data includes:

[0020] Determining a target protocol corresponding to the data to be stored using the asynchronous input / output algorithm in the customized persistent key-value storage component; wherein the target protocol is a protocol corresponding to a target storage device;

[0021] The asynchronous input-output algorithm is used to perform serialization processing on the data to be stored based on the target protocol to obtain the target write data; the serialization processing includes data verification, data compression and metadata encapsulation.

[0022] On the one hand, storing the target data in a storage device corresponding to the storage location based on the customized persistent key-value storage component includes:

[0023] The target write data is directly transferred to the storage device by utilizing a direct read / write mode flag for bypassing kernel cache in the direct input / output algorithm.

[0024] On the one hand, determining a storage location corresponding to the data to be stored based on the customized persistent key-value storage component includes:

[0025] During a write operation, determining an optimal physical write location based on the semantics and access pattern of the data to be stored;

[0026] Using the optimal writing physical location as the storage location;

[0027] The self-defined persistent key-value storage component is used to map the logical key and the storage location to obtain a mapping structure, so as to write data based on the mapping structure.

[0028] On the one hand, during a write operation, determining the optimal physical write location based on the semantics and access pattern of the data to be stored includes:

[0029] Obtaining a semantic type and an access mode type of the data to be stored; wherein the semantic type includes at least one of a data type and a business scenario, and the access mode type is determined based on read and write behavior characteristics of the data on the storage device;

[0030] A physical location that maximizes input and output efficiency is determined based on the semantic type and the access mode type to obtain the optimal write physical location.

[0031] On the one hand, before obtaining the semantic type and access mode type of the data to be stored, the method further includes:

[0032] Obtaining access frequency statistics, sequential statistics, and concurrent pressure statistics corresponding to the data to be stored;

[0033] The access pattern type is determined based on the access frequency statistical information, sequential statistical information and concurrent pressure statistical information; wherein the access pattern type includes low-frequency access mode, high-frequency access mode, sequential access mode, random access mode, high concurrency mode and low concurrency mode.

[0034] In one aspect, the data storage method further includes:

[0035] Get partition information;

[0036] Based on the partition information, the storage device is divided into corresponding partitions using a partition strategy; wherein the partition strategy is a strategy that enables data to be distributed according to a set rule;

[0037] Accordingly, storing the target data in a storage device corresponding to the storage location based on the customized persistent key-value storage component includes:

[0038] Determining a target partition corresponding to the target write data based on the customized persistent key-value storage component;

[0039] The target write data is written into a storage device corresponding to the storage location in the target partition.

[0040] On the one hand, the partition information includes at least one of a storage device identifier, an access mode, and whether the data is hot data.

[0041] On the one hand, after the storage device is divided into corresponding partitions using a partition strategy based on the partition information, the method further includes:

[0042] Detect workload information;

[0043] The divided partitions are dynamically updated based on the workload information.

[0044] On the one hand, after the target is written into the storage device corresponding to the storage location based on the customized persistent key-value storage component, the method further includes:

[0045] Obtain the status, load and wear level of each partition according to the set time period;

[0046] Determine garbage collection parameters for each partition using a machine learning model based on the state, load, and wear level of each partition;

[0047] The garbage collection parameter is compared with the garbage collection threshold, and the partition data is cleaned up.

[0048] On the one hand, the partitioning strategy includes at least one of a cold and hot data separation strategy, a dynamic data distribution strategy, and a storage device wear leveling strategy.

[0049] On the one hand, the storage device includes at least one of a non-volatile memory high-speed protocol hard disk, a serial transmission solid-state hard disk, a mechanical hard disk and an object storage device.

[0050] In one aspect, the data storage method further includes:

[0051] Loading the target data to be restored in the target snapshot using the snapshot directory based on the customized persistent key-value storage component;

[0052] A data mapping relationship in the customized persistent key-value storage component is obtained, and the target data to be restored is restored to a corresponding storage device using the customized persistent key-value storage component based on the data mapping relationship.

[0053] On the one hand, based on the customized persistent key-value storage component, using the snapshot directory, loading the target data to be restored in the target snapshot includes:

[0054] When a storage device fails, a snapshot directory at a recent time point is selected based on the customized persistent key-value storage component, and the target data to be restored in the target snapshot is loaded.

[0055] In one aspect, the data storage method further includes:

[0056] Get data merging conditions;

[0057] The customized persistent key-value storage component performs data merging using the data merging condition.

[0058] On the one hand, the data merging condition includes at least one of a data merging condition based on a data access pattern, a data merging condition based on a data update frequency, and a data merging condition based on a data size, wherein the data merging condition based on a data access pattern is triggered based on an access frequency threshold.

[0059] On the one hand, obtaining a data merging condition, and merging data based on the customized persistent key-value storage component using the data merging condition, including:

[0060] When the data merging condition is a data merging condition based on a data access pattern, determining whether the data access frequency of the data to be merged is lower than a preset data access frequency threshold; if it is lower than the data access frequency threshold, merging the data to be merged;

[0061] When the data merging condition is based on data update frequency, determining whether the update frequency of the data to be merged is lower than a preset minimum update frequency; if it is lower than the minimum update frequency, merging the data to be merged;

[0062] When the data merging condition is a data merging condition based on data size, it is determined whether the size of the data to be merged is lower than a preset minimum number of bytes; when the size of the data to be merged is lower than the minimum number of bytes, the data to be merged is merged.

[0063] An embodiment of the present invention further provides a data storage device, comprising:

[0064] A component acquisition module, configured to acquire a custom persistent key-value storage component; wherein the custom persistent key-value storage component is a component including an interface for directly interacting with an underlying storage device;

[0065] a target write data determination module, configured to determine a storage location corresponding to the data to be stored based on the user-defined persistent key-value storage component, and convert the data format of the data to be stored into the target write data;

[0066] A storage module is used to store the target write data in a storage device corresponding to the storage location based on the customized persistent key-value storage component.

[0067] An embodiment of the present invention further provides a data storage system, comprising:

[0068] A custom persistent key-value storage component is used to determine a storage location corresponding to the data to be stored, convert the data format of the data to be stored into target write data, and store the target write data in a storage device corresponding to the storage location;

[0069] A storage device is used to store the target write data.

[0070] An embodiment of the present invention further provides a data storage device, comprising:

[0071] memory for storing computer programs;

[0072] A processor is used to execute the computer program to implement the steps of the above data storage method.

[0073] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned data storage method are implemented.

[0074] An embodiment of the present invention further provides a program product (the program product is a computer program product), including a computer program / instruction, which implements the steps of the above-mentioned data storage method when executed by a processor.

[0075] To solve the above technical problems, an embodiment of the present invention provides a data storage method, including: obtaining a customized persistent key-value storage component; wherein the customized persistent key-value storage component is a component including an interface for directly interacting with an underlying storage device; determining a storage location corresponding to the data to be stored based on the customized persistent key-value storage component, and converting the data format of the data to be stored into target write data; and storing the target write data to the storage device corresponding to the storage location based on the customized persistent key-value storage component.

[0076] It can be seen from the above technical solution that the beneficial effect of the present invention is: compared with the current need to convert the data format into a storage form suitable for the storage device, serialize the data structure into a byte stream suitable for the storage device, and create a storage layer abstraction to interact with the storage device, the present invention designs a custom persistent key-value storage component that can directly interact with the storage device, so that the persistent key-value storage component can directly control the underlying storage device without conversion, thereby improving the performance of data storage. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0078] Figure 1 A flowchart of a data storage method provided by an embodiment of the present invention;

[0079] Figure 2 A flowchart of another data storage method provided by an embodiment of the present invention;

[0080] Figure 3 A flowchart of a data storage method provided by an embodiment of the present invention;

[0081] Figure 4 A structural framework diagram of a distributed storage system provided by an embodiment of the present invention;

[0082] Figure 5 A schematic diagram of the structural framework of a data storage device provided by an embodiment of the present invention;

[0083] Figure 6 A schematic diagram of the structural framework of a data storage system provided by an embodiment of the present invention;

[0084] Figure 7 A schematic diagram of the structural framework of a data storage device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0085] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0086] The terms "including" and "having," as used in the present description and accompanying drawings, and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements and may include steps or elements that are not listed.

[0087] In order to enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0088] Next, a data storage method provided by an embodiment of the present invention is described in detail. Figure 1 A flowchart of a data storage method provided in an embodiment of the present invention may include:

[0089] S101, obtaining a customized persistent key-value storage component; wherein the customized persistent key-value storage component is a component including an interface for directly interacting with an underlying storage device.

[0090] The execution subject in this embodiment is an electronic device, which can be a storage pool, a distributed storage system, a storage backend engine, etc. The customized persistent key-value storage component in this embodiment refers to a component that can directly interact with the underlying storage device. This component is an improvement on the existing persistent key-value storage component. When the execution subject is a customized persistent key-value storage component, it can directly determine the storage location corresponding to the data to be stored, convert the data format of the data to be stored into the target write data, and store the target write data in the storage device corresponding to the storage location. The reason why this embodiment adds the step of obtaining the customized persistent key-value storage component is to describe the customized persistent key-value storage component as a whole. The customized persistent key-value storage component in this embodiment is a customized RocksDB (persistent key-value storage component), which processes the address mapping and data storage format of the physical disk in RocksDB, so that it can directly interact with the underlying storage device. Among them, the address mapping of the physical disk refers to the mapping of the storage location corresponding to the data to be stored, and the data storage form refers to the determination of the data format of the data to be stored so as to meet the format required by the storage device to store the data; or the data storage form in this embodiment can dynamically adjust the data storage format based on the access frequency, for example, hot data is not compressed, and cold data is compressed. Hot data is data with high frequency access, strong real-time performance, and sensitive to delay, and cold data is data with an access frequency lower than the minimum access threshold.

[0091] It should be further explained that, in order to improve the accuracy of determining the custom persistent key-value storage component, before obtaining the custom persistent key-value storage component, the following steps may also be included:

[0092] S1011, customizing a target storage interface; wherein the target storage interface is used to directly control hardware resources of each storage device;

[0093] S1012: Integrate the target storage interface with the persistent key-value storage component to obtain a customized persistent key-value storage component.

[0094] This embodiment first customizes the target storage interface and determines algorithms for controlling hardware resources, such as asynchronous I / O algorithms, direct I / O algorithms, and data mapping algorithms. These algorithms directly manipulate the resources of the underlying hardware devices through software logic. The hardware resource control algorithms are then integrated into the interface to obtain the target storage interface, enabling the target storage interface to directly control the hardware resources of each storage device. This target storage interface is then integrated with a persistent key-value storage component (RocksDB) to obtain a customized persistent key-value storage component. In this embodiment, interface integration enables the persistent key-value storage component to directly control the hardware resources of each storage device, thereby improving the efficiency of obtaining the customized persistent key-value storage component.

[0095] It should be further explained that, to improve the feasibility of directly controlling the hardware resources of each storage device, the implementation process of the above-mentioned custom target storage interface may include: determining an algorithm for controlling hardware resources; wherein the algorithm for controlling hardware resources is an algorithm that directly manipulates the resources of the underlying hardware device through software logic; and integrating the algorithm for controlling hardware resources into the interface to obtain the target storage interface. This embodiment determines the algorithm for controlling hardware resources. Its core goal is to achieve refined scheduling of hardware resources and efficient data distribution without relying on operating system abstraction layers (such as the file system), thereby enabling the custom persistent key-value storage component to interact directly with the underlying storage device. In this embodiment, the algorithm for controlling hardware resources directly manipulates the underlying storage device through software logic. This embodiment does not limit the specific algorithm for controlling hardware resources; the algorithm may involve asynchronous I / O (input and output), direct I / O, memory-mapped files, and other mechanisms. Compared to the current reliance on operating system file systems or block device drivers, the present invention eliminates the performance loss of intermediate layers and directly interacts with storage devices.

[0096] It should be further explained that the algorithms for controlling hardware resources include at least one of an asynchronous I / O algorithm, a direct I / O algorithm, and a data mapping algorithm. The asynchronous I / O algorithm serializes key-value pairs and then directly submits asynchronous requests to the storage device. The direct I / O algorithm bypasses the kernel cache and directly controls the storage device. The data mapping algorithm maps the physical address space of the storage device to the virtual address space of a custom persistent key-value storage component, allowing direct reading and writing of data through pointers. To facilitate understanding, the following is a pseudocode example showing how to implement asynchronous I / O-based operations based on the asynchronous I / O algorithm:

[0097] class AsyncWrite{

[0098] / / Assume there is a native I / O interface

[0099] NativeIOInterface nativeIO;

[0100] constructor(){

[0101] / / Initialize the native I / O interface

[0102] nativeIO = new NativeIOInterface();

[0103] }

[0104] / / Asynchronously write data to disk

[0105] function writeAsync(key, value) {

[0106] / / Prepare data

[0107] dataToWrite = serialize(key, value);

[0108] / / Generate a write request

[0109] writeRequest = new WriteRequest(dataToWrite);

[0110] / / Call the native storage interface for asynchronous writing

[0111] nativeIO.writeAsync(writeRequest,onWriteComplete);

[0112] }

[0113] / / Write completed callback function

[0114] function onWriteComplete(status){

[0115] if(status.isSuccess()){

[0116] log("Write operation completed successfully."); / / Check whether the write operation is successful

[0117] }else{

[0118] log("Error occurred during write operation:"+status.errorMessage); / / Record error information

[0119] }

[0120] }

[0121] }

[0122] / / Native I / O interface class

[0123] class NativeIOInterface{

[0124] function writeAsync(WriteRequest request,callback){

[0125] / / Use the operating system's asynchronous I / O function to write

[0126] / / This can be libaio or io_uring library

[0127] result=nativeAioWriteFunction(request.data);

[0128] / / Listen for write completion events

[0129] if(result.isComplete()){

[0130] callback(new Status(success=true));

[0131] }else{

[0132] callback(new Status(success=false,errorMessage="I / OError"));

[0133] }

[0134] }

[0135] }

[0136] / / Example main program

[0137] function main(){

[0138] asyncWrite=new AsyncWrite();

[0139] / / Asynchronously write a pair of key values

[0140] key="exampleKey";

[0141] value="exampleValue";

[0142] asyncWrite.writeAsync(key,value);

[0143] }

[0144] The AsyncWrite class encapsulates asynchronous write operations and initializes a native I / O interface for low-level I / O operations.

[0145] The writeAsync method receives the data to be written (key-value pairs), serializes it into a writable data format, creates a write request, and initiates asynchronous writing using the writeAsync method of the native I / O interface.

[0146] onWriteComplete method: As a callback function, it handles the status of writing completion, checks whether the writing is successful, and records the corresponding log.

[0147] NativeIOInterface class: Responsible for direct interaction with the underlying hardware. It uses a hypothetical nativeAioWriteFunction method to perform asynchronous write operations and determines the callback based on the operation results.

[0148] Main program: Create an instance of AsyncWrite and send a write request to it.

[0149] Based on the above approach, we describe an implementation strategy based on asynchronous I / O, which allows RocksDB to directly control hardware resources through native storage interfaces. In this way, RocksDB can optimize data write performance and reduce latency, thereby providing higher efficiency in storage-intensive applications.

[0150] It should be further explained that, based on the above embodiment, the above-mentioned customized persistent key-value storage component determines the storage location corresponding to the data to be stored, and converts the data format of the data to be stored into the target write data, which may include: using the asynchronous input and output algorithm in the customized persistent key-value storage component to determine the target protocol corresponding to the data to be stored; wherein the target protocol is the protocol corresponding to the target storage device; using the asynchronous input and output algorithm to serialize the data to be stored based on the target protocol to obtain the target write data; the serialization process includes data verification, data compression and metadata encapsulation. The purpose of data verification in this embodiment is to ensure that the data is not damaged during transmission and processing. The purpose of data compression in this embodiment is to reduce storage space occupancy and improve transmission efficiency. The purpose of metadata encapsulation in this embodiment is to encapsulate metadata content in a form suitable for storage on a storage device.

[0151] It should be further explained that the above-mentioned custom-based persistent key-value storage component stores the target write data in the storage device corresponding to the storage location, which may include: utilizing the direct read-write mode flag that bypasses the kernel cache in the direct input-output algorithm to pass the target write data directly to the storage device. The direct input-output algorithm in this embodiment enables the kernel's cache layer to be skipped when processing data, and the data is directly transferred between the custom persistent key-value storage component and the storage device. This embodiment does not limit the specific direct read-write mode flag. For example, when an application calls read() or write(), if the file is opened in O_DIRECT mode, the data will go directly to the storage device. This embodiment forces the kernel cache to be bypassed by enabling the direct read-write mode flag and passes the data directly to the storage device. This mode ensures the immediate persistence of data by enforcing alignment rules and transmission.

[0152] S102: Determine a storage location corresponding to the data to be stored based on the customized persistent key-value storage component, and convert the data format of the data to be stored into target write data.

[0153] The customized persistent key-value storage component in this embodiment determines the storage location corresponding to the data to be stored, and converts the data format of the data to be stored into the target write data. This embodiment does not limit the specific storage location. In this embodiment, the storage location corresponding to the data to be stored can be determined based on the data structure, that is, maintaining a mapping interface to track the relationship between the physical address of the storage location and the logical key; or a suitable storage location can be selected for the data to be stored, as long as the storage location corresponding to the data to be stored is determined by the customized persistent key-value storage component. The target write data in this embodiment is the data to be stored in the data format required by the storage device.

[0154] It should be further explained that, based on any of the above embodiments, in order to improve the accuracy of determining the storage location, the above-mentioned determination of the storage location corresponding to the to-be-stored data based on the customized persistent key-value storage component may include:

[0155] S1021, during a write operation, determining an optimal physical write location based on the semantics and access pattern of the data to be stored;

[0156] S1022, using the optimal write physical location as the storage location;

[0157] S1023: Map the logical key and the storage location using a custom persistent key-value storage component to obtain a mapping structure, and write data based on the mapping structure.

[0158] The semantics of the data to be stored in this embodiment refers to the business meaning, purpose and life cycle characteristics of the data. This embodiment does not limit the specific semantics. For example, the semantics may include transaction logs, hot data, cold data, streaming data, etc. The access mode in this embodiment describes the read and write behavior characteristics of the data, including frequency, sequentiality, concurrency, etc. This embodiment does not limit the specific access mode. For example, the access mode in this embodiment may be sequential, concurrency, etc. The semantic type and access mode in this embodiment can determine the optimal write physical location. The optimal write physical location may be the optimal write physical location with the best efficiency; or the optimal write physical location in this embodiment may be the optimal write physical location with the best performance. For ease of understanding, by integrating the address mapping into a custom persistent key-value storage component, data placement on the underlying storage device can be controlled. This integration allows the system to utilize information about data semantics, statistics and access patterns (such as the required I / O parallelism level) to achieve efficient data placement. The following is a code example provided by an embodiment of the present invention:

[0159] lass AddressMapper{

[0160] / / Maintain a mapping table to map logical keys to physical storage addresses

[0161] map<string,PhysicalAddress> addressMap;

[0162] / / Add mapping relationship

[0163] functionaddMapping(key,physicalAddress){

[0164] addressMap[key]=physicalAddress;

[0165] }

[0166] / / Find the physical address

[0167] function getPhysicalAddress(key){

[0168] return addressMap[key];

[0169] }

[0170] }

[0171] lass StorageManager{

[0172] AddressMapper addressMapper; / / Member variable: address mapper instance

[0173] constructor(){

[0174] addressMapper=new AddressMapper(); / / Initialize address mapper

[0175] }

[0176] / / Write data and control physical data placement

[0177] function writeData(key,value){

[0178] / / Select physical address based on access mode and data semantics

[0179] physicalAddress=determineOptimalPlacement(key, value);

[0180] / / Map logical keys to physical addresses

[0181] addressMapper.addMapping(key,physicalAddress);

[0182] / / Actually write the operation to the underlying storage

[0183] writeToStorage(physicalAddress,value);

[0184] }

[0185] / / Determine the physical address based on data semantics and access mode

[0186] function determineOptimalPlacement(key,value){

[0187] / / Get access mode information, data statistics, etc.

[0188] accessPattern=getAccessPattern(key);

[0189] semanticInfo=getSemanticInfo(key,value);

[0190] / / Use this information to select the optimal physical address

[0191] return optimalPhysicalAddress;

[0192] }

[0193] / / Actually write to the underlying storage (native storage interface)

[0194] function writeToStorage(physicalAddress,value){

[0195] / / Here you can call the underlying storage interface to perform I / O operations

[0196] nativeStorageInterface.writeTo(physicalAddress,value);

[0197] }

[0198] }

[0199] function main(){

[0200] storageManager=new StorageManager();

[0201] / / Write data

[0202] key="exampleKey"; / / Logical key

[0203] value="exampleValue"; / / data to be stored

[0204] storageManager.writeData(key,value);

[0205] }

[0206] The unique identifier of the data in this embodiment (such as an object ID or file path) is used for business-layer access. The physical address in this embodiment refers to the actual location of the data on the storage device. The access mode in this embodiment refers to the usage characteristics of the data (such as read-write ratio, sequentiality, and latency sensitivity). The data semantics in this embodiment refers to the business meaning of the data (such as transaction logs, hot data, and archived data). The present invention is designed to define a data structure, which requires maintaining a mapping structure to track the relationship between physical addresses and logical keys, maintaining a mapping table to map logical keys to physical storage addresses, adding mapping relationships, and finding physical address methods.

[0207] It should be further explained that, in order to improve the accuracy of determining the optimal write location, the above-mentioned determination of the optimal write physical location according to the semantics and access mode of the data to be stored during the write operation may include: obtaining the semantic type and access mode type of the data to be stored; wherein the semantic type includes at least one of the data type and the business scenario, and the access mode type is determined according to the read and write behavior characteristics of the data on the storage device; determining the physical location that maximizes the input and output efficiency based on the semantic type and the access mode type, and obtaining the optimal write physical location. This embodiment has the following core advantages by dynamically determining the optimal write physical location through the data type, business scenario and access mode type: (1) Maximizing performance and efficiency: selecting low-latency media (such as NVMe SSD) based on data semantics (such as frequently accessed hot data) to improve read and write throughput. Optimizing large block I / O for sequential writes (such as log streams) and optimizing 4K small I / O for random access (such as database indexes). Example: Transaction logs (high-frequency append writes) are preferentially written to the local SSD to reduce network latency. Reducing network overhead: Writing nearby: writing data to the nearest physical node based on the business scenario (such as edge computing) to reduce cross-region transmission. Bandwidth savings: Cold data (low-frequency access) is compressed and written to a remote storage pool, reducing transmission costs. (2) Improved storage resource utilization: Hot data: Use high-performance but high-cost SSDs, occupying only the necessary capacity. Cold data: Use high-density HDDs or erasure coded storage pools to reduce unit storage costs. (3) Business scenario scalability: Allocate independent storage pools to different tenants and configure policies on demand (e.g., SSD pools for e-commerce tenants and HDD pools for logging service tenants). The local storage pool processes real-time data, while the cloud storage pool processes archived data.

[0208] It should be further explained that in order to improve the accuracy of determining the access model type, before obtaining the semantic type and access mode type of the data to be stored, the following may be included: obtaining access frequency statistics, sequential statistics, and concurrency pressure statistics corresponding to the data to be stored; determining the access mode type based on the access frequency statistics, sequential statistics, and concurrency pressure statistics; wherein the access mode types include low-frequency access mode, high-frequency access mode, sequential access mode, random access mode, high-concurrency mode, and low-concurrency mode. This embodiment can achieve refined management and optimization for different scenarios by subdividing the access mode into low-frequency access, high-frequency access, sequential access, random access, high-concurrency, and low-concurrency modes. For example, data to be stored in a high-frequency access mode is preferentially cached in memory or SSD (Solid State Drive) to reduce access latency; data in a low-frequency access mode is migrated to low-performance media (such as HDD) to reduce storage costs.

[0209] S103: writing the target data into a storage device corresponding to the storage location based on the customized persistent key-value storage component.

[0210] The custom persistent key-value store component in this embodiment stores the target write data on a storage device corresponding to the storage location. This embodiment is not limited to a specific storage device. For example, the storage device in this embodiment may be an OSD (Object Storage Device), or an HDD (Hard Disk Drive).

[0211] It should be further explained that, based on any of the above embodiments, the data storage method may further include: loading the target data to be recovered from the target snapshot using a snapshot directory based on a customized persistent key-value storage component; obtaining a data mapping relationship in the customized persistent key-value storage component, and restoring the target data to be recovered to the corresponding storage device using the customized persistent key-value storage component based on the data mapping relationship. This embodiment can determine the target snapshot to be recovered based on user needs or system configuration. The target snapshot refers to the data state captured at a specific point in time and stored in the snapshot directory. The data to be recovered refers to the data content recorded in the target snapshot, which needs to be restored to the storage device. The present invention uses a customized persistent key-value storage component to manage data storage and recovery. This component is capable of storing key-value pairs of data and maintaining mapping relationships between data. The data mapping relationship is obtained from the customized persistent key-value storage component. The data mapping relationship refers to the location information of data in the storage device, including the data key, value, and corresponding storage location. The customized persistent key-value storage component is used to restore the target data to the corresponding storage device based on the obtained data mapping relationship. This process ensures that data can be accurately restored to its original location and that the restored data is consistent with the data state in the target snapshot. This invention enables rapid data loading and recovery, significantly improving data recovery speed. The custom persistent key-value storage component provides flexible data management capabilities that can adapt to a variety of complex data structures and storage requirements.

[0212] It should be further explained that, based on any of the above embodiments, the above-mentioned custom persistent key-value storage component utilizes a snapshot directory to load the target data to be recovered from the target snapshot, which may include: when a storage device fails, the custom persistent key-value storage component selects a snapshot directory at the most recent time point and loads the target data to be recovered from the target snapshot. This embodiment can detect the status of the storage device in real time through a monitoring mechanism. When a storage device failure is detected, the data recovery process is triggered, and the snapshot at the most recent time point is selected from the snapshot directory. The snapshot at the most recent time point refers to the most recently successfully created snapshot before the failure occurs, which can minimize data loss.

[0213] It should be further explained that, based on any of the above embodiments, the data storage method may further include: obtaining data merge conditions; and merging data using the data merge conditions based on a custom persistent key-value storage component. This embodiment does not limit the specific data merge conditions, as long as the data merge conditions can improve the efficiency of data merging. For example, the data merging method in this embodiment may be based on data access patterns. For hot data, the merge operation is appropriately delayed to reduce frequent access interference to hot data; for cold data, the merge is prioritized to free up storage space. Alternatively, the data merging method may be based on a tiered merge strategy. When the merged data is files, the files may be divided into multiple tiers, with the file size of each tier gradually increasing. During the merging process, merging is prioritized at lower tiers, and the merged files are merged upwards in tiers. For example, files may be divided into tiers such as L0, L1, and L2. The files at tier L0 are smaller and merge more frequently; the files at tiers L1 and L2 gradually increase in size and merge less frequently. This tiered merge strategy can effectively reduce the number of merges for large files and improve merge efficiency. The number and size of tiers can be dynamically adjusted based on the data write speed and query pattern. When the write speed is high, increase the size of the L0 layer to reduce the frequency of merge operations; when the query speed requirement is high, increase the size of the L1, L2 and other layers to reduce the number of files that need to be scanned during the query.

[0214] It should be further explained that, based on any of the above embodiments, the above-mentioned data merge conditions include at least one of a data merge condition based on data access patterns, a data merge condition based on data update frequency, and a data merge condition based on data size, wherein the data merge condition based on data access patterns is triggered based on an access frequency threshold. It is understood that the data merge condition based on data access patterns can identify hot data and cold data by analyzing the data access patterns. For hot data, the merge operation is appropriately delayed to reduce interference with frequent access to hot data; for cold data, the merge operation is prioritized to free up storage space. For example, an access frequency threshold can be set, and the merge operation is triggered when the data access frequency falls below the threshold. For the data merge condition based on data update frequency, the merge operation is delayed for frequently updated data to reduce duplicate merges caused by frequent updates. Conversely, for data with a lower update frequency, the merge is performed promptly to maintain data cleanliness. This strategy can be implemented by counting the number of updates and timestamps for each key. For the data merge condition based on data size, the merge operation is triggered when the file size exceeds a certain threshold. At the same time, considering the size distribution of files, small files are merged first to reduce the number of times large files are merged, thereby reducing the I / O (input / output) overhead during the merging process.

[0215] It should be further explained that, based on any of the above embodiments, obtaining data merging conditions and performing data merging based on a customized persistent key-value storage component using the data merging conditions may include: when the data merging condition is a data merging condition based on a data access pattern, determining whether the data access frequency of the data to be merged is lower than a preset data access frequency threshold; when it is lower than the data access frequency threshold, merging the data to be merged; when the data merging condition is data based on a data update frequency, determining whether the update frequency of the data to be merged is lower than a preset minimum update frequency; when it is lower than the minimum update frequency, merging the data to be merged; when the data merging condition is a data merging condition based on data size, determining whether the size of the data to be merged is lower than a preset minimum number of bytes; when the size of the data to be merged is lower than the minimum number of bytes, merging the data to be merged. This embodiment provides the timing for data merging based on different data merging conditions, thereby improving the accuracy of data merging.

[0216] A data storage method provided by an embodiment of the present invention may include: S101, obtaining a custom persistent key-value storage component; wherein the custom persistent key-value storage component is a component including an interface for directly interacting with an underlying storage device; S102, determining a storage location corresponding to the data to be stored based on the custom persistent key-value storage component, and converting the data format of the data to be stored into target write data; S103, storing the target write data into a storage device corresponding to the storage location based on the custom persistent key-value storage component. Compared with the current method of converting the data format into a storage form suitable for the storage device and serializing the data structure into a byte stream suitable for the storage device, which requires creating a storage layer abstraction and interacting with the storage device, the present invention improves the performance of data storage by designing a custom persistent key-value storage component that can interact directly with the storage device, making it possible to directly control the underlying storage device without conversion.

[0217] In order to make the present invention easier to understand, please refer to Figure 2 , Figure 2 A flowchart of another data storage method provided in an embodiment of the present invention, which is applied to a custom persistent key-value storage component, may specifically include:

[0218] S201, obtaining partition information.

[0219] The partition information of this embodiment refers to the partition configuration information, including the attributes of each partition (such as storage device identification, access mode, hot and cold status, garbage collection threshold, etc.). The partition information in this embodiment includes at least one of the storage device identification, access mode, and whether it is hot data. The storage device identification in this embodiment represents the storage device bound to the partition, and maps the logical partition (such as "hot data partition") to a specific physical device or storage pool to facilitate the unified allocation of storage resources. The access mode in this embodiment describes the read and write characteristics of data on the storage device, and may include: sequential access, random access, and append access. The access mode can be used to determine the optimal storage device partition corresponding to each access mode. Whether it is hot data in this embodiment can determine whether the target write data is accessed frequently, and hot data can be allocated to high-performance storage devices.

[0220] S202 , dividing the storage device into corresponding partitions using a partition strategy based on the partition information; wherein the partition strategy is a strategy that enables data to be distributed according to a set rule.

[0221] The core purpose of partitioning in this embodiment is to logically divide physical storage devices to optimize storage performance, resource utilization, and operational flexibility. In this embodiment, each partition is a logical abstraction of one or more storage devices. Storage devices are assigned to different partitions based on a partitioning strategy, which focuses on optimizing data distribution patterns and access patterns. This embodiment does not limit specific partitioning strategies. For example, the partitioning strategy in this embodiment may include a hot-cold data separation strategy, where hot data (highly accessed) is assigned to high-performance devices (such as SSDs) and cold data (lowly accessed) is assigned to low-cost devices (such as HDDs). Alternatively, the partitioning strategy in this embodiment may include an access pattern strategy, where sequential access data is assigned to HDDs (optimizing large-block I / O) and random access data is assigned to SSDs (optimizing small IOPS). The partitioning strategy in this embodiment may also include a load balancing strategy, which dynamically adjusts partition assignment based on device load (e.g., a heavily loaded OSD migrates some data to an idle OSD). It should be further noted that the aforementioned partitioning strategies include at least one of a hot-cold data separation strategy, a dynamic data distribution strategy, and a storage device wear leveling strategy. Hot and cold data separation strategy: Based on data access frequency (hot and cold), data is stored in storage devices or partitions of different performance tiers to optimize access performance and reduce costs. The dynamic data distribution strategy in this embodiment: Dynamically adjusts data distribution based on real-time load (such as I / O pressure and storage utilization) to achieve load balancing and efficient resource utilization. The storage device wear leveling strategy in this embodiment: For flash memory devices (such as SSDs), the erase and write cycles of each storage block are balanced to extend device life and reduce performance degradation.

[0222] It should be further explained that, in order to improve the accuracy of partitioning, after partitioning the storage devices into corresponding partitions using a partitioning strategy based on the partitioning information, the following steps may also be performed: detecting workload information; and dynamically updating the partitions based on the workload information. This embodiment can dynamically update the partitions based on the workload information of the storage devices in each partition to achieve load balancing. The partitions in this embodiment can be dynamically updated, thereby improving the accuracy of the partitioning. It should be noted that the attributes of the partitions (such as hot and cold status) can change with the workload, but the binding relationship of the physical devices is dynamic.

[0223] S203 , determining a storage location corresponding to the data to be stored, and converting the data format of the data to be stored into target write data.

[0224] In this embodiment, the data to be stored is converted into target write data that can be recognized by the storage device.

[0225] S204: Determine the target partition corresponding to the target write data.

[0226] The target partition in this embodiment may include multiple storage devices. The partition in this embodiment is responsible for allocating data to the optimal physical storage location according to the strategy.

[0227] S205: Write the target write data into the storage device corresponding to the storage location in the target partition.

[0228] Through the above steps, the system of this embodiment realizes intelligent mapping from logical data to physical storage.

[0229] It should be further explained that, in order to optimize the partition, after the target data is written to the storage device corresponding to the storage location based on the custom persistent key-value storage component, the following steps may also be included:

[0230] S1: Obtain the status, load, and wear level of each partition according to a set time period.

[0231] The status information in this embodiment includes the health status of the partition (e.g., online / offline) and current capacity utilization (used space / free space). The load in this embodiment may include the number of read / write operations per second (IOPS), throughput (MB / s), request queue depth, and average latency. The wear level in this embodiment may include the number of program / erase (PE) cycles, remaining lifespan percentage, number of bad sectors, number of motor starts, and operating time.

[0232] S2: Based on the status, load, and wear level of each partition, a machine learning model is used to determine the garbage collection parameters corresponding to each partition.

[0233] The machine learning model training process in this embodiment may include inputting the status, load, and wear level of the partition, outputting training data of the corresponding garbage collection parameters, and using the training data to train the model.

[0234] S3: Compare the garbage collection parameters with the garbage collection threshold and clean up the partition data.

[0235] The partition's garbage collection threshold is a critical condition used to control when garbage collection operations are triggered. Its core purpose is to achieve a balance between storage space utilization, system performance, and resource overhead. In this embodiment, if the current partition's garbage collection parameter is greater than the garbage collection threshold, the partition's data can be cleaned up. If the current partition's garbage collection parameter is not greater than the garbage collection threshold, the partition's data is not cleaned up. This embodiment uses garbage collection logic based on garbage collection parameters and garbage collection thresholds to promptly clean up partitioned garbage data, improving cleaning efficiency.

[0236] To facilitate understanding, this embodiment provides a detailed explanation of the partitioning process described above. The partitioning method serves as an abstraction of physical storage across multiple OSDs (i.e., parallel units of storage devices). Partitions effectively optimize for different access patterns (sequential, random, and append) within various key-value store components, such as different levels of the LSM tree and the log manager. Partitions support flexible physical storage management, with their definition including parameters such as hot and cold data separation and garbage collection. Furthermore, partition definitions are not static but can evolve over time to reflect workload characteristics.

[0237] 1. First, define a Partition class to manage physical storage partition information, including hot and cold data separation and garbage collection parameters.

[0238] class Partition{

[0239] string id;

[0240] string storageDevice; / / storage device identifier

[0241] AccessPattern accessPattern; / / Access mode (sequential, random, append)

[0242] bool isHot; / / Is it hot data?

[0243] int gcThreshold; / / Garbage collection threshold

[0244] constructor(id,storageDevice){

[0245] this.id=id;

[0246] this.storageDevice=storageDevice;

[0247] this.accessPattern=AccessPattern.SEQUENTIAL; / / Default access mode

[0248] this.isHot=false; / / Default cold data

[0249] this.gcThreshold=100; / / Default garbage collection threshold

[0250] }

[0251] / / Update access mode

[0252] function updateAccessPattern(newPattern){

[0253] this.accessPattern=newPattern;

[0254] }

[0255] / / Update hot and cold data status

[0256] function updateHotColdStatus(isHot){

[0257] this.isHot=isHot;

[0258] }

[0259] / / Handle garbage collection

[0260] function performGarbageCollection(){

[0261] if(this.isHot&&dataCount>gcThreshold){

[0262] / / Execute garbage collection logic

[0263] }

[0264] }

[0265] }

[0266] 2. Manage partitions

[0267] class StorageManager{

[0268] List <partition>partitions;

[0269] constructor(){

[0270] partitions=new List <partition>();

[0271] / / You can initialize multiple partitions here

[0272] }

[0273] / / Add a new partition

[0274] function addPartition(partition){

[0275] partitions.add(partition);

[0276] }

[0277] / / Write data

[0278] function writeData(key,value){

[0279] / / Determine the target partition (can be based on key value)

[0280] targetPartition=determineTargetPartition(key);

[0281] / / Check partition type and update status

[0282] targetPartition.updateAccessPattern(determineAccessPattern(key));

[0283] targetPartition.performGarbageCollection();

[0284] / / Write data to the target physical storage

[0285] writeToStorage(targetPartition.storageDevice,key,value);

[0286] }

[0287] / / Read data

[0288] function readData(key){

[0289] targetPartition=determineTargetPartition(key);

[0290] return readFromStorage(targetPartition.storageDevice,key);

[0291] }

[0292] / / Dynamically adjust partitions (detect workload changes)

[0293] function dynamicPartitioning(){

[0294] / / Collect statistics and update partition settings

[0295] for each(partition in partitions){

[0296] if(detectWorkloadChanges(partition)){

[0297] partition.updateHotColdStatus(evaluateHotColdData(partition));

[0298] / / Other dynamic adjustment logic

[0299] }

[0300] }

[0301] }

[0302] / / Determine the target partition, which can be customized according to the algorithm

[0303] function determineTargetPartition(key){

[0304] / / Implement a custom algorithm to determine the target partition

[0305] return partitions[0]; / / Example

[0306] }

[0307] }

[0308] 3. Use the above definitions to create partitions and perform data operations.

[0309] function main(){

[0310] storageManager=newStorageManager();

[0311] / / Create partition and add it to storage manager

[0312] partition1=new Partition("partition1","OSD1");

[0313] partition2=new Partition("partition2","OSD2");

[0314] storageManager.addPartition(partition1);

[0315] storageManager.addPartition(partition2);

[0316] / / Write data

[0317] key="exampleKey";

[0318] value="exampleValue";

[0319] storageManager.writeData(key,value);

[0320] / / Dynamically adjust partitions

[0321] storageManager.dynamicPartitioning();

[0322] }

[0323] In the embodiments of the present invention, partitions are logically abstracted units of storage devices. Their core purpose is to optimize data storage and physical resource management through dynamic policies (such as access patterns, hot-cold separation, and garbage collection). Each partition serves as both a logical mapping of the physical device and a container for storage policies, enabling dynamic adjustments based on workload characteristics to achieve high performance, high resource utilization, and flexible operations and maintenance management. The present invention proposes this flexible partition management solution, which can rationally optimize for different access patterns and support the separation and dynamic adjustment of hot and cold data. This approach effectively improves storage efficiency and performance, providing powerful support for key-value storage systems in handling diverse workloads.

[0324] In order to make the present invention easier to understand, please refer to Figure 3 , Figure 3 A flowchart of a data storage method provided in an embodiment of the present invention is applied to a custom persistent key-value storage component, which may specifically include:

[0325] S301 , determining a storage location corresponding to the data to be stored, and converting the data format of the data to be stored into target write data.

[0326] S302: Determine the optimal physical write location based on the semantics and access pattern of the data to be stored.

[0327] S303: Using the optimal write physical location as the storage location, and using local input and output to store the target write data in a storage device corresponding to the storage location.

[0328] It's important to note that to enable automatic data cleanup within a custom persistent key-value store, the following modification solution was designed, combining its built-in mechanisms with distributed storage requirements: This automatically filters data to be cleaned based on timestamps, version numbers, or business tags (such as soft-delete tags). Cleanup is triggered based on the invalid data ratio, storage pressure, or a scheduled policy.

[0329] The embodiment of the present invention can directly control the underlying physical storage disk for optimization, and place the address mapping and data storage format of the physical disk in a customized persistent key-value storage component for processing, which can solve the problems of storage characteristics, device parallelism and wear leveling.

[0330] Because existing technologies use RocksDB as the storage engine for OSDs, multiple storage abstraction layers exist. Different data storage access patterns often lead to suboptimal physical I / O patterns. For example, frequent random reads and writes can increase disk seek times. Many persistent storage systems fail to apply intelligent data placement strategies based on current workloads, resulting in uneven data distribution and unnecessary I / O operations. In existing I / O paths, some functions are redundant, leading to additional computing and storage overhead, further exacerbating write amplification and performance issues.

[0331] This paper proposes a distributed storage-based key-value storage optimization method and system to address the physical storage management limitations of the traditional distributed storage system, the persistent RocksDB key-value store (KV store). These limitations arise from the compatibility of different architectural layers (e.g., RocksDB interfaces get and put with Linux server storage device interfaces write and read), the interface abstraction of hardware HDDs and NVMe SSDs, and I / O interface access patterns. Because existing technologies employ multiple storage abstraction layers, varying data storage access patterns often result in suboptimal physical I / O patterns. For example, frequent random reads and writes can increase disk seek times. Many persistent storage systems fail to apply intelligent data placement strategies based on current workloads, resulting in uneven data distribution and unnecessary I / O operations. In existing I / O paths, some functions are redundant, resulting in additional computational and storage overhead, further exacerbating write amplification and performance issues.

[0332] The optimization module of the present invention is the RocksDB module in the distributed storage architecture. OSD (Object storage daemon) is the component responsible for storing actual data, which manages the storage and retrieval of data objects. RocksDB is an embeddable persistent key-value storage KV-store for fast data storage. RocksDB provides OSD with high-performance key-value storage capabilities, enabling OSD to efficiently handle read and write operations and ensure data reliability and performance. Using RocksDB as the storage engine for OSD and managing metadata enables distributed storage to provide high-performance data storage services in different application scenarios. For example, in object storage, RocksDB, as the underlying storage engine, supports fast data reading and writing and efficient index management, which is crucial for processing a large number of data read and write requests. In addition, RocksDB also supports data persistence and rapid recovery, which is crucial for ensuring data security and system availability.

[0333] This invention optimizes the underlying physical storage disks by directly controlling them. This process, which handles physical disk address mapping, data storage, and garbage collection within the distributed storage system RocksDB, addresses storage characteristics, device parallelism, and wear leveling. By directly interacting with the underlying storage devices, the invention can better manage the physical layout of data and I / O operations. This direct control approach enables distributed storage systems to fully exploit the characteristics of HDDs or NVMe SSDs, such as parallel read and write, low latency, and durability.

[0334] In order to make the present invention easier to understand, please refer to Figure 4 , Figure 4 A structural framework diagram of a distributed storage system provided in an embodiment of the present invention may specifically include:

[0335] The design concept of the distributed storage system is based on scalability and fault tolerance, eliminating performance bottlenecks such as single-point metadata. It adopts a distributed, decentralized architecture, slicing data into objects and distributing them across multiple nodes. This ensures balanced load and capacity across each node, enabling data access and management via network connections while providing a unified storage interface. Due to its decentralized architecture, the distributed storage system is highly scalable; a single storage server node failure does not render the entire distributed storage cluster unavailable. It supports dynamic expansion, adding or removing storage server nodes, and automatically relocates data objects within the cluster to balance storage space utilization across storage servers. The distributed storage system architecture includes multiple components, including an object gateway service, a block device service, a file system service, a unified, self-controlled, and scalable distributed storage consistency management system, a storage pool, a controllable, scalable, distributed data balancing algorithm, a metadata cluster and a monitoring service cluster, and the optimized storage backend engine (RocksDB) of the present invention. Distributed storage clients associated with the distributed storage system include objects, blocks, virtual machines, containers, and files.

[0336] The distributed storage system architecture includes the following key components:

[0337] 1) Object gateway service, block device service, and file system service: Provide a unified storage interface to provide access interface services for distributed storage clients.

[0338] 2) A unified, self-controlled, and scalable distributed storage consistency management system: This component is the core component of the distributed storage cluster and provides distributed object-based storage services. In this component, data is stored as objects, each with a unique identifier and associated data. This component is responsible for distributing objects across the nodes of the storage cluster and provides data replication, recovery, and load balancing to ensure data reliability and high-performance access.

[0339] 3) Storage Pool: A distributed storage cluster consists of multiple storage nodes, each of which can contain multiple hard drives or storage devices. Storage pools divide storage resources into different pools to meet different storage needs.

[0340] 4) Controllable, scalable, distributed data balancing placement algorithm: This component evenly distributes data objects across the nodes of the storage cluster to avoid data hotspots and improve system performance.

[0341] 5) Metadata Cluster and Monitoring Service Cluster: The metadata cluster manages file system metadata, including file and directory attributes and locations. The monitoring service cluster monitors cluster status and configuration. Its technical principles include maintaining cluster status, configuration, and health information to ensure cluster consistency and availability.

[0342] 6) In the storage backend engine of the distributed storage system, RocksDB serves as the underlying storage engine to handle I / O metadata operations. RocksDB is primarily used to implement efficient key-value storage solutions and incorporates a variety of optimization techniques. RocksDB is a data storage engine based on the LSM (Log-Structured Merge) tree. It achieves efficient sequential writes and fast random reads by caching write operations in memory (MemTable) and periodically merging and writing data to the SSD or HDD (SSTable files). Data transmission between different storage media (such as SSD or HDD) is ensured through interaction between the interface that interacts with the underlying storage device (such as the Object Store API) and RocksDB. The present invention optimizes the conversion interface logic and implements it in RocksDB to avoid frequent interface conversions and redundant compatibility operations.

[0343] The core of this invention is to seamlessly integrate the data storage interfaces (interfaces that interact with the underlying storage devices) of disks and NVMe SSDs into the distributed storage system RocksDB key-value store module to improve performance and efficiency. The following details the key technical methods of this invention and how it optimizes I / O operations and resource management through native storage interfaces and physical storage abstraction.

[0344] The present invention proposes a design method for directly controlling hardware resources in RocksDB using native storage interfaces: directly interacting with the underlying storage hardware, first customizing RocksDB's storage layer, and adjusting the execution mode of I / O operations according to specific workloads and hardware characteristics. This provides flexible support for different storage media (such as NVMe (Non-Volatile Memory Host Controller Interface Specification), SATA SSD (Serial ATA Solid-State Drive), traditional HDD, etc.), and using memory-mapped I / O (mmap), RocksDB can access large-scale data sets more efficiently, reduce context switching between user space and kernel space, and improve I / O performance. For flash-based storage (such as NVMe SSD), the present invention optimizes RocksDB's write method to reduce write amplification and uses a wear leveling mechanism to extend the life of the storage medium. (1) The present invention uses a design algorithm for directly controlling hardware resources using native storage interfaces in a customized persistent key-value storage component of the backend storage engine, involving mechanisms such as asynchronous I / O, direct I / O, and memory-mapped files. Therefore, the present invention optimizes by directly controlling the underlying physical storage disk, and places the address mapping, data storage form, and garbage data collection of the physical disk in the distributed storage system RocksDB for processing, which can solve the problems of storage characteristics, device parallelism, and loss balancing. By directly interacting with the underlying storage device, the management of the data storage interface (interface for interacting with the underlying storage device) of the disk and NVmeSSD is seamlessly placed in the distributed storage system RocksDB key-value storage module to improve performance and efficiency. A design algorithm for directly controlling hardware resources using native storage interfaces in the backend storage engine RocksDB is proposed, involving mechanisms such as asynchronous I / O, direct I / O, and memory-mapped files. (2) By integrating address mapping into a custom persistent key-value storage component, the key-value storage system is able to control the physical data placement on the underlying storage device. This integration allows the system to utilize information about data semantics, statistics, and access patterns (such as the required I / O parallelism level) to achieve efficient data placement. Therefore, the present invention proposes a method for enabling the key-value storage system to control the physical data placement on the underlying storage device by integrating address mapping into the storage manager of the RocksDB key-value storage. (3) This paper proposes a partitioning approach as a physical storage abstraction across multiple OSDs (i.e., parallel units of storage devices). Partitions can be effectively optimized for different access patterns (sequential, random, append) in a custom persistent key-value store component. Partitions support flexible physical storage management, with their definition including parameters such as hot and cold data separation and garbage collection. Furthermore, the definition of partitions is not static but can evolve over time to reflect the characteristics of the workload.

[0345] Beneficial effects of the present invention:

[0346] 1. Performance. Given equivalent hardware resources, this invention can better manage data physical layout and I / O operations in highly concurrent, multi-cloud distributed object storage environments. This direct control approach enables distributed storage systems to fully leverage the characteristics of HDDs or NVMe SSDs, such as parallel read and write, low latency, and durability. By improving workload adaptability and eliminating redundancy, this invention aims to provide a solution for modern storage needs.

[0347] 2. Stability: The RocksDB key-value storage optimization method of the present invention is decoupled from the distributed storage system, transparent to business clients, and stable.

[0348] 3. Security: The components and method modules of the present invention are packaged and safe.

[0349] 4. Low cost: The RocksDB key-value storage optimization method proposed in this invention can improve the competitiveness and maintenance cost of distributed file storage.

[0350] 5. Compatibility: The RocksDB key-value storage optimization method according to the embodiment of the present invention is portable and universal across different hardware devices.

[0351] The data storage device provided by an embodiment of the present invention is introduced below. The data storage device described below and the data storage method described above can be referenced to each other.

[0352] Figure 5 A schematic diagram of a structural framework of a data storage device provided in an embodiment of the present invention may include:

[0353] The component acquisition module 100 is used to acquire a custom persistent key-value storage component; wherein the custom persistent key-value storage component is a component including an interface for directly interacting with an underlying storage device;

[0354] The target write data determination module 200 is configured to determine a storage location corresponding to the data to be stored based on the user-defined persistent key-value storage component, and convert the data format of the data to be stored into the target write data;

[0355] The storage module 300 is configured to store the target write data in a storage device corresponding to the storage location based on the customized persistent key-value storage component.

[0356] Furthermore, based on the above embodiment, the data storage device may further include:

[0357] An interface customization module, used to customize a target storage interface; wherein the target storage interface is used to directly control the hardware resources of each storage device;

[0358] The integration module is used to integrate the target storage interface with the persistent key-value storage component to obtain the customized persistent key-value storage component.

[0359] Furthermore, based on the above embodiment, the data storage device may further include:

[0360] An algorithm determination module for controlling hardware resources, used to determine an algorithm for controlling hardware resources; wherein the algorithm for controlling hardware resources is an algorithm for directly manipulating the resources of the underlying hardware device through software logic;

[0361] An algorithm integration module is used to integrate the algorithm for controlling hardware resources into the interface to obtain the target storage interface.

[0362] Further, based on any of the above embodiments, the algorithm for controlling hardware resources includes at least one of an asynchronous input and output algorithm, a direct input and output algorithm, and a data mapping algorithm;

[0363] The asynchronous input and output algorithm is an algorithm that serializes key-value pairs and directly submits asynchronous requests to the storage device;

[0364] The direct input and output algorithm is an algorithm that bypasses the kernel cache and directly controls the storage device;

[0365] The data mapping algorithm is an algorithm for mapping the physical address space of the storage device to the virtual address space of the customized persistent key-value storage component and directly reading and writing data through a pointer.

[0366] Further, based on any of the above embodiments, the target write data determination module 200 may include:

[0367] a target protocol determination module, configured to determine a target protocol corresponding to the data to be stored by utilizing the asynchronous input / output algorithm in the customized persistent key-value storage component; wherein the target protocol is a protocol corresponding to a target storage device;

[0368] A serialization processing module is used to use the asynchronous input and output algorithm to serialize the data to be stored based on the target protocol to obtain the target write data; the serialization processing includes data verification, data compression and metadata encapsulation.

[0369] Further, based on any of the above embodiments, the storage module 300 may include:

[0370] The storage unit is used to directly transfer the target write data to the storage device by utilizing the direct read / write mode flag for bypassing the kernel cache in the direct input / output algorithm.

[0371] Further, based on any of the above embodiments, the target write data determination module 200 may include:

[0372] An optimal write physical location determination unit, configured to determine an optimal write physical location according to the semantics and access mode of the data to be stored during a write operation;

[0373] a storage location determining unit, configured to use the optimal writing physical location as the storage location;

[0374] A mapping unit is used to map the logical key and the storage location using the customized persistent key-value storage component to obtain a mapping structure, so as to write data based on the mapping structure.

[0375] Further, based on any of the above embodiments, the above optimal writing physical position determination unit may include:

[0376] a semantic type and access mode type determination subunit, configured to obtain the semantic type and access mode type of the data to be stored; wherein the semantic type includes at least one of a data type and a business scenario, and the access mode type is determined based on the read and write behavior characteristics of the data on the storage device;

[0377] The optimal write physical location determination subunit is used to determine the physical location that maximizes input and output efficiency based on the semantic type and the access mode type, so as to obtain the optimal write physical location.

[0378] Furthermore, based on any of the above embodiments, the data storage device may further include:

[0379] A statistical information acquisition module, configured to acquire access frequency statistical information, sequential statistical information, and concurrent pressure statistical information corresponding to the data to be stored;

[0380] An access pattern type determination module is used to determine the access pattern type based on the access frequency statistical information, sequential statistical information and concurrency pressure statistical information; wherein the access pattern type includes low-frequency access mode, high-frequency access mode, sequential access mode, random access mode, high concurrency mode and low concurrency mode.

[0381] Furthermore, based on any of the above embodiments, the data storage device may further include:

[0382] Partition information acquisition module, used to obtain partition information;

[0383] A partition division module, configured to divide the storage device into corresponding partitions using a partition strategy based on the partition information; wherein the partition strategy is a strategy for distributing data according to a set rule;

[0384] Accordingly, the storage module 300 may include:

[0385] Determining a target partition corresponding to the target write data based on the customized persistent key-value storage component;

[0386] The target partition-based writing storage device unit is used to write the target write data into the storage device corresponding to the storage location in the target partition.

[0387] Further, based on any of the above embodiments, the partition information includes at least one of a storage device identifier, an access mode, and whether the data is hot data.

[0388] Furthermore, based on any of the above embodiments, the data storage device may further include:

[0389] A workload information determination module, configured to detect workload information;

[0390] The dynamic partition adjustment module is used to dynamically update the divided partitions based on the workload information.

[0391] Furthermore, based on any of the above embodiments, the data storage device may further include:

[0392] An information acquisition module is used to obtain the status, load, and wear level of each partition according to a set time period;

[0393] A garbage collection parameter determination module is used to determine the garbage collection parameters corresponding to each partition based on the status, load and wear level of each partition using a machine learning model;

[0394] The cleaning module is used to compare the garbage collection parameters with the garbage collection threshold and clean up the data of the partition.

[0395] Furthermore, based on any of the above embodiments, the partitioning strategy includes at least one of a cold and hot data separation strategy, a dynamic data distribution strategy, and a storage device wear leveling strategy.

[0396] Furthermore, based on any of the above embodiments, the storage device includes at least one of a non-volatile memory high-speed protocol hard disk, a serial transmission solid-state hard disk, a mechanical hard disk, and an object storage device.

[0397] Furthermore, based on any of the above embodiments, the data storage device may further include:

[0398] a target data determination module to be restored, configured to load the target data to be restored in the target snapshot using the snapshot directory based on the customized persistent key-value storage component;

[0399] The data recovery module is used to obtain the data mapping relationship in the customized persistent key-value storage component, and use the customized persistent key-value storage component to restore the target data to be recovered to the corresponding storage device based on the data mapping relationship.

[0400] Further, based on any of the above embodiments, the module for determining target data to be restored may include:

[0401] The target data to be restored determining unit is used for selecting a snapshot directory at the latest time point based on the customized persistent key-value storage component and loading the target data to be restored in the target snapshot when a storage device fails.

[0402] Furthermore, based on any of the above embodiments, the data storage device may further include:

[0403] A data merging condition determination module is used to obtain data merging conditions;

[0404] A data merging module is used to merge data based on the customized persistent key-value storage component using the data merging condition.

[0405] Further, based on any of the above embodiments, the data merging condition includes at least one of a data merging condition based on a data access pattern, a data merging condition based on a data update frequency, and a data merging condition based on a data size, wherein the data merging condition based on a data access pattern is triggered based on an access frequency threshold.

[0406] Furthermore, based on any of the above embodiments, the data merging module may include:

[0407] a first merging unit, configured to, when the data merging condition is a data merging condition based on a data access pattern, determine whether a data access frequency of the data to be merged is lower than a preset data access frequency threshold; and when the data access frequency is lower than the data access frequency threshold, merge the data to be merged;

[0408] A second merging unit is configured to, when the data merging condition is data based on data update frequency, determine whether the update frequency of the data to be merged is lower than a preset minimum update frequency; and when the update frequency is lower than the minimum update frequency, merge the data to be merged;

[0409] The third merging unit is used to determine whether the size of the data to be merged is lower than a preset minimum number of bytes when the data merging condition is a data merging condition based on data size; when the size of the data to be merged is lower than the minimum number of bytes, merge the data to be merged.

[0410] It should be noted that the order of the modules and units in the above data storage device can be changed without affecting the logic.

[0411] Figure 5 The description of the features in the corresponding embodiment can be found in Figure 5 The relevant descriptions of the corresponding embodiments will not be repeated here one by one.

[0412] A data storage device provided by an embodiment of the present invention may include: a component acquisition module 100 for acquiring a custom persistent key-value storage component; wherein the custom persistent key-value storage component is a component including an interface for directly interacting with an underlying storage device; a target write data determination module 200 for determining a storage location corresponding to the data to be stored based on the custom persistent key-value storage component, and converting the data format of the data to be stored into target write data; and a storage module 300 for storing the target write data to the storage device corresponding to the storage location based on the custom persistent key-value storage component. Compared with the current need to convert the data format into a storage form suitable for the storage device, serialize the data structure into a byte stream suitable for the storage device, and create a storage layer abstraction to interact with the storage device, the present invention improves the performance of data storage by designing a custom persistent key-value storage component that can interact directly with the storage device, so that the underlying storage device can be directly controlled without conversion.

[0413] A data storage system provided by an embodiment of the present invention is introduced below. The data storage system described below and the data storage method described above can be referenced to each other.

[0414] Figure 6 A schematic diagram of the structural framework of a data storage system provided by an embodiment of the present invention is shown in FIG. Figure 6 As shown, this may include:

[0415] The customized persistent key-value storage component 400 is used to determine a storage location corresponding to the data to be stored, convert the data format of the data to be stored into target write data, and store the target write data in a storage device corresponding to the storage location;

[0416] The storage device 500 is configured to store the target write data.

[0417] The customized persistent key-value storage component in the data storage system in the embodiment of the present invention is the customized persistent key-value storage component in the data storage method.

[0418] A data storage device provided by an embodiment of the present invention is introduced below. The data storage device described below and the data storage method described above can refer to each other.

[0419] Figure 7 A schematic diagram of a structural framework of a data storage device provided by an embodiment of the present invention is shown in FIG. Figure 7 As shown, the data storage device includes: a memory 60 for storing computer programs;

[0420] The processor 61 is configured to implement the steps of the data storage method in the above embodiment when executing a computer program.

[0421] The data storage device provided in this embodiment may include but is not limited to a smart phone, a tablet computer, a laptop computer, or a desktop computer.

[0422] The processor 61 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 61 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 61 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 61 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing content required to be displayed on the display screen. In some embodiments, the processor 61 may also include an artificial intelligence (AI) processor for handling computational operations related to machine learning.

[0423] The memory 60 may include one or more computer-readable storage media, which may be non-transitory. The memory 60 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 60 is at least used to store the following computer program 601, wherein, after the computer program is loaded and executed by the processor 61, it can implement the relevant steps of the data storage method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 60 may also include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 602 may include Windows, Unix, Linux, etc. The data 603 may include but is not limited to data required for the data storage method, etc.

[0424] In some embodiments, the data storage device may further include a display screen 62 , an input / output interface 63 , a communication interface 64 , a power supply 65 , and a communication bus 66 .

[0425] Those skilled in the art will understand that Figure 7 The structure shown in the figure does not constitute a limitation of the data storage device, and may include more or fewer components than shown in the figure.

[0426] It is understood that if the data storage method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the current technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and performs all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable ROM, a register, a hard drive, a removable disk, a CD-ROM, a magnetic disk, or an optical disk, and other media that can store program code.

[0427] Based on this, an embodiment of the present invention further provides a medium (computer-readable storage medium), on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned data storage method are implemented.

[0428] The above describes in detail a data storage method provided by an embodiment of the present invention. The various embodiments are described in a progressive manner throughout this specification, with each embodiment focusing on its differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. The device disclosed in the embodiments corresponds to the method disclosed in the embodiments, so the description is relatively brief. For relevant details, refer to the method description.

[0429] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0430] The above describes in detail a data storage method, device, system, equipment, medium and program product provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.< / partition> < / partition>

Claims

1. A data storage method, characterized in that: include: Determining an algorithm for controlling hardware resources; wherein the algorithm for controlling hardware resources is an algorithm for directly manipulating resources of underlying hardware devices through software logic; integrating the algorithm for controlling hardware resources into an interface to obtain a target storage interface; wherein the target storage interface is used to directly control the hardware resources of each storage device; Integrating the target storage interface with the persistent key-value storage component to obtain a customized persistent key-value storage component; Obtaining the custom persistent key-value storage component; wherein the custom persistent key-value storage component is a component including an interface for directly interacting with an underlying storage device; Determine a storage location corresponding to the data to be stored based on the customized persistent key-value storage component, and convert the data format of the data to be stored into target write data; Storing the target write data in a storage device corresponding to the storage location based on the customized persistent key-value storage component; Determining a storage location corresponding to the data to be stored based on the customized persistent key-value storage component includes: During a write operation, determining an optimal physical write location based on the semantics and access pattern of the data to be stored; Using the optimal writing physical location as the storage location; The self-defined persistent key-value storage component is used to map the logical key and the storage location to obtain a mapping structure, so as to write data based on the mapping structure.

2. The data storage method according to claim 1, wherein: The algorithm for controlling hardware resources includes at least one of an asynchronous input and output algorithm, a direct input and output algorithm, and a data mapping algorithm; The asynchronous input and output algorithm is an algorithm that serializes key-value pairs and directly submits asynchronous requests to the storage device; The direct input and output algorithm is an algorithm that bypasses the kernel cache and directly controls the storage device; The data mapping algorithm is an algorithm for mapping the physical address space of the storage device to the virtual address space of the customized persistent key-value storage component and directly reading and writing data through a pointer.

3. The data storage method according to claim 2, wherein: Determining a storage location corresponding to the data to be stored based on the customized persistent key-value storage component, and converting the data format of the data to be stored into target write data, including: Determining a target protocol corresponding to the data to be stored using the asynchronous input / output algorithm in the customized persistent key-value storage component; wherein the target protocol is a protocol corresponding to a target storage device; The asynchronous input-output algorithm is used to perform serialization processing on the data to be stored based on the target protocol to obtain the target write data; the serialization processing includes data verification, data compression and metadata encapsulation.

4. The data storage method according to claim 2, wherein: Storing the target data in a storage device corresponding to the storage location based on the customized persistent key-value storage component includes: The target write data is directly transferred to the storage device by utilizing a direct read / write mode flag for bypassing kernel cache in the direct input / output algorithm.

5. The data storage method according to claim 1, wherein: During a write operation, determining the optimal physical write location based on the semantics and access pattern of the data to be stored includes: Obtaining a semantic type and an access mode type of the data to be stored; wherein the semantic type includes at least one of a data type and a business scenario, and the access mode type is determined based on read and write behavior characteristics of the data on the storage device; A physical location that maximizes input and output efficiency is determined based on the semantic type and the access mode type to obtain the optimal write physical location.

6. The data storage method according to claim 5, characterized in that: Before obtaining the semantic type and access mode type of the data to be stored, the method further includes: Obtaining access frequency statistics, sequential statistics, and concurrent pressure statistics corresponding to the data to be stored; The access pattern type is determined based on the access frequency statistical information, the sequential statistical information and the concurrency pressure statistical information; wherein the access pattern type includes a low-frequency access pattern, a high-frequency access pattern, a sequential access pattern, a random access pattern, a high-concurrency mode and a low-concurrency mode.

7. The data storage method according to claim 1, wherein: Also includes: Get partition information; Based on the partition information, the storage device is divided into corresponding partitions using a partition strategy; wherein the partition strategy is a strategy that enables data to be distributed according to a set rule; Accordingly, storing the target data in a storage device corresponding to the storage location based on the customized persistent key-value storage component includes: Determining a target partition corresponding to the target write data based on the customized persistent key-value storage component; The target write data is written into a storage device corresponding to the storage location in the target partition.

8. The data storage method according to claim 7, characterized in that: The partition information includes at least one of a storage device identifier, an access mode, and whether the data is hot data.

9. The data storage method according to claim 7, characterized in that: After dividing the storage device into corresponding partitions using a partition strategy based on the partition information, the method further includes: Detect workload information; The divided partitions are dynamically updated based on the workload information.

10. The data storage method according to claim 7, wherein: After storing the target write data in the storage device corresponding to the storage location based on the customized persistent key-value storage component, the method further includes: Obtain the status, load and wear level of each partition according to the set time period; Determine garbage collection parameters for each partition using a machine learning model based on the state, load, and wear level of each partition; The garbage collection parameter is compared with the garbage collection threshold, and the partition data is cleaned up.

11. The data storage method according to claim 7, wherein: The partitioning strategy includes at least one of a cold and hot data separation strategy, a dynamic data distribution strategy, and a storage device wear leveling strategy.

12. The data storage method according to claim 1, wherein: The storage device includes at least one of a non-volatile memory high-speed protocol hard disk, a serial transmission solid-state hard disk, a mechanical hard disk and an object storage device.

13. The data storage method according to claim 1, wherein: Also includes: Loading the target data to be restored in the target snapshot using the snapshot directory based on the customized persistent key-value storage component; A data mapping relationship in the customized persistent key-value storage component is obtained, and the target data to be restored is restored to a corresponding storage device using the customized persistent key-value storage component based on the data mapping relationship.

14. The data storage method according to claim 13, wherein: Based on the customized persistent key-value storage component, the target data to be restored in the target snapshot is loaded using the snapshot directory, including: When a storage device fails, a snapshot directory at a recent time point is selected based on the customized persistent key-value storage component, and the target data to be restored in the target snapshot is loaded.

15. The data storage method according to claim 1, wherein: Also includes: Get data merging conditions; The customized persistent key-value storage component performs data merging using the data merging condition.

16. The data storage method according to claim 15, characterized in that: The data merging condition includes at least one of a data merging condition based on a data access pattern, a data merging condition based on a data update frequency, and a data merging condition based on a data size, wherein the data merging condition based on a data access pattern is triggered based on an access frequency threshold.

17. The data storage method according to claim 16, characterized in that: Acquiring a data merging condition, and merging the data based on the customized persistent key-value storage component using the data merging condition, including: When the data merging condition is a data merging condition based on a data access pattern, determining whether the data access frequency of the data to be merged is lower than a preset data access frequency threshold; if it is lower than the data access frequency threshold, merging the data to be merged; When the data merging condition is based on data update frequency, determining whether the update frequency of the data to be merged is lower than a preset minimum update frequency; if it is lower than the minimum update frequency, merging the data to be merged; When the data merging condition is a data merging condition based on data size, it is determined whether the size of the data to be merged is lower than a preset minimum number of bytes; when the size of the data to be merged is lower than the minimum number of bytes, the data to be merged is merged.

18. A data storage device, characterized in that: include: An interface customization module is configured to determine an algorithm for controlling hardware resources, wherein the algorithm for controlling hardware resources is an algorithm for directly manipulating resources of underlying hardware devices through software logic; the algorithm for controlling hardware resources is integrated into the interface to obtain a target storage interface, wherein the target storage interface is configured to directly control the hardware resources of each storage device; An integration module, configured to integrate the target storage interface with the persistent key-value storage component to obtain a customized persistent key-value storage component; A component acquisition module, configured to acquire the custom persistent key-value storage component; wherein the custom persistent key-value storage component is a component including an interface for directly interacting with an underlying storage device; a target write data determination module, configured to determine a storage location corresponding to the data to be stored based on the user-defined persistent key-value storage component, and convert the data format of the data to be stored into the target write data; a storage module, configured to store the target write data in a storage device corresponding to the storage location based on the custom persistent key-value storage component; The target write data determination module includes: An optimal write physical location determination unit, configured to determine an optimal write physical location according to the semantics and access mode of the data to be stored during a write operation; a storage location determining unit, configured to use the optimal writing physical location as the storage location; A mapping unit is used to map the logical key and the storage location using the customized persistent key-value storage component to obtain a mapping structure, so as to write data based on the mapping structure.

19. A data storage system, characterized in that: include: A custom persistent key-value storage component is used to determine the storage location corresponding to the data to be stored and convert the data format of the data to be stored into the target write data; storing the target write data in a storage device corresponding to the storage location; The process of obtaining the customized persistent key-value storage component includes: determining an algorithm for controlling hardware resources; wherein the algorithm for controlling hardware resources is an algorithm for directly manipulating resources of underlying hardware devices through software logic; integrating the algorithm for controlling hardware resources into an interface to obtain a target storage interface; wherein the target storage interface is used to directly control the hardware resources of each storage device; and integrating the target storage interface with the persistent key-value storage component to obtain the customized persistent key-value storage component. Determining the storage location corresponding to the data to be stored includes: during a write operation, determining an optimal physical write location based on the semantics and access mode of the data to be stored; using the optimal physical write location as the storage location; and mapping the logical key and the storage location using the custom persistent key-value storage component to obtain a mapping structure, and writing data based on the mapping structure. A storage device is used to store the target write data.

20. A data storage device, characterized in that include: memory for storing computer programs; A processor, configured to execute the computer program to implement the steps of the data storage method according to any one of claims 1 to 17.

21. A storage medium, characterized in that The storage medium stores a computer program, which, when executed by a processor, implements the steps of the data storage method according to any one of claims 1 to 17.

22. A program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, implements the steps of the data storage method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Persistent key value storage method, device and system based on OCSSD

    CN114741028A

  • Data storage method, apparatus and system, device and medium

    WO2024051109A1