Distributed storage method, data indexing method, device and storage medium

Through the new data unique identification generation rules and inner_logid structure reorganization, seq-id is used to determine the data location for MAX_ID_CNT, which solves the problem of high overhead in the data location determination in the prior art, and improves the performance and efficiency of the distributed storage system.

CN117311605BActive Publication Date: 2025-05-23SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Patent Information

Application Number
CN202311070279.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-23
Publication Date
2025-05-23
Estimated Expiration
2043-08-23

AI Technical Summary

Technical Problem

When determining data locations based on data identifiers, existing distributed storage technologies have high computing overhead, resulting in limited performance and response speed.

Method used

By formulating new data unique identification generation rules, reorganizing inner_logid structure, dividing n bits into seq-id and m bits as partition identifiers, and using seq-id to perform retrieval operations on MAX_ID_CNT to determine the logical storage space and logical block address of the data.

Benefits of technology

It reduces the overhead of hash calculation, improves data indexing efficiency, avoids the case where LBAs of different data are associated in the same bucket, and realizes that the location of the target data can be found in one memory access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117311605B_ABST
    Figure CN117311605B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer technology, and discloses a distributed storage method, a data indexing method, a device and a storage medium. Among them, the distributed storage method includes: if a storage request for the first data is detected, a first internal identifier in the first identifier corresponding to the first data is determined based on the first storage partition corresponding to the first data; then, the distributed system determines the second internal identifier in the first identifier corresponding to the first data, wherein the second internal identifier corresponding to the first data is different from the modulus result of the second internal identifier corresponding to other data in the first storage node relative to the maximum capacity of single storage node data in the first storage partition; finally, the second internal identifier is associated with the logical block address corresponding to the logical storage space of the first data in the first storage partition. Based on the above method, it is possible to improve the efficiency of storage and indexing data in the distributed storage system by formulating the generation rules of unique data identifiers in the distributed system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a distributed storage method, a data indexing method, a device and a storage medium. Background Art

[0002] At present, with the explosive growth of a large amount of data, in order to alleviate the rising storage pressure, distributed storage technology has gradually emerged. Distributed storage is to store data in multiple independent devices (for example, hosts, servers, etc.). Each storage device is equivalent to a storage node. Using multiple storage nodes to share the storage load can improve the scalability and utilization of storage resources. For example, a 500G data can be divided into multiple data slices, and the divided data slices are stored in multiple servers respectively, which can solve the problem of insufficient available storage space on a server.

[0003] When storing data through distributed storage technology, a data identifier is generated inside the distributed storage system to indicate the mapping relationship between the data and the storage location; if the user needs to access the data, the data identifier can be used as an index to determine the storage node (for example, a host, server, or other storage device), query the logical block address of the data on the storage node hard disk, and obtain the required data fragment from the data block pointed to by the logical block address. At present, minimizing the computational overhead (for example, shortening the computation time) when determining the data location based on the data identifier is an urgent problem to be solved in current distributed storage technology. Summary of the invention

[0004] To solve the above problems, the present application provides a distributed storage method, a data indexing method, a device and a storage medium.

[0005] In a first aspect, the present application provides a distributed storage method, which is applied to a distributed system, and the method includes: detecting a storage request for first data, wherein the storage request for the first data is used to store the first data in a first storage node in a first storage partition in the distributed system; determining a first internal identifier in a first identifier corresponding to the first data based on the first storage partition corresponding to the first data, wherein the first internal identifier is used to determine the first storage partition corresponding to the first data when querying the first data; determining a second internal identifier in the first identifier corresponding to the first data, wherein a first modulo result of the second internal identifier corresponding to the first data relative to a maximum data capacity of a single storage node in the first storage partition is different from a second modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node in the first storage partition, wherein the second internal identifier corresponding to the first data is used to determine the logical storage space of the first data in the first storage partition when querying the first data; and establishing an association between the second internal identifier and the logical block address corresponding to the logical storage space of the first data in the first storage partition.

[0006] In the present application, the distributed system may be the distributed storage system mentioned in the present application; the first storage partition may be the partition where the first storage node storing the first data mentioned in the present application is located; the first storage node may be the server node storing the data mentioned in the present application; the first identifier may be the data unique identifier (inner_logid) mentioned in the present application; the first internal identifier may be the identifier of the partition where the data is located divided into m bits (bits) as mentioned in the present application; the second internal identifier may be the internal unique identifier (seq-id) in the inner_logid mentioned in the present application; the maximum capacity of single storage node data in the first storage partition may be the maximum amount of data that can be accommodated in a single storage node in the partition mentioned in the present application (MAX_ID_CNT); the first remainder result may be the remainder result of the seq-id corresponding to the first data mentioned in the present application relative to the MAX_ID_CNT; the second remainder result may be the remainder result of the seq-id corresponding to other data in the first storage node mentioned in the present application relative to the MAX_ID_CNT.

[0007] It can be understood that in the present application, when a storage request to store first data on the first storage node of the first storage partition is detected, the distributed storage system will generate a first identifier corresponding to the first data, wherein the first identifier includes a first internal identifier and a second internal identifier; next, the distributed storage system will associate the second internal identifier with the logical storage space of the first data in the first storage partition, and at the same time, the logical storage space will be associated with the logical block address corresponding to the first data, thereby completing the distributed storage operation of the data.

[0008] In some embodiments, the first identifier divides n bits as the second internal identifier, and at the same time uses the remaining m bits as the identifier of the partition where the data is located (i.e., the first internal identifier).

[0009] In some embodiments, the remainder result of the second internal identifier corresponding to the first data with respect to the maximum data capacity that a single storage node in the storage partition can accommodate is different from the remainder result of the second internal identifier corresponding to other data in the first storage node with respect to the maximum data capacity that a single storage node in the storage partition can accommodate.

[0010] Based on the above method, it is possible to implement the generation rule of the new data unique identifier in the distributed storage system, and improve the efficiency of storing and indexing data in the distributed storage system.

[0011] In a possible implementation, the method for obtaining the maximum data capacity that a single storage node in the first storage partition can accommodate includes: obtaining the maximum data capacity corresponding to each storage node in the first storage partition; based on the maximum data capacity corresponding to each storage node, selecting the largest data volume as the maximum data capacity that a single storage node in the first storage partition can accommodate.

[0012] In this application, the maximum data capacity that a single storage node in the first storage partition can be the maximum data quantity (MAX_ID_CNT) that can be accommodated in a single storage node within the partition mentioned in this application.

[0013] It can be understood that in this application, the distributed storage system can send a query message to all storage nodes in the partition to obtain the maximum data capacity that each storage node can accommodate. After that, the distributed storage system selects the largest value among the values as the MAX_ID_CNT of the partition.

[0014] In a possible implementation, the second internal identifiers corresponding to the respective data in the first storage node are in an increasing state based on the storage order of the respective data.

[0015] In this application, the first storage node can be the server node for storing data mentioned in this application; the second internal identifier can be the internal unique identifier (seq-id) in the data unique identifier (inner_logid) mentioned in this application.

[0016] It can be understood that in this application, for each stored data in the storage node, the seq-id is in an increasing order. For example, the distributed storage system can generate the seq-id corresponding to the first data as 0, automatically generate the seq-id corresponding to the second data as 1, and automatically generate the seq-id corresponding to the third data as 2. In this way, it can be ensured that the seq-ids corresponding to the respective data in each storage node will never be repeated.

[0017] In a possible implementation, the above-mentioned determination of the second internal identifier in the first identifier corresponding to the first data includes: determining a set of second internal identifiers corresponding to each storage node in the first storage partition; determining a set of second internal identifiers for the first storage partition based on the set of second internal identifiers corresponding to each storage node; determining a maximum second internal identifier for the first storage partition based on the set of second internal identifiers for the first storage partition; and obtaining the second internal identifier corresponding to the first data based on the maximum second internal identifier for the first storage partition.

[0018] In the present application, the first identifier can be the data unique identifier (inner_logid) mentioned in the present application; the second internal identifier can be the internal unique identifier (seq-id) in the data unique identifier (inner_logid) mentioned in the present application; the second internal identifier set corresponding to each storage node can be the seq-id set of each storage node in the partition queried by the distributed storage system mentioned in the present application; the second internal identifier set of the first partition can be the seq-id set existing in the partition mentioned in the present application.

[0019] It can be understood that in the present application, when the distributed storage system generates the next inner_logid through the internal inner_logid generator, the distributed storage system will send a query message to all storage nodes in the partition, query the current seq-id set of each storage node, obtain the seq-id set existing in the current partition, and obtain the maximum value in the seq-id set. After that, the distributed storage system will generate the next seq-id for the partition based on the maximum value.

[0020] In a possible implementation, the above-mentioned obtaining of the second internal identifier corresponding to the first data based on the maximum second internal identifier of the first storage partition includes: generating a first increasing value that increases sequentially relative to the maximum second internal identifier based on the maximum second internal identifier of the first storage partition; if the third remainder result of the first increasing value relative to the maximum data capacity of a single storage node of the first storage partition is different from the second remainder result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, the first increasing value is determined as the second internal identifier corresponding to the first data.

[0021] In the present application, the maximum second internal identifier of the first storage partition can be the maximum internal unique identifier (seq-id) of the partition mentioned in the present application; the maximum capacity of single storage node data of the first storage partition can be the maximum amount of data that can be accommodated in a single storage node in the partition mentioned in the present application (MAX_ID_CNT); the third remainder result can be the remainder result of the incremental value mentioned in the present application relative to the MAX_ID_CNT; the second remainder result can be the remainder result of the seq-id corresponding to the existing data in the first storage node mentioned in the present application relative to the MAX_ID_CNT.

[0022] It can be understood that in the present application, when the distributed storage system generates the next inner_logid through the internal inner_logid generator, the distributed storage system will send a query message to all storage nodes in the partition, query the current seq-id set of each storage node, obtain the seq-id set existing in the current partition, and obtain the maximum value in the seq-id set. When the new value is in an increasing state on the existing seq-id set in the partition, that is, the new value increases sequentially relative to the maximum value in the seq-id set, and the new value and the seq-id existing in the partition are different in the modulo MAX_ID_CNT, the new value can be used as the new seq-id.

[0023] In a possible implementation, the above-mentioned acquisition of the second internal identifier corresponding to the first data based on the maximum second internal identifier of the first storage partition also includes: if the third remainder result of the above-mentioned first increasing value relative to the maximum data capacity of a single storage node of the first storage partition is the same as the second remainder result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, then a second increasing value that increases sequentially relative to the first increasing value is generated; based on the fourth remainder result of the second increasing value relative to the maximum data capacity of a single storage node of the first storage partition, which is different from the second remainder result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, the second increasing value is determined as the second internal identifier corresponding to the first data.

[0024] In the present application, the maximum second internal identifier of the first storage partition may be the maximum internal unique identifier (seq-id) of the partition mentioned in the present application; the maximum capacity of single storage node data of the first storage partition may be the maximum amount of data that can be accommodated in a single storage node in the partition mentioned in the present application (MAX_ID_CNT); the third remainder result may be the remainder result of the first incremental value mentioned in the present application relative to MAX_ID_CNT; the fourth remainder result may be the remainder result of the second incremental value mentioned in the present application relative to MAX_ID_CNT; the second remainder result may be the remainder result of the seq-id corresponding to the existing data in the first storage node mentioned in the present application relative to MAX_ID_CNT.

[0025] It can be understood that in the present application, since the second internal identifier corresponding to each data in the storage node is in an ascending state based on the storage order of each data, when the first incremental value and the seq-id existing in the partition have the same modulus result with respect to MAX_ID_CNT, it is necessary to generate a second incremental value. When the second incremental value and the seq-id existing in the partition have different modulus results with respect to MAX_ID_CNT, the second incremental value can be used as a new seq-id.

[0026] In a possible implementation, the first identifier includes a first internal identifier and a second internal identifier; the byte length of the first identifier is the sum of the byte length of the first internal identifier and the byte length of the second internal identifier.

[0027] In the present application, the first identifier can be the data unique identifier (inner_logid) mentioned in the present application; the first internal identifier can be the identifier of the partition where the data is located by dividing the first identifier into m bits as mentioned in the present application; the second internal identifier can be the internal unique identifier (seq-id) in the inner_logid mentioned in the present application.

[0028] It can be understood that in the present application, the inner_logid structure within the distributed storage system is reorganized, wherein n bits (bit) are divided out of the new inner_logid structure as the unique identifier (seq-id) of the inner_logid, and the remaining m bits are used as the partition identifier (ptid) of the current data (log), and the m bits plus the n bits are the overall length of the inner_logid.

[0029] In a second aspect, the present application provides a data indexing method, which is applied to a distributed system, and the method includes: detecting a query request for first data, wherein the query request includes a first identifier corresponding to the first data; determining a first storage partition corresponding to the first data based on a first internal identifier in the first identifier; determining a logical block address corresponding to a logical storage space of the first data in the first storage partition based on a second internal identifier in the first identifier, wherein the logical storage space is located at a first storage node, and a first modulo result of a second internal identifier corresponding to the first data relative to a maximum data capacity of a single storage node of the first storage partition is different from a second modulo result of a second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition; and obtaining the first data based on the logical block address.

[0030] In the present application, the distributed system may be the distributed storage system mentioned in the present application; the first storage partition may be the partition where the first storage node storing the first data mentioned in the present application is located; the first storage node may be the server node storing the data mentioned in the present application; the first identifier may be the data unique identifier (inner_logid) mentioned in the present application; the first internal identifier may be the identifier of the partition where the data is located divided into m bits as mentioned in the present application; the second internal identifier may be the internal unique identifier (seq-id) in the inner_logid mentioned in the present application; the maximum capacity of single storage node data in the first storage partition may be the maximum amount of data that can be accommodated in a single storage node in the partition mentioned in the present application (MAX_ID_CNT); the first modulus result may be the modulus result of the seq-id corresponding to the first data mentioned in the present application relative to the MAX_ID_CNT; the second modulus result may be the modulus result of the seq-id corresponding to other data in the first storage node mentioned in the present application relative to the MAX_ID_CNT.

[0031] It can be understood that in the present application, when a query request for querying the first data is detected, the distributed storage system will determine the storage partition where the first storage node storing the first data is located based on the first internal identifier corresponding to the first data, and then determine the logical storage space corresponding to the first data in the first storage node based on the second internal identifier, and then obtain the logical block address corresponding to the first data based on the logical storage space. Finally, the client can obtain the required data fragment from the data block pointed to by the logical block address, or send read and write commands to the data storage address pointed to by the logical block address.

[0032] In some embodiments, the first identifier includes a first internal identifier and a second internal identifier. The second internal identifier is associated with a logical storage space of the first data in the first storage partition, and the logical storage space is associated with a logical block address corresponding to the first data.

[0033] In some embodiments, the first identifier is divided into n bits as the second internal identifier, and the remaining m bits are used as the identifier of the partition where the data is located (ie, the first internal identifier).

[0034] In some embodiments, the modulo result of the second internal identifier corresponding to the first data relative to the maximum data capacity of a single storage node in the partition is different from the modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node in the partition.

[0035] Based on the above method, data can be searched according to the new data unique identifier in the distributed storage system, thereby improving the efficiency of indexing data in the distributed storage system.

[0036] In a possible implementation, the above-mentioned determination of the logical block address corresponding to the logical storage space of the first data in the first storage partition based on the second internal identifier in the first identifier includes: obtaining a first modulo result of the second internal identifier in the first identifier relative to the maximum data capacity of a single storage node in the first storage partition; determining the logical storage space of the first data in the first storage partition based on the first modulo result; and determining the logical block address corresponding to the first data based on the logical storage space.

[0037] In the present application, the first identifier can be the unique identifier of the data mentioned in the present application (inner_logid); the second internal identifier can be the internal unique identifier (seq-id) in the inner_logid mentioned in the present application; the maximum capacity of single storage node data in the first storage partition can be the maximum amount of data that can be accommodated in a single storage node in the partition mentioned in the present application (MAX_ID_CNT); the first modulo result can be the modulo result of the seq-id corresponding to the first data mentioned in the present application relative to MAX_ID_CNT.

[0038] It can be understood that in the present application, when searching for the data storage location, the local storage engine in the storage node can directly use the seq-id in the inner_logid corresponding to the data to modulo MAX_ID_CNT to obtain the location of the logical storage space, and then obtain the logical block address of the data from the metadata information associated with the logical storage space. The client can obtain the required data fragment from the data block pointed to by the logical block address, or send read and write commands to the data storage address pointed to by the logical block address.

[0039] It can be understood that in this application, the logical storage space corresponding to the inner_logid of different existing data can never be the same, that is, the logical block addresses corresponding to different data will not be associated in the same logical storage space, so the logical block addresses corresponding to different data will not form a linked list, and the logical block address position of the target data can be obtained with one modulo operation, and the local storage engine only needs one memory access to find the location of the target data.

[0040] In one possible implementation, the method for obtaining the maximum data capacity of a single storage node of the first storage partition includes: obtaining the maximum data capacity corresponding to each storage node in the first storage partition; based on the maximum data capacity corresponding to each storage node, selecting the data capacity with the largest value as the maximum data capacity of a single storage node of the first storage partition.

[0041] In the present application, the maximum data capacity of a single storage node of the first storage partition may be the maximum amount of data (MAX_ID_CNT) that can be accommodated in a single storage node in the partition mentioned in the present application.

[0042] It can be understood that in the present application, the distributed storage system can send query messages to all storage nodes in the partition to obtain the maximum amount of data that each storage node can accommodate, and then the distributed storage system selects the maximum value among the values ​​as the MAX_ID_CNT of the partition.

[0043] In a possible implementation, the second internal identifiers corresponding to the data in the first storage node are in an ascending state based on the storage order of the data.

[0044] In the present application, the first storage node may be a server node for storing data mentioned in the present application; the second internal identifier may be an internal unique identifier (seq-id) in the data unique identifier (inner_logid) mentioned in the present application.

[0045] It can be understood that in the present application, for each stored data in the storage node, the seq-id is in a sequentially increasing state. For example, the distributed storage system can generate a seq-id corresponding to the first data as 0, automatically generate a seq-id corresponding to the second data as 1, and automatically generate a seq-id corresponding to the third data as 2. In this way, it can be achieved that the seq-id corresponding to each data in each storage node will never be repeated.

[0046] In one possible implementation, a method for obtaining the second internal identifier corresponding to the above-mentioned first data includes: determining a set of second internal identifiers corresponding to each storage node in the first storage partition; determining a set of second internal identifiers for the first storage partition based on the set of second internal identifiers corresponding to each storage node; determining a maximum second internal identifier for the first storage partition based on the set of second internal identifiers for the first storage partition; and obtaining the second internal identifier corresponding to the first data based on the maximum second internal identifier for the first storage partition.

[0047] In the present application, the second internal identifier can be the internal unique identifier (seq-id) in the data unique identifier (inner_logid) mentioned in the present application; the second internal identifier set corresponding to each storage node can be the seq-id set of each storage node in the partition queried by the distributed storage system mentioned in the present application; the second internal identifier set of the first partition can be the seq-id set existing in the partition mentioned in the present application.

[0048] It can be understood that in the present application, when the distributed storage system generates the next inner_logid through the internal inner_logid generator, the distributed storage system will send a query message to all storage nodes in the partition, query the current seq-id set of each storage node, obtain the seq-id set existing in the current partition, and obtain the maximum value in the seq-id set. After that, the distributed storage system will generate the next seq-id for the partition based on the maximum value.

[0049] In a possible implementation, the above-mentioned obtaining of the second internal identifier corresponding to the first data based on the maximum second internal identifier of the first storage partition includes: generating a first increasing value that increases sequentially relative to the maximum second internal identifier based on the maximum second internal identifier of the first storage partition; if the third remainder result of the first increasing value relative to the maximum data capacity of a single storage node of the first storage partition is different from the second remainder result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, the first increasing value is determined as the second internal identifier corresponding to the first data.

[0050] In the present application, the maximum second internal identifier of the first storage partition can be the maximum internal unique identifier (seq-id) of the partition mentioned in the present application; the maximum capacity of single storage node data of the first storage partition can be the maximum amount of data that can be accommodated in a single storage node in the partition mentioned in the present application (MAX_ID_CNT); the third remainder result can be the remainder result of the incremental value mentioned in the present application relative to the MAX_ID_CNT; the second remainder result can be the remainder result of the seq-id corresponding to the existing data in the first storage node mentioned in the present application relative to the MAX_ID_CNT.

[0051] It can be understood that in the present application, when the distributed storage system generates the next inner_logid through the internal inner_logid generator, the distributed storage system will send a query message to all storage nodes in the partition, query the current seq-id set of each storage node, obtain the seq-id set existing in the current partition, and obtain the maximum value in the seq-id set. When the new value is in an increasing state on the existing seq-id set in the partition, that is, the new value increases sequentially relative to the maximum value in the seq-id set, and the new value and the seq-id existing in the partition are different in the modulo MAX_ID_CNT, the new value can be used as the new seq-id.

[0052] In a possible implementation, the above-mentioned acquisition of the second internal identifier corresponding to the first data based on the maximum second internal identifier of the first storage partition also includes: if the third remainder result of the first incremental value relative to the maximum data capacity of a single storage node of the first storage partition is the same as the second remainder result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, then generating a second incremental value that increases sequentially relative to the first incremental value; based on the fourth remainder result of the second incremental value relative to the maximum data capacity of a single storage node of the first storage partition, which is different from the second remainder result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, determining the second incremental value as the second internal identifier corresponding to the first data.

[0053] In the present application, the maximum second internal identifier of the first storage partition may be the maximum internal unique identifier (seq-id) of the partition mentioned in the present application; the maximum capacity of single storage node data of the first storage partition may be the maximum amount of data that can be accommodated in a single storage node in the partition mentioned in the present application (MAX_ID_CNT); the fourth remainder result may be the remainder result of the incremental value mentioned in the present application relative to the MAX_ID_CNT; the second remainder result may be the remainder result of the seq-id corresponding to the existing data in the first storage node mentioned in the present application relative to the MAX_ID_CNT.

[0054] It can be understood that in the present application, since the second internal identifier corresponding to each data in the storage node is in an ascending state based on the storage order of each data, when the first incremental value and the modulo result of the seq-id existing in the partition with respect to MAX_ID_CNT have the same modulo result, it is necessary to generate a second incremental value. When the second incremental value and the modulo result of the seq-id existing in the partition with respect to MAX_ID_CNT are different, the second incremental value can be used as the new seq-id.

[0055] In one possible implementation, the method also includes: detecting a deletion request for the second data, wherein the deletion request for the second data is used to delete the second data from the first storage node in the first storage partition in the distributed system; obtaining a second identifier corresponding to the second data; determining the first storage partition corresponding to the second data based on the first internal identifier in the second identifier; determining the logical storage space of the second data in the first storage partition based on the second internal identifier in the second identifier, wherein the logical storage space is located in the first storage node, wherein the fifth modulo result of the second internal identifier corresponding to the second data relative to the maximum data capacity of a single storage node in the first storage partition is different from the sixth modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node in the first storage partition; based on the logical storage space of the second data in the first storage partition, deleting the association between the logical block address corresponding to the second data and the logical storage space.

[0056] In the present application, the second identifier can be the unique identifier of the data (inner_logid) mentioned in the present application; the first internal identifier can be the identifier of the partition where the data is located by dividing the second identifier into m bits as mentioned in the present application; the second internal identifier can be the internal unique identifier (seq-id) in the inner_logid mentioned in the present application; the maximum capacity of single storage node data in the first storage partition can be the maximum number of data that can be accommodated in a single storage node in the partition mentioned in the present application (MAX_ID_CNT); the fifth remainder result can be the remainder result of the seq-id corresponding to the second data mentioned in the present application relative to the MAX_ID_CNT; the sixth remainder result can be the remainder result of the seq-id corresponding to other data in the first storage node mentioned in the present application relative to the MAX_ID_CNT.

[0057] It can be understood that in this application, when deleting data, it is necessary to delete the association between the logical storage space and the logical block address at the same time. First, the local storage engine in the storage node obtains the inner_logid automatically generated inside the distributed storage system when storing the data, and then the seq-id in the inner_logid corresponding to the data is used to perform a modulo calculation on the MAX_ID_CNT of the storage node. The calculation result is the unique logical storage space of the data. Finally, the association between the logical storage space and the logical block address of the data is deleted to complete the deletion operation.

[0058] It can be understood that in this application, since the logical block addresses corresponding to the inner_logid of different data currently existing in the current storage node can never be the same, there will be no scenario of logical block address conflict during the deletion phase, and the deletion operation can be completed without traversing the linked list.

[0059] In a third aspect, the present application provides an electronic device comprising: a memory and a processor, the memory being used to store instructions executed by one or more processors of the electronic device, and the processor being one of the one or more processors of the electronic device, and being used to execute the distributed storage method or data indexing method mentioned in the present application.

[0060] In a fourth aspect, the present application provides a readable storage medium, wherein the readable storage medium stores instructions, and when the instructions are executed on an electronic device, the electronic device executes the distributed storage method or data indexing method mentioned in the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1A According to some embodiments of the present application, a schematic diagram of basic components of a distributed storage system is shown;

[0062] Figure 1B According to some embodiments of the present application, a specific internal implementation scenario of querying data in a distributed storage system is shown;

[0063] Figure 2 According to some embodiments of the present application, a schematic diagram of a scenario of determining a data storage location by a hash algorithm is shown;

[0064] Figure 3 According to some embodiments of the present application, a structure of an internal identifier of a distributed storage system is shown;

[0065] Figure 4 According to some embodiments of the present application, an implementation of an internal identifier in a distributed storage method is shown;

[0066] Figure 5 According to some embodiments of the present application, a schematic diagram of a scenario of determining a data storage location by a new internal identifier generation rule is shown;

[0067] Fig. 6A According to some embodiments of the present application, the implementation of the first distributed storage method is shown;

[0068] Figure 6B According to some embodiments of the present application, a schematic diagram of a scenario for creating a mapping relationship between a logical block address and a logical storage space is shown;

[0069] Fig. 7AAccording to some embodiments of the present application, the implementation of the second distributed storage method is shown;

[0070] Figure 7B According to some embodiments of the present application, a schematic diagram of a scenario of deleting a mapping relationship between a logical block address and a logical storage space is shown;

[0071] Figure 8 According to some embodiments of the present application, a specific flow chart of a new internal identifier generation rule is shown;

[0072] Fig. 9 According to some embodiments of the present application, a specific flow chart of determining a data storage location by using a new internal identifier generation rule is shown;

[0073] Fig.10 According to some embodiments of the present application, a specific flow diagram of creating a mapping relationship between a logical block address and a logical storage space is shown;

[0074] Fig.11 According to some embodiments of the present application, a specific flow field schematic diagram of deleting a mapping relationship between a logical block address and a logical storage space is shown;

[0075] Fig.12 According to some embodiments of the present application, a schematic diagram of the hardware structure of an electronic device is shown. DETAILED DESCRIPTION

[0076] The illustrative embodiments of the present application include, but are not limited to, a distributed storage method, a data indexing method, a device, and a storage medium.

[0077] In order to more clearly understand the solution of the present application, the terms mentioned in the present application are first introduced.

[0078] Logical Block Address (LBA): It is a general mechanism for describing the block where data is located on a computer storage device. It can refer to the address of a data block or the data block represented by a certain address.

[0079] Bucket: a logical storage space. In a distributed storage system, a bucket is a container for organizing and managing data. It can organize data objects into logical collections and can contain any number of data objects. A distributed storage system can use buckets to distinguish different users, applications, or projects.

[0080] Metadata service: It is an independent component or service in a distributed storage system, responsible for managing metadata information of all data objects in the distributed storage system. Metadata includes properties, permissions, location information, etc. related to data objects. In addition, metadata service can also provide relevant query, retrieval and management interfaces for metadata. In a distributed storage system, LBA is usually stored in the metadata service.

[0081] Partition: An area consisting of storage nodes or disks where data is stored for reliability, for example, by saving replicas.

[0082] The following is a brief introduction to the principles of accessing data in distributed storage technology.

[0083] It can be understood that distributed storage is to store data in multiple independent devices (for example, hosts, servers, etc.), where each storage device is equivalent to a storage node. Using multiple storage nodes to share the storage load can improve the scalability and utilization of storage resources. Figure 1A The basic components of a distributed storage system are shown, in which the distributed storage system manages each storage node in the distributed storage system through a distributed control plane, including storage node registration, fault detection, status monitoring, etc. It can ensure that each storage node in the distributed storage system can work normally and promptly identify and handle storage node failures.

[0084] In addition, the distributed storage system has multiple client and server nodes, where the client provides a data access interface to the upper-level application services. For example, the upper-level application services include but are not limited to block storage services (Elastic Volume Service, EVS), file storage services (Scalable File Service, SFS), object storage services (Object Storage Service, OBS), distributed cache services (Distributed Cache Service, DCS), cloud databases (nosql), and database services (GaussDB), etc.

[0085] In addition, the distributed storage system also has multiple server nodes, each of which is equivalent to a storage node, which is a local data management system in the distributed storage system. The server node has a local storage engine that is responsible for managing the storage space on the storage node, including data allocation, storage and release, etc. At the same time, the data will be divided into multiple blocks or objects and stored in the storage node's disk (for example, a mechanical hard disk (Hard Disk Drive, HDD) or a solid-state drive (Solid State Drives, SSD), etc.) or memory. The local storage engine also provides a data read and write interface so that upper-level application services can access and modify data through specific internal identifiers.

[0086] When the client accesses data, the unique identifier (logid) of the data is used as an index to find the data storage location, and the partition where the data is located in the distributed storage system and the server node in the partition are queried. After determining the storage location, the client sends a read and write request to the server node. Specifically, for example, in some embodiments, the online photo album website saves pictures through a distributed storage system. Among them, the browser of the online photo album website will act as a client. When the user uses the browser and enters a search request into the browser (for example, searching for a picture named SCA1), the browser will send a request to view the picture to a specific website server in the distributed storage system (for example, the distributed control plane mentioned above). The website server will find the location information of the picture in the distributed storage system (i.e., the partition and the server node storing the picture) based on the unique identifier of the picture (for example, the unique name). The browser sends a read request to the server node through the website server, and then the server node returns the picture data to the browser and displays it on the user interface.

[0087] Specifically, Figure 1B The specific implementation scenario of accessing data in a distributed storage system is shown. When the upper-layer application service accesses data, the client will query the partition where the data is located in the distributed storage system and the server node in the partition. The server node will use the internal identifier (inner_logid) in the logid as an index to query the logical block address (Logical Block Address, LBA) of the data on the storage node hard disk, and obtain the required data fragment from the data block pointed to by the LBA, or send a read or write command to the data storage address pointed to by the LBA. For example, Figure 1BThe server node in the query uses inner_logid as an index to find the logical block addresses LBA0 and LBA1 corresponding to a piece of data on the storage node hard disk. The server node sends read and write requests to the HDD or SSD pointed to by LBA0 and LBA1 to enable the upper-level application services in the distributed storage system to access the data.

[0088] It can be understood that the internal identifier (inner_logid) of the distributed storage system plays a key role in determining the data location when storing and accessing data. Among them, inner_logid itself is an identification method generated internally by the distributed storage system. The upper-level application service can use the inner_logid as an index to store or access data.

[0089] At present, the method of hashing inner_logid is the main technical means to determine the data storage location in the current distributed storage technology. Among them, the hash algorithm refers to a mathematical computer program that receives any set of input information of any length and can directly map the input data to the corresponding storage space through calculation rules, which can provide the fastest single-point query for distributed storage technology.

[0090] In a distributed storage system, such as Figure 2 As shown, the local storage engine in the storage node calculates the hash value based on the inner_logid automatically generated inside the distributed storage system when storing data, that is, inner_logid is used as the input of the hash function (Hash()) to obtain the output result Hash(inner_logid), wherein Hash(inner_logid) is used to find the logical storage space (bucket) associated with the logical block address (Logical Block Address, LBA). Usually, LBA is stored in the metadata service in the distributed storage system. When obtaining LBA, the corresponding bucket must be found according to the output result of the hash function. After that, the distributed storage system will interact with the metadata service and return the metadata information associated with the bucket, wherein the metadata information includes the LBA of the data, and LBA refers to the physical storage address of the data block. This method of obtaining the data storage location through a one-step hash calculation improves the data indexing efficiency. However, for complex hash functions, the time overhead is large, and the calculation time will affect the performance and response speed of the entire system.

[0091] In some embodiments, the hash algorithm maps an infinite input set to a finite output set, so sometimes different data input to the hash algorithm will produce the same output result, that is, the LBAs corresponding to different data will be associated in the same bucket. Figure 2 The LBA corresponding to the data (log)0, the LBA corresponding to log1, and the LBA corresponding to log2 are all associated with bucket1. The LBAs corresponding to different data in the same bucket are connected in series in the form of a linked list. If the result cannot be hit in one access, multiple memory accesses are required. For example, if you want to access log2 data, you need to find the LBA corresponding to log2 by hashing the inner_logid corresponding to log2. However, when searching for LBA, you may need to traverse the entire linked list and access the memory three times to get the correct result, which reduces the efficiency of data indexing.

[0092] In order to solve the above problems, the embodiment of the present application provides a distributed storage method. In the distributed storage method of the present application, a new set of inner_logid generation rules is formulated. In the inner_logid generation rules mentioned in the present application, the structure of inner_logid is first reorganized, and n bits (bits) are divided out of the new inner_logid structure as the unique identifier (seq-id) of inner_logid, and the remaining m bits are used as the identifier of the partition where the data is located. Each time a new data index is added, the seq-id in the inner_logid corresponding to the data needs to be increased sequentially (for example, it can be increased sequentially from 0), which ensures that the corresponding data automatically generated inside the distributed storage system is inner_logid will not be repeated; in addition, for example, if it is determined that the data needs to be stored in the first storage node of the first partition (for example, server1), when the distributed storage system generates inner_logid, it is necessary to perform a modulo operation on the maximum amount of data (MAX_ID_CNT) that can be accommodated in a single storage node in the first partition based on seq-id, to ensure that the results of the modulo operation of MAX_ID_CNT for the seq-id corresponding to each data in the first storage node cannot be the same, and then the modulo result is associated with the actual stored data logical storage space (bucket). It can be understood that since bucket is associated with LBA, the modulo result is also associated with LBA. For example, if MAX_ID_CNT is 10 and seq-id is 6, seq-id (value 6) is divided by MAX_ID_CNT (value 10), and the quotient is 0, the modulo result is 6, and the modulo result 6 is associated with the data logical storage space where the data is actually stored.

[0093] Correspondingly, the present application provides a data indexing method, that is, when using inner_logid to locate the data storage location, the partition is first determined based on the partition identifier in the inner_logid corresponding to the data query request, and then the seq-id in the inner_logid is used to obtain the location of the logical storage space (bucket) corresponding to the data by taking the modulus of the MAX_ID_CNT of the partition corresponding to the partition identifier in the inner_logid, and then the LBA of the data is obtained from the metadata information associated with the bucket. The client can obtain the required data fragment from the data block pointed to by the LBA, or send a read and write command to the data storage address pointed to by the LBA. Since the inner_logid generator ensures that the modulus result of the seq-id existing in the storage node to the MAX_ID_CNT is never the same, the buckets corresponding to the inner_logid of different data can never be the same, that is, the LBAs of different data will not be associated in the same bucket, so the LBAs corresponding to different data will not form a linked list, and the LBA position of the target data can be obtained by a single modulus operation, that is, the local storage engine only needs one memory access to find the location of the target data. In this way, through the above method, the overhead caused by hash calculation that may require traversing the memory multiple times can be avoided, thereby improving indexing efficiency.

[0094] In some embodiments, Figure 3 As shown, the inner_logid structure inside the distributed storage system is reorganized, and n bits (bits) are divided out of the new inner_logid structure as the unique identifier (seq-id) of the inner_logid, and the remaining m bits are used as the partition identifier (ptid) of the current data (log), and the m bits plus the n bits are the overall length of the inner_logid. Among them, for the storage nodes in each partition, the seq-id is in a sequentially increasing state and never repeats. For example, the distributed storage system can generate the seq-id corresponding to the first segment of data log0 as 0 according to the newly formulated inner_logid generation rule, and automatically generate the seq-id corresponding to the second segment of data log1 as 1, and automatically generate the seq-id corresponding to the third segment of data log2 as 2; if new data is stored later, the distributed storage system will increase the seq-id corresponding to each data from 3 in sequence, and judge whether the incremented value meets the inner_logid generation rule, so that the inner_logid corresponding to each data automatically generated inside the distributed storage system is never the same and has unique identification.

[0095] In some embodiments, when the distributed storage system generates the seq-id in inner_logid internally, it is necessary to ensure that the results of the modulo operation of the seq-id corresponding to each data in all storage nodes in the partition with respect to MAX_ID_CNT cannot be the same, wherein MAX_ID_CNT is the maximum number of logs that can be accommodated in a single storage node. When the inner_logid generator generates the next inner_logid, the distributed storage system will send a message to all storage nodes in the partition, query the current seq-id set, obtain the seq-id set existing in the current partition, and obtain the maximum value in the seq-id set. When the new value is in an increasing state on the existing seq-id set in the partition, that is, the new value increases sequentially relative to the maximum value in the seq-id set, and the modulo operation of the new value and the seq-id existing in the partition with respect to MAX_ID_CNT is different, the new value can be used as the new seq-id.

[0096] In some embodiments, the distributed storage system may send a query message to all storage nodes in the partition to obtain the maximum amount of data that each storage node can accommodate, and then the distributed storage system selects the maximum value among the values ​​as the MAX_ID_CNT of the partition.

[0097] In some embodiments, when searching for the data storage location, for example, when a data query request is detected, the local storage engine in the storage node can directly use the seq-id in the inner_logid corresponding to the data to take the modulus of MAX_ID_CNT to obtain the location of the bucket, and then obtain the LBA of the data from the metadata information associated with the bucket. The client can obtain the required data fragment from the data block pointed to by the LBA, or send a read and write command to the data storage address pointed to by the LBA. Since the inner_logid generator mentioned in this application ensures that the existing seq-id modulo MAX_ID_CNT in the same partition is never the same, the buckets corresponding to the inner_logid of different existing data can never be the same, that is, the LBAs corresponding to different data will not be associated in the same bucket, so the LBAs corresponding to different data will not form a linked list, and the LBA position of the target data can be obtained by a single modulo operation, and the local storage engine only needs one memory access to find the location of the target data.

[0098] In addition, it can be understood that for the above hash algorithm, since the LBAs corresponding to different data associated in the same bucket are connected in series in the form of a linked list, if you want to delete or add an LBA corresponding to a log in the middle of the linked list, concurrency problems will arise. For example, if the current order in the linked list is LBA1→LBA2, if LBA3 and LBA4 exist at the same time and need to be saved in the linked list, they will simultaneously obtain the current linked list order LBA1→LBA2, resulting in LBA3 being saved after LBA2, and LBA4 being saved after LBA2, causing a conflict. In order to solve the concurrency problem, additional protection measures such as memory barriers need to be added, so that the entire linked list is locked when adding or deleting LBAs, and other operations can only be started after a complete operation is completed, but memory barriers will increase the consumption of computing resources and reduce the performance of distributed storage system processing programs.

[0099] The method provided in the present application ensures that inner_logid can point to a unique location in the storage structure, avoiding concurrency problems caused by associating LBAs of different data with the same bucket.

[0100] For example, in some embodiments, if a storage request is detected, that is, when new data needs to be stored, the logical block address indicating the physical address of the data, i.e., the LBA, needs to be associated with the logical storage space (bucket). In which, when determining the location of the bucket, the distributed storage system first generates a new inner_logid, and then needs to perform a modulo calculation on the MAX_ID_CNT of the storage node through the seq-id in the new inner_logid. The calculated result is the unique bucket of the data, and finally the bucket and LBA are associated to complete the operation of creating new data storage. Since the buckets corresponding to the inner_logid of different data in the current storage node can never be the same, there will be no bucket conflict scenario when creating the association information between the new bucket and LBA, and the creation operation can be completed without building a linked list.

[0101] In some embodiments, when a deletion request is detected, that is, when data needs to be deleted, the distributed storage system will simultaneously delete the association between the bucket and the LBA. First, the local storage engine in the storage node obtains the inner_logid automatically generated inside the distributed storage system when storing the data, and then the MAX_ID_CNT of the storage node needs to be modulo calculated by the seq-id in the inner_logid corresponding to the data. The calculated result is the unique bucket of the data, and finally the association between the bucket and the LBA of the data is deleted to complete the deletion operation. Since the buckets corresponding to the inner_logid of different data currently existing in the current storage node can never be the same, there will be no bucket conflict scenario in the deletion stage, and the deletion operation can be completed without traversing the linked list.

[0102] In some embodiments, since the buckets corresponding to the inner_logid of different data currently existing in the current storage node can never be the same, that is, the LBAs of different data will not be associated in the same bucket, there will not be a situation where the LBAs corresponding to different data in the same bucket are connected in series in the form of a linked list. The operations on the inner_logid of each data are completely independent, and the operations on different data will not cause related problems of concurrency control. For example, if you want to add or delete the association information between the LBA corresponding to a data and the bucket, it will not affect the subsequent association between the LBA corresponding to other data and the bucket.

[0103] It can be understood that in some embodiments, the LBA is stored in the metadata service in the distributed storage system. When obtaining the LBA, the corresponding bucket must be found first. Then the distributed storage system interacts with the metadata service and returns the metadata information associated with the bucket, where the metadata information includes the LBA of the data. LBA refers to the physical storage address of the data block.

[0104] It can be understood that through the method in the present application, the generation rules of inner_logid can be reformulated, and the local storage engine of the distributed storage system can be optimized based on this rule. In this way, the hash algorithm in the distributed storage can be removed to avoid the overhead generated by the hash calculation. At the same time, it is ensured that inner_logid can point to a unique location in the storage structure, avoiding concurrency problems caused by associating LBAs of different data in the same bucket.

[0105] The above-mentioned distributed storage method of the present application can be executed by an electronic device, which can be a server, or a terminal device such as a mobile phone, a computer, a virtual reality (VR) device, a tablet computer, an augmented reality (AR) device, a laptop computer, etc. The present application does not make any limitation here.

[0106] Combine the following Figure 4 , briefly introduces the generation rules of inner_logid mentioned in this application.

[0107] like Figure 4 As shown in the figure, there are three storage nodes in partition 0, namely Server0, Server1 and Server2. There are two seq-ids corresponding to data in Server0, and the current seq-id set of Server0 is {0,2}; there are two seq-ids corresponding to data in Server1, and the current seq-id set of Server1 is {0,2}; there is one seq-id corresponding to data in Server2, and the current seq-id set of Server2 is {0}. At the same time, set the maximum number of logs that a single storage node in partition 0 can accommodate MAX_ID_CNT. For example, MAX_ID_CNT can be set to 3, that is, each storage node in partition 0 can only store up to three pieces of data.

[0108] When the inner_logid generator generates the next inner_logid, the distributed storage system sends a message to all storage nodes in partition 0 to query the seq-id set of the current partition 0, and obtains the seq-id set existing in the current partition 0 as {0,2}. Since the seq-id is in a sequentially increasing state, when the values ​​0 and 2 exist, it means that the seq-id has previously used the value 1, and the data corresponding to the seq-id 1 may have been deleted.

[0109] Since the maximum value in the current seq-id set in partition 0 is 2, and the seq-id is in an increasing state, the new seq-id must start from 3 to determine whether it meets the inner_logid generation rule. In this distributed storage system, the result set of the modulus of each seq-id in the seq-id set existing in partition 0 with the above-preset MAX_ID_CNT is {0,2}. Since it is necessary to ensure that the modulus of the new value and the seq-id existing in the partition with respect to MAX_ID_CNT is different, the result of the modulus of the newly added seq-id with respect to MAX_ID_CNT cannot be 0 and 2. Therefore, the modulus result of the value 3 with respect to MAX_ID_CNT is 0, which does not meet the generation rule. The modulus result of the next increasing value 4 with respect to MAX_ID_CNT is 1, and 1 is not in the current modulus result set of partition 0, so the next seq-id of partition 0 is 4. In this way, a new inner_logid can be generated by the method in this application.

[0110] Combine the following Figure 5 , briefly introduces the scenario in which the distributed storage system determines the data storage location through the inner_logid generation rule.

[0111] like Figure 5 As shown, there are data log0 and data log1 in the storage node, where the seq-id in inner_logid corresponding to log0 is seq-id0, and the seq-id in inner_logid corresponding to log1 is seq-id1. If the storage location of data log1 needs to be found, the local storage engine in the storage node will directly use the seq-id1 corresponding to log1 to modulo MAX_ID_CNT to obtain the location of the bucket. For example, the bucket corresponding to log1 is bucket1, and then the LBA of data log1 can be obtained from the metadata information associated with bucket1. The client can obtain the required data fragment from the data block pointed to by the LBA, or send a read and write command to the data storage address pointed to by the LBA.

[0112] For another example, if you need to find the storage location of data log0, the local storage engine in the storage node will directly use the seq-id0 corresponding to log0 to modulo MAX_ID_CNT to obtain the bucket location. For example, the bucket corresponding to log0 is bucket0, and then the LBA of the data log0 can be obtained from the metadata information associated with bucket0. The client can obtain the required data fragment from the data block pointed to by the LBA, or send a read and write command to the data storage address pointed to by the LBA. If the data of log0 is backed up, multiple LBAs will be generated (for example, LBA head, LBA1, LBA2, etc.), and these LBAs will be associated in bucket0 in the form of a linked list. If the LBA head fails, the data log0 can also be found through the address pointed to by LBA1.

[0113] Since the inner_logid generator ensures that the modulo results of seq-id existing in the storage node with respect to MAX_ID_CNT in the same partition are never the same, the buckets corresponding to the inner_logid of different existing data can never be the same, that is, the LBAs corresponding to different data will not be associated in the same bucket, so the LBAs corresponding to different data will not form a linked list. A modulo operation can obtain the LBA position of the target data, and the local storage engine only needs one memory access to find the location of the target data.

[0114] Combine the following Fig. 6A and Figure 6B , briefly introduces the scenario of creating an association between buckets and LBAs when storing data.

[0115] like Fig. 6A As shown, there is data log0 in the storage node, wherein the seq-id in inner_logid corresponding to log0 is seq-id0, and the bucket corresponding to log0 is bucket0, that is, the position of bucket0 can be obtained by taking the modulus of MAX_ID_CNT by seq-id0.

[0116] At this time, if Figure 6BAs shown, if a new piece of data log1 is to be stored in the storage node, the distributed storage system needs to generate a new inner_logid according to the above inner_logid generation rule, and then perform a modulo calculation on the MAX_ID_CNT of the storage node through the seq-id in the new inner_logid (for example, seq-id1), and the calculation result is the unique bucket of the data (that is, bucket1), and finally the association of bucket1 and the logical block address LBA1 indicating the physical address of the data log1 can complete the creation operation of the new data storage.

[0117] Since the buckets corresponding to the inner_logid of different data in the current storage node can never be the same, when creating a scenario where there will be no bucket conflict, the creation operation can be completed without building a linked list.

[0118] Combine the following Fig. 7A and Figure 7B , briefly introduces the scenario of deleting the association between bucket and LBA when deleting data.

[0119] like Fig. 7A As shown, there are data log0 and data log1 in the storage node, wherein the seq-id in inner_logid corresponding to log0 is seq-id0, and the bucket corresponding to log0 is bucket0; the seq-id in inner_logid corresponding to log1 is seq-id1, and the bucket corresponding to log1 is bucket1.

[0120] At this time, if you want to delete data log0 in the storage node, you need to delete the association between bucket and LBA at the same time. First, the local storage engine in the storage node obtains the inner_logid automatically generated inside the distributed storage system when storing log0, and then needs to perform a modulo calculation on the MAX_ID_CNT of the storage node through the seq-id0 in the inner_logid corresponding to log0. The calculation result is the unique bucket of the data (i.e. bucket0). Finally, delete the association between bucket0 and the logical block address LBA0 indicating the physical address of data log0 to complete the deletion operation.

[0121] Since the buckets corresponding to the inner_logid of different data currently stored in the current storage node can never be the same, there will be no bucket conflict scenario during the deletion phase, and the deletion operation can be completed without traversing the linked list.

[0122] The following is based on Figure 8The flow chart shown in FIG. 1 describes the inner_logid generation rule in the embodiment of the present application. Figure 8 Specifically, the method comprises the following steps:

[0123] S801: Organize the structure of inner_logid.

[0124] It can be understood that in the embodiment of the present application, a new set of inner_logid generation rules is formulated, and in the new inner_logid generation rules, the structure of inner_logid needs to be reorganized. In the new inner_logid structure, n bits are divided as the unique identifier (seq-id) of inner_logid, and the remaining m bits are used as the partition identifier (ptid) of the current data. The m bits plus the n bits are the overall length of the inner_logid.

[0125] S802: Obtain the maximum amount of data MAX_ID_CNT that a single storage node in the partition can accommodate.

[0126] It can be understood that in the embodiment of the present application, it is necessary to set the maximum amount of data (log) that a single storage node in a partition can accommodate. For example, Figure 4 As shown, the maximum amount of data that can be accommodated by a single storage node in partition 0 is set to 3; for another example, in a large storage server, the maximum amount of data that can be accommodated by a single storage node in a partition can be set to 50000. This application does not limit this.

[0127] In some embodiments, the distributed storage system may send a query message to all storage nodes in the partition to obtain the maximum amount of data that each storage node can accommodate, and then the distributed storage system selects the maximum value among the values ​​as the MAX_ID_CNT of the partition.

[0128] S803: Send a message to all storage nodes in the partition to obtain the seq-id set of each storage node.

[0129] It can be understood that in the embodiment of the present application, the distributed storage system will send an inquiry message to all storage nodes in the partition to obtain the current seq-id set of each storage node.

[0130] S804: Take the union of all seq-id sets.

[0131] It can be understood that in the embodiment of the present application, after obtaining the current seq-id set of each storage node in the partition, the union of all sets is taken to obtain the seq-id set of the partition.

[0132] In some embodiments, for example, Figure 4 As shown, the distributed storage system sends messages to all storage nodes in partition 0, queries the seq-id set of the current partition 0, and obtains that the seq-id set existing in the current partition 0 is {0,2}.

[0133] S805: The seq-id set in the partition is modulo MAX_ID_CNT to obtain the modulo set of the partition.

[0134] It can be understood that in the embodiment of the present application, after the partition seq-id set is obtained through step S804, it is necessary to perform a modulo operation on MAX_ID_CNT for each seq-id in the set to obtain the partition modulo set.

[0135] S806: Get the next incrementing value.

[0136] It can be understood that in the embodiment of the present application, the seq-id of each storage node in the partition is in an increasing state, for example, it can be increased sequentially starting from 0. When a new seq-id is generated, it is necessary to obtain the next increasing value.

[0137] In some embodiments, for example, the distributed storage system can generate a seq-id of 0 corresponding to the first segment of data log0, automatically generate a seq-id of 1 corresponding to the second segment of data log1, and automatically generate a seq-id of 2 corresponding to the third segment of data log2 according to the newly formulated inner_logid generation rule. If new data is stored thereafter, the distributed storage system will increase the seq-id corresponding to each data in sequence from 3 to determine whether it meets the inner_logid generation rule.

[0138] In some embodiments, the distributed storage system obtains the maximum value in the seq-id set. When the new value is in an increasing state on the existing seq-id set in the partition, that is, the new value increases sequentially relative to the maximum value in the seq-id set, the distributed storage system will then determine whether the increasing value satisfies the inner_logid generation rule.

[0139] S807: Determine whether the result of taking the modulus of the increasing value to MAX_ID_CNT is the same as the data in the partition modulus set. If so, proceed to S806; if not, proceed to S808.

[0140] It can be understood that in the embodiment of the present application, it is necessary to perform a modulo operation on the MAX_ID_CNT of the incremental value obtained through S806. If the calculation result is the same as a data in the modulo set of the partition, it does not comply with the inner_logid generation rule mentioned in the present application, and go to S806 to continue selecting the next incremental value; if the calculation result of the modulo operation on the incremental value is completely different from the data in the modulo set of the partition, it complies with the inner_logid generation rule mentioned in the present application, and go to S808 to use the incremental value as the next seq-id of the partition.

[0141] S808: Use the value as the next seq-id of the partition to obtain a new inner_logid.

[0142] It can be understood that in the embodiment of the present application, the new value is in an increasing state on the existing seq-id set in the partition, and when the new value and the seq-id existing in the partition are different in the modulo MAX_ID_CNT, the new value can be used as the new seq-id, thereby obtaining a new inner_logid.

[0143] Through the above method, a new inner_logid generation rule can be obtained. The distributed storage system can perform adaptive optimization based on this channel generation rule, so that the inner_logid corresponding to each data automatically generated inside the distributed storage system is never the same and has unique identification.

[0144] The following is based on Fig. 9 The flowchart shown in FIG. 1 describes a method for determining a data storage location by using the inner_logid generation rule in an embodiment of the present application. Fig. 9 Specifically, the method comprises the following steps:

[0145] S901: Obtain inner_logid corresponding to the data.

[0146] It can be understood that in the embodiment of the present application, when accessing data, the local storage engine in the storage node obtains the corresponding inner_logid automatically generated inside the distributed storage system when storing the data.

[0147] In some embodiments, the inner_logid corresponding to each piece of data automatically generated within the distributed storage system is never the same and has a unique identifier, and the storage location of the data can be found through the inner_logid as an index.

[0148] S902: Obtain the maximum amount of data MAX_ID_CNT that a single storage node in the partition can accommodate.

[0149] It can be understood that in the embodiment of the present application, it is necessary to set the maximum amount of data (log) that a single storage node in a partition can accommodate. For example, Figure 4 As shown, the maximum amount of data that can be accommodated by a single storage node in partition 0 is set to 3; for another example, in a large storage server, the maximum amount of data that can be accommodated by a single storage node in a partition can be set to 50000. This application does not limit this.

[0150] In some embodiments, the distributed storage system may send a query message to all storage nodes in the partition to obtain the maximum amount of data that each storage node can accommodate, and then the distributed storage system selects the maximum value among the values ​​as the MAX_ID_CNT of the partition.

[0151] S903: Use the seq-id in inner_logid to modulo MAX_ID_CNT to obtain the bucket position corresponding to the data.

[0152] It can be understood that in the embodiment of the present application, after obtaining the inner_logid corresponding to the data and the MAX_ID_CNT of the partition, it is necessary to use the seq-id in the inner_logid to perform a modulo operation on the MAX_ID_CNT to obtain the location of the unique bucket corresponding to the data.

[0153] In some embodiments, the inner_logid generation rule ensures that the modulo MAX_ID_CNT result of the seq-id existing in the storage node is never the same, so the bucket corresponding to the inner_logid of different existing data can never be the same, that is, a unique bucket position can be obtained.

[0154] S904: Obtain the LBA corresponding to the data in the bucket.

[0155] It can be understood that in the embodiment of the present application, there is an association relationship between the bucket and the LBA. After confirming the bucket position corresponding to the data, the LBA corresponding to the data can be found in the bucket.

[0156] In some embodiments, LBA may refer to the address of a data block or the data block represented by an address.

[0157] In some embodiments, the LBA is stored in a metadata service in a distributed storage system. When obtaining the LBA, the corresponding bucket must be found first. The distributed storage system then interacts with the metadata service and returns metadata information associated with the bucket, where the metadata information includes the LBA of the data. LBA refers to the physical storage address of the data block.

[0158] S905: Send a read or write command to the data storage address pointed to by the LBA.

[0159] It can be understood that after finding the LBA corresponding to the data, the required data fragments can be obtained from the data block pointed to by the LBA, or a read or write command can be sent to the data storage address pointed to by the LBA.

[0160] In the embodiment of the present application, since the inner_logid generator ensures that the modulo results of the existing seq-id with respect to MAX_ID_CNT in the same partition are never the same, the buckets corresponding to the inner_logids of different existing data can never be the same, that is, the LBAs corresponding to different data will not be associated in the same bucket, and therefore the LBAs corresponding to different data will not form a linked list, and the LBA position of the target data can be obtained with one modulo operation, and the local storage engine only needs one memory access to find the position of the target data.

[0161] The following is based on Fig.10 The flowchart shown in FIG. 1 describes a method for creating an association relationship between a bucket and an LBA when storing data in an embodiment of the present application. Fig.10 Specifically, the method comprises the following steps:

[0162] S1001: Organize the structure of inner_logid.

[0163] It can be understood that in the embodiment of the present application, a new set of inner_logid generation rules is formulated, and in the new inner_logid generation rules, the structure of inner_logid needs to be reorganized. In the new inner_logid structure, n bits are divided as the unique identifier (seq-id) of inner_logid, and the remaining m bits are used as the partition identifier (ptid) of the current data. The m bits plus the n bits are the overall length of the inner_logid.

[0164] S1002: Obtain the maximum amount of data MAX_ID_CNT that a single storage node in the partition can accommodate.

[0165] It can be understood that in the embodiment of the present application, it is necessary to set the maximum amount of data (log) that a single storage node in a partition can accommodate. For example, Figure 4 As shown, the maximum amount of data that can be accommodated by a single storage node in partition 0 is set to 3; for another example, in a large storage server, the maximum amount of data that can be accommodated by a single storage node in a partition can be set to 50000. This application does not limit this.

[0166] In some embodiments, the distributed storage system may send a query message to all storage nodes in the partition to obtain the maximum amount of data that each storage node can accommodate, and then the distributed storage system selects the maximum value among the values ​​as the MAX_ID_CNT of the partition.

[0167] S1003: Send a message to all storage nodes in the partition to obtain the seq-id set of each storage node.

[0168] It can be understood that in the embodiment of the present application, the distributed storage system will send an inquiry message to all storage nodes in the partition to obtain the current seq-id set of each storage node.

[0169] S1004: Take the union of all seq-id sets.

[0170] It can be understood that in the embodiment of the present application, after obtaining the current seq-id set of each storage node in the partition, the union of all sets is taken to obtain the seq-id set of the partition.

[0171] In some embodiments, for example, Figure 4 As shown, the distributed storage system sends messages to all storage nodes in partition 0, queries the seq-id set of the current partition 0, and obtains that the seq-id set existing in the current partition 0 is {0,2}.

[0172] S1005: The seq-id set in the partition is modulo MAX_ID_CNT to obtain the modulo set of the partition.

[0173] It can be understood that in the embodiment of the present application, after the partition seq-id set is obtained through step S1004, it is necessary to perform a modulo operation on MAX_ID_CNT for each seq-id in the set to obtain the partition modulo set.

[0174] S1006: Get the next incrementing value.

[0175] It can be understood that in the embodiment of the present application, the seq-id of each storage node in the partition is in an increasing state, for example, it can be sequentially increased starting from 0. When a new seq-id is generated, it is necessary to obtain the next increasing value.

[0176] In some embodiments, for example, the distributed storage system can generate a seq-id of 0 corresponding to the first segment of data log0, automatically generate a seq-id of 1 corresponding to the second segment of data log1, and automatically generate a seq-id of 2 corresponding to the third segment of data log2 according to the newly formulated inner_logid generation rule. If new data is stored thereafter, the distributed storage system will increase the seq-id corresponding to each data from 3 in sequence to determine whether the inner_logid generation rule of step S1007 is met.

[0177] In some embodiments, the distributed storage system obtains the maximum value in the seq-id set. When the new value is in an increasing state on the existing seq-id set in the partition, that is, the new value increases sequentially relative to the maximum value in the seq-id set, the distributed storage system will then determine whether the increasing value satisfies the inner_logid generation rule.

[0178] S1007: Determine whether the result of taking the modulus of the increasing value to MAX_ID_CNT is the same as the data in the partition modulus set. If so, proceed to S1006; if not, proceed to S1008.

[0179] It can be understood that in the embodiment of the present application, it is necessary to perform a modulo operation on the MAX_ID_CNT of the incremental value obtained through S1006. If the calculation result is the same as a data in the modulo set of the partition, it does not comply with the inner_logid generation rule mentioned in the present application, and go to S1006 to continue selecting the next incremental value; if the calculation result of the modulo operation on the incremental value is different from all the data in the modulo set of the partition, it complies with the inner_logid generation rule mentioned in the present application, and go to S1008 to use the incremental value as the next seq-id of the partition.

[0180] S1008: Use the value as the next seq-id of the partition to obtain a new inner_logid.

[0181] It can be understood that in the embodiment of the present application, the new value is in an increasing state on the existing seq-id set in the partition, and when the new value and the seq-id existing in the partition are different in the modulo MAX_ID_CNT, the new value can be used as the new seq-id, thereby obtaining a new inner_logid.

[0182] S1009: Use the seq-id in the newly generated inner_logid to modulo MAX_ID_CNT to obtain the bucket position corresponding to the data.

[0183] It can be understood that in the embodiment of the present application, after the distributed storage system needs to generate a new inner_logid according to the inner_logid generation rule shown in the above steps S1001 to S1008, it is necessary to perform a modulo calculation on the partition's MAX_ID_CNT using the seq-id in the new inner_logid, and the calculation result is the unique bucket of the data.

[0184] S1010: Associate the bucket and LBA.

[0185] It can be understood that in the embodiment of the present application, after obtaining the bucket corresponding to the data through the above step S1009, the association information between the bucket and the LBA can be established.

[0186] In some embodiments, in some embodiments, LBA may refer to an address of a data block or a data block represented by an address.

[0187] In some embodiments, the LBA is stored in a metadata service in a distributed storage system. When obtaining the LBA, the corresponding bucket must be found first. The distributed storage system then interacts with the metadata service and returns metadata information associated with the bucket, where the metadata information includes the LBA of the data. LBA refers to the physical storage address of the data block.

[0188] In the embodiment of the present application, since the buckets corresponding to the inner_logid of different data existing in the current storage node can never be the same, there will be no bucket conflict scenario when creating the association relationship between the bucket and the LBA, and the creation operation can be completed without building a linked list.

[0189] The following is based on Fig.11 The flowchart shown in FIG. 1 describes a method for deleting the association relationship between the bucket and LBA corresponding to the data when deleting the data in an embodiment of the present application. Fig.11 Specifically, the method comprises the following steps:

[0190] S1101: Obtain inner_logid corresponding to the data.

[0191] It can be understood that in the embodiment of the present application, when accessing data, the local storage engine in the storage node obtains the corresponding inner_logid automatically generated inside the distributed storage system when storing the data.

[0192] In some embodiments, the inner_logid corresponding to each piece of data automatically generated within the distributed storage system is never the same and has a unique identifier, and the storage location of the data can be found through the inner_logid as an index.

[0193] S1102: Obtain the maximum amount of data MAX_ID_CNT that a single storage node in the partition can accommodate.

[0194] It can be understood that in the embodiment of the present application, it is necessary to set the maximum amount of data (log) that a single storage node in a partition can accommodate. For example, Figure 4 As shown, the maximum amount of data that can be accommodated by a single storage node in partition 0 is set to 3; for another example, in a large storage server, the maximum amount of data that can be accommodated by a single storage node in a partition can be set to 50000. This application does not limit this.

[0195] In some embodiments, the distributed storage system may send a query message to all storage nodes in the partition to obtain the maximum amount of data that each storage node can accommodate, and then the distributed storage system selects the maximum value among the values ​​as the MAX_ID_CNT of the partition.

[0196] S1103: Use the seq-id in inner_logid to modulo MAX_ID_CNT to obtain the bucket position corresponding to the data.

[0197] It can be understood that in the embodiment of the present application, after obtaining the inner_logid corresponding to the data and the MAX_ID_CNT of the partition, it is necessary to use the seq-id in the inner_logid to perform a modulo operation on the MAX_ID_CNT to obtain the location of the unique bucket corresponding to the data.

[0198] In some embodiments, the inner_logid generation rule ensures that the modulo MAX_ID_CNT result of the seq-id existing in the storage node is never the same, so the bucket corresponding to the inner_logid of different existing data can never be the same, that is, a unique bucket position can be obtained.

[0199] S1104: Delete the correspondence between the bucket and the LBA corresponding to the data.

[0200] It can be understood that in the embodiment of the present application, the association relationship between LBA exists in the bucket. After confirming the bucket position corresponding to the data, the LBA of the data can be obtained from the metadata information associated with the bucket, and then the association relationship between the bucket and the LBA of the data can be deleted to complete the deletion operation.

[0201] In an embodiment of the present application, since the buckets corresponding to the inner_logid of different data currently existing in the current storage node can never be the same, there will be no bucket conflict scenario when deleting the association relationship between the bucket corresponding to the data and the LBA, and the deletion operation can be completed without traversing the linked list.

[0202] The above-mentioned distributed storage method of the present application can be executed by an electronic device, which can be a server, or a terminal device such as a mobile phone, a computer, a virtual reality (VR) device, a tablet computer, an augmented reality (AR) device, a laptop computer, etc. The present application does not make any limitation here.

[0203] like Fig.12 As shown, a hardware structure diagram of an electronic device 1200 according to an embodiment of the present application is exemplified. Fig.12 As shown, the electronic device 1200 may include one or more processors 1202, a system control logic 1201 connected to at least one of the processors 1202, a system memory 1205 connected to the system control logic 1201, a storage 1203 connected to the system control logic 1201, and a network interface 1208 connected to the system control logic 1201.

[0204] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute the only possible implementation method for the electronic device 1200. In other embodiments of the present application, the electronic device 1200 may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0205] The processor 1202 may include one or more single-core or multi-core processors. In some embodiments, the processor 1202 may include any combination of a general-purpose processor and a dedicated processor (e.g., an application processor, a baseband processor, etc.). It is understood that in the embodiment of the present application, the processor 1202 may be configured to execute the executable instructions 1204 stored in the memory 1203 to implement the distributed storage method of the embodiment of the present invention.

[0206] The system control logic 1201 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1202 and / or any suitable device or component in communication with the system control logic 1201. The system control logic 1201 may include one or more memory controllers to provide an interface to the system memory 1205. The system memory 1205 may be used to load and store data and / or instructions. In some embodiments, the system memory 1205 of the electronic device 1200 may include any suitable volatile memory, such as a suitable dynamic random access memory.

[0207] The memory 1203 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the memory 1203 may include any suitable volatile memory and / or any suitable non-volatile storage device, for example, the memory 1203 may include: a random access memory unit (Random Access Memory, RAM) and / or a cache memory unit, and may further include a read-only memory unit (Read Only memory, ROM).

[0208] The memory 1203 may include a portion of storage resources on a device on which the electronic device 1200 is installed, or it may be accessible by the device but is not necessarily a portion of the device. For example, the memory 1203 may be accessed via a network via the network interface 1208 .

[0209] In particular, the system memory 1205 and the storage 1203 may respectively include: a temporary copy and a permanent copy of the instruction 1206 and a temporary copy and a permanent copy of the instruction 1204. The instruction 1206 may include: an instruction that causes the electronic device 1200 to implement the distributed storage method of the embodiment of the present invention when executed by at least one of the processors 1202. In some embodiments, the instruction 1206, hardware, firmware and / or its software components may be additionally / alternatively placed in the system control logic 1201, the network interface 1208 and / or the processor 1202.

[0210] The network interface 1208 may include a transceiver for providing a radio interface for the electronic device 1200, and then communicating with any other suitable device (such as a front-end module, an antenna, etc.) through one or more networks. In some embodiments, the network interface 1208 may be integrated with other components of the electronic device 1200. For example, the network interface 1208 may be integrated with at least one of the processor 1202, the system memory 1205, the storage 1203, and a firmware device (not shown) having instructions. When at least one of the processors 1202 executes the instructions, the electronic device 1200 implements the distributed storage method of the embodiment of the present invention.

[0211] The network interface 1208 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 1208 may be a network adapter, a wireless network adapter, a telephone modem and / or a wireless modem.

[0212] The electronic device 1200 may further include an input / output (I / O) device 1207. The I / O device 1207 may include a user interface to enable a user to interact with the electronic device 1200; the design of the peripheral component interface enables the peripheral components to interact with the electronic device 1200. In some embodiments, the electronic device 1200 further includes a sensor for determining at least one of an environmental condition and location information related to the electronic device 1200.

[0213] In some embodiments, the user interface may include, but is not limited to, a display (e.g., an LCD display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., an LED flash), and a keyboard.

[0214] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.

[0215] In some embodiments, the sensors may include, but are not limited to, gyroscope sensors, accelerometers, proximity sensors, ambient light sensors, and positioning units. The positioning unit may also be part of or interact with the network interface 1208 to communicate with components of a positioning network (e.g., Global Positioning System (GPS) satellites).

[0216] The various embodiments disclosed in the present application may be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application may be implemented as a computer program or program code executed on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0217] Program code can be applied to input instructions to perform the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application specific integrated circuit, or a microprocessor.

[0218] Program code can be implemented with high-level programming language or object-oriented programming language to communicate with the processing system. When necessary, program code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.

[0219] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable media. Therefore, a machine-readable medium may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including but not limited to, a floppy disk, an optical disk, an optical disk, a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), a magnetic card or an optical card, an erasable programmable read-only memory (EPROM), a flash memory, an electrically erasable programmable read-only memory (EEPROM), or a tangible machine-readable memory for transmitting information (e.g., a carrier wave, an infrared signal digital signal, etc.) using the Internet in an electrical, optical, acoustic, or other form of propagation signal. Accordingly, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer).

[0220] In the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be required. Instead, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of structural or method features in a particular figure does not mean that such features are required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.

[0221] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation method of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above-mentioned device embodiments.

[0222] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including one" do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0223] Although the present application has been illustrated and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the present application.

Claims

1. A distributed storage method, It is characterized in that Applied to a distributed system, the method comprises: detecting a storage request for first data, where the storage request for the first data is used to store the first data in a first storage node in a first storage partition in the distributed system; Determine, based on the first storage partition corresponding to the first data, a first internal identifier in the first identifier corresponding to the first data, wherein the first internal identifier is used to determine the first storage partition corresponding to the first data when querying the first data; Determine the second internal identifier in the first identifier corresponding to the first data, wherein the first modulo result of the second internal identifier corresponding to the first data relative to the maximum data capacity of a single storage node of the first storage partition is different from the second modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, and the second internal identifier corresponding to the first data is used to query the first data by using the first modulo result of the second internal identifier corresponding to the first data relative to the maximum data capacity of a single storage node of the first storage partition as the logical storage space of the first data in the first storage partition; Associating the logical storage space of the first data in the first storage partition with the logical block address corresponding to the first data; The method for obtaining the maximum storage capacity of a single storage node of the first storage partition includes: Obtaining the maximum amount of data that can be accommodated by each storage node in the first storage partition; Based on the maximum data capacity corresponding to each storage node, the data capacity with the largest value is selected as the maximum data capacity of a single storage node of the first storage partition.

2. The method according to claim 1, It is characterized in that The second internal identifiers corresponding to the data in the first storage node are in increasing order based on the storage order of the data.

3. The method according to any one of claims 1 to 2, It is characterized in that The determining the second internal identifier in the first identifier corresponding to the first data includes: Determine a second internal identifier set corresponding to each storage node in the first storage partition; Determine a second internal identifier set of the first storage partition based on the second internal identifier set corresponding to each storage node; Determine a maximum second internal identifier of the first storage partition based on the second internal identifier set of the first storage partition; Based on the maximum second internal identifier of the first storage partition, the second internal identifier corresponding to the first data is obtained.

4. The method according to claim 3, It is characterized in that The acquiring the second internal identifier corresponding to the first data based on the maximum second internal identifier of the first storage partition includes: Based on the maximum second internal identifier of the first storage partition, generate a first increasing value that increases sequentially relative to the maximum second internal identifier; If the third modulo result of the first incremental value relative to the maximum data capacity of a single storage node of the first storage partition is different from the second modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, the first incremental value is determined as the second internal identifier corresponding to the first data.

5. The method according to claim 4, It is characterized in that The acquiring, based on the maximum second internal identifier of the first storage partition, the second internal identifier corresponding to the first data further includes: If the third modulo result of the first incremental value relative to the maximum data capacity of a single storage node of the first storage partition is the same as the second modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, then a second incremental value that increases sequentially relative to the first incremental value is generated; Based on the fourth modulo result of the second incremental value relative to the maximum data capacity of a single storage node of the first storage partition, which is different from the second modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, the second incremental value is determined as the second internal identifier corresponding to the first data.

6. The method according to claim 1, It is characterized in that The first identification includes a first internal identification and a second internal identification; The byte length of the first identifier is the sum of the byte length of the first internal identifier and the byte length of the second internal identifier.

7. A data indexing method, It is characterized in that Applied to a distributed system, the method comprises: A query request for first data is detected, where the query request includes a first identifier corresponding to the first data; Determine a first storage partition corresponding to the first data based on a first internal identifier in the first identifier; Determine the logical storage space of the first data in the first storage partition based on the first modulo result of the second internal identifier in the first identifier relative to the maximum data capacity of a single storage node of the first storage partition, and determine the logical block address corresponding to the first data based on the logical storage space of the first data in the first storage partition, the logical storage space is located in the first storage node, wherein the first modulo result of the second internal identifier corresponding to the first data relative to the maximum data capacity of a single storage node of the first storage partition is different from the second modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition; Acquire the first data based on the logical block address; The method for obtaining the maximum storage capacity of a single storage node of the first storage partition includes: Obtaining the maximum amount of data that can be accommodated by each storage node in the first storage partition; Based on the maximum data capacity corresponding to each storage node, the data capacity with the largest value is selected as the maximum data capacity of a single storage node of the first storage partition.

8. The method according to claim 7, It is characterized in that The second internal identifiers corresponding to the data in the first storage node are in increasing order based on the storage order of the data.

9. The method according to claim 7, It is characterized in that The method for acquiring the second internal identifier corresponding to the first data includes: Determine a second internal identifier set corresponding to each storage node in the first storage partition; Determine a second internal identifier set of the first storage partition based on the second internal identifier set corresponding to each storage node; Determine a maximum second internal identifier of the first storage partition based on the second internal identifier set of the first storage partition; Based on the maximum second internal identifier of the first storage partition, the second internal identifier corresponding to the first data is obtained.

10. The method according to claim 9, It is characterized in that The acquiring the second internal identifier corresponding to the first data based on the maximum second internal identifier of the first storage partition includes: Based on the maximum second internal identifier of the first storage partition, generate a first increasing value that increases sequentially relative to the maximum second internal identifier; If the third modulo result of the first incremental value relative to the maximum data capacity of a single storage node of the first storage partition is different from the second modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, the first incremental value is determined as the second internal identifier corresponding to the first data.

11. The method according to claim 10, It is characterized in that The acquiring the second internal identifier corresponding to the first data based on the maximum second internal identifier of the first storage partition also includes: If the third modulo result of the first incremental value relative to the maximum data capacity of a single storage node of the first storage partition is the same as the second modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, then a second incremental value that increases sequentially relative to the first incremental value is generated; Based on the fourth modulo result of the second incremental value relative to the maximum data capacity of a single storage node of the first storage partition, which is different from the second modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition, the second incremental value is determined as the second internal identifier corresponding to the first data.

12. The method according to claim 7, It is characterized in that The method further comprises: detecting a deletion request for second data, where the deletion request for the second data is used to delete the second data from a first storage node in a first storage partition in the distributed system; Obtaining a second identifier corresponding to the second data; Determine a first storage partition corresponding to the second data based on the first internal identifier in the second identifier; Determine a logical storage space of the second data in the first storage partition based on the second internal identifier in the second identifier, the logical storage space being located at the first storage node, wherein a fifth modulo result of the second internal identifier corresponding to the second data relative to a maximum data capacity of a single storage node of the first storage partition is different from a sixth modulo result of the second internal identifier corresponding to other data in the first storage node relative to the maximum data capacity of a single storage node of the first storage partition; Based on the logical storage space of the second data in the first storage partition, the association between the logical block address corresponding to the second data and the logical storage space is deleted.

13. An electronic device, It is characterized in that include: A memory and a processor, wherein the memory is used to store instructions executed by one or more processors of the electronic device, and the processor is one of the one or more processors of the electronic device, and is used to execute the distributed storage method described in any one of claims 1-6 or the data indexing method described in any one of claims 7-12.

14. A readable storage medium, It is characterized in that The readable storage medium stores instructions, and when the instructions are executed on an electronic device, the electronic device executes the distributed storage method described in any one of claims 1 to 6 or the data indexing method described in any one of claims 7 to 12.

Citation Information

Patent Citations

  • Distributed metadata management method for distributed file system

    CN111597148A

Cited By

  • Storage system, request processing method, and switch

    US12613636B2