Data storage method and device, equipment, medium and product

By configuring multiple main structures and ScyllaDB clusters in the Ceph object storage system, the hash storage of metadata information is solved, and the data rebalancing problem caused by Ceph expansion is improved, and the storage capacity and system stability are improved.

CN120045602APending Publication Date: 2025-05-27CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510020995.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When existing Ceph object storage systems expand, data rebalancing will lead to data rebalancing, affecting the system's service quality, and failing to realize metadata deployment across data centers, resulting in limited metadata storage capacity.

Method used

By configuring multiple main structures, each main structure includes at least one substructure, at least one set of ScyllaDB clusters are deployed in the substructure, and after receiving user access requests, analyzing and determining the access path, hashing metadata information into different ScyllaDB clusters, realizing linear expansion of object storage metadata.

Benefits of technology

The linear expansion of object storage metadata information is realized, the object storage capacity scale and system stability are improved, data rebalancing is avoided caused by expansion, and the system service quality is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045602A_ABST
    Figure CN120045602A_ABST
Patent Text Reader

Abstract

The invention provides a data storage method and device, electronic equipment, a computer storage medium and a computer program product. The method comprises the steps of obtaining a plurality of main structural bodies configured for different types of metadata information; when an access request of a target user is received, analyzing the access request to obtain user information, bucket information and bucket index information corresponding to the target user; determining a first access path of the user information in the corresponding main structure body and a second access path of the bucket information in the corresponding main structure body; under the condition that the user information and the bucket information are determined to be stored according to the first access path and the second access path, determining a target sub-structure body for storing the bucket index information from a main structure body corresponding to the bucket index information according to a sub-structure body for storing the bucket information; and storing the bucket index information to the ScylaDB cluster in the target substructure body.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a data storage method, apparatus, electronic device, computer storage medium, and computer program product. Background Art

[0002] Ceph is a current mainstream open-source distributed storage system that supports object storage, block storage, and file storage. Among them, the object storage service is implemented based on the interface provided by a reliable, autonomic, distributed object storage system (RADOS). All metadata information related to object storage is stored in RADOS. For example, user secret key information, user bucket information, object size, and upload date information are all stored in RADOS. A set of RADOS consists of monitor (MON) nodes and object storage devices (OSD) nodes. All metadata information related to object storage is stored on the OSD node disks. A set of RADOS clusters generally does not support cross-data center deployment. Moreover, as user data grows, the metadata will also increase, and Ceph needs to be expanded. However, the expansion of Ceph will cause data rebalancing, affecting the service quality of the system. Summary of the Invention

[0003] This application provides a data storage method, apparatus, electronic device, computer storage medium, and computer program product.

[0004] The technical solution of this application is implemented as follows:

[0005] This application provides a data storage method applied to a Ceph object storage system. The method includes:

[0006] Obtain multiple main structures configured for different types of metadata information; each main structure includes at least one sub-structure, and each sub-structure includes at least one set of deployed ScyllaDB clusters;

[0007] When receiving an access request from a target user, parse the access request to obtain user information, bucket information, and bucket index information corresponding to the target user;

[0008] Determine a first access path of the user information in the corresponding main structure and a second access path of the bucket information in the corresponding main structure;

[0009] When it is determined that the user information and the bucket information have been stored according to the first access path and the second access path, a target sub-structure for storing the bucket index information is determined from the main structure corresponding to the bucket index information according to the sub-structure storing the bucket information;

[0010] Store the bucket index information into the ScyllaDB cluster in the target sub-structure.

[0011] This application provides a data storage device, which is applied to a Ceph object storage system. The device includes:

[0012] An acquisition module, configured to acquire a plurality of main structures configured for different types of metadata information; each main structure includes at least one sub-structure, and each sub-structure includes at least one set of deployed ScyllaDB clusters;

[0013] A parsing module, configured to parse the access request when receiving an access request of a target user, so as to obtain user information, bucket information, and bucket index information corresponding to the target user;

[0014] A determination module, configured to determine a first access path of the user information in the corresponding main structure and a second access path of the bucket information in the corresponding main structure; when it is determined that the user information and the bucket information have been stored according to the first access path and the second access path, a target sub-structure for storing the bucket index information is determined from the main structure corresponding to the bucket index information according to the sub-structure storing the bucket information;

[0015] A storage module, configured to store the bucket index information into the ScyllaDB cluster in the target sub-structure.

[0016] This application provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the data storage method provided by the foregoing one or more technical solutions is implemented.

[0017] This application provides a computer storage medium, which stores a computer program; after the computer program is executed, the data storage method provided by the foregoing one or more technical solutions can be implemented.

[0018] This application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the data storage method provided by the foregoing one or more technical solutions is implemented.

[0019] An embodiment of the present application provides a data storage method, apparatus, electronic device, computer storage medium, and computer program product, which are applied to a Ceph object storage system. The method includes: obtaining a plurality of main structures configured for different types of metadata information; each main structure includes at least one sub-structure, and each sub-structure includes at least one set of deployed ScyllaDB clusters; when an access request from a target user is received, parsing the access request to obtain user information, bucket information, and bucket index information corresponding to the target user; determining a first access path of the user information in the corresponding main structure, and a second access path of the bucket information in the corresponding main structure; when it is determined that the user information and the bucket information have been stored according to the first access path and the second access path, determining a target sub-structure for storing the bucket index information from the main structure corresponding to the bucket index information according to the sub-structure storing the bucket information; storing the bucket index information in the ScyllaDB cluster in the target sub-structure.

[0020] It can be seen that in the embodiment of the present application, since the configured multiple main structures are correspondingly configured for different types of metadata information, when an access request from a user is received, all the metadata information corresponding to the access request can be hashed into different main structures according to their respective types; since each sub-structure in each main structure includes several sets of ScyllaDB clusters for storing metadata, different types of metadata information related to the request will be stored in different ScyllaDB clusters, that is, all the metadata originally stored in RADOS can be hashed into different ScyllaDB clusters; in this way, linear expansion of object storage metadata information can be achieved, improving the object storage capacity scale and system stability; further, when subsequent expansion is required due to a large amount of user data, only additional ScyllaDB clusters need to be deployed, and the entire process will not cause data rebalancing in the system, which can improve the service quality of the system. Description of the Drawings

[0021] Figure 1 It is a flowchart of a data storage method provided by an embodiment of the present application;

[0022] Figure 2 It is a schematic structural diagram of a main structure provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic structural diagram of determining a first storage path provided by an embodiment of the present application;

[0024] Figure 4 It is a schematic structural diagram of determining a second storage path provided by an embodiment of the present application;

[0025] Figure 5 A structural schematic diagram for determining the storage path of bucket index information provided by an embodiment of the present application;

[0026] Figure 6 A structural schematic diagram for expanding a sub-structure provided by an embodiment of the present application;

[0027] Figure 7 A flowchart for reading and writing bucket index information provided by an embodiment of the present application;

[0028] Figure 8 A structural schematic diagram of the composition of a data storage device provided by an embodiment of the present application;

[0029] Figure 9 A structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0030] Next, the technical solutions in the present application will be clearly and completely described in conjunction with the accompanying drawings in the present application.

[0031] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the embodiments provided herein are only used to explain the present application and are not used to limit the present application. In addition, the embodiments provided below are partial embodiments for implementing the present application, rather than all embodiments for implementing the present application. Without conflict, the technical solutions described in the present application can be implemented in any combination manner.

[0032] It should be noted that, in the embodiments of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a method or system including a series of elements not only includes the clearly recited elements, but also includes other elements not explicitly listed, or further includes elements inherent to the implementation of the method or system. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of other related elements in the method or system including the element (such as steps in the method or units in the system, for example, the unit may be a part of a circuit, a part of a processor, a part of a program or software, etc.).

[0033] For example, the data storage method provided by the present application includes a series of steps, but the data storage method provided by the present application is not limited to the recited steps. Similarly, the data storage device provided by the present application includes a series of modules, but the data storage device provided by the present application is not limited to including the clearly recited modules, and may further include modules required for obtaining relevant information or processing based on the information.

[0034] In the related art, it is impossible to deploy metadata information related to object storage across regions to achieve disaster tolerance of data, and all metadata information can only be stored in one RADOS cluster, resulting in a limited number of metadata information entries that can be stored; if Ceph is expanded, the expansion process will affect the service quality due to data rebalancing in Ceph.

[0035] In view of the above technical problems, the following embodiments are provided.

[0036] In some embodiments of the present application, the data storage method can be implemented according to a processor in a data storage device. The above processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor.

[0037] Figure 1 The flowchart of a data storage method provided by an embodiment of the present application is shown in Figure 1 As shown, the process may include:

[0038] Step 100: Obtain a plurality of main structures configured for different types of metadata information.

[0039] In an embodiment of the present application, the data storage method can be applied to a Ceph object storage system; the Ceph object storage system is a distributed storage system capable of providing object storage services.

[0040] Exemplarily, the Ceph object storage system includes a plurality of main structures configured for different types of metadata information. Here, different types of metadata information refer to various different types of metadata information related to object storage.

[0041] Exemplarily, metadata information refers to data describing object attributes and is used to indicate functions such as the storage location of an object, historical data, resource search, and file records. In one embodiment, different types of metadata information may include three different types of metadata information: user information, bucket information, and bucket index information. Here, the bucket index information is also referred to as object index information in the bucket.

[0042] In the embodiments of the present application, each main structure may include at least one sub-structure, and each sub-structure may include at least one set of deployed ScyllaDB clusters.

[0043] For ease of understanding, refer to Figure 2 , the main structure can be represented as TableSets, and the sub-structure can be represented as TableSet; each TableSets can include one or more TableSet, and each TableSet can include one or more sets of deployed ScyllaDB clusters. Here, each set of deployed ScyllaDB clusters can be defined as indexset.

[0044] In the embodiments of the present application, each set of deployed ScyllaDB clusters is a NoSQL database that can be deployed across regions. Exemplarily, assume that the first set of ScyllaDB clusters is defined as indexset1, the second set of ScyllaDB clusters is defined as indexset2, and so on. When the Nth set of ScyllaDB clusters is defined as indexsetN, each TableSet can be represented as an ordered combination of indexset{K} (1 <= K <= N).

[0045] It should be noted that the embodiments of the present application can configure different TableSets for different types of metadata information. This is because in the classification of metadata information, the object index information in the bucket is the most. An object storage that supports hundreds of billions of files needs to store hundreds of billions of object index information, while the user information is relatively less, and the number of buckets that each user can create is also limited.

[0046] Exemplarily, when different types of metadata information include user information type, bucket information type, and bucket index information type, the main structure configured for the user information type can be represented as: user_info_TableSets = TableSets1; the main structure configured for the bucket information type can be represented as: bucket_info_TableSets = TableSets2; the main structure configured for the bucket index information type can be represented as: bucket_info_TableSets = TableSets3.

[0047] Further, considering the data volume of different types of metadata information, fewer TableSets can be configured in TableSets1. For example, TableSet11 and TableSet12 can be configured; fewer TableSets can also be configured in TableSets2. For example, TableSet21 and TableSet22 can be configured; more TableSets can be configured in TableSets3. For example, TableSet31, TableSet32, TableSet33, TableSet34,....TableSet3N can be configured. After the above configuration operations, three main structures configured for three different types of metadata information can be obtained, namely TableSets1, TableSets2, and TableSets3.

[0048] Step 101: When receiving a user's access request, parse the access request to obtain the user information, bucket information, and bucket index information corresponding to the target user.

[0049] Here, the access request can include a read request or a write request. Here, the read request is also called a download request, and the write request is also called an upload request or a storage request; since the embodiments of the present application mainly involve the storage of metadata data, the subsequent embodiments will take the access request as an upload request as an example for description.

[0050] In the embodiments of the present application, when the system receives the access request of the target user, it can parse the access request to obtain the user information, bucket information, and bucket index information corresponding to the target user.

[0051] Exemplarily, the access request can be parsed by parsing the request protocol. The purpose of parsing the request protocol is to understand the content and purpose of the access request so that subsequent processing can be carried out in the correct manner. By parsing the access request through the request protocol, the user information, bucket information, and bucket index information corresponding to the target user can be obtained.

[0052] Exemplarily, the user information corresponding to the target user, also called the target user metadata, refers to the information that describes and identifies the target user, and can include but is not limited to information such as user identification (uid), age, address, gender, etc.; the bucket information corresponding to the target user refers to the description of the bucket, and can include but is not limited to information such as bucket name, bucket owner, permissions, storage level, etc.; the bucket index information corresponding to the target user refers to the index information of all objects in the bucket, which can be implemented based on the data structure of the primary key and the sort key. Among them, the primary key represents the object name, such as 1.txt, and the sort key is the bucket shard number.

[0053] Step 102: Determine the first access path of the user information in the corresponding main structure and the second access path of the bucket information in the corresponding main structure.

[0054] Here, the user information is the user information corresponding to the target user; the main structure corresponding to the user information refers to the main structure configured for the user information type above, that is, TableSets1. Similarly, the bucket information is the bucket information corresponding to the target user, and the main structure corresponding to the bucket information refers to the main structure configured for the bucket information type above, that is, TableSets2.

[0055] In some embodiments, determining the first access path of the user information in the corresponding main structure may include: performing a hash operation on the user information to obtain a first hash value; obtaining the total number of sub-structures in the main structure corresponding to the user information; performing a modulo operation on the first hash value and the total number of sub-structures to obtain a first remainder value; determining the sub-structure for storing the user information from the main structure corresponding to the user information according to the first remainder value; the first access path includes the sub-structure for storing the user information.

[0056] Exemplarily, a preset hash algorithm may be used to perform a hash operation on the user information to obtain a first hash value; here, the type of the preset hash algorithm is not specifically limited. For example, it may be an MD5 value algorithm or other types of hash algorithms.

[0057] Exemplarily, assume that the user information includes the uid of the target user and uid = user1. Then, a preset hash algorithm may be used to perform a hash operation on user1 to obtain the corresponding hash value, that is, the first hash value; then, obtain the total number of sub-structures in the main structure corresponding to the user information; next, perform a modulo operation on the first hash value and the total number of sub-structures to obtain the corresponding remainder value, that is, the first remainder value.

[0058] Further, after obtaining the first remainder value, the sub-structure for storing the user information may be determined from the main structure corresponding to the user information; wherein, the first access path includes the sub-structure for storing the user information.

[0059] Here, the determination method of the sub-structure can be: determining the sub-structure with the arrangement serial number being the same as the first remainder value as the sub-structure for storing user information; it should be noted that the arrangement serial numbers of the sub-structures included in the main structure start from 0; exemplarily, assuming that TableSets1 corresponding to user information includes TableSet11 and TableSet12, then the arrangement serial numbers of TableSet11 and TableSet12 are 0 and 1 in sequence. If the first remainder value is 0, it indicates that TableSet11 is the sub-structure for storing user information; if the first remainder value is 1, it indicates that TableSet12 is the sub-structure for storing user information; at this time, the sub-structure for storing user information can be determined.

[0060] Exemplarily, the Ceph object storage system further includes an object storage gateway service; it should be noted that the above reception and parsing of the access request, as well as the hash operation and remainder operation on the parsed content, can be implemented through the object storage gateway service.

[0061] In the embodiment of the present application, after determining the first access path of the user information in the corresponding main structure, the second access path of the bucket information in the corresponding main structure can be further determined; it should be noted that the method for determining the second access path is similar to the method for determining the first access path, and for the sake of avoiding repeated description, it will not be elaborated here.

[0062] Step 103: When it is determined that the user information and the bucket information have been stored according to the first access path and the second access path, determine the target sub-structure for storing the bucket index information from the main structure corresponding to the bucket index information according to the sub-structure for storing the bucket information.

[0063] In the embodiment of the present application, after determining the first access path and the second access path according to the above steps, it can be determined whether the user information and the bucket information have been stored according to the first access path and the second access path; here, the user information and the bucket information refer to the user information and the bucket information corresponding to the target user.

[0064] Exemplarily, whether the user information has been stored refers to whether the user information has been stored in a certain sub-structure of its corresponding main structure (such as TableSets1); whether the bucket information has been stored refers to whether the bucket information has been stored in a certain sub-structure of its corresponding main structure (such as TableSets2).

[0065] Understandably, after determining the first access path, that is, the sub-structure for storing user information, user information can be queried from the ScyllaDB cluster of this sub-structure, and the queried user information is compared with the user information obtained by the above parsing. When it is determined that the two are consistent, it indicates that the system has previously created the target user, and the queried user information was stored in the ScyllaDB cluster of the sub-structure when the target user was created. At this time, it can be determined that the user information has been stored; when it is determined that the two are inconsistent, it indicates that the system has not created the target user, and the current process ends at this time.

[0066] Exemplarily, refer to Figure 3 , assuming that the main structure corresponding to the user information is TableSets1, when the storage gateway service receives an access request from the target user, after parsing, hash operation, and modulo operation, it is determined that the TableSet for storing user information in TableSets1 is TableSet12, and the ScyllaDB cluster in this TableSet12 is defined as indexset2. At this time, user information can be queried from indexset2 and compared with the parsed user information to determine whether the user information has been stored.

[0067] In the embodiment of the present application, when it is determined that the user information has been stored, it can continue to determine whether the bucket information has been stored; this determination process is similar to the determination process corresponding to the above user information, and to avoid repeated description, it will not be elaborated here.

[0068] Exemplarily, refer to Figure 4 , assuming that the main structure corresponding to the bucket information is TableSets2, when the storage gateway service receives an access request from the target user, after parsing, hash operation, and modulo operation, it is determined that the TableSet for storing bucket information in TableSets2 is TableSet23, and the ScyllaDB cluster in this TableSet23 is defined as indexset6. At this time, bucket information can be queried from indexset6 and compared with the parsed bucket information to determine whether the bucket information has been stored.

[0069] Exemplarily, when creating a target user, for example, the uid of the target user is user1. After performing a hash operation on user1, the remainder is taken with respect to the number of TableSets in TableSets1. For example, TableSet12 is obtained, indicating that the user information corresponding to the target user user1 is stored in TableSet12. There is an attribute named bi_TableSets_name in the user information corresponding to the target user user1, and this attribute defaults to the value of bi_TableSets, which is TableSets3. In this way, when the target user creates a bucket, it will default to using the TableSet in TableSets3 (selecting one TableSet according to the weights of each TableSet in TableSets3). After the bucket is successfully created, the selected TableSet will be stored as a part (bi_TableSet) of the bucket information in the ScyllaDB cluster in TableSets2. When creating a bucket, it is calculated according to the weights (TableSetWeight) of each TableSet in TableSets3. For example, the bucket bucket1 is created for the user user1, and the bi_TableSet_name attribute of this bucket is TableSet33. For TableSet33, {indexset11, indexset12} is configured, indicating that the index information of all objects uploaded in this bucket is stored in these two sets of ScyllaDB clusters, indexset11 and indexset12.

[0070] In the embodiments of the present application, when it is determined that the user information and the bucket information have been stored according to the first access path and the second access path, according to the sub-structure of the stored bucket information, the target sub-structure for accessing the bucket index information is determined from the main structure corresponding to the bucket index information.

[0071] Exemplarily, the bucket information includes the sub-structure information of the selected sub-structure for accessing the bucket index information. Therefore, the sub-structure information can be obtained from the sub-structure of the stored bucket information, and then the target sub-structure for accessing the bucket index information can be determined from the main structure corresponding to the bucket index information according to the sub-structure information.

[0072] Exemplarily, refer to Figure 5, assume that the main structure corresponding to the bucket index information is TableSets3; the sub-structure TableSet23 that stores the bucket information records the sub-structure information selected by bucket bucket1; since this sub-structure information is TableSets33; at this time, it can be determined that the target sub-structure for accessing the bucket index information in TableSets3 is TableSet33, where TableSet33 includes indexset11 and indexset12, indicating that the index information of all objects uploaded in bucket bucket1 will be written into these two sets of ScyllaDB clusters, indexset11 and indexset12; that is, the bucket index information can be accessed based on indexset11 and indexset12.

[0073] Step 104: Store the bucket index information into the ScyllaDB cluster in the target sub-structure.

[0074] In the embodiment of the present application, when the access request is a write request, if the target sub-structure for accessing the bucket index information is determined according to the above steps, the bucket index information can be stored into the ScyllaDB cluster in the target sub-structure.

[0075] Exemplarily, when the target sub-structure includes multiple sets of ScyllaDB clusters, storing the bucket index information into the ScyllaDB cluster in the target sub-structure includes: obtaining the total number of ScyllaDB clusters in the target sub-structure; determining the target ScyllaDB cluster from the multiple sets of ScyllaDB clusters included in the target sub-structure according to the total number of ScyllaDB clusters; storing the bucket index information into the target ScyllaDB cluster.

[0076] In the embodiment of the present application, first, the total number of ScyllaDB clusters in the target sub-structure is obtained, that is, the total number of indexsets in the target sub-structure TableSet; then, according to this total data volume, the target ScyllaDB cluster for storing the bucket index information is determined from the multiple sets of ScyllaDB clusters included in the target sub-structure.

[0077] It should be noted that the above process of storing bucket index information can be implemented in two different ways. Regardless of which way is used, it is necessary to perform shard splitting on the bucket corresponding to the bucket information. Exemplarily, assuming that the unique identifier of bucket1 is abcde, it can be split into 1021 shards, including abcde.0, abcde.1, abcde.2,..., abcde.1020, a total of 1021 shards. For example, if the remainder after taking the modulus of 1021 for the hash operation of the uploaded object name 1.txt is 120, the primary key and sorting key of the index information of the object 1.txt in the table of the ScyllaDB cluster are 1.txt and abcde.120. The following is an exemplary description of the two implementation methods.

[0078] In some embodiments, according to the total number of ScyllaDB clusters, determining a target ScyllaDB cluster from multiple sets of ScyllaDB clusters included in the target sub-structure may include: performing a hash operation on the bucket index information to obtain a second hash value; performing a modulus operation on the first hash value and the total number of ScyllaDB clusters to obtain a second remainder value; and determining the target ScyllaDB cluster from multiple sets of ScyllaDB clusters included in the target sub-structure according to the second remainder value.

[0079] Exemplarily, after obtaining the total number m of ScyllaDB clusters in the target sub-structure, a hash operation can be performed on the primary key of the bucket index information, that is, the object name {object_name}, to obtain a second hash value, and then the modulus operation is performed on the second hash value and the total number m. The obtained second remainder value is the serial number indexset.x of the indexset to which the bucket index information is to be written in the TableSet of the target sub-structure. This implementation method can be called Method 1, and Method 1 can be represented by formula (1):

[0080] indexset.x = hash({object_name}) % m (1)

[0081] In some embodiments, according to the total number of ScyllaDB clusters, determining a target ScyllaDB cluster from multiple sets of ScyllaDB clusters included in the target sub-structure may include: performing a hash operation on the bucket index information to obtain a third hash value; obtaining a preset number of bucket shards, performing a modulus operation on the third hash value and the preset number of bucket shards to obtain a third remainder value; performing a modulus operation on the third remainder value and the total number of ScyllaDB clusters to obtain a fourth remainder value; and determining the target ScyllaDB cluster from multiple sets of ScyllaDB clusters included in the target sub-structure according to the fourth remainder value.

[0082] Exemplarily, similar to the above-mentioned Method 1, after obtaining the total number m of ScyllaDB clusters in the target sub-structure, the primary key of the bucket index information, i.e., the object name {object_name}, can be hash-operated to obtain a third hash value; here, the third hash value is the same as the above-mentioned second hash value, just for distinction; then, the remainder operation is performed on the third hash value and the preset number of bucket shards to obtain a third remainder value; assuming that the preset number of bucket shards is 1021, the remainder operation is performed on 1021, and then the remainder operation is performed on the third remainder value and the total number m of ScyllaDB clusters to obtain a fourth remainder value; the obtained fourth remainder value is the sequence number indexset.x of the indexset to be written into the bucket index information in the target sub-structure TableSet. This implementation method can be called Method 2, and Method 2 can be represented by Formula (2) and Formula (3):

[0083] bucket_shard = hash({object_name}) % 1021 (2)

[0084] indexset.x = bucket_shard % m (3)

[0085] Exemplarily, in the case of a large amount of data, the bucket index information of all indexsets in the target sub-structure TableSet in Method 1 will contain the sorted keys from abcde.0 to abcde.1020. In Method 2, the sorted keys of different indexsets in the target sub-structure TableSet will only include different bucket shards. For example, indexset1 only contains abcde.{0,1,2,3,4...333}, indexset2 only contains abcde.{334,335,336....666}, and indexset3 only contains abcde.{667,668,...1020}.

[0086] Exemplarily, in the case where the access request is a read request, performing a read operation on the bucket index information of the ScyllaDB clusters in the target sub-structure means reading the bucket index information from the ScyllaDB clusters in the target sub-structure; it should be noted that in the case where the target sub-structure includes multiple sets of ScyllaDB clusters, the process of performing the read operation on the bucket index information is similar to the process of performing the write operation on the bucket index information described above. To avoid repeated description, it will not be elaborated here.

[0087] It can be seen that in the embodiments of the present application, since the multiple main structures are configured for different types of metadata information, when receiving a user access request, all the metadata information corresponding to the access request can be hashed into different main structures according to different types; since each sub-structure in each main structure includes several sets of ScyllaDB clusters for storing metadata, different types of metadata information related to the request will be stored in different ScyllaDB clusters, that is, all the metadata originally stored in RADOS can be hashed into different ScyllaDB clusters; in this way, linear expansion of object storage metadata information can be achieved, improving the object storage capacity scale and system stability; further, when subsequent expansion is required due to a large amount of user data, only additional ScyllaDB clusters need to be deployed, and the whole process will not cause data rebalancing in the system, which can improve the service quality of the system.

[0088] In some embodiments, each sub-structure further includes the weight of each set of ScyllaDB clusters deployed, and the above method may further include: obtaining the performance parameters of multiple sets of ScyllaDB clusters to be deployed; determining the weights of each set of ScyllaDB clusters included in each sub-structure according to the performance parameters.

[0089] Exemplarily, the performance parameter refers to a parameter that can characterize the performance of the ScyllaDB cluster, which may include but is not limited to storage capacity, number of CPU cores, and network card bandwidth, etc.

[0090] Exemplarily, when deploying ScyllaDB clusters, due to different physical hardware specifications, for example, the network card bandwidth of some physical nodes is 50 Gbp / s, the network card bandwidth of some physical nodes is 25 Gbp / s, the CPU cores of some physical nodes are 96 cores, and the CPU cores of some physical nodes are 32 cores. Usually, physical nodes with the same configuration are grouped into a set of indexset, that is, a ScyllaDB cluster, to avoid the bucket effect. The result is that the performance parameters of indexset1, indexset2, and indexset3 may be different. To ensure that an indexset with better performance can be used, the weights of each set of ScyllaDB clusters included in each sub-structure can be determined according to the performance parameters of multiple sets of ScyllaDB clusters to be deployed.

[0091] Exemplarily, taking storage capacity as an example, assume that indexset1 can accommodate 10 billion object index information, indexset2 can accommodate 5 billion, and indexset3 can accommodate 5 billion. Then, the sub-structure TableSet including only these 3 indexsets can be correspondingly configured as {indexset1, indexset1, indexset2, indexset3}, that is, the weight of indexset1 is set to twice that of indexset2 and indexset3. In this way, the proportion of indexset1 in the bucket shard of this sub-structure TableSet (way 2) is twice that of indexset2 and indexset3. On this basis, if we further consider that the number of CPU cores of indexset1 is 3 times that of indexset2 and indexset3, then the sub-structure TableSet can be correspondingly configured as {indexset1, indexset1, indexset1, indexset1, indexset1, indexset1, indexset2, indexset3}, that is, the weight of indexset1 is set to six times that of indexset2 and indexset3. In this way, the proportion of indexset1 in the bucket shard of this sub-structure TableSet (way 2) is six times that of indexset2 and indexset3. For the sake of simplicity in description, it can be expressed as {indexset1:6,indexset2:1,indexset3:1}.

[0092] Exemplarily, when expanding the bucket index information, the weight of the TableSetWeight corresponding to the TableSet with full storage capacity can be set to 0. In this way, these TableSets with full capacity will not be used when creating new buckets.

[0093] In some embodiments, the above method may further include: detecting the total storage capacity of each set of ScyllaDB clusters included in each sub-structure to obtain a detection result; and when it is determined according to the detection result that the total storage capacity does not meet the storage requirements, expanding the number of ScyllaDB clusters included in each sub-structure.

[0094] In the embodiments of the present application, the system detects the total storage capacity of each set of ScyllaDB clusters included in each sub-structure to obtain a detection result; exemplarily, the detection result can be compared with a preset capacity threshold. If it is determined according to the comparison result that the detection result is greater than or equal to the preset capacity threshold, it indicates that the total storage capacity does not meet the storage requirements. At this time, the number of ScyllaDB clusters included in each sub-structure can be expanded; conversely, if it is determined according to the comparison result that the detection result is less than the preset capacity threshold, it indicates that the total storage capacity meets the storage requirements, and no expansion is required at this time; wherein, the expansion operation increases the storage space by increasing the number of ScyllaDB clusters.

[0095] Exemplarily, refer to Figure 6 , in the case where the sub-structure bi_TableSet includes indexset1, indexset2, and indexset3, that is, bi_TableSet = {indexset1, indexset2, indexset3}, if it is determined that the total storage capacity of indexset1, indexset2, and indexset3 does not meet the storage requirements, the sub-structure bi_TableSet can be expanded to obtain an expanded sub-structure, which is correspondingly represented as bi_TableSet = {indexset1, indexset2, indexset3, indexset4, indexset5, indexset6}.

[0096] Under normal circumstances, it is recommended to use a new bucket for expansion. Refer to Figure 6, if the existing bucket capacity is almost full, expand the existing sub-structure bi_TableSet. For example, expand bi_TableSet = {indexset1, indexset2, indexset3} to bi_TableSet = {indexset1, indexset2, indexset3, indexset4, indexset5, indexset6}. In this case, the historical information of the sub-structure TableSet is stored in the bucket information. At this time, the following download logic can be modified. For the uploaded object 2.txt, calculate the bucket shard abcde.150. The mod1 obtained by taking the remainder of 150 divided by 6 is the serial number of the indexset to be newly written. At the same time, take the remainder of 150 divided by the number of indexsets included in the previous TableSet. For example, if the previous TableSet is {indexset1, indexset2, indexset3}, mod2 can be obtained by taking the remainder of 150 divided by 3. If the mod1-th and mod2-th indexsets are not the same indexset, for example, the mod1-th is indexset1 and the mod2-th is indexset3, then the indexset corresponding to the historical mod2 needs to be deleted. For the downloaded object, if the bucket information has the historical TableSet, calculate multiple mods for the historical TableSet, and download the object from the indexsets corresponding to the multiple mods. Only when there is no corresponding object in all the indexsets corresponding to the mods does it indicate that this object does not exist.

[0097] Figure 7 A flowchart of reading and writing bucket index information provided by an embodiment of the present application is shown as Figure 7 shown. The process may include: when the object storage gateway service of the system receives an access request from a target user, it can query the user information corresponding to the target user from the main structure TableSets1 corresponding to the user information, and query the bucket information corresponding to the target user from the main structure TableSets2 corresponding to the bucket information. When it is determined that the user information and bucket information corresponding to the target user have been stored, according to the sub-structure TableSet5 storing the bucket information, determine the target sub-structure for reading and writing the bucket index information from the main structure TableSets3 corresponding to the bucket index information, and perform the reading and writing operation of the bucket index information on the ScyllaDB cluster in the target sub-structure.

[0098] Figure 8 A schematic diagram of the composition structure of the data storage device according to an embodiment of the present application is shown as Figure 8As shown, the device includes: an acquisition module 300, a parsing module 301, a determination module 302, and a storage module 303, where:

[0099] The acquisition module 300 is used to acquire multiple main structures configured for different types of metadata information; each main structure includes at least one sub-structure, and each sub-structure includes at least one set of deployed ScyllaDB clusters;

[0100] The parsing module 301 is used to parse the access request when receiving the access request from the target user, and obtain the user information, bucket information, and bucket index information corresponding to the target user;

[0101] The determination module 302 is used to determine the first access path of the user information in the corresponding main structure, and the second access path of the bucket information in the corresponding main structure; when it is determined that the user information and the bucket information have been stored according to the first access path and the second access path, according to the sub-structure storing the bucket information, determine the target sub-structure for storing the bucket index information from the main structure corresponding to the bucket index information;

[0102] The storage module 303 is used to store the bucket index information into the ScyllaDB cluster in the target sub-structure.

[0103] In some embodiments, the determination module 302 is further used to:

[0104] Perform a hash operation on the user information to obtain a first hash value;

[0105] Obtain the total number of sub-structures in the main structure corresponding to the user information;

[0106] Perform a modulo operation on the first hash value and the total number of sub-structures to obtain a first remainder value;

[0107] Determine the sub-structure for storing the user information from the main structure corresponding to the user information according to the first remainder value; the first access path includes the sub-structure for storing the user information.

[0108] In some embodiments, when the target sub-structure includes multiple sets of ScyllaDB clusters, the storage module 303 is further used to:

[0109] Obtain the total number of ScyllaDB clusters in the target sub-structure;

[0110] Determine the target ScyllaDB cluster from the multiple sets of ScyllaDB clusters included in the target sub-structure according to the total number of ScyllaDB clusters;

[0111] Store the bucket index information into the target ScyllaDB cluster.

[0112] In some embodiments, the storage module 303 is further configured to:

[0113] Perform a hashing operation on the bucket index information to obtain a second hash value;

[0114] Perform a modulo operation on the second hash value and the total number of ScyllaDB clusters to obtain a second remainder value;

[0115] Determine a target ScyllaDB cluster from multiple sets of ScyllaDB clusters included in the target sub-structure according to the second remainder value.

[0116] In some embodiments, the storage module 303 is further configured to:

[0117] Perform a hashing operation on the bucket index information to obtain a third hash value;

[0118] Obtain a preset number of bucket shards, and perform a modulo operation on the third hash value and the preset number of bucket shards to obtain a third remainder value;

[0119] Perform a modulo operation on the third remainder value and the total number of ScyllaDB clusters to obtain a fourth remainder value;

[0120] Determine a target ScyllaDB cluster from multiple sets of ScyllaDB clusters included in the target sub-structure according to the fourth remainder value.

[0121] In some embodiments, each sub-structure further includes the weights of each set of ScyllaDB clusters deployed, and the acquisition module 300 is further configured to:

[0122] Obtain the performance parameters of multiple sets of ScyllaDB clusters to be deployed;

[0123] Determine the weights of each set of ScyllaDB clusters included in each sub-structure according to the performance parameters.

[0124] In some embodiments, the above device further includes a detection module, and the detection module is configured to:

[0125] Detect the total storage capacity of each set of ScyllaDB clusters included in each sub-structure to obtain a detection result;

[0126] In the case where it is determined according to the detection result that the total storage capacity does not meet the storage requirements, expand the number of ScyllaDB clusters included in each sub-structure.

[0127] In practical applications, the above-mentioned acquisition module 300, parsing module 301, determination module 302, storage module 303, and detection module can all be implemented by a processor located in an electronic device, and the processor can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor.

[0128] In addition, each functional module in this embodiment can be integrated in a processing unit, can exist separately physically for each unit, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional module.

[0129] If the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method of this embodiment. The foregoing storage medium includes: USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc and other media that can store program codes.

[0130] Specifically, the computer program instructions corresponding to a data storage method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the computer program instructions corresponding to a data storage method in the storage medium are read or executed by an electronic device, any one of the data storage methods of the foregoing embodiments is implemented.

[0131] Based on the same technical concept as the foregoing embodiments, refer to Figure 9 , which shows an electronic device 400 provided by an embodiment of the present application, and can include: a memory 401 and a processor 402; wherein,

[0132] The memory 401 is used to store computer programs and data;

[0133] The processor 402 is used to execute the computer program stored in the memory to implement any one of the data storage methods of the foregoing embodiments.

[0134] In practical applications, the above-mentioned memory 401 may be a volatile memory, such as RAM; or a non-volatile memory, such as ROM, flash memory, Hard Disk Drive (HDD), or Solid-State Drive (SSD); or a combination of the above types of memories, and provide instructions and data to the processor 402.

[0135] The above-mentioned processor 402 may be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor. It can be understood that for different data storage devices, the electronic devices used to implement the above-mentioned processor functions may also be others, and the embodiments of the present application do not make specific limitations.

[0136] In some embodiments, the embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor implements any one of the data storage methods in the foregoing embodiments.

[0137] In some embodiments, the functions or modules included in the device provided by the embodiments of the present application can be used to execute the methods described in the foregoing method embodiments. Its specific implementation can refer to the description of the foregoing method embodiments. For the sake of brevity, it will not be repeated here.

[0138] The descriptions of the foregoing embodiments tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated here.

[0139] The methods disclosed in the method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.

[0140] The features disclosed in the product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.

[0141] The features disclosed in the method or device embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0142] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.

[0143] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0145] The above is only the preferred embodiment of the present application, and is not used to limit the protection scope of the present application.

Claims

1. A data storage method, characterized in that: Applied to a Ceph object storage system, the method comprises: Obtain multiple main structures configured for different types of metadata information; each main structure includes at least one substructure, and each substructure includes at least one deployed ScyllaDB cluster; When receiving an access request from a target user, the access request is parsed to obtain user information, bucket information, and bucket index information corresponding to the target user; Determine a first access path of the user information in the corresponding main structure, and a second access path of the bucket information in the corresponding main structure; When it is determined according to the first access path and the second access path that the user information and the bucket information have been stored, determining a target substructure for storing the bucket index information from a main structure corresponding to the bucket index information according to a substructure storing the bucket information; Store the bucket index information to the ScyllaDB cluster in the target substructure.

2. The method according to claim 1, characterized in that The determining of a first access path of the user information in the corresponding main structure includes: Performing a hash operation on the user information to obtain a first hash value; Obtain the total number of substructures in the main structure corresponding to the user information; Performing a modulo operation on the first hash value and the total number of the substructures to obtain a first modulo value; According to the first remainder value, a substructure for storing the user information is determined from a main structure corresponding to the user information; and the first access path includes the substructure for storing the user information.

3. The method according to claim 1, characterized in that In the case where the target substructure includes multiple ScyllaDB clusters, storing the bucket index information in the ScyllaDB cluster in the target substructure includes: Get the total number of ScyllaDB clusters in the target substructure; Determine a target ScyllaDB cluster from the multiple ScyllaDB clusters included in the target substructure according to the total number of the ScyllaDB clusters; Store the bucket index information to the target ScyllaDB cluster.

4. The method according to claim 3, characterized in that Determining a target ScyllaDB cluster from the multiple ScyllaDB clusters included in the target substructure according to the total number of the ScyllaDB clusters includes: Performing a hash operation on the bucket index information to obtain a second hash value; Performing a modulo operation on the second hash value and the total number of the ScyllaDB clusters to obtain a second modulo value; Based on the second remainder value, determine the target ScyllaDB cluster from the multiple ScyllaDB clusters included in the target substructure.

5. The method according to claim 3, characterized in that: Determining a target ScyllaDB cluster from the multiple ScyllaDB clusters included in the target substructure according to the total number of the ScyllaDB clusters includes: Performing a hash operation on the bucket index information to obtain a third hash value; Obtain a preset number of bucket shards, and perform a modulo operation on the third hash value and the preset number of bucket shards to obtain a third modulo value; Performing a modulo operation on the third remainder value and the total number of the ScyllaDB clusters to obtain a fourth remainder value; According to the fourth remainder value, the target ScyllaDB cluster is determined from the multiple ScyllaDB clusters included in the target substructure.

6. The method according to claim 1, characterized in that Each substructure also includes the weight of each deployed ScyllaDB cluster, and the method further includes: Get the performance parameters of multiple ScyllaDB clusters that need to be deployed; Based on the performance parameters, determine the weights of the ScyllaDB clusters included in each substructure.

7. The method according to claim 1, characterized in that The method further comprises: Detect the total storage capacity of each ScyllaDB cluster included in each substructure to obtain a detection result; When it is determined according to the detection results that the total storage capacity does not meet the storage requirements, the number of ScyllaDB clusters included in each substructure is expanded.

8. A data storage device, characterized in that: Applied to a Ceph object storage system, the device comprises: An acquisition module is used to acquire multiple main structures configured for different types of metadata information; each main structure includes at least one substructure, and each substructure includes at least one deployed ScyllaDB cluster; A parsing module, configured to parse an access request from a target user upon receiving the access request, and obtain user information, bucket information, and bucket index information corresponding to the target user; a determination module, configured to determine a first access path of the user information in a corresponding main structure, and a second access path of the bucket information in a corresponding main structure; and when it is determined that the user information and the bucket information have been stored according to the first access path and the second access path, determining a target substructure for storing the bucket index information from the main structure corresponding to the bucket index information according to the substructure storing the bucket information; A storage module is used to store the bucket index information in the ScyllaDB cluster in the target substructure.

9. An electronic device, characterized in that: The device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the method according to any one of claims 1 to 7 is implemented when the processor executes the program.

10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that The computer program implements the method according to any one of claims 1 to 7 when executed by a processor.