A data migration method and unified coordination system

By establishing the mapping relationship between the hash bucket and the storage actual bucket in a unified coordination system, as well as the mapping relationship between the storage actual bucket and the cluster, the problem of low data migration efficiency in multi-heterogeneous clusters with hyperscale storage is solved, and efficient data migration across clusters is achieved.

CN115460230BActive Publication Date: 2025-05-16UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211084130.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-06
Publication Date
2025-05-16
Estimated Expiration
2042-09-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively coordinate and migrate data in multi-heterogeneous clusters stored at super-large scale, resulting in low data migration efficiency and overall system slowdown.

Method used

Efficient data migration across clusters is achieved by establishing a mapping relationship between the hash bucket and the storage actual bucket in a unified coordination system, and storing the mapping relationship between the actual bucket and the cluster. The specific steps include obtaining migration instructions, creating a new storage actual bucket and hash bucket, updating the mapping table, and migrating data based on the mapping table.

Benefits of technology

Efficient data migration across clusters is realized, avoiding hashing algorithm calculations for each data, improving migration efficiency, and reducing the overall risk of slowing down the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115460230B_ABST
    Figure CN115460230B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a data migration method and a unified coordination system. In the data migration method, a hash bucket is established as a theoretical storage location, and an actual storage bucket is established as an actual storage location. By configuring the mapping relationship between the hash bucket and the actual storage bucket, and the mapping relationship between the actual storage bucket and the cluster, during data migration, it is only necessary to determine the first data to be migrated in the target hash bucket, and there is no need to calculate each data through a hash algorithm. Moreover, the cross-cluster migration of the first data to be migrated can be achieved directly through the above two mapping relationships, thereby realizing efficient data migration across clusters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cloud storage, and more specifically, to a data migration method and a unified coordination system. Background Art

[0002] A distributed storage cluster usually uses multiple servers to form a storage cluster. The storage cluster is usually divided into storage nodes, metadata nodes, entry service nodes, etc. (which may vary depending on the type of storage cluster). The existing technology has designed various data balancing fault migration capabilities for disks and storage nodes in the cluster. Most of them are based on the hash distribution of data to redistribute the data compass position and copy and move it within the cluster, that is, for each data, a hash algorithm is used to determine the disk or storage node in the data cluster.

[0003] At present, most of the existing technical solutions are aimed at the coordination, data balancing and migration within a single cluster (PB level), and cannot cope with the data balancing and coordination between large-scale clusters composed of multiple clusters, especially multiple heterogeneous clusters. In addition, a hash algorithm must be used to determine the disk or storage node to be migrated in the current cluster for each data. A large amount of data will be moved during balancing or fault migration, which takes a long time and will cause the overall system to slow down.

[0004] It can be seen that the existing technology has the defects of being unable to cope with the coordination, data balancing and migration of multi-heterogeneous cluster structures of ultra-large-scale storage, and the problem of low data migration efficiency. Summary of the invention

[0005] The purpose of this application is to provide a data migration method and a unified coordination system, based on executing the data migration method to achieve efficient data migration across clusters.

[0006] To achieve the above object, the present application provides a data migration method, which is performed by a unified coordination system and includes:

[0007] Obtain a first migration instruction; the first migration instruction includes a first target migration cluster, a target storage actual bucket, and a target hash bucket corresponding to the target storage actual bucket; the first target migration cluster is any cluster among the multiple clusters included in the unified coordination system except the current cluster corresponding to the target storage actual bucket;

[0008] Creating a first storage actual bucket in the first target migration cluster and a first hash bucket corresponding to the first storage actual bucket;

[0009] Adding the correspondence between the first hash bucket and the first actual storage bucket to the pre-stored first mapping table to obtain a second mapping table; adding the correspondence between the first actual storage bucket and the first target migration cluster to the pre-stored third mapping table to obtain a fourth mapping table;

[0010] Determine the first data to be migrated in the target hash bucket, and migrate the first data to be migrated in the target hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table and the fourth mapping table.

[0011] Optionally, the method further includes: obtaining a migration status of the first data to be migrated in the target hash bucket; then, migrating the first data to be migrated in the target hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table and the fourth mapping table, including:

[0012] Setting the migration status of the first data to be migrated in the target hash bucket to being migrated;

[0013] Migrating the first data to be migrated in the target hash bucket from the target hash bucket to the first hash bucket to obtain the data to be migrated in the first hash bucket;

[0014] A correspondence between the first hash bucket and the target storage actual bucket is added to the row where the correspondence between the first hash bucket and the first storage actual bucket is located in the second mapping table, so as to generate a second mapping table in a migration state;

[0015] A correspondence between the first actual storage bucket and the current cluster is added to the row of the correspondence between the first actual storage bucket and the first target migration cluster in the fourth mapping table, so as to generate a fourth mapping table in the migration state;

[0016] Migrating the storage actual bucket corresponding to the data to be migrated in the first hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table in the migration state and the fourth mapping table in the migration state;

[0017] The corresponding relationship between the first hash bucket and the target storage actual bucket in the second mapping table in the migration state is deleted, and the corresponding relationship between the first storage actual bucket and the current cluster in the fourth mapping table in the migration state is deleted.

[0018] Optionally, the method further comprises:

[0019] Obtain a second migration instruction; the second migration instruction includes the target storage actual bucket and the target hash bucket;

[0020] Creating a second storage actual bucket in the current cluster according to the second migration instruction and a second hash bucket corresponding to the second storage actual bucket;

[0021] Adding the correspondence between the second hash bucket and the second storage actual bucket to the first mapping table to obtain a fifth mapping table;

[0022] Determine the second data to be migrated in the target hash bucket, and migrate the storage actual bucket corresponding to the second data to be migrated in the second target hash bucket from the second target storage actual bucket to the second storage actual bucket according to the fifth mapping table.

[0023] Optionally, the method further comprises:

[0024] Acquire a third migration instruction, where the third migration instruction includes the second target migration cluster, where the second target migration cluster is any cluster other than the current cluster among the multiple clusters included in the unified coordination system;

[0025] Obtain the third mapping table, wherein the third mapping table includes a correspondence between the target storage actual bucket and the current cluster;

[0026] According to the third migration instruction, the correspondence between the target storage actual bucket and the current cluster in the third mapping table is updated to the correspondence between the target storage actual bucket and the second target migration cluster, so as to obtain a sixth mapping table;

[0027] The target storage actual bucket is migrated from the current cluster to the second target migration cluster according to the sixth mapping table.

[0028] Optionally, the method further comprises:

[0029] Obtaining a cluster expansion instruction, where the cluster expansion instruction is generated according to a received user request to expand the cluster, or is generated when it is determined that the storage capacity of the current cluster reaches a first preset threshold;

[0030] Adding a first newly added cluster in the unified coordination system according to the cluster expansion instruction, acquiring cluster information of the first newly added cluster, storing the cluster information of the first newly added cluster in a database in the unified coordination system, and sending the cluster information of the first newly added cluster to a metadata server in the unified coordination system;

[0031] Creating a third storage actual bucket in the first newly added cluster and a third hash bucket corresponding to the third storage actual bucket;

[0032] Adding the correspondence between the third hash bucket and the third actual storage bucket to the first mapping table to obtain a seventh mapping table; adding the correspondence between the third actual storage bucket and the first newly added cluster to the third mapping table to obtain an eighth mapping table;

[0033] The hash bucket data to be migrated is determined, and according to the seventh mapping table and the eighth mapping table, the hash bucket data to be migrated is migrated from the storage actual bucket corresponding to the hash bucket data to be migrated to the third storage actual bucket.

[0034] Optionally, if the cluster expansion instruction is generated according to a received user cluster expansion request, determining the hash bucket data to be migrated, and migrating the hash bucket data to be migrated from the storage actual bucket corresponding to the hash bucket data to be migrated to the third storage actual bucket according to the seventh mapping table and the eighth mapping table, includes:

[0035] Acquire a weight of each cluster in the plurality of clusters included in the unified coordination system, where the weight of each cluster in the unified coordination system is determined according to a version type of each cluster;

[0036] Determining a cluster of data to be migrated according to the weight of each cluster in the unified coordination system;

[0037] According to the remaining storage space of each of the multiple storage actual buckets included in the cluster of the data to be migrated;

[0038] Determine the actual storage bucket for the data to be migrated according to the remaining storage space of each actual storage bucket in the cluster to be migrated;

[0039] According to the first pre-stored mapping table, a hash bucket corresponding to the actual storage bucket of the data to be migrated is acquired to obtain the hash bucket of the data to be migrated;

[0040] Determine the hash bucket data to be migrated from the hash bucket to be migrated according to a preset algorithm;

[0041] According to the seventh mapping table and the eighth mapping table, the hash bucket data to be migrated is migrated from the storage actual bucket corresponding to the hash bucket data to be migrated to the third storage actual bucket.

[0042] Optionally, the method further comprises:

[0043] Get the number of files in the target hash bucket;

[0044] Determine whether the number of files in the target hash bucket reaches a second preset threshold;

[0045] If so, a fourth hash bucket corresponding to the target storage actual bucket is newly created, and a corresponding relationship between the fourth hash bucket and the target storage actual bucket is added to the first mapping table to obtain a ninth mapping table.

[0046] Optionally, before receiving the first migration instruction, the method further includes:

[0047] Accepting a user creation request sent by a user; the user creation request includes user information;

[0048] Creating account information corresponding to the user information in the database of the unified coordination system according to the user information in the user creation request, storing the user information in the database, and generating a user information creation success instruction;

[0049] Generate a container creation request according to the user information creation success instruction; the container creation request includes user container information;

[0050] storing the user container information in a database of the unified coordination system, and writing the user container information into a newly created user bucket according to the container creation request to generate a first user bucket;

[0051] Determine the number of actual buckets stored in the first user bucket according to the container creation request;

[0052] Determine at least one cluster among the multiple clusters included in the unified coordination system for creating a plurality of storage actual buckets included in the first user bucket; create a new storage actual bucket in at least one cluster among the multiple storage actual buckets included in the first user bucket;

[0053] Determine the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket according to the container creation request; the bit width is used to determine the number of files that can be stored in the hash bucket;

[0054] The first mapping table is created according to the number of actual storage buckets and the number of hash buckets in the first user bucket; and a hash bucket corresponding to each actual storage bucket in the multiple actual storage buckets included in the first user bucket is newly created according to the first mapping table; the third mapping table is created according to the number of actual storage buckets in the first user bucket and the cluster where each actual storage bucket in the multiple actual storage buckets included in the first user bucket is located;

[0055] The first mapping table and the third mapping table are stored in the unified storage access layer OSS of the unified coordination system.

[0056] Optionally, determining the number of actual buckets stored in the first user bucket according to the container creation request includes:

[0057] Determining whether the container creation request includes an estimated number;

[0058] If the container creation request includes the estimated number, determining the number of actual buckets stored in the first user bucket according to the estimated number;

[0059] If the container creation request does not include the estimated number, determining the number of actual buckets stored in the first user bucket according to a default estimated number;

[0060] The determining, according to the container creation request, the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket includes:

[0061] Determining whether the container creation request includes an estimated number;

[0062] If the container creation request includes the estimated number, determining the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket according to the estimated number;

[0063] If the container creation request does not include the estimated number, the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket are determined according to a default estimated number.

[0064] The present application also provides a unified coordination system, the unified coordination system comprising:

[0065] A management scheduling interface module is used to obtain a first migration instruction; the first migration instruction includes a first target migration cluster, a target storage actual bucket, and a target hash bucket corresponding to the target storage actual bucket; the first target migration cluster is any cluster among the multiple clusters included in the unified coordination system except the current cluster corresponding to the target storage actual bucket;

[0066] A scheduling migration module, configured to create a first storage actual bucket and a first hash bucket corresponding to the first storage actual bucket in the first target migration cluster;

[0067] A unified metadata center module, used to add the correspondence between the first hash bucket and the first storage actual bucket in a pre-stored first mapping table to obtain a second mapping table; add the correspondence between the first storage actual bucket and the first target migration cluster in a pre-stored third mapping table to obtain a fourth mapping table;

[0068] The scheduling migration module is further used to determine the first data to be migrated in the target hash bucket, and migrate the first data to be migrated in the target hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table and the fourth mapping table.

[0069] The embodiment of the present application provides a data migration method and a unified coordination system. In the data migration method, based on obtaining a first migration instruction; creating a first storage actual bucket and a first hash bucket corresponding to the first storage actual bucket in the first target migration cluster; adding the correspondence between the first hash bucket and the first storage actual bucket in a pre-stored first mapping table to obtain a second mapping table; adding the correspondence between the first storage actual bucket and the first target migration cluster in a pre-stored third mapping table to obtain a fourth mapping table; determining the first data to be migrated from the target hash bucket, and migrating the first data to be migrated in the target hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table and the fourth mapping table; it can be seen that in the embodiment of the present application, based on establishing a hash bucket as a theoretical storage location and an actual bucket as an actual storage location, by configuring a mapping relationship between a hash bucket and a storage actual bucket, and a mapping relationship between a storage actual bucket and a cluster, during data migration, it is only necessary to determine the first data to be migrated in the target hash bucket, and it is not necessary to calculate each data through a hash algorithm, and the cross-cluster migration of the first data to be migrated can be achieved directly through the above two mapping relationships, thereby achieving efficient data migration across clusters. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0071] Figure 1 A flowchart of a data migration method provided in an embodiment of the present application;

[0072] Figure 2 A schematic diagram of the relationship between user accounts and user containers provided in an embodiment of the present application;

[0073] Figure 3 A flowchart of creating a user container provided in an embodiment of the present application;

[0074] Figure 4 A flowchart for creating a user account provided in an embodiment of the present application;

[0075] Figure 5 Another flowchart for creating a user container provided in an embodiment of the present application;

[0076] Figure 6 The mapping representation scheme under the migration state provided by the embodiment of the present application;

[0077] Figure 7A flowchart of another data migration method provided in an embodiment of the present application;

[0078] Figure 8 A flowchart of metadata update provided in an embodiment of the present application;

[0079] Fig. 9 A flowchart of another data migration method provided in an embodiment of the present application;

[0080] Fig.10 A flowchart of another data migration method provided in an embodiment of the present application;

[0081] Fig.11 A flowchart of another data migration method provided in an embodiment of the present application;

[0082] Fig.12 A flowchart of another data migration method provided in an embodiment of the present application;

[0083] Fig.13 A schematic diagram of a unified coordination system provided in an embodiment of the present application;

[0084] Fig.14 A framework diagram of the unified coordination system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0085] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0086] At present, most of the existing technical solutions are aimed at the coordination, data balancing and migration within a single cluster (PB level), and cannot cope with the data balancing and coordination between large-scale clusters composed of multiple clusters, especially multiple heterogeneous clusters. In addition, it is necessary to determine the disk or storage node to be migrated in the current cluster through a hash algorithm for each data. During balancing or fault migration, a large amount of data will be moved, which takes a long time and will cause the overall system to slow down. It can be seen that the existing technology has the defect of being unable to cope with the coordination, data balancing and migration of multiple heterogeneous cluster structures for ultra-large-scale storage, and the problem of low data migration efficiency.

[0087] An embodiment of the present application provides a data migration method and a unified coordination system. In the data migration method, based on establishing a hash bucket as a theoretical storage location and an actual bucket as an actual storage location, by configuring a mapping relationship between the hash bucket and the actual bucket, and a mapping relationship between the actual bucket and the cluster, during data migration, it is only necessary to determine the first data to be migrated in the target hash bucket, and there is no need to calculate each data through a hash algorithm. Moreover, the cross-cluster migration of the first data to be migrated can be achieved directly through the above two mapping relationships, thereby achieving efficient data migration across clusters.

[0088] It should be noted that the data migration method of the present application can be implemented on a container cluster (i.e., a K8S cluster), a physical machine or a virtual machine. The only difference between the deployment methods on physical machines and those based on containerization and virtual machines is whether to use processes or container instances when monitoring migration. Containerization and physical machines use processes to monitor migration, and virtual machines use container instances to monitor migration.

[0089] Below, a data migration method in this application is described in detail:

[0090] Figure 1 A flow chart of a data migration method provided in an embodiment of the present application. Figure 1 As shown, a data migration method in an embodiment of the present application includes:

[0091] S100: Obtain a first migration instruction; the first migration instruction includes a first target migration cluster, a target storage actual bucket, and a target hash bucket corresponding to the target storage actual bucket; the first target migration cluster is any cluster among the multiple clusters included in the unified coordination system except the current cluster corresponding to the target storage actual bucket.

[0092] It should be noted that the first migration instruction is generated when it is determined that the system has a fault, or when the number of files in any storage real bucket RealBucket exceeds the first preset value, or is generated according to a request from a user to create a new RealBucket. The number of files in the RealBucket is obtained by the monitoring module in the unified coordination system. The first preset value can be 80,000, 100,000, etc., which can be determined based on the experience of people in this field and is not limited here.

[0093] When the first migration instruction is generated based on the user's request to create a new RealBucket, the user can specify the storage actual bucket that needs to be migrated, thereby migrating the data in the storage actual bucket, or determine the storage actual bucket of the data that needs to be migrated based on the remaining storage space of each storage actual bucket, thereby migrating the data in the storage actual bucket. Priority is given to migrating storage actual buckets with small remaining storage space. The user can also specify the file to be migrated, obtain the identifier of the file to be migrated sent by the user, and according to the pre-stored index table, the index table stores the index relationship between the file identifier and the file URL. According to the identifier of the file to be migrated, the URL of the file to be migrated can be determined, and then according to the type relationship between the URL of the file to be migrated and the version setting, it is calculated to obtain which HashBucket the URL of the file to be migrated belongs to; then the specified file can be migrated.

[0094] It should be noted that the HashBucket to which the file belongs is calculated based on the file URL; the RealBucket to which the file belongs is found based on the HashBucket and RealBucket mapping table; and the cluster where the file is located can be found based on the RealBucket and cluster mapping table. In this way, the real physical location of the file can be obtained.

[0095] It should be noted that the target hash bucket is the hash bucket where the data to be migrated is currently stored, the target storage actual bucket is the storage actual bucket where the data to be migrated is currently stored, the first target migration cluster is the target cluster for data migration, the storage actual bucket (RealBucket) is distributed on each cluster in the unified coordination system, and is the actual physical storage location of the file; the hash bucket (HashBucket) is the logical location of the user file.

[0096] It should be noted that before implementing data migration, data corresponding to the user account needs to be stored, and the user account and the user bucket need to be configured. Specifically, before receiving the first migration instruction, the method further includes:

[0097] Accepting a user creation request sent by a user; the user creation request includes user information; creating account information corresponding to the user information in the database of the unified coordination system according to the user information in the user creation request, storing the user information in the database, and generating a user information creation success instruction; generating a container creation request according to the user information creation success instruction; the container creation request includes user container information; storing the user container information in the database of the unified coordination system, and writing the user container information into a newly created user bucket according to the container creation request to generate a first user bucket; determining the number of actual buckets stored in the first user bucket according to the container creation request; determining at least one cluster of multiple clusters included in the unified coordination system for creating a new first user bucket group; create a new storage actual bucket in at least one cluster of the multiple storage actual buckets included in the first user bucket; determine the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket according to the container creation request; the bit width is used to determine the number of files that the hash bucket can store; create the first mapping table according to the number of storage actual buckets and the number of hash buckets in the first user bucket; and create a hash bucket corresponding to each storage actual bucket in the multiple storage actual buckets included in the first user bucket according to the first mapping table; create the third mapping table according to the number of storage actual buckets in the first user bucket and the cluster where each storage actual bucket in the multiple storage actual buckets included in the first user bucket is located; store the first mapping table and the third mapping table in the unified storage access layer OSS of the unified coordination system.

[0098] Specifically, determining the number of actual buckets stored in the first user bucket according to the container creation request includes: judging whether the container creation request includes an estimated number; if the container creation request includes the estimated number, determining the number of actual buckets stored in the first user bucket according to the estimated number; if the container creation request does not include the estimated number, determining the number of actual buckets stored in the first user bucket according to a default estimated number;

[0099] Specifically, determining the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket according to the container creation request includes: judging whether the container creation request includes an estimated number; if the container creation request includes the estimated number, determining the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket according to the estimated number; if the container creation request does not include the estimated number, determining the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket according to a default estimated number.

[0100] It should be noted that a user bucket (UserBucket) or user container (UserContainer) must belong to a certain user account (Account), and a user account can have multiple user buckets or user containers; each UserBucket / UserContainer must have one or more groups of different versions of hash buckets HashBuckets, and each UserBucket / UserContainer has multiple storage actual buckets (RealBucket) and / or storage actual containers (RealContainer) set in the cluster, where RealBucket can be set in multiple different clusters and / or the same cluster under the corresponding cluster structure in the unified coordination system, and RealContainer can also be set in multiple different clusters and / or the same cluster under the corresponding cluster structure in the unified coordination system.

[0101] Among them, in the embodiment of the present application, data migration can be implemented for different cluster structures at the same time. The multi-heterogeneous cluster structure in the present application includes two structures, CephCluster and Swiftcluster. RealBucket is created for CephCluster, and RealContainer is created for Swiftcluster.

[0102] Specifically, as attached Figure 2 As shown, Account 1 (Account1) and Account 2 (Account2) are each configured with a corresponding user bucket / user container (UserBucket / UserContainer), and each UserBucket / UserContainer is configured with corresponding RealBucket, RealContainer, and HashBuckets. The RealBucket is set in the corresponding CephCluster, and the RealContainer is set in the corresponding Swiftcluster; the HashBuckets under each UserBucket / UserContainer include one or more groups of version setting types (Version offset). The drawings of the embodiments of the present application only show that three groups of version setting types are included, but it does not mean that the implementation of the present application can only include 3 groups. The corresponding number of version setting types can be set according to actual needs, wherein each version setting type corresponds to a hash bucket, and the version setting type of each hash bucket is different.

[0103] As attached Figure 3 The overall process of creating a new user account and a user container in the embodiment of the present application is described as follows:

[0104] First, create a new UserBucket, select the estimated number in the user information (UB) to evaluate the level of the newly created account, generate the corresponding RealBucket (RB) according to the evaluation level, and determine whether the number of storage real buckets RB under the current user bucket (realUser) reaches the threshold. If the threshold is reached, create a new user bucket (realUser) and update the user bucket information table (realUser information table); then generate a mapping table between the hash bucket (HB) and the storage real bucket (RB), and generate the storage real bucket information table (RB information table).

[0105] As attached Figure 4 , Attachment Figure 5 As shown, the specific method of creating users and user containers in the embodiment of the present application is introduced:

[0106] As attached Figure 4 In the unified coordination system, the management and scheduling interface module (Interface) receives the user creation request sent by the user. The Interface processes and parses the request to obtain the user information in the user request, which may include the user ID, the estimated number (user preset storage space), etc. The user preset storage space is used to obtain the user container information, specifically including the number of actual buckets stored in each user bucket, the number of hash buckets, and the bit width corresponding to each hash bucket; thereby generating a container creation request and creating the corresponding actual storage bucket and hash bucket. After parsing and obtaining the user information in the user request, the Interface will assign the task of creating a user account to the creation module of the unified coordination system. The creation module of the unified coordination system will create the account information corresponding to the user information in the database, store the user information in the database (DataBase), generate a user information creation success instruction, and return the user information creation success instruction to the Interface in the unified coordination system. The Interface generates a successful creation message and returns it to the user; if any of the above processes fails, a creation failure message is returned to the user, and the message will carry the reason for the failure, such as duplicate user ID, failed parsing of the creation user request, etc.

[0107] As attached Figure 5After generating a user information creation success instruction, Interface will generate a container creation request including user container information according to the user information creation success instruction, and distribute it to the creation module of the unified coordination system. The creation module of the unified coordination system will store the user container information in the database according to the container creation request. The creation module of the unified coordination system selects several clusters according to the estimation information (estimated number) and cluster weights and creates realbuckets in these clusters; wherein, the cluster weight will be configured when the unified coordination system creates the cluster, and the cluster weight can be adjusted according to the actual storage situation; then, according to the estimated number or cluster weight in the creation request, The default estimated number calculation selects the mapping bit width of HashBucket, and maps HashBucket into realbucket; then the first mapping table is created according to the number of actual storage buckets and the number of hash buckets in the first user bucket, and a hash bucket corresponding to each of the multiple actual storage buckets included in the first user bucket is newly created according to the first mapping table; a third mapping table is created according to the number of actual storage buckets in the first user bucket and the cluster where each of the multiple actual storage buckets included in the first user bucket is located; the first mapping table and the third mapping table are stored in the unified storage access layer OSS of the unified coordination system.

[0108] S101: Create a first storage actual bucket in the first target migration cluster and a first hash bucket corresponding to the first storage actual bucket.

[0109] It should be noted that the creation of a new first storage actual bucket in the target migration cluster may be the creation of one or more storage actual buckets. In the embodiment of the present application, a new storage actual bucket is used for illustration. The hash bucket corresponding to the new first storage actual bucket may be one or more. In the embodiment of the present application, a new hash bucket corresponding to the new storage actual bucket is used for illustration. The number of new ones created in this application is not limited here.

[0110] Specifically, S102: adding the correspondence between the first hash bucket and the first actual storage bucket to a pre-stored first mapping table to obtain a second mapping table; adding the correspondence between the first actual storage bucket and the first target migration cluster to a pre-stored third mapping table to obtain a fourth mapping table.

[0111] It should be noted that the first mapping table and the third mapping table are stored in a unified original data center module in the unified coordination system.

[0112] Specifically, the first mapping table is a mapping table of hash buckets and actual storage buckets, storing the correspondence between target hash buckets and target actual storage buckets, and the third mapping table is a mapping table of actual storage buckets and clusters, storing the correspondence between target actual storage buckets and current clusters.

[0113] S103: Determine the first data to be migrated in the target hash bucket, and migrate the first data to be migrated in the target hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table and the fourth mapping table.

[0114] It should be noted that the first data to be migrated in the target hash bucket can be all the data in the target hash bucket or part of the data in the target hash bucket. If it is part of the data in the target hash bucket, the part of the data in the target hash bucket to be migrated is determined according to the version of the target hash bucket and the hash algorithm.

[0115] Specifically, the method also includes: obtaining the migration status of the first data to be migrated in the target hash bucket; then, according to the second mapping table and the fourth mapping table, migrating the first data to be migrated in the target hash bucket from the target storage actual bucket to the first storage actual bucket, including: setting the migration status of the first data to be migrated in the target hash bucket to migrating; migrating the first data to be migrated in the target hash bucket from the target hash bucket to the first hash bucket to obtain the data to be migrated in the first hash bucket; adding the correspondence between the first hash bucket and the target storage actual bucket in the row where the correspondence between the first hash bucket and the first storage actual bucket is located in the second mapping table, Generate a second mapping table in the migration state; add a correspondence between the first storage actual bucket and the current cluster in the row where the correspondence between the first storage actual bucket and the first target migration cluster is located in the fourth mapping table, and generate a fourth mapping table in the migration state; according to the second mapping table in the migration state and the fourth mapping table in the migration state, migrate the storage actual bucket corresponding to the data to be migrated in the first hash bucket from the target storage actual bucket to the first storage actual bucket; delete the correspondence between the first hash bucket and the target storage actual bucket in the second mapping table in the migration state, and delete the correspondence between the first storage actual bucket and the current cluster in the fourth mapping table in the migration state.

[0116] Specifically, as attached Figure 6 For the first mapping table in the migration state, that is, the mapping table of hash buckets and actual storage buckets, the corresponding relationship between the first hash bucket and the first actual storage bucket in the mapping table of hash buckets and actual storage buckets is temporarily configured, that is, the original storage actual bucket (target actual bucket) of the first data to be migrated is temporarily added to the row where the corresponding relationship between the first hash bucket and the first actual storage bucket is located ... Figure 6 Once the data migration of the target hash bucket is completed, the RC-old will be deleted from the row where the correspondence between the first hash bucket and the first storage actual bucket is located.

[0117] Specifically, as attached Figure 6 For the third mapping table in the migration state, i.e., the mapping table of the storage actual bucket and the cluster, the row where the correspondence between the first storage actual bucket and the first target migration cluster is located in the mapping table of the storage actual bucket and the cluster will temporarily configure the correspondence between the first storage actual bucket and the current cluster, i.e., the cluster where the first data to be migrated originally was located (the current cluster corresponding to the target storage actual bucket) is temporarily added to the row where the correspondence between the first storage actual bucket and the first target migration cluster is located, i.e., the attached Figure 6 Once the data migration is completed, Cluster-old will be deleted from the row where the correspondence between the first storage actual bucket and the first target migration cluster is located.

[0118] Below, as attached Figure 7 As shown, the cross-cluster data migration method in the embodiment of the present application is generally introduced:

[0119] A first migration instruction is obtained, which includes a first target migration cluster, a target storage actual bucket, and a target hash bucket corresponding to the target storage actual bucket. According to the first migration instruction, a cluster is selected to create a new realbucket, and corresponding hashbuckets are created. Then, the hash bucket to be migrated is determined, and the migration status of the hash bucket to be migrated is set to migrating. The first mapping table and the third mapping table are modified to obtain an updated mapping table. The updated mapping table is uploaded to the unified storage access layer OSS. The unified storage access layer OSS will store the updated mapping table in the unified metadata center module. When the scheduling migration module migrates data, the updated mapping table file will be downloaded from the unified metadata center module through the metadata service. The file specifies which hashbuckets are to be migrated. The scheduling migration module starts the data migration plug-in to migrate the data in the hashbuckets according to the updated mapping table file, and the monitoring module in the unified coordination system starts the migration tracking plug-in to monitor the progress of the migration.

[0120] It should be noted that after the data migration is completed, the unified coordination system will modify the metadata information and synchronize it to all metadata servers involved.

[0121] It should be noted that the unified coordination system includes a unified operation and maintenance center, a unified metadata center module, and a unified storage access layer OSS; the above-mentioned management and scheduling interface module, monitoring module, and scheduling and migration module are all modules of the unified operation and maintenance center in the unified coordination system.

[0122] As attached Figure 8 As shown, the metadata update process of the mapping table before the operation and maintenance center of the unified coordination system in this embodiment of the present application migrates the HashBucket is introduced:

[0123] After obtaining the first migration instruction, the unified operation and maintenance center will modify the metadata of the change point, and then send it to the unified storage access layer OSS, and store it on the disk in a unified manner; after updating the metadata of the change point, the unified operation and maintenance center will send an update command to let all metadata servers update the metadata stored on the disk, and the unified metadata center will send a lock command to lock the output put instruction operation of the corresponding hash bucket (HashBucket, hb) and the storage actual bucket (Realbucket, RC); after the unified metadata center has completed the locking of the output put instruction operations of the hash bucket (HashBucket, hb) and the storage actual bucket (Realbucket, RC), the new metadata will be activated, and the service instance of the activated metadata will automatically unlock and return the instruction of the successful update, so that the unified operation and maintenance center can start migrating the HashBucket according to the instruction of the successful update. Among them, the service instance that failed in the metadata update process needs to stop the service and wait for the operation and maintenance inspection to start.

[0124] Furthermore, in this embodiment of the present application, when the number of HashBuckets files is close to the expected number, a group of HashBuckets can be added by adding a new version of HashBuckets, and the number of the new group of HashBuckets can be increased to increase the upper limit of the number of files that can be stored. The method also includes:

[0125] Obtain the number of files in the target hash bucket; determine whether the number of files in the target hash bucket reaches a second preset threshold; if so, create a fourth hash bucket corresponding to the target storage actual bucket, add the correspondence between the fourth hash bucket and the target storage actual bucket in the first mapping table, and obtain a ninth mapping table.

[0126] It should be noted that a group of HashBuckets corresponds to one version, and the number of files in the current version can be controlled by controlling the number of a group of HashBuckets.

[0127] It should be noted that when a cluster reaches the capacity limit (threshold), it can be balanced by scheduling the RealBucket on the cluster to other clusters; it can also schedule some HashBuckets on the cluster to the RealBucket of other clusters. The corresponding mapping table needs to be updated and data migration is performed according to the mapping table.

[0128] The embodiment of the present application provides a data migration method, in which the first migration instruction is obtained; a first storage actual bucket and a first hash bucket corresponding to the first storage actual bucket are newly created in the first target migration cluster; the correspondence between the first hash bucket and the first storage actual bucket is added to a pre-stored first mapping table to obtain a second mapping table; the correspondence between the first storage actual bucket and the first target migration cluster is added to a pre-stored third mapping table to obtain a fourth mapping table; the first data to be migrated is determined from the target hash bucket, and the first data to be migrated in the target hash bucket is migrated from the target storage actual bucket to the first storage actual bucket according to the second mapping table and the fourth mapping table; it can be seen that in the embodiment of the present application, based on establishing a hash bucket as a theoretical storage location and a storage actual bucket as an actual storage location, by configuring a mapping relationship between a hash bucket and a storage actual bucket, and a mapping relationship between a storage actual bucket and a cluster, during data migration, it is only necessary to determine the first data to be migrated in the target hash bucket, and it is not necessary to calculate each data through a hash algorithm, and the cross-cluster migration of the first data to be migrated can be achieved directly through the above two mapping relationships, thereby achieving efficient data migration across clusters.

[0129] Below, as attached Fig. 9 As shown, the scheme for data migration and balancing in the current cluster in the embodiment of the present application is described: the data migration method also includes:

[0130] S200: Acquire a second migration instruction; the second migration instruction includes the target storage actual bucket and the target hash bucket;

[0131] S201: obtaining a new second storage actual bucket in the current cluster according to the second migration instruction, and a second hash bucket corresponding to the second storage actual bucket;

[0132] S202: Obtain a correspondence between the second hash bucket and the second storage actual bucket in the first mapping table to obtain a fifth mapping table;

[0133] S203: Determine the second data to be migrated in the target hash bucket, and migrate the storage actual bucket corresponding to the second data to be migrated in the second target hash bucket from the second target storage actual bucket to the second storage actual bucket according to the fifth mapping table.

[0134] Specifically, as attached Fig.10As shown, the overall interactive process obtains the second migration instruction, creates a new realbucket in the current cluster, and creates corresponding hashbuckets, then determines the hash bucket to be migrated, sets the migration status of the hash bucket to be migrated to migrating, modifies the first mapping table and the third mapping table, obtains the updated mapping table, and uploads the updated mapping table to the unified storage access layer OSS. The unified storage access layer OSS will store the updated mapping table in the unified metadata center module. When the scheduling migration module migrates data, it will download the updated mapping table file from the unified metadata center module through the metadata service. The file specifies which hashbuckets are to be migrated. The scheduling migration module starts the data migration plug-in to migrate the data in the hashbuckets according to the updated mapping table file, and the monitoring module in the unified coordination system starts the migration tracking plug-in to monitor the progress of the migration.

[0135] In this embodiment of the present application, when the current cluster performs data migration and balancing, based on establishing a hash bucket as a theoretical storage location and a storage actual bucket as an actual storage location, by configuring the mapping relationship between the hash bucket and the storage actual bucket, as well as the mapping relationship between the storage actual bucket and the cluster, when migrating data, it is only necessary to determine the first data to be migrated in the target hash bucket, and there is no need to calculate each data through a hash algorithm, thereby improving the efficiency of migrating data in the same cluster.

[0136] Below, as attached Fig.11 As shown, the scheme of migrating the entire storage actual bucket to the new cluster in the embodiment of the present application is described: the data migration method also includes:

[0137] S300: Acquire a third migration instruction, where the third migration instruction includes the second target migration cluster, where the second target migration cluster is any cluster other than the current cluster among the multiple clusters included in the unified coordination system;

[0138] S301: Acquire the third mapping table, wherein the third mapping table includes a correspondence between the target storage actual bucket and the current cluster;

[0139] S302: According to the third migration instruction, the correspondence between the target storage actual bucket and the current cluster in the third mapping table is updated to the correspondence between the target storage actual bucket and the second target migration cluster, to obtain a sixth mapping table;

[0140] S303: Migrate the target storage actual bucket from the current cluster to the second target migration cluster according to the sixth mapping table.

[0141] In the embodiment of the present application, the entire actual storage bucket can be migrated across clusters. Based on the mapping relationship between the actual storage bucket and the cluster, during data migration, the entire actual storage bucket can be directly migrated across clusters through the above mapping relationship, thereby achieving efficient data migration across clusters.

[0142] Below, as attached Fig.12 As shown, in the embodiment of the present application, the cluster in the unified coordination system can also be expanded, and the data migration method also includes:

[0143] S400: Acquire a cluster expansion instruction, where the cluster expansion instruction is generated according to a received user cluster expansion request, or is generated when it is determined that the storage capacity of the current cluster reaches a first preset threshold;

[0144] S401: adding a first newly added cluster in the unified coordination system according to the cluster expansion instruction, acquiring cluster information of the first newly added cluster, storing the cluster information of the first newly added cluster in a database in the unified coordination system, and sending the cluster information of the first newly added cluster to a metadata server in the unified coordination system;

[0145] S402: creating a third storage actual bucket in the first newly added cluster, and a third hash bucket corresponding to the third storage actual bucket;

[0146] S403: Add the correspondence between the third hash bucket and the third actual storage bucket to the first mapping table to obtain a seventh mapping table; add the correspondence between the third actual storage bucket and the first newly added cluster to the third mapping table to obtain an eighth mapping table;

[0147] S404: Determine the hash bucket data to be migrated, and migrate the hash bucket data to be migrated from the storage actual bucket corresponding to the hash bucket data to be migrated to the third storage actual bucket according to the seventh mapping table and the eighth mapping table.

[0148] Specifically, if the cluster expansion instruction is generated according to a received user cluster expansion request, the determining of the hash bucket data to be migrated, and migrating the hash bucket data to be migrated from the storage actual bucket corresponding to the hash bucket data to be migrated to the third storage actual bucket according to the seventh mapping table and the eighth mapping table, includes:

[0149] Obtain the weight of each cluster among the multiple clusters included in the unified coordination system, wherein the weight of each cluster in the unified coordination system is determined according to the version type of each cluster; determine the cluster of the data to be migrated according to the weight of each cluster in the unified coordination system; determine the storage actual bucket of the data to be migrated according to the remaining storage space of each storage actual bucket among the multiple storage actual buckets included in the cluster of the data to be migrated; determine the storage actual bucket of the data to be migrated according to the remaining storage space of each storage actual bucket in the cluster to be migrated; obtain the hash bucket corresponding to the storage actual bucket of the data to be migrated according to the first pre-stored mapping table, and obtain the hash bucket of the data to be migrated; determine the hash bucket data to be migrated from the hash bucket to be migrated according to a preset algorithm; and migrate the hash bucket data to be migrated from the storage actual bucket corresponding to the hash bucket data to be migrated to the third storage actual bucket according to the seventh mapping table and the eighth mapping table.

[0150] It should be noted that in the embodiment of the present application, the migration status of data can be monitored by a monitoring module in a unified coordination system, and a migration completion report will be generated when the migration is completed.

[0151] The following is an introduction to a unified coordination system in an embodiment of the present application. Fig.13 Based on a data migration method in the above embodiment, the embodiment of the present application implements the data migration method through the unified coordination system. The unified coordination system in the embodiment of the present application includes:

[0152] The management scheduling interface module 10 is used to obtain a first migration instruction; the first migration instruction includes a first target migration cluster, a target storage actual bucket, and a target hash bucket corresponding to the target storage actual bucket; the first target migration cluster is any cluster among the multiple clusters included in the unified coordination system except the current cluster corresponding to the target storage actual bucket;

[0153] The scheduling migration module 11 is used to create a first storage actual bucket in the first target migration cluster and a first hash bucket corresponding to the first storage actual bucket;

[0154] The unified metadata center module 12 is used to add the correspondence between the first hash bucket and the first storage actual bucket in the pre-stored first mapping table to obtain a second mapping table; add the correspondence between the first storage actual bucket and the first target migration cluster in the pre-stored third mapping table to obtain a fourth mapping table;

[0155] The scheduling migration module 11 is further used to determine the first data to be migrated in the target hash bucket, and migrate the first data to be migrated in the target hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table and the fourth mapping table.

[0156] Specifically, the scheduling migration module is specifically used to:

[0157] Obtain the migration status of the first data to be migrated in the target hash bucket; set the migration status of the first data to be migrated in the target hash bucket to being migrated; migrate the first data to be migrated in the target hash bucket from the target hash bucket to the first hash bucket to obtain the data to be migrated in the first hash bucket; add the correspondence between the first hash bucket and the target storage actual bucket in the row where the correspondence between the first hash bucket and the first storage actual bucket is located in the second mapping table, and generate a second mapping table in the migration status; add the correspondence between the first storage actual bucket and the current cluster in the row where the correspondence between the first storage actual bucket and the first target migration cluster is located in the fourth mapping table, and generate a fourth mapping table in the migration status; according to the second mapping table in the migration status and the fourth mapping table in the migration status, migrate the storage actual bucket corresponding to the data to be migrated in the first hash bucket from the target storage actual bucket to the first storage actual bucket; delete the correspondence between the first hash bucket and the target storage actual bucket in the second mapping table in the migration status, and delete the correspondence between the first storage actual bucket and the current cluster in the fourth mapping table in the migration status.

[0158] Specifically, in the unified coordination system:

[0159] The management scheduling interface module is further used to obtain a second migration instruction; the second migration instruction includes the target storage actual bucket and the target hash bucket;

[0160] The scheduling migration module is further used to create a second storage actual bucket in the current cluster according to the second migration instruction, and a second hash bucket corresponding to the second storage actual bucket;

[0161] The unified metadata center module is further used to add the corresponding relationship between the second hash bucket and the second storage actual bucket in the first mapping table to obtain a fifth mapping table;

[0162] The scheduling migration module is also used to determine the second data to be migrated in the target hash bucket, and migrate the storage actual bucket corresponding to the second data to be migrated in the second target hash bucket from the second target storage actual bucket to the second storage actual bucket according to the fifth mapping table.

[0163] Specifically, in the unified coordination system:

[0164] The management scheduling interface module is further used to obtain a third migration instruction, where the third migration instruction includes the second target migration cluster, where the second target migration cluster is any cluster other than the current cluster among the multiple clusters included in the unified coordination system;

[0165] The unified metadata center module is further used to obtain the third mapping table, which includes the correspondence between the target storage actual bucket and the current cluster; according to the third migration instruction, the correspondence between the target storage actual bucket and the current cluster in the third mapping table is updated to the correspondence between the target storage actual bucket and the second target migration cluster, so as to obtain a sixth mapping table;

[0166] The scheduling migration module is further used to migrate the target storage actual bucket from the current cluster to the second target migration cluster according to the sixth mapping table.

[0167] Specifically, in the unified coordination system:

[0168] The scheduling migration module is further used to obtain a cluster expansion instruction, where the cluster expansion instruction is generated according to a received user cluster expansion request, or is generated when it is determined that the storage capacity of the current cluster reaches a first preset threshold;

[0169] The scheduling migration module is further used to add a first newly added cluster in the unified coordination system according to the cluster expansion instruction, obtain cluster information of the first newly added cluster, store the cluster information of the first newly added cluster in a database in the unified coordination system, and send the cluster information of the first newly added cluster to a metadata server in the unified coordination system;

[0170] The unified metadata center module is further used to create a third storage actual bucket in the first newly added cluster, and a third hash bucket corresponding to the third storage actual bucket; add the correspondence between the third hash bucket and the third storage actual bucket in the first mapping table to obtain a seventh mapping table; add the correspondence between the third storage actual bucket and the first newly added cluster in the third mapping table to obtain an eighth mapping table;

[0171] The scheduling migration module is also used to determine the hash bucket data to be migrated, and migrate the hash bucket data to be migrated from the storage actual bucket corresponding to the hash bucket data to be migrated to the third storage actual bucket according to the seventh mapping table and the eighth mapping table.

[0172] Specifically, if the cluster expansion instruction is generated according to a received user cluster expansion request, the scheduling migration module is specifically used to:

[0173] Obtain the weight of each cluster among the multiple clusters included in the unified coordination system, wherein the weight of each cluster in the unified coordination system is determined according to the version type of each cluster; determine the cluster of the data to be migrated according to the weight of each cluster in the unified coordination system; determine the storage actual bucket of the data to be migrated according to the remaining storage space of each storage actual bucket among the multiple storage actual buckets included in the cluster of the data to be migrated; determine the storage actual bucket of the data to be migrated according to the remaining storage space of each storage actual bucket in the cluster to be migrated; obtain the hash bucket corresponding to the storage actual bucket of the data to be migrated according to the first pre-stored mapping table, and obtain the hash bucket of the data to be migrated; determine the hash bucket data to be migrated from the hash bucket to be migrated according to a preset algorithm; and migrate the hash bucket data to be migrated from the storage actual bucket corresponding to the hash bucket data to be migrated to the third storage actual bucket according to the seventh mapping table and the eighth mapping table.

[0174] Specifically, the unified coordination system also includes:

[0175] A monitoring module, used to obtain the number of files in a target hash bucket; and determine whether the number of files in the target hash bucket reaches a second preset threshold;

[0176] If so, the scheduling migration module is also used to create a fourth hash bucket corresponding to the target storage actual bucket; the unified metadata center module is also used to add the correspondence between the fourth hash bucket and the target storage actual bucket in the first mapping table to obtain a ninth mapping table.

[0177] Specifically, in the unified coordination system:

[0178] The management scheduling interface module is also used to accept a user creation request sent by a user; the user creation request includes user information;

[0179] A creation module, configured to create account information corresponding to the user information in the database of the unified coordination system according to the user information in the user creation request, store the user information in the database, and generate a user information creation success instruction;

[0180] The management scheduling interface module is further used to generate a container creation request according to the user information creation success instruction; the container creation request includes user container information;

[0181] The creation module is further used to store the user container information in the database of the unified coordination system, and write the user container information into the newly created user bucket according to the container creation request to generate a first user bucket; determine the number of actual storage buckets in the first user bucket according to the container creation request; determine at least one cluster for creating multiple actual storage buckets included in the first user bucket among the multiple clusters included in the unified coordination system; create a new actual storage bucket in at least one cluster of the multiple actual storage buckets included in the first user bucket; determine the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket according to the container creation request; the bit width is used to determine the number of files that can be stored in the hash bucket;

[0182] The unified metadata center module is further used to create the first mapping table according to the number of actual storage buckets and the number of hash buckets in the first user bucket; create the third mapping table according to the number of actual storage buckets in the first user bucket and the cluster where each actual storage bucket in the multiple actual storage buckets included in the first user bucket is located;

[0183] The creation module is also used to create a hash bucket corresponding to each of the multiple storage actual buckets included in the first user bucket according to the first mapping table; and store the first mapping table and the third mapping table in the unified storage access layer OSS of the unified coordination system.

[0184] The creation module is further configured to determine whether the container creation request includes an estimated number; if the container creation request includes the estimated number, determine the number of actual buckets stored in the first user bucket according to the estimated number; if the container creation request does not include the estimated number, determine the number of actual buckets stored in the first user bucket according to a default estimated number;

[0185] The creation module is further used to determine whether the container creation request includes an estimated number; if the container creation request includes the estimated number, the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket are determined according to the estimated number; if the container creation request does not include the estimated number, the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket are determined according to a default estimated number.

[0186] The embodiment of the present application discloses a unified coordination system, which is based on establishing a hash bucket as a theoretical storage location and an actual storage bucket as an actual storage location, and by configuring a mapping relationship between the hash bucket and the actual storage bucket, and a mapping relationship between the actual storage bucket and the cluster. During data migration, it is only necessary to determine the first data to be migrated in the target hash bucket, and there is no need to calculate each data through a hash algorithm. Moreover, the cross-cluster migration of the first data to be migrated can be achieved directly through the above two mapping relationships, thereby achieving efficient data migration across clusters.

[0187] Below, as attached Fig.14 As shown, this embodiment of the present application introduces the framework of the unified coordination system:

[0188] The unified coordination system includes:

[0189] Two cluster architectures: Swift architecture and Ceph architecture; Swift architecture can have multiple Swift clusters, Ceph architecture can have multiple Ceph clusters, clusters are used to store user file data

[0190] Unified metadata center module (Metadata group), the unified coordination system can include multiple unified metadata center modules; the reading and storage of mapping table files are realized through the unified storage access layer OSS;

[0191] Log service module, used to record requests and coordinate system operations;

[0192] Management scheduling interface module (Interface, OR), the docking interface with the user, is used to receive and process requests sent by the front end;

[0193] Migration tools: The coordination center in the unified coordination system uses the migration tools to schedule data;

[0194] Management center (Bypass TTL), used to coordinate the management of system configuration;

[0195] The coordination center, also known as the unified operation and maintenance center, is used to implement data scheduling. Specifically, the coordination center determines whether balancing is required by monitoring the number of files in hashbucket and RealBucket. The coordination center usually does not allow the number of files in each RealBucket to exceed N million, where N is customizable and is recommended to be around 100,000. The coordination center needs to coordinate the data update, progress monitoring, and management of the mapping table (the mapping table of Realbucket and HashBucket) of data migration designed in the metadata cluster.

[0196] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data migration method, characterized in that: The method is performed by a unified coordination system, and the method includes: Obtain a first migration instruction; the first migration instruction includes a first target migration cluster, a target storage actual bucket, and a target hash bucket corresponding to the target storage actual bucket; the first target migration cluster is any cluster among the multiple clusters included in the unified coordination system except the current cluster corresponding to the target storage actual bucket; Creating a first storage actual bucket in the first target migration cluster and a first hash bucket corresponding to the first storage actual bucket; Adding the correspondence between the first hash bucket and the first actual storage bucket to the pre-stored first mapping table to obtain a second mapping table; adding the correspondence between the first actual storage bucket and the first target migration cluster to the pre-stored third mapping table to obtain a fourth mapping table; Determine the first data to be migrated in the target hash bucket, and migrate the first data to be migrated in the target hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table and the fourth mapping table; The method further includes: obtaining a migration status of the first data to be migrated in the target hash bucket; and migrating the first data to be migrated in the target hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table and the fourth mapping table, including: Setting the migration status of the first data to be migrated in the target hash bucket to being migrated; Migrating the first data to be migrated in the target hash bucket from the target hash bucket to the first hash bucket to obtain the data to be migrated in the first hash bucket; In the second mapping table, the row where the correspondence between the first hash bucket and the first actual storage bucket is located is added with the correspondence between the first hash bucket and the target actual storage bucket, so as to generate the second mapping table in the migration state; in the fourth mapping table, the row where the correspondence between the first actual storage bucket and the first target migration cluster is located is added with the correspondence between the first actual storage bucket and the current cluster, so as to generate the fourth mapping table in the migration state; Uploading the second mapping table in the migration state and the fourth mapping table in the migration state to a unified metadata center module via a unified storage access layer; The unified operation and maintenance center modifies the metadata of the second mapping table and the fourth mapping table and sends them to the unified storage access layer, and after the metadata of the second mapping table and the fourth mapping table are modified, sends an update command to the metadata server to instruct the metadata server to update the metadata stored in the metadata server; The unified metadata operation and maintenance center sends a lock command to lock the target storage actual bucket and the target hash bucket corresponding to the target storage actual bucket; after all the output instruction operations of the target storage actual bucket and the target storage actual bucket are locked, the unified metadata center activates the new metadata, the service instance of the activated metadata is automatically unlocked, and an update success instruction is sent to the unified operation and maintenance center, so that the unified operation and maintenance center starts to migrate the hash bucket according to the update success instruction; Obtaining the second mapping table in the migration state and the fourth mapping table in the migration state from the unified metadata center module through the metadata service, and migrating the storage actual bucket corresponding to the data to be migrated in the first hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table in the migration state and the fourth mapping table in the migration state; Deleting the corresponding relationship between the first hash bucket and the target storage actual bucket in the second mapping table in the migration state, and deleting the corresponding relationship between the first storage actual bucket and the current cluster in the fourth mapping table in the migration state; The method also includes: obtaining the number of files in the target hash bucket; determining whether the number of files in the target hash bucket reaches a second preset threshold; if so, creating a fourth hash bucket corresponding to the target storage actual bucket, adding the corresponding relationship between the fourth hash bucket and the target storage actual bucket in the first mapping table, and obtaining a ninth mapping table.

2. The method according to claim 1, characterized in that: The method further comprises: Obtain a second migration instruction; the second migration instruction includes the target storage actual bucket and the target hash bucket; Creating a second storage actual bucket in the current cluster according to the second migration instruction and a second hash bucket corresponding to the second storage actual bucket; Adding the correspondence between the second hash bucket and the second storage actual bucket to the first mapping table to obtain a fifth mapping table; The second data to be migrated in the target hash bucket is determined, and the storage actual bucket corresponding to the second data to be migrated in the second target hash bucket is migrated from the second target storage actual bucket to the second storage actual bucket according to the fifth mapping table.

3. The method according to claim 1, characterized in that: The method further comprises: Acquire a third migration instruction, where the third migration instruction includes a second target migration cluster, where the second target migration cluster is any cluster other than the current cluster among the multiple clusters included in the unified coordination system; Obtain the third mapping table, wherein the third mapping table includes a correspondence between the target storage actual bucket and the current cluster; According to the third migration instruction, the correspondence between the target storage actual bucket and the current cluster in the third mapping table is updated to the correspondence between the target storage actual bucket and the second target migration cluster, so as to obtain a sixth mapping table; The target storage actual bucket is migrated from the current cluster to the second target migration cluster according to the sixth mapping table.

4. The method according to claim 1, characterized in that: The method further comprises: Obtaining a cluster expansion instruction, where the cluster expansion instruction is generated according to a received user request to expand the cluster, or is generated when it is determined that the storage capacity of the current cluster reaches a first preset threshold; Adding a first newly added cluster in the unified coordination system according to the cluster expansion instruction, acquiring cluster information of the first newly added cluster, storing the cluster information of the first newly added cluster in a database in the unified coordination system, and sending the cluster information of the first newly added cluster to a metadata server in the unified coordination system; Creating a third storage actual bucket in the first newly added cluster and a third hash bucket corresponding to the third storage actual bucket; Adding the correspondence between the third hash bucket and the third actual storage bucket to the first mapping table to obtain a seventh mapping table; adding the correspondence between the third actual storage bucket and the first newly added cluster to the third mapping table to obtain an eighth mapping table; The hash bucket data to be migrated is determined, and according to the seventh mapping table and the eighth mapping table, the hash bucket data to be migrated is migrated from the storage actual bucket corresponding to the hash bucket data to be migrated to the third storage actual bucket.

5. The method according to claim 4, characterized in that: If the cluster expansion instruction is generated according to a received user cluster expansion request, determining the hash bucket data to be migrated, and migrating the hash bucket data to be migrated from the storage actual bucket corresponding to the hash bucket data to be migrated to the third storage actual bucket according to the seventh mapping table and the eighth mapping table, includes: Acquire a weight of each cluster in the plurality of clusters included in the unified coordination system, where the weight of each cluster in the unified coordination system is determined according to a version type of each cluster; Determining a cluster of data to be migrated according to the weight of each cluster in the unified coordination system; According to the remaining storage space of each of the multiple storage actual buckets included in the cluster of the data to be migrated; Determine the actual storage bucket for the data to be migrated according to the remaining storage space of each actual storage bucket in the cluster to be migrated; According to the first pre-stored mapping table, a hash bucket corresponding to the actual storage bucket of the data to be migrated is acquired to obtain the hash bucket of the data to be migrated; Determine the hash bucket data to be migrated from the hash bucket to be migrated according to a preset algorithm; According to the seventh mapping table and the eighth mapping table, the hash bucket data to be migrated is migrated from the storage actual bucket corresponding to the hash bucket data to be migrated to the third storage actual bucket.

6. The method according to claim 1, characterized in that: Before receiving the first migration instruction, the method further includes: Accepting a user creation request sent by a user; the user creation request includes user information; Creating account information corresponding to the user information in the database of the unified coordination system according to the user information in the user creation request, storing the user information in the database, and generating a user information creation success instruction; Generate a container creation request according to the user information creation success instruction; the container creation request includes user container information; storing the user container information in a database of the unified coordination system, and writing the user container information into a newly created user bucket according to the container creation request to generate a first user bucket; Determine the number of actual buckets stored in the first user bucket according to the container creation request; Determine at least one cluster among the multiple clusters included in the unified coordination system for creating a plurality of storage actual buckets included in the first user bucket; create a new storage actual bucket in at least one cluster among the multiple storage actual buckets included in the first user bucket; Determine the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket according to the container creation request; the bit width is used to determine the number of files that can be stored in the hash bucket; The first mapping table is created according to the number of actual storage buckets and the number of hash buckets in the first user bucket; and a hash bucket corresponding to each actual storage bucket in the multiple actual storage buckets included in the first user bucket is newly created according to the first mapping table; the third mapping table is created according to the number of actual storage buckets in the first user bucket and the cluster where each actual storage bucket in the multiple actual storage buckets included in the first user bucket is located; The first mapping table and the third mapping table are stored in the unified storage access layer OSS of the unified coordination system.

7. The method according to claim 6, characterized in that: The determining, according to the container creation request, the number of actual buckets stored in the first user bucket includes: Determining whether the container creation request includes an estimated number; If the container creation request includes the estimated number, determining the number of actual buckets stored in the first user bucket according to the estimated number; If the container creation request does not include the estimated number, determining the number of actual buckets stored in the first user bucket according to a default estimated number; The determining, according to the container creation request, the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket includes: Determining whether the container creation request includes an estimated number; If the container creation request includes the estimated number, determining the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket according to the estimated number; If the container creation request does not include the estimated number, the number of hash buckets in the first user bucket and the bit width corresponding to each hash bucket are determined according to a default estimated number.

8. A unified coordination system, characterized in that: The unified coordination system includes: A management scheduling interface module is used to obtain a first migration instruction; the first migration instruction includes a first target migration cluster, a target storage actual bucket, and a target hash bucket corresponding to the target storage actual bucket; the first target migration cluster is any cluster among the multiple clusters included in the unified coordination system except the current cluster corresponding to the target storage actual bucket; A scheduling migration module, configured to create a first storage actual bucket and a first hash bucket corresponding to the first storage actual bucket in the first target migration cluster; A unified metadata center module, used to add the correspondence between the first hash bucket and the first storage actual bucket in a pre-stored first mapping table to obtain a second mapping table; add the correspondence between the first storage actual bucket and the first target migration cluster in a pre-stored third mapping table to obtain a fourth mapping table; The scheduling migration module is further used to determine the first data to be migrated in the target hash bucket, and migrate the first data to be migrated in the target hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table and the fourth mapping table; The scheduling migration module is specifically used for: Acquiring a migration status of the first data to be migrated in the target hash bucket; and migrating the first data to be migrated in the target hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table and the fourth mapping table, including: Setting the migration status of the first data to be migrated in the target hash bucket to being migrated; Migrating the first data to be migrated in the target hash bucket from the target hash bucket to the first hash bucket to obtain the data to be migrated in the first hash bucket; In the second mapping table, the row where the correspondence between the first hash bucket and the first actual storage bucket is located is added with the correspondence between the first hash bucket and the target actual storage bucket, so as to generate the second mapping table in the migration state; in the fourth mapping table, the row where the correspondence between the first actual storage bucket and the first target migration cluster is located is added with the correspondence between the first actual storage bucket and the current cluster, so as to generate the fourth mapping table in the migration state; Uploading the second mapping table in the migration state and the fourth mapping table in the migration state to the unified metadata center module via a unified storage access layer; The unified operation and maintenance center modifies the metadata of the second mapping table and the fourth mapping table and sends them to the unified storage access layer, and after the metadata of the second mapping table and the fourth mapping table are modified, sends an update command to the metadata server to instruct the metadata server to update the metadata stored in the metadata server; The unified metadata operation and maintenance center sends a lock command to lock the target storage actual bucket and the target hash bucket corresponding to the target storage actual bucket; after all the output instruction operations of the target storage actual bucket and the target storage actual bucket are locked, the unified metadata center activates the new metadata, the service instance of the activated metadata is automatically unlocked, and an update success instruction is sent to the unified operation and maintenance center, so that the unified operation and maintenance center starts to migrate the hash bucket according to the update success instruction; Obtaining the second mapping table in the migration state and the fourth mapping table in the migration state from the unified metadata center module through the metadata service, and migrating the storage actual bucket corresponding to the data to be migrated in the first hash bucket from the target storage actual bucket to the first storage actual bucket according to the second mapping table in the migration state and the fourth mapping table in the migration state; Deleting the corresponding relationship between the first hash bucket and the target storage actual bucket in the second mapping table in the migration state, and deleting the corresponding relationship between the first storage actual bucket and the current cluster in the fourth mapping table in the migration state; A monitoring module is used to obtain the number of files in a target hash bucket; determine whether the number of files in the target hash bucket reaches a second preset threshold; if so, create a fourth hash bucket corresponding to the target storage actual bucket, add the corresponding relationship between the fourth hash bucket and the target storage actual bucket in the first mapping table, and obtain a ninth mapping table.

Citation Information

Patent Citations

  • Storage method for distributed storage system, electronic equipment and storage medium

    CN112162707A

  • Object positioning method for distributed storage system and electronic equipment

    CN112261097A