Data migration method and data migration system

By updating the main server of the metadata subtree after it is migrated to the target server, the problem of metadata migration affecting the access of the data storage system in the prior art is solved, and load balancing and normal access are achieved.

CN120216458APending Publication Date: 2025-06-27BEIJING DIDI INFINITY TECH & DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311833137.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the case of unbalanced load of metadata servers, the direct migration of metadata causes the client to fail to obtain the migrated metadata normally, affecting the access of the data storage system.

Method used

After the metadata subtree is migrated to the target server, the main server of the metadata subtree is updated. Through steps such as splitting the metadata tree, generating configuration information, and migrating the replication subtree, we ensure that the access of the data storage system does not affect the metadata migration process.

Benefits of technology

It effectively reduces the negative impact on the access of the data storage system during metadata migration, and ensures normal access and load balancing of the data storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216458A_ABST
    Figure CN120216458A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data migration method and a data migration system. According to the embodiment of the invention, a management server sends a metadata tree splitting instruction to a source master server, so that the source master server splits a metadata tree to obtain metadata sub-trees, generates configuration information of the metadata sub-trees, and sends a migration instruction to a source slave server; the method comprises the following steps: sending a source slave server to a source slave server to migrate a replicated sub-tree of a metadata sub-tree of the source slave server to a target server, sending a master server transfer instruction to a source master server, updating the source master server as a source slave server of the metadata sub-tree, determining the replicated sub-tree of the metadata sub-tree, and sending a migration instruction to the source master server, and the source master server migrates the replica sub-tree to the target slave server of the metadata sub-tree. Therefore, in the embodiment of the invention, after the metadata sub-tree is migrated to the target server, the main server of the metadata sub-tree is updated, so that the negative influence on the access to the data storage system in the metadata migration process is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and more particularly, to a data migration method and a data migration system. Background Art

[0002] In order to store and access data, data files are stored in a data storage system, and the data storage system stores and manages the metadata of different data files through a metadata server cluster to provide metadata services. The continuous growth of the number of data files makes the number of metadata also increase continuously. However, the growth of the scale of the metadata trees of different metadata servers in the metadata server cluster and the access volume are different, which causes the phenomenon of load imbalance in the metadata server cluster.

[0003] In order to solve the problem of load imbalance of the metadata server, the prior art directly migrates the metadata in the metadata server with a higher load to the metadata server with a lower load. However, this method will cause the client to be unable to normally obtain the migrated metadata during the metadata migration process, thus having a negative impact on the access to the data storage system. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a data migration method and a data migration system to update the master server of the metadata subtree after the metadata subtree is migrated to the target server, reducing the negative impact on the access to the data storage system during the metadata migration process.

[0005] In a first aspect, an embodiment of the present invention provides a data migration method, which is applicable to a source master server, and the method includes:

[0006] Responding to receiving a metadata tree splitting instruction, splitting the target metadata tree to obtain a metadata subtree;

[0007] Generating configuration information of the metadata subtree;

[0008] Responding to receiving a master server transfer instruction, updating to be the source slave server of the metadata subtree;

[0009] Determining the replication subtree of the metadata subtree;

[0010] Responding to receiving a second migration instruction, migrating the replication subtree to the target slave server corresponding to the metadata subtree.

[0011] Optionally, the method further includes:

[0012] Responding to receiving a removal instruction, deleting the metadata subtree.

[0013] Optionally, the method further includes:

[0014] The sending server group sends a request to obtain the storage area identifier of the metadata subtree;

[0015] The generation of the configuration information of the metadata subtree includes:

[0016] Generating the configuration information according to the received storage area identifier.

[0017] Optionally, the configuration information of the metadata subtree includes at least one of the node change information and the node range information of the metadata subtree and the storage area identifier of the metadata subtree.

[0018] In a second aspect, an embodiment of the present invention provides a data migration method, which is applicable to a management server. The method includes:

[0019] Sending a metadata tree splitting instruction to the source master server to split the target metadata tree and obtain a metadata subtree;

[0020] Sending a first migration instruction to the source slave server to migrate the replication subtree of the metadata subtree corresponding to the source slave server to the target server;

[0021] Sending a master server transfer instruction to the source master server to update the target server as the target master server of the metadata subtree;

[0022] Sending a second migration instruction to the source master server to migrate the replication subtree of the metadata subtree corresponding to the source master server to the corresponding target slave server.

[0023] Optionally, the sending of the master server transfer instruction to the source master server includes:

[0024] Obtaining the migration result sent by the source slave server;

[0025] In response to the migration result indicating that the replication subtree transfer is successful, sending the master server transfer instruction to the source master server, so that the source master server sends a master server transfer request to each target server to determine the target master server.

[0026] Optionally, the sending of the first migration instruction to the source slave server includes:

[0027] In response to the received splitting result indicating that the target metadata tree splitting is successful, sending the first migration instruction to the source slave server.

[0028] In a third aspect, an embodiment of the present invention provides a data migration method, which is applicable to a target server. The method includes:

[0029] The receiving source receives a replicated subtree of the metadata subtree sent by the server, where the metadata subtree is obtained by splitting the target metadata tree;

[0030] Determine the target master server of the metadata subtree and receive a data operation request for processing the metadata subtree.

[0031] Optionally, the method further includes:

[0032] Perform data synchronization to update the replicated subtree.

[0033] In a fourth aspect, an embodiment of the present invention provides a data storage system, where the system includes:

[0034] A management server configured to send a metadata tree splitting instruction, a master server transfer instruction, and a second migration instruction to the source master server, and send a first migration instruction to the source slave server;

[0035] A source master server configured to, in response to receiving a metadata tree splitting instruction, split the target metadata tree to obtain a metadata subtree, generate configuration information of the metadata subtree, update to the source slave server of the metadata subtree in response to receiving a master server transfer instruction, determine the replicated subtree of the metadata subtree, and migrate the replicated subtree to the target slave server corresponding to the metadata subtree in response to receiving a second migration instruction;

[0036] A target server configured to determine the target master server of the metadata subtree and receive a data operation request for processing the metadata subtree.

[0037] In a fifth aspect, an embodiment of the present invention provides a data migration device applicable to a source master server, and the device includes:

[0038] A subtree generation unit configured to, in response to receiving a metadata tree splitting instruction, split the target metadata tree to obtain a metadata subtree;

[0039] A configuration information generation unit configured to generate configuration information of the metadata subtree;

[0040] A server update unit configured to update to the source slave server of the metadata subtree in response to receiving a master server transfer instruction;

[0041] A subtree replication unit configured to determine the replicated subtree of the metadata subtree;

[0042] A migration unit configured to migrate the replicated subtree to the target slave server corresponding to the metadata subtree in response to receiving a second migration instruction.

[0043] Sixth aspect, an embodiment of the present invention provides a data migration device, applicable to a management server. The device includes:

[0044] A first instruction sending unit, configured to send a metadata tree splitting instruction to a source master server to split a target metadata tree and obtain metadata sub-trees;

[0045] A second instruction sending unit, configured to send a first migration instruction to a source slave server to migrate a replicated sub-tree of the metadata sub-tree corresponding to the source slave server to a target server;

[0046] A third instruction sending unit, configured to send a master server transfer instruction to the source master server to update the target server as the target master server of the metadata sub-tree;

[0047] A fourth instruction sending unit, configured to send a second migration instruction to the source master server to migrate a replicated sub-tree of the metadata sub-tree corresponding to the source master server to a corresponding target slave server.

[0048] Seventh aspect, an embodiment of the present invention provides a data migration device, applicable to a target server. The device includes:

[0049] A sub-tree receiving unit, configured to receive a replicated sub-tree of a metadata sub-tree sent by a source slave server, where the metadata sub-tree is obtained by splitting a target metadata tree;

[0050] A master server transfer unit, configured to determine the target master server of the metadata sub-tree and receive and process data operation requests for the metadata sub-tree.

[0051] Eighth aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor. The memory is used to store one or more computer program instructions. Among them, the one or more computer program instructions are executed by the processor to implement the method according to any one of the first aspect to the third aspect.

[0052] Ninth aspect, an embodiment of the present invention provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the method according to any one of the first aspect to the third aspect is implemented.

[0053] The management server in the embodiment of the present invention sends a metadata tree splitting instruction to the source master server, so that the source master server splits the metadata tree to obtain a metadata subtree, generates configuration information of the metadata subtree, and sends a migration instruction to the source slave server to migrate the replication subtree of the metadata subtree of the source slave server to the target server. Furthermore, a master server transfer instruction is sent to the source master server to update the source master server to the source slave server of the metadata subtree, and the replication subtree of the metadata subtree is determined, and then a migration instruction is sent to the source master server, so that the source master server migrates the replication subtree to the target slave server of the metadata subtree. Thus, in the embodiment of the present invention, the target server is updated to the master server of the metadata subtree to process the data operation request of the client for the metadata subtree only after the metadata subtree of the source slave server is migrated to the target server, so as to reduce the negative impact on the access to the data storage system during the metadata migration process. Description of the Drawings

[0054] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0055] Figure 1 is a schematic diagram of the data storage system in the embodiment of the present invention;

[0056] Figure 2 is a schematic diagram of the metadata server cluster in the embodiment of the present invention;

[0057] Figure 3 is a flowchart of a data migration method in the embodiment of the present invention;

[0058] Figure 4 is a schematic diagram of the target metadata tree and the metadata subtree in the embodiment of the present invention;

[0059] Figure 5 is a data flow diagram of the data migration method in the embodiment of the present invention;

[0060] Figure 6 is a flowchart of another data migration method in the embodiment of the present invention;

[0061] Figure 7 is a schematic diagram of determining the first load state of the metadata tree according to multiple first load scores of the metadata tree within a predetermined time period in the embodiment of the present invention;

[0062] Figure 8 is a schematic diagram of the target metadata tree in the embodiment of the present invention;

[0063] Figure 9 is a schematic diagram of the data migration device in the embodiment of the present invention;

[0064] Figure 10It is a schematic diagram of the data migration device according to an embodiment of the present invention;

[0065] Figure 11 It is a schematic diagram of the data migration device according to an embodiment of the present invention;

[0066] Figure 12 It is a schematic diagram of the electronic device according to an embodiment of the present invention. Detailed implementation manners

[0067] The following describes the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. In order to avoid obscuring the essence of the present application, well-known methods, processes, procedures, components and circuits are not described in detail.

[0068] In addition, those of ordinary skill in the art should understand that the drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.

[0069] Unless the context clearly requires otherwise, words such as "including" and "comprising" in the entire application document should be interpreted as having an inclusive meaning rather than an exclusive or exhaustive meaning; that is, it is the meaning of "including but not limited to".

[0070] In the description of the present application, it should be understood that terms such as "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0071] This embodiment provides a data storage system, which provides a unified data storage entry for various computing applications in the data lake ecosystem, and integrates various storage protocols to support storage semantic fusion of various storage services such as Hadoop storage services (such as storage services based on the HDFS distributed file system and the HBase distributed NoSQL database), S3 (S3 Simple Storage Service), K8S CSI storage service, and Posix (Portable Operating System Interface of UNIX), so as to be applied to various data ecosystems and application scenarios, improve the performance of data access, and reduce the storage space required for data applications.

[0072] Figure 1 It is a schematic diagram of the data storage system according to an embodiment of the present invention. As Figure 1As shown, the data storage system 20 in this embodiment is connected to the service layer 10 and the storage layer 30. In this embodiment, the data storage system 20 performs semantic fusion on various storage services so that each service layer 10 can call an appropriate storage interface to operate on data files or the data therein (such as read, write, delete, etc.). Among them, the data storage system 20 includes a storage interface service module 21 with multiple storage interfaces, a metadata service module 22, and a configuration module (not shown in the figure). Optionally, the storage interface service module 21 of the data storage system 20 may include interface services such as Posix file interface, HDFS SDK, CSI, S3 interface, S2 interface, and image processing interface. It should be understood that this embodiment is not limited to the above storage interfaces, and other storage interfaces for implementing storage services can also be integrated into this embodiment.

[0073] In this embodiment, regardless of the type of storage protocol (such as S3 or file system, etc.), the data file consists of two parts: metadata and data. Among them, the metadata is stored in the metadata service, and the data is stored in the corresponding data storage space (such as GIFT DFS storage system, S3 storage system, OSS storage system, COS storage system, etc.).

[0074] Furthermore, in this embodiment, the metadata service module 22 is configured to store and manage the metadata of the file. Among them, the metadata of the file is stored in a data volume and stored in the form of a metadata tree or a metadata subtree.

[0075] The configuration module is configured to store and manage the configuration information of the metadata tree and the configuration information of the metadata subtree. Among them, the configuration information of the metadata tree includes at least one of the node change information (SubTreeEpoch) and the node range information (SubtreeTreeRange) of the metadata tree and the storage area identifier (Raft Group ID) of the metadata tree, and the configuration information of the metadata subtree includes at least one of the node change information and the node range information of the metadata subtree and the server group identifier of the metadata subtree.

[0076] In this embodiment, the metadata service module 22 is implemented by a metadata server cluster, and a metadata server (Meta Data Server, MDS) in the metadata server cluster is a node in the metadata server cluster. To facilitate clients to access the metadata of various files in the data storage system, each metadata server stores the corresponding metadata tree or metadata subtree in the local storage module or the remote storage module. As the number of files continues to grow, the scale of the metadata tree or metadata subtree also continues to grow. Moreover, there are differences in the load status information such as the access volume between the metadata trees or metadata subtrees, which results in differences in the loads of different metadata servers in the metadata server cluster, thus causing the phenomenon of load imbalance.

[0077] However, the existing method will cause the client to be unable to perform data operations on the migrated metadata (i.e., the metadata subtree to be migrated) during the metadata migration process, thus having a negative impact on the access to the data storage system.

[0078] Figure 2 It is a schematic diagram of the metadata server cluster according to an embodiment of the present invention. As Figure 2 shown, the metadata service module 22 includes two metadata server groups. One group is server 221, server 222, and server 223, and the other group is server 224, server 225, and server 226. In the embodiment of the present invention, taking the case where the load of the target metadata tree stored in server 221, server 222, and server 223 is too large and needs to be split to ensure the load balance of the metadata server cluster as an example, the migration process of the metadata subtree will be described.

[0079] Among them, server 221 is the primary metadata server (i.e., the primary node leader) in the group where the target metadata tree is located, that is, the source primary server, and server 222 and server 223 are the two secondary metadata servers (i.e., the secondary nodes follower) of server 221, that is, the source secondary servers; server 224, server 225, and server 226 are all target metadata servers corresponding to the target metadata tree. Among them, server 221, server 222, and server 223 store and manage the same metadata tree and the metadata in the metadata tree, and server 224, server 225, and server 226 store and manage the same metadata tree and the metadata in the metadata tree. Further, server 221 is configured to read / write the target metadata tree and the metadata in the target metadata tree, and server 222 and server 223 are responsible for reading the target metadata tree and the metadata in the target metadata tree. It is easy to understand that Figure 2 the structure of the metadata server cluster and the number of metadata servers shown are only illustrative.

[0080] To solve the above problems, the management server (masterserver) in the data management system according to the embodiments of the present invention sends a metadata tree splitting instruction to server 221. After receiving the metadata tree splitting instruction, server 221 splits the target metadata tree, obtains metadata subtrees, and generates configuration information of the metadata subtrees. After server 221 finishes splitting the target metadata tree, the management server sends a first migration instruction to server 222 and / or server 223. Server 222 and / or server 223 migrate the replica of the metadata subtree to server 224, server 225, or server 226 according to the received first migration instruction. After the replica migration is completed, the management server sends a master server transfer instruction to server 221. Server 221 is updated from the source master server of the metadata subtree to the source slave server, and at the same time, one of servers 224, 225, and 226 is updated to the target master server of the metadata subtree. Further, after the master server of the metadata subtree is transferred, server 221 determines the replica of the metadata subtree. Then, the management server sends a second migration instruction to server 221. Server 221 migrates the replica of the metadata subtree to the target slave server among servers 224, 225, and 226.

[0081] Thus, during the process of server 222 or server 223 migrating the replica of the metadata subtree, server 221 can still normally receive data operation requests from the client for the metadata subtree. After the replica migration is completed, the master server of the metadata subtree is updated to one of servers 224, 225, and 226 to process data operation requests from the client through the new master server. This method enables the data storage system to still normally receive access requests from the client during the migration of the metadata subtree, so that it can improve the load balance of the metadata server cluster while significantly reducing the negative impact on the access to the data storage system.

[0082] The following is described through method embodiments. Figure 3 It is a flowchart of a data migration method according to an embodiment of the present invention. As Figure 3 shown, the method of this embodiment includes the following steps:

[0083] Step S301, send a metadata tree splitting instruction to the source master server.

[0084] In this embodiment, the management server is used to maintain the load balancing of the metadata server cluster. Each server in the metadata cluster reports its heartbeat to the management server, and carries the load status information of each metadata tree managed by the server through the heartbeat, so that the management server can obtain the load status information of each metadata tree and thus determine the load status of each metadata tree.

[0085] Therefore, in an optional implementation manner, if the load status of the metadata tree (hereinafter also referred to as the first load status) indicates that the metadata tree is in an overloaded state, the management server will send a metadata tree splitting instruction to the source server group corresponding to the overloaded metadata tree.

[0086] In another optional implementation manner, the administrator can control the load status of the metadata tree manually. Specifically, the administrator can send a metadata tree splitting instruction to the management server through the client. After receiving the metadata tree splitting instruction, the management server will forward the metadata tree splitting instruction to the source server group of the target metadata tree corresponding to the metadata tree splitting instruction.

[0087] Optionally, according to actual needs, the management server can also send a metadata tree splitting instruction to the source master server in other cases, and this embodiment does not make specific restrictions. For example, the management server can determine the load status of each metadata tree according to a predetermined period, and send a metadata tree splitting instruction to the source server group where the metadata tree with the largest load status ranking is located.

[0088] Step S302: Split the target metadata tree to obtain metadata subtrees.

[0089] After receiving the metadata tree splitting instruction, the source master server will determine the metadata tree corresponding to the metadata tree splitting instruction as the target metadata tree, and split the target metadata tree to obtain at least one metadata subtree.

[0090] The metadata tree splitting instruction carries the splitting method of the target metadata tree, which may specifically include the index identifiers of at least one node. Therefore, optionally, after receiving the metadata tree splitting instruction, the source master server can split the metadata subtree with the node as the root node from the target metadata tree corresponding to the metadata tree splitting instruction according to the index identifiers of each node.

[0091] In this embodiment, only the source master server can perform data modification operations such as writing / deleting on the metadata tree. Therefore, before splitting the target metadata tree, the source server in the source server group can check whether it is the source master server. If it is the source master server, the source master server will split the target metadata tree to obtain the metadata subtree; if it is the source slave server, the source slave server will forward the metadata tree splitting instruction to the source master server, enabling the source master server to split the target data tree to obtain the metadata subtree.

[0092] Meanwhile, to prevent the target metadata tree from being split in the same way or in different ways multiple times, the source master server can check whether the node change information in the configuration information of the target metadata tree has changed. In this embodiment, the node change information of the target metadata tree may include the node addition / deletion information (conf_version) and subtree splitting information (stru_version) of the target metadata tree. Specifically, when each node is added or deleted in the target metadata tree, the value of the node addition / deletion information of the target metadata tree will increase by 1; when each metadata subtree is split from the target metadata tree, the subtree splitting information of the target metadata tree increases by 1.

[0093] Step S303: Generate the configuration information of the metadata subtree.

[0094] After obtaining the metadata subtree, the source master server will generate the configuration information of the metadata subtree. The configuration information of the metadata subtree includes at least one of the node update information and node range information of the metadata subtree.

[0095] Among them, the node change information of the metadata subtree is the same as the node change information of the target metadata tree, including the node addition / deletion information and subtree splitting information of the metadata subtree; the node range information of the metadata subtree includes the included subtree information (include_tree) and excluded subtree information (exclude_tree) of the metadata subtree. The included subtree information is the index identifier (inode) of the root node of the metadata subtree, and the excluded subtree information is the index identifier of the root node of the metadata subtree split from the metadata subtree.

[0096] Figure 4 is a schematic diagram of the target metadata tree and metadata subtree of the embodiment of the present invention. As Figure 4As shown in the figure, the metadata tree 40 is the target metadata tree, and the metadata tree includes nodes 41 (i.e., the root node of the metadata tree 40) and node 42. The source master server splits the metadata tree 40 to obtain a metadata subtree, that is, the metadata tree 43. The root node of the metadata tree 43 is node 42. The source server generates configuration information of the metadata subtree, including the included subtree information and the non-included subtree information of the metadata tree 43. The included subtree information is the index identifier of node 42, and the non-included subtree information is null.

[0097] When splitting the target metadata tree into a metadata subtree, the initial node addition and deletion information and the initial subtree splitting information of the metadata subtree are usually 0. The initial included subtree information is the index identifier of the root node of the metadata subtree, and the initial non-included subtree information is usually a null value (null).

[0098] Step S304: Send a first migration instruction to the source slave server to migrate the replicated subtree of the metadata subtree corresponding to the source slave server to the target server.

[0099] The target server group of the metadata subtree is allocated by the management server according to the load status of each server group. Therefore, the management server can determine the first migration instruction according to the storage area identifier of the corresponding storage area of the target server group and send the first migration instruction to the source slave server.

[0100] In the same server group, the slave server will keep data synchronized with the master server. Therefore, after the source master server splits the target metadata tree to obtain at least one metadata subtree, the source slave server can synchronously obtain the metadata subtree.

[0101] In order to enable the source master server to still normally process the data operations of the client on the metadata subtree during the migration process of the metadata subtree, the source slave server will generate a replicated subtree of the metadata subtree, and after receiving the first migration instruction sent by the management server, migrate the replicated subtree to any target server in the target server group according to the storage area identifier corresponding to the target server group, so that the target server stores the replicated subtree of the metadata subtree sent by the source slave server in the corresponding storage area.

[0102] Take Figure 2 the metadata server cluster shown in the figure as an example. After server 221 splits the target metadata tree to obtain a metadata subtree, servers 222 and 223 can synchronously obtain the metadata subtree and generate a replicated subtree of the metadata subtree. After server 222 receives the first migration instruction sent by the management server, it migrates the replicated subtree to server 224, server 225, or server 226.

[0103] Step S305: Send a master server transfer instruction to the source master server.

[0104] After the source slave server successfully migrates the replication subtree of the metadata subtree to the target server, the management server sends a master server migration instruction to the source master server to update a target server as the target master server of the metadata subtree.

[0105] Step S306, update to the source slave server of the metadata subtree.

[0106] After receiving the master server transfer instruction sent by the management server, the source master server updates itself to the source slave server of the metadata subtree.

[0107] Step S307, determine the replication subtree of the metadata subtree.

[0108] During the process of the source slave server migrating the replication subtree of the metadata subtree to the target server, the source master server remains the master server of the metadata subtree. Therefore, the data operation requests of the client for the metadata subtree are still processed by the source master server, and the source master server will modify the metadata in the metadata subtree according to the data operation requests.

[0109] Therefore, after the source master server is updated to the source slave server of the metadata subtree, the source master server will stop processing the data operation requests for the metadata subtree and generate an updated replication subtree of the metadata subtree to migrate the updated metadata subtree to the target server.

[0110] Step S308, send a second migration instruction to the source master server.

[0111] After updating the target server of the metadata subtree to the target master server, the management server determines the second migration instruction according to the storage area identifier corresponding to the target server group and sends the second migration instruction to the source master server, so that the source master server migrates the replication subtree to the target slave server.

[0112] Step S309, migrate the replication subtree to the target slave server corresponding to the metadata subtree.

[0113] After receiving the second migration instruction, the source master server migrates its own replication subtree to the target slave server corresponding to the metadata subtree according to the storage area identifier corresponding to the target server group, so that the target server stores the replication subtree of the metadata subtree sent by the source master server in the corresponding storage area.

[0114] Figure 5 It is the data flow diagram of the data migration method of the embodiment of the present invention. As Figure 5As shown, a management server (not shown in the figure) sends a metadata tree splitting instruction to server 221 to obtain at least one metadata subtree of the target metadata tree. Metadata tree 51 is the metadata subtree obtained by server 221 splitting the target metadata tree. The management server sends a first migration instruction to server 222. Server 222 generates a replicated subtree 52 of metadata tree 51 and migrates the replicated subtree 52 to server 225. At the same time, the management server sends a first migration instruction to server 223. Server 223 generates a replicated subtree 53 of metadata tree 51 and migrates the replicated subtree 53 to server 226. During the migration process of replicated subtree 52 and replicated subtree 53, server 221 remains the source master server of metadata tree 51, processes the data operation requests of the client for metadata tree 51, and updates metadata tree 51 to metadata tree 54 according to the data operation requests.

[0115] After the migration of replicated subtree 52 and replicated subtree 53 is completed, the management server sends a master server transfer instruction to server 221, causing server 221 to be updated to the source slave server of metadata tree 51, and updating server 225 from the target server of metadata tree 51 to the target master server to process the data operation requests sent by the client. After server 221 updates itself to the source slave server of metadata tree 51, it generates a replicated subtree 55 of metadata tree 54. After the master server transfer is successful, the management server sends a second migration instruction to server 221. After receiving the second migration instruction, server 221 migrates the replicated subtree 55 to server 224 to synchronize the updated replicated subtree 52 and replicated subtree 55 in the target server group composed of server 224, server 225, and server 226, completing the migration process of the metadata subtree.

[0116] The management server in the embodiment of the present invention sends a metadata tree splitting instruction to the source master server, so that the source master server splits the metadata tree to obtain a metadata subtree, generates configuration information of the metadata subtree, and sends a migration instruction to the source slave server to migrate the replicated subtree of the metadata subtree of the source slave server to the target server, and then sends a master server transfer instruction to the source master server to update the source master server to the source slave server of the metadata subtree, and determines the replicated subtree of the metadata subtree, thereby sending a migration instruction to the source master server to cause the source master server to migrate the replicated subtree to the target slave server of the metadata subtree.

[0117] In an embodiment of the present invention, during the process of the source slave server migrating the replication subtree of the metadata subtree, the source master server can still normally receive data operation requests from the client for the metadata subtree. After the replication subtree of the source slave server is migrated, the target server of the metadata subtree is updated to the target master server of the metadata subtree, and the source master server is updated to the source slave server of the metadata subtree. The source slave server generates a replication subtree of the updated metadata subtree and migrates the replication subtree to the target server to synchronize data for the metadata subtree in the target server group and process operation requests sent by the client. Therefore, this method enables the data storage system to still normally receive access requests from the client during the migration process of the metadata subtree, thereby significantly reducing the negative impact on the access to the data storage system while improving the load balancing of the metadata server cluster.

[0118] Figure 6 is a flowchart of another data migration method according to an embodiment of the present invention. As Figure 6 shown, the method of this embodiment includes the following steps:

[0119] Step S601, send a metadata tree splitting instruction to the source master server.

[0120] In this embodiment, the management server can obtain multiple load status information of each metadata tree within a predetermined time period based on the heartbeats reported by each metadata server.

[0121] The load status information of the metadata tree is used to reflect the current load situation of the metadata tree, and specifically may include at least one of the storage resource consumption amount and operation request parameters of the metadata tree. Among them, the storage resource consumption amount represents the size of the storage space occupied by the metadata tree, and can be specifically determined according to the data volume of the metadata written into the metadata tree and the data volume of the metadata deleted from the metadata; the operation request parameters include the request quantities of read / write / delete and other operation requests received by the metadata tree, and can specifically be the queries-per-second (QPS) of the operation requests, that is, the number of operation requests per 1 second, and may also include the access latency of a single read / write / delete and other operation requests received by the metadata tree.

[0122] The greater the storage resource consumption of the metadata tree, the larger the amount of metadata in the metadata tree, which makes the time for the metadata server group to respond to the client operation request to find the corresponding metadata longer, and the time to return the result is also longer accordingly. The operation requests of the client to the metadata server group are usually stored in the request queue and processed in the order of being stored in the request queue. Therefore, the larger the operation request parameter of the metadata tree, the longer the response time of the metadata server group to the operation request, and the longer the time to return the result. It is easy to understand that the load status information can also include other information, such as the usage rate of the metadata tree on the Central Processing Unit (CPU), etc., which is not specifically limited in this embodiment.

[0123] The write operation of the client to the metadata will be written into a predetermined database (such as RocksDB) in the form of key-value pairs. Therefore, the management server can determine the storage resource consumption of the metadata tree according to the amount of data of the values written into / removed from the metadata tree. The key-value pairs written into the metadata tree will increase the storage resource consumption of the metadata tree, while the key-value pairs removed from the metadata tree will reduce the storage resource consumption of the metadata tree. Therefore, for each metadata tree, the management server can calculate the difference between the sum of the amounts of data of the key-value pairs written into the metadata tree and the sum of the amounts of data of the key-value pairs removed from the metadata tree to obtain the storage resource consumption of the metadata tree.

[0124] After obtaining multiple load status information of each metadata tree within a predetermined time period, the management server will determine the metadata trees whose multiple load status information meets the predetermined load conditions as target metadata trees.

[0125] Specifically, for each metadata tree, the management server will determine the first load status of the metadata tree according to the multiple load status information of the metadata tree, and when the first load status of the metadata tree indicates that the metadata tree is overloaded, the metadata tree will be determined as the target metadata tree.

[0126] Furthermore, the management server can determine the first load score corresponding to each load status information according to the multiple load status information of each metadata tree within a predetermined time period. If there are a first number of first load scores among the multiple first load scores within the predetermined time period that meet the first condition, the management server will determine the first load status of the metadata tree as overloaded.

[0127] In an alternative implementation, the storage resource consumption and the priority of operation request parameters can be preset, where the priority indicates the importance of the storage resource consumption and the operation request parameters. The management server can determine the first load score of the server group based on the load status information with a higher priority. If there are server groups with the same first load score, the management server can further update the first load score based on the load status information with a lower priority.

[0128] Taking the case where the priority of the storage resource consumption is higher than that of the operation request parameters as an example, the management server will determine the first load score of the metadata tree based on the storage resource consumption of the metadata tree. If there are multiple metadata trees with the same storage resource consumption, the management server will simultaneously determine the first load score of the metadata tree based on the storage resource consumption and the operation request parameters of the metadata tree.

[0129] In another alternative implementation, the management server can determine the first load score of the metadata tree only based on the storage resource consumption of the metadata tree. For example, if the storage resource consumption of metadata tree 1 is 1 GB, the management server can convert the measurement unit of this resource consumption to bytes, obtaining that the storage resource consumption of metadata tree 1 is 1073741824 bytes, and determine the first load score of metadata tree 1 as 1073741824. Another example, if the storage resource consumption of metadata tree 1 is 1 GB, the management server can convert the measurement unit of this resource consumption to bytes, obtaining that the storage resource consumption of metadata tree 1 is 1073741824 bytes, and multiply the storage resource consumption of metadata tree 1 by 0.0000001 (i.e., the predetermined coefficient) and take the integer part to obtain that the first load score of metadata tree 1 is 107 points. Another example, if the resource storage consumption of metadata tree 1 is 1 GB, the management server can directly determine this resource consumption as the first load score of metadata tree 1, that is, the first load score of metadata tree 1 is 1.

[0130] In another alternative implementation, the management server can determine the first load score of the metadata tree only based on the operation request parameters of the metadata tree. Specifically, the management server can determine it based on the operation request parameters of the metadata tree and the corresponding operation request scores.

[0131] Different operation requests sent by the client to the data storage system will cause different degrees of load on the metadata server group. Therefore, in this embodiment, different operation request categories correspond to different operation request scores. For example, for operation request categories such as StatFS, Open, Close, Access, GetAttr, Resolve, ReadLink, GetXattr, ListXattr, and RemoveXAttr, the corresponding operation request scores can be set to 1 point; for operation request categories such as Lookup, Read, NextSlice, and NextINode, the corresponding operation request scores can be set to 2 points; for operation request categories such as Write, Create, MKnod, Mkdir, Rename, SetAttr, Rmdir, Unlink, Truncate, Fallocate, Flock, SetLlk, Link, Symlink, AppendFile, CommitCompact, NewSession, and SessionHeartbeat, the corresponding operation request scores can be set to 6 points; for operation request categories such as ReadDir and CommitAppend, the corresponding operation request scores can be set to 10 points.

[0132] For each metadata tree, after determining the product of the number of operation requests generated by each operation request category and / or the access latency caused by the operation request and the operation request score corresponding to the operation request category, the management server can determine the sum of the above products as the first load score corresponding to the metadata tree.

[0133] In another alternative implementation, the management server can determine the first load score of the metadata tree according to the storage resource consumption of the metadata tree and the operation request parameters. Specifically, for each metadata tree, the management server can determine the first score according to the storage resource consumption of the metadata tree and determine the second score according to the operation request parameters, so as to determine the first load score of the metadata tree according to the first score and the second score. Among them, the determination methods of the first score and the second score are similar to the method of determining the first load score in the above alternative implementation, and will not be elaborated here. Optionally, the first score and the second score can be set with the same or different weights respectively, and the management server can calculate the weighted sum of the first score and the second score of the metadata tree to determine the first load score of the metadata tree.

[0134] Further, the first score and the weight of the first score can be determined according to the load conditions caused by the storage resource consumption and operation request parameters on the metadata server cluster in actual applications. A larger weight is set for the load status information with a greater impact on the load, such as 0.6, and a smaller weight is set for the load status information with a smaller impact on the load, such as 0.4.

[0135] To prevent large short-term load fluctuations from having a greater impact on the load status evaluation of the metadata tree, after determining the first load scores of the same metadata tree at different times within a predetermined duration, the management server will determine the metadata tree as the target metadata tree when there are a first number of first load scores among these multiple first load scores that meet the first condition.

[0136] In this embodiment, both the first number and the first condition can be determined according to the actual load conditions of the metadata server cluster. For example, if the metadata tree has a high update frequency and the server cluster needs to perform load balancing every 60 minutes, the predetermined duration can be set to 30 minutes, 45 minutes, etc.; if the number of first load scores in the predetermined duration is 30, the first number can be set to 10, 15, etc.; when the first load score is not higher than 10000000 points, the metadata tree is not overloaded, then the first condition can be set to the first load score not being higher than 10000000 points.

[0137] Optionally, the first condition can be set to the first load score ≥ a predetermined load parameter and the load score ≥ the minimum value of all load scores, where the load score of the metadata tree is the median of the first load scores in the sliding window sequence.

[0138] The frequency at which the server reports heartbeats is usually fixed. Therefore, within the same time period, the number of heartbeats reported by the server is the same. In an alternative implementation, for any metadata tree, when the management server determines that the first load score of the metadata tree meets the first condition, this first load score can be added to a sliding window queue of a predetermined length, and it is determined that the consecutive (n - 1) first load scores after this first load score of the metadata tree are added to the sliding window queue. Then, these n first load scores are counted one by one from the 1st to the nth. If the i-th (1 ≤ i ≤ n) first load score meets the first condition, the management server increments the hot degree value of the metadata tree by 1 and sets the counter value of the server to 0. If the i-th first load score does not meet the first condition, the management server decrements the hot degree value of the metadata tree by 1 and increments the counter value of the metadata tree by 1. When the counter value of the metadata tree is not less than the minimum elimination parameter, that is, when there are a first number of first load scores among the multiple first load scores corresponding to the metadata tree that do not meet the first condition, the management server can determine that the first load status of the metadata tree is not overloaded. If the hot degree value of the metadata tree is not less than the minimum hot degree parameter and the load score is not less than the minimum of all load scores, the management server can determine that the first load status of the metadata tree is overloaded. Here, n is the number of load status information of the server obtained within the predetermined time period.

[0139] Figure 7 It is a schematic diagram for determining the first load status of a metadata tree according to multiple first load scores of the metadata tree within a predetermined time period in an embodiment of the present invention. Figure 7 The first load score shown is the first load status determined according to the load status information reported at different times of the same metadata tree within a predetermined time period. As Figure 7 As shown, the management server sequentially stores the n first load scores of metadata tree 1, that is, the first load score 1 - the first load score n, into the sliding window queue 70. Among them, the window 71 in the sliding window queue 70 stores the first load score 1, the window 72 stores the first load score 2,..., and the window 7n stores the first load score n. When the management server determines that the first load score 1 does not meet the first condition, the first load score 1 is eliminated. At this time, the window 71 is emptied and changes from the first window in the sliding window queue 70 to the last window in the sliding window queue 70. The management server will obtain the (n + 1)-th first load score of metadata tree 1, that is, the first load score (n + 1), and store the first load score (n + 1) in the window 71 to continue determining the first load status of metadata tree 1.

[0140] In this embodiment, the management server determines the load scores of each server in the server cluster according to the multiple load status information corresponding to each metadata tree, and determines the second load status of the corresponding server group according to the load scores of each server (that is, the sum of the load scores of each server in the same server group). Then, according to the second load status of each server group, the target status corresponding to each metadata subtree and the target server group corresponding to each metadata subtree are determined, so as to migrate each metadata subtree to the corresponding target server group. In this embodiment, the target server group can be an existing metadata server group in the metadata server cluster or a newly added server group in the metadata server cluster, and this embodiment does not make specific limitations.

[0141] In an alternative implementation, for each server group, the management server determines multiple second load scores of the server group within a predetermined time period according to the load status information of each metadata tree corresponding to the server group, and when there are a second number of second load scores among the multiple second load scores that meet the second condition, determines the second load status of the server group as overloaded.

[0142] Specifically, the management server can determine the sum of the first load scores of each metadata tree corresponding to the same server group at the same moment as the second load score of the server group at that moment. The management server can also determine the second load status of the server group by counting the multiple second load scores corresponding to each server group in the sliding window sequence. Specifically, it can refer to the determination method of the first load status of the above metadata tree and will not be elaborated here.

[0143] Optionally, for each server group, the management server determines the second condition according to the average value of the first load scores of each metadata tree corresponding to the server group and the minimum balance parameter of the metadata tree. The minimum balance parameter is the minimum value at which the metadata tree can maintain an equilibrium state. Specifically, the management server can set the second condition as the second load score ≤ the average value of the second load scores of each metadata tree * (1 + the minimum balance parameter). Alternatively, similar to the first condition and the first number, the second condition and the second number can also be determined according to the actual load situation of the metadata server cluster.

[0144] After determining the second load status of each server group, the management server determines the server groups with the second load status indicating not overloaded as candidate server groups, then determines the expected migration amount of the target metadata tree and the expected reception amount of the candidate server groups, and further determines at least one metadata subtree corresponding to the target metadata tree and the target server group corresponding to each metadata subtree according to the expected migration amount of each target metadata tree and the expected reception amount of the candidate server groups, so as to migrate each metadata subtree to the corresponding target server group and improve the load balance of the metadata server cluster.

[0145] Specifically, the management server can obtain the first balance parameter of each target metadata tree and the second balance parameter of each candidate server group, and determine the expected migration amount of the target metadata tree according to the first load score and the first balance parameter of each target metadata tree, and determine the expected reception amount of the candidate server group according to the second load score and the second balance parameter of each candidate server group. Among them, the first balance parameter is the load status score corresponding to the target metadata tree when it reaches the balanced state, which can specifically be the maximum load status score, and the second balance parameter is the load status score corresponding to the candidate server group when it reaches the balanced state, which can specifically be the maximum load status score.

[0146] The management server can determine the expected migration amount of the target metadata tree according to the difference between the first load score and the first balance parameter of each target metadata tree. For the candidate server group including the target metadata tree, the management server can determine the expected reception amount of the candidate server group according to the expected migration amounts of the target metadata trees corresponding to the candidate server group and the difference between the second load score and the second balance parameter of the candidate server group; for the candidate server group not including the target metadata tree, the management server can determine the expected reception amount of the candidate server group according to the difference between the second load score and the second balance parameter of the candidate server group.

[0147] For example, the target metadata trees include metadata tree T1 and metadata tree T2, and the candidate server groups include server group MDS1, server group MDS2, and server group MDS3, where server group MDS2 includes metadata tree T1. The first load scores of metadata tree T1 and metadata tree T2 are the storage resource consumption amounts of the corresponding target metadata trees. Among them, the first load score of metadata tree T1 is 580 (GB), and the first load score of metadata tree T2 is 620 (GB). The second load scores of server group MDS1, server group MDS2, and server group MDS3 are the sum of the storage resource consumption amounts of the corresponding metadata trees. Among them, the second load score of server group MDS1 is 1860 (GB), the second load score of server group MDS2 is 2010 (GB), and the second load score of server group MDS3 is 1980 (GB). The second balance parameter of the candidate server group is 2048 (GB), and the first balance parameter of the target metadata tree is 500 GB.

[0148] The management server determines that the expected migration volume of metadata tree T1 is 80 (GB) based on the difference between the first load score of metadata tree T1 and the first balancing parameter, and the expected migration volume of metadata tree T2 is 120 (GB) based on the difference between the first load score of metadata tree T2 and the first balancing parameter. The expected receiving volume of server group MDS1 is determined to be 188 (GB) according to the difference between the second load score of server group MDS1 and the second balancing parameter. The expected receiving volume of server group MDS2 is determined to be 2048 - (2010 - 80) = 118 (GB) according to the expected migration volume of metadata tree T1 and the difference between the second load score of server group MDS2 and the second balancing parameter. The expected receiving volume of server group MDS3 is determined to be 68 (GB) according to the difference between the second load score of server group MDS3 and the second balancing parameter.

[0149] After determining the expected migration volume corresponding to each target metadata tree and the expected receiving volume corresponding to each candidate server group, the management server matches the target metadata tree and the candidate server groups according to the expected migration volume and the expected receiving volume, and then determines at least one candidate server group that matches the target metadata tree as the target server group corresponding to the metadata subtree of the target metadata tree. Then, according to the expected receiving volume of each target server group, the metadata subtree corresponding to each target server group is determined.

[0150] Figure 8 It is a schematic diagram of the target metadata tree of the embodiment of the present invention. Figure 8 The expected migration volume of the metadata tree 80 shown is 100 (GB). The management server matches the expected migration volume of the metadata tree 80 with the expected migration volumes of each candidate server group in the metadata server cluster, and determines that the candidate server groups that match the metadata tree 80 are server group S1 and server group S2 respectively. Among them, the expected receiving volume of server group S1 is 60 (GB), and the expected receiving volume of server group S2 is 40 (GB). The management server determines that the load score of the metadata subtree t1 with node 81 as the root node is 60 (GB) and the load score of the metadata subtree t2 with node 82 as the root node is 40 (GB) by calculating the load scores of each metadata node in the metadata tree 80. Therefore, it can be determined that server group S1 is the target server group of metadata subtree t1, and server group S2 is the target server group of metadata subtree t2.

[0151] After determining each metadata subtree and the target server group of each metadata subtree, the management server determines a metadata tree splitting instruction (SplitSubTree) according to the index identifier of the root node of the metadata subtree, and sends the metadata tree splitting instruction to the source server group of the target metadata tree.

[0152] Step S602: Split the target metadata tree to obtain metadata subtrees.

[0153] To facilitate the processing of data operation requests from clients, the server group usually caches the metadata tree in the local storage module. Therefore, after the source master server splits the metadata subtree from the target metadata tree, both the source master server and the source slave server will cache the metadata subtree in the local storage module.

[0154] Step S603: Send a server group acquisition request to obtain the storage area identifier of the metadata subtree.

[0155] After splitting at least one metadata subtree from the target metadata tree, the source master server will send a server group acquisition request to the management server to obtain the storage area identifiers of the target servers corresponding to each metadata subtree. The same server group implements metadata consistency based on the distributed consistency algorithm, and both the metadata tree and the metadata subtree are stored in the data volume. Therefore, in this embodiment, the storage area identifier can be the identifier of the data volume used to store the metadata subtree.

[0156] Step S604: Generate the configuration information of the metadata subtree.

[0157] After the source master server obtains the storage area identifiers of each metadata subtree fed back by the management server, it generates the configuration information of each metadata subtree according to the storage area identifiers.

[0158] Optionally, after generating the configuration information of each metadata subtree, the source master server will also send the configuration information of the metadata subtree to the configuration module of the data storage system (such as the root server RootServer) so that the configuration module stores the configuration information of the metadata subtree.

[0159] Step S605: Modify the configuration information of the target metadata tree.

[0160] After splitting at least one metadata subtree from the target metadata tree, the configuration information of the target metadata tree will also change synchronously. Therefore, the source master server will modify the configuration information of the target metadata tree. In this embodiment, in addition to the node change information, the configuration information of the target metadata tree can also include the node range information of the target metadata tree. Similar to the node range information of the metadata subtree, the node range information of the target metadata tree can also include the included subtree information and the non-included subtree information of the target metadata tree.

[0161] Still taking Figure 4Taking the target metadata tree and metadata subtree shown as an example for illustration. The configuration information of the metadata tree 40 before splitting includes included subtree information and non-included subtree information. Among them, the non-included subtree information is the index identifier of the node 41, and the non-included subtree information is null. After the source master server splits the metadata tree 40 to obtain the metadata tree 43, it will modify the non-included subtree information of the metadata tree 40 to the index identifier of the node 42.

[0162] Optionally, the source master server may send configuration update information of the target metadata subtree to the configuration module to modify the configuration information of the target metadata tree stored in the configuration module.

[0163] It is easy to understand that in this embodiment, step S605 and step S603 may be executed simultaneously or sequentially, and this embodiment does not make any restrictions.

[0164] Step S606, send a first migration instruction to the source slave server to migrate the replication subtree of the metadata subtree corresponding to the source slave server to the target server.

[0165] In this embodiment, the management server will send a subtree removal (RemovePeer) instruction to the source slave server and a subtree addition (AddPeer) instruction to the target server group to create a replication subtree of the metadata subtree in the target server group and remove the replication subtree of the metadata subtree in the source slave server, reducing the load on the source slave server.

[0166] Step S607, receive the replication subtree of the metadata subtree sent by the source slave server.

[0167] After receiving the subtree addition instruction sent by the management server, the target server will receive the replication subtree of the metadata subtree sent by the source slave server and store the replication subtree in the local storage module or the remote storage module.

[0168] Step S608, send a master server transfer instruction to the source master server.

[0169] In this embodiment, the source slave server will feedback the migration result of the replication subtree to the management server. When the migration result indicates that the replication subtree migration is successful, the management server will send a master server migration (TransferLeader) instruction to the source master server to update the target server as the target master server of the metadata subtree.

[0170] Step S609, determine the target master server of the metadata subtree and receive a data operation request for processing the metadata subtree.

[0171] After receiving the main server transfer instruction sent by the management server, the source main server sends the main server transfer instruction to the target server. After receiving the main server transfer instruction, each target server in the target server group elects the target main server of the metadata subtree based on the main server election mechanism and sends a reply message to the source main server.

[0172] Specifically, each target server can receive the election requests of other target servers. The election requests carry the configuration information of the corresponding target servers, including priority, metadata version, etc. Then, it compares the configuration information of each target server with its own configuration information and sends a reply message carrying its own voting result to other target servers and the source main server to vote for the target server with the highest configuration. Furthermore, the target server determines the target main server according to the election result feedback by the source main server and transfers the main server of the metadata subtree to this target main server.

[0173] After transferring the main server of the metadata subtree to the target server, other target servers will automatically change to target slave servers. The target main server and each target slave server will receive and process the data processing requests of the client for the metadata subtree. Specifically, when the target server is the target main server, it will process various data operation requests such as read / write / delete of the client for the metadata subtree; when the target server is the target slave server, it will process the read operation requests of the client for the metadata subtree and forward other types of requests to the target main server.

[0174] Step S610, update to the source slave server of the metadata subtree.

[0175] In this embodiment, after receiving the main server transfer instruction sent by the management server, the source main server sends the main server transfer instruction to the target server group, and after receiving the reply message of the target server group, determines the target main server of the metadata subtree with the highest number of votes according to the reply message, then updates itself to the source slave server of the metadata subtree, and sends the election result of the target server group to the management server and the target server group to update the target server to the corresponding main server of the metadata subtree.

[0176] Step S611, determine the replication subtree of the metadata subtree.

[0177] In this embodiment, the implementation manner of step S611 is similar to that of step S307 and will not be elaborated here.

[0178] Step S612, send a second migration instruction to the source main server.

[0179] In this embodiment, the implementation manner of step S612 is similar to that of step S308, and will not be elaborated here.

[0180] Step S613: Migrate the replicated subtree to the target slave server corresponding to the metadata subtree.

[0181] In this embodiment, the implementation manner of step S613 is similar to that of step S309, and will not be elaborated here.

[0182] Step S614: Delete the metadata subtree in response to receiving the removal instruction.

[0183] After migrating the replicated subtree of its own metadata subtree to the target slave server, the source master server will feedback the migration result to the management server. After the migration result indicates that the replicated subtree migration is successful, the management server sends a removal instruction to the source master server. After receiving the removal instruction, the source master server will delete the metadata subtree in the local storage module and the remote storage module to reduce the load on itself and the server group.

[0184] Step S615: Perform data synchronization.

[0185] To avoid the possibility that the metadata of the metadata subtree migrated from the source master server is inconsistent with the metadata in the metadata subtree stored in the target master server, data synchronization will be performed between the target servers to update the metadata subtree.

[0186] The management server in the embodiment of the present invention sends a metadata tree splitting instruction to the source master server so that the source master server splits the metadata tree to obtain the metadata subtree. The source master server sends a server group acquisition request to the management server and generates the configuration information of the metadata subtree, and at the same time modifies the configuration information of the target metadata tree. After the target metadata tree is successfully split, the management server sends a migration instruction to the source slave server to migrate the replicated subtree of the metadata subtree of the source slave server to the target server, and after the replicated subtree of the source slave server is successfully migrated, sends a master server transfer instruction to the source master server. The target server determines the target master server of the metadata subtree to update the master server of the metadata subtree, and at the same time the source master server updates itself to the source slave server of the metadata subtree and determines the replicated subtree of the metadata subtree. After the master server change of the metadata subtree is successful, the target server starts to receive and process the data processing requests for the metadata subtree, and the management server will send a migration instruction to the source master server to make the source master server migrate the replicated subtree to the target slave server of the metadata subtree and delete the metadata subtree stored in the source master server.

[0187] In an embodiment of the present invention, during the process of the source slave server migrating the replication subtree of the metadata subtree, the source master server can still normally receive data operation requests from the client for the metadata subtree. After the replication subtree of the source slave server is migrated, the target server of the metadata subtree will be updated to the target master server of the metadata subtree, and the source master server will be updated to the source slave server of the metadata subtree. The source slave server generates a replication subtree of the updated metadata subtree and migrates the replication subtree to the target server to synchronize data for the metadata subtree in the target server group and process operation requests sent by the client. Therefore, this method enables the data storage system to still normally receive access requests from the client during the migration process of the metadata subtree, thereby significantly reducing the negative impact on the access to the data storage system while improving the load balancing of the metadata server cluster.

[0188] Figure 9 is a schematic diagram of the data migration device according to an embodiment of the present invention. As Figure 9 shown, the data migration device according to an embodiment of the present invention is applicable to the source master server and includes a subtree generation unit 901, a configuration information generation unit 902, a server update unit 903, a subtree replication unit 904, and a migration unit 905.

[0189] Among them, the subtree generation unit 901 is configured to split the target metadata tree in response to receiving a metadata tree splitting instruction to obtain a metadata subtree. The configuration information generation unit 902 is configured to generate configuration information for the metadata subtree. The server update unit 903 is configured to update to the source slave server of the metadata subtree in response to receiving a master server transfer instruction. The subtree replication unit 904 is configured to determine the replication subtree of the metadata subtree. The migration unit 905 is configured to migrate the replication subtree to the target slave server corresponding to the metadata subtree in response to receiving a second migration instruction.

[0190] In an optional implementation manner, the data migration device further includes a subtree deletion unit.

[0191] Among them, the subtree deletion unit is configured to delete the metadata subtree in response to receiving a removal instruction.

[0192] In an optional implementation manner, the data migration device further includes a request sending unit.

[0193] Among them, the request sending unit is configured to send a server group acquisition request to obtain the storage area identifier of the metadata subtree. The configuration information generation unit 902 is further configured to generate the configuration information according to the received storage area identifier.

[0194] In an alternative implementation, the configuration information of the metadata subtree includes at least one of the node change information and the node range information of the metadata subtree and the storage area identifier of the metadata subtree.

[0195] After receiving the metadata tree splitting instruction sent by the management server, the source master server in the embodiment of the present invention splits the metadata tree to obtain a metadata subtree, generates the configuration information of the metadata subtree, and updates itself to the source slave server of the metadata subtree after receiving the master server transfer instruction sent by the management server, and determines the replication subtree of the metadata subtree. Further, after receiving the master server transfer instruction sent by the management server, the source master server migrates the replication subtree to the target slave server of the metadata subtree. Thus, in the embodiment of the present invention, the target server is updated to the master server of the metadata subtree to process the data operation request of the client for the metadata subtree only after the metadata subtree of the source slave server is migrated to the target server, so as to reduce the negative impact on the access to the data storage system during the metadata migration process.

[0196] Figure 10 is a schematic diagram of the data migration device in the embodiment of the present invention. As Figure 10 shown, the data migration device in the embodiment of the present invention is applicable to a management server and includes a first instruction sending unit 1001, a second instruction sending unit 1002, a third instruction sending unit 1003, and a fourth instruction sending unit 1004.

[0197] Among them, the first instruction sending unit 1001 is used to send a metadata tree splitting instruction to the source master server to split the target metadata tree to obtain a metadata subtree. The second instruction sending unit 1002 is used to send a first migration instruction to the source slave server to migrate the replication subtree of the metadata subtree corresponding to the source slave server to the target server. The third instruction sending unit 1003 is used to send a master server transfer instruction to the source master server to update the target server to the target master server of the metadata subtree. The fourth instruction sending unit 1004 is used to send a second migration instruction to the source master server to migrate the replication subtree of the metadata subtree corresponding to the source master server to the corresponding target slave server.

[0198] In an alternative implementation, the third instruction sending unit 1003 includes a result obtaining subunit and a third instruction sending subunit.

[0199] Among them, the result acquisition subunit is used to acquire the migration result sent by the source slave server. The third instruction sending subunit is used to send the master server transfer instruction to the source master server in response to the migration result indicating that the replication subtree transfer is successful, so that the source master server sends a master server transfer request to each of the target slave servers to determine the target master server.

[0200] In an optional implementation manner, the second instruction sending unit 1002 is further configured to send the first migration instruction to the source slave server in response to the received split result indicating that the target metadata tree split is successful.

[0201] The management server in the embodiment of the present invention sends a metadata tree split instruction to the source master server, so that the source master server splits the metadata tree to obtain a metadata subtree, generates configuration information of the metadata subtree, and sends a migration instruction to the source slave server to migrate the replication subtree of the metadata subtree of the source slave server to the target server. Furthermore, a master server transfer instruction is sent to the source master server to update the source master server to the source slave server of the metadata subtree, so as to send a migration instruction to the source master server, so that the source master server migrates the replication subtree of the metadata subtree to the target slave server of the metadata subtree. Thus, in the embodiment of the present invention, the target server is updated to the master server of the metadata subtree to process the data operation request of the client for the metadata subtree only after the metadata subtree of the source slave server is migrated to the target server, so as to reduce the negative impact on the access to the data storage system during the metadata migration process.

[0202] Figure 11 is a schematic diagram of the data migration device in the embodiment of the present invention. As Figure 11 shown, the data migration device in the embodiment of the present invention is applicable to a target server and includes a subtree receiving unit 1101 and a request processing unit 1102.

[0203] Among them, the subtree receiving unit 1101 is used to receive the replication subtree of the metadata subtree sent by the source slave server, and the metadata subtree is obtained by splitting the target metadata tree. The request processing unit 1102 is used to receive the master server transfer request, determine the target master server of the metadata subtree, and receive and process the data operation request of the metadata subtree.

[0204] In an optional implementation manner, the data migration device further includes a data synchronization unit.

[0205] The data synchronization unit is used to perform data synchronization to update the replication subtree.

[0206] The target server in the embodiment of the present invention receives a replicated subtree of the metadata subtree sent by the source slave server, determines the target master server of the metadata subtree after receiving the master server transfer request, and receives and processes the data operation request for the metadata subtree after the master server of the metadata subtree is updated to the target master server. Thus, in the embodiment of the present invention, the target server is updated to the master server of the metadata subtree to process the data operation request of the client for the metadata subtree only after the metadata subtree of the source slave server is migrated to the target server, so as to reduce the negative impact on the access to the data storage system during the metadata migration process.

[0207] Figure 12 It is a schematic diagram of the electronic device in the embodiment of the present invention. Figure 12 The electronic device shown is a general data processing device, which includes a general computer hardware structure, and at least includes a processor 1201 and a memory 1202. The processor 1201 and the memory 1202 are connected through a bus 1203. The memory 1202 is suitable for storing instructions or programs executable by the processor 1201. The processor 1201 can be an independent microprocessor or a set of one or more microprocessors. Thus, the processor 1201 processes data and controls other devices by executing the instructions stored in the memory 1202 to implement the method flow of the embodiment of the present invention as described above. The bus 1203 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to a display controller 1204, a display device, and an input / output (I / O) device 1205. The input / output (I / O) device 1205 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a body sensing input device, a printer, and other devices well known in the art. Typically, the input / output (I / O) device 1205 is connected to the system through an input / output (I / O) controller 1206.

[0208] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device (equipment), or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be implemented as a computer program product on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0209] The present application is described with reference to the flowcharts of the method, device (equipment), and computer program product according to the embodiments of the present application. It should be understood that each process in the flowchart can be implemented by computer program instructions.

[0210] These computer program instructions can be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the function specified in one process Figure 1 or functions in a plurality of processes.

[0211] These computer program instructions can also be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the function specified in one process Figure 1 or functions in a plurality of processes.

[0212] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, which is used for a computer to execute some or all of the above method embodiments.

[0213] That is, those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by specifying relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., all of which can store program codes.

[0214] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data migration method, applicable to a source master server, characterized in that, The method includes: In response to receiving a metadata tree splitting instruction, splitting the target metadata tree to obtain metadata subtrees; Generating configuration information for the metadata subtrees; In response to receiving a primary server transfer instruction, updating to the source slave server for the metadata subtrees; Determining the replicated subtrees of the metadata subtrees; In response to receiving a second migration instruction, migrating the replicated subtrees to the target slave servers corresponding to the metadata subtrees.

2. The method according to claim 1, wherein The method further includes: In response to receiving a removal instruction, deleting the metadata subtrees.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Sending a server group acquisition request to obtain the storage area identifier of the metadata subtrees; The generating the configuration information for the metadata subtrees includes: Generating the configuration information according to the received storage area identifier.

4. The method according to claim 1, wherein The configuration information of the metadata subtrees includes at least one of the node change information and the node range information of the metadata subtrees and the storage area identifier of the metadata subtrees.

5. A data migration method, applicable to a management server, characterized in that, The method includes: Sending a metadata tree splitting instruction to the source primary server to split the target metadata tree to obtain metadata subtrees; Sending a first migration instruction to the source slave server to migrate the replicated subtrees corresponding to the source slave server to the target server; Sending a primary server transfer instruction to the source primary server to update the target server to the target primary server for the metadata subtrees; Sending a second migration instruction to the source primary server to migrate the replicated subtrees corresponding to the source primary server to the corresponding target slave servers.

6. The method according to claim 5, wherein The sending the primary server transfer instruction to the source primary server includes: Obtaining the migration result sent by the source slave server; In response to the migration result indicating that the replication subtree transfer is successful, sending the primary server transfer instruction to the source primary server, so that the source primary server sends a primary server transfer request to each of the target servers to determine the target primary server.

7. The method according to claim 5, wherein The sending the first migration instruction to the source slave server includes: In response to the received splitting result indicating that the target metadata tree splitting is successful, sending the first migration instruction to the source slave server.

8. A data migration method, applicable to a target server, characterized in that, The method includes: Receiving the replicated subtrees of the metadata subtrees sent by the source slave server, where the metadata subtrees are obtained by splitting the target metadata tree; Determining the target primary server of the metadata subtrees and receiving data operation requests for processing the metadata subtrees.

9. The method according to claim 8, wherein The method further includes: Performing data synchronization to update the replicated subtrees.

10. A data storage system, characterized in that, The system includes: A management server configured to send a metadata tree splitting instruction, a primary server transfer instruction, and a second migration instruction to the source primary server, and send a first migration instruction to the source slave server; The source master server is configured to split a target metadata tree in response to receiving a metadata tree splitting instruction, obtain a metadata subtree, generate configuration information for the metadata subtree, update to the source slave server for the metadata subtree in response to receiving a master server transfer instruction, determine the replicated subtree of the metadata subtree, and migrate the replicated subtree to the target slave server corresponding to the metadata subtree in response to receiving a second migration instruction; The target server is configured to receive the replicated subtree of the metadata subtree sent by the source slave server, determine the target master server of the metadata subtree, and receive and process data operation requests for the metadata subtree.

11. A data migration device, applicable to a source master server, characterized in that, The device includes: The subtree generation unit is configured to split a target metadata tree in response to receiving a metadata tree splitting instruction to obtain a metadata subtree; The configuration information generation unit is configured to generate configuration information for the metadata subtree; The server update unit is configured to update to the source slave server for the metadata subtree in response to receiving a master server transfer instruction; The subtree replication unit is configured to determine the replicated subtree of the metadata subtree; The migration unit is configured to migrate the replicated subtree to the target slave server corresponding to the metadata subtree in response to receiving a second migration instruction.

12. A data migration device, applicable to a management server, characterized in that, The device includes: The first instruction sending unit is configured to send a metadata tree splitting instruction to the source master server to split a target metadata tree to obtain a metadata subtree; The second instruction sending unit is configured to send a first migration instruction to the source slave server to migrate the replicated subtree of the metadata subtree corresponding to the source slave server to the target server; The third instruction sending unit is configured to send a master server transfer instruction to the source master server to update the target server to the target master server of the metadata subtree; The fourth instruction sending unit is configured to send a second migration instruction to the source master server to migrate the replicated subtree of the metadata subtree corresponding to the source master server to the corresponding target slave server.

13. A data migration device, applicable to a target server, characterized in that, The device includes: The subtree receiving unit is configured to receive the replicated subtree of the metadata subtree sent by the source slave server, where the metadata subtree is obtained by splitting a target metadata tree; The master server transfer unit is configured to determine the target master server of the metadata subtree and receive and process data operation requests for the metadata subtree.

14. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, where the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1-9.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method according to any one of claims 1-9.