Data migration method and device

By building a cyclable red and black tree in a distributed database, identifying candidate nodes for downtime nodes and migrating data to these nodes, the system avalanche caused by downtime is solved and data migration and query efficiency is improved.

CN120277048APending Publication Date: 2025-07-08BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410020834.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In a distributed database, the downtime node caused by a database crash causes data migration and query requests to be overloaded, which in turn causes a system avalanche, which is difficult to effectively solve in the existing technology.

Method used

By building a cyclable red and black tree on the hash ring, identify multiple candidate nodes for the downtime node, and migrate data to these candidate nodes, the data migration and query process is optimized using the hierarchical relationship of the red and black tree.

Benefits of technology

It improves the efficiency of node query, reduces the probability of system avalanche, and improves the efficiency of data migration and query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277048A_ABST
    Figure CN120277048A_ABST
Patent Text Reader

Abstract

The invention discloses a data migration method and device, and relates to the technical field of computers. A specific embodiment of the method comprises the steps of determining a crash node corresponding to a crash database from virtual nodes, corresponding to a plurality of databases, on a hash ring; according to a preset data migration strategy, determining a plurality of candidate nodes corresponding to the downtime node from a recyclable red-black tree constructed according to the virtual nodes; the data migration strategy indicates a hierarchical relationship between the plurality of candidate nodes and the downtime node in the recyclable red-black tree; and migrating data stored in the crash database to databases corresponding to the plurality of candidate nodes. Therefore, the candidate nodes are determined through the red-black tree, and the node query efficiency is improved; according to the method, the data corresponding to the downtime node is dispersed and transferred to the candidate nodes in parallel, so that the data migration and query efficiency is improved, and the probability of system avalanche is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and in particular, to a data migration method and apparatus. Background Art

[0002] Currently, load balancing of a distributed database usually adopts a consistent hashing algorithm to evenly disperse data into each database. In some special scenarios, such as a promotion festival, a certain database is accessed by a large number of users, causing the database to crash. Correspondingly, the node corresponding to this database on the hash ring is in a down state, which is equivalent to the deletion of this node. Since this node is deleted, all the data stored in it will be migrated to the next node. The next node not only has to store the data migrated from the deleted node, but also has to process a large number of data access requests corresponding to the migrated data, resulting in the next node crashing. Subsequent nodes repeat the above process in turn, ultimately leading to a system avalanche. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a data migration method and apparatus. A down node corresponding to a crashed database is determined from virtual nodes respectively corresponding to multiple databases on a hash ring; according to a preset data migration strategy, multiple candidate nodes corresponding to the down node are determined from a cyclic red-black tree constructed based on the virtual nodes; the data migration strategy indicates a hierarchical relationship between the multiple candidate nodes and the down node in the cyclic red-black tree; and the data stored in the crashed database is migrated to the databases corresponding to the multiple candidate nodes. Thus, candidate nodes are determined through a red-black tree, improving the node query efficiency; the data corresponding to the down node is dispersed and transferred in parallel to multiple candidate nodes, thereby improving the data migration and query efficiency, and at the same time reducing the probability of a system avalanche.

[0004] To achieve the above object, according to one aspect of embodiments of the present invention, a data migration method is provided.

[0005] A data migration method according to an embodiment of the present invention includes: determining a down node corresponding to a crashed database from virtual nodes respectively corresponding to multiple databases on a hash ring; according to a preset data migration strategy, determining multiple candidate nodes corresponding to the down node from a cyclic red-black tree constructed based on the virtual nodes; the data migration strategy indicates a hierarchical relationship between the multiple candidate nodes and the down node in the cyclic red-black tree; and migrating the data stored in the crashed database to the databases corresponding to the multiple candidate nodes.

[0006] Optionally, determining multiple candidate nodes corresponding to the crashed node from the circular red - black tree constructed according to the virtual nodes according to the preset data migration policy includes: determining the crashed node from the circular red - black tree; and according to the hierarchical relationship, starting from the crashed node and along a preset direction, finding multiple candidate nodes corresponding to the crashed node.

[0007] Optionally, finding multiple candidate nodes corresponding to the crashed node from the crashed node along a preset direction according to the hierarchical relationship includes: according to the hierarchical relationship, starting from the crashed node and along the direction towards the leaf nodes of the circular red - black tree, finding a first node corresponding to the crashed node; according to the hierarchical relationship, starting from the crashed node and along the direction towards the root node of the circular red - black tree, finding a second node corresponding to the crashed node; and taking the first node and the second node as the multiple candidate nodes.

[0008] Optionally, the method provided by the present invention further includes: determining the position number corresponding to the virtual node on the hash ring; the position number is determined according to the addresses of the multiple databases; constructing an initial red - black tree corresponding to the virtual node according to the position number and the structural characteristics of the red - black tree; taking the root node of the initial red - black tree as the child node of the last leaf node in the pre - order traversal of the initial red - black tree; for each leaf node except the last leaf node in the initial red - black tree: according to the pre - order traversal order, determining the next node whose traversal order is adjacent to the leaf node and after the leaf node, and taking the next node as the child node of the leaf node.

[0009] Optionally, the method provided by the present invention further includes: in response to receiving a user access request corresponding to the crashed database, determining the target data indicated by the user access request; determining adjacent nodes of the crashed node on the hash ring; and querying the target data from the databases corresponding to the adjacent nodes.

[0010] Optionally, the method provided by the present invention further includes: in the case that the target data is not queried from the databases corresponding to the adjacent nodes, determining whether the adjacent nodes belong to the multiple candidate nodes; in the case that the adjacent nodes belong to the multiple candidate nodes, determining other candidate nodes except the adjacent nodes from the multiple candidate nodes; and querying the target data from the databases corresponding to the other candidate nodes.

[0011] Optionally, the method provided by the present invention further includes: in response to receiving a data storage request corresponding to the crash database, determining the data to be stored indicated by the data storage request; and storing the data to be stored in the database corresponding to the adjacent node.

[0012] Optionally, the method provided by the present invention further includes: in response to receiving a join request for a new database, determining a hash value corresponding to the new database address; according to the hash value, determining a new position number corresponding to the virtual node corresponding to the new database on the hash ring; and updating the circular red-black tree and the hierarchical relationship between the nodes in the circular red-black tree according to the new position number.

[0013] Optionally, the migrating the data stored in the crash database to the databases corresponding to the multiple candidate nodes includes: determining an identification value of the data to be migrated corresponding to the crash database; according to the identification value, determining multiple candidate nodes corresponding to the data to be transferred; and correspondingly migrating the data to be transferred to the databases corresponding to the multiple candidate nodes.

[0014] To achieve the above object, according to another aspect of the embodiments of the present invention, there is provided a data migration apparatus.

[0015] A data migration apparatus according to an embodiment of the present invention includes: a monitoring module, configured to determine a down node corresponding to the crash database from the virtual nodes respectively corresponding to multiple databases on the hash ring; a processing module, configured to determine multiple candidate nodes corresponding to the down node from a circular red-black tree constructed according to the virtual nodes according to a preset data migration strategy, where the data migration strategy indicates a hierarchical relationship between the multiple candidate nodes and the down node in the circular red-black tree; and a data migration module, configured to migrate the data stored in the crash database to the databases corresponding to the multiple candidate nodes.

[0016] To achieve the above object, according to another aspect of the embodiments of the present invention, there is provided a server.

[0017] A server according to an embodiment of the present invention includes: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement a data migration method according to an embodiment of the present invention.

[0018] To achieve the above object, according to still another aspect of the embodiments of the present invention, there is provided a computer-readable storage medium.

[0019] A computer-readable storage medium according to an embodiment of the present invention stores a computer program thereon, and when the program is executed by a processor, it implements a data migration method according to an embodiment of the present invention.

[0020] One embodiment of the above invention has the following advantages or beneficial effects: By determining the down nodes corresponding to the crashed databases from the virtual nodes respectively corresponding to multiple databases on the hash ring; according to a preset data migration strategy, determining multiple candidate nodes corresponding to the down nodes from the circular red-black tree constructed based on the virtual nodes; the data migration strategy indicates the hierarchical relationship between the multiple candidate nodes and the down nodes in the circular red-black tree; migrating the data stored in the crashed databases to the databases corresponding to the multiple candidate nodes. Thus, by determining candidate nodes through the red-black tree, the node query efficiency is improved; the data corresponding to the down nodes is scattered and transferred in parallel to multiple candidate nodes, thereby improving the data migration and query efficiency, and at the same time reducing the probability of system avalanche.

[0021] The further effects of the above non-conventional optional methods will be described below in combination with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:

[0023] Figure 1 is a schematic diagram of the main steps of a data migration method according to an embodiment of the present invention;

[0024] Figure 2 is a schematic diagram of the mapping relationship between a database and virtual nodes on a hash ring according to an embodiment of the present invention;

[0025] Figure 3 is a schematic diagram of the main steps of constructing a circular red-black tree according to an embodiment of the present invention;

[0026] Figure 4a is a schematic diagram of constructing a circular red-black tree according to virtual nodes on a hash ring according to an embodiment of the present invention;

[0027] Figure 4b is a decomposition diagram of a two-way circular link of a circular red-black tree according to an embodiment of the present invention;

[0028] Figure 5 is a schematic diagram of the main steps of finding candidate nodes in a circular red-black tree according to an embodiment of the present invention;

[0029] Figure 6 is a schematic diagram of data migration of a down node according to an embodiment of the present invention;

[0030] Figure 7 is a schematic diagram of querying data corresponding to a down node according to an embodiment of the present invention;

[0031] Figure 8 is a schematic diagram of the main steps of a data migration method according to an embodiment of the present invention;

[0032] Figure 9 is a schematic diagram of the main modules of a data migration device according to an embodiment of the present invention;

[0033] Figure 10 is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;

[0034] Figure 11 is a schematic diagram of the structure of a computer system of a database server or a server suitable for implementing the embodiments of the present invention. Detailed implementation manners

[0035] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following.

[0036] It should be noted that, without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0037] Figure 1 is a schematic diagram of the main steps of the data migration method according to the embodiment of the present invention.

[0038] As Figure 1 shown, the data migration method according to the embodiment of the present invention mainly includes the following steps:

[0039] Step S101: Determine the down node corresponding to the crashed database from the virtual nodes corresponding to multiple databases on the hash ring. Among them, the hash ring (Hash Ring) is a data structure, usually used for data sharding, load balancing, and node selection in a distributed system. For multiple databases in a distributed system, the hash values corresponding to the multiple databases are calculated through a hash function, and the multiple databases are mapped to a ring structure according to the hash values, and multiple virtual nodes are correspondingly formed on the ring structure. Through the hash ring, the corresponding virtual node can be quickly determined on the hash ring through a given key value, and the data can be stored in the database corresponding to the virtual node, or the data can be quickly read from the database corresponding to the virtual node. The specific steps to determine the corresponding virtual node on the hash ring according to the key value are: calculate the hash value of the key value through the hash function, and find the first virtual node greater than or equal to the hash value in the clockwise direction, which is the virtual node corresponding to the key value.

[0040] According to the URLs (Uniform Resource Locators) of multiple databases, map the multiple databases to a hash ring respectively to form virtual nodes. For example, take the hash values of the URLs of multiple databases, and determine the maximum number of virtual nodes that can be accommodated on the hash ring, such as 2^32. Take the remainder of the hash values corresponding to the multiple databases divided by this maximum number, and the obtained remainder values can be used as the position numbers of the virtual nodes corresponding to the multiple databases on the hash ring.

[0041] For database A including n sub-databases, add i to the URL of database A in turn, where 0 <= i < n, to form an initial string. Take the hash value of this initial string, and take the remainder of the hash values corresponding to the n sub-databases divided by this maximum number. The obtained remainder values can be used as the position numbers of the virtual nodes corresponding to the n sub-databases on the hash ring. Therefore, conversely, according to the position numbers of the virtual nodes, the corresponding hash values can be determined, and then the URLs of the databases and sub-databases can be determined. Figure 2 Shows the mapping relationship between the position numbers of the virtual nodes, the hash values corresponding to the databases, and the database URLs, where the position numbers, hash values, and URLs are stored in the mapping table in a one-to-one correspondence.

[0042] After receiving the message that database B crashes, according to the mapping relationship between the URLs of multiple databases and their virtual nodes on the hash ring, the position number of the virtual node corresponding to the crashed database B on the hash ring can be determined, and this virtual node is the down node.

[0043] Step S102: According to a preset data migration strategy, determine multiple candidate nodes corresponding to the down node from the circular red-black tree constructed based on the virtual nodes; the data migration strategy indicates the hierarchical relationship between the multiple candidate nodes and the down node in the circular red-black tree. According to the data migration strategy, it can be determined which nodes having a certain relationship with the down node can be used as its corresponding candidate nodes. Therefore, once the down node is determined, according to the hierarchical relationship between the candidate nodes indicated by the data migration strategy and the down node, such as the parent node, grandparent node, son node, grandson node, etc. of the down node, according to this hierarchical relationship, the nodes having this hierarchical relationship with the down node are used as its corresponding candidate nodes. It should be noted that the number of candidate nodes is 2 or an integer multiple of 2.

[0044] In an optional embodiment of the present invention, the construction of the circular red-black tree mainly includes steps S301 - S304, as Figure 3 shown:

[0045] Step S301: Determine the position number corresponding to the virtual node on the hash ring; the position number is determined according to the addresses of the multiple databases;

[0046] Step S302: Construct an initial red-black tree corresponding to the virtual node according to the position number and the structural characteristics of the red-black tree;

[0047] Step S303: Use the root node of the initial red-black tree as the child node of the last leaf node in the pre-order traversal of the initial red-black tree;

[0048] Step S304: For each leaf node in the initial red-black tree except the last leaf node: Determine the next node whose traversal order is adjacent to the leaf node and after the leaf node according to the pre-order traversal order, and use the next node as the child node of the leaf node. In order to quickly find the corresponding candidate node regardless of the position of the downed node in the red-black tree, it is necessary to be able to traverse all nodes along the link of its child nodes at any node position in one pass. In a normal red-black tree, the leaf nodes are the nodes in the bottommost layer, so there are no child nodes. To implement a circular red-black tree, child nodes are determined for each leaf node, and the root node of the red-black tree is determined as the child node of the last leaf node. It should be noted that after determining the child nodes of each node in the circular red-black tree, the parent node can be determined conversely according to the child nodes, that is, the parent node of each node can be determined, so bidirectional circulation of the nodes in the red-black tree can be realized. Figure 4a Schematic diagram of constructing a circular red-black tree according to the virtual node of the hash ring, where the gray balls represent red nodes and the black balls represent black nodes; Figure 4b Decomposition diagram of the bidirectional circular link included in the circular red-black tree, where the gray balls represent red nodes and the black balls represent black nodes.

[0049] In order to determine candidate nodes corresponding to the crashed node, in an alternative embodiment of the present invention, the step of determining multiple candidate nodes corresponding to the crashed node from the circular red-black tree constructed according to the virtual nodes according to the preset data migration strategy includes: determining the crashed node from the circular red-black tree; according to the hierarchical relationship, starting from the crashed node and along a preset direction, finding multiple candidate nodes corresponding to the crashed node. After determining the crashed node and the hierarchical relationship between the candidate nodes and the crashed node, nodes having a preset relationship with the crashed node can be found starting from the crashed node. For example, the parent node and the grandparent node of the crashed node can be found as its corresponding candidate nodes; or the son node or the grandson node of the crashed node can be found as its corresponding candidate nodes. It can be understood that the parent node, the grandparent node, the son node, and the grandson node are only examples and are not used for limitation. Nodes having any relationship with the crashed node can be preset as candidate nodes, and the number of candidate nodes can also be flexibly set. The number of candidate nodes needs to be an integer multiple of 2, such as 2, 4, 6, 8...

[0050] In order to improve the flexibility of candidate node search and improve the search efficiency, search can be performed simultaneously in the directions of the root node and the child nodes starting from the crashed node. Therefore, in an alternative embodiment of the present invention, the step of finding multiple candidate nodes corresponding to the crashed node starting from the crashed node along a preset direction according to the hierarchical relationship includes S501 - S503, as Figure 5 shown:

[0051] Step S501: According to the hierarchical relationship, starting from the crashed node, find the first node corresponding to the crashed node along the direction towards the leaf nodes of the circular red-black tree;

[0052] Step S502: According to the hierarchical relationship, starting from the crashed node, find the second node corresponding to the crashed node along the direction towards the root node of the circular red-black tree;

[0053] Step S503: Use the first node and the second node as the multiple candidate nodes. For example, starting from the crashed node, the parent node and the grandparent node of the crashed node can be found in the direction of the root node, and the son node and the grandson node of the crashed node can be found in the direction of the leaf nodes. Finally, the parent node, the grandparent node, the son node, and the grandson node are jointly used as candidate nodes of the crashed node.

[0054] Step S103: Migrate the data stored in the crash database to the databases corresponding to the multiple candidate nodes. In order to disperse and parallelly migrate the data in the crash database to multiple candidate nodes. In an optional embodiment of the present invention, the migrating the data stored in the crash database to the databases corresponding to the multiple candidate nodes includes: determining the identification value of the data to be migrated corresponding to the crash database; determining multiple candidate nodes corresponding to the data to be transferred according to the identification value; and correspondingly migrating the data to be transferred to the databases corresponding to the multiple candidate nodes. The identification value of the data to be migrated is an identifier indicating the identity of the data to be migrated. In the embodiment of the present invention, the storage address of the data can be used as its identification value. Perform a hash operation on the identification value of each data to be migrated to obtain the corresponding hash value, and take the remainder of the hash value divided by the number of candidate nodes. For example, the candidate nodes are the parent node, the grandfather node, the son node, and the grandson node, and the number is 4, that is, take the remainder of the hash value corresponding to the data to be migrated divided by 4, and determine the corresponding candidate node according to the obtained remainder. Number the candidate nodes as follows:

[0055] Number Node 0 v_nodei (grandfather node) 1 v_nodej (parent node) 2 v_nodek (son node) 3 v_nodel (grandson node) …… ……

[0056] After determining the corresponding candidate nodes for the data to be migrated, migrate the data to be migrated to the corresponding candidate nodes. Figure 6 It is a schematic diagram of data migration, where v_node represents the downed node; v_nodei, v_nodej, v_nodek, and v_nodel represent candidate nodes, and the dashed line represents the direction of data transfer.

[0057] After a node goes down, there is still a data storage requirement for this node. At this time, the database corresponding to the downed node can no longer meet the data storage requirement. Therefore, store the data to be stored in the adjacent node corresponding to the downed node. In an optional embodiment of the present invention, the method provided in the embodiment of the present invention further includes: in response to receiving a data storage request corresponding to the crash database, determining the data to be stored indicated by the data storage request; and storing the data to be stored in the database corresponding to the adjacent node.

[0058] In order to continue to provide data query services after a node fails, in an alternative embodiment of the present invention, the method provided by the embodiments of the present invention further includes: in response to receiving a user access request corresponding to the crashed database, determining the target data indicated by the user access request; determining the adjacent nodes of the crashed node on the hash ring; querying the target data from the databases corresponding to the adjacent nodes. As described above, after a node fails, the new data corresponding to the crashed node is stored in its corresponding adjacent node. Therefore, for the query of new data, it is necessary to query from the corresponding adjacent node. Additionally, before the disaster recovery function of the system is started, the database corresponding to the crashed node may already have problems with degraded service performance. According to the processing mechanism of the consistent hashing algorithm, it may have migrated some data to the adjacent nodes corresponding to the crashed node on the hash ring. After the disaster recovery function is started, the remaining data is migrated to the candidate nodes. Therefore, when a user queries the database corresponding to the crashed node, it is necessary to first query the databases corresponding to the adjacent nodes of the crashed node, and then query the databases corresponding to the candidate nodes. Only in this way can the target data be ensured to be queried. It can be understood that it is also possible to first query the databases corresponding to the candidate nodes and then query the databases corresponding to the adjacent nodes, and the query order is not limited. It can be flexibly processed according to the requirements of query speed and query efficiency.

[0059] There may also be a situation where the adjacent nodes corresponding to the crashed node are already included in the candidate nodes. Therefore, when the target data is not found in the adjacent nodes, it only needs to be queried in the other candidate nodes except the adjacent nodes. Therefore, in an alternative embodiment of the present invention, the method provided by the embodiments of the present invention further includes: in the case where the target data is not found in the databases corresponding to the adjacent nodes, determining whether the adjacent nodes belong to the multiple candidate nodes; in the case where the adjacent nodes belong to the multiple candidate nodes, determining the other candidate nodes except the adjacent nodes from the multiple candidate nodes; querying the target data from the databases corresponding to the other candidate nodes. Among them, the data query corresponding to the crashed node is as Figure 7 shown, where node 1 represents the mapping of the target data on the hash ring; v_node represents the crashed node; v_nodei, v_nodej, v_nodek, and v_nodel represent candidate nodes; v_node_next represents the adjacent node; the dotted line represents querying the candidate nodes; and the solid line represents querying the adjacent nodes.

[0060] When a new database or sub-database is added to the distributed storage system, it is necessary to insert the virtual node corresponding to the new database into the hash ring, and update the cyclic red-black tree corresponding to the virtual nodes of the hash ring according to the newly inserted node. In an optional embodiment of the present invention, the method provided by the embodiment of the present invention further includes: in response to receiving a join request for a new database, determining a hash value corresponding to the address of the new database; according to the hash value, determining a new position number corresponding to the virtual node corresponding to the new database on the hash ring; according to the new position number, updating the cyclic red-black tree and the hierarchical relationship between the nodes in the cyclic red-black tree. As described above, the position of the new database on the hash ring can determine its hash value according to its corresponding URL, take the remainder of the hash value with respect to the maximum number of virtual nodes that the hash ring can accommodate, and determine its position on the hash ring according to the remainder, and this remainder can be used as the position number. Insert this position number into the original cyclic red-black tree. According to the structural characteristics of the red-black tree, the nodes therein may change color and / or rotate, which may involve updating the hierarchical relationship between the nodes.

[0061] The following further exemplarily describes the data migration method. The data migration method of the embodiment of the present invention mainly includes the following steps, as Figure 8 shown:

[0062] Step S801: Determine the down node corresponding to the crashed database from the virtual nodes corresponding to multiple databases on the hash ring.

[0063] Step S802: According to a preset data migration strategy, determine multiple candidate nodes corresponding to the down node from the cyclic red-black tree constructed according to the virtual nodes.

[0064] Step S803: Migrate the data stored in the crashed database to the databases corresponding to the multiple candidate nodes.

[0065] Step S804: In response to receiving a user access request corresponding to the crashed database, query the target data indicated by the user access request from the database corresponding to the adjacent node of the down node.

[0066] Step S805: In the case where the target data is not queried from the adjacent node, query the target data from the databases corresponding to the candidate nodes.

[0067] Step S806: In response to receiving a data storage request corresponding to the crashed database, store the data to be stored indicated by the data storage request in the database corresponding to the adjacent node.

[0068] Step S807: In response to receiving a request to add a new database, determine the new position number corresponding to the new database address on the hash ring, and update the red-black tree and the hierarchical relationships among the nodes in the red-black tree according to the new position number.

[0069] As can be seen from the data migration method according to the embodiments of the present invention, by determining the down nodes corresponding to the crashed database from the virtual nodes respectively corresponding to multiple databases on the hash ring; according to a preset data migration strategy, determining multiple candidate nodes corresponding to the down nodes from the circular red-black tree constructed based on the virtual nodes; the data migration strategy indicates the hierarchical relationships between the multiple candidate nodes and the down nodes in the circular red-black tree; migrating the data stored in the crashed database to the databases corresponding to the multiple candidate nodes. Thus, by determining candidate nodes through the red-black tree, the node query efficiency is improved; the data corresponding to the down nodes is scattered and transferred in parallel to multiple candidate nodes, thereby improving the data migration and query efficiency, and at the same time reducing the probability of system avalanche.

[0070] Figure 9 It is a schematic diagram of the main modules of the data migration device according to the embodiments of the present invention.

[0071] As Figure 9 shown, the data migration device 900 according to the embodiments of the present invention includes:

[0072] A monitoring module 901, configured to determine the down nodes corresponding to the crashed database from the virtual nodes respectively corresponding to multiple databases on the hash ring;

[0073] A processing module 902, configured to determine multiple candidate nodes corresponding to the down nodes from the circular red-black tree constructed based on the virtual nodes according to a preset data migration strategy; the data migration strategy indicates the hierarchical relationships between the multiple candidate nodes and the down nodes in the circular red-black tree;

[0074] A data migration module 903, configured to migrate the data stored in the crashed database to the databases corresponding to the multiple candidate nodes.

[0075] In an alternative embodiment of the present invention, the processing module 902 is further configured to determine the down nodes from the circular red-black tree; according to the hierarchical relationships, starting from the down nodes, search for multiple candidate nodes corresponding to the down nodes along a preset direction.

[0076] In an alternative embodiment of the present invention, the processing module 902 is further configured to, according to the hierarchical relationship, starting from the crashed node, search for a first node corresponding to the crashed node in the direction towards the leaf node of the circular red-black tree; according to the hierarchical relationship, starting from the crashed node, search for a second node corresponding to the crashed node in the direction towards the root node of the circular red-black tree; and use the first node and the second node as the multiple candidate nodes.

[0077] In an alternative embodiment of the present invention, the processing module 902 is further configured to determine a position number corresponding to the virtual node on the hash ring; the position number is determined according to the addresses of the multiple databases; construct an initial red-black tree corresponding to the virtual node according to the position number and the structural characteristics of the red-black tree; and determine child nodes for each leaf node in the initial red-black tree to obtain the circular red-black tree.

[0078] In an alternative embodiment of the present invention, the processing module 902 is further configured to, according to the preorder traversal order, determine the next node that is adjacent to the leaf node and whose traversal order is after the leaf node, and use the next node as the child node of the leaf node; and use the root node of the initial red-black tree as the child node of the last leaf node in the preorder traversal of the initial red-black tree.

[0079] In an alternative embodiment of the present invention, the processing module 902 is further configured to, in response to receiving a user access request corresponding to the crashed database, determine the target data indicated by the user access request; determine adjacent nodes of the crashed node on the hash ring; and query the target data from the databases corresponding to the adjacent nodes.

[0080] In an alternative embodiment of the present invention, the processing module 902 is further configured to, in the case where the target data is not queried from the databases corresponding to the adjacent nodes, determine whether the adjacent nodes belong to the multiple candidate nodes; in the case where the adjacent nodes belong to the multiple candidate nodes, determine other candidate nodes except the adjacent nodes from the multiple candidate nodes; and query the target data from the databases corresponding to the other candidate nodes.

[0081] In an alternative embodiment of the present invention, the processing module 902 is further configured to, in response to receiving a data storage request corresponding to the crashed database, determine the data to be stored indicated by the data storage request; and store the data to be stored in the databases corresponding to the adjacent nodes.

[0082] In an alternative embodiment of the present invention, the processing module 902 is further configured to, in response to receiving a request to add a new database, determine a hash value corresponding to the address of the new database; determine a new position number corresponding to the virtual node of the new database on the hash ring according to the hash value; and update the cyclic red-black tree and the hierarchical relationship between the nodes in the cyclic red-black tree according to the new position number.

[0083] In an alternative embodiment of the present invention, the migration module 903 is further configured to determine an identification value of the data to be migrated corresponding to the crashed database; determine a plurality of candidate nodes corresponding to the data to be transferred according to the identification value; and migrate the data to be transferred to the databases corresponding to the plurality of candidate nodes correspondingly.

[0084] As can be seen from the data migration apparatus according to the embodiment of the present invention, a down node corresponding to the crashed database is determined from the virtual nodes corresponding to a plurality of databases on the hash ring respectively; according to a preset data migration policy, a plurality of candidate nodes corresponding to the down node are determined from the cyclic red-black tree constructed according to the virtual nodes; the data migration policy indicates the hierarchical relationship between the plurality of candidate nodes and the down node in the cyclic red-black tree; and the data stored in the crashed database is migrated to the databases corresponding to the plurality of candidate nodes. Therefore, the candidate nodes are determined by the red-black tree, improving the node query efficiency; the data corresponding to the down node is scattered and transferred to a plurality of candidate nodes in parallel, thereby improving the data migration and query efficiency and reducing the probability of system avalanche at the same time.

[0085] Figure 10 An exemplary system architecture 1000 is shown to which the data migration method or the data migration apparatus according to the embodiment of the present invention can be applied.

[0086] As Figure 10 shown, the system architecture 1000 may include database servers 1001, 1002, 1003, a network 1004, and a data migration server 1005. The network 1004 is used to provide a medium for communication links between the database servers 1001, 1002, 1003 and the data migration server 1005. The network 1004 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0087] Database servers 1001, 1002, and 1003 interact with the data migration server 1005 via the network 1004 to receive or send messages, etc. The data migration server 1005 obtains the addresses corresponding to multiple databases from the database servers 1001, 1002, and 1003, determines the corresponding virtual nodes on the hash ring according to the database addresses, and constructs a red-black tree structure based on the virtual nodes; in addition, it obtains database downtime messages from the database servers 1001, 1002, and 1003, determines the virtual nodes and candidate nodes corresponding to the downtime databases, and migrates the data in the downtime databases to the databases corresponding to the candidate nodes.

[0088] The data migration server 1005 can be a server that provides various services, such as a background management server that supports hash operations and modulo operations on the addresses corresponding to the database servers 1001, 1002, and 1003. The background management server can analyze and process data such as user access requests received, and query the target data indicated by the user access request from the databases corresponding to the adjacent nodes according to the processing results (such as adjacent nodes).

[0089] It should be noted that the data migration method provided by the embodiments of the present invention is generally executed by the data migration server 1005. Correspondingly, the data migration device is generally set in the data migration server 1005.

[0090] It should be understood that Figure 10 the numbers of the database servers, the network, and the data migration server in

[0091] are merely illustrative. According to the implementation requirements, there can be any number of database servers, networks, and data migration servers. Figure 11 The following refers to Figure 11 which shows a schematic structural diagram of a computer system 1100 of a database server suitable for implementing the embodiments of the present invention.

[0092] As Figure 11 shown, the computer system 1100 includes a central processing unit (CPU) 1101, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1102 or the program loaded from the storage section 1108 into the random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the system 1100 are also stored. The CPU 1101, the ROM 1102, and the RAM 1103 are connected to each other via the bus 1104. The input / output (I / O) interface 1105 is also connected to the bus 1104.

[0093] The following components are connected to the I / O interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the I / O interface 1105 as required. A removable medium 1111 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 1110 as required so that a computer program read therefrom is installed into the storage section 1108 as required.

[0094] Specifically, according to an embodiment disclosed by the present invention, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment disclosed by the present invention includes a computer program product which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1109, and / or installed from the removable medium 1111. When the computer program is executed by a central processing unit (CPU) 1101, the above functions defined in the system of the present invention are executed.

[0095] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0097] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a monitoring module, a processing module, and a data migration module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the data migration module can also be described as "the module that migrates the data stored in the crashed database to the databases corresponding to the multiple candidate nodes".

[0098] As another aspect, the present invention also provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: determining a down node corresponding to the crashed database from the virtual nodes respectively corresponding to multiple databases on the hash ring; determining multiple candidate nodes corresponding to the down node from the cyclic red-black tree constructed according to the virtual nodes according to a preset data migration strategy; the data migration strategy indicates the hierarchical relationship between the multiple candidate nodes and the down node in the cyclic red-black tree; and migrating the data stored in the crashed database to the databases corresponding to the multiple candidate nodes.

[0099] According to the technical solution of the embodiments of the present invention, by determining a down node corresponding to the crashed database from the virtual nodes respectively corresponding to multiple databases on the hash ring; determining multiple candidate nodes corresponding to the down node from the cyclic red-black tree constructed according to the virtual nodes according to a preset data migration strategy; the data migration strategy indicates the hierarchical relationship between the multiple candidate nodes and the down node in the cyclic red-black tree; and migrating the data stored in the crashed database to the databases corresponding to the multiple candidate nodes. Thus, by using the red-black tree to determine candidate nodes, the node query efficiency is improved; the data corresponding to the down node is scattered and transferred to multiple candidate nodes in parallel, thereby improving the data migration and query efficiency and reducing the probability of system avalanche at the same time.

[0100] The above specific embodiments do not limit the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data migration method, characterized in that, including: determining a down node corresponding to the crashed database from virtual nodes respectively corresponding to multiple databases on the hash ring; determining, according to a preset data migration strategy, multiple candidate nodes corresponding to the down node from a cyclic red-black tree constructed based on the virtual nodes; the data migration strategy indicates a hierarchical relationship between the multiple candidate nodes and the down node in the cyclic red-black tree; migrating the data stored in the crashed database to the databases corresponding to the multiple candidate nodes.

2. The method according to claim 1, wherein The determining, according to a preset data migration strategy, multiple candidate nodes corresponding to the down node from a cyclic red-black tree constructed based on the virtual nodes includes: determining the down node from the cyclic red-black tree; according to the hierarchical relationship, starting from the down node and along a preset direction, finding multiple candidate nodes corresponding to the down node.

3. The method according to claim 2, characterized in that, The finding, according to the hierarchical relationship, starting from the down node and along a preset direction, multiple candidate nodes corresponding to the down node includes: according to the hierarchical relationship, starting from the down node and along the direction towards the leaf nodes of the cyclic red-black tree, finding a first node corresponding to the down node; according to the hierarchical relationship, starting from the down node and along the direction towards the root node of the cyclic red-black tree, finding a second node corresponding to the down node; taking the first node and the second node as the multiple candidate nodes.

4. The method according to claim 1, characterized in that, further including: determining a position number corresponding to the virtual node on the hash ring; the position number is determined according to the addresses of the multiple databases; constructing an initial red-black tree corresponding to the virtual node according to the position number and the structural characteristics of the red-black tree; taking the root node of the initial red-black tree as the child node of the last leaf node in the preorder traversal of the initial red-black tree; for each leaf node in the initial red-black tree except the last leaf node: according to the preorder traversal order, determining a next node whose traversal order is adjacent to the leaf node and after the leaf node, and taking the next node as the child node of the leaf node.

5. The method according to claim 1, wherein further including: in response to receiving a user access request corresponding to the crashed database, determining target data indicated by the user access request; determining adjacent nodes of the down node on the hash ring; querying the target data from the databases corresponding to the adjacent nodes.

6. The method according to claim 5, characterized in that, further including: in the case that the target data is not queried from the databases corresponding to the adjacent nodes, determining whether the adjacent nodes belong to the multiple candidate nodes; in the case that the adjacent nodes belong to the multiple candidate nodes, determining other candidate nodes except the adjacent nodes from the multiple candidate nodes; querying the target data from the databases corresponding to the other candidate nodes.

7. The method according to claim 5, wherein further including: in response to receiving a data storage request corresponding to the crashed database, determining data to be stored indicated by the data storage request; storing the data to be stored into the databases corresponding to the adjacent nodes.

8. The method according to claim 1, characterized in that, further including: Upon receiving a request to add a new database, determine the hash value corresponding to the address of the new database; Based on the hash value, determine the new position number corresponding to the virtual node of the new database on the hash ring; Based on the new position number, update the circular red-black tree and the hierarchical relationship between the nodes in the circular red-black tree.

9. The method according to claim 1, wherein The migrating the data stored in the crashed database to the databases corresponding to the multiple candidate nodes includes: Determine the identification value of the data to be migrated corresponding to the crashed database; Based on the identification value, determine the multiple candidate nodes corresponding to the data to be transferred; Migrate the data corresponding to the data to be transferred to the databases corresponding to the multiple candidate nodes.

10. A data migration device, characterized in that, Includes: A monitoring module, configured to determine the downed node corresponding to the crashed database from the virtual nodes respectively corresponding to multiple databases on the hash ring; A processing module, configured to determine, from the circular red-black tree constructed according to the virtual nodes, multiple candidate nodes corresponding to the downed node according to a preset data migration strategy; the data migration strategy indicates the hierarchical relationship between the multiple candidate nodes and the downed node in the circular red-black tree; A data migration module, configured to migrate the data stored in the crashed database to the databases corresponding to the multiple candidate nodes.

11. A server, characterized in that, Includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-9.

12. A computer-readable medium having a computer program stored thereon, characterized in that, The program, when executed by the processor, implements the method according to any one of claims 1-9.