Data migration method, medium, device and program product

By using data from other data nodes in the distributed storage system to restore the data to be migrated, the problem of bandwidth bottleneck of nodes with high storage utilization is solved and the efficiency of data migration is improved.

CN120780685APending Publication Date: 2025-10-14ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410424111.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-09
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

In a distributed storage system, the bandwidth capacity of data nodes with high storage utilization can easily become a bottleneck in the data migration process, resulting in reduced migration efficiency.

Method used

By using the data on other data nodes to recover the data to be migrated by the third data node with better storage performance instead of directly copying the data from the first data node with poor storage performance, the dependence on the bandwidth capacity of the first data node is reduced.

Benefits of technology

Improves the efficiency of the data migration process and reduces bottlenecks caused by bandwidth limitations, especially when there are fewer data nodes with high storage utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780685A_ABST
    Figure CN120780685A_ABST
Patent Text Reader

Abstract

A data migration method, medium, device and program product are applied to a distributed storage system comprising a plurality of data nodes, and data stored on at least one of the plurality of data nodes can be recovered based on data stored on other data nodes in the plurality of data nodes. The method comprises the following steps: determining first data stored on a first data node in the plurality of data nodes; determining a second data node where second data used for performing data recovery on the first data is located, and sending the second data stored on the second data node to a third data node, so that the third data node performs data recovery on the first data based on the second data; the storage performance of the third data node is superior to that of the first data node; and deleting the first data on the first data node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of communication technology, and in particular to a data migration method, medium, device, and program product. Background Art

[0002] To ensure a more even distribution of data across data nodes in a distributed storage system, data stored on data nodes with higher storage utilization is typically migrated to data nodes with lower storage utilization. In related technologies, this approach typically uses data nodes with higher storage utilization as source nodes and copies data from the source nodes directly to data nodes with lower storage utilization. However, the bandwidth capacity of data nodes with higher storage utilization can easily become a bottleneck in the data replication process, reducing the efficiency of the data migration process. Summary of the Invention

[0003] In a first aspect, an embodiment of the present disclosure provides a data migration method, which is applied to a distributed storage system including multiple data nodes, wherein data stored on at least one data node among the multiple data nodes can be recovered based on data stored on other data nodes among the multiple data nodes; the method comprises: determining first data stored on a first data node among the multiple data nodes; determining a second data node where second data used to recover the first data is located, and sending the second data stored on the second data node to a third data node, so that the third data node recovers the first data based on the second data; the storage performance of the third data node is better than the storage performance of the first data node; and deleting the first data on the first data node.

[0004] In some embodiments, the storage performance of a data node is represented by storage utilization; if the difference between the storage utilization of the first data node and the average storage utilization of the multiple data nodes is greater than or equal to a first preset threshold, the sending of the second data stored on the second data node to the third data node includes: if the difference between the storage utilization of the second data node and the average storage utilization of the multiple data nodes is less than the first preset threshold, the sending of the second data stored on the second data node to the third data node.

[0005] In some embodiments, the method further includes: if the difference between the storage utilization of the second data node and the average storage utilization of the multiple data nodes is greater than or equal to the first preset threshold, copying the first data stored on the first data node to the third data node.

[0006] In some embodiments, the storage performance of a data node is represented by storage utilization; the method also includes: after the third data node recovers the first data based on the second data, if the multiple data nodes include a fourth data node whose storage utilization is greater than or equal to the average storage utilization of the multiple data nodes and a fifth node whose storage utilization is less than the average storage utilization of the multiple data nodes, migrating the data stored on the fourth data node to the fifth data node.

[0007] In some embodiments, the first data node and the second data node are both original data nodes of the distributed storage system, and the third data node is a data node newly added after the distributed storage system is expanded.

[0008] In some embodiments, the storage performance of the data node is represented by storage utilization; the difference between the storage utilization of the first data node and the average storage utilization of the multiple data nodes is greater than or equal to a first preset threshold, the difference between the storage utilization of the third data node and the average storage utilization of the multiple data nodes is less than the first preset threshold, and the number of the first data nodes is less than the number of the third data nodes.

[0009] In some embodiments, the second data is a data copy of the first data; sending the second data stored on the second data node to the third data node so that the third data node performs data recovery on the first data based on the second data includes: sending the second data stored on the second data node to the third data node so that the third data node uses the second data as the recovered first data.

[0010] In some embodiments, the second data is a data block used to reconstruct the first data; sending the second data stored on the second data node to the third data node so that the third data node can recover the first data based on the second data includes: sending the second data stored on the second data node to the third data node so that the third data node can reconstruct the first data based on the second data to obtain the recovered first data.

[0011] In some embodiments, the method further includes: determining third data stored on the first data node; copying the third data to the third data node, and deleting the third data on the first data node.

[0012] In some embodiments, the number of the first data nodes is greater than or equal to 1; the determining of the third data stored on the first data node includes: respectively obtaining the total data volume of the data to be migrated on the multiple first data nodes; the data to be migrated includes the first data and the third data; obtaining the ratio between the number of the first data nodes and the total number of the multiple data nodes, and determining the total data volume of the third data on the multiple first data nodes based on the total data volume of the data to be migrated on the multiple first data nodes and the ratio; for any first data node, based on the average storage utilization of the multiple first data nodes, the storage utilization of the first data node and the total data volume of the third data on the multiple first data nodes, obtaining the data volume of the third data on the first data node, and determining the third data stored on the first data node based on the data volume of the third data on the first data node.

[0013] In some embodiments, the method further includes: if the amount of the second data sent by the second data node reaches a preset maximum amount of data, stopping sending the second data stored on the second data node to the third data node.

[0014] In some embodiments, the method further includes: if the amount of the third data sent by the first data node reaches a preset maximum amount of data, stopping copying the third data to the third data node.

[0015] In some embodiments, the first data node is a data node among the multiple data nodes, the difference between the storage utilization and the average storage utilization of the multiple data nodes being greater than or equal to a first preset threshold; after sending the second data stored on the second data node for data recovery of the first data to the third data node, so that the third data node recovers the first data based on the second data, the method further includes: re-determining the storage utilization of the multiple data nodes; if the multiple data nodes still include the first data node whose storage utilization and the average storage utilization of the multiple data nodes are greater than or equal to the first preset threshold, returning to the step of determining the first data stored on the first data node among the multiple data nodes.

[0016] In a second aspect, an embodiment of the present disclosure provides a data migration method, which is applied to a distributed storage system including a metadata node and multiple data nodes, wherein the data stored on at least one data node among the multiple data nodes can be recovered based on the data stored on other data nodes among the multiple data nodes; the method is executed by the metadata node; the method includes: determining the first data stored on the first data node among the multiple data nodes; determining the second data node where the second data used to recover the first data is located; sending a data migration indication to a third data node, so that the third data node obtains the second data from the second data node in response to the data migration indication, and recovers the first data based on the second data; the storage performance of the third data node is better than the storage performance of the first data node; sending a data deletion indication to the first data node, so that the first data node deletes the first data in response to the data deletion indication.

[0017] In a third aspect, an embodiment of the present disclosure provides a data migration method, which is applied to a distributed storage system including a metadata node and multiple data nodes, wherein the data stored on at least one of the multiple data nodes can be recovered based on the data stored on other data nodes among the multiple data nodes; the method is executed by the metadata node; the method includes: determining the first data stored on the first data node among the multiple data nodes; determining the second data node where the second data used to recover the first data is located; sending a data migration indication to the second data node, so that the second data node sends the second data to a third data node in response to the data migration indication, and the second data is used to recover the first data on the third data node; the storage performance of the third data node is better than the storage performance of the first data node; sending a data deletion indication to the first data node, so that the first data node deletes the first data in response to the data deletion indication.

[0018] In a fourth aspect, the embodiments of the present disclosure provide a data migration method, applied to a distributed storage system including a metadata node and a plurality of data nodes, data stored on at least one data node of the plurality of data nodes being capable of data recovery based on data stored on other data nodes of the plurality of data nodes; the method is executed by any one data node of the plurality of data nodes; the method includes: in response to obtaining a migration instruction sent by the metadata node, obtaining second data from a second data node, the second data being used for data recovery of first data stored on a first data node of the plurality of data nodes; performing data recovery on the first data based on the second data; the storage performance of the node is better than the storage performance of the first data node; and the first data on the first data node is deleted after the first data recovery on the node is successful.

[0019] In a fifth aspect, the embodiments of the present disclosure provide a computer readable storage medium, having a computer program stored thereon, the program being executed by a processor to implement the method of any one of the embodiments of the present disclosure.

[0020] In a sixth aspect, the embodiments of the present disclosure provide a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor implementing the method of any one of the embodiments of the present disclosure when executing the program.

[0021] In a seventh aspect, the embodiments of the present disclosure provide a computer program product, including a computer program, the computer program being executed by a processor to implement the method of any one of the embodiments of the present disclosure.

[0022] In the embodiments of the present disclosure, when migrating the first data on the first data node with poor storage performance to the third data node with better storage performance, the second data stored on the second data node is sent to the third data node, and the first data is recovered on the third data node, so that the first data does not need to be directly copied from the first data node to the third data node, thereby reducing the influence of the bandwidth capability of the first data node on the data migration process, and higher data migration efficiency can be obtained when the bandwidth capability of the first data node is poor.

[0023] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, rather than limiting the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0024] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the technical solutions of the present disclosure.

[0025] Figure 1is a schematic diagram of a distributed storage system of an embodiment of the present disclosure.

[0026] Figure 2 is a schematic diagram of a data migration process in the related art.

[0027] Figure 3 is a flowchart of a data migration method of an embodiment of the present disclosure.

[0028] Figure 4A is a schematic diagram of data recovery by copying data replicas of an embodiment of the present disclosure.

[0029] Figure 4B is a schematic diagram of data recovery by data reconstruction of an embodiment of the present disclosure.

[0030] Figure 5A and Figure 5B are schematic diagrams of the number of overloaded nodes and the number of idle nodes under different numbers of expansion nodes of an embodiment of the present disclosure.

[0031] Figure 6 is a flowchart of a data migration method of another embodiment of the present disclosure.

[0032] Figure 7 is a flowchart of a data migration method of yet another embodiment of the present disclosure.

[0033] Figure 8 is a flowchart of a data migration method of still another embodiment of the present disclosure.

[0034] Figure 9 is a schematic diagram of a computer device of an embodiment of the present disclosure. DETAILED DESCRIPTION

[0035] The exemplary embodiments will be described in detail herein with reference to the attached drawings. In the following description, the same numbers are used to indicate the same or similar elements, unless otherwise represented. The embodiments described in the following exemplary embodiments do not represent all the implementations in accordance with the present disclosure. Rather, they are merely examples in accordance with some aspects of the present disclosure as detailed in the appended claims.

[0036] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used in the present disclosure and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. In addition, the term "at least one of' as used herein means any one of or any combination of at least two of the listed possibilities.

[0037] It should be understood that, although the terms first, second, third, etc. can be employed in this disclosure to describe various information, these information should not be limited to these terms. These terms are only used to distinguish one type of information from another type of information. For example, without departing from the scope of the disclosure, first information can also be referred to as second information, and similarly, second information can also be referred to as first information. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining" or "in response to ascertaining".

[0038] It should be noted that the various information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0039] In order to enable the person skilled in the art to better understand the technical solutions in the embodiments of the present disclosure, and make the above-mentioned purposes, characteristics and advantages of the embodiments of the present disclosure more apparent and easy to understand, the technical solutions in the embodiments of the present disclosure will be further described in detail below with reference to the drawings.

[0040] The distributed storage system refers to a system that stores data on multiple independent data nodes to realize the distributed storage and management of data. The distributed storage system needs to ensure that the state presented to the outside is consistent and does not occur state rollback. As shown in FIG. 1, the distributed storage system 10 includes a metadata node (MetaNode) 102 and multiple data nodes (ChunkServer) 104. The metadata node 102 is a centralized meta-information storage node, which is usually used to store state information, storage location information, length information of data, etc. of data. The data node 104 is a node in the distributed storage system 10 for storing data, which is usually responsible for writing, storing, reading, deleting, etc. of data copies. Figure 1

[0041] In the distributed storage system 10, in order to ensure that the data is evenly distributed in the entire cluster and keep load balance, in order to ensure that the data is evenly distributed in the cluster, the distributed storage system 10 often performs data migration and reallocates data on each data node 104 in the cluster. When performing data migration, the data stored on the data node 104 with high storage utilization can be migrated to the data node 104 with low storage utilization. The above process is also called Rebalance of the cluster.

[0042] ​In the related art, a data node 104 with a higher storage utilization rate is generally directly taken as a source node, and data on the source node is directly copied to a data node 104 with a lower storage utilization rate. As shown in Figure 2 For ease of distinction, the two data nodes 104 are denoted as data node CS1 and data node CS2. Both data node CS1 and data node CS2 can store data blocks. Assuming that data node CS1 has a higher storage utilization rate and data node CS2 has a lower storage utilization rate, data node CS1 can be taken as a source node, data node CS2 can be taken as a destination node, at least one data block (for example, data block 1) stored on data node CS1 is sent to data node CS2 for storage, and after the sending is successful, data block 1 on data node CS1 is deleted.

[0043] However, the bandwidth capability of the data node 104 with a higher storage utilization rate is likely to become a bottleneck of the data replication process, resulting in a reduced efficiency of the data replication process. For example, in a case where the number of data nodes 104 with a higher storage utilization rate is small and the number of data nodes 104 with a lower storage utilization rate is large, a large amount of data needs to be transmitted from the same data node 104 with a higher storage utilization rate, which can cause the amount of data to be transmitted to be much larger than the transmission bandwidth of the data node 104 with a higher storage utilization rate, so that the bandwidth capability of the data node 104 with a higher storage utilization rate becomes a bottleneck of the data replication process.

[0044] Based on this, the data migration method provided in the embodiments of the present disclosure can not take the data node 104 with a higher storage utilization rate as a source node when migrating data on the data node 104, but take other data nodes 104 as source nodes, restore the data to be migrated by using data on the other data nodes 104, and then send the restored data to a destination node with a lower storage utilization rate. In this way, the situation that the transmission bottleneck of the data node 104 with a higher storage utilization rate leads to a reduced efficiency of the data replication process can be reduced.

[0045] The data migration method provided in the embodiments of the present disclosure is applied to a distributed storage system 10 including a plurality of data nodes 104, and data stored on at least one data node 104 in the plurality of data nodes 104 can be restored based on data stored on other data nodes 104 in the plurality of data nodes 104. Referring to Figure 3 , the method comprises:

[0046] Step S12: determining first data stored on a first data node in the plurality of data nodes 104;

[0047] Step S14: determining a second data node where second data used for data recovery of the first data is stored, and sending the second data stored on the second data node to the third data node, so that the third data node performs data recovery of the first data based on the second data; the storage performance of the third data node is superior to that of the first data node;

[0048] Step S16: deleting the first data on the first data node.

[0049] The method of the embodiments of the present disclosure can be performed by the metadata node 102. When data is migrated from the first data node with poor storage performance to the third data node with superior storage performance, the second data stored on the second data node is sent to the third data node, and the third data node can perform data recovery of the first data based on the second data. In this way, it is not necessary to directly copy the first data from the first data node to the third data node, thereby reducing the influence of the bandwidth capability of the first data node on the data copying process, and high data copying efficiency can be obtained even when the bandwidth capability of the first data node is poor.

[0050] The specific implementation of the embodiments of the present disclosure is illustrated below.

[0051] In step S12, the metadata node 102 can determine the storage performance of each data node 104. The storage performance can be represented by the storage utilization, the remaining storage space, or the occupied storage space. Taking the storage utilization as an example, the storage utilization of the data node 104 is used to represent the occupancy of the storage space on the data node 104, which can be the storage space provided by a persistent storage medium such as a disk. The storage utilization of the data node 104 can be determined by the ratio between the occupied storage space on the data node 104 and the total storage space on the data node 104. The greater the storage utilization of the data node 104, the more the occupied storage space on the data node 104; on the contrary, the smaller the storage utilization of the data node 104, the less the occupied storage space on the data node 104. When the storage utilization of the data node 104 is too large, it can cause the read-write performance to decrease and the storage space to be insufficient, resulting in problems such as being unable to store new data. Therefore, it is necessary to migrate the data on the data node 104 with a large storage utilization to reduce the storage utilization of the data node 104. Similarly, when the storage performance is represented by the remaining storage space, it is necessary to migrate the data on the data node 104 with a small remaining storage space to increase the remaining storage space of the data node 104. When the storage performance is represented by the occupied storage space, it is necessary to migrate the data on the data node 104 with a large occupied storage space to reduce the occupied storage space of the data node 104. For ease of description, the data migration process is described below by taking the storage utilization as an example.

[0052] In some embodiments, each data node 104 can report the storage utilization of the node to the metadata node 102, and the metadata node 102 can determine the data node with a high storage utilization as the first data node. The high storage utilization can mean that the storage utilization of the data node 104 is greater than or equal to a threshold value, which can be the average storage utilization of each data node 104 in the distributed storage system 10 or a fixed value. Taking the average storage utilization of each data node 104 in the distributed storage system 10 as the threshold value as an example, assuming that the number of data nodes 104 in the distributed storage system 10 is 4, and each data node 104 is denoted as data node CS1, CS2, CS3, and CS4, then for the data node CS1, if the following conditions are met, the data node CS1 can be determined as the first data node:

[0053]

[0054] wherein R CS1 represents the storage utilization of the data node CS1, Indicates the average storage utilization of each data node 104 in the distributed storage system 10. Average storage utilization The storage utilization R of data node CS1 can be CS1 , the storage utilization rate R of data node CS2 CS2 , storage utilization R of data node CS3 CS3 and the storage utilization R of data node CS4 CS4 The average value can also be determined based on the ratio of the total occupied storage space on each data node 104 to the total storage space on each data node 104. A data node 104 that meets the above conditions can also be called a high-watermark node, indicating that the water level of the occupied storage space on the data node 104 is high.

[0055] The opposite of the high-watermark node is the low-watermark node. If the storage utilization of a data node 104 is less than the average storage utilization of each data node 104 in the distributed storage system 10, the data node is a low-watermark node, indicating that the water level of the occupied storage space on the data node 104 is low.

[0056] Furthermore, only the overloaded node can be determined as the first data node. If the difference between the storage utilization of a data node 104 and the average storage utilization of multiple data nodes 104 of the distributed storage system 10 is greater than or equal to the first preset threshold, the data node 104 is an overloaded node. Among them, the first preset threshold is greater than 0. In other words, if a data node 104 is an overloaded node, it means that a lot of storage space has been occupied on the data node 104. Therefore, the data migration demand of the data node 104 is greater than that of a general high-water mark node, and the overloaded node can be used as the first data node for data migration. The relationship between the storage utilization of the overloaded node and the average storage utilization of multiple data nodes 104 of the distributed storage system 10 can be expressed as:

[0057]

[0058] Among them, R ov represents the storage utilization of the overloaded node, and δ1 represents the first preset threshold.

[0059] The opposite of the above-mentioned overloaded nodes is an idle node. If the difference between the storage utilization of a data node 104 and the average storage utilization of multiple data nodes 104 in the distributed storage system 10 is less than a second preset threshold, the data node 104 is considered an overloaded node. The second preset threshold is less than 0. The relationship between the storage utilization of an idle node and the average storage utilization of multiple data nodes 104 in the distributed storage system 10 can be expressed as:

[0060]

[0061] Among them, R un represents the storage utilization rate of the idle node, and δ2 represents the second preset threshold.

[0062] When characterizing storage performance by the remaining storage space, if the remaining storage space of a data node 104 is less than the average remaining storage space of each data node 104 in the distributed storage system 10, then the data node 104 is a high-water mark node; otherwise, the data node 104 is a low-water mark node. If the difference between the remaining storage space of a data node 104 and the average remaining storage space of each data node 104 in the distributed storage system 10 is less than the first threshold upper limit (the first threshold upper limit is less than 0), then the data node 104 is an overloaded node. If the difference between the remaining storage space of a data node 104 and the average remaining storage space of each data node 104 in the distributed storage system 10 is greater than or equal to the first threshold lower limit (the first threshold lower limit is greater than 0), then the data node 104 is an idle node. When characterizing storage performance by other parameters, high-water mark nodes, low-water mark nodes, overloaded nodes, and idle nodes can also be determined accordingly, which will not be repeated here.

[0063] In some embodiments, when there are overloaded nodes and high-water-mark non-overloaded nodes at the same time, the overloaded node can be prioritized as the first data node. After the overloaded node recovers to a non-overloaded state, the high-water-mark non-overloaded node can be used as the first data node. In this way, data migration to the overloaded node can be prioritized to avoid the overloaded node being unavailable due to excessive storage utilization.

[0064] After determining the first data node, the amount of data that needs to be migrated to get the first data node out of the overload state or the high watermark state can be determined. Taking the first data node as an example where the first data node is an overloaded node, the amount of data that needs to be migrated can be determined based on the storage utilization rate of the first data node, the average storage utilization rate of the multiple data nodes 104 of the distributed storage system 10, and the total storage space C of the first data node. ov , determine the amount of data D that needs to be migrated for the first data node to escape the overload state:

[0065]

[0066] Then, the first data to be migrated on the first data node is determined based on the data volume D. Assuming that the data stored on the first data node is organized in data blocks, and assuming that the size of each data block is d, the number of data blocks to be migrated on the first data node is D / d. D / d data blocks can be selected from the first data node as the first data according to a certain data selection strategy, and the specific data selection strategy is not limited in this disclosure.

[0067] In step S14, the second data node where the second data used to recover the first data is located can be determined, and the second data stored on the second data node can be sent to the third data node. In some embodiments, the second data node can be a data node 104 whose storage utilization is less than the storage utilization of the first data node. For example, when the first data node is a high-water mark node, the second data node can be a low-water mark node. When the first data node is an overloaded node, the second data node can be an idle node. Using the second data on the second data node with lower storage utilization to recover the first data on the first data node with higher storage utilization can reduce the situation where the transmission process of the first data is slow due to the higher storage utilization of the first data node, and at the same time, it also reduces the situation where the transmission process of the second data is slow due to the higher storage utilization of the second data node.

[0068] Since in actual applications, the second data used to restore the first data may also be stored on a high-water mark node or even an overloaded node, in this step, it is possible to first determine whether the second data node is suitable for transmitting the second data based on the storage utilization of the second data node storing the second data. If so, the second data stored on the second data node is sent to the third data node; otherwise, the first data is directly migrated from the first data node. Specifically, when the storage utilization of the first data node is higher than the average storage utilization of the multiple data nodes 104 of the distributed storage system 10, if the storage utilization of the second data node is lower than the average storage utilization of the multiple data nodes 104 of the distributed storage system 10, the second data stored on the second data node is sent to the third data node; otherwise, the first data is directly migrated from the first data node. When the difference between the storage utilization of the first data node and the average storage utilization of multiple data nodes 104 of the distributed storage system 10 is greater than or equal to the first preset threshold, if the difference between the storage utilization of the second data node and the average storage utilization of multiple data nodes 104 of the distributed storage system 10 is less than the first preset threshold, the second data stored on the second data node is sent to the third data node; otherwise, the first data is directly migrated from the first data node.

[0069] Migrating the first data from the first data node may be copying the first data on the first data node to the third data node. The metadata node 102 may send a data request to the first data node so that the first data node sends the first data to the third data node in response to the data request. Sending the second data stored on the second data node to the third data node may be sending a data request to the second data node so that the second data node sends the second data to the third data node in response to the data request. The data request sent to the first data node and the second data node may carry indication information of the data and indication information of the third data node. The indication information of the data is used to enable the data node that receives the data request to determine which data on the node needs to be sent out, and the indication information of the third data node is used to enable the data node that receives the data request to determine to which data node 104 the data on the node needs to be sent.

[0070] After the second data stored on the second data node is sent to the third data node, the third data node can recover the first data based on the second data and store the recovered first data. The data recovery disclosed in this embodiment can be achieved by copying the data copy or by reconstructing the data.

[0071] In the embodiment of data recovery by replicating data copies, the data stored in the distributed storage system 10 includes multiple data copies, and the data content of each data copy can be the same. The second data can be a data copy of the first data, and the third data node can directly use the received second data as the first data after recovery. Figure 4A As shown, it is assumed that the first data node is data node CS1, the second data node is data node CS2, and the third data node is data node CS3. Among them, data block 1 and data block 2 stored on data node CS1 are first data, and data block 1 stored on data node CS1 has a corresponding data copy on data node CS2, that is, data block 1 stored on data node CS2. When the storage utilization rate of data node CS1 is high, data block 1 stored on data node CS2 can be sent to data node CS3. Data node CS3 can directly use the data block 1 received from data node CS2 as the recovered data block 1 stored on data node CS1 and store it. For data block 2 in data node CS1, since it does not have a data copy on other data nodes, it can be directly copied from data node CS1 to data node CS3.

[0072] In an embodiment of realizing data recovery through data reconstruction, multiple data blocks for reconstructing the original data can be stored on multiple data nodes 104 of the distributed storage system 10. The second data is a data block used to reconstruct the first data, and the third data node can reconstruct the first data based on the second data to obtain the recovered first data. The purpose of data reconstruction is generally to improve the reliability of the data in the distributed storage system 10. When some data is lost or damaged, the distributed storage system 10 can obtain data blocks from multiple data nodes 104 to reconstruct the lost or damaged data, and write the reconstructed data to other data nodes 104, thereby ensuring data security. Figure 4B As shown, it is assumed that the first data node is data node CS1, the second data node includes data nodes CS2, CS4 and CS5, and the third data node is data node CS3. Data block 1 on data node CS1 can be reconstructed through data blocks 1x on data nodes CS2, CS4 and CS5. When the storage utilization rate of data node CS1 is high, data blocks 1x on data nodes CS2, CS4 and CS5 can be sent to data node CS3 respectively, and data block 1 on data node CS1 can be reconstructed through data node CS3. The reconstructed data block 1 is the recovered data block 1, and data node CS3 can store the reconstructed data block 1. For data block 2 in data node CS1, if other data blocks used to reconstruct the data block are not stored in the distributed storage system 10, data block 2 can be directly copied from data node CS1 to data node CS3.

[0073] In the above example, the storage utilization rate of the third data node is less than the storage utilization rate of the first data node. For example, when the first data node is a high-water-level node, the third data node may be a low-water-level node. When the first data node is an overloaded node, the third data node may be an idle node. When the distributed storage system 10 includes both idle nodes and low-water-level non-idle nodes, the idle nodes may be preferentially used as the third data nodes. After the idle nodes are in a non-idle state due to data migration, if there is still data to be migrated, the low-water-level non-idle nodes may be used as the third data nodes. In this way, the situation in which the storage utilization rate of the third data node increases due to data migration and the need for data migration again may be reduced. Each data node 104 can report its own storage utilization rate to the metadata node 102. The metadata node can determine the amount of data that the idle node can receive when it leaves the idle state based on the storage utilization rate reported by each data node 104, and determine the amount of data that the node can receive when it leaves the idle state based on the amount of data that the overloaded node needs to transmit when it leaves the overloaded state and the amount of data that the idle node can receive when it leaves the idle state, and whether it is necessary to use the low-water-level non-idle node as the third data node for data migration.

[0074] In some embodiments, the number of first data nodes is less than the number of third data nodes. Furthermore, the first data node is an overloaded node, the third data node is an idle node, and the number of first data nodes is less than the number of third data nodes. When the number of first data nodes is small and the number of third data nodes is large, a large amount of first data needs to be migrated from a small number of first data nodes. Therefore, the transmission performance of the first data node often becomes a bottleneck for data migration. In addition, when the storage utilization rate of the first data node is high, data read and write operations may need to be performed frequently, which may lead to an increase in the bandwidth demand for the first data node, resulting in a competition for bandwidth between read and write operations and data migration. In particular, when the first data node is an overloaded node, the above problem is particularly prominent. Therefore, when the number of first data nodes is small and the number of third data nodes is large, there is often a more urgent need to reduce the impact of the transmission performance of the first data node on the data migration process. The embodiment of the present disclosure recovers the first data through the second data on the second data node, so that the first data node does not need to be used as the source node for data migration, which can effectively improve the data migration efficiency in the above application scenario.

[0075] In some embodiments, the first data node and the second data node are both original data nodes of the distributed storage system 10, and the third data node is a data node newly added after the distributed storage system 10 is expanded, which is called an expansion node. The expansion node is usually an idle node. In an actual distributed storage system 10, the system can usually ensure data balance, but when the new node is expanded, this balance will be broken and Rebalance will be initiated. Among them, the scale of the distributed storage system 10 and the number of expansion nodes are different, and the distribution of overloaded nodes and idle nodes will also be different. For example, under the configuration of an average storage utilization rate of 50%, a first preset threshold of 4%, and a second preset threshold of -4%, assuming that the amount of stock data in the distributed storage system 10 is normally distributed (99.5% of the data nodes have a storage utilization rate within 46%-54%), after a distributed storage system 10 including 200 data nodes adds 1 to 50 expansion nodes, the number of overloaded nodes and the number of idle nodes generated are as follows: Figure 5A It can be seen that in most cases the number of overloaded nodes is greater than the number of idle nodes, or even greater than the number of remaining nodes. In this case, overloaded nodes often do not become a transmission bottleneck. Figure 5B In the illustrated embodiment, assume that the distributed storage system 10 includes 2,000 data nodes. After adding 1 to 50 expansion nodes, the number of overloaded nodes is significantly smaller than the number of idle nodes. Directly transferring data from the overloaded nodes will cause the overloaded nodes to become transmission bottlenecks. In this case, the method of the embodiment of the present disclosure can be used to perform data migration.

[0076] In some embodiments, an overloaded node is preferentially designated as the first data node, and an idle node is designated as the third data node. Data migration is performed on the third data node based on the second data sent by the second data node. After the third data node recovers and stores the first data based on the second data, if the distributed storage system 10 still includes a fourth data node whose storage utilization is greater than or equal to the average storage utilization of the multiple data nodes 104 and a fifth node whose storage utilization is less than the average storage utilization of the multiple data nodes, the data stored on the fourth data node can be migrated to the fifth data node. The fourth data node can be a data node 104 that meets the following conditions: its storage utilization is greater than or equal to the average storage utilization, and the difference between its storage utilization and the average storage utilization is less than a first preset threshold. That is, the fourth data node is a high-watermark, non-overloaded node. The fifth data node can be a data node 104 that meets the following conditions: its storage utilization is less than the average storage utilization, and the difference between its storage utilization and the average storage utilization is greater than a second preset threshold. That is, the fifth data node is a low-watermark, non-idle node. Through the above approach, data on overloaded nodes can be preferentially migrated to idle nodes. If the distributed storage system 10 also includes high-watermark nodes and low-watermark nodes, data on the high-watermark nodes can be migrated to the low-watermark nodes. This allows for further improvement in data balance across data nodes 104 while prioritizing data balance on overloaded nodes.

[0077] In step S16, after the third data node successfully obtains the second data, it can delete the first data stored on the first data node. After successfully obtaining the second data, the third data node can send a message to metadata node 102. In response to the message, metadata node 102 can send a data deletion instruction to the first data node, causing the first data node to delete the first data. By deleting the first data on the first data node, the storage utilization rate of the first data node can be reduced.

[0078] In some embodiments, third data stored on the first data node can also be determined, the third data is copied to the third data node, and the third data on the first data node is deleted. In this embodiment, the data migration is implemented by simultaneously using the way of obtaining data from the low-water-level node and recovering the data, and the way of directly copying data from the high-water-level node. Although the transmission bandwidth of the high-water-level node can be affected due to its high storage utilization, the high-water-level node itself can still provide a certain throughput for data migration. By copying the third data directly from the first data node to the third data node, the throughput of the first data node can be fully utilized, thereby further improving the data migration efficiency.

[0079] In some embodiments, the amount of third data that needs to be directly copied from the first data node can be determined according to the total amount of data to be migrated in the distributed storage system 10, the number M of the first data nodes, and the total number N of the data nodes 104 in the distributed storage system 10. The data to be migrated includes the first data and the third data.

[0080] First, a ratio between the number M of the first data nodes and the total number N of the data nodes 104 in the distributed storage system 10 can be obtained, and according to the ratio, a proportion of the data migration tasks undertaken by the first data nodes and the second data nodes can be determined, that is, a proportion of the amount of data transmitted by the data recovery mode and the amount of data migrated by the data replication of the first data nodes, which can reflect the probability of the data migration by the third data replication of the first data nodes and the data migration by the second data transmission of the second data nodes. According to the total amount of the data to be migrated and the ratio, the total amount of the third data required to be transmitted by each first data node can be determined. For any first data node, the amount of the third data on the first data node can be obtained based on the average storage utilization of the plurality of first data nodes, the storage utilization of the first data node, and the total amount of the third data on the plurality of first data nodes, and the third data stored on the first data node can be determined based on the amount of the third data on the first data node. Specifically, according to the ratio between the average storage utilization of the plurality of first data nodes and the storage utilization of the single first data node, the proportion of the amount of the third data responsible for transmission by each first data node can be determined, and the total amount of the third data on the plurality of first data nodes can be multiplied by the above-mentioned proportion to obtain the amount of the third data required to be transmitted on each first data node. In this way, the allocation proportion of the data migration task on each first data node can be further adjusted according to the average storage utilization of the first data nodes / the storage utilization of each first data node, so that the amount of data transmitted by each first data node is as same as possible. The first data node can replicate the third data corresponding to the above-mentioned amount of data from the node according to the above-mentioned amount of data.

[0081] In the embodiment of the data migration by the replication of the data copy, the probability of the data migration by the third data replication of the first data nodes can be denoted as:

[0082]

[0083] The probability of the data migration by the second data transmission of the second data nodes can be denoted as:

[0084]

[0085] wherein D[i] represents the data amount of the data to be migrated on the ith first data node. It can be seen that the probability of selecting the way of copying the third data through the first data node to perform data migration is inversely related to the data amount of the data to be migrated on the first data node. The more the data amount of the data to be migrated on the first data node, the more data migration tasks need to be allocated to the second data node, thereby reducing the influence of the high storage utilization of the first data node on the data migration efficiency.

[0086] In the embodiment of realizing data migration through data reconstruction, assuming that K second data are needed for reconstructing each first data, the data amount transmitted by the first data node is 1 / K of the data amount transmitted by the second data node, and the allocation ratio needs to be multiplied by K, the probability of selecting the way of copying the third data through the first data node to perform data migration can be denoted as:

[0087]

[0088] The probability of selecting the way of transmitting the second data through the second data node to perform data migration can be denoted as:

[0089]

[0090] In some embodiments, in order to reduce the influence of the data migration process on the data read and write on the data node 104, an upper limit of the data amount that each data node 104 can transmit during the data migration process can be set. If the data amount transmitted by the data node 104 for data migration reaches the above data amount upper limit, the data migration task of the node is suspended.

[0091] Specifically, if the data amount of the second data sent by the second data node reaches the preset maximum data amount, and the data amount of the third data sent by the first data node does not reach the preset maximum data amount, the sending of the second data stored on the second data node to the third data node is stopped, and the data migration of the first data on the first data node is performed only by copying the third data on the first data node to the third data node. If the data amount of the third data sent by the first data node reaches the preset maximum data amount, and the data amount of the second data sent by the second data node does not reach the preset maximum data amount, the copying of the third data to the third data node is stopped, and the data migration of the first data on the first data node is performed only by sending the second data on the second data node to the third data node for data reconstruction. If the data amount of the second data sent by the second data node reaches the preset maximum data amount, and the data amount of the third data sent by the first data node reaches the preset maximum data amount, the data transmission of the first data node and the second data node can be suspended first, and then retry after a preset time period.

[0092] In some embodiments, multiple rounds of data migration can be performed, in each round of data migration, data migration can be performed by sending the second data on the second data node to the third data node for data recovery, and data migration can also be performed by directly copying the third data on the first data node to the third data node. After each round of data migration, the storage utilization of each data node 104 can be re-determined, and the next round of data migration can be performed according to the re-determined storage utilization of each data node 104. Specifically, after each round of data migration is completed (after the second data on the second data node is sent to the third data node, and / or after the third data on the first data node is copied to the third data node), the storage utilization of each data node 104 can be re-determined. If the high-water mark node (overloaded node or high-water mark non-overloaded node) is still included in the distributed storage system 10, return to step S12.

[0093] Referring to Figure 6 The embodiments of the present disclosure also provide a data migration method, applied to a distributed storage system 10 including a metadata node 102 and a plurality of data nodes 104, data stored on at least one data node 104 in the plurality of data nodes 104 can be recovered based on data stored on other data nodes 104 in the plurality of data nodes 104; the method is performed by the metadata node 102; and the method includes:

[0094] Step S22: determining first data stored on a first data node in the plurality of data nodes 104;

[0095] Step S24: determining a second data node in which second data used for data recovery of the first data is stored;

[0096] Step S26: sending a data migration instruction to the third data node, so that the third data node acquires the second data from the second data node and performs data recovery on the first data based on the second data in response to the data migration instruction; the storage performance of the third data node is better than that of the first data node;

[0097] Step S28: sending a data deletion instruction to the first data node, so that the first data node deletes the first data in response to the data deletion instruction.

[0098] In the embodiments of the present disclosure, after determining the first data node with poor storage performance, the metadata node 102 can send a data migration indication to the third data node. The data migration indication can carry identification information of the second data node, wherein the second data node stores the second data, and the second data can perform data recovery on the first data of the first data node. In this way, the third data node can obtain the second data from the second data node and perform data recovery on the first data, so that the first data node does not need to directly transmit the first data to the third data node. After the first data recovery is successful, the metadata node 102 can also send a data deletion indication to the first data node to instruct the first data node to delete the first data. The data deletion indication can carry identification information or a storage location of the first data, so that the first data node deletes the first data including the corresponding identification information or stored in the corresponding storage location.

[0099] For other details of the embodiments of the present disclosure, please refer to the foregoing method embodiments, which will not be described here.

[0100] Referring to Figure 7 The embodiments of the present disclosure also provide a data migration method, applied to a distributed storage system 10 including a metadata node 102 and a plurality of data nodes 104, data stored on at least one data node 104 in the plurality of data nodes 104 can be recovered based on data stored on other data nodes 104 in the plurality of data nodes 104; the method is performed by the metadata node 102; and the method includes:

[0101] Step S32: determining first data stored on a first data node in the plurality of data nodes 104;

[0102] Step S34: determining a second data node where second data for data recovery of the first data is located;

[0103] Step S36: sending a data migration indication to the second data node, so that the second data node sends the second data to a third data node in response to the data migration indication, the second data being used for data recovery of the first data on the third data node; and the storage performance of the third data node is better than that of the first data node;

[0104] Step S38: sending a data deletion indication to the first data node, so that the first data node deletes the first data in response to the data deletion indication.

[0105] Different from the previous embodiment, in the present embodiment, the metadata node 102 sends the data migration indication to the second data node instead of the third data node. The data migration indication can carry the identification information of the third data node, so that the second data node determines the destination of the second data. After receiving the data migration indication, the second data node can send the second data stored on the node to the third data node, so that the third data node recovers the first data. Other details of the present embodiment are described in the foregoing method embodiments, which are not described here again.

[0106] With reference to Figure 8 The present embodiment also provides a data migration method, applied to a distributed storage system 10 including a metadata node 102 and a plurality of data nodes 104, data stored on at least one data node 104 in the plurality of data nodes 104 can be recovered based on data stored on other data nodes 104 in the plurality of data nodes 104; the method is executed by any one data node 104 in the plurality of data nodes 104; and the method includes:

[0107] Step S42: in response to obtaining the migration indication sent by the metadata node 102, obtaining second data from the second data node, the second data being used for data recovery of first data stored on a first data node in the plurality of data nodes 104;

[0108] Step S44: recovering the first data based on the second data; the storage performance of the node is better than the storage performance of the first data node; and the first data on the first data node is deleted after the first data recovery on the node is successful.

[0109] The present embodiment also provides a computer device, which at least includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of any of the foregoing embodiments when executing the program.

[0110] Figure 9 A more specific hardware structure schematic diagram of a computer device provided by the present embodiment is shown, which can include a processor 20, a memory 22, an input / output interface 24, a communication interface 26, and a bus 28. The processor 20, the memory 22, the input / output interface 24, and the communication interface 26 are connected to each other through the bus 28 for communication within the device.

[0111] The processor 20 can be implemented using a general-purpose central processing unit, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure. The processor 20 may also include a graphics card, such as an Nvidia Titan X graphics card or an 1080Ti graphics card.

[0112] The memory 22 can be implemented in the form of a read-only memory (ROM), a random access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 22 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present disclosure are implemented through software or firmware, the relevant program codes are stored in the memory 22 and are called and executed by the processor 20.

[0113] The input / output interface 24 is used to connect input / output modules to enable information input and output. The input / output modules can be configured as components within the device (not shown) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0114] The communication interface 26 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WIFI, Bluetooth, etc.).

[0115] The bus 28 comprises a pathway for transmitting information between the various components of the device, such as the processor 20 , the memory 22 , the input / output interface 24 , and the communication interface 26 .

[0116] It should be noted that although the above device only shows the processor 20, memory 22, input / output interface 24, communication interface 26, and bus 28, in a specific implementation, the device may also include other components necessary for normal operation. In addition, those skilled in the art will understand that the above device may only include the components necessary to implement the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.

[0117] An embodiment of the present disclosure provides a computer program product, including a computer program, which implements the method described in any embodiment of the present disclosure when executed by a processor.

[0118] An embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in any of the aforementioned embodiments when the program is executed by a processor.

[0119] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0120] Through the description of the above implementation methods, it can be seen that those skilled in the art can clearly understand that the embodiments of the present disclosure can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the embodiments of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments of the present disclosure.

[0121] The systems, devices, modules, or units described in the above embodiments may be implemented by a computer device or entity, or by a product having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0122] Each embodiment in the present disclosure is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely illustrative, wherein the modules described as separate components may or may not be physically separated, and when implementing the embodiment of the present disclosure, the functions of each module can be implemented in the same one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the embodiment. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0123] The above is only a specific implementation of the embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the embodiment of the present disclosure. These improvements and modifications should also be regarded as the scope of protection of the embodiment of the present disclosure.

Claims

1. A data migration method, applied to a distributed storage system comprising a plurality of data nodes, wherein data stored on at least one of the plurality of data nodes can be restored based on data stored on other data nodes of the plurality of data nodes; the method comprising: Determining first data stored on a first data node among the plurality of data nodes; Determining a second data node where second data used to recover the first data is located, and sending the second data stored on the second data node to a third data node, so that the third data node recovers the first data based on the second data; the storage performance of the third data node is better than that of the first data node; Delete the first data on the first data node.

2. The method according to claim 1, wherein the storage performance of the data node is represented by storage utilization; and the difference between the storage utilization of the first data node and the average storage utilization of the plurality of data nodes is greater than or equal to a first preset threshold, wherein the sending the second data stored on the second data node to the third data node comprises: If the difference between the storage utilization of the second data node and the average storage utilization of the multiple data nodes is less than the first preset threshold, the second data stored on the second data node is sent to the third data node.

3. The method according to claim 2, further comprising: If the difference between the storage utilization of the second data node and the average storage utilization of the multiple data nodes is greater than or equal to the first preset threshold, the first data stored on the first data node is copied to the third data node.

4. The method according to claim 1, wherein the storage performance of the data node is represented by storage utilization; the method further comprising: After the third data node recovers the first data based on the second data, if the multiple data nodes include a fourth data node whose storage utilization is greater than or equal to the average storage utilization of the multiple data nodes and a fifth node whose storage utilization is less than the average storage utilization of the multiple data nodes, the data stored on the fourth data node is migrated to the fifth data node.

5. According to the method of claim 1, the first data node and the second data node are both original data nodes of the distributed storage system, and the third data node is a data node newly added after the distributed storage system is expanded.

6. According to the method according to claim 1, the storage performance of the data node is represented by storage utilization; the difference between the storage utilization of the first data node and the average storage utilization of the multiple data nodes is greater than or equal to a first preset threshold, the difference between the storage utilization of the third data node and the average storage utilization of the multiple data nodes is less than the first preset threshold, and the number of the first data nodes is less than the number of the third data nodes.

7. The method according to claim 1, wherein the second data is a data copy of the first data; and sending the second data stored on the second data node to a third data node so that the third data node recovers the first data based on the second data, comprising: The second data stored on the second data node is sent to a third data node, so that the third data node uses the second data as the restored first data.

8. The method according to claim 1, wherein the second data is a data block used to reconstruct the first data; and sending the second data stored on the second data node to a third data node so that the third data node recovers the first data based on the second data, comprises: The second data stored on the second data node is sent to a third data node, so that the third data node reconstructs the first data based on the second data to obtain restored first data.

9. The method according to claim 1, further comprising: Determining third data stored on the first data node; The third data is copied to the third data node, and the third data on the first data node is deleted.

10. The method according to claim 9, wherein the number of the first data nodes is greater than or equal to 1; and determining the third data stored on the first data node comprises: Obtaining the total amount of data to be migrated on each of the first data nodes; The data to be migrated includes the first data and the third data; Obtaining a ratio between the number of the first data nodes and the total number of the plurality of data nodes, and determining a total amount of the third data on the plurality of first data nodes based on a total amount of the data to be migrated on the plurality of first data nodes and the ratio; For any first data node, based on the average storage utilization of multiple first data nodes, the storage utilization of the first data node and the total data volume of the third data on multiple first data nodes, the data volume of the third data on the first data node is obtained, and based on the data volume of the third data on the first data node, the third data stored on the first data node is determined.

11. The method according to claim 1 , further comprising: If the amount of the second data sent by the second data node reaches a preset maximum amount of data, stop sending the second data stored on the second data node to the third data node.

12. The method according to claim 9, further comprising: If the amount of the third data sent by the first data node reaches a preset maximum amount of data, copying the third data to the third data node is stopped.

13. The method according to claim 1, wherein the first data node is a data node among the multiple data nodes, the difference between the storage utilization rate and the average storage utilization rate of the multiple data nodes being greater than or equal to a first preset threshold; after sending second data stored on the second data node for data recovery of the first data to a third data node, so that the third data node recovers the first data based on the second data, the method further comprises: Re-determining storage utilization of the plurality of data nodes; If the multiple data nodes still include a first data node whose storage utilization rate differs from the average storage utilization rate of the multiple data nodes by more than or equal to a first preset threshold, return to the step of determining the first data stored on the first data node among the multiple data nodes.

14. A data migration method, applied to a distributed storage system comprising a metadata node and a plurality of data nodes, wherein data stored on at least one of the plurality of data nodes can be restored based on data stored on other data nodes of the plurality of data nodes; The method is performed by the metadata node; The method comprises: Determining first data stored on a first data node among the plurality of data nodes; Determining a second data node where second data used to recover the first data is located; sending a data migration instruction to a third data node, so that the third data node, in response to the data migration instruction, obtains the second data from the second data node and performs data recovery on the first data based on the second data; the storage performance of the third data node is better than the storage performance of the first data node; A data deletion instruction is sent to the first data node, so that the first data node deletes the first data in response to the data deletion instruction.

15. A data migration method, applied to a distributed storage system comprising a metadata node and a plurality of data nodes, wherein data stored on at least one of the plurality of data nodes can be restored based on data stored on other data nodes of the plurality of data nodes; The method is performed by the metadata node; The method comprises: Determining first data stored on a first data node among the plurality of data nodes; Determining a second data node where second data used to recover the first data is located; sending a data migration instruction to the second data node, so that the second data node sends the second data to a third data node in response to the data migration instruction, where the second data is used to perform data recovery on the first data on the third data node; the storage performance of the third data node is better than the storage performance of the first data node; A data deletion instruction is sent to the first data node, so that the first data node deletes the first data in response to the data deletion instruction.

16. A data migration method, applied to a distributed storage system comprising a metadata node and a plurality of data nodes, wherein data stored on at least one of the plurality of data nodes can be recovered based on data stored on other data nodes of the plurality of data nodes; The method is executed by any one of the multiple data nodes; the method includes: In response to obtaining the migration instruction sent by the metadata node, obtaining second data from the second data node, where the second data is used to perform data recovery on the first data stored on the first data node among the multiple data nodes; The first data is restored based on the second data; the storage performance of this node is better than the storage performance of the first data node; the first data on the first data node is deleted after the first data on this node is successfully restored.

17. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 16 is implemented.

18. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 16 when executing the program.

19. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 16 is implemented.