Data migration method and device, equipment and medium

By automatically determining the data density difference in the federated cluster and performing data migration, the problem of business impact due to data imbalance is solved, and the effectiveness of the federated cluster is improved.

CN120010750AActive Publication Date: 2025-05-16CHINA UNITED NETWORK COMM GRP CO LTD

Patent Information

Application Number
CN202311533266.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16
Estimated Expiration
2043-11-16

AI Technical Summary

Technical Problem

In the prior art, manual balance is only started after the business is affected due to extremely uneven data, which affects the use effect of the federal cluster.

Method used

By determining the data density difference value of each subcluster in the federated cluster, setting the first threshold and the second threshold to characterize the security range of the data density difference, automatically determine the subclusters that need to be migrated out of data and the subclusters that need to be migrated in data, and migrate the data to the corresponding subcluster.

Benefits of technology

It realizes automatic triggering of sub-cluster data migration, avoids manual judgment, and timely adjusts the amount of stored data between each sub-cluster, which improves the overall use effect of the federated cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010750A_ABST
    Figure CN120010750A_ABST
Patent Text Reader

Abstract

The invention provides a data migration method and device, equipment and a medium. The method comprises the steps that a data density difference value in each sub-cluster in a federated cluster is determined, a first sub-cluster and a second sub-cluster in the federated cluster are determined according to the data density difference value of each sub-cluster in the federated cluster, the data density difference value of the first sub-cluster is larger than a first threshold value, and the data density difference value of the second sub-cluster is larger than a second threshold value. The data density difference value of the second sub-cluster is smaller than a second threshold, and the first threshold is larger than the second threshold; and migrating the data of the first sub-cluster to the second sub-cluster. According to the method, the data density difference value of each sub-cluster is compared with the first threshold value and the second threshold value, the state of the data density of each sub-cluster is confirmed in real time, and the data migration program is triggered according to the state of the data density of each sub-cluster, so that manual judgment is avoided, the storage data volume among the sub-clusters can be adjusted in time, and the data migration efficiency is improved. And the overall use effect of the federated cluster is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data migration, and in particular to a data migration method, device, equipment and medium. Background Art

[0002] The federated cluster is a distributed cluster management system that includes multiple subclusters. Due to differences in business, each directory or file has access hot spots, resulting in storage imbalance among nodes and subclusters.

[0003] In the prior art, data migration between sub-clusters is performed manually, but manual balancing is often performed only after the data is extremely unbalanced and affects the business, which affects the use effect of the federated cluster. Summary of the invention

[0004] The present application provides a data migration method, device, equipment and medium, which are used to solve the problem in the prior art that manual balancing is not started until the business is affected due to extreme data imbalance, resulting in the use effect of the federated cluster being affected.

[0005] In a first aspect, a data migration method is provided, which is applied to a data migration system, and the method includes:

[0006] determining a data density difference in each subcluster in the federated cluster, the subcluster data density difference being determined based on a target data density and a subcluster data density of the subcluster, the target data density being determined based on a total storage amount and an overall storage usage of the federated cluster, and the subcluster data density being determined based on a total subcluster storage amount and a subcluster storage usage of the subcluster;

[0007] Determine a first subcluster and a second subcluster in the federated cluster according to a data density difference of each subcluster in the federated cluster, wherein the data density difference of the first subcluster is greater than a first threshold, the data density difference of the second subcluster is less than a second threshold, and the first threshold is greater than the second threshold;

[0008] Migrate data from the first subcluster to the second subcluster.

[0009] In this application, determining the data density difference in each sub-cluster in the federated cluster includes:

[0010] The target data density is obtained by calculating the ratio of the total storage usage to the total storage volume.

[0011] The sub-cluster storage usage and the sub-cluster storage total amount are ratioed to obtain the sub-cluster data density;

[0012] The sub-cluster data density and the target data density are subtracted to determine the data density difference in each sub-cluster in the federated cluster.

[0013] In this application, migrating data of the first subcluster to the second subcluster includes:

[0014] Determine the source mount point of the first subcluster, the data to be migrated from the source mount point, and the target mount point of the second subcluster;

[0015] If the amount of data to be migrated is greater than the preset migration data amount, the data to be migrated is migrated to the target mount point of the second sub-cluster according to the preset batch migration data amount.

[0016] In the present application, if the amount of data to be migrated is greater than the preset migration data amount, the data to be migrated is migrated to the target mount point of the second sub-cluster according to the preset batch migration data amount, including:

[0017] If the amount of data to be migrated is greater than the preset migration data amount, the data to be migrated is copied to the target mount point of the second sub-cluster according to the preset batch migration data amount, wherein the state of the source mount point is not writable;

[0018] When the migration of the data to be migrated is completed, the migrated data of the target mount point in the second sub-cluster is determined;

[0019] Compare the data to be migrated with the migrated data to obtain the comparison result;

[0020] According to the comparison result, the source mount point status is adjusted from a non-writable state to a writable state;

[0021] Delete the data to be migrated in the first subcluster.

[0022] In the present application, if the amount of data to be migrated is greater than the preset migration data amount, the data to be migrated is copied to the target mount point of the second sub-cluster according to the preset batch migration data amount, including:

[0023] Determine the current batch data of the data to be migrated;

[0024] According to the source mount point and the target mount point, copy the current batch data of the data to be migrated to the target mount point of the second sub-cluster to obtain first migrated data;

[0025] After the current batch of data is copied, the source mount point status is adjusted to a non-writable state;

[0026] Determine the remaining data to be migrated;

[0027] According to the source mount point and the target mount point, copy the remaining data of the data to be migrated to the target mount point of the second sub-cluster to obtain second migrated data;

[0028] The migrated data is obtained according to the first migrated data and the second migrated data.

[0029] In this application, deleting the data to be migrated in the first sub-cluster includes:

[0030] Determine the mount table corresponding to the data to be migrated;

[0031] The mount table corresponding to the data to be migrated updates the source mount point to the target mount point;

[0032] Set the waiting time for deleting the data to be migrated;

[0033] The data to be migrated is deleted from the first subcluster according to the waiting time for deleting the data to be migrated.

[0034] In the present application, after determining the source mount point of the first sub-cluster, the data to be migrated of the source mount point, and the target mount point of the second sub-cluster, the method further includes:

[0035] If the amount of data to be migrated is less than or equal to the preset migration data amount, the source mount point state is adjusted to a non-writable state;

[0036] According to the source mount point and the target mount point, the data to be migrated is copied to the target mount point of the second sub-cluster to obtain the migrated data.

[0037] In this application, the method further comprises:

[0038] When the data density difference of the first sub-cluster is less than a first threshold, stop migrating the data in the first sub-cluster to the second sub-cluster;

[0039] When the cluster data density difference of the second subcluster is greater than the third threshold, the data in the first subcluster is migrated to the third subcluster, the subcluster data density difference of the third subcluster is less than the fourth threshold, and the third threshold is greater than the fourth threshold.

[0040] In the present application, the method further includes: sending real-time execution status information of the data migration system to the state memory according to a preset reporting mechanism, the real-time execution status information being used to characterize the execution status of the data migration method;

[0041] receiving an operation instruction generated by the state memory according to the real-time execution state information, the operation instruction including a continue execution instruction indicating that the execution state of the data migration method is normal and a stop execution instruction indicating that the execution state of the data migration method is abnormal;

[0042] Control the data migration system according to operating instructions.

[0043] In a second aspect, the present application provides a data migration device, which is applied to a data migration system, and the device includes:

[0044] a first determination module, configured to determine a data density difference in each subcluster in the federated cluster, the subcluster data density difference being determined based on a target data density and a subcluster data density of the subcluster, the target data density being determined based on a total storage volume and an overall storage usage of the federated cluster, and the subcluster data density being determined based on a total subcluster storage volume and a subcluster storage usage of the subcluster;

[0045] A second determination module is used to determine a first subcluster and a second subcluster in the federated cluster according to a data density difference of each subcluster in the federated cluster, wherein the data density difference of the first subcluster is greater than a first threshold, the data density difference of the second subcluster is less than a second threshold, and the first threshold is greater than the second threshold;

[0046] The migration module is used to migrate the data of the first sub-cluster to the second sub-cluster.

[0047] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0048] Memory stores computer-executable instructions;

[0049] The processor executes the computer-executable instructions stored in the memory to implement any one of the aforementioned data migration methods.

[0050] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement a method such as any one of the aforementioned data migration methods.

[0051] The present application provides a data migration method, device, equipment and medium, by determining the data density difference in each subcluster in a federated cluster, the data density difference of the subcluster is determined according to the target data density and the subcluster data density of the subcluster, the target data density is determined according to the total storage amount and the overall storage usage of the federated cluster, and the subcluster data density is determined according to the total subcluster storage amount and the subcluster storage usage of the subcluster; according to the data density difference of each subcluster in the federated cluster, the first subcluster and the second subcluster in the federated cluster are determined, wherein the data density difference of the first subcluster is greater than the first threshold, the data density difference of the second subcluster is less than the second threshold, and the first threshold is greater than the second threshold; means for migrating the data of the first subcluster to the second subcluster. The present application compares the data density difference of each subcluster with the first threshold and the second threshold to confirm the data density status of each subcluster, thereby determining the subcluster that needs to be migrated out of the data and the subcluster that needs to be migrated in, and triggering the data migration program when the data volume of the subcluster exceeds the first threshold, avoiding manual judgment, and being able to timely adjust the storage data volume between each subcluster, thereby improving the overall use effect of the federated cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0053] Figure 1 A schematic diagram of a data migration scenario provided for this application;

[0054] Figure 2 A flowchart of a data migration method provided for this application;

[0055] Figure 3 A flowchart of another data migration method provided for this application;

[0056] Figure 4 A schematic diagram of the structure of a data migration device provided by the present application;

[0057] Figure 5 A schematic diagram of the structure of the electronic device provided in this application.

[0058] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0059] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0060] In order to clearly understand the technical solution of the present application, the solution of the prior art is first introduced in detail.

[0061] The federated cluster is a distributed cluster management system that includes multiple subclusters. Due to differences in business, each directory or file has access hot spots, resulting in storage imbalance among nodes and subclusters.

[0062] In the prior art, data migration between sub-clusters is performed manually, but manual balancing is often performed only after the data is extremely unbalanced and affects the business, which affects the use effect of the federated cluster.

[0063] In response to the problem that manual balancing is not started until the business is affected due to extreme data imbalance, resulting in the use effect of the federated cluster being affected, the inventors have discovered in their research that the data density of the federated cluster can be used as a reference value to calculate the difference between the data density of the sub-cluster and the data density of the federated cluster. A first threshold and a second threshold can be set to characterize the safety range of the data density difference of the sub-clusters. Sub-clusters above the safety range can be used as sub-clusters that require data migration, and sub-clusters below the safety range can be used as sub-clusters to which data can be migrated. This can achieve the purpose of automatically triggering sub-cluster data migration, avoid affecting the business use of the sub-clusters, and thus improve the use effect of the federated cluster.

[0064] The following introduces the application scenarios of the data migration method provided in the embodiments of the present application.

[0065] Figure 1 A schematic diagram of a data migration scenario provided for this application, such as Figure 1 As shown, the scenario includes a federated cluster and a data migration system, wherein the federated cluster includes multiple subclusters. In order to maintain the balance of data in the subclusters of the federated cluster, the data migration system calculates whether the data density difference of each subcluster is within the range of a first threshold and a second threshold according to a preset time. When it is found that the data density difference of the subcluster exceeds the first threshold, the data migration program is started, and the subcluster is used as the first subcluster. Then, a second subcluster whose data density difference is lower than the second threshold is determined, and the data of the first subcluster is migrated to the second subcluster, thereby automatically completing the data balance of the federated cluster, reducing the impact on business use, and improving the use effect of the subclusters.

[0066] The federated cluster refers to the HDFS (Hadoop Distributed File System) RBF (Router Bbased Federation) cluster. The HDFS RBF cluster combines multiple subclusters and solves the problem of limited cluster storage capacity. It can disperse data storage and complete efficient storage and management of massive data.

[0067] A subcluster may refer to a member cluster in a federated cluster, such as an HDFS cluster, a common federated cluster, another federated cluster, or a hybrid cluster of independent clusters.

[0068] It should be noted that the above application scenarios are merely illustrative. For example, the data migration system may also be any other system with data migration functions. The data migration methods, devices, equipment and media provided in the embodiments of the present application include but are not limited to the above application scenarios.

[0069] Figure 2 A flow chart of a data migration method provided for this application is shown in FIG. Figure 2 As shown, the method includes:

[0070] S201. Determine the data density difference in each subcluster in the federated cluster. The subcluster data density difference is determined based on the target data density and the subcluster data density of the subcluster. The target data density is determined based on the total storage volume and the overall storage usage of the federated cluster. The subcluster data density is determined based on the total subcluster storage volume and the subcluster storage usage of the subcluster.

[0071] Among them, data density can refer to storage usage rate. In the embodiment of the present application, data density is the ratio of total storage usage to total storage, wherein the total storage usage is all used storage in the sub-cluster or federated cluster, and the total storage is the total storage capacity of the sub-cluster or federated cluster. For example, if the storage capacity of a sub-cluster is 10T and the used storage is 1T, then the total storage usage of the sub-cluster is 1T, the total storage is 10T, and the data density of the sub-cluster is 1T / 10T=0.1.

[0072] The data density difference may refer to the difference between the data density of the sub-cluster and the target data density. In the embodiment of the present application, the target data density is the data density of the federated cluster. The data density difference may be a positive number or a negative number. A positive number indicates that the data density of the sub-cluster is higher than the target data density, and a negative number indicates that the data density of the sub-cluster is lower than the target data density. By calculating the data density difference of the sub-cluster, the degree of difference in the data density of the sub-cluster compared to the target data density can be determined, thereby providing data support for determining whether the sub-cluster needs to be migrated in the subsequent steps.

[0073] The method for determining the data density difference in each subcluster in the federated cluster may include: regularly collecting the total storage amount and the total storage usage of each subcluster in the federated cluster through a data migration system, and calculating the data density difference of each subcluster.

[0074] In the embodiment of the present application, determining the data density difference in each sub-cluster in the federated cluster includes:

[0075] The target data density is obtained by calculating the ratio of the total storage usage to the total storage volume.

[0076] The sub-cluster storage usage and the sub-cluster storage total amount are ratioed to obtain the sub-cluster data density;

[0077] The sub-cluster data density and the target data density are subtracted to determine the data density difference in each sub-cluster in the federated cluster.

[0078] S202. Determine a first subcluster and a second subcluster in the federated cluster according to a data density difference of each subcluster in the federated cluster, wherein the data density difference of the first subcluster is greater than a first threshold, the data density difference of the second subcluster is less than a second threshold, and the first threshold is greater than the second threshold.

[0079] The first subcluster may refer to a subcluster whose data density difference is greater than a first threshold. When the data density difference of a subcluster is greater than the first threshold, it indicates that the data density of the subcluster exceeds a preset range and data migration is required.

[0080] The second subcluster may refer to a subcluster whose data density difference is less than a second threshold. When the data density difference of a subcluster is less than the second threshold, it indicates that the data density of the subcluster is lower than a preset range, and more storage capacity is available for data storage. Therefore, the migrated data of the first subcluster can be received.

[0081] The first threshold may refer to the upper limit of the preset range, and the second threshold may refer to the lower limit of the preset range. For example, if the preset range is -10 to 10, then the first threshold is 10 and the second threshold is -10.

[0082] The method for determining the first subcluster and the second subcluster in a federated cluster may include: when the data migration system detects that the data density difference of a subcluster in the federated cluster exceeds a first threshold, the subcluster is determined as the first subcluster; if the data density difference of other subcluster in the federated cluster is lower than the second threshold, the other subcluster may be determined as the second subcluster; in addition, when the other subcluster is a plurality of subcluster, any subcluster may be selected from the other subcluster as the second subcluster, or the size of the data density difference of other subcluster may be compared, and the subcluster with the smallest data density difference among the other subcluster may be selected as the second subcluster. If there is no other subcluster in the federated cluster with a data density difference lower than the second threshold, the administrator may be notified to perform data migration by adding a subcluster or manually selecting the second subcluster.

[0083] S203: Migrate data of the first sub-cluster to the second sub-cluster.

[0084] The data migration may refer to moving the data of the first sub-cluster to the second sub-cluster. After the data migration is completed, the migrated data in the first sub-cluster will no longer exist.

[0085] In the embodiment of the present application, migrating the data of the first sub-cluster to the second sub-cluster includes:

[0086] Determine the source mount point of the first subcluster, the data to be migrated from the source mount point, and the target mount point of the second subcluster;

[0087] If the amount of data to be migrated is greater than the preset migration data amount, the data to be migrated is migrated to the target mount point of the second sub-cluster according to the preset batch migration data amount.

[0088] Among them, the mount point can refer to the mount point in the Linux (operating system kernel) file system. The mount point is the entry directory of the disk file system in Linux. In the embodiment of the present application, the source mount point is the mount point path of the first subcluster, and the target mount point is the mount point path of the second subcluster.

[0089] The data to be migrated may refer to data in a mount point path. In the embodiment of the present application, it specifically refers to data in a source mount point path. For example, if the source mount point is the first subcluster / data, the data to be migrated is data in the first subcluster / data.

[0090] The method for determining a source mount point may include: determining an excess data amount exceeding a preset range according to a data density difference value of a first sub-cluster, selecting a mount point having a data amount greater than the excess data amount from the first sub-cluster, and after determining that there is no sub-mount point under the mount point, using the mount point as a source mount point.

[0091] The method for determining the data to be migrated includes: according to the determined source mount point, determining the data in the source mount point as the data to be migrated.

[0092] The method for determining the target mount point includes: creating a blank directory with the same path as the source mount point in the second sub-cluster. For example, if the source mount point is: first sub-cluster / data, then the target mount point is: second sub-cluster / data.

[0093] In the embodiment of the present application, if the amount of data to be migrated is greater than the preset migration data amount, the data to be migrated is migrated to the target mount point of the second sub-cluster according to the preset batch migration data amount, including:

[0094] If the amount of data to be migrated is greater than the preset migration data amount, the data to be migrated is copied to the target mount point of the second sub-cluster according to the preset batch migration data amount, wherein the state of the source mount point is not writable;

[0095] When the migration of the data to be migrated is completed, the migrated data of the target mount point in the second sub-cluster is determined;

[0096] Compare the data to be migrated with the migrated data to obtain the comparison result;

[0097] According to the comparison result, the source mount point status is adjusted from a non-writable state to a writable state;

[0098] Delete the data to be migrated in the first subcluster.

[0099] The unwritable state may refer to a state in which data cannot be written to the source mount point.

[0100] The method for comparing the data to be migrated with the migrated data is as follows: using the diff function in the distributed replication program distcp provided by Hadoop, a snapshot difference report is generated to check whether the data of the source mount point and the target mount point are consistent.

[0101] The method for adjusting the state of a source mount point from a non-writable state to a writable state may include: adjusting configuration data of the source mount point to adjust the state of the source mount point from a non-writable state to a writable state.

[0102] In the embodiment of the present application, if the amount of data to be migrated is greater than the preset migration data amount, the data to be migrated is copied to the target mount point of the second sub-cluster according to the preset batch migration data amount, including:

[0103] Determine the current batch data of the data to be migrated;

[0104] According to the source mount point and the target mount point, copy the current batch data of the data to be migrated to the target mount point of the second sub-cluster to obtain first migrated data;

[0105] After the current batch of data is copied, the source mount point status is adjusted to a non-writable state;

[0106] Determine the remaining data to be migrated;

[0107] According to the source mount point and the target mount point, copy the remaining data of the data to be migrated to the target mount point of the second sub-cluster to obtain second migrated data;

[0108] The migrated data is obtained according to the first migrated data and the second migrated data.

[0109] The current batch data may refer to the existing data in the current source mount point.

[0110] The method of copying the current batch of data to be migrated to the target mount point of the second sub-cluster may include: copying the current batch of data to the target mount point of the second sub-cluster by copying and pasting.

[0111] The method for adjusting the source mount point state to a non-writable state may include: adjusting the configuration data of the source mount point, adjusting the source mount point state from a read-write state to a non-writable state, and the purpose of adjusting the source mount point state is that all the source mounted data needs to be copied, so the source mount point needs to be set to a non-writable state before copying to ensure that no data is missed.

[0112] The method for determining the remaining data of the data to be migrated may include: determining the remaining data of the data to be migrated by using diff / rdiff / append of distcp.

[0113] The method of copying the remaining data of the data to be migrated to the target mount point of the second sub-cluster may include: performing incremental copying based on the newly added changes of the data copied for the first time by using the diff / rdiff / append / update function of distcp.

[0114] According to the first migrated data and the second migrated data, a method for obtaining the migrated data may include: adding the first migrated data and the second migrated data to obtain the migrated data.

[0115] By migrating the data to be migrated in preset batches, the source mount point can still be written to during the first copy to avoid affecting the use of the business. After the first copy is completed, the remaining data is copied. At this time, the amount of remaining data is much smaller than that of the first copy, and the copying time will be shortened accordingly. Therefore, the time for setting the source mount point to a non-writable state can be further shortened, further reducing the impact of data migration on business use.

[0116] In the embodiment of the present application, deleting the data to be migrated in the first sub-cluster includes:

[0117] Determine the mount table corresponding to the data to be migrated;

[0118] The mount table corresponding to the data to be migrated updates the source mount point to the target mount point;

[0119] Set the waiting time for deleting the data to be migrated;

[0120] The data to be migrated is deleted from the first subcluster according to the waiting time for deleting the data to be migrated.

[0121] The mount table may refer to a mount table in an HDFS RBF federated cluster, which is used to configure the data write path of a router. The router is a router in an HDFS RBF federated cluster. A user connects to the router through a client. The router determines the mount point path of the file through the mount table and performs read and write operations on the mount point path.

[0122] The waiting time for deleting the data to be migrated can be a countdown. When the countdown is up, the data to be migrated is deleted. Although the source mount point cannot be written, it can be read. Other applications need to read data from the source mount point. Therefore, setting the waiting time for deleting the data to be migrated can enable other applications to read the data of the source mount point normally, avoiding affecting the use of services.

[0123] The method for updating the source mount point to the target mount point may include: in the mount table, updating the source mount point information to the target mount point information, so that the router stores the newly arrived data to the target mount point through the updated mount table.

[0124] According to the waiting time for deleting the data to be migrated, the method for deleting the data to be migrated from the first subcluster may include: after setting the waiting time for deleting the data to be migrated, deleting the data to be migrated after the specified time is reached through cronservice (scheduled processing service) in Hadoop (distributed system infrastructure).

[0125] In the embodiment of the present application, after determining the source mount point of the first sub-cluster, the data to be migrated of the source mount point, and the target mount point of the second sub-cluster, the method further includes:

[0126] If the amount of data to be migrated is less than or equal to the preset migration data amount, the source mount point state is adjusted to a non-writable state;

[0127] According to the source mount point and the target mount point, the data to be migrated is copied to the target mount point of the second sub-cluster to obtain the migrated data.

[0128] Among them, in the embodiment of the present application, it also includes:

[0129] When the data density difference of the first sub-cluster is less than a first threshold, stop migrating the data in the first sub-cluster to the second sub-cluster;

[0130] When the cluster data density difference of the second subcluster is greater than the third threshold, the data in the first subcluster is migrated to the third subcluster, the subcluster data density difference of the third subcluster is less than the fourth threshold, and the third threshold is greater than the fourth threshold.

[0131] Among them, in the embodiment of the present application, it also includes:

[0132] According to a preset reporting mechanism, real-time execution status information of the data migration system is sent to the state storage, where the real-time execution status information is used to characterize the execution status of the data migration method;

[0133] receiving an operation instruction generated by the state memory according to the real-time execution state information, the operation instruction including a continue execution instruction indicating that the execution state of the data migration method is normal and a stop execution instruction indicating that the execution state of the data migration method is abnormal;

[0134] Control the data migration system according to operating instructions.

[0135] The reporting mechanism may refer to a heartbeat mechanism, which is a mechanism for sending a custom structure (heartbeat packet) at regular intervals to let the other party know that he is still alive to ensure the validity of the connection. In the embodiment of the present application, the heartbeat packet of the heartbeat mechanism is real-time execution status information.

[0136] The state store (StateStore) may refer to the StateStore in Hadoop, which is used to collect resource information of each process distributed on the cluster. In the embodiment of the present application, the StateStore is used to collect information and then supervise the data migration program.

[0137] The real-time execution status information may include execution success and execution failure. In each step of data migration, the execution result of a step is sent to the state storage after the step is executed.

[0138] The execution continuation instruction may refer to an instruction fed back to the data migration system when the real-time execution status information received by the status memory is information of successful execution.

[0139] The stop execution instruction may refer to an instruction fed back to the data migration system when the real-time execution status information received by the status memory is information of execution failure, or the real-time execution status information is not received within a preset time.

[0140] The present application provides a data migration method, which determines the data density difference in each subcluster in a federated cluster, the data density difference of the subcluster is determined according to the target data density and the subcluster data density of the subcluster, the target data density is determined according to the total storage volume and the overall storage usage of the federated cluster, and the subcluster data density is determined according to the total subcluster storage volume and the subcluster storage usage of the subcluster; according to the data density difference of each subcluster in the federated cluster, determines the first subcluster and the second subcluster in the federated cluster, wherein the data density difference of the first subcluster is greater than a first threshold, the data density difference of the second subcluster is less than a second threshold, and the first threshold is greater than the second threshold; the first subcluster is migrated to the The method uses a method to migrate the data from the first sub-cluster to the second sub-cluster, and compares the data density difference of each sub-cluster with the first threshold and the second threshold to confirm the data density status of each sub-cluster, so as to determine the sub-cluster that needs to migrate data out and the sub-cluster that needs to migrate data in. When the data volume of the sub-cluster exceeds the first threshold, the data migration program is triggered to avoid manual judgment, and the storage data volume between each sub-cluster can be adjusted in time, which improves the overall use effect of the federated cluster. The batch migration method is used to migrate data that exceeds the preset data volume, and the time that the source mount point of large data volume cannot be written during the migration process is reduced, thereby reducing the impact of federated cluster data rebalancing on data reading and writing.

[0141] Figure 3 A flowchart of another data migration method provided for this application is shown in FIG. Figure 3 As shown, the method includes:

[0142] S301 : Determine data distribution uniformity based on sub-cluster data density, and determine a sub-cluster to be migrated and a sub-cluster to be migrated into.

[0143] The determination of data distribution uniformity is calculated in the following way:

[0144] Assume that the total storage capacity of the federated cluster is total_capacity and the total storage usage of the federated cluster is used_capacity. Then the ideal data density (ideal_storage) is:

[0145] ideal_storage=used_capacity / total_capacity

[0146] Assume that the storage usage of the sub-cluster is sub_used_radio and the data density of the sub-cluster is sub_data_density:

[0147] sub_data_density=|ideal_storage-sub_used_radio|

[0148] The cluster provides the configuration parameter rbf.balance.threshold.percent, which triggers the data migration process when the following conditions are met:

[0149] |ideal_storage-sub_data_density|>rbf.balance.threshold.percent.

[0150] S302: Determine the source mount point and target NameService of the subcluster to be migrated, and the target path of the subcluster to be migrated, according to the subcluster to be migrated and the subcluster to be migrated.

[0151] The source mount point is the source path (mount point) where data migration is expected.

[0152] The target NameService is the NameService of the subcluster that needs to perform data migration;

[0153] The target path is the full path of the subcluster to be migrated.

[0154] S303, creating a pipeline process according to the source mount point, the target NameService, and the target path, and storing the source mount point in the StateStore;

[0155] Among them, for one source mount point, only one rebalance process is allowed to be executed. If multiple rebalance processes are started, except for the first one, the other processes will be disconnected automatically.

[0156] S304: Verify the source mount point according to a preset verification scheme.

[0157] The verification scheme may refer to verifying whether a source mount point exists, whether a sub-mount point exists for a source mount point, and whether a source path and a target path are different sub-clusters.

[0158] S305: If the verification result is passed, the data of the sub-cluster to be migrated is migrated to the sub-cluster to be migrated into according to the source mount point, the target NameService, and the target path.

[0159] Among them, in the embodiment of the present application, if the verification result is passed, the data of the sub-cluster to be migrated is migrated to the sub-cluster to be migrated into according to the source mount point, the target NameService, and the target path, including:

[0160] If the verification result is passed, the data of the subcluster to be migrated is copied to the subcluster to be migrated based on the source mount point, target NameService, and target path;

[0161] Verify that the data of the source mount point and the target path are consistent;

[0162] If the data of the source mount point and the target path are consistent, the pipeline process closing instruction is executed.

[0163] Among them, in the embodiment of the present application, if the verification result is passed, the data of the sub-cluster to be migrated is copied to the sub-cluster to be migrated according to the source mount point, the target NameService, and the target path, including:

[0164] This step sets two options to ensure data consistency during data balancing and business awareness:

[0165] 1) Option 1 (NO_WRITE): This parameter can be used for paths with small data volumes. This parameter sets the source mount point directory to read-only in the StateStore before data migration begins, ensuring that no write operations are performed during data migration and data consistency is guaranteed.

[0166] 2) Option 2 (COPY_TWICE): If a data directory has a large amount of data and a data copy takes more than several hours, if the source mount point directory is directly set to read-only, the business cannot be carried out normally during the data migration. To solve this problem, this option is introduced. The function of this parameter is to divide the data migration into two steps. In the first step, the business can read and write the source directory normally until the first data copy is completed; in the second step, the source directory is set to read-only in the StateStore, and the diff / rdiff / append / update function of distcp is used to perform incremental copy based on the new changes of the first copied data. Through the above two steps, the data migration can minimize the impact on the business;

[0167] At the same time, an expiration time is set for the read-only lock during data replication, and the expiration time of the read-only lock needs to be continuously updated during the execution of the rebalance process. This prevents the data balancing process from crashing and causing the source path mount point to remain in a read-only state.

[0168] In the embodiment of the present application, verifying whether the data of the source mount point and the target path are consistent includes:

[0169] When this step is executed, the table name data has been successfully copied from the source path to the target path, and the read-only lock of the source mount point is still not released; at this time, use the diff function of distcp again to generate a snapshot difference report to check whether the data of the source mount point and the target path are consistent. If they are consistent, release the read-only lock, otherwise repeat the diff / rdiff / append / update operation.

[0170] Among them, in the embodiment of the present application, if the data of the source mount point and the target path are consistent, the pipeline process closing instruction is executed, including:

[0171] 1) Update the mount point: Update the MountTable information in the StateStore, and update the source mount point to the target path; wait for all Routers to complete synchronization before proceeding to the next step;

[0172] 2) Delete data in the source directory: In this step, you can set a waiting time for deleting source data to prevent the failure of the ongoing read operation of the business. After the set waiting time is reached, the data will be deleted. After the deletion waiting time is set, the data will be retained and the pipeline process will be set to the DELETE PENDING state. At the same time, the cron service will be started to delete the data after the specified time is reached.

[0173] 3) Delete the mount information of the source mount point: After the data is deleted, delete the mount information of the source mount point in the StateStore.

[0174] In the embodiment of the present application, the method further includes:

[0175] The execution status and results of each step will be reported to StateStore in the form of heartbeats. If StateStore receives a failure message or loses the heartbeat, it is considered that the data migration has failed. At this time, the Rollback process is started to restore the data to its original state.

[0176] Another data migration method provided by the present application determines the uniformity of data distribution through the sub-cluster data density, determines the sub-cluster to be migrated and the sub-cluster to be migrated into, determines the source mount point and target NameService of the sub-cluster to be migrated and the target path of the sub-cluster to be migrated into according to the sub-cluster to be migrated and the sub-cluster to be migrated into, creates a pipeline process according to the source mount point, the target NameService, and the target path, stores the source mount point in the StateStore, verifies the source mount point according to a preset verification scheme, and if the verification result is passed, migrates the data of the sub-cluster to be migrated to the sub-cluster to be migrated into according to the source mount point, the target NameService, and the target path, thereby achieving the purpose of automatically triggering the sub-cluster data migration, avoiding affecting the business use of the sub-cluster, and thus improving the use effect of the sub-cluster.

[0177] Figure 4 A schematic diagram of the structure of a data migration device provided by the present application, the device includes a first determination module 401 、 Second determination module 402 、 以及 Migration module 403, wherein:

[0178] A first determination module 401 is used to determine a data density difference in each sub-cluster in the federated cluster, the sub-cluster data density difference is determined according to a target data density and a sub-cluster data density of the sub-cluster, the target data density is determined according to a total storage amount and an overall storage usage of the federated cluster, and the sub-cluster data density is determined according to a total sub-cluster storage amount and a sub-cluster storage usage of the sub-cluster;

[0179] A second determination module 402 is used to determine a first subcluster and a second subcluster in the federated cluster according to a data density difference of each subcluster in the federated cluster, wherein the data density difference of the first subcluster is greater than a first threshold, the data density difference of the second subcluster is less than a second threshold, and the first threshold is greater than the second threshold;

[0180] The migration module 403 is configured to migrate the data of the first sub-cluster to the second sub-cluster.

[0181] In this embodiment of the present application, the first determining module 401 is further used to:

[0182] The target data density is obtained by calculating the ratio of the total storage usage to the total storage volume.

[0183] The sub-cluster storage usage and the sub-cluster storage total amount are ratioed to obtain the sub-cluster data density;

[0184] The sub-cluster data density and the target data density are subtracted to determine the data density difference in each sub-cluster in the federated cluster.

[0185] In this embodiment of the present application, the migration module 403 is also used for:

[0186] Determine the source mount point of the first subcluster, the data to be migrated from the source mount point, and the target mount point of the second subcluster;

[0187] If the amount of data to be migrated is greater than the preset migration data amount, the data to be migrated is migrated to the target mount point of the second sub-cluster according to the preset batch migration data amount.

[0188] In this embodiment of the present application, the migration module 403 is also used for:

[0189] If the amount of data to be migrated is greater than the preset migration data amount, the data to be migrated is copied to the target mount point of the second sub-cluster according to the preset batch migration data amount, wherein the state of the source mount point is not writable;

[0190] When the migration of the data to be migrated is completed, the migrated data of the target mount point in the second sub-cluster is determined;

[0191] Compare the data to be migrated with the migrated data to obtain the comparison result;

[0192] According to the comparison result, the source mount point status is adjusted from a non-writable state to a writable state;

[0193] Delete the data to be migrated in the first subcluster.

[0194] In this embodiment of the present application, the migration module 403 is also used for:

[0195] Determine the current batch data of the data to be migrated;

[0196] According to the source mount point and the target mount point, copy the current batch data of the data to be migrated to the target mount point of the second sub-cluster to obtain first migrated data;

[0197] After the current batch of data is copied, the source mount point status is adjusted to a non-writable state;

[0198] Determine the remaining data to be migrated;

[0199] According to the source mount point and the target mount point, copy the remaining data of the data to be migrated to the target mount point of the second sub-cluster to obtain second migrated data;

[0200] The migrated data is obtained according to the first migrated data and the second migrated data.

[0201] In this embodiment of the present application, the migration module 403 is also used for:

[0202] Determine the mount table corresponding to the data to be migrated;

[0203] The mount table corresponding to the data to be migrated updates the source mount point to the target mount point;

[0204] Set the waiting time for deleting the data to be migrated;

[0205] The data to be migrated is deleted from the first subcluster according to the waiting time for deleting the data to be migrated.

[0206] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 5 As shown, the electronic device 50 includes:

[0207] The electronic device 50 may include a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a communication component 503 and other components. The processor 501 , the memory 502 and the communication component 503 are connected via a bus 504 .

[0208] In a specific implementation process, at least one processor 501 executes the computer execution instructions stored in the memory 502, so that at least one processor 501 executes the above data migration method.

[0209] The specific implementation process of the processor 501 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0210] In the above Figure 5 In the illustrated embodiment, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the invention may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0211] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.

[0212] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.

[0213] In some embodiments, a computer program product is also proposed, including a computer program or instructions, which implement the steps in any of the above-mentioned data migration methods when executed by a processor.

[0214] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0215] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0216] To this end, an embodiment of the present application provides a computer-readable storage medium, in which multiple instructions are stored. The instructions can be loaded by a processor to execute the steps in any data migration method provided in the embodiment of the present application.

[0217] The storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0218] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program comprises computer instructions stored in a computer-readable storage medium.

[0219] Since the instructions stored in the storage medium can execute the steps in any data migration method provided in the embodiments of the present application, the beneficial effects that can be achieved by any data migration method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0220] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0221] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A data migration method, characterized in that: Applied to a data migration system, the method includes: Determine a data density difference in each subcluster in the federated cluster, the subcluster data density difference being determined based on a target data density and a subcluster data density of the subcluster, the target data density being determined based on a total storage amount and an overall storage usage of the federated cluster, the subcluster data density being determined based on a total subcluster storage amount and a subcluster storage usage of the subcluster; Determine a first subcluster and a second subcluster in the federated cluster according to a data density difference of each subcluster in the federated cluster, wherein the data density difference of the first subcluster is greater than a first threshold, the data density difference of the second subcluster is less than a second threshold, and the first threshold is greater than the second threshold; Migrate the data of the first subcluster to the second subcluster.

2. The method according to claim 1, characterized in that The determining of the data density difference in each sub-cluster in the federated cluster includes: Performing ratio processing on the overall storage usage and the total storage amount to obtain the target data density; Calculating the ratio of the sub-cluster storage usage to the sub-cluster storage total amount to obtain the sub-cluster data density; The sub-cluster data density and the target data density are subjected to difference processing to determine the data density difference in each sub-cluster in the federated cluster.

3. The method according to claim 1, characterized in that The migrating the data of the first subcluster to the second subcluster includes: Determine a source mount point of the first subcluster, data to be migrated of the source mount point, and a target mount point of the second subcluster; If the data volume of the data to be migrated is greater than the preset migration data volume, the data to be migrated is migrated to the target mounting point of the second sub-cluster according to the preset batch migration data volume.

4. The method according to claim 3, characterized in that If the amount of the data to be migrated is greater than the preset migration data amount, migrating the data to be migrated to the target mount point of the second sub-cluster according to the preset batch migration data amount, including: If the amount of the data to be migrated is greater than the preset migration data amount, the data to be migrated is copied to the target mount point of the second sub-cluster according to the preset batch migration data amount, wherein the state of the source mount point is a non-writable state; When the migration of the to-be-migrated data is completed, determining the migrated data of the target mount point in the second sub-cluster; Comparing the data to be migrated with the migrated data to obtain a comparison result; According to the comparison result, adjusting the state of the source mount point from a non-writable state to a writable state; The data to be migrated in the first subcluster is deleted.

5. The method according to claim 4, characterized in that If the amount of the data to be migrated is greater than the preset migration data amount, copying the data to be migrated to the target mount point of the second sub-cluster according to the preset batch migration data amount includes: Determine the current batch data of the data to be migrated; According to the source mount point and the target mount point, copy the current batch data of the data to be migrated to the target mount point of the second sub-cluster to obtain first migrated data; After the current batch of data is copied, the source mount point state is adjusted to a non-writable state; Determining remaining data of the data to be migrated; According to the source mount point and the target mount point, copy the remaining data of the data to be migrated to the target mount point of the second sub-cluster to obtain second migrated data; The migrated data is obtained according to the first migrated data and the second migrated data.

6. The method according to claim 4, characterized in that The deleting the to-be-migrated data in the first sub-cluster includes: Determine a mount table corresponding to the data to be migrated; The mount table corresponding to the data to be migrated updates the source mount point to the target mount point; Set the waiting time for deleting the data to be migrated; The data to be migrated is deleted from the first sub-cluster according to the waiting time for deleting the data to be migrated.

7. The method according to claim 3, characterized in that After determining the source mount point of the first sub-cluster, the data to be migrated of the source mount point, and the target mount point of the second sub-cluster, the method further includes: If the amount of the data to be migrated is less than or equal to the preset amount of migration data, adjusting the state of the source mount point to a non-writable state; According to the source mount point and the target mount point, the data to be migrated is copied to the target mount point of the second sub-cluster to obtain migrated data.

8. The method according to claim 1, characterized in that The method further comprises: When the data density difference of the first sub-cluster is less than the first threshold, stop migrating the data in the first sub-cluster to the second sub-cluster; When the cluster data density difference of the second subcluster is greater than a third threshold, data in the first subcluster is migrated to a third subcluster, the subcluster data density difference of the third subcluster is less than a fourth threshold, and the third threshold is greater than the fourth threshold.

9. The method according to claim 1, characterized in that: The method further comprises: According to a preset reporting mechanism, real-time execution status information of the data migration system is sent to a status memory, wherein the real-time execution status information is used to represent the execution status of the data migration method; receiving an operation instruction generated by the state memory according to the real-time execution state information, the operation instruction comprising a continue execution instruction indicating that the execution state of the data migration method is normal and a stop execution instruction indicating that the execution state of the data migration method is abnormal; According to the operation instruction, the data migration system is controlled.

10. A data migration device, characterized in that: Applied to a data migration system, the device comprises: a first determination module, configured to determine a data density difference in each subcluster in a federated cluster, wherein the subcluster data density difference is determined according to a target data density and a subcluster data density of the subcluster, wherein the target data density is determined according to a total storage amount and an overall storage usage of the federated cluster, and the subcluster data density is determined according to a total subcluster storage amount and a subcluster storage usage of the subcluster; A second determination module is configured to determine a first subcluster and a second subcluster in the federated cluster according to a data density difference of each subcluster in the federated cluster, wherein the data density difference of the first subcluster is greater than a first threshold, the data density difference of the second subcluster is less than a second threshold, and the first threshold is greater than the second threshold; A migration module is used to migrate the data of the first sub-cluster to the second sub-cluster.

11. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 9 when executed by a processor.

Citation Information

Patent Citations

  • Migration strategy adjustment method as well as capacity change suggestion method and migration strategy adjustment device

    CN106502576A

  • HDFS-based data equalization optimization method, system terminal and storage medium

    CN110928836A

  • Data migration method based on big data

    CN112328539A

  • Data storage method and device, storage medium and electronic equipment

    CN117008817A

  • Determining hardware requirements for a wireless network event using crowdsourcing

    US10356552B1

Cited By

  • File management system and file management method

    US20250328502A1