Data migration method and device, computer program product and electronic equipment
By obtaining migration configuration information and metadata analysis results to generate a migration catalog, perform verification, and create migration tasks, the problems of low accuracy and efficiency of cross-cluster data migration tools are solved, an automated and standardized data migration process is implemented, and data integrity and security are ensured.
Patent Information
- Application Number
- CN202510808566.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-30
AI Technical Summary
Existing cross-cluster data migration tools have low accuracy and efficiency, and require extremely high labor costs, which affects the application of cross-cluster data migration.
By obtaining migration configuration information and predetermined metadata analysis results, generating a migration catalog, and performing correctness and data uniqueness verification, a migration task is created. The server is responsible for task scheduling and verification, and the client executes the migration task, thus achieving an automated and standardized data migration process.
It improves the accuracy and efficiency of data migration, reduces the risk of errors caused by human intervention, ensures the integrity and security of data after migration, and adapts to cluster environments of different sizes.
Smart Images

Figure CN120723745A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and more particularly, to a data migration method, a data migration device, a computer program product, and an electronic device. Background Art
[0002] With the rapid development of computers, the storage cost of business data in big data businesses continues to rise, along with increasing business complexity and the deepening application of the lake-warehouse integrated architecture. For example, in HDFS (Hadoop Distributed File System), the value and access frequency of data show significant decline over time. In the lake-warehouse integrated architecture, hot clusters typically use SSDs (Solid State Drives) or high-speed disks to ensure low-latency response for frequently accessed data. However, as the data lifecycle evolves, large amounts of low-access data occupy high-performance storage, resulting in significant resource waste. Warm clusters achieve cost optimization through storage media costs, changes in the number of HDFS replicas, and data compression methods.
[0003] Therefore, migrating low-access data from hot clusters to warm clusters can significantly reduce total storage costs while maintaining business continuity. Currently, cross-cluster data migration often uses distributed data copy tools. However, using these tools directly can be inaccurate and inefficient, and requires significant labor costs, which hinders the application of cross-cluster data migration.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0005] The purpose of the present disclosure is to provide a data migration method and apparatus, a computer program product, and an electronic device, thereby improving the accuracy and efficiency of data migration and reducing migration costs, at least to a certain extent.
[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.
[0007] According to one aspect of the present disclosure, a data migration method is provided, which is applied to a server, and the method includes: obtaining migration configuration information; obtaining a migration directory based on the migration configuration information and predetermined metadata analysis results, where the metadata analysis results are obtained by analyzing and processing the metadata of the source cluster; verifying the correctness and data uniqueness of the migration directory, and if the migration directory passes the verification, creating a migration task based on the migration directory; sending the migration task to a client, so that the client executes the migration task and migrates the data in the source cluster to the target cluster.
[0008] In an exemplary embodiment of the present disclosure, obtaining a predetermined metadata analysis result includes: obtaining the metadata of the source cluster stored in the name node; analyzing and processing the directory file information in the metadata of the source cluster, obtaining the metadata analysis result and storing it in the target database, and the target database provides an interface for data query.
[0009] In an exemplary embodiment of the present disclosure, the method further includes: reconstructing metadata based on a metadata snapshot to obtain reconstructed metadata; obtaining an edit log corresponding to the client operation; comparing and analyzing the metadata of the source cluster based on the reconstructed metadata and / or the edit log to obtain operation change information, and updating the metadata analysis results based on the operation change information.
[0010] In an exemplary embodiment of the present disclosure, the migration configuration information includes the path to be migrated and the corresponding time range information; based on the migration configuration information and the predetermined metadata analysis results, a migration directory is obtained, including: based on the metadata analysis results, obtaining the sub-path information under the path to be migrated, and determining the modification time of the sub-path information; based on the modification time and time range information, obtaining the target sub-path information corresponding to the path to be migrated to obtain the migration directory.
[0011] In an exemplary embodiment of the present disclosure, the correctness and data uniqueness of the migration directory are verified, including: detecting whether the directory of the source cluster includes the migration directory, and obtaining a detection result; detecting whether there is a migration task with the migration directory in the historical migration tasks, and obtaining a detection result; detecting whether there are other migration tasks that have a path inclusion relationship with the migration directory, and obtaining a detection result; determining a first verification result based on the detection result, wherein the first verification result is used to indicate whether the correctness and data uniqueness verification of the migration directory have passed.
[0012] In an exemplary embodiment of the present disclosure, it is detected whether there are other migration tasks that have a path inclusion relationship with the migration directory to obtain a detection result, including: dividing the migration directory according to the directory hierarchy to obtain multiple target directories; and querying whether there are other migration tasks that have a path inclusion relationship according to each target directory to obtain a detection result.
[0013] In an exemplary embodiment of the present disclosure, if the migration directory passes the verification, a migration task is created according to the migration directory, including: determining the task to be migrated according to the migration directory; supplementing the content of the task to be migrated according to at least one of the task identification information, cluster identification information, and scheduling cycle information in the migration configuration information to obtain the supplemented task to be created; adding task status information to the supplemented task to be created to obtain the migration task, wherein the task status information is used to characterize the current task status of the migration task.
[0014] In an exemplary embodiment of the present disclosure, the migration task includes multiple sub-directories, and the method further includes: determining the current available bandwidth based on the total bandwidth and the occupied bandwidth; comparing the total bandwidth required for a single migration of the multiple sub-directories with the current available bandwidth to obtain a first comparison result, and the first comparison result is used to indicate whether the migration task is currently allowed to be executed; and scheduling the client to execute the migration task based on the first comparison result.
[0015] In an exemplary embodiment of the present disclosure, the method also includes: storing the data in the source cluster to be migrated to the target cluster in a logical deletion buffer, the logical deletion buffer is used to store data to be cleaned up during the migration process but the possibility of recovery must be retained; based on a preset period, the logical deletion buffer is cleaned up regularly.
[0016] In an exemplary embodiment of the present disclosure, the method further includes: determining a directory to be restored in response to a data recovery request; when no data exists in the directory to be restored under the directory of the source cluster, logically deleting the path of the directory to be restored in the buffer and moving it to the directory of the source cluster.
[0017] According to one aspect of the present disclosure, a data migration method is provided, which is applied to a client. The method includes: obtaining a migration task, where the migration task is constructed according to any of the above-mentioned data migration methods, and the migration task includes multiple sub-directories; constructing migration plans for the multiple sub-directories respectively to obtain multiple migration plans; and executing the multiple migration plans to migrate data in a source cluster to a target cluster.
[0018] In an exemplary embodiment of the present disclosure, multiple migration plans are executed to migrate data in a source cluster to a target cluster, including: for each migration plan, creating a temporary storage path for the target cluster according to the migration plan; obtaining migration constraint information according to the migration configuration information, and migrating the data corresponding to the migration plan to the temporary storage path according to the migration constraint information, wherein the migration constraint information is used to limit the transmission bandwidth and / or transmission volume when executing the migration plan; performing a data consistency check on the migration plan, and if the migration plan passes the check, moving the data in the temporary storage path to the target path of the target cluster, wherein the data consistency check is used to characterize the consistency of the data before and after the execution of the migration plan.
[0019] In an exemplary embodiment of the present disclosure, a data consistency check is performed on the migration plan, including: for each migration plan, obtaining whether the logical storage space occupied by the data before and after the execution of the migration plan is consistent, and obtaining a verification result; and / or obtaining whether the total hash value of the data before and after the execution of the migration plan is consistent, and obtaining a verification result; determining a second check result based on the verification result, and the second check result is used to indicate whether the data consistency check of the migration plan has passed.
[0020] In an exemplary embodiment of the present disclosure, if a target migration plan fails the data consistency check, the source cluster and the target cluster are rolled back to restore to the state before the migration task is executed; and the temporary storage path is cleared.
[0021] In an exemplary embodiment of the present disclosure, multiple migration plans are executed to migrate data in a source cluster to a target cluster, including: for the migration plan currently to be migrated, the required single migration bandwidth is compared with the currently available bandwidth to obtain a second comparison result, and the second comparison result is used to indicate whether the migration plan currently to be migrated is allowed to be executed; and the migration plan currently to be migrated is controlled to wait or stop according to the second comparison result.
[0022] According to one aspect of the present disclosure, a data migration device is provided, which is applied to a server, and the device includes: a path configuration module for obtaining migration configuration information; a data processing module for obtaining a migration directory based on the migration configuration information and a predetermined metadata analysis result, where the metadata analysis result is obtained by analyzing and processing the metadata of the source cluster; a verification module for verifying the correctness and data uniqueness of the migration directory, and if the migration directory passes the verification, creating a migration task based on the migration directory; and a task generation module for sending the migration task to the client, so that the client executes the migration task and migrates the data in the source cluster to the target cluster.
[0023] According to one aspect of the present disclosure, a data migration device is provided, which is applied to a client. The device includes: a task acquisition module, which is used to obtain a migration task, where the migration task is constructed by any of the above-mentioned data migration methods, and the migration task includes multiple sub-directories; a plan construction module, which is used to construct migration plans for the multiple sub-directories respectively, to obtain multiple migration plans; and a data migration module, which is used to execute the multiple migration plans to migrate data in the source cluster to the target cluster.
[0024] According to one aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, any one of the above methods is implemented.
[0025] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any one of the above methods by executing the executable instructions.
[0026] The data migration method in the exemplary embodiment of the present disclosure obtains migration configuration information, obtains a migration directory based on the migration configuration information and predetermined metadata analysis results, the metadata analysis results are obtained by analyzing and processing the metadata of the source cluster, and then the migration directory is verified for correctness and data uniqueness. If the migration directory passes the verification, a migration task is created based on the migration directory; finally, the migration task is sent to the client to enable the client to execute the migration task and migrate the data in the source cluster to the target cluster.
[0027] On the one hand, users only need to provide migration configuration information, and the migration catalog is automatically generated by matching the migration configuration information with the metadata analysis results, avoiding the tedious manual configuration. The automated creation and distribution of migration tasks standardizes the migration process, reducing the risk of errors caused by human intervention, thereby ensuring the efficiency and accuracy of cross-cluster data migration. On the other hand, the migration catalog is generated based on the metadata analysis results of the source cluster, ensuring that the migration scope accurately corresponds to the source data structure, avoiding omissions or redundancies. Correctness and data uniqueness verification prevent duplicate data or data corruption in the target cluster, ensuring the complete availability of data after migration. Furthermore, the pre-emptive verification phase (completed before migration) can identify potential issues in advance, reduce abnormal operations during the migration process, and further ensure the security and effectiveness of data migration. Furthermore, the server is responsible for task scheduling and verification, while the client performs the actual migration, achieving load balancing, avoiding server-side performance bottlenecks, and adapting to migration tasks in cluster environments of varying sizes.
[0028] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:
[0030] Figure 1 A system architecture diagram of an exemplary embodiment of the present disclosure is shown.
[0031] Figure 2 A flow chart of data migration based on a distributed data copy tool according to an exemplary embodiment of the present disclosure is shown.
[0032] Figure 3 A flow chart of a data migration method according to an exemplary embodiment of the present disclosure is shown.
[0033] Figure 4A flowchart of updating metadata analysis results according to an exemplary embodiment of the present disclosure is shown.
[0034] Figure 5 A schematic diagram showing an analysis result of updated source data according to an exemplary embodiment of the present disclosure is shown.
[0035] Figure 6 A flowchart illustrating an implementation method for obtaining a migration directory according to an exemplary embodiment of the present disclosure is shown.
[0036] Figure 7 A flowchart of checking the correctness and data uniqueness of a migration directory according to an exemplary embodiment of the present disclosure is shown.
[0037] Figure 8 A flow chart of a server-side current limiting method according to an exemplary embodiment of the present disclosure is shown.
[0038] Figure 9 A flow chart of yet another data migration method according to an exemplary embodiment of the present disclosure is shown.
[0039] Figure 10 A flowchart of executing a migration plan according to an exemplary embodiment of the present disclosure is shown.
[0040] Figure 11 A schematic diagram of the composition of a data migration device according to an exemplary embodiment of the present disclosure is shown.
[0041] Figure 12 A schematic diagram showing the composition of yet another data migration device according to an exemplary embodiment of the present disclosure is shown.
[0042] Figure 13 A block diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown.
[0043] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION
[0044] The exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these examples are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the exemplary embodiments to those skilled in the art. Identical reference numerals in the figures represent identical or similar structures, and thus detailed descriptions thereof will be omitted.
[0045] In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known structures, methods, devices, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0046] The blocks shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. Specifically, these functional entities may be implemented in software, or in one or more software-hardened modules, or in different networks and / or processor devices and / or microcontroller devices.
[0047] like Figure 1 Shown is a system architecture diagram of an exemplary embodiment of the present disclosure. The source cluster is a cluster of data to be migrated, and its number is not limited to 1, such as a hot cluster, which is used to process the latest and most active data. These data are often written and read frequently, so high-performance hardware is required to support fast indexing and search operations. The target cluster is a cluster used to receive source cluster data, such as a warm cluster, which stores relatively less active data that still needs to be accessed regularly. This type of data no longer changes frequently, but still needs to be queried. Cluster routing refers to the combination of multiple router instances, such as a Router cluster, to jointly process network traffic through specific management and scheduling mechanisms to improve the overall performance and reliability of the system. The data middle platform refers to the use of data technology to collect, calculate, store, and process massive amounts of data while unifying standards and calibers.
[0048] Of course, the data migration method of the exemplary embodiment of the present disclosure is applicable to cross-cluster data migration, and does not limit the number and type of source clusters and target clusters. The description is only given by taking hot clusters and warm clusters as examples.
[0049] When using distributed data copy tools (such as DistCp) to copy data across clusters, the process may include initialization and job submission, MR (MapReduce job, a programming model) job execution, target path processing, etc. Figure 2 The figure shows a flow chart of data migration based on the distributed data copy tool, including initialization and job submission, MR job execution, and target path processing.
[0050] However, in actual production, the volume of data generated in tables or log files, the accumulation of historical data files, and the frequency of use of hot and cold data all require different amounts of data at different business stages. Relying solely on manual methods to control and record these complex paths is extremely labor-intensive and costly. Furthermore, this tool only replicates data from the source cluster to the target cluster. In actual production, this tool fails to implement a complete, standardized migration process to guarantee migration principles, which in turn impacts the efficiency and accuracy of cross-cluster migrations.
[0051] Based on this, exemplary embodiments of the present disclosure provide a data migration method that can be applied to cross-cluster data migration. For example, it can implement cross-cluster warm data migration in an HDFS system to facilitate data storage optimization and resource management, or multi-cloud / data cloud synchronization. Exemplary embodiments of the present disclosure are not specifically limited to this.
[0052] First, terms or concepts related to exemplary embodiments of the present disclosure are explained.
[0053] HDFS: Hadoop distributed file storage system, designed as a file system suitable for running on general-purpose hardware. It has high fault tolerance and can be deployed on cheap hardware. It is suitable for applications with large data sets.
[0054] HDFS NameNode: HDFS name node, the master server that manages HDFS metadata, is used to manage client operations on files and determine the specific locations of files and storage blocks.
[0055] Hive: A Hadoop-based data warehouse tool used to extract, transform, and load large-scale data in Hadoop. It can map structured data files into a database table and provide comprehensive SQL functions.
[0056] DisctCp: Also known as Distributed Copy, it is a high-performance copy tool for large-scale clusters or between clusters. It uses Map / Reduce to implement file distribution, error handling and recovery, and report generation.
[0057] refer to Figure 3 The flowchart of the data migration method of the exemplary embodiment of the present disclosure is shown, which is applied to the server. The data migration method includes steps S310 to S340, which are specifically as follows:
[0058] In step S310, migration configuration information is obtained.
[0059] In the exemplary embodiments of this disclosure, the server side may refer to the service corresponding to the data middleware. This service is user-facing, allowing users to configure data using an interface to obtain migration configuration information. The server side then constructs migration tasks based on this migration configuration information. This migration configuration information can be understood as the user's migration requirements for cross-cluster data migration. This configurability of data migration meets the user's personalized needs for cross-cluster data migration.
[0060] Among them, the migration configuration information may include but is not limited to the path to be migrated and the corresponding time range information, task identification information, cluster identification information (also known as Router Uri), scheduling cycle information, etc. The time range information is used to filter the target that the user needs to migrate. The cluster identification information may include the source cluster identification and the target cluster identification. It should be noted that the cluster identification information here is not the identification of the actual cluster, but can correspond to each actual cluster respectively. In other words, when operating a certain data, the data may be distributed in multiple paths. According to the cluster identification information, by querying the metadata mapping table, the specific cluster can be queried, and then the data of this part can be queried in the specific cluster. Therefore, in actual operation, by introducing the Router framework of Hadoop, there is no need to pay attention to the actual cluster, and it is only necessary to access the data based on the cluster identification information.
[0061] In step S320 , a migration directory is obtained according to the migration configuration information and a predetermined metadata analysis result, where the metadata analysis result is obtained by analyzing and processing the metadata of the source cluster.
[0062] In an exemplary embodiment of the present disclosure, the predetermined metadata analysis result may include the metadata of the source cluster itself and result data reconstructed by performing statistical processing based on the metadata of the source cluster.
[0063] HDFS natively provides a query interface for directory and file storage status information and the directory-file hierarchical relationship. However, this native interface requires recursive node computation of statistical information. The NameNode's read and write operations share a read-write lock. When there are many nodes, this can result in long-term blocking of other user requests, causing internal thread timeouts and cluster jitter. To address this issue, exemplary embodiments of the present disclosure analyze and process the metadata of the source cluster to obtain predetermined metadata analysis results. These results can then be directly utilized when performing required queries, improving data query capabilities.
[0064] The migration directory is a directory that is selected based on user requirements (migration configuration information) and the current latest path information of the source cluster (predetermined metadata analysis results) as the final directory to be migrated.
[0065] In step S330 , the correctness and data uniqueness of the migration directory are verified. If the migration directory passes the verification, a migration task is created based on the migration directory.
[0066] In the exemplary embodiments of the present disclosure, a correctness check verifies that entries in the migration catalog are legal and executable, while a data uniqueness check verifies whether there are any data conflicts between historical migration tasks and the migration catalog, including path and task conflicts. After passing both checks, a migration task can be created based on the migration catalog.
[0067] Optionally, the creation of migration tasks can be based on the permissions of the data middle platform. According to the specified core tasks and operators in the service context, it can be confirmed whether their roles or users have relevant operation permissions, that is, user permission verification. Only after passing the user permission verification can the migration task be created according to the migration directory.
[0068] Multiple checks can help avoid invalid migration and data contamination, and improve the accuracy of creating migration tasks.
[0069] In step S340 , the migration task is sent to the client, so that the client executes the migration task and migrates the data in the source cluster to the target cluster.
[0070] In an exemplary embodiment of the present disclosure, the client is the actual operator of cross-cluster migration and executes the migration process based on specific task scheduling of the server.
[0071] If there are multiple source clusters, migration tasks can be generated for each source cluster separately. These migration tasks can then be asynchronously and concurrently executed by multiple clients, achieving asynchronous cross-cluster data migration. During asynchronous execution, clients can be started asynchronously without waiting for task completion, and each migration task can be executed separately by concurrently launching multiple task instances. Corresponding thread pools can also be established for different source clusters, and data from different clusters can be executed concurrently through different thread pools.
[0072] In an exemplary embodiment of the present disclosure, on the one hand, the user only needs to provide migration configuration information, and the migration directory can be generated by automatically matching the migration configuration information with the metadata analysis results, thus avoiding the tedious manual configuration. The migration process is standardized through the automated creation and distribution of migration tasks, reducing the risk of errors caused by human intervention, thereby ensuring the efficiency and accuracy of cross-cluster data migration. On the other hand, the migration directory is generated based on the metadata analysis results of the source cluster to ensure that the migration scope accurately corresponds to the source data structure, avoiding omissions or redundancies, and through correctness verification and data uniqueness verification, duplicate data or data corruption can be prevented in the target cluster, ensuring that the data is complete and available after migration. At the same time, since the verification link is pre-placed (completed before migration), potential problems can be identified in advance, reducing abnormal operations during the migration process, and further ensuring the safety and effectiveness of data migration. In addition, the server is responsible for task scheduling and verification, and the client performs the actual migration, realizing load separation, avoiding server performance bottlenecks, and adapting to migration tasks in cluster environments of different sizes.
[0073] In an exemplary embodiment, a method for obtaining a predetermined metadata analysis result is provided, including:
[0074] First, the metadata of the source cluster stored in the name node is obtained, and then the directory file information in the metadata of the source cluster is analyzed and processed to obtain the metadata analysis results and store them in the target database. The target database provides an interface for data query.
[0075] The source cluster's metadata, originally stored on the name node, can be stored in the target database. This target database can be a distributed database, such as HBase, where read and write operations do not require shared read-write locks. Storing the source cluster's data in the target database can prevent prolonged service unavailability.
[0076] Furthermore, the target database's metadata directory file information is analyzed and processed. This target file information includes information about files within the directory, including but not limited to the number of files within the directory (calculated recursively and non-recursively), and the file sizes within the directory (calculated recursively and non-recursively). The target database provides an interface that allows direct data query during service queries, avoiding further analysis, processing, or locking operations, thereby improving query performance.
[0077] In an exemplary embodiment, considering that the read and write speed of memory may be higher than that of the target database (such as HBase), if the metadata is stored in the target database, the speed of analyzing and processing the original data may be lower than the production speed of log editing. Therefore, the following steps are optimized:
[0078] Step S410: Reconstruct metadata according to the metadata snapshot to obtain reconstructed metadata.
[0079] The metadata snapshot (also called fsimage) is a complete metadata snapshot (file directory tree, block mapping, etc.). It is a persistent metadata snapshot saved by NameNode, which records the complete directory structure of HDFS and the mapping relationship between files and data blocks, including static and full data. When the service starts, you can check whether it is necessary to obtain a metadata snapshot. If necessary, execute this step to obtain the metadata snapshot and rebuild the metadata accordingly. Among them, whether it is necessary to obtain a metadata snapshot depends on the scenario and task requirements. For example, metadata snapshots are obtained regularly according to a preset period, or when the Active NameNode is down and the metadata of the Standby NameNode is incomplete (for example, the edits log is lost or damaged), it is necessary to pull a metadata snapshot, or when the metadata of HDFS needs to be backed up offline, it is necessary to pull a metadata snapshot, and so on. Exemplary embodiments of the present disclosure include but are not limited to the timing of pulling metadata snapshots as mentioned above.
[0080] The metadata directory in the memory can be rebuilt according to the metadata snapshot, and the metadata state at this time is equal to the state when the snapshot was generated.
[0081] Step S420: Obtain the editing log corresponding to the client operation.
[0082] The edit log, also known as the edits log, records file system metadata changes to ensure data consistency and recoverability. When a client performs a write operation on HDFS (such as creating a file), in addition to recording the metadata of the operation in local memory, the operation is also submitted to the JournalNode cluster as an edits log. Therefore, the edit log is available from the JournalNode cluster.
[0083] Step S430: performing a comparison analysis on the metadata of the source cluster based on the reconstructed metadata and / or the edit log to obtain operation change information, and updating the metadata analysis result based on the operation change information.
[0084] Operational change information is metadata differences obtained through comparative analysis. When both reconstructed metadata and edit logs are acquired, all edit logs subsequent to the metadata snapshot can be scanned—that is, replayed—to obtain operational change information. Operational change information includes changes corresponding to file creation, deletion, modification, renaming, and movement. Furthermore, metadata analysis results can be updated based on this operational change information to ensure accuracy. If a metadata snapshot is not required, operational change information can be directly analyzed using edit logs.
[0085] Among them, you can set the HBase write method to submit a write for the target number (such as 2000) rowkeys. By creating a thread pool for each cluster, data from different clusters can be executed concurrently through different thread pools, further improving the real-time consumption capability of the edit log.
[0086] As an example, Figure 5 This figure shows a schematic diagram of updating source data analysis results. The JournalNode cluster is a log node cluster that replaces traditional shared storage in the HDFS architecture to manage edit logs. The Active NameNode is the master NameNode that handles all client requests and is responsible for writing edit logs to the JournalNode. HBase is an example of a target database.
[0087] Specifically, the Active NameNode submits metadata operations submitted by the client (such as creating a file) as edit logs to the log node cluster. When the NameNode starts (such as fault recovery) through the data migration device, it first loads the fsimage, then replays all unmerged edit logs (edits) to restore the latest status, and updates the metadata analysis results in HBase accordingly.
[0088] By combining metadata snapshots and edit logs to obtain operational change information and update metadata analysis results, we can avoid migration omissions or conflicts caused by source cluster metadata changes, ensuring consistency between the migration directory and the real-time state of the source cluster. Furthermore, by optimizing HBase read and write, rowkey involvement, and multi-threaded processing, we further improve the real-time consumption of edit logs.
[0089] In an exemplary embodiment, a method for obtaining a migration directory is provided, wherein the migration configuration information includes the path to be migrated and the corresponding time range information. Figure 6 As shown, obtaining the migration catalog based on the migration configuration information and predetermined metadata analysis results includes:
[0090] Step S610: According to the metadata analysis result, the sub-path information under the path to be migrated is obtained, and the modification time of the sub-path information is determined.
[0091] As mentioned above, the metadata analysis results are pre-generated structured data containing metadata for all paths in the source cluster, including the user-configured root path of the path to be migrated. This step aims to extract all sub-paths (files / directories) under the path to be migrated from the metadata analysis results and obtain their last modification time, providing the basis for subsequent filtering by time range. This process avoids the performance overhead of a full path scan and directly utilizes pre-analyzed metadata to quickly locate the subset to be migrated.
[0092] To do this, you can first recursively traverse the subpaths, extracting all nested subpaths (files and directories) under the path to be migrated from the metadata. Then, extract the modification time by reading the target field of each subpath (such as the last_modified field) and converting it to a timestamp to obtain the modification time of the subpath information. Then, you can output a list of subpaths containing the modification time.
[0093] Step S620: According to the modification time and time range information, the target sub-path information corresponding to the path to be migrated is obtained to obtain a migration directory.
[0094] This step filters eligible paths from the sub-path list based on the user-configured time range and generates the final migration directory. This process enables incremental migration or point-in-time recovery, migrating only data that changed within a specific time period. It is configurable and can migrate the required data based on user-configured migration requirements, providing greater flexibility.
[0095] The time range information input by the user can be converted into a timestamp interval, and then the subpaths are filtered according to the timestamp interval and the modification time of each subpath information to obtain the target subpath information within the timestamp interval, so as to obtain the migration directory.
[0096] After the user configures the parent path (the path to be migrated) and the migration cycle (time range information), the system automatically obtains the new path to be migrated, generates the corresponding migration task instance, and starts the client to perform the migration, thereby improving the configurability and automation of cross-cluster data migration.
[0097] In an exemplary embodiment, Figure 7 , checking the correctness and data uniqueness of the migration directory may include:
[0098] Step S710: Detect whether the directory of the source cluster includes the migration directory, and obtain a detection result.
[0099] Verify that all paths in the migration directory exist and are accessible in the source cluster to avoid migration failures due to incorrect paths or insufficient permissions. You can use the migration directory to check whether the source cluster actually has the directory. For example, you can call the source cluster's exists() method to perform batch verification. If a path does not exist, it can be marked as a verification failure and removed from the migration directory.
[0100] Step S720: Check whether there is a migration task for the migration directory in the historical migration tasks, and obtain a detection result.
[0101] Historically created tasks (if not deleted) are recorded in the database and can be used to verify migration paths. This step checks whether the paths in the migration directory have been processed by historical migration tasks to avoid duplicate data migrations. This can be understood as task uniqueness, avoiding the creation of meaningless tasks.
[0102] Among them, you can obtain the path list and historical task database in the migration directory, and query the database to see whether there are task records with the same source path + target path.
[0103] Step S730: Detect whether there are other migration tasks that have a path inclusion relationship with the migration directory, and obtain a detection result.
[0104] This step identifies whether there are nested paths in the migration directory to prevent duplicate migration or omission of some files. This can be understood as a directory mutual exclusivity check. Data migration must ensure the mutual exclusivity of paths. That is, migration task paths must not contain each other, which would lead to empty paths during the actual migration.
[0105] The migration directory may be firstly divided according to the directory hierarchy to obtain multiple target directories, and then, based on each target directory, it is respectively queried whether there are other migration tasks with path inclusion relationships to obtain detection results.
[0106] This process can identify whether there are path nesting conflicts between the current migration directory and other unfinished or historical migration tasks, preventing the same file from being processed by multiple tasks, or parent path tasks from overwriting child path tasks, or vice versa. It can also avoid data duplication, overwriting, or omissions when performing multi-task parallel migration, incremental migration, and full migration in a mixed manner.
[0107] Specifically, the splitting strategy can be to split each migration path into all its parent directory levels to detect partial overlap or complete containment relationships. Then, for each split target directory, check whether other tasks meet one of the following conditions: parent contains child (the path of the other task is the parent directory of the current path), child contains parent (the path of the other task is the subdirectory of the current path), or partial overlap (the paths have a common prefix but are not completely contained). If at least one of the above conditions exists, the test fails.
[0108] Based on this, it is possible to avoid the same file being repeatedly migrated by multiple tasks, prevent partial data omission due to nested path ranges, and reduce manual intervention costs through automated conflict detection.
[0109] Step S740: determining a first verification result according to the detection result, wherein the first verification result is used to indicate whether the correctness of the migration directory and the data uniqueness verification have passed.
[0110] The test results of steps S710-S730 are combined to determine whether the migration directory passes verification and obtain a first verification result. It should be understood that if at least one of the test results of steps S710-S730 fails, the migration directory is determined to have failed verification.
[0111] Real-time path detection ensures the accessibility of migration data and reduces the generation of invalid tasks. Task uniqueness detection avoids duplicate migrations and saves storage and computing resources. Directory mutual exclusivity detection prevents data omissions or redundancies caused by overlapping path ranges, providing guaranteed data verification for building migration tasks and improving the accuracy and effectiveness of migration task construction.
[0112] In an exemplary embodiment, if the migration catalog passes verification, creating a migration task according to the migration catalog may include:
[0113] First, the task to be migrated is determined based on the migration catalog. Then, the content of the task to be migrated is supplemented based on at least one of the task identification information, cluster identification information, and scheduling period information in the migration configuration information to obtain the supplemented task to be created. Finally, task status information is added to the supplemented task to be created to obtain the migration task. The task status information indicates the current task status of the migration task.
[0114] After ensuring data accuracy and interface security (permission verification) through directory screening and data verification, basic data is supplemented and tasks are created. Specifically, the server can automatically supplement and generate task-related information, including but not limited to: Task Identify, Source Cluster URI, Target Cluster URI, Schedule Cycle, Task Status, Target Filter Condition, and other supplementary information. The target filter condition can be understood as follows: migration tasks are overview information, recording parent directory information, while subdirectory tasks are based on the "target filter condition," filtering out subdirectories that meet the conditions. Migration does not migrate the parent directory all at once, but rather migrates subdirectories incrementally, determining the migration directory based on the subdirectories being migrated. For example, when obtaining the migration directory, filtering can be performed based on a user-specified template or time range. Supplementation can be based on user input.
[0115] Task status information describes the operational status of a task. It can include: created, indicating that the task has been created; submitted, indicating that the task has been submitted and can be scheduled; running, indicating that the corresponding instance task is running; to_cancel, indicating that the task is pending cancellation; canceled, indicating that the task has been canceled; deleted, indicating that the task has been deleted; and completed, indicating that the task has been completed. This status can be flexibly adjusted based on the current task status of the migration task.
[0116] Specifically, tasks to be migrated are determined based on the migration directory. Paths in the migration directory are converted into schedulable minimum task units, supporting parallelization and fault tolerance. For example, tasks can be split by file shards, partitions / tables, or data blocks, with no specific restrictions. Based on this, a basic task structure, such as a JSON structure, is output.
[0117] Furthermore, at least one of the task identification information, cluster identification information, and scheduling period information can be identified from the migration configuration information to supplement the content of the task to be migrated, mark the task status, and persist the task to obtain the created migration task.
[0118] Diverse migration requirements are supported by configuration drivers (such as task identification and scheduling cycles). By adding states, tasks can be tracked and recovered based on state machine management.
[0119] In an exemplary embodiment, the migration task includes multiple sub-directories and also provides a server-side flow limiting method. Figure 8 As shown, the current limiting method includes:
[0120] Step S810: Determine the current available bandwidth based on the total bandwidth and the occupied bandwidth.
[0121] Total bandwidth is the total available bandwidth, and occupied bandwidth is the bandwidth already occupied by other tasks. Subtracting the two will give the current available bandwidth.
[0122] Step S820: Compare the total bandwidth required for a single migration of multiple sub-directories with the current available bandwidth to obtain a first comparison result. The first comparison result is used to indicate whether the migration task is currently allowed to be executed.
[0123] The server limits bandwidth at the task level, so the total bandwidth required for a single migration of multiple subdirectories is calculated and compared with the currently available bandwidth. The total bandwidth required for a single migration of multiple subdirectories can be calculated based on the number of subdirectories and the expected completion time.
[0124] Step S830: Scheduling the client to execute the migration task according to the first comparison result.
[0125] A migration task is allowed only when the currently available bandwidth is greater than or equal to the total bandwidth for a single migration. If the bandwidth is insufficient, the currently scheduled migration task is put into a dormant state and waits for the next scheduled schedule.
[0126] Through real-time bandwidth calculation and task scheduling, network congestion during multi-subdirectory migration is avoided, migration efficiency is balanced with business system stability, and cluster jitter or network unavailability caused by migration tasks preempting bandwidth is prevented.
[0127] In an exemplary embodiment, data in the source cluster that is migrated to the target cluster may also be stored in a logical deletion buffer, so that the logical deletion buffer is cleaned up regularly based on a preset period.
[0128] The logical delete buffer is used to store data that needs to be cleaned during the migration process but must remain recoverable. This can be understood as a grayscale space for storing source cluster data (such as hot data) after logical cleanup. This logical delete buffer is invisible to users. Data in the logical delete buffer can be recorded at the directory level, with each piece of data tagged and its age monitored. This allows for configuration-controlled physical cleanup times, achieving both automated processing and data recoverability.
[0129] Specifically, when migrating data from the source cluster to the target cluster, the source data is not physically deleted immediately. Instead, the source data is marked as logically deleted, and the data is moved to a logical delete buffer (such as an independent directory or database). The logical delete buffer can retain data such as the original path, permissions, and timestamps, and the data content can be stored in a compressed and encrypted form to save space. The preset period can be user-configured (dynamically adjustable) or configured by default (fixed period), so that overdue data in the logical delete buffer can be physically deleted according to the preset period, and accidentally deleted data can also be restored from this area to the source cluster or target cluster.
[0130] Based on this, a directory to be restored can also be determined in response to a data restoration request. When the directory to be restored does not have data under the directory of the source cluster, the path of the directory to be restored in the buffer is logically deleted and moved to the directory of the source cluster.
[0131] In some cases, such as when the source cluster data has been cleaned but the target cluster data is incomplete or corrupted, it is necessary to quickly restore the data from the logical delete buffer to the source cluster. This data restoration request can be user-initiated or automatically triggered when a target cluster validation failure is detected.
[0132] Specifically, the source cluster is checked for the existence of the path to be restored. If so, the restore is terminated (to avoid overwriting). If not, the restore process can continue. The path data is searched in the logical delete buffer. If found, the restore is executed; otherwise, an error is reported. During the restore, the path of the directory to be restored in the logical delete buffer is moved to the directory in the source cluster. Path-level restore avoids the redundant cost of a full rollback, embodying the atomicity of data migration. The buffer is moved rather than copied, avoiding double storage overhead.
[0133] Furthermore, the exemplary embodiments of the present disclosure also provide a data migration method, which is applied to a client, such as Figure 9 , the data migration method includes:
[0134] Step S910: Acquire a migration task, where the migration task is constructed according to any one of the data migration methods in the above exemplary embodiments, and includes multiple sub-directories.
[0135] Step S920: construct migration plans for each of the multiple sub-directories to obtain multiple migration plans.
[0136] Step S930: Execute multiple migration plans to migrate data in the source cluster to the target cluster.
[0137] It should be noted that the creation of the migration task is executed by the server. This process has been described in the above exemplary embodiment and will not be described in detail here.
[0138] Migration plans are constructed for multiple subdirectories, resulting in multiple migration plans. This can be understood as generating a separate migration plan for each subdirectory. A migration task is a high-level task message (based on the migration directory) created based on user-specified information, but the actual migration involves migrating subdirectories within the parent directory. Based on the high-level task information, a migration plan is generated for each subdirectory and the migration is executed.
[0139] When building a migration plan based on the subdirectory of the migration task, the contents of the migration plan may include but are not limited to: SourceClusterUri, indicating the source cluster identifier; TargetClusterUri, indicating the target cluster identifier; SourceClusterPath, indicating the source cluster path; TargetClusterTempPath, indicating the temporary path of the target cluster; TargetClusterPath, indicating the target path of the target cluster;
[0140] SourceClusterTrashPath, which indicates the grayscale space path (logical deletion buffer) of the source cluster;
[0141] path: indicates the user-entered path (migration configuration information) and does not need to include the cluster URI;
[0142] TargetClusterTrashPath indicates the cleanup path of the target cluster.
[0143] Specifically, TargetClusterTempPath is used to migrate source cluster data to a temporary directory in the target cluster. Data can only be moved to TargetClusterPath after passing data verification. The generation rule is TargetClusterUri / .migration / timestamp-target directory-UUID. TargetClusterPath is the target path. Its structure should be consistent with SourceClusterPath except for the cluster URI. SourceClusterTrashPath is used for delayed deletion of source cluster data. Source cluster data needs to be temporarily moved to this directory. The generation rule is SourceClusterUri / warmstandy / timestamp-UUID / path. TargetClusterTrashPath is the cleanup path of the target cluster. Its purpose is to move existing empty directories to avoid affecting data writing. The generation rule is: TargetClusterPath / warmstandy / yyyy-MM-dd / timestamp-UUID.
[0144] Optionally, before generating a migration plan and actually migrating data, perform parameter validation on the data in the plan. Verification items should include at least the following: TargetClusterTempPath does not exist; TargetClusterPath does not exist; SourceClusterTrashPath does not exist; SourceClusterPath exists; TargetClusterTrashPath does not exist. This further ensures the accuracy of the migration plan. Of course, the type of parameter validation can be adjusted flexibly during implementation.
[0145] In actual implementation, the migration plan can also be implemented based on DistCp. Specifically, Figure 10 As shown, executing multiple migration plans to migrate data from a source cluster to a target cluster may include:
[0146] Step S1010: For each migration plan, create a temporary storage path of the target cluster according to the migration plan.
[0147] Create a temporary storage path according to the migration plan, such as TargetClusterTempPath. Data to be migrated to the target cluster will be temporarily migrated to this path and finally migrated to the target path of the target cluster after subsequent verification succeeds.
[0148] Step S1020: obtaining migration constraint information according to the migration configuration information, and migrating the data corresponding to the migration plan to a temporary storage path according to the migration constraint information, wherein the migration constraint information is used to limit the transmission bandwidth and / or transmission volume when executing the migration plan.
[0149] After creating the temporary path, you can also pass in relevant migration configurations, such as the specified queue, refresh policy, bandwidth for the DistCp task, and migration strategy. This configuration is read as a configuration file and stored in a key:value format. In practice, this configuration (migration constraints) is read when migrating each subdirectory. If the relevant configuration changes or new Hadoop parameters are added, a migration plan for the DistCp subdirectory is recreated using the new content.
[0150] Migration constraints can also include information such as the total migration bandwidth and the number of migration tasks, to constrain the transmission bandwidth and / or transmission volume when executing the migration plan. By adjusting the migration constraints, tasks can be dynamically adjusted without hard-coding them into the code.
[0151] As an example, you can first create a temporary storage path for the target cluster according to the migration plan, then pass in the Hadoop configuration, and build a dynamic adjustment configuration for sub-directory migration based on the transmission restrictions (migration constraint information) across cluster keys. Then, build the attribute retention items during the actual transmission process to preserve the consistency of the actual file metadata attributes. Finally, realize the atomicity of the DistCp process migration. In the event of partial success, recovery and cleanup are performed at the business level to ensure the complete and successful migration of the data.
[0152] It should be noted that the migration plan process is based on DistCp and there are no excessive restrictions on this.
[0153] Step S1030: performing a data consistency check on the migration plan. If the migration plan passes the check, the data in the temporary storage path is moved to the target path of the target cluster.
[0154] Among them, data consistency check is used to characterize the consistency of data before and after the execution of the migration plan. Considering that network jitter or node failure during the migration process may cause data corruption or loss, data consistency check is performed to ensure that the data in the temporary storage path is completely consistent with the source cluster data before switching to the target path to avoid business unavailability due to partial data errors.
[0155] Optionally, for each migration plan, the consistency of the logical storage space occupied by the data before and after the migration plan is executed can be obtained to obtain a verification result. Alternatively, the consistency of the total hash value of the data before and after the migration plan is executed can be obtained to obtain a verification result. Furthermore, a second verification result is determined based on the verification result. The second verification result is used to indicate whether the migration plan's data consistency check has passed.
[0156] Specifically, based on du calculation, the logical storage space occupied by files or directories in the source and target paths can be calculated to determine whether they are consistent. Alternatively, based on count and hash calculation, the number of files in the source and target paths and the total hash value of the files can be calculated to determine whether they are consistent. Optionally, if both verification results indicate data consistency before and after executing the migration plan, the second verification result is determined to indicate that the migration plan's data consistency verification has passed.
[0157] Furthermore, if any target migration plan fails the data consistency check, the source cluster and the target cluster are rolled back to restore to the state before the migration task was executed; and the temporary storage path is cleared.
[0158] Since data consistency verification is required after each migration plan is executed, if a target migration plan fails the verification, an atomic rollback is performed. That is, through logical business guarantees, if a migration plan encounters an exception, it will be restored to the state before the initial operation and the temporary storage path will be cleared.
[0159] For example, you can delete newly added files during the migration and restore the original versions of modified files. For the target cluster, you can delete all data that fails validation. If the target cluster has updated metadata, you need to roll back to the previous version.
[0160] This means that after a directory or file is migrated from the source cluster to the target cluster, its data content, size, and related data attributes must remain unchanged and must remain consistent. With atomic rollback, the directory or file migration process can be completed or incomplete, preventing data loss or corruption caused by a missed step.
[0161] In an exemplary embodiment, a client flow limiting method is also provided. Executing multiple migration plans to migrate data from a source cluster to a target cluster may include:
[0162] First, for the migration plan currently being migrated, the required single migration bandwidth is compared with the currently available bandwidth to obtain a second comparison result. The second comparison result indicates whether the migration plan is currently allowed to be executed. Furthermore, the migration plan is controlled to wait or stop based on the second comparison result.
[0163] As can be seen above, the server limits traffic at the task level, while the client implements finer-grained traffic control for each subdirectory of the migration plan that actually migrates. The bandwidth required for a single migration plan is defined as the bandwidth required for that migration plan. If the currently available bandwidth meets (is greater than or equal to) the bandwidth required for the migration plan to be migrated, the migration plan is executed. Otherwise, a loop waits. If sufficient bandwidth is still unavailable after the preset duration, the task is terminated and the next server-side scheduling is requested.
[0164] As an example, when migrating to the migration plan currently to be migrated, if the currently available bandwidth does not meet the bandwidth required by the migration plan currently to be migrated, you can use a while loop to wait and make a judgment every 5 minutes. If it is still not met after more than 30 minutes, exit the waiting and stop process.
[0165] In one exemplary embodiment, if the migration plan passes the consistency check, it can be understood that the data in the temporary storage path is final data that has successfully undergone data verification. The data in the temporary storage path can then be moved to the target path in the target cluster. For example, the temporary storage path can be mounted, thereby completing the entire data migration process.
[0166] In an exemplary embodiment, to ensure the atomicity, consistency, and durability of cross-cluster data migration, exceptions during the migration process can also be classified and captured. Specifically, the exception types of the exemplary embodiment of the present disclosure may include, but are not limited to: source cluster configuration exceptions; target cluster configuration exceptions; task operation exceptions; task submission and launch exceptions; source cluster source path does not exist; target cluster source path already exists; terminal device task DistCp execution exceptions; data verification exceptions after data migration DistCp; data movement exceptions; and data restoration exceptions.
[0167] If an exception occurs in any of the above steps, an alarm message can be generated based on the exception type and the relevant person in charge can be notified for investigation or remediation. For example, if a DistCP task fails, the administrator can check the task log on YARN and resolve the issue. Alternatively, if a client or server exception occurs, it may not be an issue with the actual task but rather a cluster or service problem, and the developer should check the application.
[0168] The data migration method in the exemplary embodiment of the present disclosure obtains migration configuration information, obtains a migration directory based on the migration configuration information and predetermined metadata analysis results, and the metadata analysis results are obtained by analyzing and processing the metadata of the source cluster. The migration directory is then verified for correctness and data uniqueness. If the migration directory passes the verification, a migration task is created based on the migration directory; finally, the migration task is sent to the client so that the client executes the migration task and migrates the data in the source cluster to the target cluster. On the one hand, the user only needs to provide the migration configuration information, and the migration directory can be generated by automatically matching the migration configuration information with the metadata analysis results, avoiding the tedious manual configuration. The standardization of the migration process is achieved through the automated creation and distribution of migration tasks, reducing the risk of errors caused by human intervention, thereby ensuring the efficiency and accuracy of cross-cluster data migration. On the other hand, a migration catalog is generated based on the metadata analysis results of the source cluster to ensure that the migration scope accurately corresponds to the source data structure, avoiding omissions or redundancies, and through correctness verification and data uniqueness verification, duplicate data or data corruption can be prevented in the target cluster, ensuring that the data is complete and available after migration. At the same time, since the verification link is pre-placed (completed before migration), potential problems can be identified in advance, abnormal operations during the migration process can be reduced, and the security and effectiveness of data migration can be further ensured. In addition, the server is responsible for task scheduling and verification, and the client performs the actual migration to achieve load separation, avoid server performance bottlenecks, and adapt to migration tasks in cluster environments of different sizes. Therefore, the exemplary embodiment of the present disclosure realizes configurable, adaptive, automated, and asynchronous data migration across clusters, thereby ensuring the atomicity, consistency, isolation (task uniqueness) and persistence of data migration from the business logic level, improving the accuracy and efficiency of high-speed data migration, reducing migration costs, and improving service reliability.
[0169] In an exemplary embodiment of the present disclosure, a data migration device is also provided, which is applied to a server. Figure 11 As shown, the apparatus 1100 may include a path configuration module 1110, a data processing module 1120, a verification module 1130, and a task generation module 1140. Specifically:
[0170] The path configuration module 1110 is used to obtain migration configuration information; the data processing module 1120 is used to obtain the migration directory based on the migration configuration information and the predetermined metadata analysis results, where the metadata analysis results are obtained by analyzing and processing the metadata of the source cluster; the verification module 1130 is used to verify the correctness and data uniqueness of the migration directory. If the migration directory passes the verification, a migration task is created based on the migration directory; the task generation module 1140 is used to send the migration task to the client so that the client executes the migration task and migrates the data in the source cluster to the target cluster.
[0171] In an exemplary embodiment of the present disclosure, the path configuration module 1110 is also configured to execute: obtaining the metadata of the source cluster stored in the name node; analyzing and processing the directory file information in the metadata of the source cluster, obtaining the metadata analysis results and storing them in the target database, and the target database provides an interface for data query.
[0172] In an exemplary embodiment of the present disclosure, the path configuration module 1110 is further configured to perform: metadata reconstruction based on the metadata snapshot to obtain reconstructed metadata; obtaining the editing log corresponding to the client operation; comparing and analyzing the metadata of the source cluster based on the reconstructed metadata and / or the editing log to obtain operation change information, and updating the metadata analysis results based on the operation change information.
[0173] In an exemplary embodiment of the present disclosure, the migration configuration information includes the path to be migrated and the corresponding time range information; the data processing module 1120 is configured to execute: according to the metadata analysis results, obtain the sub-path information under the path to be migrated, and determine the modification time of the sub-path information; according to the modification time and time range information, obtain the target sub-path information corresponding to the path to be migrated to obtain the migration directory.
[0174] In an exemplary embodiment of the present disclosure, the verification module 1130 is configured to perform: detecting whether the directory of the source cluster includes a migration directory to obtain a detection result; detecting whether there is a migration task for the migration directory in the historical migration tasks to obtain a detection result; detecting whether there are other migration tasks that have a path inclusion relationship with the migration directory to obtain a detection result; determining a first verification result based on the detection result, wherein the first verification result is used to indicate whether the correctness of the migration directory and the data uniqueness verification have passed.
[0175] In an exemplary embodiment of the present disclosure, it is detected whether there are other migration tasks that have a path inclusion relationship with the migration directory to obtain a detection result, including: dividing the migration directory according to the directory hierarchy to obtain multiple target directories; and querying whether there are other migration tasks that have a path inclusion relationship according to each target directory to obtain a detection result.
[0176] In an exemplary embodiment of the present disclosure, if the migration directory passes the verification, a migration task is created according to the migration directory, including: determining the task to be migrated according to the migration directory; supplementing the content of the task to be migrated according to at least one of the task identification information, cluster identification information, and scheduling cycle information in the migration configuration information to obtain the supplemented task to be created; adding task status information to the supplemented task to be created to obtain the migration task, wherein the task status information is used to characterize the current task status of the migration task.
[0177] In an exemplary embodiment of the present disclosure, the migration task includes multiple sub-directories, and the verification module 1130 is further configured to perform: determining the current available bandwidth based on the total bandwidth and the occupied bandwidth; comparing the total bandwidth required for a single migration of multiple sub-directories with the current available bandwidth to obtain a first comparison result, and the first comparison result is used to indicate whether the migration task is currently allowed to be executed; and scheduling the client to execute the migration task based on the first comparison result.
[0178] In an exemplary embodiment of the present disclosure, the device also includes a data cleaning module, which is configured to store the data migrated from the source cluster to the target cluster in a logical deletion buffer, and the logical deletion buffer is used to store data to be cleaned during the migration process but the possibility of recovery must be retained; based on a preset period, the logical deletion buffer is cleaned regularly.
[0179] In an exemplary embodiment of the present disclosure, the data cleaning module is further configured to execute: in response to a data recovery request, determine the directory to be restored; when there is no data in the directory to be restored under the directory of the source cluster, logically delete the path of the directory to be restored in the buffer and move it to the directory of the source cluster.
[0180] In an exemplary embodiment of the present disclosure, a data migration device is also provided, which is applied to a client. Figure 12 As shown, the apparatus 1200 may include a task acquisition module 1210, a plan construction module 1220, and a data migration module 1230. Specifically:
[0181] The task acquisition module 1210 is used to acquire a migration task, where the migration task is constructed according to any one of the data migration methods described above, and the migration task includes multiple sub-directories; the plan construction module 1220 is used to construct migration plans for the multiple sub-directories respectively to obtain multiple migration plans; the data migration module 1230 is used to execute the multiple migration plans to migrate the data in the source cluster to the target cluster.
[0182] In an exemplary embodiment of the present disclosure, the data migration module 1230 is configured to perform: for each migration plan, create a temporary storage path for the target cluster according to the migration plan; obtain migration constraint information according to the migration configuration information, and migrate the data corresponding to the migration plan to the temporary storage path according to the migration constraint information, wherein the migration constraint information is used to limit the transmission bandwidth and / or transmission volume when executing the migration plan; perform data consistency verification on the migration plan, and if the migration plan passes the verification, move the data in the temporary storage path to the target path of the target cluster, wherein the data consistency verification is used to characterize the consistency of the data before and after the execution of the migration plan.
[0183] In an exemplary embodiment of the present disclosure, a data consistency check is performed on the migration plan, including: for each migration plan, obtaining whether the logical storage space occupied by the data before and after the execution of the migration plan is consistent, and obtaining a verification result; and / or obtaining whether the total hash value of the data before and after the execution of the migration plan is consistent, and obtaining a verification result; determining a second check result based on the verification result, and the second check result is used to indicate whether the data consistency check of the migration plan has passed.
[0184] In an exemplary embodiment of the present disclosure, if a target migration plan fails the data consistency check, the source cluster and the target cluster are rolled back to restore to the state before the migration task is executed; and the temporary storage path is cleared.
[0185] In an exemplary embodiment of the present disclosure, the data migration module 1230 is configured to perform: for the migration plan to be migrated currently, compare the required single migration bandwidth with the currently available bandwidth to obtain a second comparison result, and the second comparison result is used to characterize whether the migration plan to be migrated currently is allowed to be executed; and control the migration plan to be migrated currently to wait or stop according to the second comparison result.
[0186] Since the details of the functional modules of the data migration apparatus of the exemplary embodiment of the present disclosure have been described in the exemplary embodiment of the data migration method described above, they will not be repeated here.
[0187] It should be noted that although the above detailed description mentions several modules or units of the data migration device, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in a single module or unit. Conversely, the features and functions of a single module or unit described above can be further divided and embodied by multiple modules or units.
[0188] The exemplary embodiments of the present disclosure further provide a computer program product, which includes a computer program, and when the computer program is executed by a processor, implements the above-mentioned data migration method.
[0189] In one embodiment, a computer program product may be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The computer-readable storage medium may be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, or other signals, including but not limited to random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory (Flash), mechanical hard disk drive (HDD), solid-state drive (SSD), and the like. Exemplarily, the computer program product may be implemented as a non-volatile storage medium storing the computer program, such as a read-only memory, NAND flash memory, and the like.
[0190] In one embodiment, the computer program product may be an intangible product containing a computer program. For example, the computer program product may be implemented as a virtual digital product, such as a digital file such as an executable file or installation package storing the computer program.
[0191] The code of the computer program can be written in one or more programming languages. Programming languages include C, Java, C++, etc. The program code can be executed entirely on the user computing device, partially on the user computing device, or as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device via any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or can be connected to an external computing device (e.g., via an Internet connection provided by a carrier).
[0192] Computer programs can be carried or transmitted via electrical, magnetic, optical, electromagnetic, infrared, or other signals. Electronic devices can convert signals carrying computer programs into digital signals to run the computer programs. When the computer program runs on an electronic device, its code is used to cause the electronic device to execute (more specifically, to cause the processor of the electronic device to execute) the method steps of various exemplary embodiments of the present disclosure, such as the data migration method described above.
[0193] In addition, in an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided. Those skilled in the art will appreciate that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to herein as a "circuit," "module," or "system."
[0194] Refer to the following Figure 13 13 to describe the electronic device 1300 according to this embodiment of the present disclosure. Figure 13 The electronic device 1300 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0195] like Figure 13As shown, electronic device 1300 is implemented as a general-purpose computing device. Components of electronic device 1300 may include, but are not limited to, the aforementioned at least one processing unit 1310, the aforementioned at least one storage unit 1320, a bus 1330 connecting various system components (including storage unit 1320 and processing unit 1310), and a display unit 1340.
[0196] The storage unit stores program codes, which can be executed by the processing unit 1310, so that the processing unit 1310 performs the steps described in the above “Exemplary Method” section of this specification according to various exemplary embodiments of the present disclosure.
[0197] The storage unit 1320 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 1321 and / or a cache memory unit 1322 , and may further include a read-only memory unit (ROM) 1323 .
[0198] The storage unit 1320 may also include a program / utility 1324 having a set (at least one) of program modules 1325, such program modules 1325 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0199] Bus 1330 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0200] The electronic device 1300 can also communicate with one or more external devices 1400 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 1300, and / or any device that enables the electronic device 1300 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 1350. Furthermore, the electronic device 1300 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 1360. As shown, the network adapter 1360 communicates with other modules of the electronic device 1300 via a bus 1330. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 1300, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0201] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0202] Furthermore, the figures above are merely illustrative of the processes included in the methods according to exemplary embodiments of the present disclosure and are not intended to be limiting. It is readily understood that the processes illustrated in the figures above do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0203] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow from the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.
Claims
1. A data migration method, characterized in that: Applied to the server, the method includes: Get migration configuration information; Obtaining a migration directory based on the migration configuration information and a predetermined metadata analysis result, wherein the metadata analysis result is obtained by analyzing and processing metadata of the source cluster; Verifying the correctness and data uniqueness of the migration directory, and if the migration directory passes the verification, creating a migration task based on the migration directory; The migration task is sent to the client, so that the client executes the migration task and migrates the data in the source cluster to the target cluster.
2. The method according to claim 1, characterized in that Obtaining the predetermined metadata analysis result includes: Get the metadata of the source cluster stored on the name node; The directory file information in the metadata of the source cluster is analyzed and processed to obtain the metadata analysis result and store it in the target database. The target database provides an interface for data query.
3. The method according to claim 1, characterized in that The method further comprises: Reconstructing metadata according to the metadata snapshot to obtain reconstructed metadata; Get the edit log corresponding to the client operation; The metadata of the source cluster is compared and analyzed based on the reconstructed metadata and / or the edit log to obtain operation change information, and the metadata analysis result is updated based on the operation change information.
4. The method according to claim 1, wherein The migration configuration information includes the path to be migrated and the corresponding time range information; The step of obtaining a migration directory according to the migration configuration information and a predetermined metadata analysis result includes: According to the metadata analysis result, obtaining sub-path information under the path to be migrated, and determining the modification time of the sub-path information; According to the modification time and the time range information, target sub-path information corresponding to the path to be migrated is acquired to obtain the migration directory.
5. The method according to claim 1, wherein The correctness and data uniqueness verification of the migration directory includes: Detecting whether the directory of the source cluster includes the migration directory, and obtaining a detection result; Detect whether there is a migration task for the migration directory in the historical migration tasks, and obtain a detection result; Detecting whether there are other migration tasks that have a path inclusion relationship with the migration directory, and obtaining a detection result; A first verification result is determined according to the detection result, wherein the first verification result is used to indicate whether the correctness and data uniqueness verification of the migration directory have passed.
6. The method according to claim 5, characterized in that The detecting whether there are other migration tasks that have a path inclusion relationship with the migration directory and obtaining a detection result includes: Splitting the migration directory according to the directory hierarchy to obtain multiple target directories; According to each target directory, check whether there are other migration tasks with the same path inclusion relationship to obtain the detection result.
7. The method according to claim 1, characterized in that If the migration directory passes the verification, creating a migration task according to the migration directory includes: Determine tasks to be migrated according to the migration catalog; Supplementing the content of the task to be migrated according to at least one of the task identification information, cluster identification information, and scheduling period information in the migration configuration information to obtain a supplemented task to be created; Task status information is added to the supplemented task to be created to obtain the migration task, wherein the task status information is used to represent the current task status of the migration task.
8. The method according to claim 1, characterized in that The migration task includes multiple sub-directories, and the method further includes: Determine the current available bandwidth based on the total bandwidth and occupied bandwidth; Comparing the total bandwidth required for a single migration of the multiple subdirectories with the currently available bandwidth to obtain a first comparison result, where the first comparison result is used to indicate whether the migration task is currently allowed to be executed; The client is scheduled to execute the migration task according to the first comparison result.
9. The method according to claim 1, characterized in that The method further comprises: Storing the data in the source cluster that is migrated to the target cluster in a logical deletion buffer, where the logical deletion buffer is used to store data that needs to be cleaned up during the migration process but needs to be able to be restored; Based on a preset period, the logical deletion buffer is cleaned up regularly.
10. A data migration method, characterized in that: Applied to a client, the method includes: Obtaining a migration task, where the migration task is constructed according to the data migration method according to any one of claims 1 to 9, and the migration task includes multiple subdirectories; Constructing migration plans for each of the multiple subdirectories to obtain multiple migration plans; The multiple migration plans are executed to migrate data in the source cluster to the target cluster.