Data migration method and related equipment
By determining the task type according to the file size and calling the MapReduce task for data migration, the problem of inefficient data migration in the existing technology is solved, and more efficient data migration and the effect of reducing storage costs is achieved.
Patent Information
- Application Number
- CN202510420923.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-27
AI Technical Summary
Existing data migration methods are inefficient, resulting in longer time to migrate from cluster to cloud servers and increasing storage costs.
Based on the file size information recorded by the data platform, determine the task type of the file, and call the MapReduce task to migrate the file to the cloud server according to the task type. Task types are divided into large tasks and small tasks, and they are migrated using different bandwidths and concurrent numbers respectively.
By dividing file migration tasks into large and small tasks, and selecting the appropriate bandwidth and concurrency quantity according to the task type, the efficiency of data migration is improved and storage costs are reduced.
Smart Images

Figure CN120215839A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technology, and particularly to a data migration method and related devices. Background Art
[0002] Existing data is often stored in a cluster composed of multiple computers, resulting in high hardware resource overhead and high storage costs. Migrating data from the cluster to a cloud server for storage can reduce storage costs. However, due to the large amount of data stored in the cluster, data migration takes a lot of time, resulting in low data migration efficiency. Summary of the Invention
[0003] In view of the above, it is necessary to provide a data migration method, an electronic device, and a storage medium to solve the technical problem of low data migration efficiency.
[0004] In a first aspect, this application provides a data migration method, which includes: determining the task type of a file according to the file size information of the file recorded by a data platform; invoking the mapping and reduction MapReduce task of a data migration tool, and migrating the file to a cloud server according to the task type of the file.
[0005] In some embodiments of this application, the task type includes a first task type and a second task type, and determining the task type of the file includes: if the file size of the file is greater than a preset byte threshold, determining that the task type of the file is the first task type; if the file size of the file is less than or equal to the preset byte threshold, determining that the task type of the file is the second task type.
[0006] In some embodiments of this application, migrating the file to the cloud server according to the task type of the file includes: if the task type of the file is the first task type, migrating the file to the cloud server based on a first bandwidth and a first concurrency number; if the task type of the file is the second task type, migrating the file to the cloud server based on a second bandwidth and a second concurrency number, where the first bandwidth is greater than the second bandwidth, and the first concurrency number is less than the second concurrency number.
[0007] In some embodiments of the present application, the method further includes: determining a file size information table according to the first audit log recorded by the data platform; determining a data partition table according to the first audit log and the second audit log recorded by the cloud server, where the data partition table includes a differential data partition in which the data platform and the cloud server are inconsistent; determining the task type of the differential data partition according to the data partition table and the file size information table; invoking the MapReduce task, and migrating the data corresponding to the differential data partition to the cloud server according to the task type of the data partition.
[0008] In some embodiments of the present application, determining the data partition table according to the first audit log and the second audit log recorded by the cloud server includes: determining a first data partition set of the data platform according to the first audit log; determining a second data partition set of the cloud server according to the second audit log; determining the differential data partition based on the partition difference between the first data partition set and the first data partition set; and recording the differential data partition in the data partition table.
[0009] In some embodiments of the present application, migrating the data corresponding to the differential data partition to the cloud server according to the task type of the data partition includes: if the task type of the differential data partition is the first task type, migrating the data corresponding to the differential data partition to the cloud server based on the first bandwidth and the first number of concurrencies; if the task type of the differential data partition is the second task type, migrating the data corresponding to the differential data partition to the cloud server based on the second bandwidth and the second number of concurrencies, where the first bandwidth is greater than the second bandwidth and the first number of concurrencies is less than the second number of concurrencies.
[0010] In some embodiments of the present application, the method further includes: determining whether the differential data partition has been successfully migrated to the cloud server; if the differential data partition has not been successfully migrated to the cloud server, migrating the data corresponding to the differential data partition to the cloud server again according to the task type of the data partition.
[0011] In some embodiments of the present application, migrating the file to the cloud server according to the task type of the file includes: storing the file in the object storage service of the cloud server according to the task type of the file.
[0012] In a second aspect, the present application provides an electronic device, where the electronic device includes a memory and a processor: the memory is used for storing program instructions; the processor is used for reading and executing the program instructions stored in the memory, and when the program instructions are executed by the processor, the electronic device executes the above data migration method.
[0013] In a third aspect, the present application provides a computer storage medium storing program instructions that, when run on an electronic device, cause the electronic device to execute the above data migration method.
[0014] In an embodiment of the present application, based on the file size information of a file recorded by a data platform, the task type of the file is determined, and a MapReduce task is invoked to migrate the file to a cloud server according to the task type of the file. In this way, the file migration task can be divided into large tasks and small tasks according to the file size, and the file is migrated to the cloud server according to the large tasks and small tasks, thus avoiding low data migration efficiency caused by an overly large file and improving the data migration efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a diagram of an application scenario of the data migration method provided by an embodiment of the present application.
[0016] Figure 2 It is a flowchart of the data migration method provided by some embodiments of the present application.
[0017] Figure 3 It is a flowchart of the data migration method provided by some other embodiments of the present application.
[0018] Figure 4 It is a schematic structural diagram of an electronic device provided by some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] In order to more clearly understand the above objects, features, and advantages of the present application, the present application will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used in the specification of the present application herein are only for the purpose of describing an embodiment and are not intended to limit the present application.
[0021] It should be noted that the terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the accompanying drawings are used to distinguish similar objects and are not used to describe a specific order or sequence. In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0022] In addition, it should be noted that the methods disclosed in the embodiments of the present application or the methods shown in the flowcharts include one or more steps for implementing the methods. Without departing from the scope of the claims, the execution order of multiple steps can be interchanged with each other, and some steps can also be deleted. Some embodiments will be described below with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0023] Existing data is often stored in a cluster composed of multiple computers, resulting in high hardware resource overhead and high storage costs. Migrating the data from the cluster to a cloud server for storage can reduce the storage cost. However, due to the large amount of data stored in the cluster, data migration takes a lot of time, resulting in low efficiency of data migration.
[0024] To solve the above technical problems, the embodiments of the present application provide a data migration method. Refer to Figure 1 As shown, it is an application scenario diagram of the data migration method provided by the embodiments of the present application. The data migration method is applied in the Figure 1 network architecture shown. In the network architecture as shown in Figure 1 the data platform 10 and the cloud server 20 are deployed in the Internet, and the data platform 10 is communicatively connected to the cloud server 20. In some embodiments of the present application, the data platform 10 can store multiple files, and the data of the files can include data of goods such as physical products, digital products, tickets, service subscriptions, etc. In some embodiments of the present application, the data platform 10 can be a machine cluster composed of multiple computers or servers. For example, the data platform 10 can be a Hadoop platform composed of a Hadoop offline machine cluster. In some embodiments of the application, the cloud server 20 includes an Object Storage Service (OSS). In some embodiments of the application, the cloud server 20 can be cloud servers built by different manufacturers, such as Alibaba Cloud and Huawei Cloud.
[0025] The data migration method provided by the present application can be programmed as a computer program product and deployed to run in a computer or a server. For example, in the exemplary application scenario of the present application, it can be deployed and implemented in a computer or a server of the Hadoop platform 10. Thus, through the interface opened after the computer program product runs, human-computer interaction can be performed with the process of the computer program product through a graphical user interface to execute the method.
[0026] The method can determine the task type of the files stored in the Hadoop platform according to the file size information of the files recorded by the Hadoop platform, call the MapReduce task of the data migration tool, and migrate the files to the cloud server according to the task type of the files (such as large tasks and small tasks). In this way, the file migration tasks can be divided into large tasks and small tasks according to the file size, and the files can be migrated to the cloud server according to the large tasks and small tasks, thus avoiding the problem of low data migration efficiency caused by overly large files.
[0027] Reference Figure 2 As shown, it is a flowchart of the data migration method provided by some embodiments of the present application. Figure 2 The exemplary method includes one or more steps, but does not constitute a limitation to the present application. In addition, the order of the steps of the method is only for example, and the order of the steps can be changed. Without departing from the content disclosed in the present application, additional steps can be added or steps can be reduced. The method includes the following steps.
[0028] Step S201: Determine the task type of the file according to the file size information of the file recorded by the Hadoop platform, where the task type includes a first task type and a second task type.
[0029] In some embodiments of the present application, the Hadoop platform generates an audit log (Audit Log) according to the stored file data every preset period (such as every hour or every day), and the audit log records the file information of all the files stored in the Hadoop platform. In some embodiments of the present application, the audit log is a log file used by the Hadoop platform to record system operations and file access information. The purpose of the audit log is to track and monitor file operations to ensure data security and traceability. For example, whenever there is a change in the stored file data (such as adding, modifying, or deleting a file), the Hadoop platform will record the relevant file operations.
[0030] In some embodiments of the present application, the file information includes file location information, file size information, and metadata information of the file. The file location information represents the storage path of the file in the Hadoop Distributed File System (HDFS) of the Hadoop platform. The file size information represents the size of the file. For example, the file size information is 1024 kb. The metadata information of the file represents the attribute information of the file. In some embodiments of the present application, the metadata information of the file includes the file creation time, file modification time, and file permissions, where the file permissions include read and write permissions.
[0031] In some embodiments of the present application, the audit logs are parsed to extract the key fields of the audit logs, and the data of the key fields is converted into a structured data format (such as a table form), and the structured data is stored in a database table. In some embodiments of the present application, the key fields include, but are not limited to, timestamp, file path, file size, file owner, permissions, creation time, and modification time.
[0032] In some embodiments of the present application, the determination of the task type of the file includes: if the file size of the file is greater than a preset byte threshold, determining that the task type of the file is the first task type; if the file size of the file is less than or equal to the preset byte threshold, determining that the task type of the file is the second task type. In some embodiments of the present application, the first task type is a large task, and the second task type is a small task. In some embodiments of the present application, the preset byte threshold can be set according to user needs. For example, the preset byte threshold can be set to 1 Gb. The following is an example. If the file size of the file is greater than 1 Gb, determining that the task type of the file is a large task; if the file size of the file is less than or equal to 1 Gb, determining that the task type of the file is a small task.
[0033] Step S202, call the mapping and reduction (MapReduce) task of the data migration tool to migrate the file to the cloud server according to the task type of the file.
[0034] In some embodiments of the present application, the data migration tool can be the Hadoop distcp data migration tool. The Hadoop distcp data migration tool uses the cluster resource management system (Yet Another Resource Negotiator, YARN) to start the MapReduce task and performs data migration through the MapReduce task. In the embodiments of the present application, starting the MapReduce task through YARN can make full use of the cluster resources and efficiently complete large-scale data processing tasks. The resource management and scheduling capabilities of YARN enable the MapReduce task to dynamically allocate resources and run stably in a distributed environment.
[0035] In some embodiments of the present application, the migration of the file to the cloud server according to the task type of the file includes: if the task type of the file is the first task type, migrating the file to the cloud server based on the first bandwidth and the first concurrent number; if the task type of the file is the second task type, migrating the file to the cloud server based on the second bandwidth and the second concurrent number, where the first bandwidth is greater than the second bandwidth and the first concurrent number is less than the second concurrent number.
[0036] In the embodiments of the present application, when the task type of the file is the first task type, at this time, the data volume of the file is large. By using the first bandwidth and the first concurrency number to migrate the file to the cloud server, the bandwidth for large tasks can be preferentially guaranteed to ensure the migration speed, and resource contention caused by multiple large tasks running and migrating simultaneously can be avoided; when the task type of the file is the second task type, at this time, the data volume of the file is small, and the bandwidth requirement for a single task is low. By using the second bandwidth and the second concurrency number to migrate the file to the cloud server, the cluster resources can be fully utilized to quickly complete the migration of a large number of small files.
[0037] In some embodiments of the present application, when performing data migration through a MapReduce task, the method further includes: determining whether the file has been successfully migrated to the cloud server; if the file has not been successfully migrated to the cloud server, migrating the file to the cloud server again according to the task type of the file.
[0038] In some embodiments of the present application, when invoking the MapReduce task of the data migration tool to migrate the file to the cloud server according to the task type of the file, record the status of the MapReduce task; if the status of the MapReduce task is a successful status, it is determined that the file has been successfully migrated to the cloud server; if the status of the MapReduce task is a failed status, it is determined that the file has not been migrated to the cloud server, and the file is migrated to the cloud server again according to the task type of the file.
[0039] By recording the successful status and failed status of the MapReduce task in the embodiments of the present application, the full life cycle management of the data migration task can be realized, ensuring the traceability and maintainability of the task.
[0040] In some embodiments of the present application, migrating the file to the cloud server according to the task type of the file includes: storing the file in the object storage service (OSS) of the cloud server according to the task type of the file.
[0041] In the embodiments of the present application, according to the file size information of the file recorded by the Hadoop platform, the task type of the file is determined, and the MapReduce task is invoked to migrate the file to the cloud server according to the task type of the file. In this way, the migration task of the file can be divided into large tasks and small tasks according to the file size, and the file is migrated to the cloud server according to the large tasks and small tasks, thus avoiding low data migration efficiency caused by too large files and improving the data migration efficiency.
[0042] In some embodiments of the present application, after migrating the file from the Hadoop platform to the cloud server, the method can also update the data that has changed in the Hadoop platform to the cloud server. Refer to Figure 3As shown in the figure, it is a flowchart of a data migration method provided by some other embodiments of the present application. Figure 3 The exemplary method includes one or more steps, but does not constitute a limitation to the present application. In addition, the order of the steps of the method is only for example, and the order of the steps can be changed. Without departing from the content disclosed in the present application, additional steps can be added or steps can be reduced. The method includes the following steps.
[0043] Step S301: Determine the task type of the file according to the file size information of the file recorded by the Hadoop platform, where the task type includes a first task type and a second task type.
[0044] Step S302: Invoke the MapReduce task of the data migration tool, and migrate the file to the cloud server according to the task type of the file.
[0045] In some embodiments of the present application, the specific implementation manners of steps S301 - S302 can refer to Figure 2 the implementation content of steps S201 - S202 therein, which will not be described in detail here.
[0046] Step S303: Determine the file size information table according to the first audit log recorded by the Hadoop platform.
[0047] In some embodiments of the present application, in some embodiments of the present application, the Hadoop platform generates a first audit log according to the file data stored in the Hadoop platform. The first audit log records the file information and data partitions of all files stored in the Hadoop platform. The file information includes file location information, file size information, and metadata information of the file. The data partition means that the file or data is divided into multiple logical parts according to a preset rule (such as time, region, business attribute, etc.), each part is a partition, and each partition can be independently stored and managed.
[0048] Step S304: Determine the data partition table according to the first audit log and the second audit log recorded by the cloud server, where the data partition table includes the differential data partitions that are inconsistent between the Hadoop platform and the cloud server.
[0049] In some embodiments of the present application, the cloud server generates a second audit log according to the stored file data every preset period, and the audit log records the file information and data partitions of all files stored in the Hadoop platform.
[0050] In some embodiments of the present application, determining the data partition table according to the first audit log and the second audit log recorded by the cloud server includes: determining a first data partition set of the Hadoop platform according to the first audit log; determining a second data partition set of the cloud server according to the second audit log; calculating the partition difference between the first data partition set and the first data partition set to obtain a differential data partition; and recording the differential data partition in the data partition table.
[0051] In some embodiments of the present application, the partition difference between the first data partition set and the first data partition set is calculated using Spark Structured Query Language (SQL) to obtain a differential data partition. For example, the preset set operations (such as except, union, intersect) of Spark SQL can be used to calculate the differential data partition between the first data partition set and the first data partition set.
[0052] Step S305: Determine the task type of the differential data partition according to the data partition table and the file size information table.
[0053] In some embodiments of the present application, the data volume of the differential data partition is determined according to the data partition table and the file size information table. If the data volume of the differential data partition is greater than the preset data volume threshold, the task type of the differential data partition is determined as the first task type. If the data volume of the differential data partition is less than or equal to the preset data volume threshold, the task type of the differential data partition is determined as the second task type. In some embodiments of the present application, the preset data volume threshold can be set according to user needs. For example, the preset data volume threshold can be set to 1 Gb.
[0054] Step S306: Invoke the MapReduce task to migrate the data corresponding to the differential data partition to the cloud server according to the task type of the data partition.
[0055] In some embodiments of the present application, if the task type of the differential data partition is the first task type, the data corresponding to the differential data partition is migrated to the cloud server based on the first bandwidth and the first concurrency number. If the task type of the differential data partition is the second task type, the data corresponding to the differential data partition is migrated to the cloud server based on the second bandwidth and the second concurrency number, where the first bandwidth is greater than the second bandwidth and the first concurrency number is less than the second concurrency number.
[0056] In the embodiments of the present application, when the task type of the differential data partition is the first task type, the data volume of the differential data partition is large at this time. By using the first bandwidth and the first concurrency number to migrate the data corresponding to the differential data partition to the cloud server, the bandwidth of the large differential data partition can be preferentially guaranteed to ensure the migration speed, and resource contention caused by the simultaneous migration of data partitions corresponding to multiple large tasks can be avoided; when the task type of the differential data partition is the second task type, at this time, the data volume of the differential data partition is small, and the bandwidth requirement of a single task is low. By using the second bandwidth and the second concurrency number to migrate the data corresponding to the differential data partition to the cloud server, the cluster resources can be fully utilized to quickly complete the migration of a large number of small differential data partitions.
[0057] In some embodiments of the present application, after migrating the data corresponding to the differential data partition to the cloud server according to the task type of the data partition, the method further includes: determining whether the differential data partition has been successfully migrated to the cloud server; if the differential data partition has not been successfully migrated to the cloud server, migrating the data corresponding to the differential data partition to the cloud server again according to the task type of the data partition.
[0058] In the embodiments of the present application, when calling the MapReduce task to migrate the data corresponding to the differential data partition to the cloud server according to the task type of the data partition, the status of the MapReduce task is recorded. If the status of the MapReduce task is a successful status, it is determined that the differential data partition has been successfully migrated to the cloud server; if the status of the MapReduce task is a failed status, it is determined that the differential data partition has not been migrated to the cloud server, and the data corresponding to the differential data partition is migrated to the cloud server again according to the task type of the differential data partition.
[0059] In the embodiments of the present application, by recording the successful status and the failed status of the MapReduce task, the data synchronization and update of the Hadoop platform can be realized to the cloud server.
[0060] Please refer to Figure 4 As shown, it is a schematic structural diagram of an electronic device provided by some embodiments of the present application.
[0061] The electronic device 40 may include at least one memory 41, a processor 42, and a communication unit 43. The memory 41 includes a computer-readable storage medium, which is used to store computer programs, such as multiple logical instructions. The communication unit 43 is used to communicate with the electronic device 40 or the server. The processor 42 can run the logical instructions to execute the above data migration method.
[0062] It can be understood that in some embodiments, the communication unit 43 includes modules with communication functions such as a network module, and the present application does not limit this.
[0063] Among them, when the logical instructions in the above computer-readable storage medium can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0064] The computer-readable storage medium can be set to store software programs and computer-executable programs, such as the program instructions corresponding to the data migration method in the embodiments of the present application. The processor 42 realizes the data migration method in the above embodiments by running the software programs, instructions or modules stored in the computer-readable storage medium.
[0065] In the embodiments of the present application, the computer-readable storage medium includes non-volatile computer-readable memories, such as magnetic disks, memories, etc. It can be understood that the computer-readable storage medium may also include other non-volatile computer-readable memories, such as plug-in hard disks, smart media cards (Smart Media Card, SMC), secure digital (Secure Digital, SD) cards, flash cards (Flash Card), at least one flash device and / or other non-volatile solid-state storage devices.
[0066] In the embodiments of the present application, the processor 42 may be a central processing unit (Central Processing Unit, CPU), or may also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), field programmable gate arrays (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 42 is the control center of the electronic device 40, and can use various interfaces and lines to connect to other external devices and / or systems / modules / units to provide base sequencing functions for the applications of other external devices and / or systems / modules / units.
[0067] This embodiment also provides a computer program product. When the computer program product runs on a computer, it causes the computer to execute the above related steps to implement the data migration method in the above embodiments.
[0068] Among them, the electronic device, computer storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.
[0069] From the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above function modules is used as an example. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above.
[0070] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the module or unit is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0071] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it can be located in one place, or it can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0072] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0073] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device or processor to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk and other various media that can store program codes.
[0074] The above embodiments are only used to illustrate the technical solutions of the present application rather than to limit it. Although the present application has been described in detail with reference to the above preferred embodiments, those of ordinary skill in the art should understand that modifications or equivalent replacements can be made to the technical solutions of the present application without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A data migration method, characterized in that: The method comprises: Determine the task type of the file according to the file size information of the file recorded by the data platform; The MapReduce task of the data migration tool is called to migrate the file to the cloud server according to the task type of the file.
2. The data migration method according to claim 1, characterized in that: The task type includes a first task type and a second task type, and determining the task type of the file includes: If the file size of the file is greater than a preset byte threshold, determining that the task type of the file is the first task type; If the file size of the file is less than or equal to the preset byte threshold, it is determined that the task type of the file is the second task type.
3. The data migration method according to claim 1 or 2, characterized in that: Migrating the file to the cloud server according to the task type of the file includes: If the task type of the file is a first task type, migrating the file to the cloud server based on a first bandwidth and a first concurrent number; If the task type of the file is a second task type, the file is migrated to the cloud server based on a second bandwidth and a second concurrent number, wherein the first bandwidth is greater than the second bandwidth, and the first concurrent number is less than the second concurrent number.
4. The data migration method according to claim 1, characterized in that: The method further comprises: Determine a file size information table according to the first audit log recorded by the data platform; Determine a data partition table according to the first audit log and the second audit log recorded by the cloud server, wherein the data partition table includes differential data partitions that are inconsistent between the data platform and the cloud server; Determine the task type of the difference data partition according to the data partition table and the file size information table; The MapReduce task is called to migrate the data corresponding to the difference data partition to the cloud server according to the task type of the data partition.
5. The data migration method according to claim 4, characterized in that: The determining of the data partition table according to the first audit log and the second audit log recorded by the cloud server includes: Determine a first data partition set of the data platform according to the first audit log; determining a second data partition set of the cloud server according to the second audit log; Determining the difference data partition based on the partition difference between the first data partition set and the second data partition set; The differential data partition is recorded in the data partition table.
6. The data migration method according to claim 4, characterized in that: Migrating the data corresponding to the difference data partition to the cloud server according to the task type of the data partition includes: If the task type of the difference data partition is the first task type, migrating the data corresponding to the difference data partition to the cloud server based on the first bandwidth and the first concurrency number; If the task type of the difference data partition is the second task type, the data corresponding to the difference data partition is migrated to the cloud server based on the second bandwidth and the second concurrent number, wherein the first bandwidth is greater than the second bandwidth, and the first concurrent number is less than the second concurrent number.
7. The data migration method according to claim 5, characterized in that: The method further comprises: Determining whether the differential data partition is successfully migrated to the cloud server; If the differential data partition is not successfully migrated to the cloud server, the data corresponding to the differential data partition is re-migrated to the cloud server according to the task type of the data partition.
8. The data migration method according to claim 1, characterized in that: Migrating the file to the cloud server according to the task type of the file includes: The file is stored in an object storage service of a cloud server according to the task type of the file.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor: The memory is used to store program instructions; The processor is used to read and execute the program instructions stored in the memory. When the program instructions are executed by the processor, the electronic device executes the data migration method as described in any one of claims 1 to 8.
10. A computer storage medium, characterized in that: The computer storage medium stores program instructions, and when the program instructions are executed on an electronic device, the electronic device executes the data migration method according to any one of claims 1 to 8.