Offline data warehouse migration method and device, equipment and storage medium

By analyzing the timed scheduling task dependencies in offline data warehouses, prioritizing and rebuilding tasks, the problems of low migration efficiency and poor accuracy in the existing technology are solved, and an efficient and stable data migration process is achieved.

CN120336280APending Publication Date: 2025-07-18BEIJING BAIJU YIXING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510322891.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, when migrating offline data warehouses, timed scheduling tasks need to be manually redeveloped and deployed, resulting in low migration efficiency and accuracy, affecting data quality and warehouse stability.

Method used

By obtaining the dependencies of each timed scheduling task in the offline data warehouse, determining the upstream and downstream dependency sets, prioritizing tasks, building a target task sorting array, and rebuilding the task to the target data warehouse.

Benefits of technology

Improve the efficiency and accuracy of offline data warehouse migration, ensuring the data quality and warehouse stability after migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336280A_ABST
    Figure CN120336280A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of computers, and discloses an off-line data warehouse migration method, device and equipment and a storage medium, and the method comprises the following steps: obtaining a first dependency relationship between each timing scheduling task in an off-line data warehouse and other timing scheduling tasks; determining an upstream dependency relationship set and a downstream dependency relationship set corresponding to the timed scheduling task based on the first dependency relationship; on the basis of the upstream dependency relationship set and the downstream dependency relationship set, task priority ranking is conducted on all the timed scheduling tasks, and a target task ranking array is obtained; and based on the target task sorting array and the dependency relationship, reconstructing each timed scheduling task to the target data warehouse so as to migrate the offline data warehouse. By applying the technical scheme of the invention, the efficiency and accuracy of migrating the off-line data warehouse can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technology, and particularly to an offline data warehouse migration method, apparatus, device, and storage medium. Background Art

[0002] In the context of the rapid development of current computer technology, offline data warehouse migration has become an important part of the data processing field. Offline data warehouse migration involves migrating a large amount of data from one storage location or format to another, usually for data backup, data archiving, data integration, or data migration to a new storage system. However, in related technologies, when performing offline data warehouse migration, it is usually necessary to manually re-develop, deploy, and go live with each scheduled task in the new environment, resulting in low migration efficiency and accuracy, thereby affecting the data quality and warehouse stability after migration. Summary of the Invention

[0003] In view of the above problems, embodiments of the present invention provide an offline data warehouse migration method, apparatus, device, and storage medium.

[0004] According to one aspect of the embodiments of the present invention, an offline data warehouse migration method is provided. The method includes: obtaining a first dependency relationship between each scheduled task in the offline data warehouse and other scheduled tasks; determining an upstream dependency relationship set and a downstream dependency relationship set corresponding to the scheduled task based on the first dependency relationship; performing task priority sorting on each scheduled task based on the upstream dependency relationship set and the downstream dependency relationship set to obtain a target task sorting array; and reconstructing each scheduled task into the target data warehouse based on the target task sorting array and the dependency relationship to migrate the offline data warehouse. Through the above process, the efficiency and accuracy of offline data warehouse migration can be improved, and at the same time, the data quality and warehouse stability after migration can be ensured.

[0005] In an optional implementation manner, performing task priority sorting on each scheduled task based on the upstream dependency relationship set and the downstream dependency relationship set to obtain a target task sorting array includes:

[0006] obtaining a second dependency relationship of each upstream scheduled task in the upstream dependency relationship set and a third dependency relationship of each downstream scheduled task in the downstream dependency relationship set;

[0007] performing priority sorting on each upstream scheduled task based on the second dependency relationship to obtain an upstream task sorting;

[0008] performing priority sorting on each downstream scheduled task based on the third dependency relationship to obtain a downstream task sorting;

[0009] obtaining a target task sorting array based on the upstream task sorting and the downstream task sorting.

[0010] In an alternative embodiment, priority sorting is performed on each upstream timing scheduling task based on the second dependency relationship to obtain an upstream task sorting, including:

[0011] If the second dependency relationship is a dependency, determine the upstream parallel priority order of the corresponding upstream timing scheduling task in the upstream dependency relationship set;

[0012] If the second dependency relationship is a non-dependency, determine the upstream serial priority order of the corresponding upstream timing scheduling task in the upstream dependency relationship set;

[0013] Based on the upstream parallel priority order and the upstream serial priority order, priority sorting is performed on each upstream timing scheduling task to obtain an upstream task sorting.

[0014] In an alternative embodiment, priority sorting is performed on each downstream timing scheduling task based on the third dependency relationship to obtain a downstream task sorting, including:

[0015] If the third dependency relationship is a dependency, determine the downstream parallel priority order of the corresponding downstream timing scheduling task in the downstream dependency relationship set;

[0016] If the second dependency relationship is a non-dependency, determine the downstream serial priority order of the corresponding upstream timing scheduling task in the upstream dependency relationship set;

[0017] Based on the downstream parallel priority order and the downstream serial priority order, priority sorting is performed on each downstream timing scheduling task to obtain a downstream task sorting.

[0018] In an alternative embodiment, based on the target task sorting array and the dependency relationship, each timing scheduling task is reconstructed into the target data warehouse to migrate the offline data warehouse, including:

[0019] Obtain the scheduling information of each timing scheduling task in the offline data warehouse;

[0020] Based on the order of the target task sorting array, the dependency relationship, and the scheduling information, each timing scheduling task is reconstructed into the target data warehouse to migrate the offline data warehouse.

[0021] In an alternative embodiment, based on the dependency relationship, determine the upstream dependency relationship set and the downstream dependency relationship set corresponding to the timing scheduling task, including:

[0022] Obtain the task type of the timing scheduling task;

[0023] If the task type indicates that the timing scheduling task is a task starting point, determine the downstream dependency relationship set corresponding to the timing scheduling task based on the first dependency relationship;

[0024] When the task type represents that the timing scheduling task is the task end point, determine the upstream dependency relationship set corresponding to the timing scheduling task based on the first dependency relationship;

[0025] When the task type represents that the timing scheduling task is neither the task start point nor the task end point, determine the upstream dependency relationship set and the downstream dependency relationship set corresponding to the timing scheduling task based on the first dependency relationship.

[0026] In an alternative embodiment, based on the upstream dependency relationship set and the downstream dependency relationship set, perform task priority sorting on each timing scheduling task to obtain a target task sorting array, and further include:

[0027] Obtain multiple timing scheduling tasks with the task type being the task start point;

[0028] Construct root tasks for multiple timing scheduling tasks;

[0029] Construct an initial task sorting array based on the root tasks;

[0030] Update the initial task sorting array based on the downstream dependency relationship set and the timing scheduling tasks to obtain a target task sorting array.

[0031] According to another aspect of the embodiments of the present invention, there is provided an offline data warehouse migration device, including: a relationship acquisition module, configured to acquire the first dependency relationship between each timing scheduling task in the offline data warehouse and other timing scheduling tasks; a data determination module, configured to determine the upstream dependency relationship set and the downstream dependency relationship set corresponding to the timing scheduling task based on the first dependency relationship; a task sorting module, configured to perform task priority sorting on each timing scheduling task based on the upstream dependency relationship set and the downstream dependency relationship set to obtain a target task sorting array; a data migration module, configured to reconstruct each timing scheduling task into the target data warehouse based on the target task sorting array and the dependency relationship to migrate the offline data warehouse. Through the above modules, the efficiency and accuracy of migrating the offline data warehouse can be improved, and at the same time, the data quality and warehouse stability after migration can be ensured.

[0032] According to another aspect of the embodiments of the present invention, there is provided a computer device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations of the foregoing offline data warehouse migration method.

[0033] According to yet another aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, in which at least one executable instruction is stored, and the executable instruction causes the computer device / device to perform the operations of the foregoing offline data warehouse migration method.

[0034] According to another aspect of the embodiments of the present invention, there is provided a computer program product including computer instructions for causing a computer to perform the operations of the offline data warehouse migration method according to the first aspect or any corresponding embodiment thereof as described above.

[0035] The above description is only an overview of the technical solutions of the embodiments of the present invention. In order to be able to understand the technical means of the embodiments of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of the embodiments of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Description of the Drawings

[0036] The drawings are only used to illustrate the embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0037] Figure 1 A flowchart showing a process of an offline data warehouse migration method provided by the present invention is shown;

[0038] Figure 2 Another flowchart showing a process of an offline data warehouse migration method provided by the present invention is shown;

[0039] Figure 3 A schematic diagram of an initial task sorting array of an offline data warehouse migration method provided by the present invention is shown;

[0040] Figure 4 A schematic diagram of a target task sorting array of an offline data warehouse migration method provided by the present invention is shown;

[0041] Figure 5 A schematic diagram of the structure of an offline data warehouse migration device provided by the present invention is shown;

[0042] Figure 6 A schematic diagram of the structure of a computer device provided by the present invention is shown. Detailed Embodiments

[0043] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] Figure 1 A flowchart of the first embodiment of an offline data warehouse migration method of the present invention is shown. As Figure 1As shown in the figure, the method includes the following steps:

[0045] Step 110: Obtain the first dependency relationship between each scheduled task and other scheduled tasks in the offline data warehouse.

[0046] As mentioned above, by deeply analyzing the dependency relationships between scheduled tasks in the offline data warehouse, it is possible to ensure that the original task execution logic is not damaged during the migration process. The first dependency relationship may include direct dependencies and indirect dependencies. A direct dependency means that one task directly depends on the execution result of another task, while an indirect dependency may involve a complex dependency chain among multiple tasks. After obtaining the first dependency relationship, the system can parse and store these dependency relationships for reconstructing the scheduled tasks according to these dependency relationships in subsequent migration steps.

[0047] In addition, to more comprehensively understand and process the above-mentioned dependency relationships, it is also possible to further analyze the execution cycle, execution time of the scheduled tasks, and the priority relationships between tasks. The execution cycle and execution time help determine the specific time window for task migration to avoid migration operations during peak business periods, thus affecting normal business operations. The priority relationships between tasks can guide the system on how to reasonably arrange the migration order of tasks when resources are limited during the migration process to ensure that critical business tasks can be processed first.

[0048] After parsing and storing these dependency relationships, a detailed migration plan can be formulated based on this information. The migration plan can include the migration schedule, the migration order of each scheduled task, and the task configuration after migration, etc. By formulating a detailed migration plan, it is possible to ensure the smooth progress of the migration process and minimize the impact on business operations.

[0049] Step 120: Determine the upstream dependency relationship set and downstream dependency relationship set corresponding to the scheduled tasks based on the first dependency relationship.

[0050] Among them, the upstream dependency relationship set refers to the set of all dependent tasks that must be completed before the current scheduled task is executed. The successful execution of these tasks is a prerequisite for the execution of the current task. The downstream dependency relationship set refers to the set of all subsequent tasks that depend on the execution result of the current scheduled task. The execution result of the current task will directly affect the execution of these subsequent tasks. By clarifying the upstream and downstream dependency relationship sets, it is possible to clearly understand the position and role of each scheduled task in the overall task chain, thereby more effectively managing and optimizing the migration process.

[0051] In some alternative embodiments, when determining the upstream dependency set and the downstream dependency set corresponding to a timed scheduling task based on dependencies, the task type of the timed scheduling task may be obtained first; if the task type indicates that the timed scheduling task is a task start point, then the downstream dependency set corresponding to the timed scheduling task is determined based on the first dependency; if the task type indicates that the timed scheduling task is a task end point, then the upstream dependency set corresponding to the timed scheduling task is determined based on the first dependency; if the task type indicates that the timed scheduling task is neither a task start point nor a task end point, then the upstream dependency set and the downstream dependency set corresponding to the timed scheduling task are determined based on the first dependency.

[0052] As described above, by specifically determining the dependency set according to the special properties of different task types during the migration process, the efficiency and accuracy of the migration can be further improved. For a task start point, it has no upstream dependent tasks, so only the downstream dependency set needs to be concerned to ensure that these subsequent tasks can correctly receive and process the execution results of the start point task during the migration process. For a task end point, it has no downstream dependent tasks, so only the upstream dependency set needs to be concerned to ensure that these pre - tasks have been successfully executed before the migration, providing a reliable premise for the migration and execution of the end point task. For tasks that are neither start points nor end points, both their upstream and downstream dependency sets need to be considered to ensure that these tasks can correctly receive the outputs of the pre - tasks and pass their own execution results to the subsequent tasks during the migration process, making the migration plan more flexible and adaptable, reducing the migration risk, increasing the success rate of the migration, and thus better meeting the migration requirements in different business scenarios.

[0053] Step 130: Based on the upstream dependency set and the downstream dependency set, perform task priority sorting on each timed scheduling task to obtain a target task sorting array.

[0054] As described above, by performing task priority sorting on each timed scheduling task based on the upstream dependency set and the downstream dependency set, a more reasonable and efficient migration order can be obtained, thereby guiding the offline data warehouse migration process more reasonably and efficiently.

[0055] In some alternative embodiments, when performing task priority sorting on each timed scheduling task based on the upstream dependency set and the downstream dependency set to obtain the target task sorting array, the second dependencies of each upstream timed scheduling task in the upstream dependency set and the third dependencies of each downstream timed scheduling task in the downstream dependency set can be obtained first; perform priority sorting on each upstream timed scheduling task based on the second dependencies to obtain the upstream task sorting; perform priority sorting on each downstream timed scheduling task based on the third dependencies to obtain the downstream task sorting; based on the upstream task sorting and the downstream task sorting, obtain the target task sorting array, which can avoid the blocking and interruption of each timed scheduling task during the migration process, and the problem that the current timed scheduling task cannot go online after import due to the incomplete import of the upstream timed scheduling task, resulting in the need to continuously repeat the import or manual intervention, thus causing bottlenecks in project progress. Among them, the target task sorting array can be stored in a data table or a local file.

[0056] Specifically, if the second dependency is a dependency, determine the upstream parallel priority order of the corresponding upstream timed scheduling task in the upstream dependency set; if the second dependency is a non-dependency, determine the upstream serial priority order of the corresponding upstream timed scheduling task in the upstream dependency set; perform priority sorting on each upstream timed scheduling task based on the upstream parallel priority order and the upstream serial priority order to obtain the upstream task sorting. Secondly, if the third dependency is a dependency, determine the downstream parallel priority order of the corresponding downstream timed scheduling task in the downstream dependency set; if the second dependency is a non-dependency, determine the downstream serial priority order of the corresponding upstream timed scheduling task in the upstream dependency set; perform priority sorting on each downstream timed scheduling task based on the downstream parallel priority order and the downstream serial priority order to obtain the downstream task sorting. Subsequently, merge the upstream task sorting and the downstream task sorting, and adjust the tasks with dependency relationships in the order of the dependency relationships to ensure that all timed scheduling tasks are arranged in the correct order. During the merging and adjustment process, factors such as the urgency of the tasks, the execution time, and the resource requirements can be considered to further optimize the task sorting. Finally, a target task sorting array that meets all dependency relationships and is efficient and reasonable is obtained.

[0057] In some alternative embodiments, when obtaining the upstream dependency set and the downstream dependency set, multiple factors such as the execution time of the task, the importance of the task, and the complexity of the task can also be comprehensively considered based on the second dependency and the third dependency. For example, for tasks with a longer execution time or higher importance, a higher priority can be assigned to ensure that these tasks can be executed as early as possible. At the same time, for tasks with a higher complexity, it can also be considered to split them into multiple subtasks and assign appropriate priorities to each subtask to improve the overall migration efficiency.

[0058] In some alternative embodiments, the sorting principle of the target task sorting array can also be comprehensively considered based on factors such as the urgency, importance, and data volume of the tasks. For example, for tasks with fewer upstream dependencies and a smaller data volume, a higher priority can be assigned to quickly complete the migration; while for tasks with numerous upstream dependencies and a large data volume, the priority needs to be appropriately reduced to ensure that all prerequisite tasks can be successfully executed before the migration. Through such a sorting method, the migration plan can be made more reasonable and efficient, further improving the success rate and stability of the migration.

[0059] In addition, to improve the sorting efficiency and accuracy, quicksort, merge sort, etc. can be used to optimize the target task sorting array. During the sorting process, parallel processing technology can also be introduced to utilize multi-core CPUs or multiple computers to process the sorting tasks simultaneously, thereby greatly shortening the sorting time and improving the migration efficiency. At the same time, to ensure the stability and reliability of the sorting results, the sorting algorithm can be fully tested and verified to ensure that it can run correctly in various complex scenarios, providing strong support for the migration of the offline data warehouse.

[0060] Step 140, based on the target task sorting array and the dependency relationship, reconstruct each scheduled task to the target data warehouse to migrate the offline data warehouse.

[0061] Specifically, during the reconstruction process, each scheduled task can be migrated from the source data warehouse to the target data warehouse one by one according to the sorted target task array and the dependency relationship. During the migration, the system will ensure that the prerequisite tasks of each task have been successfully migrated and can run normally in the target data warehouse. For complex dependency relationships, the system will perform in-depth parsing to ensure that the reconstructed task chain can maintain the original logical relationship and timing requirements. Through such a migration method, the integrity and consistency of the offline data warehouse can be maximally guaranteed, providing a reliable basis for subsequent data analysis and applications.

[0062] In some alternative embodiments, when reconstructing each scheduled task to the target data warehouse based on the target task sorting array and the dependency relationship to migrate the offline data warehouse, the scheduling information of each scheduled task in the offline data warehouse can be obtained first; based on the order, dependency relationship, and scheduling information of the target task sorting array, each scheduled task can be reconstructed to the target data warehouse to migrate the offline data warehouse.

[0063] In the process of obtaining scheduling information, by traversing all the scheduled tasks in the offline data warehouse, key information such as the scheduling time, execution frequency, and task type of each scheduled task is extracted. This information is crucial for subsequent task reconstruction because it determines the execution method and timing of tasks in the target data warehouse. Subsequently, according to the order of the target task sorting array, combined with the dependency relationship and scheduling information, the scheduled tasks are gradually reconstructed in the target data warehouse. During the reconstruction process, by identifying and processing the dependency relationships between scheduled tasks, it is ensured that each scheduled task starts to be constructed only after the scheduled tasks it depends on are completed. At the same time, according to the scheduling information, the correct scheduling time and execution frequency are set for each scheduled task to ensure that the scheduled tasks can run in the target data warehouse in the expected manner and timing. This not only ensures the correctness and timeliness of the scheduled tasks, but also greatly improves the efficiency and accuracy of data migration, so that the scheduled tasks in the offline data warehouse can be seamlessly migrated to the target data warehouse, providing a solid foundation for subsequent data processing and analysis.

[0064] The offline data warehouse migration method according to the embodiment of the present invention obtains the first dependency relationship between each scheduled task in the offline data warehouse and other scheduled tasks; determines the upstream dependency relationship set and the downstream dependency relationship set corresponding to the scheduled task based on the first dependency relationship; sorts the priorities of each scheduled task based on the upstream dependency relationship set and the downstream dependency relationship set to obtain a target task sorting array; and reconstructs each scheduled task into the target data warehouse based on the target task sorting array and the dependency relationship to migrate the offline data warehouse. Through the above process, the efficiency and accuracy of offline data warehouse migration can be improved, and at the same time, the data quality and warehouse stability after migration can be ensured.

[0065] Figure 2 The flowchart of another embodiment of the offline data warehouse migration method of the present invention is shown. As Figure 2 shown, the method includes the following steps:

[0066] Step 210, obtain the first dependency relationship between each scheduled task in the offline data warehouse and other scheduled tasks.

[0067] For details, please refer to Figure 1 Step 110 of the embodiment shown, which will not be elaborated here.

[0068] Step 220, determine the upstream dependency relationship set and the downstream dependency relationship set corresponding to the scheduled task based on the first dependency relationship.

[0069] For details, please refer to Figure 1 Step 120 of the embodiment shown, which will not be elaborated here.

[0070] Step 230: Based on the upstream dependency set and the downstream dependency set, perform task priority sorting on each timed scheduling task to obtain a target task sorting array.

[0071] Specifically, the above Step 230 includes:

[0072] Step 2301: Obtain multiple timed scheduling tasks whose task type is the task starting point.

[0073] Among them, the above task starting points are usually timed scheduling tasks in the data warehouse without upstream dependencies, and they can be used as the starting points of the migration process. The determination of these task starting points is crucial for the subsequent migration process because they do not depend on the completion of other tasks and can start execution independently.

[0074] Step 2302: Construct root tasks for multiple timed scheduling tasks.

[0075] Specifically, the process of constructing root tasks is mainly based on the timed scheduling tasks at the task starting points. For each task starting point, we mark it as a root task and store these root tasks in a specific data structure for subsequent management and operation. These root tasks are the starting nodes of the migration process. They do not depend on the completion of other tasks and can therefore start execution independently. In this way, we can clearly define the starting points of the offline data warehouse migration and provide a solid foundation for subsequent task scheduling and execution.

[0076] In some optional implementation manners, when constructing root tasks based on the timed scheduling tasks at the task starting points, virtual root tasks can also be constructed based on the priorities of the timed scheduling tasks at the task starting points, so that the timed scheduling tasks at multiple task starting points depend on the virtual root task, enabling users to more flexibly adjust the execution order of the task starting points according to actual needs. For example, in some cases, although multiple task starting points can start execution independently logically, for business considerations, it is desired that they can be executed in a specific priority order. At this time, we can create a virtual root task and set these task starting points to depend on the virtual root task. Then, by adjusting the priorities of the task starting points under the virtual root task, precise control over the execution order of the entire migration process can be achieved, not only improving the flexibility of the migration process but also enhancing its adaptability to complex business requirements.

[0077] Step 2303: Based on the root tasks, construct an initial task sorting array.

[0078] Among them, the above initial task sorting array is used to record the execution order of the starting points of each timed scheduling task. When constructing the initial task sorting array, the root task can be used as the first element of the array, and then the task starting points are added to the array in turn according to the priority and dependency relationship of the task starting points. If there is a virtual root task, all task starting points that depend on the virtual root task are added to the array in the order of their priorities, and the virtual root task itself is not used as an element of the array. In this way, it can be ensured that during the subsequent task execution process, tasks can be carried out in a predetermined order, thus meeting the business requirements. In addition, when constructing the initial task sorting array, the parallel and serial relationships between tasks also need to be considered. For task starting points that can be executed in parallel, appropriate positions in the array can be marked to indicate that the system allows these tasks to start simultaneously during execution. For tasks that must be executed serially, they are arranged strictly in the order of their dependency relationships and priorities to ensure that the next task can only start after the previous task is completed. Such a design not only improves the efficiency of task execution but also ensures the consistency and accuracy of data migration.

[0079] Step 2304, update the initial task sorting array based on the downstream dependency set and the timed scheduling tasks to obtain the target task sorting array.

[0080] During the update process, check the downstream dependency set of each timed scheduling task to ensure that all downstream tasks are executed after their dependent upstream tasks. For timed scheduling tasks without downstream dependencies, they are regarded as leaf nodes and placed at the end of the sorting array during traversal. Through recursive traversal and dependency check, a complete and ordered task execution sequence, that is, the target task sorting array, is constructed. This sorting array ensures the correct execution order of tasks during the data migration process, thereby improving the efficiency and accuracy of data migration.

[0081] In specific implementation, an initial task sorting array arr_res can be constructed first. As Figure 3 shown, starting from the root task (if there are multiple timed scheduling tasks with first-level task starting points, a virtual root task is constructed virtually so that multiple timed scheduling tasks with first-level task starting points depend on the virtual root task), traverse based on the obtained upstream dependency set A. Judge that if all upstream timed scheduling tasks of the current timed scheduling task exist in the initial task sorting array arr_res, or there are no upstream timed scheduling tasks and the ID of this timed scheduling task does not exist in the initial task sorting array arr_res, then add the ID of this timed scheduling task to the initial task sorting array arr_res, continue to traverse the downstream dependency set of this timed scheduling task, and recursively repeat this logical calculation; otherwise, return to the previous layer and end the recursive logic of this layer. Finally, obtain the sorted result array of the target task sorting array arr_rest, asFigure 4 As shown, the IDs of the scheduled tasks are sorted in sequence to ensure that all upstream scheduled tasks of each scheduled task are before this scheduled task.

[0082] Step 240: Based on the target task sorting array and the dependency relationship, reconstruct each scheduled task to the target data warehouse to migrate the offline data warehouse.

[0083] For details, please refer to Figure 1 Step 140 of the embodiment shown, which will not be elaborated here.

[0084] In summary, the offline data warehouse migration method of the embodiment of the present invention obtains the first dependency relationship between each scheduled task and other scheduled tasks in the offline data warehouse; determines the upstream dependency relationship set and the downstream dependency relationship set corresponding to the scheduled task based on the first dependency relationship; performs task priority sorting on each scheduled task based on the upstream dependency relationship set and the downstream dependency relationship set to obtain the target task sorting array; based on the target task sorting array and the dependency relationship, reconstruct each scheduled task to the target data warehouse to migrate the offline data warehouse, which can improve the efficiency and accuracy of migrating the offline data warehouse, and at the same time ensure the data quality and warehouse stability after migration.

[0085] Figure 5 shows a schematic structural diagram of an embodiment of an offline data warehouse migration device of the present invention. As Figure 5 shown, the device includes:

[0086] A relationship acquisition module 510, configured to obtain the first dependency relationship between each scheduled task and other scheduled tasks in the offline data warehouse;

[0087] A data determination module 520, configured to determine the upstream dependency relationship set and the downstream dependency relationship set corresponding to the scheduled task based on the first dependency relationship;

[0088] A task sorting module 530, configured to perform task priority sorting on each scheduled task based on the upstream dependency relationship set and the downstream dependency relationship set to obtain the target task sorting array;

[0089] A data migration module 540, configured to reconstruct each scheduled task to the target data warehouse based on the target task sorting array and the dependency relationship to migrate the offline data warehouse.

[0090] In an optional implementation manner, the task sorting module 530 includes:

[0091] A dependency relationship acquisition sub-module, configured to obtain the second dependency relationship of each upstream scheduled task in the upstream dependency relationship set and the third dependency relationship of each downstream scheduled task in the downstream dependency relationship set;

[0092] An upstream task sorting sub-module, configured to perform priority sorting on each upstream timed scheduling task based on the second dependency relationship to obtain an upstream task sorting;

[0093] A downstream task sorting sub-module, configured to perform priority sorting on each downstream timed scheduling task based on the third dependency relationship to obtain a downstream task sorting;

[0094] A target array sub-acquisition module, configured to obtain a target task sorting array based on the upstream task sorting and the downstream task sorting.

[0095] In an alternative embodiment, the upstream task sorting sub-module includes:

[0096] An upstream parallel order unit, configured to determine the upstream parallel priority order of the corresponding upstream timed scheduling task in the upstream dependency relationship set if the second dependency relationship is a dependency;

[0097] An upstream serial order unit, configured to determine the upstream serial priority order of the corresponding upstream timed scheduling task in the upstream dependency relationship set if the second dependency relationship is a non-dependency;

[0098] An upstream task sorting unit, configured to perform priority sorting on each upstream timed scheduling task based on the upstream parallel priority order and the upstream serial priority order to obtain an upstream task sorting.

[0099] In an alternative embodiment, the downstream task sorting sub-module includes:

[0100] A downstream parallel order unit, configured to determine the downstream parallel priority order of the corresponding downstream timed scheduling task in the downstream dependency relationship set if the third dependency relationship is a dependency;

[0101] A downstream serial order unit, configured to determine the downstream serial priority order of the corresponding upstream timed scheduling task in the upstream dependency relationship set if the second dependency relationship is a non-dependency;

[0102] A downstream task sorting unit, configured to perform priority sorting on each downstream timed scheduling task based on the downstream parallel priority order and the downstream serial priority order to obtain a downstream task sorting.

[0103] In an alternative embodiment, the data migration module 540 includes:

[0104] A scheduling information acquisition sub-module, configured to acquire the scheduling information of each timed scheduling task in the offline data warehouse;

[0105] A data warehouse reconstruction sub-module, which is used to reconstruct each timed scheduling task to a target data warehouse based on the order, dependency relationship, and scheduling information of the target task sorting array, so as to migrate the offline data warehouse.

[0106] In some alternative embodiments, the data determination module 520 further includes:

[0107] A task type acquisition sub-module, which is used to acquire the task type of the timed scheduling task;

[0108] A downstream relationship set acquisition sub-module, which is used to, if the task type indicates that the timed scheduling task is a task starting point, determine a downstream dependency relationship set corresponding to the timed scheduling task based on the first dependency relationship;

[0109] An upstream relationship set acquisition sub-module, which is used to, if the task type indicates that the timed scheduling task is a task ending point, determine an upstream dependency relationship set corresponding to the timed scheduling task based on the first dependency relationship;

[0110] An upstream and downstream relationship set acquisition sub-module, which is used to, if the task type indicates that the timed scheduling task is neither a task starting point nor a task ending point, determine an upstream dependency relationship set and a downstream dependency relationship set corresponding to the timed scheduling task based on the first dependency relationship.

[0111] In some alternative embodiments, the task sorting module 530 is further used to acquire multiple timed scheduling tasks with the task type of task starting point; construct root tasks for the multiple timed scheduling tasks; construct an initial task sorting array based on the root tasks; and update the initial task sorting array based on the downstream dependency relationship set and the timed scheduling tasks to obtain a target task sorting array.

[0112] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding method embodiments described above, and will not be elaborated here.

[0113] Through the above-mentioned device and its components, the technical solution provided by the embodiments of the present invention has the following advantages:

[0114] The offline data warehouse migration device according to the embodiments of the present invention acquires the first dependency relationship between each timed scheduling task in the offline data warehouse and other timed scheduling tasks; determines the upstream dependency relationship set and the downstream dependency relationship set corresponding to the timed scheduling task based on the first dependency relationship; performs task priority sorting on each timed scheduling task based on the upstream dependency relationship set and the downstream dependency relationship set to obtain a target task sorting array; and reconstructs each timed scheduling task to the target data warehouse based on the target task sorting array and the dependency relationship to migrate the offline data warehouse, which can improve the efficiency and accuracy of migrating the offline data warehouse, and at the same time ensure the data quality and warehouse stability after migration.

[0115] Please refer to Figure 6 ,Figure 6 FIG. Figure 6 is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As Figure 6 shown, the computer device includes: one or more processors 610, a memory 620, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 6 Here, one processor 610 is taken as an example.

[0116] The processor 610 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 610 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.

[0117] Among them, the memory 620 stores instructions executable by at least one processor 610, so that at least one processor 610 executes the method shown in the above embodiment.

[0118] The memory 620 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device presented by a kind of landing page of a small program, etc. In addition, the memory 620 can include a high-speed random access memory and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 620 can optionally include a memory remotely set relative to the processor 610, and these remote memories can be connected to the computer device through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a server cluster, a mobile communication network, and combinations thereof.

[0119] The memory 620 can include a volatile memory, for example, a random access memory; the memory can also include a non-volatile memory, for example, a flash memory, a hard disk, or a solid-state drive; the memory 620 can also include a combination of the above types of memories.

[0120] The computer device further includes a communication interface 630 for the computer device to communicate with other devices or a communication network.

[0121] An embodiment of the present invention also provides a computer-readable storage medium storing at least one executable instruction. When the executable instruction runs on a computer device / offline data warehouse migration device, the computer device / offline data warehouse migration device is caused to execute the offline data warehouse migration method in any of the above method embodiments.

[0122] An embodiment of the present invention also provides a computer program product including computer instructions for causing a computer to execute the offline data warehouse migration method in the first aspect or any corresponding implementation manner thereof above.

[0123] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. In addition, embodiments of the present invention are not directed to any particular programming language.

[0124] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that embodiments of the present invention may be practiced without these specific details. Similarly, in order to streamline the present invention and assist in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of the present invention above, the various features of the embodiments of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. Among them, the claims following the specific implementation manner are hereby expressly incorporated into the specific implementation manner, where each claim itself serves as a separate embodiment of the present invention.

[0125] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into a module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive.

[0126] It should be noted that the above embodiments illustrate the present invention rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. An offline data warehouse migration method, characterized in that, The method includes: Obtaining the first dependencies between each scheduled task and other scheduled tasks in the offline data warehouse; Determining the upstream dependency set and the downstream dependency set corresponding to the scheduled task based on the first dependencies; Performing task priority sorting on each of the scheduled tasks based on the upstream dependency set and the downstream dependency set to obtain a target task sorting array; Reconstructing each of the scheduled tasks into the target data warehouse based on the target task sorting array and the dependencies to migrate the offline data warehouse.

2. The method according to claim 1, wherein The performing task priority sorting on each of the scheduled tasks based on the upstream dependency set and the downstream dependency set to obtain a target task sorting array includes: Obtaining the second dependencies of each upstream scheduled task in the upstream dependency set and the third dependencies of each downstream scheduled task in the downstream dependency set; Performing priority sorting on each of the upstream scheduled tasks based on the second dependencies to obtain an upstream task sorting; Performing priority sorting on each of the downstream scheduled tasks based on the third dependencies to obtain a downstream task sorting; Obtaining the target task sorting array based on the upstream task sorting and the downstream task sorting.

3. The method according to claim 2, characterized in that, The performing priority sorting on each of the upstream scheduled tasks based on the second dependencies to obtain an upstream task sorting includes: If the second dependency is a dependency, determining the upstream parallel priority order of the corresponding upstream scheduled task in the upstream dependency set; If the second dependency is a non-dependency, determining the upstream serial priority order of the corresponding upstream scheduled task in the upstream dependency set; Performing priority sorting on each of the upstream scheduled tasks based on the upstream parallel priority order and the upstream serial priority order to obtain the upstream task sorting.

4. The method according to claim 2, wherein The performing priority sorting on each of the downstream scheduled tasks based on the third dependencies to obtain a downstream task sorting includes: If the third dependency is a dependency, determining the downstream parallel priority order of the corresponding downstream scheduled task in the downstream dependency set; If the second dependency is a non-dependency, determining the downstream serial priority order of the corresponding upstream scheduled task in the upstream dependency set; Performing priority sorting on each of the downstream scheduled tasks based on the downstream parallel priority order and the downstream serial priority order to obtain the downstream task sorting.

5. The method according to claim 1, wherein The reconstructing each of the scheduled tasks into the target data warehouse based on the target task sorting array and the dependencies to migrate the offline data warehouse includes: Obtaining the scheduling information of each of the scheduled tasks in the offline data warehouse; Reconstructing each of the scheduled tasks into the target data warehouse based on the order of the target task sorting array, the dependencies, and the scheduling information to migrate the offline data warehouse.

6. The method according to claim 1 or 2, characterized in that, Determining an upstream dependency set and a downstream dependency set corresponding to the timing scheduling task based on the dependency relationship includes: Obtaining the task type of the timing scheduling task; If the task type indicates that the timing scheduling task is a task starting point, determining the downstream dependency set corresponding to the timing scheduling task based on the first dependency relationship; If the task type indicates that the timing scheduling task is a task ending point, determining the upstream dependency set corresponding to the timing scheduling task based on the first dependency relationship; If the task type indicates that the timing scheduling task is neither the task starting point nor the task ending point, determining the upstream dependency set and the downstream dependency set corresponding to the timing scheduling task based on the first dependency relationship.

7. The method according to claim 6, wherein Based on the upstream dependency set and the downstream dependency set, performing task priority sorting on each timing scheduling task to obtain a target task sorting array, and further includes: Obtaining multiple timing scheduling tasks whose task type is a task starting point; Constructing root tasks for the multiple timing scheduling tasks; Constructing an initial task sorting array based on the root tasks; Updating the initial task sorting array based on the downstream dependency set and the timing scheduling task to obtain the target task sorting array.

8. An offline data warehouse migration device, characterized in that The apparatus includes: A relationship acquisition module, configured to acquire a first dependency relationship between each timing scheduling task in an offline data warehouse and other timing scheduling tasks; A data determination module, configured to determine an upstream dependency set and a downstream dependency set corresponding to the timing scheduling task based on the first dependency relationship; A task sorting module, configured to perform task priority sorting on each timing scheduling task based on the upstream dependency set and the downstream dependency set to obtain a target task sorting array; A data migration module, configured to reconstruct each timing scheduling task into a target data warehouse based on the target task sorting array and the dependency relationship to migrate the offline data warehouse.

9. A computer device, characterized in that, Including: A processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete mutual communication through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations of the offline data warehouse migration method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, At least one executable instruction is stored in the storage medium, and when the executable instruction runs on a computer device, it causes the computer device to execute the operations of the offline data warehouse migration method according to any one of claims 1-7.