File Migration Method and Device

Through large language model and syntax tree analysis, DAG files are automatically migrated from one scheduling system to another scheduling system, solving the inefficiency and error problems caused by manual configuration in the existing technology, and achieving efficient and accurate file migration.

CN120104589BActive Publication Date: 2025-07-11SHANGHAI HODE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510578015.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-11
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

When migrating files between different scheduling systems, the existing technology requires manual comparison of old task configuration, which consumes a lot of manpower and time, and is prone to inaccurate configuration and repeated migration problems, which are low in migration efficiency.

Method used

Use the large language model to convert the objective functions and target syntax in the DAG file, generate a syntax tree, and traverse the syntax tree to obtain DAG parameters and Operator parameters, convert them into parameters adapted to the target scheduling system through the parameter mapping relationship table, and call the target scheduling system interface to create scheduling tasks, realizing automatic migration.

Benefits of technology

It improves the efficiency and convenience of file migration, reduces manual operation costs and error probability, and ensures compatibility and executability of migrated files in the target scheduling system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104589B_ABST
    Figure CN120104589B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a file migration method, apparatus, computer device, medium, and program product. It relates to the technical field of data processing. The method includes: obtaining a DAG file to be migrated, where the DAG file runs in a first scheduling system; using a large language model to convert the target function and target syntax in the DAG file to obtain the corresponding function and syntax; parsing the DAG file processed by the large language model conversion to generate a syntax tree; traversing the syntax tree to obtain the DAG parameters and Operator parameters in the DAG file; converting the DAG parameters and Operator parameters into a first parameter and a second parameter respectively that are adapted to the second scheduling system; calling a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to achieve migration. The technical solution of the embodiment of the present application can improve the efficiency and convenience of file migration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of data processing, and in particular, to a file migration method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] In the field of big data, task scheduling systems are usually used to manage and execute data processing, analysis, and other scheduled tasks. Currently, relatively popular systems include Azkaban (batch workflow task scheduler), Airflow (workflow platform), etc.

[0003] Different scheduling systems have their own characteristics and advantages. Enterprises and organizations may choose different scheduling systems according to the needs of different projects, which leads to an increasing demand for migrating files between different scheduling systems.

[0004] When it is necessary to migrate tasks from one scheduling system to another, it is often necessary to manually generate new tasks by referring to the configurations of the old tasks. This not only consumes a large amount of manpower and time, but also easily causes problems such as inaccurate migration task configurations and duplicate migrations, resulting in low migration efficiency.

[0005] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention

[0006] Embodiments of the present application provide a file migration method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the above technical problems.

[0007] One aspect of the embodiments of the present application provides a file migration method, the method including:

[0008] Obtain a DAG file to be migrated, where the DAG file runs in a first scheduling system;

[0009] Use a large language model to convert the target functions and target syntax in the DAG file to obtain functions and syntax corresponding to the target functions and the target syntax;

[0010] Parse the DAG file processed by the large language model to generate a syntax tree;

[0011] Traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file;

[0012] Convert the DAG parameters and the Operator parameters into first parameters and second parameters respectively adapted to a second scheduling system;

[0013] Call the preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and send the created scheduling task to the second scheduling system to achieve migration.

[0014] Optionally, traverse the syntax tree to obtain the DAG parameters and Operator parameters in the DAG file, including:

[0015] Traverse the syntax tree using the visitor pattern to obtain the DAG parameters and Operator parameters in the DAG file.

[0016] Optionally, the converting the DAG parameters and the Operator parameters into a first parameter and a second parameter respectively that are adapted to the second scheduling system includes:

[0017] Obtain the parameter mapping relationship table between the first scheduling system and the second scheduling system;

[0018] Convert the DAG parameters and the Operator parameters into a first parameter and a second parameter respectively that are adapted to the second scheduling system according to the parameter mapping relationship table.

[0019] Optionally, the calling the preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to achieve migration includes:

[0020] Call the preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter;

[0021] Verify the created scheduling task;

[0022] After the verification passes, send the created scheduling task to the second scheduling system to achieve migration.

[0023] Optionally, the verifying the created scheduling task includes:

[0024] Obtain the first execution code generated when the first scheduling system executes the scheduling task of the first preset type in the DAG file;

[0025] Obtain the second execution code generated when the second scheduling system executes the scheduling task corresponding to the scheduling task of the first preset type;

[0026] When the first execution code is the same as the second execution code, the verification passes.

[0027] Optionally, the verification of the created scheduling task includes:

[0028] Obtain the first execution code generated when the first scheduling system executes the scheduling task of the second preset type in the DAG file and the first parameter information used for executing the scheduling task of the second preset type in the DAG file;

[0029] Obtain the second execution code generated when the second scheduling system executes the scheduling task corresponding to the scheduling task of the second preset type and the second parameter information used for executing the scheduling task corresponding to the scheduling task of the second preset type;

[0030] If the first execution code is the same as the second execution code, and the first parameter information is the same as the second parameter information, the verification passes.

[0031] Another aspect of the embodiments of the present application provides a file migration device, and the device includes:

[0032] An acquisition module, configured to acquire a DAG file to be migrated, and the DAG file runs in a first scheduling system;

[0033] A first conversion module, configured to convert the target function and target syntax in the DAG file by using a large language model to obtain a function and syntax corresponding to the target function and the target syntax;

[0034] An analysis module, configured to analyze the DAG file processed by the large language model conversion to generate a syntax tree;

[0035] A traversal module, configured to traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file;

[0036] A second conversion module, configured to convert the DAG parameters and the Operator parameters into first parameters and second parameters adapted to a second scheduling system respectively;

[0037] A creation module, configured to call a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and send the created scheduling task to the second scheduling system to achieve migration.

[0038] Another aspect of the embodiments of the present application provides a computer device, including:

[0039] At least one processor; and

[0040] A memory communicatively connected to the at least one processor;

[0041] Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0042] Another aspect of the embodiments of the present application provides a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions are executed by a processor, the method as described above is implemented.

[0043] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0044] The embodiments of the present application adopting the above technical solutions may include the following advantages: First, obtain the DAG file to be migrated running in the first scheduling system, and then use the large language model to convert the target function and target syntax in the DAG file, so as to utilize the powerful language understanding and generation ability of the large language model to deeply understand the complex logic and semantic information in the DAG file, and accurately convert the target function and target syntax into functions and syntax convenient for the parsing tool to parse. Then, parse the DAG file processed by the large language model conversion to generate a syntax tree, and traverse the syntax tree to obtain DAG parameters and Operator parameters. Since the syntax tree can clearly display the structure and hierarchical relationship of the file, by traversing the syntax tree, the key parameters in the file can be systematically and comprehensively extracted, avoiding the risk of missing important information, and providing a solid foundation for subsequent parameter conversion and the creation of scheduling tasks. Then, convert the extracted DAG parameters and Operator parameters into first parameters and second parameters adapted to the second scheduling system respectively. This parameter adaptation strategy fully considers the differences and requirements of different scheduling systems, enabling the migrated file to perfectly conform to the specifications and standards of the second scheduling system, thereby improving the compatibility and executability of the file in the target scheduling system and avoiding migration failures or abnormal task executions caused by parameter mismatches. Finally, call the preset interface of the second scheduling system, create corresponding scheduling tasks based on the converted parameters, and send the created scheduling tasks to the second scheduling system to achieve automatic migration of the file. This entire migration process requires little manual intervention, greatly reducing the manual operation cost and error probability, improving the efficiency and convenience of file migration, and enabling tasks to run quickly in the new scheduling system. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings exemplarily illustrate embodiments and form part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0046] Figure 1 Schematically shows an operating environment diagram of a file migration method according to Embodiment 1 of the present application;

[0047] Figure 2 Schematically shows a flowchart of a file migration method according to Embodiment 1 of the present application;

[0048] Figure 3 Schematically shows a refined flowchart of the steps of converting the DAG parameters and the Operator parameters into first parameters and second parameters respectively adapted to a second scheduling system;

[0049] Figure 4 Schematically shows a refined flowchart of the steps of calling a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to achieve migration;

[0050] Figure 5 Schematically shows a refined flowchart of the steps of verifying the created scheduling task;

[0051] Figure 6 Schematically shows a refined flowchart of the steps of verifying the created scheduling task;

[0052] Figure 7 Schematically shows a block diagram of a file migration device according to Embodiment 2 of the present application; and

[0053] Figure 8 Schematically shows a schematic diagram of the hardware architecture of a computer device according to Embodiment 3 of the present application. Detailed implementation manners

[0054] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0055] It should be noted that in the embodiments of the present application, the descriptions involving "first", "second", etc. are only for descriptive purposes, and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0056] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and to distinguish each step. Therefore, it should not be construed as a limitation to the present application.

[0057] First, the following provides the term explanations involved in the present application:

[0058] DAG file: Generally known as the "task flow definition file" or "workflow configuration file", it is used to define the task flow, configure the scheduling rules, manage task dependencies, etc., and is the core basis for the scheduling system to execute tasks. It not only describes the dependency relationships and execution order between tasks, but also provides rich configuration options to meet complex scheduling requirements. Through the DAG file, the scheduling system can efficiently manage the execution, resource allocation, and fault recovery of tasks to ensure the smooth progress of tasks.

[0059] To facilitate the understanding of the technical solutions provided in the embodiments of the present application by those skilled in the art, the related technologies are described below:

[0060] Different scheduling systems have their own characteristics and advantages. Enterprises and organizations may choose different scheduling systems according to the requirements of different projects, which has led to an increasing demand for migrating files between different scheduling systems.

[0061] When it is necessary to migrate tasks from one scheduling system to another, it is often necessary to manually generate new tasks by referring to the configurations of the old tasks. This not only consumes a large amount of manpower and time, but also easily leads to problems such as inaccurate configuration of migrated tasks and duplicate migrations, resulting in low migration efficiency.

[0062] To this end, the embodiments of the present application provide a file migration technical solution. In this technical solution, first, obtain the DAG file to be migrated running in the first scheduling system, and then use the large language model to convert the target functions and target syntax in the DAG file, so as to utilize the powerful language understanding and generation capabilities of the large language model to deeply understand the complex logic and semantic information in the DAG file, and accurately convert the target functions and target syntax into functions and syntax that are easy for the parsing tool to parse. Next, parse the DAG file processed by the large language model conversion to generate a syntax tree, and traverse the syntax tree to obtain DAG parameters and Operator parameters. Since the syntax tree can clearly show the structure and hierarchical relationship of the file, by traversing the syntax tree, the key parameters in the file can be systematically and comprehensively extracted, avoiding the risk of missing important information, and providing a solid foundation for subsequent parameter conversion and the creation of scheduling tasks. Then, convert the extracted DAG parameters and Operator parameters into the first parameter and the second parameter that are adapted to the second scheduling system respectively. This parameter adaptation strategy fully considers the differences and requirements of different scheduling systems, enabling the migrated file to perfectly conform to the specifications and standards of the second scheduling system, thereby improving the compatibility and executability of the file in the target scheduling system and avoiding migration failures or task execution anomalies caused by parameter mismatches. Finally, call the preset interface of the second scheduling system, create the corresponding scheduling task based on the converted parameters, and send the created scheduling task to the second scheduling system to achieve automatic file migration. This entire migration process does not require a large amount of manual intervention, greatly reducing the manual operation cost and error probability, improving the efficiency and convenience of file migration, and enabling the task to run quickly in the new scheduling system. See the following for details.

[0063] Finally, for ease of understanding, an exemplary operating environment is provided below.

[0064] As Figure 1 shown, the environmental schematic diagram includes a service platform 2, a network 4, and a client 6, where:

[0065] The service platform 2 can be composed of a single or multiple computing devices. The multiple computing devices can include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing device can load the virtual machine based on a virtual image and / or other data that defines a specific software (e.g., operating system, dedicated application, server) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices. A hypervisor can be implemented to manage the use of different virtual machines on the same computing device.

[0066] The service platform 2 can be configured to communicate with the client 6 etc. via the network 4. The network 4 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links, combinations thereof, etc., or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.

[0067] The service platform 2 can provide services such as storage, reading, writing, querying, deleting, etc., such as providing a file migration service for the client.

[0068] The client 6 can be an electronic device running an operating system such as Windows, Android™, or iOS, such as a smartphone, tablet device, laptop computer, virtual reality device, gaming device, set-top box, in-vehicle terminal, smart TV. Based on the above operating systems, various application programs can be run, such as a file migration program.

[0069] The client 6 can provide / configure a user access page for manipulating the service platform 2 or uploading objects, etc.

[0070] It should be noted that the above devices are exemplary, and in different scenarios or according to different requirements, the number and types of devices can be adjusted.

[0071] The technical solutions of the present application will be introduced below through multiple embodiments. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.

[0072] Embodiment 1

[0073] Figure 2 A flowchart of the file migration method according to Embodiment 1 of the present application is schematically shown.

[0074] As Figure 2 shown, the file migration method may include steps S200 to S210, where:

[0075] Step S200, obtain the DAG file to be migrated, and the DAG file runs in the first scheduling system.

[0076] Step S202, use a large language model to convert the target function and target syntax in the DAG file to obtain the function and syntax corresponding to the target function and the target syntax.

[0077] Step S204, parse the DAG file processed by the large language model conversion to generate a syntax tree.

[0078] Step S206: Traverse the syntax tree to obtain the DAG parameters and Operator parameters in the DAG file.

[0079] Step S208: Convert the DAG parameters and the Operator parameters into a first parameter and a second parameter respectively that are adapted to the second scheduling system.

[0080] Step S210: Call the preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and send the created scheduling task to the second scheduling system to achieve migration.

[0081] For the file migration method provided in this embodiment, first, obtain the DAG file to be migrated running in the first scheduling system, and then use the large language model to convert the target function and target syntax in the DAG file, so as to utilize the powerful language understanding and generation ability of the large language model to deeply understand the complex logic and semantic information in the DAG file, and accurately convert the target function and target syntax into functions and syntax that are easy for the parsing tool to parse. Then, parse the DAG file after being processed by the large language model conversion to generate a syntax tree, and traverse the syntax tree to obtain the DAG parameters and Operator parameters. Since the syntax tree can clearly show the structure and hierarchical relationship of the file, by traversing the syntax tree, the key parameters in the file can be systematically and comprehensively extracted, avoiding the risk of missing important information, and providing a solid foundation for subsequent parameter conversion and scheduling task creation. Then, convert the extracted DAG parameters and Operator parameters into a first parameter and a second parameter respectively that are adapted to the second scheduling system. This parameter adaptation strategy fully considers the differences and requirements of different scheduling systems, making the migrated file perfectly conform to the specifications and standards of the second scheduling system, thereby improving the compatibility and executability of the file in the target scheduling system and avoiding migration failures or task execution anomalies caused by parameter mismatches. Finally, call the preset interface of the second scheduling system, create a corresponding scheduling task based on the converted parameters, and send the created scheduling task to the second scheduling system to achieve automatic file migration. This entire migration process requires little manual intervention, greatly reducing the manual operation cost and error probability, improving the efficiency and convenience of file migration, and enabling the task to run quickly in the new scheduling system.

[0082] The following will combine Figure 2 to elaborate in detail on each step in steps S200 to S210 and other optional steps.

[0083] Step S200 , obtain the DAG file to be migrated, and the DAG file runs in the first scheduling system.

[0084] Among them, the first scheduling system is a task scheduling system that uses a distributed visualization DAG (Directed Acyclic Graph) workflow. The first scheduling systems include Oozie (task scheduling framework), Azkaban (batch workflow task scheduler), Airflow (workflow platform), etc.

[0085] In the first scheduling system, the DAG file is used to define the task process, configure scheduling rules, manage task dependencies, etc., and is the core basis for the first scheduling system to execute tasks.

[0086] In one embodiment, when it is necessary to migrate the DAG file, a migration request can be sent to the first scheduling system. When the first scheduling system receives the migration request, it can return the DAG file running in the first scheduling system to the initiator of the migration request, so as to obtain the DAG file to be migrated.

[0087] In another embodiment, the DAG file running in the first scheduling system can also be stored in a preset location in advance. When it is necessary to migrate the DAG file, the DAG file can be obtained from this preset location.

[0088] Step S202 , use a large language model to convert the target functions and target syntax in the DAG file to obtain functions and syntax corresponding to the target functions and the target syntax.

[0089] The large language model can convert the target functions and target syntax in the DAG file into functions and syntax that can be easily parsed by a parsing tool later.

[0090] Among them, the large language model can be obtained by training and learning existing open-source large language models by collecting a large number of code examples and corresponding DAG files. The large language model can accurately identify the target functions and target syntax in the DAG file, and will convert the target functions and target syntax into corresponding functions and syntax that can be parsed by a parsing tool.

[0091] Among them, the target function is a function that cannot be parsed by a parsing tool, and the target function can be determined in advance according to the parsing function of the parsing tool. For example, if the parsing tool cannot parse function a and function b, then function a and function b can be used as target functions.

[0092] Similarly, the target syntax is a syntax that cannot be parsed by a parsing tool, and the target syntax can also be determined in advance according to the parsing function of the parsing tool. For example, if the parsing tool cannot parse syntax A and syntax B, then syntax A and syntax B can be used as target functions.

[0093] It should be noted that when the large language model performs conversion processing on the DAG file, in addition to converting the target function and target syntax, it will not perform conversion processing on other content, so it will remain unchanged.

[0094] Step S204 Parse the DAG file that has undergone the conversion processing by the large language model to generate a syntax tree.

[0095] In one embodiment, an open-source LibCST parsing tool can be used to parse the DAG file that has undergone the conversion processing by the large language model to generate a syntax tree.

[0096] Among them, the LibCST parsing tool is a Concrete SyntaxTree (CST) library for parsing and serializing Python code, which can parse Python source code into a CST tree and can retain all format details.

[0097] In another embodiment, other parsing tools can also be used to parse the DAG file that has undergone the conversion processing by the large language model to generate a syntax tree. For example, the ANTLR parsing tool can be used.

[0098] Among them, the ANTLR parsing tool is a powerful parser generator tool that can generate parser code according to a grammar file to parse the DAG file and generate an Abstract Syntax Tree (AST).

[0099] Step S206 Traverse the syntax tree to obtain the DAG parameters and Operator parameters in the DAG file.

[0100] The DAG parameters are used to control the scheduling, running, and management of the DAG as a whole.

[0101] The Operator parameters are used to define the behaviors and characteristics of specific tasks.

[0102] Among them, the DAG parameters can include parameters such as dag_id, schedule_interval, start_date, end_date, default_args, dagrun_timeout, tags, params, catchup, etc.

[0103] The dag_id is used to uniquely identify a DAG, and in Airflow, a specific DAG can be found and managed through this ID.

[0104] schedule_interval, which is used to determine the running frequency of the DAG and controls the periodic execution of the workflow.

[0105] start_date, which is used to define the start time of the DAG. The DAG will only be scheduled for execution when the current time is later than this time.

[0106] end_date, which is used to limit the running time period of the DAG. If set, the DAG will no longer be scheduled when the current time exceeds this date.

[0107] default_args, which is used to provide default parameters for all tasks in the DAG. By using this parameter, the task definition is simplified, and the repeated setting of common parameters is avoided.

[0108] dagrun_timeout, which is used to limit the maximum running time of the DAG, ensuring that DAGs that have not been completed for a long time are terminated in a timely manner to avoid resource waste.

[0109] tags, which are used to classify and label the DAG, facilitating the filtering and searching of DAGs in the Airflow Web UI.

[0110] params, which are used to define parameters at the DAG level. These parameters can be shared and accessed by all tasks within the DAG, facilitating the transfer and use of common data.

[0111] catchup, which is used to determine whether the DAG will backfill the previously unexecuted task instances, affecting the historical task execution of the DAG.

[0112] Among them, the Operator parameters can include parameters such as task_id, dag, trigger_rule, depends_on_past, email, email_on_retry, execution_timeout, priority_weight, queue, owner, retries, retry_delay, etc.

[0113] task_id, which serves as the unique identifier of a task in the DAG, facilitating the identification, management, and scheduling of specific tasks.

[0114] dag, which is used to specify the DAG to which the task belongs, ensuring that the task is executed in the correct DAG.

[0115] trigger_rule, which is used to define the trigger condition of the task, controlling under what circumstances the task starts to execute and providing flexible task execution control.

[0116] depends_on_past, which is used to determine whether a task depends on the successful completion of the previous task instance with the same name and is used for task dependency management.

[0117] emai, which is used to specify the email address to send notifications when the task fails, facilitating timely understanding of the situation where the task execution fails.

[0118] email_on_retry, which is used to determine whether to send an email notification when the task is retried. Through this parameter, the task notification mechanism can be further refined.

[0119] execution_timeout, which is used to limit the maximum execution time of the task, ensuring that the task does not run indefinitely and guaranteeing the timeliness of task execution.

[0120] priority_weight, which is used to define the priority weight of the task. Through this parameter, the priority of the task in the queue can be affected, determining the execution order of the task.

[0121] queue, which is used to specify the queue name where the task runs. It can cooperate with the executor to control the task to run in a specific queue, realizing the reasonable allocation of resources.

[0122] owner, which is used to record the person in charge of the task, facilitating the tracking of the ownership and management of the task.

[0123] retries, which is used to set the number of retries after the task fails. Through this parameter, the reliability of the task can be enhanced, ensuring that the task has the opportunity to be re-executed when encountering temporary errors.

[0124] retry_delay, which is used to define the interval time between two retries, avoiding the task from being retried too frequently and giving the system a certain buffer time.

[0125] In this embodiment, during the process of traversing the syntax tree, the definition nodes of the DAG and Operator are found, and then the DAG parameters are extracted from the DAG node and the Operator parameters are extracted from the Operator node.

[0126] In an alternative embodiment, the traversing the syntax tree to obtain the DAG parameters and Operator parameters in the DAG file may include:

[0127] Traversing the syntax tree using the visitor pattern to obtain the DAG parameters and Operator parameters in the DAG file.

[0128] Among them, the visitor pattern is a behavioral design pattern that can traverse a complex object structure during operations without changing the classes of the elements.

[0129] In this embodiment, by traversing the syntax tree using the visitor pattern, the definition nodes of DAG and Operator can be more accurately identified from the syntax tree, and then the DAG parameters and Operator parameters can be extracted.

[0130] In other embodiments, other patterns can also be used to traverse the syntax tree. For example, the listener pattern can be used to traverse the syntax tree.

[0131] Step S208 , the DAG parameters and the Operator parameters are respectively converted into a first parameter and a second parameter adapted to the second scheduling system.

[0132] Since different scheduling systems may have different task definition methods and parameter requirements, the extracted parameters need to be adapted and converted. Therefore, in this embodiment, after obtaining the DAG parameters and the Operator parameters, the DAG parameters need to be converted into a first parameter that can be recognized and used by the second scheduling system, and at the same time, the Operator parameters are converted into a second parameter that can be recognized and used by the second scheduling system.

[0133] In practical applications, parameter conversion can be performed in various ways. The following provides an exemplary way.

[0134] In an alternative embodiment, refer to Figure 3 , converting the DAG parameters and the Operator parameters into a first parameter and a second parameter adapted to the second scheduling system respectively may include:

[0135] Step S300, obtain the parameter mapping relationship table between the first scheduling system and the second scheduling system.

[0136] Step S302, convert the DAG parameters and the Operator parameters into a first parameter and a second parameter adapted to the second scheduling system respectively according to the parameter mapping relationship table.

[0137] In one embodiment, the mapping relationships of various parameters between the first scheduling system and the second scheduling system can be predefined. For example, the parameter corresponding to DAG parameter a is parameter A, the parameter corresponding to DAG parameter b is parameter B. The parameter corresponding to Operator parameter c is parameter C. The parameter corresponding to Operator parameter d is parameter D.

[0138] In this embodiment, after obtaining the DAG parameters and the Operator parameters, the parameter mapping relationship table predefined between the first scheduling system and the second scheduling system can be queried, so as to find the first parameter corresponding to the DAG parameters and the second parameter corresponding to the Operator parameters.

[0139] In this embodiment, based on the parameter mapping relationship table, the conversion between parameters can be quickly and accurately realized.

[0140] In another implementation manner, the conversion between parameters can also be realized in other ways. For example, the DAG parameters and the Operator parameters can be first converted into parameters in a general format, and then the parameters in the general format are converted into the first parameter and the second parameter.

[0141] Step S210 , call the preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and send the created scheduling task to the second scheduling system to achieve migration.

[0142] The preset interface is an interface for creating a scheduling task based on the first parameter and the second parameter.

[0143] In one implementation manner, when creating a scheduling task, the preset interface first reads the first parameter and the second parameter, and then sets the parameters of the task according to the requirements of the second scheduling system based on the first parameter and the second parameter, so as to realize the creation of the scheduling task.

[0144] In an alternative embodiment, refer to Figure 4 , the calling the preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to achieve migration includes:

[0145] Step S400, call the preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter.

[0146] Step S402, verify the created scheduling task.

[0147] Step S404, after the verification is passed, send the created scheduling task to the second scheduling system to achieve migration.

[0148] In this embodiment, after creating the scheduling task, to ensure that the created scheduling task is problem-free and can be used normally for the direct use of the second scheduling system, the created scheduling task can be first sent to a preset verification terminal so that the verification terminal verifies the created scheduling task.

[0149] After the verification passes, the created scheduling task is then sent to the second scheduling system to achieve migration.

[0150] When the verification fails, corresponding prompt information can be generated so that the staff can make further modifications and verifications. Until the verification passes, the created scheduling task can be sent to the second scheduling system to achieve migration.

[0151] Of course, it can be understood that in specific implementation, the created scheduling task can be directly sent to the second scheduling system, and the second scheduling system verifies it. After the verification passes, it is released for use in the production environment.

[0152] By the above method, by pre-verifying the created scheduling task to ensure that the created scheduling task is problem-free, it can be released for use in the production environment, thus facilitating the direct use of the second scheduling system.

[0153] In practical applications, the created scheduling task can be verified in various ways. The following provides an exemplary method.

[0154] In an alternative embodiment, refer to Figure 5 , the verification of the created scheduling task includes:

[0155] Step S500, obtain the first execution code generated when the first scheduling system executes the scheduling task of the first preset type in the DAG file.

[0156] Step S502, obtain the second execution code generated when the second scheduling system executes the scheduling task corresponding to the scheduling task of the first preset type.

[0157] Step S504, when the first execution code is the same as the second execution code, the verification passes.

[0158] The first preset type is a task type set in advance according to the actual situation. For example, the scheduling task of the first preset type is an sql task.

[0159] In this embodiment, when a scheduling task of a first preset type is included in the DAG file, one or more scheduling tasks of the first preset type can be executed by the first scheduling system, and then, the first execution code generated when the first scheduling system executes one or more scheduling tasks of the first preset type is obtained. For example, an SQL task is executed regularly at 2 o'clock every day, and its execution code is: select a from table where log_date = 20250422.

[0160] Meanwhile, one or more scheduling tasks of the first preset type corresponding to the preset scheduling task can be executed by the second scheduling system, and then, the second execution code generated when the second scheduling system executes one or more scheduling tasks corresponding to the preset scheduling task is obtained.

[0161] After obtaining the first execution code and the second execution code, it can be compared whether the two are the same. If the two are the same, it indicates that the verification passes. If the two are different, it indicates that the verification fails.

[0162] It should be noted that the scheduling tasks corresponding to the preset scheduling task mentioned above refer to the scheduling tasks under the same scheduling time.

[0163] In this embodiment, verification is achieved by comparing codes. Compared with the existing technology that uses the double-run method for verification, it can reduce the consumption of resources and improve the verification efficiency.

[0164] Another exemplary method is provided below.

[0165] In an alternative embodiment, refer to Figure 6 , the verification of the created scheduling task includes:

[0166] Step S600, obtain the first execution code generated when the first scheduling system executes the scheduling task of the second preset type in the DAG file and the first parameter information used for executing the scheduling task of the second preset type in the DAG file.

[0167] Step S602, obtain the second execution code generated when the second scheduling system executes the scheduling task corresponding to the scheduling task of the second preset type and the second parameter information used for executing the scheduling task corresponding to the scheduling task of the second preset type.

[0168] Step S604, if the first execution code is the same as the second execution code, and the first parameter information is the same as the second parameter information, the verification passes.

[0169] The second preset type is a task type set in advance according to the actual situation. For example, the scheduling tasks of the second preset type are generally non-sql tasks. For example, they are tasks for executing Shell commands or scripts.

[0170] In this embodiment, when the DAG file contains scheduling tasks of the second preset type, one or more scheduling tasks of the second preset type can be executed by the second scheduling system. Then, the first execution code generated when the first scheduling system executes one or more scheduling tasks of the second preset type and the first parameter information used for executing one or more scheduling tasks of the second preset type in the DAG file are obtained.

[0171] Meanwhile, one or more scheduling tasks corresponding to the scheduling tasks of the second preset type can be executed by the second scheduling system. Then, the second execution code generated when the second scheduling system executes one or more scheduling tasks corresponding to the scheduling tasks of the second preset type, and the second parameter information used for executing one or more scheduling tasks corresponding to the scheduling tasks of the second preset type by the second scheduling system are obtained.

[0172] After obtaining the first execution code and the second execution code, it can be compared whether they are the same. If they are the same, then it can be further compared whether the first parameter information and the second parameter information are the same. If the first parameter information and the second parameter information are also the same, it indicates that the verification passes. If they are different, it indicates that the verification fails.

[0173] In this embodiment, verification is achieved by comparing codes and parameters. Compared with the prior art where double-running is used for verification, it can reduce the consumption of resources and improve the verification efficiency.

[0174] Embodiment 2

[0175] Figure 7 A block diagram of a file migration device 700 according to Embodiment 2 of the present application is schematically shown. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 7 shown, the device 700 may include: an acquisition module 710, a first conversion module 720, a parsing module 730, a traversal module 740, a second conversion module 750, and a creation module 760, where:

[0176] An acquisition module 710, configured to acquire a DAG file to be migrated, where the DAG file runs in a first scheduling system;

[0177] A first conversion module 720, configured to use a large language model to convert the target functions and target syntax in the DAG file to obtain functions and syntax corresponding to the target functions and the target syntax;

[0178] A parsing module 730, configured to parse the DAG file processed by the large language model conversion to generate a syntax tree;

[0179] A traversal module 740, configured to traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file;

[0180] A second conversion module 750, configured to convert the DAG parameters and the Operator parameters into a first parameter and a second parameter respectively adapted to a second scheduling system;

[0181] A creation module 760, configured to call a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and send the created scheduling task to the second scheduling system to implement migration.

[0182] As an optional embodiment, the traversing the syntax tree to obtain DAG parameters and Operator parameters in the DAG file includes:

[0183] Traversing the syntax tree using the visitor pattern to obtain DAG parameters and Operator parameters in the DAG file.

[0184] As an optional embodiment, the converting the DAG parameters and the Operator parameters into a first parameter and a second parameter respectively adapted to a second scheduling system includes:

[0185] Obtaining a parameter mapping relationship table between the first scheduling system and the second scheduling system;

[0186] Converting the DAG parameters and the Operator parameters into a first parameter and a second parameter respectively adapted to a second scheduling system according to the parameter mapping relationship table.

[0187] As an optional embodiment, the calling a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to implement migration includes:

[0188] Create a corresponding scheduling task based on the first parameter and the second parameter by invoking a preset interface of the second scheduling system;

[0189] Verify the created scheduling task;

[0190] After the verification passes, send the created scheduling task to the second scheduling system to achieve migration.

[0191] As an optional embodiment, the verification of the created scheduling task includes:

[0192] Obtain the first execution code generated when the first scheduling system executes a scheduling task of a first preset type in the DAG file;

[0193] Obtain the second execution code generated when the second scheduling system executes a scheduling task corresponding to the scheduling task of the first preset type;

[0194] If the first execution code is the same as the second execution code, the verification passes.

[0195] As an optional embodiment, the verification of the created scheduling task includes:

[0196] Obtain the first execution code generated when the first scheduling system executes a scheduling task of a second preset type in the DAG file and the first parameter information used for executing the scheduling task of the second preset type in the DAG file;

[0197] Obtain the second execution code generated when the second scheduling system executes a scheduling task corresponding to the scheduling task of the second preset type and the second parameter information used for executing the scheduling task corresponding to the scheduling task of the second preset type;

[0198] If the first execution code is the same as the second execution code, and the first parameter information is the same as the second parameter information, the verification passes.

[0199] Embodiment III

[0200] Figure 8 Schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing the file migration method according to Embodiment III of the present application. In some embodiments, the computer device 10000 may be a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc. AsFigure 8 As shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can be communicatively linked to each other through a system bus. Among them:

[0201] The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 10010 can be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 can also be an external storage device of the computer device 10000, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 can also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the file migration method. In addition, the memory 10010 can also be used to temporarily store various data that have been output or will be output.

[0202] In some embodiments, the processor 10020 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0203] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, the Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0204] It should be noted that Figure 8 only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be alternatively implemented.

[0205] In this embodiment, the file migration method stored in the memory 10010 may also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.

[0206] Embodiment 4

[0207] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the file migration method in the embodiments are implemented.

[0208] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the file migration method in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various data that have been output or will be output.

[0209] Embodiment 5

[0210] The embodiment of the present application further provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0211] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general-purpose computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device, so that they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0212] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A file migration method, characterized in that, The method includes: Obtain a DAG file to be migrated, where the DAG file runs in a first scheduling system; Use a large language model to convert the target function and target syntax in the DAG file to obtain a function and syntax corresponding to the target function and the target syntax; Parse the DAG file that has been processed by the large language model conversion to generate a syntax tree; Traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file; Convert the DAG parameters and the Operator parameters into a first parameter and a second parameter respectively that are adapted to a second scheduling system; Call a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and send the created scheduling task to the second scheduling system to achieve migration.

2. The method according to claim 1, wherein The traversing the syntax tree to obtain DAG parameters and Operator parameters in the DAG file includes: Use the visitor pattern to traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file.

3. The method according to claim 1, wherein The converting the DAG parameters and the Operator parameters into a first parameter and a second parameter respectively that are adapted to a second scheduling system includes: Obtain a parameter mapping relationship table between the first scheduling system and the second scheduling system; Convert the DAG parameters and the Operator parameters into a first parameter and a second parameter respectively that are adapted to the second scheduling system according to the parameter mapping relationship table.

4. The method according to any one of claims 1 to 3, characterized in that, The calling a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to achieve migration includes: Call a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter; Verify the created scheduling task; After the verification passes, send the created scheduling task to the second scheduling system to achieve migration.

5. The method according to claim 4, characterized in that, The verifying the created scheduling task includes: Obtain the first execution code generated when the first scheduling system executes the first preset type of scheduling task in the DAG file; Obtain the second execution code generated when the second scheduling system executes a scheduling task corresponding to the first preset type of scheduling task; If the first execution code is the same as the second execution code, the verification passes.

6. The method according to claim 4, wherein The verifying the created scheduling task includes: Obtain the first execution code generated when the first scheduling system executes the second preset type of scheduling task in the DAG file and the first parameter information used for executing the second preset type of scheduling task in the DAG file; Obtain the second execution code generated when the second scheduling system executes a scheduling task corresponding to the second preset type of scheduling task and the second parameter information used for executing the scheduling task corresponding to the second preset type of scheduling task; If the first execution code is the same as the second execution code and the first parameter information is the same as the second parameter information, the verification passes.

7. A file migration device, characterized in that, The device includes: an acquisition module configured to acquire a DAG file to be migrated, where the DAG file runs in a first scheduling system; a first conversion module configured to use a large language model to convert a target function and target syntax in the DAG file to obtain a function and syntax corresponding to the target function and the target syntax; a parsing module configured to parse the DAG file that has been processed by the large language model conversion to generate a syntax tree; a traversal module configured to traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file; a second conversion module configured to convert the DAG parameters and the Operator parameters into first parameters and second parameters respectively adapted to a second scheduling system; a creation module configured to call a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameters and the second parameters, and send the created scheduling task to the second scheduling system to achieve migration.

8. A computer device, characterized in that, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Deep learning scheduling configuration system and method

    CN113127182A

  • Code migration method and device for front-end framework upgrading, equipment and medium

    CN118409793A