File migration method and device

By using a large language model to convert and parse DAG files and automatically extract and adapt parameters, task migration from one scheduling system to another scheduling system is realized, solving the problem of low manual migration efficiency and improving migration efficiency and convenience.

CN120104589AActive Publication Date: 2025-06-06SHANGHAI HODE INFORMATION TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510578015.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

When migrating tasks between different scheduling systems, new tasks need to be generated manually compared with the old task configuration, which consumes a lot of manpower and time, and is prone to inaccurate migration configuration and repeated migration problems, resulting in low migration efficiency.

Method used

By obtaining the DAG file to be migrated, using a large language model to convert the objective function and the target syntax, generate a syntax tree and traversal to extract DAG parameters and Operator parameters, then convert these parameters into parameters that adapt to the second scheduling system, call the interface of the second scheduling system to create and send scheduling tasks, realizing automatic migration.

Benefits of technology

Without a large amount of manual intervention, it greatly reduces the cost of manual operation and the probability of error, improves the efficiency and convenience of file migration, and can quickly run tasks in the new scheduling system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104589A_ABST
    Figure CN120104589A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a file migration method and device, computer equipment, a medium and a program product. Relates to the technical field of data processing. The method comprises the steps of obtaining a to-be-migrated DAG file, wherein the DAG file runs in a first scheduling system; the target function and the target grammar in the DAG file are converted through a large language model, and a function and a grammar corresponding to the target function and the target grammar are obtained; analyzing the DAG file subjected to large language model conversion processing to generate a syntax tree; traversing the syntax tree to obtain a DAG parameter and an Operator parameter in the DAG file; respectively converting the DAG parameter and the Operator parameter into a first parameter and a second parameter which are adaptive to a second scheduling system; and calling a preset interface of a second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to realize migration. According to the technical scheme, the file migration efficiency and convenience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of data processing technology, and in particular, to a file migration method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] In the field of big data, task scheduling systems are often used to manage and execute data processing, analysis, and other scheduled tasks. Currently, popular systems include Azkaban (batch workflow task scheduler) and Airflow (workflow platform).

[0003] Different scheduling systems have their own characteristics and advantages. Enterprises and organizations may choose different scheduling systems according to the needs of different projects, which leads to an increasing demand for migrating files between different scheduling systems.

[0004] When tasks need to be migrated from one scheduling system to another, it is often necessary to manually generate new tasks based on the configuration of the old tasks. This not only consumes a lot of manpower and time, but also easily leads to inaccurate configuration of the migration tasks and repeated migration, resulting in low migration efficiency.

[0005] It should be noted that the above content is not necessarily prior art, nor is it intended to limit the scope of patent protection of this application. Summary of the invention

[0006] Embodiments of the present application provide a file migration method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the technical problems raised above.

[0007] One aspect of an embodiment of the present application provides a file migration method, the method comprising: Obtaining a DAG file to be migrated, where the DAG file runs in the first scheduling system; Using a large language model to convert the target function and the target grammar in the DAG file to obtain a function and a grammar corresponding to the target function and the target grammar; Parsing the DAG file converted by the large language model to generate a syntax tree; Traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file; Convert the DAG parameter and the Operator parameter into a first parameter and a second parameter respectively adapted to the second scheduling system; The preset interface of the second scheduling system is called to create a corresponding scheduling task based on the first parameter and the second parameter, and the created scheduling task is sent to the second scheduling system to achieve migration.

[0008] Optionally, the syntax tree is traversed to obtain DAG parameters and Operator parameters in the DAG file, including: The syntax tree is traversed using the visitor pattern to obtain DAG parameters and Operator parameters in the DAG file.

[0009] Optionally, converting the DAG parameter and the Operator parameter into a first parameter and a second parameter respectively adapted to the second scheduling system includes: Obtaining a parameter mapping relationship table between the first scheduling system and the second scheduling system; The DAG parameters and the Operator parameters are respectively converted into first parameters and second parameters that are compatible with the second scheduling system according to the parameter mapping relationship table.

[0010] Optionally, the calling of a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to achieve migration includes: Calling a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter; Verify the created scheduling tasks; After the verification is passed, the created scheduling task is sent to the second scheduling system to achieve migration.

[0011] Optionally, the verifying the created scheduling task includes: Obtaining a first execution code generated when the first scheduling system executes a scheduling task of a first preset type in the DAG file; Acquire a second execution code generated by the second scheduling system executing a scheduling task corresponding to the scheduling task of the first preset type; When the first execution code is identical to the second execution code, the verification succeeds.

[0012] Optionally, the verifying the created scheduling task includes: Acquire a first execution code generated when the first scheduling system executes a scheduling task of a second preset type in the DAG file and first parameter information used to execute the scheduling task of the second preset type in the DAG file; Acquire a second execution code generated by the second scheduling system when executing a scheduling task corresponding to the second preset type of scheduling task and second parameter information used when executing the scheduling task corresponding to the second preset type of scheduling task; If the first execution code is the same as the second execution code, and the first parameter information is the same as the second parameter information, the verification is successful.

[0013] Another aspect of an embodiment of the present application provides a file migration device, the device comprising: An acquisition module, used to acquire a DAG file to be migrated, where the DAG file runs in the first scheduling system; A first conversion module, configured to convert the target function and the target grammar in the DAG file by using a large language model to obtain a function and a grammar corresponding to the target function and the target grammar; A parsing module, used to parse the DAG file converted by the large language model to generate a syntax tree; A traversal module, used to traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file; A second conversion module, used to convert the DAG parameter and the Operator parameter into a first parameter and a second parameter respectively adapted to a second scheduling system; A creation module is used to call a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and send the created scheduling task to the second scheduling system to achieve migration.

[0014] Another aspect of an embodiment of the present application provides a computer device, including: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein: the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0015] Another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method described above is implemented.

[0016] Another aspect of an embodiment of the present application provides a computer program product, including a computer program, which implements the method described above when executed by a processor.

[0017] The embodiment of the present application adopts the above technical solution, which may include the following advantages: First, obtain the DAG file to be migrated running in the first scheduling system, and then use the large language model to convert the target function and target syntax in the DAG file, so as to use the powerful language understanding and generation capabilities of the large language model to deeply understand the complex logic and semantic information in the DAG file, and accurately convert the target function and target syntax into functions and syntax that are easy to parse with the parsing tool. Then, parse the DAG file after the large language model conversion to generate a syntax tree, and traverse the syntax tree to obtain DAG parameters and Operator parameters. Since the syntax tree can clearly display the structure and hierarchical relationship of the file, the key parameters in the file can be systematically and comprehensively extracted by traversing the syntax tree, avoiding the risk of missing important information, and providing a solid foundation for subsequent parameter conversion and scheduling task creation. Then, the extracted DAG parameters and Operator parameters are converted into first parameters and second parameters that are compatible with the second scheduling system, respectively. This parameter adaptation strategy fully considers the differences and requirements of different scheduling systems, so that the migrated files can perfectly meet the specifications and standards of the second scheduling system, thereby improving the compatibility and executability of the files in the target scheduling system, and avoiding migration failures or task execution exceptions caused by parameter mismatches. Finally, the preset interface of the second scheduling system is called, the corresponding scheduling tasks are created based on the converted parameters, and the created scheduling tasks are sent to the second scheduling system to realize automatic file migration. This entire migration process does not require a lot of manual intervention, which greatly reduces the cost of manual operation and the probability of error, improves the efficiency and convenience of file migration, and can quickly run tasks in the new scheduling system. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0019] Figure 1 The following schematically shows an operating environment diagram of the file migration method according to the first embodiment of the present application; Figure 2 A flowchart of a file migration method according to Embodiment 1 of the present application is schematically shown; Figure 3 A detailed flowchart of steps for converting the DAG parameters and the Operator parameters into first parameters and second parameters respectively adapted to the second scheduling system is schematically shown; Figure 4A flowchart schematically shows the detailed steps of calling the preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to achieve migration; Figure 5 A detailed flowchart of the steps for verifying the created scheduling task is schematically shown; Figure 6 A detailed flowchart of the steps for verifying the created scheduling task is schematically shown; Figure 7 A block diagram schematically shows a file migration device according to Embodiment 2 of the present application; and Figure 8 The hardware architecture diagram of the computer device according to the third embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0021] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0022] In the description of the present application, it should be understood that the numerical labels before the steps do not indicate the order in which the steps are executed, but are only used to facilitate the description of the present application and to distinguish each step, and therefore should not be understood as a limitation on the present application.

[0023] First, the following terms are explained: DAG file: generally called "task flow definition file" or "workflow configuration file", it is used to define task flow, configure scheduling rules, manage task dependencies, etc. It is the core basis for the scheduling system to execute tasks. It not only describes the dependencies and execution order between tasks, but also provides a wealth of configuration options to meet complex scheduling requirements. Through DAG files, the scheduling system can efficiently manage task execution, resource allocation, and fault recovery to ensure the smooth progress of tasks.

[0024] To facilitate those skilled in the art to understand the technical solutions provided in the embodiments of the present application, the relevant technologies are described below: Different scheduling systems have their own characteristics and advantages. Enterprises and organizations may choose different scheduling systems according to the needs of different projects, which leads to an increasing demand for migrating files between different scheduling systems.

[0025] When tasks need to be migrated from one scheduling system to another, it is often necessary to manually generate new tasks based on the configuration of the old tasks. This not only consumes a lot of manpower and time, but also easily leads to inaccurate configuration of the migration tasks and repeated migration, resulting in low migration efficiency.

[0026] To this end, an embodiment of the present application provides a technical solution for file migration. In this technical solution, first, the DAG file to be migrated running in the first scheduling system is obtained, and then the target function and target grammar in the DAG file are converted using a large language model, so as to use the powerful language understanding and generation capabilities of the large language model to deeply understand the complex logic and semantic information in the DAG file, and accurately convert the target function and target grammar into functions and grammar that are easy for the parsing tool to parse. Next, the DAG file converted by the large language model is parsed to generate a syntax tree, and the syntax tree is traversed to obtain the DAG parameters and Operator parameters. Since the syntax tree can clearly display the structure and hierarchical relationship of the file, the key parameters in the file can be systematically and comprehensively extracted by traversing the syntax tree, avoiding the risk of missing important information, and providing a solid foundation for the subsequent parameter conversion and the creation of scheduling tasks. Then, the extracted DAG parameters and Operator parameters are respectively converted into first parameters and second parameters that are compatible with the second scheduling system. This parameter adaptation strategy fully considers the differences and requirements of different scheduling systems, so that the migrated files can perfectly meet the specifications and standards of the second scheduling system, thereby improving the compatibility and executability of the files in the target scheduling system, and avoiding migration failures or task execution exceptions caused by parameter mismatches. Finally, the preset interface of the second scheduling system is called, the corresponding scheduling tasks are created based on the converted parameters, and the created scheduling tasks are sent to the second scheduling system to realize automatic file migration. This entire migration process does not require a lot of manual intervention, which greatly reduces the cost of manual operation and the probability of error, improves the efficiency and convenience of file migration, and can quickly run tasks in the new scheduling system. See below for details.

[0027] Finally, for ease of understanding, an exemplary operating environment is provided below.

[0028] like Figure 1 As shown, the environment diagram includes a service platform 2, a network 4, and a client 6, wherein: The service platform 2 may be composed of a single or multiple computing devices. The multiple computing devices may include virtualized computing instances. Virtualized computing instances may include virtual machines, such as simulations of computer systems, operating systems, servers, etc. The computing device may load a virtual machine based on a virtual image and / or other data defining specific software (e.g., operating system, dedicated application, server) for simulation. As the demand for different types of processing services changes, different virtual machines may be loaded and / or terminated on one or more computing devices. A hypervisor may be implemented to manage the use of different virtual machines on the same computing device.

[0029] The service platform 2 may be configured to communicate with the client 6 or the like via a network 4. The network 4 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 may include physical links, such as coaxial cable links, twisted pair cable links, optical fiber links, combinations thereof, and the like, or wireless links, such as cellular links, satellite links, Wi-Fi links, and the like.

[0030] The service platform 2 can provide storage, reading, writing, querying, deleting and other services, such as providing file migration services for clients.

[0031] The client 6 may be an electronic device running an operating system such as Windows, Android™ or iOS, such as a smart phone, a tablet device, a laptop computer, a virtual reality device, a gaming device, a set-top box, a vehicle terminal, or a smart TV. Based on the above operating system, various application programs such as a file migration program may be run.

[0032] The client 6 may provide / configure a user access page for manipulating the service platform 2 or uploading objects, etc.

[0033] It should be noted that the above devices are exemplary, and the number and type of devices are adjustable in different scenarios or according to different needs.

[0034] The technical solutions of the present application are described below through multiple embodiments. It should be noted that these embodiments can be implemented in a variety of different forms and should not be interpreted as being limited to the embodiments described here.

[0035] Embodiment 1 Figure 2 The flowchart of the file migration method according to the first embodiment of the present application is schematically shown.

[0036] like Figure 2 As shown, the file migration method may include steps S200 to S210, wherein: Step S200: obtaining a DAG file to be migrated, where the DAG file runs in a first scheduling system.

[0037] Step S202: Use a large language model to convert the target function and target grammar in the DAG file to obtain a function and grammar corresponding to the target function and the target grammar.

[0038] Step S204: parsing the DAG file converted by the large language model to generate a syntax tree.

[0039] Step S206: traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file.

[0040] Step S208: Convert the DAG parameters and the Operator parameters into first parameters and second parameters respectively that are compatible with the second scheduling system.

[0041] Step S210, calling a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to achieve migration.

[0042] The file migration method provided in this embodiment first obtains the DAG file to be migrated running in the first scheduling system, and then uses the large language model to convert the target function and target syntax in the DAG file, thereby using the powerful language understanding and generation capabilities of the large language model to deeply understand the complex logic and semantic information in the DAG file, and accurately convert the target function and target syntax into functions and syntax that are easy to parse with the parsing tool. Next, the DAG file converted by the large language model is parsed to generate a syntax tree, and the syntax tree is traversed to obtain the DAG parameters and Operator parameters. Since the syntax tree can clearly display the structure and hierarchical relationship of the file, the key parameters in the file can be systematically and comprehensively extracted by traversing the syntax tree, avoiding the risk of missing important information, and providing a solid foundation for subsequent parameter conversion and scheduling task creation. Then, the extracted DAG parameters and Operator parameters are converted into first parameters and second parameters that are compatible with the second scheduling system, respectively. This parameter adaptation strategy fully considers the differences and requirements of different scheduling systems, so that the migrated files can perfectly meet the specifications and standards of the second scheduling system, thereby improving the compatibility and executability of the files in the target scheduling system, and avoiding migration failures or task execution exceptions caused by parameter mismatches. Finally, the preset interface of the second scheduling system is called, the corresponding scheduling tasks are created based on the converted parameters, and the created scheduling tasks are sent to the second scheduling system to realize automatic file migration. This entire migration process does not require a lot of manual intervention, which greatly reduces the cost of manual operation and the probability of error, improves the efficiency and convenience of file migration, and can quickly run tasks in the new scheduling system.

[0043] The following combination Figure 2 , each step in steps S200~S210 and other optional steps are explained in detail.

[0044] Step S200 , obtaining a DAG file to be migrated, where the DAG file runs in the first scheduling system.

[0045] The first scheduling system is a task scheduling system using a distributed visual DAG (Directed acyclic graph) workflow. The first scheduling system includes Oozie (task scheduling framework), Azkaban (batch workflow task scheduler), Airflow (workflow platform), etc.

[0046] In the first scheduling system, the DAG file is used to define task processes, configure scheduling rules, manage task dependencies, etc., and is the core basis for the first scheduling system to execute tasks.

[0047] In one embodiment, when the DAG file needs to be migrated, a migration request can be sent to the first scheduling system. When the first scheduling system receives the migration request, it can return the DAG file running in the first scheduling system to the initiator of the migration request, thereby obtaining the DAG file to be migrated.

[0048] In another implementation, the DAG file running in the first scheduling system may be stored in a preset location in advance, and when the DAG file needs to be migrated, the DAG file may be obtained from the preset location.

[0049] Step S202 , using a large language model to convert the target function and target grammar in the DAG file to obtain a function and grammar corresponding to the target function and the target grammar.

[0050] The large language model can convert the target function and target grammar in the DAG file into functions and grammar that can be easily parsed by the parsing tool later.

[0051] The large language model can be obtained by collecting a large number of code examples and corresponding DAG files to train and learn the existing open source large language model. The large language model can accurately identify the target function and target grammar in the DAG file, and convert the target function and target grammar into corresponding functions and grammars that can be parsed by the parsing tool.

[0052] The target function is a function that cannot be analyzed by the analysis tool, and the target function can be determined in advance according to the analysis function of the analysis tool. For example, if the analysis tool cannot analyze function a and function b, function a and function b can be used as the target function.

[0053] Similarly, the target grammar is a grammar that cannot be parsed by the parsing tool, and the target grammar can also be determined in advance according to the parsing function of the parsing tool. For example, if the parsing tool cannot parse grammar A and grammar B, grammar A and grammar B can be used as the target function.

[0054] It should be noted that when the large language model converts the DAG file, in addition to converting the target function and the target grammar, other contents will not be converted, so they will be kept as they are.

[0055] Step S204 , parse the DAG file that has been converted by the large language model to generate a syntax tree.

[0056] In one implementation, an open source LibCST parsing tool may be used to parse the DAG file that has been converted by the large language model, thereby generating a syntax tree.

[0057] Among them, the LibCST parsing tool is a Concrete SyntaxTree (CST) library for parsing and serializing Python code. It can parse Python source code into a CST tree and retain all format details.

[0058] In another embodiment, other parsing tools may be used to parse the DAG file after the large language model conversion process to generate a syntax tree. For example, an ANTLR parsing tool may be used.

[0059] Among them, the ANTLR parsing tool is a powerful parser generation tool that can generate parser code according to the grammar file to parse the DAG file and generate an abstract syntax tree (AST).

[0060] Step S206 , traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file.

[0061] DAG parameters are used to control the scheduling, operation, and management of DAG as a whole.

[0062] Operator parameters are used to define the behavior and characteristics of specific tasks.

[0063] Among them, DAG parameters can include parameters such as dag_id, schedule_interval, start_date, end_date, default_args, dagrun_timeout, tags, params, catchup, etc.

[0064] dag_id is used to uniquely identify a DAG. This ID can be used to find and manage a specific DAG in Airflow.

[0065] schedule_interval is used to determine the running frequency of the DAG and control the periodic execution of the workflow.

[0066] start_date is used to define the start time of the DAG. The DAG will be scheduled for execution only when the current time is later than this time.

[0067] end_date is used to limit the running time period of the DAG. If set, when the current time exceeds this date, the DAG will no longer be scheduled.

[0068] default_args is used to provide default parameters for all tasks in the DAG. This parameter simplifies task definitions and avoids repeated settings of common parameters.

[0069] dagrun_timeout is used to limit the maximum running time of a DAG, ensuring that a DAG that has not been completed for a long time is terminated in time to avoid wasting resources.

[0070] Tags are used to classify and mark DAGs, making it easier to filter and find DAGs in the Airflow Web UI.

[0071] params is used to define parameters at the DAG level. These parameters can be shared and accessed by all tasks in the DAG, making it easy to pass and use common data.

[0072] catchup is used to determine whether the DAG will catch up with previously unexecuted task instances, affecting the historical task execution status of the DAG.

[0073] Among them, Operator parameters can include task_id, dag, trigger_rule, depends_on_past, email, email_on_retry, execution_timeout, priority_weight, queue, owner, retries, retry_delay and other parameters.

[0074] task_id, as a unique identifier of a task in DAG, facilitates identification, management, and scheduling of specific tasks.

[0075] dag is used to specify the DAG to which the task belongs, ensuring that the task is executed in the correct DAG.

[0076] trigger_rule is used to define the triggering conditions of a task and control the circumstances under which the task starts to execute, providing flexible task execution control.

[0077] depends_on_past is used to determine whether a task depends on the successful completion of a previous task instance with the same name, and is used for task dependency management.

[0078] emai is used to specify the email address to send notifications when a task fails, so that you can promptly understand the failure of the task.

[0079] email_on_retry is used to determine whether to send an email notification when a task is retried. This parameter can be used to further refine the task notification mechanism.

[0080] execution_timeout is used to limit the maximum time for task execution, ensuring that the task does not run indefinitely and guaranteeing the timeliness of task execution.

[0081] priority_weight is used to define the priority weight of the task. This parameter can affect the priority of the task in the queue and determine the execution order of the tasks.

[0082] queue, which is used to specify the name of the queue where the task runs. It can be used in conjunction with the executor to control the task to run in a specific queue and achieve reasonable allocation of resources.

[0083] owner is used to record the person in charge of the task, which facilitates tracking of the ownership and management of the task.

[0084] retries is used to set the number of retries after a task fails. This parameter can enhance the reliability of the task and ensure that the task has a chance to be re-executed when encountering a temporary error.

[0085] retry_delay is used to define the interval between two retries to avoid the task from being retried too frequently and to give the system a certain buffer time.

[0086] In this embodiment, in the process of traversing the syntax tree, the definition nodes of DAG and Operator are found, and then the DAG parameters are extracted from the DAG nodes, and the Operator parameters are extracted from the Operator nodes.

[0087] In an optional embodiment, traversing the syntax tree to obtain the DAG parameters and Operator parameters in the DAG file may include: The syntax tree is traversed using the visitor pattern to obtain DAG parameters and Operator parameters in the DAG file.

[0088] Among them, the visitor pattern is a behavioral design pattern that can traverse complex object structures in operations without changing the class of the element.

[0089] In this embodiment, by adopting the visitor pattern to traverse the syntax tree, the definition nodes of the DAG and the Operator can be more accurately identified from the syntax tree, and then the DAG parameters and the Operator parameters can be extracted.

[0090] In other implementations, other modes may be used to traverse the syntax tree. For example, a listener mode may be used to traverse the syntax tree.

[0091] Step S208 , converting the DAG parameters and the Operator parameters into first parameters and second parameters respectively adapted to the second scheduling system.

[0092] Since different scheduling systems may have different task definition methods and parameter requirements, the extracted parameters need to be adapted and converted. Therefore, in this embodiment, after obtaining the DAG parameters and the Operator parameters, the DAG parameters need to be converted into first parameters that can be recognized and used by the second scheduling system, and at the same time, the Operator parameters need to be converted into second parameters that can be recognized and used by the second scheduling system.

[0093] In practical applications, parameter conversion can be performed in a variety of ways, and an exemplary way is provided below.

[0094] In an alternative embodiment, see Figure 3 , converting the DAG parameter and the Operator parameter into a first parameter and a second parameter respectively adapted to the second scheduling system may include: Step S300: Obtain a parameter mapping relationship table between the first scheduling system and the second scheduling system.

[0095] Step S302: Convert the DAG parameter and the Operator parameter into a first parameter and a second parameter respectively adapted to the second scheduling system according to the parameter mapping relationship table.

[0096] In one embodiment, the mapping relationship between various parameters of the first scheduling system and the second scheduling system may be predefined. For example, the parameter corresponding to DAG parameter a is parameter A, and the parameter corresponding to DAG parameter b is parameter B. The parameter corresponding to Operator parameter c is parameter C. The parameter corresponding to Operator parameter d is parameter D.

[0097] In this embodiment, after obtaining the DAG parameters and the Operator parameters, a predefined parameter mapping relationship table between the first scheduling system and the second scheduling system may be queried to find the first parameter corresponding to the DAG parameter and the second parameter corresponding to the Operator parameter.

[0098] In this embodiment, the conversion between parameters can be achieved quickly and accurately based on the parameter mapping relationship table.

[0099] In another implementation, the conversion between parameters may also be achieved in other ways. For example, the DAG parameters and the Operator parameters may be first converted into parameters in a common format, and then the parameters in the common format may be converted into the first parameters and the second parameters.

[0100] Step S210 , calling the preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to achieve migration.

[0101] The preset interface is an interface for creating a scheduling task based on the first parameter and the second parameter.

[0102] In one embodiment, when creating a scheduling task, the preset interface will first read the first parameter and the second parameter, and then set the task parameters based on the first parameter and the second parameter according to the requirements of the second scheduling system, thereby realizing the creation of the scheduling task.

[0103] In an alternative embodiment, see Figure 4 , the calling of the preset interface of the second scheduling system creates a corresponding scheduling task based on the first parameter and the second parameter, and sends the created scheduling task to the second scheduling system to realize the migration, including: Step S400: calling a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter.

[0104] Step S402: verify the created scheduling task.

[0105] Step S404: after verification, the created scheduling task is sent to the second scheduling system to achieve migration.

[0106] In this embodiment, after the scheduling task is created, in order to ensure that the created scheduling task has no problem and can be used normally, so as to facilitate direct use by the second scheduling system, the created scheduling task can be sent to a preset verification terminal so that the verification terminal verifies the created scheduling task.

[0107] After the verification is passed, the created scheduling task is sent to the second scheduling system to achieve migration.

[0108] When the verification fails, a corresponding prompt message can be generated to enable the staff to make further modifications and verifications. When the verification passes, the created scheduling task can be sent to the second scheduling system to achieve migration.

[0109] Of course, it can be understood that in specific implementation, the created scheduling task can be directly sent to the second scheduling system, and the second scheduling system verifies it. After the verification is passed, it is released to the production environment for use.

[0110] Through the above method, the created scheduling tasks are verified in advance to ensure that there are no problems with the created scheduling tasks and they can be released to the production environment for use, so that they can be directly used by the second scheduling system.

[0111] In practical applications, the created scheduling tasks can be verified in a variety of ways. An exemplary way is provided below.

[0112] In an alternative embodiment, see Figure 5 , the verification of the created scheduling task includes: Step S500: obtaining a first execution code generated when the first scheduling system executes a scheduling task of a first preset type in the DAG file.

[0113] Step S502: obtaining a second execution code generated by the second scheduling system executing a scheduling task corresponding to the scheduling task of the first preset type.

[0114] Step S504: If the first execution code is the same as the second execution code, verification is successful.

[0115] The first preset type is a task type that is preset according to actual conditions. For example, the scheduling task of the first preset type is a SQL task.

[0116] In this embodiment, when the DAG file contains a first preset type of scheduling task, one or more first preset type of scheduling tasks can be executed by the first scheduling system, and then the first execution code generated when the first scheduling system executes one or more first preset type of scheduling tasks is obtained. For example, a SQL task is executed at 2 o'clock every day, and its execution code is: select a from table where log_date=20250422.

[0117] At the same time, the second scheduling system can execute one or more scheduling tasks of the first preset type corresponding to the preset scheduling task, and then obtain the second execution code generated by the second scheduling system executing one or more scheduling tasks corresponding to the preset scheduling task.

[0118] After obtaining the first execution code and the second execution code, the two can be compared to see if they are the same. If they are the same, it indicates that the verification is successful. If they are different, it indicates that the verification is unsuccessful.

[0119] It should be noted that the above-mentioned scheduling task corresponding to the preset scheduling task refers to the scheduling task under the same scheduling time.

[0120] In this embodiment, verification is achieved by code comparison, which can reduce resource consumption and improve verification efficiency compared to the dual-run verification method in the prior art.

[0121] Another exemplary approach is provided below.

[0122] In an alternative embodiment, see Figure 6 , the verification of the created scheduling task includes: Step S600: obtaining a first execution code generated when the first scheduling system executes a scheduling task of a second preset type in the DAG file and first parameter information used to execute the scheduling task of the second preset type in the DAG file.

[0123] Step S602: obtaining a second execution code generated by the second scheduling system when executing a scheduling task corresponding to the second preset type of scheduling task and second parameter information used when executing the scheduling task corresponding to the second preset type of scheduling task.

[0124] Step S604: If the first execution code is the same as the second execution code, and the first parameter information is the same as the second parameter information, verification is successful.

[0125] The second preset type is a task type that is preset according to actual conditions. For example, the scheduling task of the second preset type is generally a non-SQL type task, such as a task for executing a Shell command or a script.

[0126] In this embodiment, when the DAG file includes a scheduling task of the second preset type, one or more scheduling tasks of the second preset type can be executed by the second scheduling system, and then, the first execution code generated by the first scheduling system when executing one or more scheduling tasks of the second preset type and the first parameter information used to execute one or more scheduling tasks of the second preset type in the DAG file are obtained.

[0127] At the same time, one or more scheduling tasks corresponding to the second preset type of scheduling tasks can be executed by the second scheduling system, and then the second execution code generated by the second scheduling system to execute one or more scheduling tasks corresponding to the second preset type of scheduling tasks is obtained, and the second parameter information used by the second scheduling system to execute one or more scheduling tasks corresponding to the second preset type of scheduling tasks is obtained.

[0128] After obtaining the first execution code and the second execution code, the two can be compared to see if they are the same. If they are the same, the first parameter information and the second parameter information can be further compared to see if they are the same. If the first parameter information and the second parameter information are also the same, it indicates that the verification is passed. If they are different, it indicates that the verification is not passed.

[0129] In this embodiment, verification is achieved by code and parameter comparison, which can reduce resource consumption and improve verification efficiency compared to the dual-run verification method in the prior art.

[0130] Embodiment 2 Figure 7 The block diagram of the file migration device 700 according to the second embodiment of the present application is schematically shown. The device can be divided into one or more program modules, one or more program modules are stored in a storage medium, and are executed by one or more processors to complete the embodiment of the present application. The program module referred to in the embodiment of the present application refers to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. Figure 7 As shown, the apparatus 700 may include: an acquisition module 710, a first conversion module 720, a parsing module 730, a traversal module 740, a second conversion module 750 and a creation module 760, wherein: An acquisition module 710 is used to acquire a DAG file to be migrated, where the DAG file runs in a first scheduling system; A first conversion module 720 is used to convert the target function and the target grammar in the DAG file by using a large language model to obtain a function and a grammar corresponding to the target function and the target grammar; A parsing module 730 is used to parse the DAG file after the large language model conversion process to generate a syntax tree; A traversal module 740 is used to traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file; A second conversion module 750, configured to convert the DAG parameter and the Operator parameter into a first parameter and a second parameter respectively adapted to a second scheduling system; The creation module 760 is used to call the preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and send the created scheduling task to the second scheduling system to achieve migration.

[0131] As an optional embodiment, traversing the syntax tree to obtain DAG parameters and Operator parameters in the DAG file includes: The syntax tree is traversed using the visitor pattern to obtain DAG parameters and Operator parameters in the DAG file.

[0132] As an optional embodiment, the converting the DAG parameter and the Operator parameter into a first parameter and a second parameter respectively adapted to the second scheduling system includes: Obtaining a parameter mapping relationship table between the first scheduling system and the second scheduling system; The DAG parameters and the Operator parameters are respectively converted into first parameters and second parameters that are compatible with the second scheduling system according to the parameter mapping relationship table.

[0133] As an optional embodiment, the calling of the preset interface of the second scheduling system creates a corresponding scheduling task based on the first parameter and the second parameter, and sends the created scheduling task to the second scheduling system to achieve migration, including: Calling a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter; Verify the created scheduling tasks; After the verification is passed, the created scheduling task is sent to the second scheduling system to achieve migration.

[0134] As an optional embodiment, the verification of the created scheduling task includes: Obtaining a first execution code generated when the first scheduling system executes a scheduling task of a first preset type in the DAG file; Acquire a second execution code generated by the second scheduling system executing a scheduling task corresponding to the scheduling task of the first preset type; When the first execution code is identical to the second execution code, the verification succeeds.

[0135] As an optional embodiment, the verification of the created scheduling task includes: Acquire a first execution code generated when the first scheduling system executes a scheduling task of a second preset type in the DAG file and first parameter information used to execute the scheduling task of the second preset type in the DAG file; Acquire a second execution code generated by the second scheduling system when executing a scheduling task corresponding to the second preset type of scheduling task and second parameter information used when executing the scheduling task corresponding to the second preset type of scheduling task; If the first execution code is the same as the second execution code, and the first parameter information is the same as the second parameter information, the verification is successful.

[0136] Embodiment 3 Figure 8 The schematic diagram of the hardware architecture of a computer device 10000 suitable for implementing the file migration method according to the third embodiment of the present application is schematically shown. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc. Figure 8 As shown, the computer device 10000 includes but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10010 can be an internal storage module of the computer device 10000, such as a hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 can also be an external storage device of the computer device 10000, such as a plug-in hard disk equipped on the computer device 10000, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the memory 10010 can also include both the internal storage module of the computer device 10000 and its external storage device. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed in the computer device 10000, such as program codes of the file migration method, etc. In addition, the memory 10010 can also be used to temporarily store various data that have been output or will be output.

[0137] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0138] The network interface 10030 may include a wireless network interface or a wired network interface, and the network interface 10030 is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and to establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0139] It should be pointed out that Figure 8 Only a computer device having components 10010 - 10030 is shown, but it should be understood that implementation of all of the components shown is not a requirement, and more or fewer components may alternatively be implemented.

[0140] In this embodiment, the file migration method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as processor 10020) to complete the embodiment of the present application.

[0141] Embodiment 4 An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the file migration method in the embodiment are implemented.

[0142] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of a computer device, such as a hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium can also be an external storage device of a computer device, such as a plug-in hard disk equipped on the computer device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the computer-readable storage medium can also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the computer-readable storage medium is generally used to store an operating system and various application software installed on the computer device, such as the program code of the file migration method in the embodiment, etc. In addition, the computer-readable storage medium can also be used to temporarily store various types of data that have been output or are to be output.

[0143] Embodiment 5 An embodiment of the present application also provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0144] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present application can be implemented by general-purpose computer devices, they can be concentrated on a single computer device, or distributed on a network composed of multiple computer devices, optionally, they can be implemented by executable program codes of computer devices, so that they can be stored in a storage device and executed by the computer device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0145] It should be noted that the above are only preferred embodiments of the present application, and the patent protection scope of the present application is not limited thereto. Any equivalent structure or equivalent process transformation made using the contents of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A file migration method, characterized in that: The method comprises: Obtaining a DAG file to be migrated, where the DAG file runs in the first scheduling system; Using a large language model to convert the target function and the target grammar in the DAG file to obtain a function and a grammar corresponding to the target function and the target grammar; Parsing the DAG file converted by the large language model to generate a syntax tree; Traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file; Convert the DAG parameter and the Operator parameter into a first parameter and a second parameter respectively adapted to the second scheduling system; The preset interface of the second scheduling system is called to create a corresponding scheduling task based on the first parameter and the second parameter, and the created scheduling task is sent to the second scheduling system to achieve migration.

2. The method according to claim 1, characterized in that The traversing of the syntax tree to obtain DAG parameters and Operator parameters in the DAG file includes: The syntax tree is traversed using the visitor pattern to obtain DAG parameters and Operator parameters in the DAG file.

3. The method according to claim 1, characterized in that The converting the DAG parameter and the Operator parameter into a first parameter and a second parameter respectively adapted to the second scheduling system comprises: Obtaining a parameter mapping relationship table between the first scheduling system and the second scheduling system; The DAG parameters and the Operator parameters are respectively converted into first parameters and second parameters that are compatible with the second scheduling system according to the parameter mapping relationship table.

4. The method according to any one of claims 1 to 3, characterized in that: The calling of the preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and sending the created scheduling task to the second scheduling system to achieve migration includes: Calling a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter; Verify the created scheduling tasks; After the verification is passed, the created scheduling task is sent to the second scheduling system to achieve migration.

5. The method according to claim 4, characterized in that The verification of the created scheduling task includes: Obtaining a first execution code generated when the first scheduling system executes a scheduling task of a first preset type in the DAG file; Acquire a second execution code generated by the second scheduling system executing a scheduling task corresponding to the scheduling task of the first preset type; When the first execution code is identical to the second execution code, the verification succeeds.

6. The method according to claim 4, characterized in that The verification of the created scheduling task includes: Acquire a first execution code generated when the first scheduling system executes a scheduling task of a second preset type in the DAG file and first parameter information used to execute the scheduling task of the second preset type in the DAG file; Acquire a second execution code generated by the second scheduling system when executing a scheduling task corresponding to the second preset type of scheduling task and second parameter information used when executing the scheduling task corresponding to the second preset type of scheduling task; If the first execution code is the same as the second execution code, and the first parameter information is the same as the second parameter information, the verification is successful.

7. A file migration device, characterized in that: The device comprises: An acquisition module, used to acquire a DAG file to be migrated, where the DAG file runs in the first scheduling system; A first conversion module, configured to convert the target function and the target grammar in the DAG file by using a large language model to obtain a function and a grammar corresponding to the target function and the target grammar; A parsing module, used to parse the DAG file converted by the large language model to generate a syntax tree; A traversal module, used to traverse the syntax tree to obtain DAG parameters and Operator parameters in the DAG file; A second conversion module, used to convert the DAG parameter and the Operator parameter into a first parameter and a second parameter respectively adapted to a second scheduling system; A creation module is used to call a preset interface of the second scheduling system to create a corresponding scheduling task based on the first parameter and the second parameter, and send the created scheduling task to the second scheduling system to achieve migration.

8. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Deep learning scheduling configuration system and method

    CN113127182A

  • Code migration method and device for front-end framework upgrading, equipment and medium

    CN118409793A

  • Method and system for automatic workflow generation by large language models

    US20250068398A1

  • Method, apparatus and device for workflow migration, and computer-readable storage medium

    WO2021068692A1

  • Applet cross-application migration method, device, terminal, system and storage medium

    WO2023035563A1