Method for arranging and running AI task flow
Through custom operators and container execution, the problems of low resource utilization and incomplete isolation in the existing technology are solved, efficient AI task flow orchestration and visual display are realized, and user operations are simplified.
Patent Information
- Application Number
- CN202510159290.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-13
AI Technical Summary
In the prior art, AI task flow orchestration tools such as Airflow and DolphinScheduler have problems with low resource utilization and incomplete resource isolation, making it difficult for users to customize operators and the orchestration is complicated.
Through user-defined operators, encapsulated into mirrors and registered to the task flow orchestration and scheduling platform, task flow is executed in container form, and orchestrated through drag and drop, visual display and result storage of task flow are realized.
Improve resource utilization and resource isolation, simplify the task flow orchestration process, and improve user work efficiency. Users can quickly create and modify task processes without complex configuration or programming.
Smart Images

Figure CN120085986A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more particularly, to a method for orchestrating and running AI task flows. Background Art
[0002] Cloud-native AI development platforms integrate mature artificial intelligence development frameworks and the ability of cloud-native tools to flexibly call cloud resources and efficiently deploy cloud applications. On the one hand, they help enterprise developers improve the development efficiency of algorithm models. On the other hand, they enhance the efficiency of the delivery, deployment, and operation and maintenance links and reduce various costs. Artificial intelligence model training requires repeated running, evaluation, and correction of the model. Through automated process components brought by cloud-native such as pipelines, the automation level of processes such as parameter selection, hyperparameter tuning, and periodic training with new data can be improved, and the model output speed can be enhanced.
[0003] In cloud-native scenario AI task flow orchestration, the carrier of tasks is generally a process, and the process methods are such as Airflow and DolphinScheduler. In this architecture, before the task starts, when the platform is deployed, the components required for scheduling are first deployed in the form of pods. For example, the worker executor pod. When deploying the worker executor, machine resources are already occupied. Airflow cannot achieve the orchestration of task flows in the form of drag-and-drop on the web interface. Users need to orchestrate Python scripts to define each task and the upstream and downstream relationships of tasks in the task flow, which is of relatively high difficulty. In addition, the tasks of Airflow are carried in the form of processes, and the tasks run in the pre-deployed workers. Even if the tasks are not running, the processes will still occupy resources on the machine. Therefore, the resource utilization rate is relatively low, and resource isolation between tasks cannot be provided. DolphinScheduler can achieve the drag-and-drop form of task flow orchestration on the web interface, but users cannot customize operators. In addition, the tasks of DolphinScheduler, like Airflow, are carried in the form of processes and run in the pre-deployed workers, also having the problems of low resource utilization rate and incomplete resource isolation. Summary of the Invention
[0004] To solve the above problems, the purpose of the present invention is to provide a method for orchestrating and running AI task flows.
[0005] A method for orchestrating and running AI task flows includes:
[0006] Step 1: The user develops the code of a custom operator by themselves;
[0007] Step 2: Package the custom operator into an image;
[0008] Step 3: Register the image into the operator library in the task flow orchestration and scheduling platform;
[0009] Step 4: The user performs task flow orchestration of operators in a drag-and-drop manner on the WEB interface and configures task execution parameters;
[0010] Step 5: Orchestrate multiple nodes of upstream and downstream tasks and the execution parameters of each task into a task flow and execute it in the form of a container;
[0011] Step 6: After the container runs to completion, return the running result to the WEB interface for visual display to the user.
[0012] Preferably, in Step 2, the custom operator is encapsulated into a container image, and the custom operator has a start time and an end time.
[0013] Preferably, in Step 5, the task flow orchestration and scheduling platform renders the corresponding operator into the web interface according to the task execution parameters, and then the user fills in the input parameters. The task flow orchestration and scheduling platform passes the user input parameters to the image of the corresponding operator, and then starts the task.
[0014] Preferably, in Step 6, the container includes a control container and a business container; start the control container, and the control container starts the business container to run the orchestrated task flow. The business container stores the visual output of the running result into the / metric.json file; the control container obtains the / metric.json file from the business container, and the control container writes the / metric.json file into the distributed storage; read and parse the / metric.json from the distributed storage and display it in the WEB interface.
[0015] Preferably, in the task flow, the variable transfer between upstream and downstream tasks is divided into key-value type variable transfer, file transfer, and cache transfer.
[0016] Preferably, for key-value type variable transfer, each task node in the task flow will start the corresponding business container and control container. The business container will output the variable, write it into the / output file, and the control container will write it into the distributed storage. The downstream node reads the upstream output as input through {{task_name.output}}.
[0017] Preferably, for file transfer, each task will automatically mount the personal distributed storage directory into the container. The container can write the variable into the file in the upstream task and read the file in the downstream task.
[0018] Preferably, for cache transfer, the container of each task will configure the address of the cache redis in the form of environment variables, and users can read and write the variables in the cache in the task by themselves.
[0019] The present invention also provides an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored on the memory and executable on the processor. The transceiver, the memory, and the processor are connected through the bus. The feature is that when the computer program is executed by the processor, it realizes the steps in the above method for orchestrating and running an AI task flow.
[0020] The present invention also provides a computer-readable storage medium, on which a computer program is stored. The feature is that when the computer program is executed by a processor, it realizes the steps in the above method for orchestrating and running an AI task flow.
[0021] The beneficial effect of the method for orchestrating and running an AI task flow provided by the present invention is that: compared with the prior art, the present invention uses containers to carry out the execution of tasks, can occupy resources on-site during task execution, achieve higher resource utilization and resource isolation, and the present invention realizes the orchestration of the task flow by means of drag-and-drop, enabling users to quickly create and modify the task process without complex configuration or programming, thus improving work efficiency.
[0022] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. Description of the Drawings
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 Shows a flowchart of a method for orchestrating and running an AI task flow provided by an embodiment of the present invention;
[0025] Figure 2 Shows a schematic diagram of task flow orchestration provided by an embodiment of the present invention;
[0026] Figure 3 Shows a schematic diagram of using a process as a task carrier provided by an embodiment of the present invention;
[0027] Figure 4 Shows a schematic diagram of using a container as a carrier provided by an embodiment of the present invention;
[0028] Figure 5 Shows the flowchart of the visualization display implementation provided by the embodiments of the present invention;
[0029] Figure 6 Shows the effect diagram of the front-end display provided by the embodiments of the present invention;
[0030] Figure 7 Shows the flowchart of the variable transfer provided by the embodiments of the present invention;
[0031] Figure 8 Shows the flowchart of the file transfer provided by the embodiments of the present invention;
[0032] Figure 9 Shows the schematic diagram of the defined constants provided by the embodiments of the present invention;
[0033] Figure 10 Shows the schematic diagram of the use of constants provided by the embodiments of the present invention;
[0034] Figure 11 Shows the schematic diagram of the flow direction control of the task flow provided by the embodiments of the present invention. Detailed implementation manners
[0035] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0036] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0037] In the present invention, unless otherwise clearly specified or limited, terms such as "installation", "connection", "linkage", "fixation", etc. shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be a direct connection or an indirect connection through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0038] Please refer to Figure 1 , a method for orchestrating and running an AI task flow, including:
[0039] Step 1: The user develops the code of a custom operator by himself.
[0040] Step 2: Package the custom operator into an image.
[0041] In Step 2, the custom operator is packaged into a container image, and the custom operator has a start time and an end time.
[0042] Step 3: Register the image in the operator library of the task flow orchestration and scheduling platform.
[0043] Step 4: The user performs the task flow orchestration of the operator in a drag-and-drop manner on the WEB interface and configures the task execution parameters.
[0044] An operator fixes the function through code and can receive the parameters of the user, so as to realize a template in which a user without a foundation can realize an algorithm or engineering processing only through simple parameter configuration. A task operator is each node in the task flow after the operator is dragged to the task flow orchestration interface, and it is an actually executable process. In order to realize the reuse of the general algorithm training function and enable a user without algorithm skills to directly perform algorithm debugging, it is necessary to abstractly define the general function as a task operator. The task operator contains some general algorithm training functions, forming an operator list as shown on the left. The user drags and drops the task operator to the canvas to form a task, configures the task start parameters, completes the definition of the task, and defines the upstream and downstream relationships between multiple tasks through connections.
[0045] Step 5: Orchestrate multiple nodes of upstream and downstream tasks and the execution parameters of each task into a task flow and execute it in the form of a container.
[0046] In the said Step 5, the task flow orchestration and scheduling platform renders the corresponding operator into a web interface according to the task execution parameters, and then the user fills in the input parameters. The task flow orchestration and scheduling platform passes the user input parameters to the image of the corresponding operator, and then starts the task.
[0047] Step 6: After the container runs to completion, return the running result to the WEB interface for visual display to the user.
[0048] In Step 6, the container includes a control container and a business container; start the control container, and the control container starts the business container to run the choreographed task flow. The business container stores the visual output of the running result in the / metric.json file; the control container obtains the / metric.json file from the business container, and the control container writes the / metric.json file to the distributed storage; read the / metric.json from the distributed storage and parse it for display in the WEB interface.
[0049] In the task flow, the variable transfer between upstream and downstream tasks is divided into variable transfer of key-value type, file transfer, and cache transfer.
[0050] For the variable transfer of key-value type, each task node in the task flow will start the corresponding business container and control container. The business container will output the variable, write it to the / output file, and the control container will write it to the distributed storage. The downstream node reads the upstream output as input through {{task_name.output}}.
[0051] For file transfer, each task will automatically mount the personal distributed storage directory to the container. The container can write variables to files in the upstream task and read files in the downstream task.
[0052] For cache transfer, the container of each task will configure the address of the cache redis in the form of an environment variable, and the user can read and write variables in the cache in the task by themselves.
[0053] The following further describes a method for choreographing and running an AI task flow in the present invention in combination with specific embodiments:
[0054] As Figure 1-2 shown, the user fills in parameters in the WEB interface according to the corresponding operators, and the parameters will be saved to the platform backend. The platform relies on the mirror to start the container according to the operators registered by the developer and the parameters filled in by the user, and the parameters are passed into the container. That is, the parameters filled in by the user are passed to the function code defined by the developer through the container, thereby realizing the reuse of the developer's function code. After the container runs to completion, return the running result to the WEB interface for visual display to the user.
[0055] To implement the above process, the following six functions need to be implemented first:
[0056] 1. Carrier of the task: The carrier of the task should be a container. Because this will make it more convenient for AI task resource scheduling and isolation in a multi-machine environment. Figure 3 It is a schematic diagram of a process as the task carrier. Before the task runs, the container starts first and the resources are occupied. Figure 4 It is a schematic diagram of a container as the carrier. The task and the container are in one-to-one correspondence. When a task needs to run, the container is started and resource occupation occurs. Therefore, using a container as the task carrier can better achieve high resource utilization and resource isolation.
[0057] 2. Formal definition of the operator: The operator is not defined by the platform but by the user himself. However, when the user defines the operator, he needs to follow the rules for operator definition by the platform. The user can implement any operator code without code language restrictions. After the code development is completed, it can be packaged into a container image and then registered as an operator. Only the relevant information of the input parameters required for running the operator image needs to be told to the platform.
[0058] Rules for operator definition:
[0059] 1) The operator must have a start time and an end time and cannot be permanently online;
[0060] Because the operator finally exists as a task node in the task flow, it must have a start time and an end time, otherwise the next task cannot be started.
[0061] 2) There is no language restriction for the written operator code;
[0062] Because the operator ultimately needs to be built into a container image, the implementation language of the operator is not concerned as long as it can be encapsulated into an image.
[0063] 3) The operator's code should be able to accept the startup parameters and startup commands for starting the code;
[0064] In this way, the operator can be reused, allowing different users to achieve different functions by setting different startup parameters. The types of startup parameters must be strings, long texts, json, etc.
[0065] 4) The operator should be able to be packaged into a container image.
[0066] Ultimately, the operator runs in a cloud-native environment, so it must exist in the form of an image because the cloud-native environment is essentially container orchestration and containers are started based on images.
[0067] The operators that users can implement by themselves are divided into 4 categories:
[0068] 1) Request the platform external interface through the api;
[0069] Send a request to the external server interface of the platform and obtain the corresponding returned data. Developers are usually developers of the external platforms of this platform. For example, data is pulled locally through a certain API, such as downloading the Hugging Face dataset, downloading models, and importing data from the annotation platform.
[0070] 2) Call the platform's own interface through the API;
[0071] This type of operator calls the internal interface. Developers are usually developers of the development platform. The operator can achieve automatic linkage of each module within the platform through the pipeline between different modules. For example, deploying an inference service through the pipeline within the platform itself and automatically releasing it online. Similar to the above as a client, except that here it is connected to the platform itself, so it can be linked with other modules of the platform to achieve automation. If it is a simple pipeline, an online service cannot be deployed, and only services with a clear end time can be deployed. However, through the operator, the linkage between the pipeline module and the inference service module can be achieved.
[0072] 3) Execute an independently running task;
[0073] Independent running means that it does not need to interact with other application platforms and can be independently executed within a single container. For example, training of traditional machine learning models, evaluation of models, and offline inference, as well as processing of image, text, and audio data.
[0074] 4) Start a distributed cluster;
[0075] Start, monitor, and recycle this distributed cluster. This type of operator needs to handle multiple machines and multiple pods, so it requires relatively high-level permissions and relatively complex parameter configurations. For example, distributed TensorFlow, PyTorch, and distributed DeepSpeed.
[0076] For example, if the present invention now wants to develop a decision tree model operator to allow users to reuse the decision tree model training code and only needs to control operator startup parameters such as the input data address and model parameters, the following steps are required:
[0077] 1) Develop the decision tree model training code;
[0078] 2) Write a Dockerfile to define the environment required to run the code in 1);
[0079] 3) Write a shell script to package the code into an image and push the image to the image repository;
[0080] 4) Register the image as an operator, configure the startup parameters, and publish the template;
[0081] 5) Drag the operator in the pipeline to form a task node, configure the startup parameters, and execute.
[0082] The page for filling in parameters in the WEB interface is as follows. These are the startup parameters related to the training of a decision tree model operator.
[0083]
[0084] The platform can render the operator usage into a web interface based on the information of the operator input parameters. Then, the user fills in the input parameters, and the platform passes the user input parameters to the operator's image, thereby starting the task.
[0085] Visualization results of the task: When the user is running the container, the operator actively writes the visualization content to the / metric.json in the container. The defined format is as follows: It can be in the form of visualizable pictures, texts, csv data, echart source codes, html source codes, iframes embedding other pages, etc.
[0086] The program flow of result visualization is as Figure 5 shown.
[0087] 1. First, the task will start the control container, and at the same time, the backend tells the control container the task id in the form of environment variables;
[0088] 2. The control container starts the business container to run the operator. The business container will generate the output to be visualized and store it in the / metric.json file in the form defined above;
[0089] 3. The control container will actively obtain / metric.json from the business container and write the file to the distributed storage by the control container;
[0090] 4. The backend reads / metric.json from the distributed storage;
[0091] 5. The front end parses / metric.json and displays it in the WEB interface, as Figure 6 shown.
[0092] Input and output of the task:
[0093] The input and output of the task, that is, the variable transfer between upstream and downstream tasks in the task flow, is divided into three forms. The first is the variable transfer of the key-value type, one is file transfer, and the other is cache transfer.
[0094] 1. Transfer of variables (for variables of the key-value type)
[0095] Such as Figure 7As shown, each task node in the task flow will start the corresponding business container and control container. The business container will output variables, write them to the / output file, and the control container will write them to the distributed storage. Downstream nodes read the upstream output as input through {{task_name.output}}.
[0096] When the system reads that the template variable represented by {{}} contains the upstream node name, it will automatically convert it to the upstream output of argo, and argo will implement the output variable transferred by the object storage.
[0097] Since the control container has completed the variable transfer before the business container starts, these variables can be used as the startup parameters of the business container, which is the biggest difference from the other two scenarios.
[0098] 2. File transfer (for large file variables)
[0099] As Figure 8 shown, each task will automatically mount the personal distributed storage directory to the container. The task container can write variables to a file in the upstream task and read the file in the downstream task.
[0100] 3. Cache transfer (for large memory variables)
[0101] The task container of each task will configure the address of the cached redis in the form of environment variables, and users can read and write variables in the cache in the task by themselves.
[0102] Global constants of the task flow: defined by {{}} by the user, supporting python objects such as datetime. Users need to configure them by themselves when using. For example, the process of defining a time variable is as follows:
[0103] As Figure 9 shown, define a constant
[0104] YYYYMMDD = {{datetime.datetime.now().strftime('%Y-%m-%d')}}
[0105] As Figure 10 shown, use the constant
[0106] v{{YYYYMMDD}}
[0107] Flow control of the task flow: As Figure 11 shown, the downstream task to be selected for execution is determined by the downstream task name included in the last line of the task log.
[0108] The present invention uses a container to carry out the execution of tasks, which can occupy resources on-site during task execution, achieving higher resource utilization and resource isolation. Moreover, the present invention realizes the orchestration of task flows through a drag-and-drop method, enabling users to quickly create and modify task processes without complex configuration or programming, thereby improving work efficiency.
[0109] The present invention also provides an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored on the memory and executable on the processor. The transceiver, the memory, and the processor are connected through the bus. It is characterized in that when the computer program is executed by the processor, it realizes the steps in the above-mentioned method for running an AI task flow orchestration. Compared with the prior art, the beneficial effects of the electronic device provided by the present invention are the same as those of the above-mentioned method for running an AI task flow orchestration, and will not be elaborated here.
[0110] The present invention also provides a computer-readable storage medium, on which a computer program is stored. It is characterized in that when the computer program is executed by a processor, it realizes the steps in the above-mentioned method for running an AI task flow orchestration. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present invention are the same as those of the above-mentioned method for running an AI task flow orchestration, and will not be elaborated here.
[0111] As mentioned above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of technical solutions of changes or substitutions, which should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for arranging and running an AI task flow, characterized in that: include: Step 1: The user develops the code of the custom operator; Step 2: Encapsulate the custom operator into an image; Step 3: Register the image to the operator library in the task flow orchestration and scheduling platform; Step 4: The user drags and drops the task flow of the operator and configures the execution parameters of the task in the WEB interface; Step 5: Arrange multiple nodes of upstream and downstream tasks and the execution parameters of each task into a task flow and execute it in the form of a container; Step 6: After the container runs, the running results are returned to the WEB interface for visual display to the user.
2. The method for arranging and running an AI task flow according to claim 1, characterized in that: In step 2, the custom operator is encapsulated into a container image, and the custom operator has a start time and an end time.
3. The method for arranging and running an AI task flow according to claim 2, characterized in that: In step 5, the task flow orchestration and scheduling platform renders the corresponding operator into a web interface according to the task execution parameters, and then the user fills in the input parameters. The task flow orchestration and scheduling platform passes the user input parameters to the image of the corresponding operator, and then starts the task.
4. The method for arranging and running an AI task flow according to claim 3, characterized in that: In step 6, the container includes a control container and a business container; the control container is started, and the control container starts the business container to run the orchestrated task flow, and the business container stores the visual output of the running results in the / metric.json file; the control container obtains the / metric.json file from the business container, and the control container writes the / metric.json file to the distributed storage; the / metric.json is read from the distributed storage and parsed and displayed in the WEB interface.
5. A method for arranging and running an AI task flow according to any one of claims 1 to 4, characterized in that: In the task flow, variable transfer between upstream and downstream tasks is divided into key-value type variable transfer, file transfer and cache transfer.
6. The method for arranging and running an AI task flow according to claim 5, characterized in that: For key-value type variable transfer, each task node in the task flow will start the corresponding business container and control container. The business container will output the variable and write it to the / output file, which will be written to the distributed storage by the control container. The downstream node reads the upstream output as input through {{task_name.output}}.
7. The method for arranging and running an AI task flow according to claim 5, characterized in that: For file delivery, each task automatically mounts its own distributed storage directory into the container. The container can write variables to files in upstream tasks and read files in downstream tasks.
8. The method for arranging and running an AI task flow according to claim 5, characterized in that: For cache delivery, the container of each task will configure the address of the cache redis in the form of environment variables, and users can read and write variables in the cache in the task.
9. An electronic device, comprising a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the transceiver, the memory, and the processor are connected via the bus, wherein: When the computer program is executed by the processor, the steps in the method for choreographing and executing an AI task flow as described in any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method for choreographing and running an AI task flow as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Cloud high-performance scientific calculation workflow design control system and user graphical interface
CN112162727A
Medical artificial intelligence reasoning method and system based on low-code programming
CN115346669A
Task scheduling method and device, electronic equipment and storage medium
CN116483531A
Application development method and system based on container technology, electronic equipment and storage medium
CN116880823A
Cross-platform data fusion service customization method and device, equipment and medium
CN117235036A