A method for processing data transfer between task nodes in a machine learning platform
By setting up a task scheduling system in the machine learning platform, analyzing task rules to generate a scheduling pipeline, and adopting the lazy delivery mode of "storage-notification", the problem of uncertainty and inflexibility of data transmission between task nodes in the existing technology is solved, and the flexibility and reliability of data transmission are achieved.
Patent Information
- Application Number
- CN202210574519.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-05-24
AI Technical Summary
There is uncertainty and inflexibility in data transmission between task nodes in existing machine learning platforms, and it is difficult to effectively coordinate data transmission in distributed training.
Set up a task scheduling system in the machine learning platform, generate a scheduling pipeline by analyzing task rules, scheduling executable task nodes and sending data, and adopting the lazy delivery mode of "storage-notification" to realize data stored in a specific location for the next task node to be pulled.
Through the coordination of the task scheduling system, the flexibility and reliability of data transmission are achieved, the uncertainty of data transmission is avoided, and the execution flexibility of task nodes is increased.
Smart Images

Figure CN114896066B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of data transfer processing, and more particularly to a method for processing data transfer between task nodes in a machine learning platform. Background Art
[0002] With the increasing popularity of artificial intelligence, machine learning has become increasingly mature. There is an endless stream of products for building machine learning platforms, and there are corresponding markets for those targeting different fields and being universal. In a machine learning platform, it is inevitable to disassemble training tasks according to a pipeline for distributed training. In existing machine learning platforms, for data transfer between nodes of training tasks, a single method is mostly adopted, either the node pulls the data over or the node is placed at the data location to work. In distributed training, data transfer between different nodes of the same task will become a problem that must be addressed. Summary of the Invention
[0003] Therefore, the embodiments of the present invention provide a method for processing data transfer between task nodes in a machine learning platform to solve the problems of uncertainty in data transfer between data nodes and inflexibility in data transfer between nodes in the prior art.
[0004] To achieve the above object, the embodiments of the present invention provide the following technical solution: A method for processing data transfer between task nodes in a machine learning platform, comprising the following steps:
[0005] Set up a task scheduling system, the task scheduling system:
[0006] Pull tasks from the task queue and parse the task rules of the tasks;
[0007] Generate a scheduling pipeline according to the task rules;
[0008] Send a start-up instruction according to the scheduling pipeline;
[0009] According to the start-up instruction, start up executable task nodes and send data to the executable task nodes;
[0010] After receiving the data, the executable task nodes execute the processing and return the generated data after the processing to the task scheduling system. At this time, the executable task nodes have completed the execution;
[0011] Receive the feedback that the executable task nodes have completed the execution, and repeat starting up other uncompleted task nodes until all task nodes in the scheduling pipeline have been completely executed;
[0012] Put the task results into the result queue.
[0013] Further, generating a scheduling pipeline according to the task rules specifically includes:
[0014] Determine the second storage location of the data generated by all task nodes.
[0015] Further, according to the invocation instruction, invoke executable task nodes and send data to the executable task nodes, specifically including:
[0016] Send an invocation instruction to invoke an executable task node, and send the data of the first storage location of the first processing data required by the executable task node and the data of the second storage location of the second processing data generated by the executable task node;
[0017] After receiving the invocation instruction, the executable task node pulls the first processing data from the first storage location for processing, and stores the second processing data generated after processing in the second storage location.
[0018] Further, the step of invoking executable task nodes according to the invocation instruction and sending data to the executable task nodes specifically includes:
[0019] Send an invocation instruction to invoke an executable task node, obtain the first processing data from the first storage location and send it to the executable task node; obtain the data of the second storage location where the executable task node generates the second processing data, and return the data of the second storage location to the task scheduling system.
[0020] Further, the step of invoking executable task nodes according to the invocation instruction and sending data to the executable task nodes specifically includes:
[0021] Invoke multiple task nodes that can be executed concurrently at one time.
[0022] Further, after the executable task node receives the data, it performs processing, specifically including:
[0023] Judge the execution capabilities of the multiple task nodes and select an execution method;
[0024] The execution methods include:
[0025] Perform data pulling;
[0026] Push the task node to the data storage location.
[0027] Further, generating a scheduling pipeline according to the task rules specifically includes:
[0028] Define the type of data in the scheduling pipeline;
[0029] Move the task node closer to the data.
[0030] The embodiments of the present invention have the following advantages: In an existing machine learning platform system, a task scheduling service system is established to record the process of tasks, issue tasks to task nodes, and coordinate the normal progress of tasks in the pipeline until completion. In terms of data transfer, a lazy transfer mode of "store - notify" is adopted, which realizes that data is not transferred from beginning to end, but stored at a specific location and pulled by the next task node, eliminating the uncertainty of data transfer. At the same time, it adds multiple choices for the execution of the next task. For example, whether to execute the task close to the data or pull the data down for use. Additionally, it increases the flexibility of task node startup, changing from passive to active. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained according to the provided drawings.
[0032] The structures, proportions, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have technical essence. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0033] Figure 1 It is a flowchart of a method for processing data transfer between task nodes in a machine learning platform provided by an embodiment of the present invention;
[0034] Figure 2 It is a schematic diagram of the relationship of a scheduling system for a method for processing data transfer between task nodes in a machine learning platform provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The following specific embodiments illustrate the implementation manners of the present invention. Those familiar with this technology can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope protected by the present invention.
[0036] Embodiment: A method for processing data transfer between task nodes in a machine learning platform, as follows Figure 1 and Figure 2 shown. Specifically, a task scheduling system is set up, and the processing method of the task scheduling system is as follows:
[0037] Pull the tasks in the task queue and parse the task rules of the tasks. First, pull the corresponding tasks from the task queue and automatically parse the processing rules of the tasks. In this embodiment, the processing rules of the tasks are included in the task types. When pulling the tasks, identify the task types and find the corresponding processing rules according to the task types. It is also possible to process according to the task rules included in the preset tasks, or traverse the task process to parse the task rules according to the process.
[0038] Generate a scheduling pipeline according to the task rules. From the above preset task rules, generate a scheduling pipeline. This scheduling pipeline calculates the data required for the task and determines each task node. This scheduling pipeline determines the storage locations of the data generated at each task node.
[0039] According to the invocation instruction, invoke the executable task nodes and send data to the executable task nodes. Here, multiple currently concurrently executable task nodes can be invoked at one time to increase the execution efficiency. The storage locations of the required data and the generated data are passed to each node as parameters.
[0040] Specifically: Send an invocation instruction to invoke the executable task nodes, and send the first storage location data of the first processing data required by the executable task nodes and the second storage location data of the second processing data generated by the executable task nodes;
[0041] It is also possible to, according to the invocation instruction, invoke the executable task nodes and send data to the executable task nodes, specifically including:
[0042] Send an invocation instruction to invoke the executable task nodes, obtain the first processing data from the first storage location and send it to the executable task nodes; obtain the second storage location data of the second processing data generated by the executable task nodes, and return the second storage location data to the task scheduling system.
[0043] After receiving the invocation instruction, the executable task nodes pull the first processing data from the first storage location for processing, and store the processed second processing data in the second storage location.
[0044] After receiving the data, the executable task nodes execute the processing and return the processed generated data to the task scheduling system. At this time, the executable task nodes have completed the execution;
[0045] Accept the feedback that the executable task node has completed its execution, and repeatedly trigger other unfinished task nodes until all task nodes in the scheduling pipeline have been executed;
[0046] Put the task result into the result queue.
[0047] The triggered task selects an execution method according to its own capabilities, either pulls data to complete the work or pushes itself to the data storage location to execute the task;
[0048] The scheduling system waits for the notification that the node has completed its execution. When a node completes a task, it will receive the completion notification and the storage location of the data produced by the node to generate subsequent scheduling parameters;
[0049] Trigger tasks in the same way as above; repeatedly execute the task triggering until all the work in the scheduling pipeline is completed; finally, put the task result into the result queue.
[0050] In the existing machine learning platform system, the present invention establishes a task scheduling service system to be responsible for recording the task process, issuing tasks to task nodes, and coordinating the normal progress of tasks in the pipeline until the end. In terms of data transfer, the "store-notify" lazy transfer mode is adopted, realizing that data is not transferred from beginning to end, but stored in a specific location and pulled by the next task node, eliminating the uncertainty of data transfer. At the same time, it adds multiple choices for the execution of the next task, such as whether to execute the task close to the data or pull the data down for use, and further increases the flexibility of task node startup, changing from passive to active.
[0051] Although the present invention has been described in detail with general descriptions and specific embodiments above, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A method for processing data transfer between task nodes in a machine learning platform, characterized in that, it includes the following steps: Set up a task scheduling system, and the task scheduling system: Pull tasks from the task queue and parse the task rules of the tasks; Generate a scheduling pipeline according to the task rules; Send a start instruction according to the scheduling pipeline; According to the start instruction, start executable task nodes and send data to the executable task nodes; After receiving the data, the executable task nodes perform processing and return the generated data after processing to the task scheduling system. At this time, the executable task nodes have completed execution; Receive the feedback that the executable task nodes have completed execution, and repeat starting other uncompleted task nodes until all task nodes in the scheduling pipeline have been executed; Put the task results into the result queue.
2. The method for processing data transfer between task nodes in a machine learning platform according to claim 1, characterized in that: Generating a scheduling pipeline according to the task rules specifically includes: Determine the second storage locations of the data generated by all task nodes.
3. The method for processing data transfer between task nodes in a machine learning platform according to claim 2, characterized in that: According to the start instruction, start executable task nodes and send data to the executable task nodes, specifically including: Send a start instruction to start executable task nodes, and send the data of the first storage location of the first processing data required by the executable task nodes and the data of the second storage location of the second processing data generated by the executable task nodes; After receiving the start instruction, the executable task nodes pull the first processing data from the first storage location for processing and store the generated second processing data in the second storage location.
4. The method for processing data transfer between task nodes in a machine learning platform according to claim 1, characterized in that: According to the start instruction, start executable task nodes and send data to the executable task nodes, specifically including: Send a start instruction to start executable task nodes, obtain the first processing data from the first storage location and send it to the executable task nodes; obtain the data of the second storage location of the second processing data generated by the executable task nodes and return the data of the second storage location to the task scheduling system.
5. The method for processing data transfer between task nodes in a machine learning platform according to claim 1, characterized in that: According to the start instruction, start executable task nodes and send data to the executable task nodes, specifically including: Start multiple task nodes that can be executed concurrently at one time.
6. The method for processing data transfer between task nodes in a machine learning platform according to claim 1, characterized in that: After receiving the data, the executable task nodes perform processing, specifically including: Judge the execution capabilities of multiple task nodes and select an execution method; The execution methods include: Perform data pulling; Push the task node to the data storage location.
7. The method for processing data transfer between task nodes in a machine learning platform according to claim 1, characterized in that: The generation of the scheduling pipeline according to the task rules specifically includes: Define the type of data in the scheduling pipeline; Move the task node closer to the data.
Citation Information
Patent Citations
Task scheduling method and system, computing device and readable storage medium
CN111949386A
Task scheduling method based on cloud computing technology
CN113238841A