Task processing method and system, model training task processing method, and computing device
By sending a request to the resource adjustment unit, determining the target data processing unit and executing the target task calculation graph, the problem of unsmooth task processing during computer resource adjustment is solved, and efficient resource utilization and smooth task execution are achieved.
Patent Information
- Application Number
- PCT/IB2025/053174
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-11
- Filing Date
- 2025-03-26
- Publication Date
- 2025-10-16
AI Technical Summary
During the computer resource adjustment process, existing technologies cannot ensure the smooth progress of task processing, resulting in a waste of computer resources and a failure to meet the actual needs of users.
By sending an operation resource adjustment request to the resource adjustment unit, the receiving unit determines the instruction, determines the target data processing unit, and executes the target task according to the target task calculation graph and processing resources, so as to achieve the matching of the target task calculation graph and the adjusted data processing unit.
It ensures the smooth execution of target tasks, avoids task processing failure caused by waste of computer resources and resource mismatch, and improves task processing efficiency.
Smart Images

Figure IB2025053174_16102025_PF_FP_ABST
Abstract
Description
[0001] The present disclosure relates to the technical field of computer technology, and particularly relates to a task processing method. Background art With the continuous development of computer technology, computer resources can be utilized in the process of task processing of a user's task. However, in the process of task processing by utilizing computer resources, the computer resources required for executing the task cannot be accurately estimated, and therefore the computer resources need to be adjusted. In the prior art, after the computer resources are adjusted, the changed computer resources do not match the calculation graph corresponding to the task, and therefore the task cannot be processed based on the calculation graph, which leads to the problems of waste of computer resources and failure to meet the actual demand of the user. Therefore, there is an urgent need to provide a method capable of ensuring smooth task processing while adjusting the computer resources. In view of this, the embodiments of the present disclosure provide a task processing method. One or more embodiments of the present disclosure also provide a model training task processing method, and a computing device. 1The application discloses a task processing method, a task processing device, a model training task processing device, another task processing device, a computing device, a computer readable storage medium and a computer program product, and aims at solving the technical problem that the prior art cannot ensure smooth task processing while adjusting computer resources. According to a first aspect of an embodiment of the present disclosure, a task processing method is provided, which comprises the following steps: sending a running resource adjustment request to a resource adjustment unit, wherein the running resource adjustment request is sent in a process of executing a target task, the target task is executed according to an initial task computing graph and processing resources of an initial data processing unit, and the initial task computing graph is established according to task data and task operation information carried in the target task; receiving a unit determination instruction sent by the resource adjustment unit, wherein the unit determination instruction is sent by the resource adjustment unit in response to the running resource adjustment request and in a case that a target data processing unit is determined, and the target data processing unit is obtained by adjusting the processing resources of the initial data processing unit by the resource adjustment unit; determining the target data processing unit based on the unit determination instruction, and determining a target task computing graph according to processing resources of the target data processing unit and the initial task computing graph; and executing the target task according to the target task computing graph and the processing resources of the target data processing unit. According to a second aspect of an embodiment of the present disclosure, a task processing device is provided, which comprises: a request sending module configured to send a running resource adjustment request to a resource adjustment unit, wherein the running resource adjustment request is sent in a process of executing a target task, the target task is executed according to an initial task computing graph and processing resources of an initial data processing unit, and the initial task computing graph is established according to task data and task operation information carried in the target task; an instruction receiving module configured to receive a unit determination instruction sent by the resource adjustment unit, wherein the unit determination instruction is sent by the resource adjustment unit in response to the running resource adjustment request and in a case that a target data processing unit is determined, and the target data processing unit is obtained by adjusting the processing resources of the initial data processing unit by the resource adjustment unit; a computing graph determining module configured to determine the target data processing unit based on the unit determination instruction, and determine a target task computing graph according to processing resources of the target data processing unit and the initial task computing graph; and a task executing module configured to execute the target task according to the target task computing graph and the processing resources of the target data processing unit.According to a third aspect of embodiments of the present disclosure, another task processing method is provided, including: determining current running resources of an initial data processing unit in response to a running resource adjustment request sent by a task computing unit, wherein the running resource adjustment request is sent in a process in which the task computing unit executes a target task, the target task is executed by the task computing unit according to an initial task computing graph and processing resources of the initial data processing unit, and the initial task computing graph is constructed by the task computing unit according to task data and task operation information carried in the target task; determining a data processing unit adjustment strategy according to the current running resources of the initial data processing unit; adjusting processing resources of the initial data processing unit according to the data processing unit adjustment strategy to obtain a target data processing unit, so that the task computing unit executes the target task according to processing resources of the target data processing unit and a target task computing graph, wherein the target task computing graph is determined according to the processing resources of the target data processing unit and the initial task computing graph. According to a fourth aspect of embodiments of the present disclosure, another task processing apparatus is provided, including: a request response module configured to determine current running resources of an initial data processing unit in response to a running resource adjustment request sent by a task computing unit, wherein the running resource adjustment request is sent in a process in which the task computing unit executes a target task, the target task is executed by the task computing unit according to an initial task computing graph and processing resources of the initial data processing unit, and the initial task computing graph is constructed by the task computing unit according to task data and task operation information carried in the target task; a strategy determination module configured to determine a data processing unit adjustment strategy according to the current running resources of the initial data processing unit; and a unit adjustment module configured to adjust processing resources of the initial data processing unit according to the data processing unit adjustment strategy to obtain a target data processing unit, so that the task computing unit executes the target task according to processing resources of the target data processing unit and a target task computing graph, wherein the target task computing graph is determined according to the processing resources of the target data processing unit and the initial task computing graph.According to a fifth aspect of embodiments of the present disclosure, a model training task processing method is provided, including: sending a running resource adjustment request to a resource adjustment unit, wherein the running resource adjustment request is sent in a process of executing a model training task, the model training task is executed according to an initial task computation graph and processing resources of an initial data processing unit, and the initial task computation graph is constructed according to task data and task operation information carried in the model training task; receiving a unit determination instruction sent by the resource adjustment unit, wherein the unit determination instruction is sent by the resource adjustment unit in a case of determining a target data processing unit in response to the running resource adjustment request, and the target data processing unit is obtained by adjusting the processing resources of the initial data processing unit by the resource adjustment unit; determining the target data processing unit based on the unit determination instruction, and determining a target task computation graph according to processing resources of the target data processing unit and the initial task computation graph; and executing the model training task according to the target task computation graph and the processing resources of the target data processing unit. According to a sixth aspect of embodiments of the present disclosure, a model training task processing apparatus is provided, including: a request sending module configured to send a running resource adjustment request to a resource adjustment unit, wherein the running resource adjustment request is sent in a process of executing a model training task, the model training task is executed according to an initial task computation graph and processing resources of an initial data processing unit, and the initial task computation graph is constructed according to task data and task operation information carried in the model training task; an instruction receiving module configured to receive a unit determination instruction sent by the resource adjustment unit, wherein the unit determination instruction is sent by the resource adjustment unit in a case of determining a target data processing unit in response to the running resource adjustment request, and the target data processing unit is obtained by adjusting the processing resources of the initial data processing unit by the resource adjustment unit; a computation graph determining module configured to determine the target data processing unit based on the unit determination instruction, and determine a target task computation graph according to processing resources of the target data processing unit and the initial task computation graph; and a task executing module configured to execute the model training task according to the target task computation graph and the processing resources of the target data processing unit.According to a seventh aspect of the embodiments of the present disclosure, a task processing system is provided, which comprises a task computing unit, an initial data processing unit and a resource adjusting unit. The task computing unit is configured to send a running resource adjusting request to the resource adjusting unit, wherein the running resource adjusting request is sent in the process of executing a target task, the target task is executed according to an initial task computing graph and processing resources of the initial data processing unit, and the initial task computing graph is constructed according to task data and task operation information carried in the target task. The task computing unit receives a unit determination instruction sent by the resource adjusting unit, wherein the unit determination instruction is sent by the resource adjusting unit in response to the running resource adjusting request and in a case that the target data processing unit is determined, and the target data processing unit is obtained by the resource adjusting unit adjusting the processing resources of the initial data processing unit. The task computing unit determines the target data processing unit based on the unit determination instruction, determines a target task computing graph according to the processing resources of the target data processing unit and the initial task computing graph, and executes the target task according to the target task computing graph and the processing resources of the target data processing unit. The resource adjusting unit is configured to determine current running resources of the initial data processing unit in response to the running resource adjusting request sent by the task computing unit, wherein the running resource adjusting request is sent by the task computing unit in the process of executing the target task, the target task is executed by the task computing unit according to the initial task computing graph and the processing resources of the initial data processing unit, and the initial task computing graph is constructed by the task computing unit according to task data and task operation information carried in the target task. The resource adjusting unit determines a data processing unit adjusting strategy according to the current running resources of the initial data processing unit, adjusts the processing resources of the initial data processing unit according to the data processing unit adjusting strategy, obtains the target data processing unit, so that the task computing unit executes the target task according to the processing resources of the target data processing unit and a target task computing graph, wherein the target task computing graph is determined according to the processing resources of the target data processing unit and the initial task computing graph. According to an eighth aspect of the embodiments of the present disclosure, a computing device is provided, which comprises a memory and a processor. The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned task processing method, another task processing method or a model training task processing method are implemented.According to a ninth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which stores computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the steps of the task processing method, another task processing method, or a model training task processing method. According to a tenth aspect of an embodiment of the present disclosure, a computer program product is provided, which includes computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the steps of the task processing method, another task processing method, or a model training task processing method. The task processing method in one or more embodiments of the present disclosure can send a running resource adjustment request to a resource adjustment unit in the process of executing a target task according to an initial task computation graph and the initial data processing unit, and in the case that the resource adjustment unit obtains a target data processing unit by processing the initial data processing unit according to the running resource adjustment request, determine a target task computation graph according to the processing resource of the target data processing unit and the initial task computation graph, thereby overcoming the problem of mismatch between the target task computation graph and the adjusted target data processing unit. Then, the task computation unit successfully executes the target task according to the target task computation graph and the target data processing unit, thereby ensuring the smooth execution of the target task while adjusting the data processing unit, avoiding the problem of computer resource waste caused by the mismatch between the computation graph and the adjusted computer resource, and the problem of being unable to meet the actual needs of users due to the inability to process the target task.BRIEF DESCRIPTION OF DRAWINGS FIG. 1 is an application diagram of a task processing method according to an embodiment of the present disclosure; FIG. 2 is a flowchart of a task processing method according to an embodiment of the present disclosure; FIG. 3 is a diagram of a training system in a task processing method according to an embodiment of the present disclosure; FIG. 4 is a diagram of creating a subgraph for a parameter in a task processing method according to an embodiment of the present disclosure; FIG. 5 is a flowchart of another task processing method according to an embodiment of the present disclosure; FIG. 6 is a diagram of a state machine in another task processing method according to an embodiment of the present disclosure; FIG. 7 is a diagram of the interaction of various modules in a task processing system in another task processing method according to an embodiment of the present disclosure; FIG. 8 is a diagram of a Gazer module in another task processing method according to an embodiment of the present disclosure; FIG. 9 is a diagram of scaling out in another task processing method according to an embodiment of the present disclosure; FIG. 10 is a diagram of scaling in in another task processing method according to an embodiment of the present disclosure; FIG. 11 is a flowchart of a processing procedure of a task processing method according to an embodiment of the present disclosure; FIG. 12 is a flowchart of a model training task processing method according to an embodiment of the present disclosure; FIG. 13 is a diagram of the structure of a task processing system according to an embodiment of the present disclosure; and FIG. 14 is a block diagram of the structure of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, the present disclosure can be practiced without the specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the present disclosure. The following description is not intended to limit the present disclosure, but to provide an appreciation of possible embodiments of the present disclosure. The terminology used in the description of the embodiments herein is not intended to be limiting in scope but is intended to help the reading person become familiar with the present disclosure. The use of "including", "comprising", or "having" and variations thereof in the description of the embodiments herein is intended to be inclusive and not exclusive. The use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive. It should be understood that the use of "and / or" in the description of the embodiments herein is intended to be inclusive and not exclusive.Depending on the context, the word "if" as used herein can be interpreted to mean "when" or "while" or "in response to determining." In addition, it is important to note that user information (including, but not limited to, user device information, user personal information, etc.) and data (including, but not limited to, data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by users or authorized by all parties, and the collection, use, and processing of related data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and provide appropriate operation portals for users to choose authorization or refusal. In one or more embodiments of the present disclosure, a large model refers to a deep learning model with a large number of model parameters, usually containing hundreds of millions, tens of billions, hundreds of billions, thousands of billions, or even tens of billions of model parameters. The large model can also be referred to as a foundation model. Through large-scale unlabeled corpus pre-training of the large model, a pre-training model with hundreds of millions of parameters is output. Such a model can adapt to a wide range of downstream tasks, and the model has good generalization ability. Examples of large models include large language models (LLM) and multi-modal pre-training models. In actual applications, the large model only needs a small amount of samples to fine-tune the pre-training model and can be applied to different tasks. The large model can be widely used in natural language processing (NLP) and computer vision fields. Specifically, the large model can be applied to computer vision field tasks such as visual question answering (VQA), image captioning (IC), and image generation, and natural language processing field tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of the large model include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc. Based on this, the model in one or more embodiments of the present disclosure can be understood as a large model, and the model training task can be understood as a task of training the large model. First, the nomenclature involved in one or more embodiments of the present disclosure is explained.
[0002] Tensorflow: A deep learning framework. Compute graph: A graph structure in computer science used in deep learning to represent the computations of a model. The nodes in the graph generally represent computational operations, and the edges represent the data flowing between the operations.
[0003] Parameter Server: PS for short, refers to the parameter server, a training mode in deep learning. It also usually refers to the character that stores the model parameters in training. Elastic training: refers to the ability to quickly expand or reduce computing processing, memory and storage resources to meet changing needs without worrying about capacity planning and engineering for peak usage.
[0004] RPC: refers to the RPC protocol, a protocol that requests services from remote computer programs over the network without understanding the underlying network technology. Global_step: represents the global step number, such as at what step should the operation be performed, how many rounds the neural network has been trained to, etc. It is similar to a clock.
[0005] PPU: PPU for short, refers to the physical processing unit.
[0006] Worker: can refer to the role of completing model computation in tensorflow.
[0007] Elastic gRPC server: refers to a gRPC server designed with elastic scaling features, which can dynamically adjust the number of instances and service capabilities to adapt to changing load demands.
[0008] DFS: Distributed File System is a file system that allows files to be shared across multiple hosts over a network, allowing multiple users on multiple machines to share files and storage space. Checkpoint: refers to the checkpoint, which captures and saves the state of a task running, so that the task can be restarted at that state.
[0009] CRD (CustomResourceDefinition) is a core feature of Kuber netes, which allows users to extend the Kubernetes API to create new resource types. With the continuous development of computer technology, it is necessary to utilize its own resources in the process of computer task processing, such as storage resources, computing resources, etc. For example, in a training task of a Tensorflow using Parameter Server mode, users need to manually allocate parameters in the computation graph to heterogeneous devices (CPU / GPU / PPU) and estimate the reasonable worker / PS resource number according to the allocation of parameters. Since users do not understand how the distributed computing of the model is actually executed on each device, it is necessary for users to have high background knowledge to plan a training task that matches the application resources and the used resources (i.e. the resources actually used by the training task). In the mainstream model, the parameters in the computation graph can be divided into two categories: dense parameters and sparse parameters. The characteristics of sparse parameters are large in size, generally GB~TB, and often account for more than 90% of the memory usage of the entire model. In the processing scheme for this sparse parameter, since its resource requirement is dynamically changing, it further increases the difficulty of user resource estimation. Therefore, an ideal training task needs to meet two goals: 1. high training throughput; 2. high resource utilization. However, if the resource allocation is not appropriate, it will affect the normal operation (OOM) of the task or waste resources. If the user does not properly allocate the parameters in the computation graph, it will lead to uneven resource utilization of different PSs in part of the task, which not only directly leads to the waste of resources of part of the PS, but also affects the performance of the training task to some extent because of the concentration of network connections on the hot PS; and another type of uncontrollable reason is that in the process of actual running of the user task, there may be slow machine nodes due to the reasons of the cluster environment, which will affect the training speed of the task. Based on this, the present disclosure provides a solution to the resource application and utilization problem of the Tensorflow distributed training task in the Parameter Server mode, which has been a pain point for users to expand distributed tasks.In order to reduce the difficulty of resource allocation for distributed training of small user groups, the scheme starts a management container for each Tensorflow task, and the functions of the container include: 1. Start the container of the entire Tensorflow task; 2. Process the data collected by the container of each Tensorflow task; 3. Monitor the state of each Tensorflow container; 4. Update the CRD of the cluster according to the updated PS resource number. In addition, when the container of each Tensorflow task (worker, PS) is started, a training process is started, which is the parent process of the actual Tensorflow task and is responsible for collecting the CPU and memory usage of the machine, whether the process exits abnormally, and reporting to the management container. During the running of the Tensorflow task, the management container processes the state of all containers in an event-driven manner, and continuously judges whether to need to expand or shrink the container according to the current task state. If an expansion event is received, the management will start the best PS instance of the current task, then let all tensorflow containers save the checkpoint, and then restart the container, and only when all containers are ready will the previous checkpoint be restored to continue training. However, the scheme has a big defect. Based on the elastic training of the checkpoint, in order to preserve the state of the current training task 3 parameters, it is necessary to complete the saving of the checkpoint, the restart of the process and the loading of the checkpoint. The save and restore of the checkpoint involve two long-time I / O (data transmission) of the file system, which is a redundant process and will cause low task processing efficiency. Based on this, in the present disclosure, a task processing method is provided, and the present disclosure simultaneously relates to a model training task processing method, another task processing method, a task processing device, a model training task processing device, another task processing device, a computing device, a computer readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.Referring to FIG. 1, FIG. 1 shows an application diagram of a task processing method according to one embodiment of the present disclosure. Based on FIG. 1, a user can send a model training task to a server 104 through a terminal 102. The server 104 is configured with a training system. A DLC service in the training system can receive the model training task sent by the user and forward the model training task to a kubeDL in the training system for scheduling. The kubeDL allocates a container set for the model training task and sends the model training task to the container set for processing. The container set includes a master container, a plurality of PS containers, and a plurality of worker containers. After receiving the model training task, the master container in the container set allocates model parameters to each PS container and each worker container according to the resource conditions of the plurality of PS containers and the plurality of worker containers. oThe worker container constructs a computation graph of the model training task based on the model parameters and executes, and stores data in the execution process to the PS container. During the execution of the model training task, the worker container sends a scaling request for the PS container to the Master container. When the Master container scales the PS container based on the scaling request, the number of PS containers changes due to the scaling operation. Therefore, after the PS scaling or the PS scaling, the worker container needs to use the "on-demand parameter redistribution mechanism" and the "dynamic rewriting of the computation graph" to optimize and adjust the computation graph. Then the optimized computation graph is executed to implement the execution of the model training task. After completing the model training task and obtaining the trained model, the Master container sends the trained model to the kubeDLo, and the kubeDL sends the trained model to the terminal 102 through the DLC service, thereby providing the trained model to the user. The problem that the computer resources cannot be reasonably allocated for task processing due to different resources required by different tasks and the inability to accurately estimate the required resources during task processing is avoided, the processing efficiency of the task is further improved, and the problems of waste of computer resources caused by the mismatch between the task computation graph and the adjusted PS container and the inability to perform model training tasks to meet the actual needs of the user are avoided. Referring to FIG. 2, FIG. 2 shows a flowchart of a task processing method according to an embodiment of the present disclosure, which can be applied to a task computing unit in a task processing system. The task computing unit can be in communication connection with an initial data processing unit and / or a resource adjusting unit in the task processing system, and specifically includes the following steps. Step 202: Send a running resource adjustment request to the resource adjusting unit, wherein the running resource adjustment request is sent during the execution of a target task, the target task is executed according to an initial task computation graph and processing resources of the initial data processing unit, and the initial task computation graph is constructed according to task data and operation information carried in the target task. The task computing unit can be understood as a unit capable of computing and processing the target task, for example, the task computing unit can be a virtual server, a virtual machine (VM), a physical server, a worker container, a container, etc. That is, the task computing unit can be understood as a server, a container, etc. capable of computing and processing the target task.In one or more embodiments provided in the present disclosure, the task processing method is explained by taking the task computing unit as a container. That is, the task computing unit can be a task computing container. For the implementation mode of the task computing unit being a virtual server, a virtual machine (VM), or a physical server, refer to the related content in the specification by taking the task computing unit as a container, which will not be repeated here. The task computing unit can be one or more. The data processing unit can be understood as a unit capable of processing task data of a target task, for example, the data processing unit can be a virtual server, a virtual machine (VM), a physical server, a PS, a PS container, a container, etc. That is, the data processing unit can be understood as a PS, a container, etc. capable of processing task data of a target task. In one or more embodiments provided in the present disclosure, the task processing method is explained by taking the data processing unit as a container. That is, the initial data processing unit can be an initial data processing container. For the implementation mode of the data processing unit being a virtual server, a virtual machine (VM), a physical server, a PS, or a PS container, refer to the related content in the specification by taking the data processing unit as a container, which will not be repeated here. The data processing unit can be one or more. In the case of the data processing unit being one, the initial data processing unit is the data processing unit. In the case of the data processing unit being multiple, the initial data processing unit can be any one of the multiple data processing units. For the explanation of the data processing units other than the initial data processing unit, refer to the corresponding or corresponding content in the explanation of the initial data processing unit, which will not be specifically limited here. In one or more embodiments provided in the present disclosure, the task computing unit can be a task computing unit in a task processing system, and the task processing system includes a task computing unit, a data processing unit, a resource adjusting unit, and a resource allocating unit. The resource adjusting unit can be understood as a node for managing the data processing unit and the task computing unit. The management of the task computing unit includes but is not limited to: determining the current running resource of the initial data processing unit in response to a running resource adjustment request sent by the task computing unit; determining a data processing unit adjustment strategy according to the current running resource of the initial data processing unit; adjusting the processing resource of the initial data processing unit according to the data processing unit adjustment strategy to obtain a target data processing unit. The resource adjusting unit can be a virtual server, a virtual machine (VM), a physical server, or a container, etc.For example, in the case of a resource adjustment unit as a container, the resource adjustment unit can be understood as a resource adjustment container for managing the task calculation unit and the data processing unit. It should be noted that in the case of a resource adjustment unit as a container, the resource adjustment container can be referred to as a Master container. The resource allocation unit can be understood as a unit for managing unit creation resources in the task processing system. The unit creation resources in the task processing system can be understood as device resources, which can be resources of entity servers, host computers, terminals and the like in the task processing system, for example, the device resources can be memory resources, CPU resources, GPU resources, video memory resources and the like of the device. The management of the unit creation resources in the task processing system can be understood as the allocation of the purpose of the unit creation resources in the task processing system, for example, the computing or storage resources of the entity servers in the task processing system are used as devices supporting the implementation of the container, or the devices in the task processing system are set to be available or unavailable, etc. The resource allocation unit can be a virtual server, a virtual machine (VM), a physical server, or a container, etc. In one or more embodiments provided by the present disclosure, in the case of a data processing unit as a container, the device resources in the task processing system can be an entity server for implementing the container. Alternatively, in one or more embodiments provided by the present disclosure, in the case of a data processing unit as an entity server, the device resources in the task processing system can be the entity server, which is subsequently allocated to the resource adjustment unit for management, and used for processing the target task. The running resource adjustment request can be understood as a request for adjusting the processing resources of the initial data processing unit. In one or more embodiments provided by the present disclosure, since the data processing unit may be in a state of high load or idling during the execution of the target task, in order to avoid the problem of low target task processing efficiency caused by the high load of the data processing unit, or the problem of insufficient resource utilization rate caused by the idling of the data processing unit, the processing resources of the data processing unit need to be adjusted through the running resource adjustment request. In one or more embodiments provided by the present disclosure, the running resource adjustment request can be a scaling request, that is, the running resource adjustment request can be a scaling request or a scaling request for the initial data processing unit. The resource adjustment module judges whether to perform scaling processing or scaling processing on the data processing unit by responding to the scaling request sent by the task calculation unit.In one or more embodiments provided in the present disclosure, the running resource adjustment request can be sent by the task computing unit according to a preset sending rule in the process of executing the target task. The preset sending rule can be understood as a rule triggering the task computing unit to send the running resource adjustment request to the resource adjustment unit. The preset sending rule can be set according to the actual application scenario. For example, the preset sending rule can be to send the running resource adjustment request at a specific time frequency, such as once every 1 minute. Alternatively, the preset sending rule can be a rule that the task computing unit sends the running resource adjustment request to the resource adjustment unit when the initial data processing unit has a high load or is idle, without specific limitation. The current running resource can be understood as a performance analysis indicator representing the current performance of the initial data processing unit, for example, the current running resource includes but is not limited to memory utilization, CPU utilization, GPU utilization, and video memory utilization, without specific limitation. In one or more embodiments provided in the present disclosure, the task processing system provides a resource detection module of Tensorflow (the resource detection module can be a data collection plug-in, which can be referred to as a Gazer module), which can support collecting different performance analysis indicators (Profiler Metrics) of each role (data processing unit or task computing unit) in the process of performing the target task, so as to obtain the fine-grained state of the task computing unit in the running process. The processing resource can be understood as a resource of the data processing unit (initial data processing unit or target data processing unit) used for executing the target task, for example, the processing resource can be a storage resource (such as memory resource, cache resource, and external storage resource) of the data processing unit, a computing resource (such as CPU computing resource and GPU computing resource), and a data transmission resource (such as bus resource). The initial task computing graph can be understood as a computing graph generated based on the task data of the target task and the task operation information of the task data. The target task can be understood as a task that needs to be processed by the task computing unit and the initial data processing unit. The target task can be set according to the actual application scenario. In one or more embodiments provided in the present disclosure, the target task can be a model training task, that is, a task of training a neural network model. The neural network model can be a large model, a language inference model, an image processing model, or a data analysis model, etc. It should be noted that when the task processing method is applied to different model training scenarios, the model training task is also different.For example, in the case where the task processing method is applied to an image processing model training scenario, the model training task can be an image processing model training task for training an image processing model, which can be a graph-to-text model for realizing graph-to-text, a face recognition model for face recognition based on a face image, etc. In the case where the task processing method is applied to a text processing model training scenario, the model training task can be a text processing model training task for training a text processing model, which can be a text-to-graph model for realizing text-to-graph, a language reasoning model, etc. In the case where the task processing method is applied to a speech processing model training scenario, the model training task can be a speech processing model training task for training a speech processing model, which can be a speech analysis model, a model for converting speech into text, etc. It should be noted that the above neural network models can all be large models. In one or more embodiments provided by the present disclosure, the above neural network models are not specifically limited. In one or more embodiments provided by the present disclosure, the target task can also be an image processing task (for example, a face recognition task for recognizing a face from an image, an object recognition task for recognizing an object in an image, a text-to-graph task for generating an image according to a text description, etc.), a language reasoning task (for example, a question and answer task for reasoning a corresponding answer according to a question raised by a user, etc.), a data analysis task, etc., and the target task is not specifically limited herein. The task data can be understood as data required to perform the target task. The task operation information of the task data can be understood as information representing what operation is performed on the task data. By processing the task data based on the task operation information, the purpose of performing the target task is achieved. For example, in the case where the target task is a model training task, the task data can be training sample data, model parameters (for example, gradient information, loss value, etc.), etc. The task operation information of the task data can be data preprocessing (for example, de-duplication, format adjustment, etc.) on the training sample data, feature extraction on the training sample data, etc., or model parameter adjustment based on the model parameters, etc. For another example, in the case where the target task is a face recognition task, the task data can be face image data, etc. The task operation information of the task data can be face image feature extraction operation on the face image, operation of inputting the face image feature into a face recognition model, etc. The task computation graph can be understood as a computation graph for the target task.The high-availability adaptive elastic training deep learning system (hereinafter referred to as a training system) can be a task processing system, and the training system includes a master container (a container for managing containers), a plurality of worker containers, and a plurality of PS containers. The task processing method can be applied to the worker container. Referring to FIG. 3, FIG. 3 is a schematic diagram of a training system in a task processing method according to an embodiment of the present disclosure. As shown in FIG. 3, the training system is a system including a DLC backend and a k8s cluster, and the training system includes a DLC backend and a plurality of containers, and a master container for managing the plurality of containers. The DLC service in the DLC backend refers to a backend service for submitting a job to the k8s cluster, the kubeDL is a resource controller of k8s, responsible for scheduling and coordination of deep learning, and the kubeDL can be a resource management node for resource management (e.g., server allocation, etc.) in the task processing system. The master container includes a resource detection module (Gazer), an elastic training controller (Elastic Training controller), and an elastic training service (Elastic Training service); the resource detection module is used to collect performance analysis reports of different roles (PS, worker), so as to realize the ability to obtain the state of the computing graph in a fine-grained manner in each role; in one or more embodiments of the present disclosure, the fine-grained state can be understood as the processing of task data in the computing graph by the PS container and the worker container during the execution of the target task based on the computing graph, such as feature extraction operations, convolution operations, etc.; correspondingly, the fine-grained state can be understood as the state of each role during the execution of the processing operation of the task data, for example, the fine-grained state can be the CPU and memory usage, operation time, etc. of each role. The elastic training controller (Elastic Training controller) is a controller for elastic model training. The elastic training service (Elastic Training service) can provide elastic scaling service. Based on this, during application, the DLC service in the training system can receive a model training task sent by a user, and send the model training task to the kubeDL for scheduling.The kubeDL assigns a container set for processing a model training task, and sends the model training task to the container set. The container set includes one master container, multiple worker containers, and multiple PS containers, and the master container is used to manage the PS containers and the worker containers. In addition, the Elastic Grpc server in Embodiment 3 is a Grpc server with elastic scaling characteristics. After receiving the model training task, the elastic training controller in the master container first determines the resource information carried in the model training task, so as to determine how many resources are needed to execute the model training task, and allocates the task resources to the worker containers and the PS containers for execution. The worker containers construct a computation graph of the model training task according to the model parameters carried in the model training task as edges, and the task operation information for the model parameters as nodes, and generate a corresponding session for the computation graph, and subsequently execute the computation graph through the session. Finally, the worker containers and the PS containers implement processing of the model training task through the computation graph and the session, thereby training the neural network model. It should be noted that in the process of processing the model training task by the worker containers and the PS containers, the Elastic Gazer module in the elastic training system can collect the Tensorflow task training information and report to the master container. In addition, it should be noted that the worker containers can send a request of sealing to the elastic training service (Elastic Training service) for the PS containers. Based on the request of sealing, if the PS containers need to be scaled, the elastic training service determines a scaling strategy through an elastic strategy module in the elastic training controller, and scales the PS containers (i.e., updates the number of PS containers, for example, adds new PS containers, or deletes idle PS containers) based on the scaling strategy. In the design of the Tensorflow training system, the computation graph, the parameters (Variable) in the computation graph, and the corresponding session (Session) are all constructed in advance (static).After the container 9 is completed, the entire model training task needs to be restarted in the elastic scaling of Tensorflow, and the saving and loading recovery of the computation graph parameters are needed, which cannot be ignored for a normal training of the Tensorflow task. In view of this problem, the task processing method in one or more embodiments of the present disclosure provides an efficient solution, and the worker container performs computation graph optimization through a "parameter on-demand redistribution mechanism" and a "dynamic rewriting of the computation graph", so as to reduce the additional system overhead caused by dynamically adjusting the number of PS. The "parameter on-demand redistribution mechanism" and the "dynamic rewriting of the computation graph" are realized through an elastic training hook (Elastic Training Hook), and specific contents are described in one or more embodiments below. In one or more embodiments provided by the present disclosure, before the resource adjustment unit sends a running resource adjustment request, the target task sent by the resource adjustment unit is received, and an initial task computation graph is constructed according to the task data and the task operation information carried in the target task, and the target task is executed according to the initial task computation graph and the processing resource of the initial data processing unit. Wherein, the initial task computation graph can be constructed by taking the task data as an edge and the task operation information for the task data as a node. Based on this, in the task processing method in one or more embodiments of the present disclosure, the task computation unit can receive the target task sent by the resource adjustment unit. After obtaining the target task, the task data and the task operation information of the task data carried in the target task are determined. Then, the initial task computation graph is constructed according to the task data and the task operation information. After the construction of the initial task computation graph is completed, the target task can be executed according to the initial task computation graph and the initial data processing unit, so as to meet the actual needs of the user by executing the target task. In the above example, the task processing system can be a model training task processing system (i.e., the training system described above), the task computation unit can be a worker container (i.e., a container for completing model computation in the model training task process), the data processing unit can be a PS container (i.e., a container serving as a PS in the model training task process), and the resource adjustment unit can be a master container.Based on this, the Master container, after receiving the model training task, allocates model parameters to the RS containers and the worker containers according to the resource conditions of the PS containers and the worker containers. The worker container constructs a computation graph of the model training task based on the model parameters and performs model calculation, and stores data and model parameters in the process of performing model calculation to the PS container; the PS container can store the data and model parameters in the process of performing model calculation to the memory, cache and / or external storage through the CPU, data transmission line (i.e. processing resource). Step 204: receiving a unit determination instruction sent by the resource adjustment unit, wherein the unit determination instruction is sent by the resource adjustment unit in response to the running resource adjustment request, and the target data processing unit is obtained by adjusting the processing resource of the initial data processing unit by the resource adjustment unit. Wherein, the target data processing unit can be understood as the data processing unit obtained by adjusting the processing resource of the initial data processing unit by the resource adjustment unit in response to the running resource adjustment request; the adjustment of the processing resource of the initial data processing unit can be understood as increasing the processing resource of the initial data processing unit or reducing the processing resource of the initial data processing unit, for example, expanding the initial data processing unit or reducing the initial data processing unit. In the above example, the worker container will send an expansion / reduction request for the PS container to the Master container in the process of performing the model training task, and when the Master container expands / reduces the PS container based on the expansion / reduction request, the worker container can determine the PS container after expansion / reduction. Step 206: determining the target data processing unit based on the unit determination instruction, and determining the target task computation graph according to the processing resource of the target data processing unit and the initial task computation graph. Wherein, the target task computation graph can be understood as a task computation graph adapted to the target data processing unit, because the target data processing unit is needed to perform the target task, in order to avoid problems such as target task execution error, unable to perform task processing, etc. caused by the inadaptation of the task computation graph to the target data processing unit, the target task computation graph can be determined according to the processing resource of the target data processing unit and the initial task computation graph.In one or more embodiments provided in the present disclosure, the determining of the target task computation graph according to the processing resources of the target data processing unit and the initial task computation graph comprises steps one to three: step one: dividing the task data into a plurality of task sub-data according to the processing resources of the target data processing unit. Wherein, the task sub-data can be understood as a plurality of sub-data obtained by dividing and processing the task data. Following the above example, the task processing system in one or more embodiments of the present disclosure can be a flexible training system (i.e., a highly available adaptive flexible training deep learning system), which comprises: a Gazer module for collecting Tensorflow training indicators and reporting to a Master container, a Master container for controlling the flexible expansion of Tensorflow jobs, and a Tensorflow task internal flexible expansion self-adaptive computation graph.
[0010] an ElasticTrainingHook module and a module for dynamically rewriting and modifying the computation graph. The ElasticTrainingHook module and the module for dynamically rewriting and modifying the computation graph can be disposed in a worker container (which can be simply referred to as a worker) and a PS container (which can be simply referred to as a PS). In the task processing method provided in the present disclosure, in order to realize flexible training for the Tensorflow task (target task), the following functions need to be completed: the re-graphing of the computation graph after the adjustment of the PS module, the distribution of the parameter quantity on different PSs, and the function of regularly interacting with the Master container. The corresponding function implementation is based on the
[0011] The SessionRunHook is implemented in the E I ast i cTra i n i ngHook and the E I ast i cTra i n i ngPass. The SessionRunHook is used to register a hook function in the session (Mon i tored Tra i n i ng Sess i on) of the tensorf I ow calculation, and the hook function will construct the subgraph (i.e., the pre-designed operator graph) related to the "graph parameter redistribution" in advance for all the computation graph parameters (i.e., the model parameters for constructing the computation graph) in the computation graph (i.e., the task computation graph) when the session starts. Referring to FIG. 4, which is a schematic diagram of creating a subgraph for a parameter in an embodiment of the task processing method provided by the present disclosure, the computation graph parameter constructs the subgraph in the following steps: first, it is determined whether the variable (i.e., the computation graph parameter) is an Embedd i ngVar i ab I e. If yes, the Embedd i ngVar i ab I e redistribution subgraph is constructed. If no, it is determined whether the variable is a distributed variable. If yes, the redistribution subgraph of the ordinary variable is constructed. If no, the subgraph of the ordinary variable transfer is constructed. Since the number of roles in the Tensorf l ow distributed task has changed, the computation graph for task running also needs to be changed accordingly. However, since the computation graph submitted by the user is fixed, it is necessary to support dynamic modification of the computation graph. Therefore, in the E I ast i cTra i n i ngPass, the computation graph is supported to dynamically modify the subgraph for the expanded PS number. In the Tensorf l ow, according to the different use characteristics of the variables in the computation graph, the variables can be divided into the following types:
[0012] 1. Embedd i ngVar i ab I e; 2. densevar i ab I e; 3. sparsevar i ab I e oBy analyzing, it is found that the corresponding forward and backward calculation subgraphs of each of them are fixed, so the rewriting rules of their corresponding fixed digraph matching are realized. Finally, there is a problem of how to determine the appropriate number of sharding of each variable, and the number of sharding of each variable and the placement of different PSs affect the memory resource utilization of PS and the throughput of training tasks. Here, a placement strategy for parameters in a distributed task computing graph based on better memory utilization is implemented in the Master module. Based on this, after creating a new container, the computing graph of the old container needs to be updated in the manner of "on-demand reallocation of parameters" and "dynamic rewriting of computing graph". Among them, the execution steps of the on-demand reallocation mechanism of parameters are: according to the actual situation of the load pressure and processing resources of the scaled PS (including the scaled PS and the scaled PS), the model parameters in the computing graph are sharded, so as to obtain multiple model parameter shards (i.e. multiple task sub-data); the multiple model parameter shards are processed by the scaled PS. Among them, the scaled PS includes the scaled PS and the original PS; the scaled PS includes multiple original PSs except the PS to be deleted (scaled) among the multiple original PSs. Among them, several Ops will be reserved in the parameter reallocation process: for example, asti scaling will reserve all the variables on the new PS, including Embedd i and V a
[0013] The "e I ast i c_subgraph_ i mport" triggers the flow of the parameter reassignment by hand, and the "e I ast i c_subgraph_move" is used to transfer the variables without sharding. Step 2: According to the preset update data corresponding to the initial task computing graph, the initial task computing graph is updated to obtain an updated task computing graph. The preset update data can be understood as pre-designed data for updating the task computing graph, and the preset update data can be a pre-designed operator graph. The updated task computing graph can be understood as a task computing graph after the updating processing based on the preset update data. In one or more embodiments provided in the present disclosure, the preset update data is a pre-designed operator graph; the initial task computing graph is updated according to the preset update data corresponding to the initial task computing graph to obtain an updated task computing graph, including: determining a plurality of computing graph parameters in the initial task computing graph and operation nodes corresponding to each computing graph parameter; determining a pre-designed operator graph corresponding to the initial task computing graph, and determining a target computing subgraph corresponding to each computing graph parameter from the pre-designed operator graph; and replacing the operation nodes corresponding to each computing graph parameter in the initial task computing graph by using the target computing subgraph corresponding to each computing graph parameter to obtain an updated task computing graph. In the above example, since the number of roles in the Tensorf l ow distributed task has changed, the computing graph for task running also needs to be changed accordingly. However, the computing graph submitted by the user is fixed, which needs to support dynamic modification of the computing graph. Based on this, in the E I ast i cTra i n i ngPass, the number of PS after expansion is supported for dynamic graph rewriting of the computing graph. For dynamic rewriting of the computing graph, the steps are as follows:
[0014] 1. Pre-design a subgraph corresponding to each computing graph parameter. Since a hook function is registered in the session (Mon i toredT ra i n i ngSess i on) of the tensorf I ow computing by using Sess i onRunHook, the hook function will construct a subgraph related to parameter reassignment in the computing graph in advance when the session starts.
[0015] 2、In the case of capacity expansion or capacity reduction, determine the subgraph corresponding to each computing graph parameter (i.e. the target computing subgraph), replace the computing graph part corresponding to the parameter with the pre-constructed subgraph, and determine the corresponding PS for execution, thereby completing the computing graph update and obtaining the updated task computing graph. Step three: determine the task computing subgraph corresponding to each task sub-data from the updated task computing graph, and determine each task computing subgraph as the target task computing graph. In the above example, after completing the computing graph update, determine the task computing subgraph corresponding to each model parameter slice, and use it as the computing graph for executing the model training task after capacity expansion or capacity reduction. Based on this, after determining the target data processing unit according to the data processing unit adjustment strategy, determine the target task computing graph according to the processing resources of the target data processing unit, so as to avoid problems such as target task execution errors and inability to process tasks caused by the inadaptation of the task computing graph and the target data processing unit. In one or more embodiments of the present disclosure, the execution of the target task according to the target task computing graph and the processing resources of the target data processing unit includes: generating a target task execution session based on the target task computing graph; and executing the target task according to the target task computing graph and the processing resources of the target data processing unit using the target task execution session. It should be noted that before determining the target data processing unit obtained by the resource adjustment unit, it also includes: deleting the initial task execution session generated based on the initial task computing graph. The target task execution session can be understood as the session required by the task computing unit and the target data processing unit in the process of executing the target task computing graph, and the initial task execution session can be understood as the session required by the task computing unit and the initial data processing unit in the process of executing the initial task computing graph. In the task processing method provided in the present disclosure, after the worker container reconstructs the computing graph after the PS capacity expansion or capacity reduction, a new session is created for the reconstructed computing graph, which facilitates the execution of the computing graph based on the newly created session, thereby completing the Tensorflow task. It should be noted that before capacity expansion or capacity reduction, the worker container needs to delete the session corresponding to the computing graph; in addition, before capacity expansion or capacity reduction, the worker container captures and saves the state of the target task running, so that the target task can be restored and restarted in this state.In one or more embodiments provided by the present disclosure, the executing the target task according to the target task computing graph and the processing resource of the target data processing unit comprises: storing the task data to the target data processing unit, wherein the target data processing unit comprises the initial data processing unit and a newly added data processing unit, and the newly added data processing unit is created by the resource adjustment unit in response to the running resource adjustment request, or the target data processing unit comprises a plurality of data processing units in addition to the initial data processing unit, and the initial data processing unit is any one of the plurality of data processing units; and performing computing processing on the task data stored in the target data processing unit based on the target task computing graph, so as to execute the target task. The data processing unit can be used to store task data and computing data generated by the task computing unit in the process of computing the task data. The target data processing unit comprises a plurality of data processing units in addition to the initial data processing unit, which can be understood as that the target data processing unit comprises one data processing unit in addition to the initial data processing unit. In the above example, the Worker container is a container for completing model computing in the process of executing the tensorflow task, and the PS container is a container for implementing PS function in the process of executing the tensorflow task, and the PS can store model parameters in the process of executing the tensorflow task. Based on this, the PS container can store model parameters in the process of executing the tensorflow task and data generated by the Worker container in the process of model computing. In one or more embodiments provided by the present disclosure, the task processing method can send a running resource adjustment request to the resource adjustment unit in the process of executing the target task according to the initial task computing graph and the initial data processing unit, and when the resource adjustment unit processes the initial data processing unit according to the running resource adjustment request to obtain a target data processing unit, the target task computing graph is determined according to the processing resource of the target data processing unit and the initial task computing graph, so as to overcome the problem of mismatch between the target task corresponding task computing graph and the adjusted target data processing unit.Then, the task computing unit successfully executes the target task according to the target task computing graph and the target data processing unit, so as to ensure the successful execution of the target task while adjusting the data processing unit, avoid the problem of waste of computer resources caused by the mismatch between the computing graph and the adjusted computer resources, and the problem of failure to process the target task and thus failure to meet the actual needs of the user. Referring to FIG. 5, FIG. 5 shows a flowchart of another task processing method according to an embodiment of the present disclosure, which can be applied to a resource adjustment unit in a task processing system, the resource adjustment unit can be in communication connection with an initial data processing unit and / or a task computing unit in the task processing system, and specifically includes the following steps. Step 502: In response to a running resource adjustment request sent by the task computing unit, the current running resource of the initial data processing unit is determined, wherein the running resource adjustment request is sent in the process of executing a target task by the task computing unit, the target task is executed by the task computing unit according to an initial task computing graph and the processing resource of the initial data processing unit, and the initial task computing graph is constructed by the task computing unit according to the task data and task operation information carried in the target task. In one or more embodiments provided in the present disclosure, the resource adjustment unit can be a resource adjustment container. Following the above example, the resource adjustment container can be a master container. The master container can include the following functions: 1. Get the state of the tensorflow task from the container cluster; 2. Receive and process the external monitoring module of Gazer; 3. Data strategy decision module; 3. Update the plan of the Tensorflow job according to the elastic expansion and contraction result. o The master container includes a k8sApiClient module to continuously monitor the state of the Tensorflow task. If the Tensorflow task is abnormal, the master container will determine whether to recover the container according to the exit code of the PS or worker container. The exit code refers to an integer value returned by the main process of a container (such as a Docker container) when it terminates running, which is used to indicate the reason for stopping running or the execution result of the main process. The exit code of the container is important for troubleshooting. By checking the exit code, it can be determined whether the container exits and whether the container needs to be recovered. For example, when the container exits due to high resource load, the container can be recovered when the resource load is low.
[0016] The Master container has a DecisionMgr module, which periodically decides the appropriate number of PS (i.e., PS containers) resources for the current Tensorflow task based on the data reported by external monitoring such as the Gazer module. If it is found that the memory used by the existing PS has reached 80% of the memory applied for by the container, the Master will apply for more PS to avoid the risk of container OOM (Out Of Memory). The risk of container OOM refers to the risk of causing a memory overflow error in the process of running the container due to the requested memory exceeding the limit of available memory resources. If it is found that the memory used by the existing PS is less than 50% of the memory applied for by the container, the Master will reduce the number of PS to reduce the wasted PS memory resources.
[0017] The Master container has a module responsible for controlling the expansion and contraction of Tensorflow. The module maintains a state machine. The state machine is shown in FIG. 6. FIG. 6 is a state machine diagram of another task processing method according to an embodiment of the present disclosure. Based on FIG. 6, the state machine flowchart includes the following stages: Running, Sealing up / down, Acquire Resource, WaitForSessionReady, and Release Resource. In the Running stage, the Master container is in a normal running state. In the Sealing up / down stage, the Master container determines whether the task needs to be expanded or contracted according to the decision of the DecisionMgr module. In the Acquire Resource stage, if the task needs to be expanded, the Master container applies for resources required for expansion through the K8sApiClient module to generate PS containers. The PS containers are pulled up to expand the information of each Tensorflow role, and the next state is entered. In the WaitForSessionReady stage, the Master container waits for each role of Tensorflow to complete the process of re-distributing parameters and return a SessionReady signal to the Master container. After being ready, the next state is entered. In the Release Resource stage, if the task needs to be contracted, the Master container releases the redundant PS resources. The interaction between the modules is shown in FIG. 7. FIG. 7 is a diagram of the interaction between the modules in the task processing system according to another task processing method.Based on Figure 7, the cluster part contains the information of training cluster. In the version of TensorFlow, a cluster can include a number of nodes (workers), but in the worker, usually there is a worker that needs to do more work, such as saving the checkpoint and TensorBoard summary file, this worker is called the chief worker. In the implementation of the MuIliWorkerMirroredStrategy of TensorFlow, the default 0 worker (the first in the worker list) will be the chief worker. For the chief worker, in the distributed training scenario of TensorFlow, Chief can refer to a special kind of worker in the distributed environment. In the distributed implementation of TensorFlow, when multiple workers or multiple GPUs are trained, one of the workers will be designated as the "Chief", also known as the "Chief Worker". The Chief can perform initialization, checkpoint management, evaluation and detection, end training and other operations in distributed training. Among them, initialization can refer to: before training starts, the Chief Worker is responsible for initializing global variables and other necessary training states. Checkpoint management can refer to: the Chief Worker is usually responsible for saving and restoring the checkpoint file of the model. This means that it controls when and how to persist the updated model parameters in the training process to the disk, so as to continue training when training fails or needs to be restored. Evaluation and detection can refer to: in some cases, the Chief Worker is also responsible for performing evaluation steps, such as periodically evaluating the performance of the model on the validation set, and recording training indicators. End of training can refer to: after training is completed, the Chief Worker may also trigger some cleanup jobs or other end-of-life operations. Based on this, according to Figure 7, 1. Gazer: periodically report measurement information.
[0018] 2. Elastic Training Controller: According to the measurement information, drive the training task resource state machine. 3. When elastic scaling needs to occur, apply for resources from kubeDL to form new PS containers (clusters).
[0019] 4. Elastic TrainingHook, periodically checks whether to perform elastic scaling. If yes, close the current session (Session) and block, and wait until step 5 is completed. 5. When all worker containers complete closing the current session, send a new PS container (cCluster) to all GrpcServer (PS container / worker container). 6. ch i ef container sends E I ast i cTra i n i ngHook, executes run op, and other worker containers block until all parameters are restored. In one or more embodiments provided by the present disclosure, the current running resource is obtained by a resource detection module corresponding to the initial data processing unit, and the resource detection module performs resource detection on the initial data processing unit through a plurality of resource detection objects to obtain the current running resource. The resource detection module can be understood as a module capable of collecting the current running resource of the initial data processing unit. The resource detection module can send the current running resource of the initial data processing unit to the resource adjustment element. The resource detection module can be a software module or a hardware module, for example, the resource detection module can be a software module such as a plug-in, an instance, a script, or a hardware module such as a device, a hardware component, without specific limitation. The resource detection object can be understood as an object managed by the resource detection module for detecting various types of current running resources of the initial data processing unit. The resource detection object can be an operator, a process, etc. Specifically, another task processing method provided by one or more embodiments of the present disclosure is applied to a resource adjustment element in a task processing system, and the resource adjustment element in the task processing system is configured with a resource detection module. The resource detection module can perform detection on various types of current running resources of the initial data processing unit through a plurality of resource detection objects in the process that the initial data processing unit executes a target task according to an initial task computation graph (for example, an operator (i.e., a resource detection object) is used to collect the memory usage of resources in the computation graph, and another operator is used to collect the CPU and memory usage of the process), to obtain a plurality of types of current running resources of the initial data processing unit. After the resource adjustment element receives a running resource adjustment request sent by the task computation element, the resource adjustment element can obtain a plurality of types of current running resources of the initial data processing unit from a data storage element.The data storage unit can be a local storage unit (for example, a hard disk, a memory, etc.) of the resource adjustment unit. Alternatively, the data storage unit can be a database connected to the resource adjustment unit, used to store the current running resources of the initial data processing unit. In the process of processing the model training task by the worker container and the PS container, the Gazer module in the elastic training system can collect the Tensorflow task training indicators and report to the Master container. For the running process of the Gazer module, refer to FIG. 8, which is a schematic diagram of the Gazer module in another task processing method according to an embodiment of the present disclosure. The visualization tool (TensorBoard) is a visualization tool of Tensorflow, which can help users understand, debug and optimize the Tensorflow program. TensorBoard provides a series of built-in tools, which can intuitively view and analyze the training process of the machine learning model by displaying various useful visual representations in the browser. GraphStatOp, ResourcelltizationOp and ReportMasterOp are all custom Tensorflow Ops (operators) built in the Gazer. Based on FIG. 8, the Gazer module is a Tensorflow plugin (Tensorflow Addons), which functions to collect different Profiler Metrics of each role in the Tensorflow task, so that the Master container can obtain the fine-grained state of the computation graph during task running. The Gazer module includes Tensorflow custom Ops (i.e., resource acquisition operators): GraphStatOp, ResourcelltizationOp, ReportMasterOp and a Graphper. The Graphper will dynamically insert the custom Ops (GraphStatOp, ResourcelltizationOp, ProcessStatOp and ReportMasterOp) into each role (worker, PS) in turn when the Tensorflow task starts.The GraphStatOp is used to collect the time of Gather subgraph and an iteration in the computation graph; the ResoureeUtilizationOp is used to collect the resource memory usage in the computation graph; the ProcessStatOp is used to collect the CPU and memory usage of the process; the ReportMasterOp is used to collect the data collected by other Gazer Op and send the data to the Master through RPC for real-time analysis of the task (i.e., the chief worker in FIG. 3 reports the performance analysis indicators (report metrics) to the master M), and meanwhile, the Op also persists the data to the DFS for Tensorboard visualization. The Gazer Op can be extended, and users can collect key indicators in the graph in a customized manner. Step 504: determining a data processing unit adjustment strategy according to the current running resource of the initial data processing unit. In one or more embodiments provided in the present disclosure, the data processing unit adjustment strategy can be understood as a strategy for adjusting the data processing unit, and in one or more embodiments provided in the present disclosure, the data processing unit adjustment strategy can be a data processing unit addition strategy or a data processing unit deletion strategy. In one or more embodiments provided in the present disclosure, the determining of the data processing unit adjustment strategy according to the current running resource of the initial data processing unit includes: in a case where the current running resource of the initial data processing unit is less than or equal to a first resource threshold, determining target resource data of the target data processing unit; and generating a data processing unit addition strategy for the initial data processing unit based on the target resource data. The first resource threshold can be understood as a threshold for judging whether the initial data processing unit is in a load state. If the current running resource is less than or equal to the first resource threshold, it is determined that the initial data processing unit is in a load state, and a data processing unit addition strategy needs to be determined for the initial data processing unit. For example, in a case where the current running resource is an idle memory resource (for example, any value in the value range [0, 1] can be used to represent the idle memory resource; the larger the value, the more the idle memory resource; the smaller the value, the less the idle memory resource), the first resource threshold can be 0.1. That is, in a case where the idle memory resource is less than or equal to 0.1, it can be determined that the initial data processing unit is in a load state. In addition, if the current running resource is of multiple types, each type of current running resource can have a corresponding first resource threshold.The target data processing unit can be understood as a data processing unit capable of processing the target task in the case of load occurring in the initial data processing unit processing the target task, and is used to share the load pressure of the initial data processing unit. The load problem of the initial data processing unit can be avoided through the target data processing unit. The target resource data corresponding to the target data processing unit can be understood as data representing the processing resources in the target data processing unit capable of processing the target task, such as CPU computing capacity, memory capacity, etc. The data processing unit addition strategy can be understood as a strategy of adding a new data processing unit. The data processing unit addition strategy includes the number of added data processing units and resource data of the added data processing units. Based on this, the resource adjustment unit generates a data processing unit addition strategy for the initial data processing unit in response to a running resource adjustment request, in the case that the current running resource of the initial data processing unit is equal to the first resource threshold, thereby adding a new data processing unit through the data processing unit addition strategy to share the load pressure of the initial data processing unit and ensure the smooth progress of the target task. In one or more embodiments of the present disclosure, the data processing unit adjustment strategy is determined according to the current running resource of the initial data processing unit, including: in the case that the current running resource of the initial data processing unit is greater than or equal to a second resource threshold, a data processing unit deletion strategy for the initial data processing unit is determined. The second resource threshold can be understood as a threshold for judging whether the initial data processing unit is idle. If the current running resource is greater than or equal to the second resource threshold, it is determined that the initial data processing unit is in an idle state, and a data processing unit deletion strategy needs to be determined for it. For example, in the case of the current running resource being idle memory resource (for example, any value in the value range [0, 1] can be used to represent it; the larger the value, the more idle memory resources; the smaller the value, the less idle memory resources), the second resource threshold can be 0.6. That is, in the case that the memory usage is greater than or equal to 0.6, it can be determined that the initial data processing unit is in an idle state. In addition, if the current running resource is of multiple types, each type of current running resource can have a corresponding second resource threshold. The data processing unit deletion strategy can be understood as a strategy of deleting the initial data processing unit.Based on this, the resource adjustment unit generates a data processing unit deletion strategy for the initial data processing unit in response to the running resource adjustment request, in a case where the current running resource of the initial data processing unit is greater than or equal to the second resource threshold, so as to delete the initial data processing unit through the data processing unit deletion strategy, thereby releasing the idle running resource of the initial data processing unit and avoiding waste of resources. Step 406: Adjust the processing resource of the initial data processing unit according to the data processing unit adjustment strategy, to obtain a target data processing unit, so that the task computing unit executes the target task according to the processing resource of the target data processing unit and a target task computing graph, wherein the target task computing graph is determined according to the processing resource of the target data processing unit and the initial task computing graph. In one or more embodiments provided in the present disclosure, adjusting the processing resource of the initial data processing unit according to the data processing unit adjustment strategy to obtain a target data processing unit includes: generating a resource acquisition request based on the data processing unit addition strategy, and sending the resource acquisition request to a resource allocation unit; receiving the unit creation resource returned by the resource allocation unit based on the resource acquisition request, and creating an added data processing unit based on the unit creation resource; and determining the initial data processing unit and the added data processing unit as the target data processing unit. In the above example, in the task processing method provided in the present disclosure, when the resource adjustment unit determines the data processing unit addition strategy, it can generate a resource acquisition request based on the data processing unit addition strategy and send the resource acquisition request to the resource allocation unit. After receiving the resource acquisition request, the resource allocation unit divides the corresponding server resource (i.e., unit creation resource) for the resource adjustment unit from the resources managed by itself based on the data processing unit addition strategy, and sends the resource information of the server resource to the resource adjustment unit. The resource adjustment unit receives the resource information returned by the resource management node, determines the corresponding server resource based on the resource information, and then generates an added data processing unit based on the server resource.After obtaining the new data processing unit, the initial data processing unit and the new data processing unit are determined as target data processing units; the processing resources of the target data processing unit include the processing resources of the initial data processing unit and the processing resources of the new data processing unit, by creating a new new data processing unit, the processing resources of the data processing unit in the execution of the target task process are increased; thereby the load pressure of the initial data processing unit is shared by the new data processing unit, and the processing resources of the initial data processing unit are increased in the execution of the target task process by using the processing resources of the initial data processing unit, so that the target task is smoothly performed. In the above example, refer to FIG. 9, FIG. 9 is a schematic diagram of another task processing method provided by an embodiment of the present disclosure. The steps of container expansion are as follows.
[0020] 1. The ElasticTrainingHook configured in each worker container (i.e., task computing unit) requests the Master container whether to perform scaling by the function "isready scaling () " at the beginning of a certain number of gIobaI_step.
[0021] 2. The Elastic Training service module in the Master container (i.e., resource adjustment unit) decides whether scaling is needed according to the seaIing_action field in the response. It should be noted that the Master container has a DecisionMgr module, and the function of the DecisionMgr is to periodically decide the appropriate number of PS resources for the current Tensorflow task according to the data reported by the external monitoring such as Gazer. If it is found that the use of the existing PS (i.e., PS container) memory reaches 80% of the memory applied by the container, the Master container will apply more PS to avoid the risk of container OOM. If it is found that the use of the existing PS memory is less than 50% of the memory applied by the container, the Master will reduce the number of PS to reduce the waste of PS memory resources.
[0022] 3. The Master container sends a session release instruction to the ElasticTrainingHook when it is determined that scaling (SCALINGJJP) is needed.
[0023] 4、 ElasticTrainingHook (refers to the ElasticTrainingHook configured in the PS container and the worker container), in response to the release session instruction, executes sess i on. c I ose () to release the original session.
[0024] 5、 ElasticTrainingHook, notifies the Master container to trigger the E I ast i cGrpcServer update through the function “ readytoupdate () ”.
[0025] 6、 Master container, according to the instruction of the worker, applies to the kubeDL to generate a container server (i.e., E I ast i cGrpcServer, a gRPC server with elastic scaling feature).
[0026] 7、 Master container, generates a new PS based on the device resources of the server, thereby realizing the expansion of the PS.
[0027] 8、 Master container, after generating the new container, provides the new PS to the worker container.
[0028] 9、 After generating the new PS, the ElasticTrainingHook configured on the chief executes the function “ sess i on. create () ” to re-create a Sess i on (session) for the calculation subgraph.
[0029]
[0030] 10、 The ElasticTrainingHook configured on the chief executes sess i on. run ( [e I ast i c_subgraph_i nit]) to initialize the variables on the PS of the chief.
[0031] 11. The ElasticTrainingHook configured on chief executes session.run([elastic_subgraph_import]) to redistribute parameters on all PSs, thereby executing the computation subgraph through the session and implementing model training. In one or more embodiments provided herein, adjusting the processing resources of the initial data processing unit according to the data processing unit adjustment policy to obtain a target data processing unit includes: determining multiple data processing units based on the data processing unit deletion policy, wherein the initial data processing unit is any one of the multiple data processing units; determining other data processing units in the multiple data processing units other than the initial data processing unit as target data processing units, and deleting the initial data processing unit. The other data processing units in the multiple data processing units other than the initial data processing unit can be understood as at least one data processing unit in the multiple data processing units other than the initial data processing unit. Continuing with the above example, referring to FIG10 , FIG10 is a schematic diagram of shrinking in another task processing method provided by an embodiment of the present disclosure. Based on FIG10 , it can be seen that the steps of shrinking the container are:
[0032] 1. After a certain number of glob_steps, the ElasticTrainingHook configured in each worker calls the "isreadyscale()" function to request the Master container whether to scale up or down.
[0033] 2. The Master container decides whether to scale up or down based on the seaLink_action field in the response.
[0034] 3. When the Master container determines that scaling down is required (SCALING_DOWN), the Master container determines the PS that needs to be deleted and sends an instruction to the ElasticTrainingHook.
[0035] 4. In response to the instruction, the Elasti cTrainingHook on the chief determines the PS to be deleted and executes session.run([elasti c_subgraph_import]) to complete the reallocation of parameters on all PSs.
[0036] 5. In the Elasti cTrainingHook on the chip, execute session.run([elasti c_subgraph_move]) to complete the migration of some parameters on the deleted PS.
[0037] 6. chief EI ast i cTrainingHook, execute session. c I ose () to release the original session.
[0038] 7. The ElasticTrainingHook on the chip notifies the Master container to trigger resource recycling (delete PS) through the function "readytoupdate()"
[0039] 8. The Master container deletes the PS container according to the instructions of ElasticTrainingHook.
[0040] 9. The Master container, according to the instructions of EAST icTrainingHook, sends a server resource recovery request to kubeDL, instructing kubeDL to recycle the three spoonfuls of resources corresponding to the deleted PS in the server (EAST icGrpcServer).
[0041] 10. After the Master container completes the deletion of the PS container, it notifies the Worker container.
[0042] 11. After resource recycling is completed, the ElasticTrainingHook on the worker container creates a Session for the computation graph through the function "session.create()".
[0043] 12. The ElasticTrainingHook configured on chief executes session.run([elastic_subgraph_init]) to initialize variables on the selected PS. Furthermore, the ElasticTrainingHook on chief redistributes parameters on the selected PS, thereby executing the computational subgraph and implementing model training. Another task processing method in one or more embodiments of the present disclosure can determine a data processing unit adjustment strategy based on the current operating resources of the initial data processing unit, and determine a target data processing unit based on the data processing unit adjustment strategy. This allows for flexible determination of the target data processing unit for processing the target task based on the current operating resources of the initial data processing unit. This avoids the problem of inappropriate allocation of computer resources for task processing due to factors such as different resource requirements for different tasks and the inability to accurately estimate the resources required during task processing, thereby further improving the processing efficiency of the target task. The following, combined with Figure 11, further illustrates the task processing method provided by the present disclosure using its application in a model training scenario as an example. Figure 11 shows a flowchart of the processing process of a task processing method provided by one embodiment of the present disclosure. It should be noted that the task processing method provided by the present disclosure is applied to the Master module of a training system (a highly available, adaptive, elastic training deep learning system). The training system comprises a Distributed Container Service (DLC) backend and a Kubernetes cluster. The DLC backend includes a DLC service and kubeDL. The DLC service is the backend service for submitting jobs to the Kubernetes cluster. kubeDL is a Kubernetes resource controller responsible for scheduling and coordinating deep learning. In the Kubernetes cluster, each Tensorflow task processed by the Kubernetes cluster corresponds to a container set consisting of N Master containers, N PS containers, and M worker containers. The Master container manages multiple PS containers and multiple worker containers; both PS and worker containers are virtual services. Based on the above content, the task processing method provided in this disclosure specifically includes the following steps.Step 1102: Train the DLC service in the system, which can receive the model training task (create job) sent by the user. Specifically, since the DLC service is a backend service that can submit jobs to the k8s cluster, the DLC service can receive the model training task sent by the user. Step 1104: The DLC service forwards the model training task to kubeDL for scheduling. Step 1106: kubeDL determines the corresponding container set for the model training task and processes the model training task through the container cluster. Step 1108: The Master container in the container set can receive the model training task (i.e., the Tensorflow task) sent by kubeDL through the Elastic Training controller module. The Master container can include the Elastic Training controller module, the Gazer module, and the Elastic Training service module. Step 1110: The Elastic Training controller module determines the required task resources for executing the Tensorflow task and calculates the current appropriate task resource configuration for the multiple PS containers and multiple worker containers in the container set based on the required task resources. Specifically, the Elastic Training controller module first determines the task resources carried in the Tensorflow task. The task resources are the resource information required by the user to execute the Tensorflow task. Second, the current PS resource information of the multiple PS containers and the current worker resource information of the multiple worker containers in the container set are determined. Finally, based on the current idle resource conditions (PS resource information and worker resource information) of the PS containers and worker containers, the model parameters of the Tensorflow task are divided into multiple model sub-parameters, where the model sub-parameters correspond to the idle resource conditions of the PS containers and worker containers; thereby calculating the current appropriate task resource configuration (i.e., the correspondence between the model sub-parameters and the PS containers and worker containers).Step 1112: The E I ast i c Training cont ro I I er module sends the model sub-parameters to the PS container and the worker container for execution, so as to improve the model training performance by using multiple PS containers and multiple worker containers. Step 1114: The Gazer module in the Master container can obtain the performance analysis indexes (Profi ler Metr ics) of the worker containers and the PS containers during the running of the PS containers and the worker containers. o Specifically, since the multiple PS containers have corresponding worker containers, the Gazer module in the Master container can collect the performance analysis indexes of the worker containers and the PS containers through multiple operators (Ops).
[0044] Gazer module: contains Tensorf I ow custom Op: GraphStatOp, Resourcellt i I zat i on0p , ReportA I MasterOp and Grapp I er o Among them, Report A I MasterOp is deployed in ch i ef PS, and the function of this operator is to collect the data collected by other Gazer Ops, and send the data (performance analysis indexes) collected by other Ops to the Master container through the ch i ef PS for real-time analysis of tasks. Step 1116: During the running of the PS containers and the worker containers, the worker containers can periodically send requests for scaling to the Master container through the E I ast i c Training hook deployed by themselves. Among them, the PS containers can also deploy the E I ast i c Training serv i ce module. Step 1118: The E I ast i c Training serv i ce module in the Master container performs scaling on the PS containers according to the performance analysis indexes detected by the Gazer module. Specifically, the E I ast i c Training serv i ce module in the Master container determines whether to perform scaling on the PS containers according to the performance analysis indexes detected by the Gazer module. Among them, the execution mode of scaling the PS containers is:
[0045] 1. The ElastTrainingHook configured in each worker container uses the function "isreadyscale()" to request the Master container whether to expand or shrink after a certain number of glob_steps.
[0046] 2. The Elastic Training service module in the Master container decides whether to scale up or down based on the seaLink_action field in the response. It should be noted that the Master container contains a DecisionMgr module, which regularly determines the appropriate number of PS resources for the current TensorFlow task based on data reported by external monitoring sources such as Gazer. If the memory usage of the existing PS (i.e., PS container) reaches 80% of the container's requested memory, ? The master container will request more PSs to avoid the risk of the container reaching 00MB. If the memory used by the existing PSs is consistently less than 50% of the container's requested memory, the master will reduce the number of PSs to reduce wasted PS memory resources.
[0047] 3. Master container, when it is determined that expansion is needed (SCALING_UP), the master container sends a session release instruction to ElasticTrainingHook.
[0048] 4. EI Ast i cTrainingHook (refers to the configuration in the PS container and worker container
[0049] EI ast i cTrainingHook), in response to the release session instruction, executes session. c Iose() to release the original session.
[0050] 5. ElasticTraini ngHook, through function u readytoupdate() notifies the Master container to trigger an EI ast i cGrpcServer update.
[0051] 6、 Master container, according to the instruction of worker, apply to kubeDL to generate container server (i.e. ElasticGrpcServer, a gRPC server with elastic scaling feature).
[0052] 7、 Master container, generate new PS based on the device resource of server, so as to realize the expansion of PS. The execution mode of the scaling of PS is:
[0053] 1、 E I ast i cTra i n i ngHook configured in each worker, after a certain number of g I oba I _step, the function " i sreadysca I i ng () ” requests Master container whether to perform expansion and scaling.
[0054] 2、 Master container, according to the sea I i ng_act i on field in the response, decide whether to need to expand and scale.
[0055] 3、 Master container, in the case of determining the need to scale down (SCAL I NG_DOWN), the master container determines the PS to be deleted, and sends an instruction to E I ast i cTra i n i ngHook.
[0056] 4. The ElasticTrainingHook on the chief, in response to the instruction, determines the PS to be deleted and executes session.run([elastic_subgraph_import]) to complete the parameter redistribution on all PSs. In the distributed training scenario of TensorFlow, Chief can refer to a special type of Worker in the distributed environment. In the distributed implementation of TensorFlow, when performing multi-worker and multi-GPU training, one of the Workers is designated as the "Chief," also known as the "Chief Worker." In distributed training, the Chief can implement operations such as initialization, checkpoint management, evaluation and testing, and training termination. Initialization can refer to: Before training begins, the Chief Worker is responsible for initializing global variables and other necessary training states. Checkpoint management can refer to: The Chief Worker is generally responsible for saving and restoring the model's checkpoint files. This means deciding how and when to persist updated model parameters to disk during training, so that training can resume in the event of training failure or recovery. Evaluation and testing can refer to: In some cases, the Chief Worker is also responsible for performing evaluation steps, such as regularly evaluating model performance on a validation set and recording training metrics. Terminating training can refer to: After training, the Chief Worker may also trigger cleanup or other time-consuming operations.
[0057] 5. chief EI ast i cTrainingHook, execute session.run([e I ast i c_subgraph_move]) to complete the migration of some parameters on the deleted PS.
[0058] 6. chief EI ast i cTrainingHook, execute session. c I ose () to release the original session.
[0059] 7. The ElasticTrainingHook on the chip notifies the Master container to trigger resource recycling (delete PS) through the function "readytoupdate()"
[0060] 8、 Master container, according to the instruction of Elast icTraini ngHook, delete PS container.
[0061] 9、 Master container, according to the instruction of Elast icTraini ngHook, send a server resource recycling request to kubeD L, and inform kubeDL to recycle the resource corresponding to the deleted PS in the server (Elasti cGrpcServer). Step 1120: After completing the scale operation of PS, because the number of PS changes, the worker container needs to optimize and adjust the computing graph by executing the "on-demand parameter reallocation mechanism" and "dynamic rewriting of computing graph". It should be noted that because the operation of scaling or scaling the PS will cause the number of PS in the container to change, after the PS scaling or PS scaling, the "on-demand parameter reallocation mechanism" and "dynamic rewriting of computing graph" are needed to optimize and adjust the computing graph. Among them, for the "on-demand parameter reallocation mechanism", the execution steps are:
[0062] 1、 Elast icTraini ngHook configured on chief, according to the performance of the scaled PS and the actual demand of the original PS load pressure, etc., the parameters are divided into multiple model parameter fragments;
[0063] 2、 Elast icTraini ngHook configured on cchief, distribute multiple model parameter fragments to corresponding worker containers and PS containers, and use the scaled PS and worker containers to process multiple model parameter fragments. Among them, for the "dynamic rewriting of computing graph", the execution steps are:
[0064] 1. The ElasticTrainingPass configured on the chief predetermines the corresponding subgraph for each parameter. To achieve flexible training, the Tensorflow job in the training system needs to reconfigure the computation graph after adjusting the number of PSs, redistribute parameters across different PSs, and regularly interact with the Master. These functions are implemented in ElasticTrainingHook and ElasticTrainingPass via the Tensorflow SessionRunHook. The SessionRunHook registers a hook function in the Tensorflow computation (MonitoredTrainingSession) session. At the start of the session, this hook function pre-builds the subgraphs related to graph parameter redistribution for all model parameters in the computation graph.
[0065] 2. When the ElasticTrainingPass configured on the chief is expanded, the pre-built subgraph is used to replace the portion of the computation graph corresponding to the parameter and the corresponding expanded PS is determined for execution, thereby completing the computation graph adjustment. Step 1122: The chief sends the optimized computation graph to each PS for execution, thereby implementing the execution of the model training task. After the PS is expanded, the process of executing the computation graph using the expanded PS includes:
[0066] 1. After generating a new PS, the EAST i cTrainingHook configured on the chief recreates the Session for the computational subgraph through the function "session.create()".
[0067] 2. The ElasticTrainingHook configured on chief executes session.run([elastic_subgraph_init]) to initialize the variables on the new PS.
[0068] 3. The ElasticTrainingHook configured on chief executes session.run([elastic_subgraph_import]) to redistribute parameters on the PS, thereby executing the computation subgraph through the session and implementing model training. After the PS is scaled down, the process of executing the computation graph using the scaled-down PS includes:
[0069] 1. After resource recycling, the ElasticTrainingHook on the worker container creates a Session for the computation graph through the function "session.create()".
[0070] 2. The ElasticTrainingHook configured on chief executes session.run([elastic_subgraph_init]) to initialize the variables on the new PS.
[0071] 3. The ElasticTrainingHook on the chief completes the redistribution of parameters on all PSs, thereby executing the computational subgraph through the session to implement model training. Step 1124: After completing the model training task and obtaining the trained model, the Master container in the container sends the trained model to kubeDL. Step 1126: kubeDL sends the trained model to the DLC service, thereby providing the trained model to the user. The task processing method in one or more embodiments of the present disclosure provides a highly available adaptive elastic training deep learning system. Considering the problem of insufficient resource utilization in the Tensorflow task in the cluster. In addition, many users are often affected by I when training new models using the Parameter Server distributed mode of Tensorflow. 1The problems of the task throughput cannot be extended and so on. The TensorflowJob (model training task) is solved by supporting the elastic expansion. Specifically, by collecting the pre-set ProfilerMetrics in the TensorflowJob through the roles periodically to the Controller, the PS in the Job is dynamically adjusted to alleviate the problem of insufficient resource utilization. Since the Graph and the Variable in the Graph and the corresponding Session are built in advance in the system design of Tensorflow, the Job needs to be reconfigured when it is scaled, and the saving and loading of the Graph and the Variable are needed, which is a non-negligible overhead for a normal TensorflowJob. Therefore, a scheme of non-restarting Session is designed, and the mechanism of re-distribution of the Variable and the dynamic rewriting of the Graph are introduced to reduce the additional overhead of the system caused by the dynamic adjustment of the number of PS. The method of dynamically networking the PS roles of the Tensorflow distributed task is implemented: by modifying the implementation of the Tensorflow distributed or GrpcSession, the PS is supported to be added and deleted, and the mechanism of on-demand re-distribution of the Variable and the dynamic rewriting of the Graph are introduced to realize the efficient expansion and contraction of the tensorflow task. In addition, the parameter placement strategy considering the memory usage: the parameters in the Tensorflow distributed task Graph are placed based on the better memory utilization strategy; thus, the user can meet the two goals of high resource utilization and high task throughput without the corresponding background knowledge.And, through the adaptive elastic training deep learning system, the users can focus on the model without additional experimental optimization to achieve good training performance, and the utilization rate of machine resources of the cluster can be improved. Referring to FIG. 12, FIG. 12 is a flowchart of a model training task processing method according to an embodiment of the present disclosure, which can be applied to a task computing unit in a task processing system, and the task computing unit can be in communication connection with an initial data processing unit and / or a resource adjusting unit in the task processing system, and specifically includes the following steps. Step 1202: sending a running resource adjustment request to the resource adjusting unit, wherein the running resource adjustment request is sent in the process of executing a model training task, the model training task is executed according to an initial task computing graph and processing resources of an initial data processing unit, and the initial task computing graph is constructed according to service data and operation information carried in the model training task. Step 1204: receiving a unit determination instruction sent by the resource adjusting unit, wherein the unit determination instruction is sent by the resource adjusting unit in response to the running resource adjustment request, and the target data processing unit is obtained by the resource adjusting unit adjusting the processing resources of the initial data processing unit. Step 1206: determining the target data processing unit based on the unit determination instruction, and determining a target task computing graph according to the processing resources of the target data processing unit and the initial task computing graph. Step 1208: executing the model training task according to the target task computing graph and the processing resources of the target data processing unit. The model training task processing method in one or more embodiments of the present disclosure can send a running resource adjustment request to the resource adjusting unit in the process of executing the model training task according to the initial task computing graph and the initial data processing unit; and when the target data processing unit is obtained by the resource adjusting unit adjusting the initial data processing unit according to the running resource adjustment request, the target task computing graph is determined according to the processing resources of the target data processing unit and the initial task computing graph, thereby overcoming the problem of mismatch between the task computing graph corresponding to the model training task and the adjusted target data processing unit. Then, the task computing unit successfully executes the model training task according to the target task computing graph and the target data processing unit, thereby ensuring the smooth execution of the model training task while adjusting the data processing unit.The problems of mismatching of the calculation graph and the adjusted computer resources or waste of computer resources are avoided, and the problem that the model training task processing cannot meet the actual needs of users is solved. The above is an illustrative scheme of the model training task processing method of the embodiment. It should be noted that the technical scheme of the model training task processing method belongs to the same concept as the technical scheme of the above-described model training task processing method or another model training task processing method. The details of the technical scheme of the model training task processing method that are not described in detail can be found in the description of the technical scheme of the above-described model training task processing method or another model training task processing method. Referring to FIG. 13, FIG. 13 shows a structural schematic diagram of a task processing system according to an embodiment of the present disclosure, which includes a task calculation unit 1302, an initial data processing unit 1304, and a resource adjustment unit 1306. The task calculation unit 1302, the initial data processing unit 1304, and the resource adjustment unit 1306 can be communicatively connected. The task calculation unit 1302 is configured to send a running resource adjustment request to the resource adjustment unit 1306. The running resource adjustment request is sent in the process of executing a target task. The target task is executed according to an initial task calculation graph and the processing resources of the initial data processing unit 1304. The initial task calculation graph is constructed according to task data and operation information carried in the target task. The unit determination instruction sent by the resource adjustment unit 1306 is received. The unit determination instruction is sent by the resource adjustment unit 1306 in response to the running resource adjustment request and in the case of a target data processing unit. The target data processing unit is obtained by adjusting the processing resources of the initial data processing unit 1304 by the resource adjustment unit 1306. The target data processing unit is determined based on the unit determination instruction. The target task calculation graph is determined according to the processing resources of the target data processing unit and the initial task calculation graph. The target task is executed according to the target task calculation graph and the processing resources of the target data processing unit.The resource adjusting unit 1306 is configured to determine a current running resource of the initial data processing unit 1304 in response to a running resource adjusting request sent by the task computing unit 1302, wherein the running resource adjusting request is sent in a process of executing a target task by the task computing unit 1302, the target task is executed by the task computing unit 1302 according to an initial task computing graph and a processing resource of the initial data processing unit 1304, and the initial task computing graph is constructed by the task computing unit 1302 according to task data and task operation information carried in the target task. The current running resource of the initial data processing unit 1304 is determined, a data processing unit adjusting strategy is determined according to the current running resource of the initial data processing unit 1304, the processing resource of the initial data processing unit 1304 is adjusted according to the data processing unit adjusting strategy, a target data processing unit is obtained, the target task is executed by the task computing unit 1302 according to a processing resource of the target data processing unit and a target task computing graph, and the target task computing graph is determined according to the processing resource of the target data processing unit and the initial task computing graph. In the task processing system provided by one or more embodiments of the present disclosure, the task computing unit can send a running resource adjusting request to the resource adjusting unit in a process of executing a target task according to an initial task computing graph and an initial data processing unit, and in the case that the target data processing unit is obtained by adjusting the initial data processing unit according to the running resource adjusting request, the target task computing graph is determined according to a processing resource of the target data processing unit and the initial task computing graph, so as to overcome the problem of mismatch between the target task computing graph and the target data processing unit after adjustment. Then, the target task is executed by the task computing unit according to the target task computing graph and the target data processing unit, so as to ensure the smooth execution of the target task while the data processing unit is adjusted.The problems of computer resource waste caused by the mismatch between the calculation graph and the adjusted computer resources, and the problem of failing to meet the actual needs of users due to the failure to process the target task are avoided. Moreover, the resource adjustment unit can determine the data processing unit adjustment strategy according to the current running resources of the initial data processing unit, and determine the target data processing unit according to the data processing unit adjustment strategy, so as to flexibly determine the target data processing unit for processing the target task according to the current running resources of the initial data processing unit, and avoid the problem of failing to reasonably allocate computer resources for task processing due to different resources required by different tasks and the inability to accurately estimate the resources required during task processing, and further improve the processing efficiency of the target task. The above is a schematic scheme of an embodiment of the task processing system. It should be noted that the technical scheme of the task processing system belongs to the same concept as the technical scheme of the above task processing method or another task processing method, and the details of the technical scheme of the task processing system that are not described in detail can be referred to the description of the technical scheme of the above task processing method or another task processing method. Corresponding to the above method embodiment, the present disclosure also provides a task processing device embodiment, which can be applied to a task calculation unit in a task processing system. The device comprises: a request sending module configured to send a running resource adjustment request to a resource adjustment unit, wherein the running resource adjustment request is sent during the execution of a target task, the target task is executed according to an initial task calculation graph and the processing resources of an initial data processing unit, and the initial task calculation graph is constructed according to the task data and operation information carried in the target task; an instruction receiving module configured to receive a unit determination instruction sent by the resource adjustment unit, wherein the unit determination instruction is sent by the resource adjustment unit in response to the running resource adjustment request, and the target data processing unit is obtained by adjusting the processing resources of the initial data processing unit by the resource adjustment unit; and a calculation graph determination module configured to determine the target data processing unit based on the unit determination instruction, and determine a target task calculation graph according to the processing resources of the target data processing unit and the initial task calculation graph.The task execution module is configured to execute the target task according to the target task computation graph and the processing resource of the target data processing unit. Optionally, the computation graph determination module is further configured to: divide the task data into a plurality of task sub-data according to the processing resource of the target data processing unit; update the initial task computation graph according to preset update data corresponding to the initial task computation graph to obtain an updated task computation graph; determine a task computation sub-graph corresponding to each task sub-data from the updated task computation graph, and determine each task computation sub-graph as the white-label task computation graph. Optionally, the preset update data is a preset computation sub-graph; the computation graph determination module is further configured to: determine a plurality of computation graph parameters in the initial task computation graph and operation nodes corresponding to each computation graph parameter; determine a preset computation sub-graph corresponding to the initial task computation graph, and determine a target computation sub-graph corresponding to each computation graph parameter from the preset computation sub-graph; and replace the operation nodes corresponding to each computation graph parameter in the initial task computation graph with the target computation sub-graph corresponding to each computation graph parameter to obtain an updated task computation graph. Optionally, the task processing device further includes a computation graph construction module configured to: receive the target task sent by the resource adjustment unit; construct an initial task computation graph according to the task data and the task operation information carried in the target task, and execute the target task according to the initial task computation graph and the processing resource of the initial data processing unit. Optionally, the task execution module is further configured to: generate a target task execution session based on the target task computation graph; and execute the target task according to the target task computation graph and the processing resource of the target data processing unit by using the target task execution session. Optionally, the task execution module is further configured to: store the task data to the target data processing unit, wherein the target data processing unit includes the initial data processing unit and an added data processing unit, and the added data processing unit is created by the resource adjustment unit in response to the running resource adjustment request, or the target data processing unit includes a plurality of data processing units in addition to the initial data processing unit, and the initial data processing unit is any one of the plurality of data processing units.Based on the target task computation graph, the task data stored in the target data processing unit is calculated and processed to execute the target task. One of the embodiments provided by the present disclosure is a task processing device which can send a running resource adjustment request to a resource adjustment unit in the process of executing a target task according to an initial task computation graph and an initial data processing unit, and according to the processing resource of the target data processing unit and the initial task computation graph, determine a target task computation graph when the resource adjustment unit processes the initial data processing unit according to the running resource adjustment request, so as to overcome the problem of mismatch between the task computation graph corresponding to the target task and the adjusted target data processing unit. Then, the task computation unit successfully executes the target task according to the target task computation graph and the target data processing unit, thereby ensuring the smooth execution of the target task while adjusting the data processing unit, avoiding the problem of waste of computer resources caused by the mismatch between the computation graph and the adjusted computer resources, and the problem of failing to process the target task and thus failing to meet the actual needs of the user. The above is a schematic scheme of the task processing device of the embodiment. It should be noted that the technical scheme of the task processing device belongs to the same concept as the technical scheme of the task processing method described above, and the details of the technical scheme of the task processing device which are not described in detail can be referred to the description of the technical scheme of the task processing method. Corresponding to the method embodiment described above, the present disclosure also provides another task processing device embodiment, which can be applied to a resource adjustment unit in a task processing system, and the device comprises: a request response module configured to determine the current running resource of an initial data processing unit in response to a running resource adjustment request sent by a task computation unit, wherein the running resource adjustment request is sent in the process of executing a target task by the task computation unit, the target task is executed by the task computation unit according to an initial task computation graph and the current running resource of the initial data processing unit, and the initial task computation graph is constructed by the task computation unit according to the task data and task operation information carried in the target task; a strategy determination module configured to determine a data processing unit adjustment strategy according to the current running resource of the initial data processing unit;The unit adjustment module is configured to adjust the processing resource of the initial data processing unit according to the data processing unit adjustment strategy, obtain a target data processing unit, and enable the task computing unit to execute the target task according to the processing resource of the target data processing unit and a target task computing graph, wherein the target task computing graph is determined according to the processing resource of the target data processing unit and the initial task computing graph. Optionally, the current running resource is obtained by a resource detection module corresponding to the initial data processing unit, and the resource detection module detects the resource of the initial data processing unit by a plurality of resource detection objects to obtain the current running resource. Optionally, the strategy determination module is further configured to: in a case where the current running resource of the initial data processing unit is determined to be less than or equal to a first resource threshold, determine target resource data of the target data processing unit; and generate a data processing unit addition strategy for the initial data processing unit based on the target resource data. Optionally, the unit adjustment module is further configured to: generate a resource acquisition request based on the data processing unit addition strategy, and send the resource acquisition request to a resource allocation unit; receive unit creation resources returned by the resource allocation unit based on the resource acquisition request, and create an additional data processing unit based on the unit creation resources; and determine the initial data processing unit and the additional data processing unit as the target data processing unit. Optionally, the strategy determination module is further configured to: in a case where the current running resource of the initial data processing unit is determined to be greater than or equal to a second resource threshold, determine a data processing unit deletion strategy for the initial data processing unit; and the unit adjustment module is further configured to: determine a plurality of data processing units based on the data processing unit deletion strategy, wherein the initial data processing unit is any one of the plurality of data processing units.The other data processing units except the initial data processing unit in the plurality of data processing units are determined as target data processing units, and the initial data processing unit is deleted. Another task processing device provided by one or more embodiments of the present disclosure can determine a data processing unit adjustment strategy according to the current running resource of the initial data processing unit, and determine the target data processing unit according to the data processing unit adjustment strategy, so as to realize flexible determination of the target data processing unit for processing the target task according to the current running resource of the initial data processing unit, and avoid the problem that computer resources cannot be reasonably allocated for task processing due to different resources required by different tasks and the fact that the required resources during task processing cannot be accurately estimated, and further improve the processing efficiency of the target task. The method can be applied to the task processing device shown in the accompanying drawings. It should be noted that the technical solutions of the above-mentioned another method for adjusting the current running resource of the initial data processing unit belong to the same technical field of the above-mentioned method for adjusting the current running resource of the initial data processing unit, and can be applied to the above-mentioned another method for adjusting the current running resource of the initial data processing unit. The target data processing unit is determined according to the computer resources required by the target task, so as to avoid the problem that the computer resources cannot be reasonably allocated for task processing, and the target data processing unit is determined according to the computer resources required by the target task, so as to avoid the problem that the computer resources cannot be reasonably allocated for task processing.According to the description of the training type, the construction of the training case is carried out, and the development and questioning of the time are analyzed. The target task is to determine the mechanism of the design model and the configuration of the training case. The model should be used to provide a single training model for the training of the training personnel. The model should also be used to provide a single training model for the training personnel. The information of the model should be provided according to the needs of the training case. The technical information of the training case should be provided according to the needs of the training project. The technical information of the training case should be provided according to the needs of the training personnel. The technical information of the training case should be provided according to the needs of the training personnel. The source of the service. In the task according to the data installation standard, the basic training and collection of the single-purpose service case, the customer service task directly obtains the actual number of training methods, and the method of using the diagram to the project model and the individual training and adjustment technology to calculate the whole unit for the purpose of adjusting the training and adjustment of the type of single-purpose service. The source of the service is multiple project models and the mathematical and technical descriptions are matched and matched with each other or the other service method system is different. The model is assigned to be processed and adjusted after the technical description is performed, and the source service method is required. The model is assigned to be initially configured and the block number is sent according to the description. The training is required to perform the technical implementation of the model. The source diagram of the model is assigned to the number of blocks, and the number of blocks is assigned to the initial calculation mechanism of the configuration of the adjustment type. The model is assigned to provide the model with the required calculation and processing standards, and the training elements are actually required to provide the correct implementation method and the initial implementation of the model and the design and processing type. FIG14 shows a block diagram of a computing device 1400 provided according to an embodiment of the present disclosure. The computing device 1400 includes, but is not limited to, a memory 1410 and a processor 1420. oThe processor 1420 is connected with the memory module 1422 and the database 1450 through the bus 1430, and the database 1450 is used to save data. The computing device 1400 further includes an access device 1440, which enables the computing device 1400 to communicate via one or more networks 1460. Examples of these networks include the public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of networks such as the Internet. The access device 1440 can include one or more of any type of network interface (for example, a network interface card (NIC)), wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface. In one embodiment of the present disclosure, the above-mentioned components of the computing device 1400 and other components not shown in FIG. 14 can also be connected with each other, for example, through a bus. It should be understood that the computing device structure block diagram shown in FIG. 14 is only for the purpose of example, and is not a limitation on the scope of the present disclosure. Those skilled in the art can add or replace other components as needed.The computing device 1400 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other type of mobile device, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1400 can also be a mobile or stationary server. The processor 1420 is configured to execute instructions / computer programs to implement the steps of the above-described task processing method, another task processing method, or a model training task processing method. Each embodiment in the present disclosure is described in a progressive manner, and the same or similar parts between embodiments can be mutually referred to. Each embodiment focuses on the differences from other embodiments. In particular, for the computing device embodiment, since it is basically similar to the task processing method embodiment, another task processing method embodiment, or a model training task processing embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the task processing method embodiment, another task processing method embodiment, or a model training task processing embodiment. An embodiment of the present disclosure further provides a computer-readable storage medium storing computer programs / instructions, which are executed by a processor to implement the steps of the above-described task processing method, another task processing method, or a model training task processing method. Each embodiment in the present disclosure is described in a progressive manner, and the same or similar parts between embodiments can be mutually referred to. Each embodiment focuses on the differences from other embodiments. In particular, for the computer-readable storage medium embodiment, since it is basically similar to the task processing method embodiment, another task processing method embodiment, or a model training task processing embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the task processing method embodiment, another task processing method embodiment, or a model training task processing embodiment. An embodiment of the present disclosure further provides a computer program product including computer programs / instructions, which are executed by a processor to implement the steps of the above-described task processing method, another task processing method, or a model training task processing method. The above is a schematic diagram of a computer program product of an embodiment of the present disclosure.It should be noted that the technical solution of the computer program product belongs to the same concept as the above-mentioned task processing method, another task processing method, or model training task processing method, and the technical solution of the computer program product is not described in detail. The details can be seen from the description of the technical solution of the above-mentioned task processing method, another task processing method, or model training task processing method. The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous. The computer instructions include computer program code, which can be in the form of source code, object code, executable code, or some intermediate form. The computer-readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals. It should be noted that for the foregoing method embodiments, in order to facilitate description, they are all expressed as a series of actions, but those skilled in the art should know that the embodiments of the present disclosure are not limited by the order of the described actions, because according to the embodiments of the present disclosure, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the embodiments of the present disclosure. In the above embodiments, the description of each embodiment is focused on, and the parts not described in detail in a certain embodiment can be seen from the related description of other embodiments. The above disclosed preferred embodiments of the present disclosure are only used to help explain the present disclosure. Optional embodiments do not describe all the details and limit the invention to the described specific embodiments. Obviously, according to the content of the embodiments of the present disclosure, many modifications and changes can be made.The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can well understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.
Claims
Claims 1. A task processing method, comprising: Sending an operation resource adjustment request to a resource adjustment unit, wherein the operation resource adjustment request is sent during the execution of a target task, the target task is executed according to an initial task calculation graph and the processing resources of an initial data processing unit, and the initial task calculation graph is constructed according to the task data and task operation information carried in the target task; receiving a unit determination instruction sent by the resource adjustment unit, wherein the unit determination instruction is sent when the resource adjustment unit determines a target data processing unit in response to the operation resource adjustment request, and the target data processing unit is obtained by the resource adjustment unit adjusting the processing resources of the initial data processing unit; determining the target data processing unit based on the unit determination instruction, and determining a target task calculation graph according to the processing resources of the target data processing unit and the initial task calculation graph; executing the target task according to the target task calculation graph and the processing resources of the target data processing unit.
2. The task processing method according to claim 1, wherein determining the target task computation graph based on the processing resources of the target data processing unit and the initial task computation graph comprises: According to the processing resources of the target data processing unit, the task data is divided into a plurality of task sub-data; according to the preset update data corresponding to the initial task calculation graph, the initial task calculation graph is updated to obtain an updated task calculation graph; the task calculation sub-graph corresponding to each task sub-data is determined from the updated task calculation graph, and each task calculation sub-graph is determined as the target task calculation graph.
3. The task processing method according to claim 2, wherein the preset update data is a preset computation subgraph; and updating the initial task computation graph according to the preset update data corresponding to the initial task computation graph to obtain an updated task computation graph comprises: Determine multiple computation graph parameters in the initial task computation graph, and operation nodes corresponding to each computation graph parameter; Determine a preset computation subgraph corresponding to the initial task computation graph, and determine a target computation subgraph corresponding to each computation graph parameter from the preset computation subgraph; The target computation subgraph corresponding to each computation graph parameter is used to replace the operation node corresponding to each computation graph parameter in the initial task computation graph to obtain an updated task computation graph.
4. The task processing method according to any one of claims 1 to 3, before sending the running resource adjustment request to the resource adjustment unit, further comprising: receiving the target task sent by the resource adjustment unit; 25 An initial task calculation graph is constructed according to the task data and the task operation information carried in the target task, and the target task is executed according to the initial task calculation graph and the processing resources of the initial data processing unit.
5. The task processing method according to any one of claims 1 to 4, wherein executing the target task according to the target task computation graph and the processing resources of the target data processing unit comprises: Generate a target task execution session based on the target task computation graph; The target task is executed using the target task execution session, and the target task is executed according to the target task computation graph and the processing resources of the target data processing unit.
6. The task processing method according to any one of claims 1 to 5, wherein executing the target task according to the target task computation graph and the processing resources of the target data processing unit comprises: The task data is stored in the target data processing unit, wherein the target data processing unit includes the initial data processing unit and a newly added data processing unit, and the newly added data processing unit is created by the resource adjustment unit in response to the running resource adjustment request, or the target data processing unit includes other data processing units among a plurality of data processing units except the initial data processing unit, and the initial data processing unit is any one of the plurality of data processing units; based on the target task computation graph, computationally processing the task data stored in the target data processing unit to execute the target task.
7. A task processing method, comprising: In response to an operating resource adjustment request sent by a task computing unit, the current operating resources of the initial data processing unit are determined, wherein the operating resource adjustment request is sent by the task computing unit during the execution of a target task, the target task is executed by the task computing unit according to an initial task calculation graph and the processing resources of the initial data processing unit, and the initial task calculation graph is constructed by the task computing unit according to the task data and task operation information carried in the target task; according to the current operating resources of the initial data processing unit, a data processing unit adjustment strategy is determined; according to the data processing unit adjustment strategy, the processing resources of the initial data processing unit are adjusted to obtain a target data processing unit, so that the task computing unit executes the target task according to the processing resources of the target data processing unit and the target task calculation graph, wherein the target task calculation graph is determined according to the processing resources of the target data processing unit and the initial task calculation graph.
8. The task processing method according to claim 7, wherein the current running resources are obtained by a resource detection module corresponding to the initial data processing unit, and the resource detection module performs resource detection on the initial data processing unit through multiple resource detection objects to obtain the current running resources.
9. The task processing method according to claim 7 or 8, wherein determining the data processing unit adjustment strategy based on the current operating resources of the initial data processing unit comprises: determining target resource data of the target data processing unit when it is determined that the current running resource of the initial data processing unit is less than or equal to a first resource threshold; A data processing unit addition strategy is generated for the initial data processing unit based on the target resource data.
10. The task processing method according to claim 9, wherein adjusting the processing resources of the initial data processing unit according to the data processing unit adjustment strategy to obtain the target data processing unit comprises: generating a resource acquisition request based on the new policy added by the data processing unit, and sending the resource acquisition request to the resource allocation unit; Receive the unit creation resource returned by the resource allocation unit based on the resource acquisition request, and create a new data processing unit based on the unit creation resource; and determine the initial data processing unit and the new data processing unit as the target data processing unit.
11. The task processing method according to any one of claims 7 to 10, wherein determining a data processing unit adjustment strategy based on current operating resources of the initial data processing unit comprises: determining a data processing unit deletion policy for the initial data processing unit when it is determined that the current running resource of the initial data processing unit is greater than or equal to a second resource threshold; The adjusting the processing resources of the initial data processing unit according to the data processing unit adjustment policy to obtain a target data processing unit includes: determining multiple data processing units based on the data processing unit deletion policy, wherein the initial data processing unit is any one of the multiple data processing units; determining other data processing units among the multiple data processing units except the initial data processing unit as target data processing units, and deleting the initial data processing unit.
12. A method for processing a model training task, comprising: Sending an operation resource adjustment request to the resource adjustment unit, wherein the operation resource adjustment request is sent during the execution of a model training task, the model training task is executed according to an initial task calculation graph and the processing resources of an initial data processing unit, and the initial task calculation graph is constructed according to the task data and task operation information carried in the model training task; receiving a unit determination instruction sent by the resource adjustment unit, wherein the unit determination instruction is sent when the resource adjustment unit determines a target data processing unit in response to the operation resource adjustment request, and the target data processing unit is obtained by adjusting the processing resources of the initial data processing unit by the resource adjustment unit; determining the target data processing unit based on the unit determination instruction, and determining the target task calculation graph according to the processing resources of the target data processing unit and the initial task calculation graph; executing the model according to the target task calculation graph and the processing resources of the target data processing unit. Training mission.
13. A task processing system, comprising a task calculation unit, an initial data processing unit, and a resource adjustment unit, wherein: The task computing unit is configured to send an operation resource adjustment request to the resource adjustment unit, wherein the operation resource adjustment request is sent during the execution of a target task, the target task is executed according to an initial task calculation graph and the processing resources of the initial data processing unit, and the initial task calculation graph is constructed according to the task data and task operation information carried by the target task, and receives a unit determination instruction sent by the resource adjustment unit, wherein the unit determination instruction is sent when the resource adjustment unit determines a target data processing unit in response to the operation resource adjustment request, and the target data processing unit is obtained by the resource adjustment unit adjusting the processing resources of the initial data processing unit, determines the target data processing unit based on the unit determination instruction, determines a target task calculation graph according to the processing resources of the target data processing unit and the initial task calculation graph, and executes the target task according to the target task calculation graph and the processing resources of the target data processing unit; and the resource adjustment unit is configured to determine the current operation resources of the initial data processing unit in response to the operation resource adjustment request sent by the task computing unit, wherein the operation resource adjustment request is sent during the execution of the target task by the task computing unit, The target task is executed by the task computing unit according to the initial task computing graph and the processing resources of the initial data processing unit, and the initial task computing graph is constructed by the task computing unit according to the task data and task operation information carried in the target task. According to the current running resources of the initial data processing unit, a data processing unit adjustment strategy is determined, and the processing resources of the initial data processing unit are adjusted according to the data processing unit adjustment strategy to obtain a target data processing unit, so that the task computing unit executes the target task according to the processing resources of the target data processing unit and the target task computing graph, wherein the target task computing graph is determined according to the processing resources of the target data processing unit and the initial task computing graph.
14. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the task processing method according to any one of claims 1 to 6, the task processing method according to any one of claims 7 to 11, or the model training task processing method according to any one of claim 12 are implemented.
15. A computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the task processing method described in any one of claims 1 to 6, the task processing method described in any one of claims 7 to 11, or the model training task processing method described in any one of claim 12.
16. A computer program product comprising a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the task processing method according to any one of claims 1 to 6, the task processing method according to any one of claims 7 to 11, 28 The task processing method or the steps of the model training task processing method described in any one of claim 12. 29
Citation Information
Patent Citations
Business processing method and device and equipment
CN113961267A
CPU-GPU heterogeneous resource-oriented task scheduling method
CN114911612A
Data processing method, device and equipment based on many-core architecture and storage medium
CN116578522A