Model training method, task processing method and model training system
By identifying the abnormal training resources in distributed model training and determining the target distribution strategy, the problems of training process blocking and cost increase caused by exception conditions in distributed model training are solved, and efficient and automatic exception recovery and cost optimization are achieved.
Patent Information
- Application Number
- CN202311813223.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing technology is difficult to effectively solve the abnormal situation in distributed model training, resulting in blocking of training processes and increasing costs.
By receiving the status information of each training resource, identifying abnormal training resources, and determining the target distribution strategy based on task parameters and resource parameters, and optimizing task execution indicators to achieve automatic recovery.
It realizes fine-grained abnormality monitoring, optimizes monitoring costs and sub-health costs, improves global training efficiency, and reduces the cost of model training and improves reliability.
Smart Images

Figure CN120218271A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the technical field of machine learning, and particularly to a model training method, a task processing method, and a model training system. Background Art
[0002] With the development of machine learning technology, the parameter specifications, training complexity, and training costs of machine learning models are rising rapidly. Distributing the model training task to multiple training resources for collaborative execution to achieve distributed training has become the mainstream method for model training. In the distributed training method, abnormal conditions that occur on the training resources seriously hinder the training process of model training and increase huge training costs.
[0003] Regarding abnormal conditions, there are currently various solutions, such as abnormal monitoring solutions, checkpoint mechanisms, and dynamic configuration solutions. However, the above solutions usually focus on optimizing a single metric. For example, the abnormal monitoring solution optimizes the monitoring cost. For another example, the checkpoint mechanism optimizes the transition cost. For still another example, the dynamic configuration solution optimizes the sub-healthy cost. It is difficult for any solution to handle various complex and diverse abnormal situations. Therefore, there is an urgent need for a comprehensively optimized model training method with the ability to solve abnormal conditions. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a model training method. One or more embodiments of this specification also relate to a task processing method, a model training system, a model training device, a task processing device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0005] In one embodiment of this specification, a model training method is provided, including:
[0006] Receiving status information sent by each training resource participating in the model training task;
[0007] Identifying abnormal training resources based on the status information of each training resource;
[0008] Determining a target distribution strategy when the task execution metric reaches a preset metric based on the task parameters of the model training task and the resource parameters of each target training resource, where the target training resources are the training resources except the abnormal training resources;
[0009] Distributing the model training task to each target training resource for execution according to the target distribution strategy to obtain a target model that has completed training.
[0010] In one embodiment of this specification, status information sent by each training resource participating in the model training task is received; based on the status information of each training resource, abnormal training resources are identified; based on the task parameters of the model training task and the resource parameters of each target training resource, a target distribution strategy is determined when the task execution metrics reach the preset metrics, where the target training resources are the training resources other than the abnormal training resources; according to the target distribution strategy, the model training task is distributed to each target training resource for execution, and a target model that has completed training is obtained. Directly receiving the status information sent by each training resource participating in the model training task and identifying abnormal training resources based on the status information of each training resource realizes fine-grained anomaly monitoring, can promptly identify abnormal training resources, and optimizes the monitoring cost; based on the task parameters of the model training task and the resource parameters of each target training resource, a target distribution strategy is determined when the task execution metrics reach the preset metrics, comprehensively considering the resource parameters of the target training resources other than the abnormal training resources and combining them with the task parameters of the model training task to determine the target distribution strategy when the task execution metrics reach the preset metrics, optimizing the sub-healthy cost and enhancing the global training efficiency; according to the target distribution strategy, the model training task is distributed to each target training resource for execution, and a target model that has completed training is obtained. According to the target distribution strategy, quickly transitioning from the current abnormal training state to a new training state, promptly resuming model training, optimizing the transition cost, realizing comprehensive, efficient, and automatic recovery of the model training in case of anomalies, reducing the cost of model training, and enhancing the reliability of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a schematic diagram of recovering training for resource anomalies during the model training process;
[0012] Figure 2 is a flowchart of a model training method provided by an embodiment of this specification;
[0013] Figure 3 is a schematic flow diagram of a model training method provided by an embodiment of this specification;
[0014] Figure 4 is a flowchart of a task processing method provided by an embodiment of this specification;
[0015] Figure 5 is a flowchart of the processing process of a model training method applied to a large language model provided by an embodiment of this specification;
[0016] Figure 6 is a schematic structural diagram of a model training system provided by an embodiment of this specification;
[0017] Figure 7It is a schematic structural diagram of a model training device provided by an embodiment of this specification;
[0018] Figure 8 It is a schematic structural diagram of a task processing device provided by an embodiment of this specification;
[0019] Figure 9 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0020] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0021] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0022] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0023] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select to authorize or reject.
[0024] In one or more embodiments of this specification, a large model refers to a machine learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. A large model can also be referred to as a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one hundred million parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.
[0025] When a large model is actually applied, only a small number of samples are needed to fine-tune the pre-trained model for application in different tasks. Large models can be widely applied in fields such as natural language processing (NLP) and computer vision. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0026] First, the noun terms involved in one or more embodiments of this specification are explained.
[0027] Cluster: A scalable computing resource composed of multiple physical or virtual servers that can provide high availability and elastic scaling.
[0028] Exception recovery: Can help users quickly recover services and reduce project interruptions.
[0029] Gradient weight: Refers to the relationship between the gradients in the model and the model parameters. The gradient describes the direction and rate of change of the loss function with respect to the model parameters. The weight, on the other hand, refers to the degree of influence of the model parameters on the loss function. It determines the speed at which the gradient descent algorithm updates the parameters. A larger gradient weight will result in faster parameter updates.
[0030] Distributed training: A method of training machine learning models using multiple distributed training resources, which can greatly improve the training speed and reduce the computing time.
[0031] Model Parallel (MP): A distributed training strategy that involves splitting a model into multiple parts and deploying each part to different distributed training resources for training. This method can effectively address the problem of excessive parameters in the model.
[0032] Data Parallel (DP): A distributed training strategy that splits the entire sample set into multiple sample subsets and assigns each subset to different distributed training resources for training. This technique can improve training efficiency and reduce memory usage.
[0033] Pipeline Parallel (PP): A distributed training strategy that lies between model parallel and data parallel. Its core idea is to break down a large model into multiple layers and organize them in a pipeline for forward and backward propagation calculations, thereby reducing the video memory occupancy of a single card and also lowering the communication overhead.
[0034] Batch: In the machine learning scenario, by dividing the entire dataset into several small datasets and training on the small datasets each time, it avoids the problem of huge computational load caused by all the data in the dataset participating in training at once. Additionally, when performing gradient updates, the gradient direction does not differ significantly from that of the entire dataset, ensuring the training effect.
[0035] Graphics Processing Unit (GPU): A microprocessor used for graphics-related calculations. With the development of machine learning technology, due to its parallel structure, it can achieve efficient matrix operations and is widely used as the computing hardware for model training.
[0036] The Central Processing Unit (CPU) can also be used for model training, but it is usually less efficient than GPUs and TPUs.
[0037] Tensor Processing Unit (TPU): A customized chip specifically designed for machine learning tasks, which can provide higher efficiency and performance than GPUs.
[0038] Field-Programmable Gate Array (FPGA): A programmable logic gate array that can be programmed to implement various functions, including convolutional operations in machine learning.
[0039] Application Specific Integrated Circuit (ASIC): An integrated circuit designed specifically for a particular application, which can provide very high performance but at a relatively high cost.
[0040] Persistent Memory Array: A data storage system that combines multiple persistent storage devices (such as hard disk drives, solid state drives, or tape drives, etc.) to provide a more reliable, efficient, and scalable storage solution.
[0041] Remote Direct Memory Access (RDMA): A technology for directly reading and writing memory between remote computers, which can greatly improve the data transfer speed in the network.
[0042] Network interface card: A hardware device mainly used to establish a physical connection between a computer and a network and complete data transmission by sending and receiving data packets.
[0043] Peripheral Component Interconnect Express (PCIe): A physical connection structure between a set of training resources, usually composed of a bus, which can be used to support data transfer between multiple devices.
[0044] High-speed channel between graphics processing units: A fast communication channel used to connect two or more graphics processing units, which can provide higher bandwidth and lower latency for data transfer between graphics processing units.
[0045] Deep self-attention model (Transformer model): A machine learning architecture based on the attention mechanism, used to process sequential data such as natural language.
[0046] Bidirectional Encoder Representations from Transformers (BERT): A special Transformer model trained using bidirectional Transformer encoders and large-scale unlabeled text data. BERT's excellent performance has made it a standard baseline for many natural language processing tasks.
[0047] Large Language Model (LLM): A machine learning model trained on a large corpus of text for natural language processing tasks. These models generally consist of multiple layers of neural networks. Their input is a sequence of text for text generation, and the output is the result text of performing a specific natural language processing task on this text sequence. Pre-training means that before a specific task, the model has been trained and has pre-learned to process a large amount of language data. By pre-training the model, it can capture more complex language and semantic rules, thus performing well in various natural language processing tasks and reducing the need for large-scale data for specific tasks.
[0048] Training efficiency: Refers to the computing efficiency of the Graphics Processing Unit (GPU) involved in training in large language models, usually measured by the number of floating-point operations per second (FLOPS).
[0049] Fast transition strategy: Quickly transition from the current abnormal training state to a new training state to reduce the cost impact during the transition period.
[0050] Network link interruption: Refers to the interruption of network link services, which can help users quickly restore network connections and reduce project interruptions.
[0051] Abnormal monitoring solution: Through operation and maintenance tools, monitor the status of each training resource in the cluster and locate abnormal training resources. However, these operation and maintenance tools only provide coarse-grained monitoring and are not specifically designed for the model training process. To reduce monitoring costs, currently, abnormal monitoring can be achieved by sending heartbeat messages to training resources regularly, which speeds up the abnormal monitoring. It can also be achieved through an automatic abnormal monitoring system that uses multi-dimensional monitoring metrics and log analysis to monitor abnormalities. However, the above solutions still rely on data outside the training framework for abnormal monitoring and cannot provide the same monitoring accuracy and efficiency as abnormal monitoring directly integrated into the training framework.
[0052] Checkpoint mechanism: The process of periodically saving model information such as the model's state (parameter weights) and optimizer state during model training. These saved checkpoints can be used to load the previous model state in case of training anomalies or when resuming training, to avoid retraining the model from scratch or losing the progress made in model training. The effectiveness of the checkpoint mechanism is limited by training resources, especially in a distributed training environment, where human intervention is required. Due to the factor of human intervention, the time to resolve resource anomalies or determine new training resources is both long and unpredictable, resulting in long waits during the waiting phase. This waiting time can cause other non-anomalous training resources to be idle, leading to a waste of a large amount of training resources and generating high transition costs. At the same time, determining hot standby training resources in advance (extra training resources prepared in a machine learning environment to ensure a continuous and highly available training process) is an option to reduce the waiting time. Such hot standby training resources increase the overall cost and resource waste in the absence of anomalies, adding sub-healthy costs, i.e., reducing the throughput achieved with the same amount of training resources during normal training without training anomalies.
[0053] Dynamic configuration scheme: A scheme for dynamically reconfiguring training resources to continue model training, also known as elastic training / scheduling. This method allows training to continue with a reduced number of worker nodes during the sub-healthy phase, shortening the waiting phase caused by resource anomalies. Although the dynamic configuration scheme is effective in reducing the waiting phase time and improving the utilization of training resources, the reconfiguration process itself incurs overhead, quantified as transition costs, and suboptimal configuration strategies may lead to a significant reduction in throughput, resulting in high sub-healthy costs. To achieve the desired results through dynamic configuration, it is necessary to fully consider the cluster topology and the parallel characteristics of the model to address related challenges such as load imbalance, communication overhead, and memory limitations. The complexity of the problem further increases in the scenario where multiple model training tasks are running on the same cluster, which is very common in distributed training scenarios.
[0054] Figure 1 It is a schematic diagram of resuming training for resource anomalies during the model training process, as Figure 1 shown:
[0055] During the model training process, first, after the normal training phase, when an anomaly occurs, it enters the interruption phase. Then, the interruption phase includes four stages: the waiting phase, the task redistribution phase, the resource reconfiguration and task re-execution phase. After completing the above four stages in sequence, it re-enters the normal training phase. Among them, the longer the waiting phase, the later the training task restarts.
[0056] Alternatively, during the model training process, first, after the normal training phase, when an anomaly occurs, it enters the interruption phase. Then, the interruption phase includes a sub-healthy phase. Before the sub-healthy phase, it is necessary to manually identify the anomaly. In the sub-healthy phase, identify abnormal training resources, repair abnormal training resources, and reconstruct the cluster in sequence, and finally re-enter the normal training phase.
[0057] By Figure 1 , it can be understood that recovering training from an anomaly is a multi-stage process. Among them, each stage will incur specific costs, including: the cost lost during monitoring, which is the monitoring cost C detection ; the cost of reconfiguration loss, that is, the transition cost C transition ; the throughput cost reduced in the sub-healthy phase, that is, the sub-healthy cost C sub-healthy . These costs together affect the overall efficiency and output of the training process. Therefore, when resource anomalies occur during the model training process, the corresponding recovery training cost C recovery is expressed by Formula 1, and Formula 1 is as follows:
[0058] C recovery = C detection + C transition + C sub-healthy Formula 1
[0059] Existing solutions such as anomaly monitoring schemes, checkpoint mechanisms, and dynamic configuration schemes all have limitations. These solutions usually optimize a single cost in Formula 1 above, but do not fully consider the overall cost of the recovery training cost C recovery in Formula 1, especially in multi-task scenarios.
[0060] Therefore, how to comprehensively reduce the recovery training cost C recovery , so as to achieve a comprehensive, efficient, and automatic recovery of the model training in case of anomalies, reduce the cost of model training, and improve the reliability of model training, is an urgent problem to be solved. An ideal solution should optimize the costs consumed in the monitoring phase and the transition phase, so as to reduce the monitoring cost C detection and the transition cost C transition respectively. In addition, it should ensure that during the normal training phase, the throughput is not reduced and the sub-healthy cost C sub-healthy generated is minimized.
[0061] To address this problem, this specification provides a model training method. This specification also relates to a task processing method, a model training system, a model training device, a task processing device, a computing device, a computer-readable storage medium, and a computer program, which will be described in detail one by one in the following embodiments.
[0062] SeeFigure 2 , Figure 2 shows a flowchart of a model training method provided by an embodiment of this specification, including the following specific steps:
[0063] Step 202: Receive the status information sent by each training resource participating in the model training task.
[0064] The embodiment of this specification is applied to the training scheduling end of the model training function and the exception recovery function. This scheduling end can be a dedicated system, function end, module, or plug-in. Taking the plug-in as an example, the exception recovery is realized during the model training process through this plug-in, ensuring the smooth compatibility of the model training. The integration ensures the retention of the model training function, thus avoiding any reduction in the model training performance.
[0065] The model training task is a distributed task for training a target model. This model training task includes multiple training subtasks. By distributing the training subtasks to distributed training resources for execution, the distributed training is completed. For example, the model training task is a model training task for a large language model, and there are 100 resource instances (nodes) for training. The model training task is divided into 100 training subtasks and distributed to 100 resource instances for execution to complete the distributed training. The model training task includes the model parameters of the target model and the sample data for training the target model. According to the distributed training strategy, the model training task is divided into multiple types, including but not limited to: the model training task of the model parallel type, the model training task of the data parallel type, and the model training task of the pipelined parallel type.
[0066] The training resource is a resource instance for distributed training, including but not limited to: computing resources, storage resources, and network resources. A training resource can be understood as a distributed node in a distributed training cluster, that is, a physical or virtual server deployed with computing resources, storage resources, and network resources. Among them, the computing resources include but not limited to: graphics processing unit, central processing unit, tensor processing unit, programmable logic gate array, and application-specific integrated circuit. The storage resources include but not limited to: cache, memory, hard disk, or persistent storage array. The network resources include but not limited to: network card, external component interconnect bus topology, and high-speed channel between graphics processing units. In the cluster, the control of multiple nodes is completed through a dedicated configuration module, and data exchange is realized through collective communication on each distributed node.
[0067] The status information of training resources is data information that describes the status of each training resource during the training process, and describes the running status, running performance, utilization rate, and health status of each training resource. The status information is used to monitor whether the training process is executed normally. The status information of training resources includes, but is not limited to: Resource availability: Reflects whether the training resource is in a usable state, such as whether the computing resource is normal, whether the storage resource is sufficient, whether the network connection is normal, etc. Performance metrics: Reflect the running efficiency and speed of the training resource. For example, the utilization rate of the computing resource, the speed of I / O operations, the occupancy of the network bandwidth, etc. Health status: Reflects the stability and reliability of the training resource, including: whether there are software and hardware abnormalities, operating temperature, etc. Load balancing: Reflects the workload distribution among various training resources. For example, in a distributed training environment with 10 distributed nodes, 4 graphics processing units are deployed on each node. Resource availability: 9 servers are online, and 1 server is offline due to a power abnormality. Performance metrics: The average utilization rate of the graphics processing units is 75%, and the utilization rate of the graphics processing units on 3 nodes reaches more than 90%; the overall occupancy rate of the network bandwidth is 60%. Health status: Server 2 reports a warning about the overheating temperature of a graphics processing unit, and there are read and write errors in the storage resources of Server 8. Load balancing: The utilization rate of the graphics processing units on Server 1 and Server 5 is relatively low (50%), while the utilization rate of the graphics processing units on Server 3 and Server 7 is relatively high (95%).
[0068] Receive the status information sent by each training resource participating in the model training task. The specific method is: through the monitoring threads pre-deployed on the training resources, receive the status information sent by each training resource participating in the model training task. Among them, the monitoring threads are monitoring threads pre-deployed on each training resource, used to obtain the status information of the training resource and send it to the scheduler through a network connection, and this network connection is a persistent network connection pre-established between each node and the scheduling end. It realizes the fine-grained monitoring of each training resource.
[0069] Exemplarily, on the model training system, through a preset training scheduling terminal plugin, the model training and anomaly recovery are realized. The model training system includes 10 distributed nodes, and each distributed node is deployed with 1 central processing unit, 4 graphics processing units, and memory. The model training task of the large language model is divided into 10 training subtasks according to the data parallel strategy, and the 10 training subtasks are distributed to 10 distributed nodes for execution. During the task execution process of the 10 distributed nodes, through the monitoring threads pre-deployed on the training resources, the status information of the training resources sent by the 10 distributed nodes is received: the average utilization rate of the central processing unit on node 1 is 85%, the utilization rate of the graphics processing unit on node 4 is 75%, the network bandwidth occupancy rate between node 3 and node 8 is 70%, the read and write speed of node 2 is (1700MB / s, 1200MB / s), the packet loss rate of node 6 is 23%, node 4 reports that the temperature of graphics processing unit -1 is 95°C...
[0070] Receive the status information sent by each training resource participating in the model training task. Directly receiving the status information sent by each training resource participating in the model training task improves the efficiency of anomaly recognition and provides information support for subsequent fine-grained anomaly monitoring and timely identification of abnormal training resources.
[0071] Step 204: Identify abnormal training resources based on the status information of each training resource.
[0072] An abnormal training resource is a training resource with abnormal resource status during the model training process. Abnormal training resources will affect the efficiency, stability, and accuracy of the entire model training task. For example, if the power supply of node 2 is unstable, node 2 is an abnormal training resource. Another example is that if node 6 has a software crash, node 6 is an abnormal training resource.
[0073] Optionally, after identifying abnormal training resources based on the status information of each training resource, the following specific steps are further included:
[0074] Interrupt the model training task on the abnormal training resource, or interrupt the model training tasks on each training resource.
[0075] Interrupting the training task specifically means interrupting the training process on the training resource.
[0076] To identify abnormal training resources based on the status information of each training resource, the specific method is: based on the status information of each training resource, determine whether each training resource is abnormal, and identify abnormal training resources from multiple training resources. Among them, determining whether each training resource is abnormal can be judged by a preset threshold, or by statistical analysis, or by analysis using machine learning algorithms, which is not limited here. Through the fine-grained monitoring of each training resource, abnormal training resources can be identified in a timely manner.
[0077] Exemplarily, based on the status information of the training resources sent by 10 distributed nodes, through a preset threshold, it is determined whether there are any abnormalities in the training resources on the 10 distributed nodes: the memory of node 2 is abnormal, the graphics processing unit of node 4 is abnormal, and the network connection of node 6 is abnormal. The abnormal training resources are determined from the 10 distributed nodes: node 2, node 4, and node 6, and the model training tasks on the 10 distributed nodes are interrupted.
[0078] Based on the status information of each training resource, abnormal training resources are identified. Fine-grained anomaly monitoring is achieved, abnormal training resources can be identified in a timely manner, and the monitoring cost is optimized.
[0079] Step 206: Based on the task parameters of the model training task and the resource parameters of each target training resource, determine the target distribution strategy when the task execution index reaches the preset index, where the target training resources are the training resources other than the abnormal training resources.
[0080] The target training resources are resource instances that can normally execute the model training task during the model training process except for the abnormal training resources. They can be resource instances participating in the model training task or resource instances not participating in the model training task. For example, among the 10 distributed nodes participating in the model training task, 1 node is an abnormal training resource, and the remaining 9 are target training resources. Another example is that among the 10 distributed nodes participating in the model training task, 1 node is an abnormal training resource, and 1 new distributed node is added to obtain 10 target training resources.
[0081] The resource parameters of the target training resources are parameters describing the resource characteristics of the target training resources, including but not limited to: the computing parameters, storage parameters, and network parameters of the target training resources. For the computing parameters, including but not limited to: the number, model, and parameters of the central processing unit, the number, model, and parameters of the graphics processing unit. For the storage parameters, including but not limited to: the number, read / write speed, and capacity of the memory, the number, read / write speed, and capacity of the hard disk. For the network parameters, including but not limited to: the upstream speed, downstream speed, bandwidth, latency, and packet loss rate, etc. For example, a target training resource may have the following resource parameters: 4 graphics processing units of model A, 1 central processing unit with 64 cores, 512 GB of memory, and a 10 Gbps network bandwidth. The resource parameters of the target training resources characterize the ability of the target training resources to execute the model training task.
[0082] The task parameters of the model training task are the parameters that define the task characteristics of the model training task, including but not limited to: model architecture, training dataset size, batch size, learning rate, and number of iterations. For example, the task parameters of the model training task include: a deep neural network architecture, a 1TB training dataset, a batch size of 64, a learning rate of 0.001, and 100 iterations. The task parameters of the model training task determine the efficiency, accuracy, and training resource requirements of the model training.
[0083] The distribution strategy is the strategy for distributing the training subtasks of the model training task to multiple training resources. The distribution strategy determines how to distribute the training subtasks to the distributed training resources for execution, and determines the task execution effect of the parallel execution of the model training task. During the execution of the model training task, a specific distribution strategy generates corresponding task execution metrics. The distribution strategy is determined by the task parameters of the model training task and the resource parameters of each target training resource. For example, Distribution Strategy 1 is: for a model training task of 64 batches, every 8 batches form a training subtask, which is distributed to 8 distributed nodes. Another example, Distribution Strategy 2 is: for a 1TB model training task, it is evenly divided into 8 training subtasks according to the data volume and distributed to 8 distributed nodes. The target distribution strategy is the distribution strategy when the task execution metrics reach the preset metrics, which can be understood as an optimized distribution strategy, achieving the optimization of the task execution effect. For example, the training cost of the above Distribution Strategy 1 does not reach the preset training cost, and the training cost of Distribution Strategy 2 reaches the preset training cost, so Distribution Strategy 2 is determined as the target distribution strategy.
[0084] The task execution metrics are quantitative metrics used to evaluate the execution effect of the model training task of the distribution strategy, including but not limited to: training cost, resource utilization rate, and validation accuracy. For the training cost, it includes but not limited to: the number of samples processed per second and the number of floating-point operations per second. For the resource utilization rate, it includes but not limited to: the usage rate of the central processing unit, the usage rate of the graphics processing unit, and the memory occupancy rate. For example, the task execution metrics may include: processing 1000 samples per second, the usage rate of the graphics processing unit is 80%, and the validation accuracy reaches 95%.
[0085] The preset metrics are the preset thresholds of the task execution metrics, serving as the criteria for evaluating whether the task execution effect meets the expectations. It includes but not limited to: training cost threshold, resource utilization rate threshold, and validation accuracy threshold. For example, the preset metrics include: processing at least 500 samples per second, the usage rate of the graphics processing unit does not exceed 90%, and the validation accuracy is higher than 90%. The preset metrics are used to guide the formulation of the distribution strategy.
[0086] Based on the task parameters of the model training task and the resource parameters of each target training resource, determine the target distribution strategy when the task execution metrics reach the preset metrics. The specific method is as follows: Based on the task parameters of the model training task and the resource parameters of each target training resource, construct at least one distribution strategy, and from at least one distribution strategy, determine the target distribution strategy when the task execution metrics reach the preset metrics. Among them, the methods for determining the target distribution strategy include but are not limited to: mathematical analysis, linear programming, particle swarm optimization, genetic algorithms, and machine learning analysis.
[0087] Exemplarily, based on the task parameters of the model training task (large language model architecture, 1TB training dataset, batch size of 64, learning rate of 0.001, number of iterations of 100) and the resource parameters of each target training resource (10 distributed nodes, with 4 A-type graphics processing units, 1 64-core central processing unit, 512GB of memory, and 10Gbps network bandwidth deployed on any one distributed node), construct 20 distribution strategies, calculate the training costs (floating-point operations per second) of the 20 distribution strategies, and determine the target distribution strategy when the training cost (floating-point operations per second) reaches the training cost threshold (floating-point operations per second higher than 300TFLOPS): For 64 batches of model training tasks, batches 1-4 form a training subtask and are distributed to one node, batches 5-10 form a training subtask and are distributed to one node... batches 60-64 form a training subtask and are distributed to one node.
[0088] Based on the task parameters of the model training task and the resource parameters of each target training resource, determine the target distribution strategy when the task execution metrics reach the preset metrics, where the target training resources are the training resources except for the abnormal training resources. The resource parameters of the target training resources except for the abnormal training resources are comprehensively considered and combined with the task parameters of the model training task to determine the target distribution strategy when the task execution metrics reach the preset metrics, optimize the sub-healthy costs, and improve the global training efficiency.
[0089] Step 208: Distribute the model training task to each target training resource for execution according to the target distribution strategy, and obtain the target model that has completed training.
[0090] The target model is a machine learning model that has completed training and has a large number of model parameters. From the perspective of model architecture, the target model includes, but is not limited to: deep self-attention models, deep self-attention models with bidirectional encoding representations, and large language models. From the perspective of data modalities processed, the target model includes, but is not limited to: language processing models, image processing models, speech processing models, code processing models, etc. Taking the language processing model as an example, the language processing model can perform one or more natural language processing tasks, including, but not limited to: machine translation tasks, speech recognition tasks, text analysis tasks, or text question-and-answer tasks.
[0091] According to the target distribution strategy, distribute the model training tasks to each target training resource for execution to obtain the target model that has completed training. The specific method is as follows: Based on the target distribution strategy, determine the sub-tasks to be distributed corresponding to each target training resource, and distribute each sub-task to be distributed to the corresponding target training resource for execution to obtain the target model that has completed training.
[0092] Exemplarily, based on the target distribution strategy, form batch 1-4 into a training sub-task and distribute it to node 1, form batch 5-10 into a training sub-task and distribute it to node 2... form batch 60-64 into a training sub-task and distribute it to node 10. Execute the model training tasks on each node to obtain the large language model that has completed training.
[0093] In the embodiments of this specification, the status information sent by each training resource participating in the model training task is received; based on the status information of each training resource, abnormal training resources are identified; based on the task parameters of the model training task and the resource parameters of each target training resource, a target distribution strategy is determined when the task execution index reaches a preset index, where the target training resources are the training resources other than the abnormal training resources; according to the target distribution strategy, the model training task is distributed to each target training resource for execution, and a target model that has completed training is obtained. Directly receiving the status information sent by each training resource participating in the model training task and identifying abnormal training resources based on the status information of each training resource realizes fine-grained anomaly monitoring, can timely identify abnormal training resources, and optimizes the monitoring cost; based on the task parameters of the model training task and the resource parameters of each target training resource, a target distribution strategy is determined when the task execution index reaches a preset index, comprehensively considering the resource parameters of the target training resources other than the abnormal training resources and combining them with the task parameters of the model training task to determine the target distribution strategy when the task execution index reaches a preset index, optimizing the sub-healthy cost and improving the global training efficiency; according to the target distribution strategy, the model training task is distributed to each target training resource for execution, and a target model that has completed training is obtained. According to the target distribution strategy, quickly transitioning from the current abnormal training state to a new training state, timely resuming the model training, optimizing the transition cost, realizing a comprehensive, efficient, and automatic recovery of the model training in case of anomalies, reducing the cost of model training, and improving the reliability of model training.
[0094] In an alternative embodiment of this specification, before step 206, the following specific steps are further included:
[0095] Based on the status information, determine the abnormal type of the abnormal training resources;
[0096] Based on the abnormal type, determine the target operation information for the abnormal training resources;
[0097] Correspondingly, step 206 includes the following specific steps:
[0098] In the case where the target operation information is to reconfigure the training resources, based on the task parameters of the model training task and the resource parameters of each target training resource, determine the target distribution strategy when the task execution index reaches a preset index.
[0099] The abnormal type of the abnormal training resource is the type of resource abnormality on the abnormal training resource, including but not limited to: lost connection, abnormal exit, connection refused / reset, illegal memory access, memory error, invalid direct memory access mapping, graphics processing unit calculation error, high-speed channel error between graphics processing units, graphics processing unit driver error, other network errors, other software errors, collective communication timeout, link fluctuation, and task suspension. Determining the abnormal type can better understand and solve the factors affecting training efficiency and stability, so as to take corresponding operations.
[0100] The target operation information for the abnormal training resource is the information of the corresponding operation executed for the abnormal type of the abnormal training resource. The target operation information includes but not limited to: retraining operation, restarting the training process operation, and reconfiguring the training resource operation. Optionally, the three operations can be progressive, that is, execute the retraining operation, execute the restarting the training process operation in case of failure, and execute the reconfiguring the training resource operation in case of failure. After determining the target operation information for the abnormal training resource, execute the corresponding target operation. It should be noted that step 106 and step 108 are exactly the reconfiguring the training resource operation.
[0101] Among them, the retraining operation is an operation to re-attempt to handle the abnormality, which is effective for solving temporary connection abnormalities (such as link fluctuation or connection refused / reset). If the retraining operation is successful, the training process will immediately resume normal. For example, in the distributed training process, a node fails to transmit data due to a short network fluctuation. In this case, the system can choose to re-attempt data transmission at the same location. If the network abnormality has been resolved, the retry operation may succeed and the training process can continue.
[0102] Among them, the restarting the training process operation is an operation to restart the training process on the abnormal training resource to handle the abnormality. The configuration of the model training system remains unchanged. If the restarting the training process operation is successful, use the backup information to restart the training process and continue to complete the training. If the restarting the training process operation fails, the configuration of the model training system needs to be modified. For example, on a distributed node, the training process crashes due to a graphics processing unit calculation error. In this case, the system can choose to restart the training process on this node and use the backup information to restore the model state and training progress. If the problem still exists after restarting, the configuration of the model training system needs to be modified.
[0103] Among them, the operation of reconfiguring training resources is an operation to reconfigure the configuration of the model training system to handle exceptions. When the abnormal training resources are successfully restored and available again, or new training resources are newly allocated, they are added to the ongoing training process. In this case, the cluster is started into the reconfiguration process to enable the model training system to adapt to and integrate the target training resources. In addition, when the existing tasks are completed or new tasks are started, the reconfiguration process can also be triggered because the distribution policy may be different from the current configuration, and it is necessary to redistribute the model training tasks to achieve overall cost optimization. For example, in a distributed training cluster, a node fails due to a hardware exception. In this case, the abnormal node is isolated, and the reconfiguration process is started. If there is an available spare node, it is added to the cluster, and the distribution policy of the model training tasks is adjusted according to the new resource allocation to maximize the overall optimized training cost.
[0104] Exemplarily, based on the status information of the training resources sent by 10 distributed nodes, determine the abnormal type of the abnormal training resources. In the case where the abnormal type is "connection refused / reset", determine that the target operation information for the abnormal training resources is the retraining operation, and execute the reconnect operation; in the case where the abnormal type is "graphics processing unit calculation error", determine that the target operation information for the abnormal training resources is the operation of restarting the training process, and execute the operation of restarting the training process; in the case where the target operation information is to reconfigure the training resources, based on the task parameters of the model training task (large language model architecture, 1TB training dataset, batch size of 64, learning rate of 0.001, number of iterations of 100) and the resource parameters of each target training resource (10 distributed nodes, 4 graphics processing units of type A, 1 64-core central processing unit, 512GB of memory, and 10Gbps network bandwidth are deployed on any one of the distributed nodes), construct 20 distribution policies, calculate the training costs (floating-point operations per second) of the 20 distribution policies, and determine the target distribution policy when the training cost (floating-point operations per second) reaches the training cost threshold (floating-point operations per second is higher than 300TFLOPS): 64 batches of model training tasks, batches 1-4 form a training subtask and are distributed to one node, batches 5-10 form a training subtask and are distributed to one node... batches 60-64 form a training subtask and are distributed to one node.
[0105] Based on the status information, determine the anomaly type of the abnormal training resource; based on the anomaly type, determine the target operation information for the abnormal training resource; when the target operation information is to reconfigure the training resource, based on the task parameters of the model training task and the resource parameters of each target training resource, determine the target distribution strategy when the task execution metric reaches the preset metric. Based on the status information, the target operation information is determined more flexibly and accurately, improving the stability of model training.
[0106] In an optional embodiment of this specification, determining the target operation information for the abnormal training resource based on the anomaly type includes the following specific steps:
[0107] Based on the anomaly type, determine the target impact level corresponding to the abnormal training resource;
[0108] Based on the target impact level, determine the target operation information for the abnormal training resource.
[0109] The target impact level is used to measure the impact degree and severity level of this anomaly on the entire distributed training system, and is determined according to the anomaly type of the abnormal training resource. The target impact level includes but is not limited to: the first impact level, the second impact level, and the third impact level. Among them, the first impact level represents a serious anomaly that may affect the stability and function of the entire system. These anomalies usually require immediate action and may cause the training to pause or restart some nodes. The first impact level includes but is not limited to: hardware anomalies of the graphics processing unit, system-level software crashes, and complete interruption of network communication. Anomalies at the first impact level may cause the entire training task to be unable to continue, and it is necessary to immediately isolate the abnormal node and reconfigure the cluster. The second impact level represents a medium-level anomaly that may have a negative impact on the training efficiency or process, but does not immediately threaten the overall stability of the system. Such problems usually require restarting the relevant process or adjusting some configuration parameters. Anomalies at the second impact level include but are not limited to: graphics processing unit calculation errors, illegal memory access, and task suspension. These problems may affect the training process of a single node, and it is necessary to restart the training process or adjust the relevant configuration on the affected node. The third impact level represents a relatively minor anomaly, usually involving temporary problems or local anomalies, and has a relatively small impact on the overall training process. Such problems can usually be resolved by simple retry operations or waiting for a period of time to recover automatically. The third impact level includes but is not limited to: temporary network connection problems (such as link fluctuations or connection rejection / reset) and collective communication timeouts. These anomalies are usually temporary, and normal training can be resumed after simple retry operations or waiting for the network to stabilize.
[0110] Target impact levels and corresponding target operation information, including but not limited to: the third impact level corresponds to the retraining operation, the second impact level corresponds to the operation of restarting the training process, and the first impact level corresponds to the operation of reconfiguring training resources.
[0111] Exemplarily, based on the status information of the training resources sent by 10 distributed nodes, determine the abnormal type of the abnormal training resources. When the abnormal type is "connection refused / reset", determine that the target impact level corresponding to the abnormal training resources is the third impact level. Based on the third impact level, determine that the target operation information for the abnormal training resources is the retraining operation, and perform the reconnection operation; when the abnormal type is "graphics processing unit calculation error", determine that the target impact level corresponding to the abnormal training resources is the second impact level. Based on the second impact level, determine that the target operation information for the abnormal training resources is the operation of restarting the training process, and perform the operation of restarting the training process; when the target operation information is to reconfigure the training resources, determine that the target impact level corresponding to the abnormal training resources is the first impact level. Based on the first impact level, based on the task parameters of the model training task (large language model architecture, 1TB training dataset, batch size of 64, learning rate of 0.001, number of iterations of 100) and the resource parameters of each target training resource (10 distributed nodes, 4 graphics processing units of model A, 1 64-core central processing unit, 512GB of memory, and 10Gbps network bandwidth deployed on any distributed node), construct 20 distribution strategies, calculate the training costs (floating-point operations per second) of the 20 distribution strategies, and determine the target distribution strategy when the training cost (floating-point operations per second) reaches the training cost threshold (floating-point operations per second higher than 300T FLOPS): 64 batches of model training tasks, batches 1-4 form a training subtask and are distributed to one node, batches 5-10 form a training subtask and are distributed to one node... batches 60-64 form a training subtask and are distributed to one node.
[0112] Based on the abnormal type, determine the target impact level corresponding to the abnormal training resources; based on the target impact level, determine the target operation information for the abnormal training resources. Based on the target impact level, the target operation information is determined more flexibly and accurately, further improving the stability of model training.
[0113] In an optional embodiment of this specification, before step 206, the following specific steps are further included:
[0114] Obtain the task backup information recorded during the execution of the model training task, where the task backup information includes task progress information and backup model information, and the backup model information is the model information of the model training task at a preset checkpoint;
[0115] Determine the task parameters of the model training task based on the task backup information.
[0116] The task backup information is the training data recorded during the execution of the model training task. It is a type of backup information to handle possible failures, interruptions, or other abnormal situations. The task backup information includes task progress information and backup model information. The task progress information is the data describing the key metrics and states during the execution of the model training task. The task progress information is used to monitor the execution process of the model training task, evaluate the model performance, and adjust or resume training when needed. The backup model information is the model information at the preset checkpoint of the model training task, which is used to restore the model state during or after the training process to prevent data loss or progress loss caused by training failure. The model information at the preset checkpoint is the model parameters and state data at the key time points or conditions preset during the model training process for saving the model state. To ensure that the model state before an anomaly can be restored when an anomaly occurs during training. For example, the task backup information includes: task progress information such as the current iteration number, prediction data of forward propagation, gradient values of backward propagation, loss function values, learning rates, etc., and backup model information such as the model parameter weights and biases saved at the preset checkpoint. Among them, the model information at the preset checkpoint may be saved at the end of each iteration or when the performance improvement on the validation set exceeds a certain threshold.
[0117] Obtain the task backup information recorded during the execution of the model training task. The specific method is: obtain the task backup information recorded during the execution of the model training task from the storage medium. Further, obtain the task backup information recorded during the execution of the model training task from the hierarchical storage medium according to the storage performance from high to low. Among them, the storage medium includes but is not limited to: cache, memory, hard disk, and persistent storage array. The cache, memory, hard disk, and persistent storage array are hierarchically arranged according to the read / write speed from high to low, forming a hierarchical storage medium. During the model training process, backup storage is completed from high to low through asynchronous storage. For example, first obtain the task backup information from the memory, and if it cannot be obtained from the memory, obtain it from the hard disk.
[0118] Exemplarily, obtain the task backup information from the hierarchical storage medium according to the storage performance from high to low (cache → memory → persistent storage array): task progress information such as the current iteration number, prediction data of forward propagation, gradient values of backward propagation, loss function values, learning rates, etc., and backup model information such as the model parameter weights and biases saved at the preset checkpoint.
[0119] Obtain the task backup information recorded during the execution of the model training task, where the task backup information includes task progress information and backup model information, and the backup model information is the model information of the model training task at a preset checkpoint; determine the task parameters of the model training task based on the task backup information. This ensures the ability to resume the training progress in case of an anomaly.
[0120] In an optional embodiment of this specification, step 208 includes the following specific steps:
[0121] Based on the target distribution strategy and the task backup information, determine the sub-tasks to be distributed corresponding to each target training resource;
[0122] Distribute each sub-task to be distributed to the corresponding target training resource for execution to obtain the target model that has completed training.
[0123] The sub-task to be distributed is a task unit that needs to be distributed to each target training resource. The sub-task to be distributed is, during the model training process, according to the target distribution strategy and the task backup information, the overall model training task is decomposed into multiple smaller, independent, and parallel-executable task units corresponding to the target training resources. The sub-task to be distributed enables the model training task to be processed in parallel on multiple target training resources, thereby improving the efficiency of model training.
[0124] Exemplarily, based on the target distribution strategy (64 batches of model training tasks, batches 1 - 4 form a training sub-task and are distributed to one node, batches 5 - 10 form a training sub-task and are distributed to one node... batches 60 - 64 form a training sub-task and are distributed to one node) and the task backup information (the current iteration number, the predicted data of the forward propagation, the gradient value of the backward propagation, the loss function value, the learning rate, etc. of the task progress information, and the backup model information such as the model parameter weights and biases saved at the preset checkpoint), determine the sub-tasks to be distributed corresponding to 10 distributed nodes, distribute batches 1 - 4 to form a training sub-task to node 1, batches 5 - 10 to form a training sub-task to node 2... batches 60 - 64 to form a training sub-task to node 10, and execute the model training task on each node to obtain the large language model that has completed training.
[0125] Based on the target distribution strategy and the task backup information, determine the sub-tasks to be distributed corresponding to each target training resource; distribute each sub-task to be distributed to the corresponding target training resource for execution to obtain the target model that has completed training. While ensuring the ability to resume the training progress in case of an anomaly, according to the target distribution strategy, quickly transition from the current abnormal training state to a new training state, promptly resume model training, process in parallel on multiple target training resources, thereby improving the efficiency of model training and further optimizing the transition cost.
[0126] In an alternative embodiment of this specification, step 206 includes the following specific steps:
[0127] Based on the task parameters of the model training task and the resource parameters of each target training resource, construct at least one distribution strategy;
[0128] Analyze at least one distribution strategy, and obtain the target distribution strategy when the task execution metric reaches the preset metric.
[0129] The distribution strategy is a strategy for distributing the model training task to multiple training resources. A specific distribution strategy generates corresponding task execution metrics during the execution of the model training task. The distribution strategy is determined by the task parameters of the model training task and the resource parameters of each target training resource. The target distribution strategy is the distribution strategy when the task execution metric reaches the preset metric, which can be understood as an optimized distribution strategy, achieving the optimization of the task execution effect.
[0130] Analyzing at least one distribution strategy includes, but is not limited to: mathematical analysis, linear programming, particle swarm optimization, genetic algorithms, and machine learning analysis.
[0131] In an alternative embodiment of this specification, analyzing at least one distribution strategy and obtaining the target distribution strategy when the task execution metric reaches the preset metric includes the following specific steps:
[0132] For at least one distribution strategy, calculate the task execution metric and the task distribution metric under various distribution strategies;
[0133] Under the constraint of the task distribution metric, determine the target distribution strategy when the task execution metric reaches the preset metric.
[0134] The task distribution metric is a quantitative metric for evaluating the distribution efficiency of the model training task of the distribution strategy, including but not limited to: distribution cost. The corresponding distribution cost includes but is not limited to: the number of training subtasks processed per second and the distribution duration.
[0135] When the task execution metric is the number of floating-point operations per second, it is illustrated by the following example:
[0136] In a cluster, there are n target training resources for distributed training, and there are m training subtasks. Our goal is to fully utilize the computing power of these target training resources while satisfying each running training subtask. By default, a graphics processing unit is regarded as a target training resource.
[0137] First, define a metric to measure the training efficiency of a task. This metric is the Weighted Achieved Aggregate FLOP per Second (WAF), which measures the weighted achieved FLOP per second of a training subtask. We define a function F: N×N→R, where F(t, x) represents the WAF of the training subtask when x target training resources are allocated to training subtask t. The most important part of the WAF is the total achieved FLOP / s, denoted as T(t, x), given the training subtask t and x target training resources. Note that T(t, x) reflects the optimized performance of the training subtask, which is obtained through a complex combination of adjusting parallelism hyperparameters and other optimization settings. To address this, we rely on calibrating the training subtask on a cluster of target training resources and using automated execution policy generation techniques to estimate the optimized parallelism settings and the associated T(t, x). In addition to the total achieved FLOP / s, we further incorporate two additional factors to model the minimum computational requirements and task priorities. The requirement condition (Tnecessary(t)) represents the minimum resources required for a given training subtask. The training subtask will only be scheduled if the available target training resources can meet this requirement. The task weight (w(t)) is used to model the priority of the training subtask. Training subtasks with higher priorities are assigned higher weights. By default, we set w(t)=1 for all training subtasks. This weight factor allows users to adjust the relative importance of different training subtasks in the problem, and the recommended range of w(t) values is between 0.5 and 2.0.
[0138] Therefore, the WAF is defined as:
[0139] F(t, x) = w(t)·T(t, x), if (t, x) ≥ Tnecessary(t); 0, otherwise Equation 2
[0140] Here, if (t, x) meets the necessary requirement Tnecessary(t), then F(t, x) is calculated as the product of the weight of the training subtask w(t) and its total achieved FLOP / s T(t, x). Otherwise, if it is below this threshold, the WAF of the training subtask is considered zero. The optimization objective is centered around a trade-off: maximizing the WAF of the reconfigured cluster while minimizing the impact on the graphics processing unit during the transition.
[0141] Our goal is to optimize the distribution strategy, construct at least one distribution strategy based on the task parameters of the model training task and the resource parameters of each target training resource. Then parse these distribution strategies to obtain the target distribution strategy when the task execution metrics reach the preset metrics. The specific parsing method can be achieved by parsing the cumulative rewards of each distribution strategy while minimizing the task distribution metrics. This requires finding a balance between maximizing the rewards during healthy operation and restricting the scope of reconfiguration to keep as many training subtasks running as possible. The specific calculation formula is shown in Formula 3:
[0142]
[0143] where G(t i , x′ i ) = F(t i , x′ i ) · D running (n′) - F(t i , x i ) · l(t i , x i → x′ i ) · D transition ,
[0144]
[0145] where the superscript ′ is used to distinguish the states before and after configuration. The subscript i represents the training subtask identifier, t i represents the i-th training subtask, and _x i is the number of initially allocated target training resources. This constraint ensures that the target training resources allocated to all training subtasks do not exceed the total available amount in the cluster. The goal is to maximize the cumulative reward of the cluster, which is represented by the sum of the rewards G(t i , x' i ) of each training subtask. The main term G(t i , x' i ) reflects the reward when the cluster is running after reconfiguration, calculated as the WAF of the training subtask t i using x' i target training resources, multiplied by the expected running duration Drunning(n0). This duration depends on the running status of the target training resources and the cluster size; the larger the cluster size of the target training resources, the higher the probability of anomalies, which may shorten the running duration. The constraint term F(t i , x i ) · 1(t i , x i → x' i ) · D transitionDetermine the WAF loss during distribution, where Dtransition is the estimated duration of this distribution model training task. The indicator function l(t i , x i →x' i ) is activated when the number of target training resources allocated to the training subtask t i changes or a target training resource anomaly occurs, as shown in Equation 4:
[0146] l(t i , x i →x' i ) = 1, if x i ≠x' i or the target training resource t i is abnormal; 0, otherwise Equation 4
[0147] By considering the WAF that can be achieved by the unaffected target training resources themselves, frequent reconfiguration is prevented, depending on the specific situation. In a cluster of target training resources with a smaller scale or higher reliability, maximizing the WAF gives priority to the rewards during healthy operation. On the contrary, in a cluster of target training resources with a larger scale or lower reliability, where anomalies are more common, the constraint term becomes more important, emphasizing the necessity of restricting the scope of reconfiguration and maintaining the operation of training subtasks as much as possible. Keep as many training subtasks running as possible.
[0148] Exemplarily, based on the task parameters of the model training task (large language model architecture, 1TB training dataset, batch size of 64, learning rate of 0.001, number of iterations of 100) and the resource parameters of each target training resource (10 distributed nodes, with 4 graphics processing units of model A, 1 64-core central processing unit, 512GB of memory, and 10Gbps network bandwidth deployed on any distributed node), 20 distribution strategies are constructed. Using Equation 2, Equation 3, and Equation 4, the weighted cumulative floating-point operations per second (WAF) under various distribution strategies are calculated, and the target distribution strategy is determined when the weighted cumulative floating-point operations per second (WAF) reaches the weighted cumulative floating-point operations per second threshold (higher than 300T FLOPS per second of floating-point operations): For a model training task of 64 batches, batches 1 - 4 form a training subtask and are distributed to one node, batches 5 - 10 form a training subtask and are distributed to one node... batches 60 - 64 form a training subtask and are distributed to one node.
[0149] Construct at least one distribution strategy based on the task parameters of the model training task and the resource parameters of each target training resource; parse at least one distribution strategy, and obtain the target distribution strategy when the task execution metric reaches a preset metric. By constructing the distribution strategy and parsing each distribution strategy, the resource parameters of the target training resources except for the abnormal training resources are comprehensively considered, combined with the task parameters of the model training task, the target distribution strategy when the task execution metric reaches the preset metric is determined, the sub-healthy cost is optimized, and the global training efficiency is further improved.
[0150] In an optional embodiment of this specification, after step 206, the following specific steps are further included:
[0151] Send the target distribution strategy to the front end;
[0152] Correspondingly, step 208 includes the following specific steps:
[0153] In the case of receiving the confirmation execution instruction fed back by the front end, distribute the model training task to each target training resource for execution according to the target distribution strategy, and obtain the target model that has completed training.
[0154] The confirmation execution instruction is an instruction signal sent by the front end for confirming the execution of the target distribution strategy. In a distributed training environment, after determining the target distribution strategy, the target distribution strategy will be sent to the front end for display and review. After the front-end user (such as an administrator or a data scientist) views and evaluates whether the target distribution strategy meets the expectations and requirements, they choose whether to confirm the execution. The confirmation execution instruction is that the user clearly distributes and executes the model training task according to the target distribution strategy.
[0155] The target distribution strategy can be sent to the front end by means of API or message queue, etc.
[0156] Optionally, after sending the target distribution strategy to the front end, the following is further included:
[0157] Receive the modified distribution strategy fed back by the front end, where the modified distribution strategy is obtained by the user modifying the target distribution strategy on the front end.
[0158] Exemplarily, the target distribution strategy is as follows: For 64 batches of model training tasks, batches 1 - 4 form a training subtask and are distributed to one node, batches 5 - 10 form a training subtask and are distributed to one node... batches 60 - 64 form a training subtask and are distributed to one node. This target distribution strategy is sent to the front - end user interface through the API. On the front - end user interface, the administrator sees the target distribution strategy. The administrator confirms that the target distribution strategy meets the expectations and requirements on the front - end user interface, clicks the "Confirm Execution" control, and sends a confirmation execution instruction to the system. In the case of receiving the confirmation execution instruction fed back from the front - end, based on the target distribution strategy, batches 1 - 4 form a training subtask and are distributed to node 1, batches 5 - 10 form a training subtask and are distributed to node 2... batches 60 - 64 form a training subtask and are distributed to node 10, and the model training tasks are executed on each node to obtain the large - language model that has completed training.
[0159] Send the target distribution strategy to the front - end; in the case of receiving the confirmation execution instruction fed back from the front - end, distribute the model training tasks to each target training resource for execution according to the target distribution strategy to obtain the target model that has completed training. Through an interactive method, the interactive communication inside and outside the system is realized, improving the user experience.
[0160] In an optional embodiment of this specification, after step 204, the following specific steps are further included:
[0161] Generate an exception prompt message based on the resource information of the abnormal training resource;
[0162] Send the exception prompt message to the front - end.
[0163] The resource information of the abnormal training resource is the identification information and parameter information of the training resource with abnormal resource status during the model training process. For example, the identifier of node 2 is "node - 2", the current temperature of the graphics processing unit exceeds the safety threshold of 80 °C, the memory usage rate exceeds 90%, the remaining disk space is less than 10 GB, and the network connection is unstable.
[0164] The exception prompt message is generated based on the resource information of the abnormal training resource and is used to notify the front - end of the abnormal situation of the training resource. The exception prompt message should clearly and accurately describe the nature of the problem, the scope of influence, and possible solutions so that users can quickly understand and respond to the problem. For example, the exception prompt message: Node 2 "node - 2" has an exception. Resource information: Temperature: 82 °C (exceeding the safety threshold of 80 °C); Memory usage rate: 95%; Remaining disk space: 9.5 GB; Network connection: Unstable (packet loss rate 5%).
[0165] Generate an exception prompt message based on the resource information of the abnormal training resource; send the exception prompt message to the front end. The exception prompt message contains the resource information of the abnormal training resource, which can help users quickly locate the abnormal training resource and abnormal recovery. The interaction and communication inside and outside the system are realized, and the user experience is improved.
[0166] In addition to the above two interaction methods, the following interaction methods are also included: integrating target training resources, and processing the submission / termination of training tasks.
[0167] Figure 3 The flowchart of a model training method provided by an embodiment of this specification is shown, as Figure 3 shown:
[0168] The model training system includes a training scheduling end, multiple nodes (Node 1, Node 2, Node 3, Node 4... Node n), a persistent storage array, and a network cloud configuration module.
[0169] Among them, the network cloud configuration module controls each node through the control flow, and collective communication is built between each node. There is a data flow between the central processing unit and memory on any node and the persistent storage array, and the task backup information is stored in the persistent storage arrangement through the data flow to obtain the persistent storage array checkpoint.
[0170] Taking Node 1 as an example, 4 graphics processing units (Image Processing Unit-1, Image Processing Unit-2, Image Processing Unit-3, and Image Processing Unit-4) are deployed on Node 1, and a training process runs on each image processing unit. Collective communication is built between the 4 graphics processing units. For each training thread on the central processing unit of Node 1, a dedicated monitoring thread is assigned to implement status information monitoring / reporting. The central processing unit writes the task backup information into the memory to obtain the memory checkpoint, and the central processing unit distributes the training subtasks to the graphics processing units to run the training process to execute the training.
[0171] The training scheduling end includes a status information monitoring module, an abnormal operation module, and a policy generation module. In the status information monitoring module, the status information of the graphics processing unit is received through the monitoring thread. In the abnormal operation module, based on the status information of the graphics processing unit, the abnormal type of the abnormal training resource is determined, and based on the abnormal type, the target operation information for the abnormal training resource is determined, and the target operation is executed. In the policy generation module, when the target operation information is to reconfigure the training resources, based on the task parameters of the model training task and the resource parameters of each target training resource, the target distribution policy when the task execution index reaches the preset index is determined. Finally, according to the target distribution policy, the model training task is distributed to each target training resource for execution to obtain the target model that has completed training.
[0172] See Figure 4 , Figure 4 which shows a flowchart of a task processing method provided by an embodiment of this specification, including the following specific steps:
[0173] Step 402: Obtain the task data of the task to be processed.
[0174] Step 404: Input the task data into the task processing model to obtain the task processing result of the task to be processed. Among them, the model training of the task processing model includes the following steps:
[0175] Receive the status information sent by each training resource participating in the model training task;
[0176] Based on the status information of each training resource, identify abnormal training resources;
[0177] Based on the task parameters of the model training task and the resource parameters of each target training resource, determine the target distribution strategy when the task execution index reaches the preset index. Among them, the target training resources are the training resources except the abnormal training resources;
[0178] Distribute the model training task to each target training resource for execution according to the target distribution strategy, and obtain the task processing model that has completed training.
[0179] The embodiments of this specification are applied to applications, websites or applets with task processing functions.
[0180] The task to be processed is a to-be-executed task implemented based on a machine learning model, including natural language processing tasks, image processing tasks, etc. For example, for natural language processing tasks, there are entity recognition (entity extraction) tasks, question answering tasks, reasoning tasks, text classification tasks, text-to-image tasks, etc. For example, for image processing tasks, there are image classification tasks, object detection tasks, image segmentation tasks, image enhancement tasks, image denoising tasks, style transfer tasks, etc.
[0181] The task data of the task to be processed is the input data for executing the task to be processed, including the input data of natural language processing tasks, the input data of image processing tasks, etc. For example, for natural language processing tasks, there are texts to be recognized, question texts, texts to be reasoned, texts to be classified, prompt texts, etc. For example, for image processing tasks, there are images to be classified, images to be detected, images to be segmented, images to be enhanced, noisy images, images with the initial style, etc.
[0182] The task processing model is a machine learning model trained based on the above model training method. The task processing model has a large number of model parameters. From the perspective of model architecture, the task processing model includes but is not limited to: deep self-attention model, deep self-attention model with bidirectional encoding representation, and large language model. From the perspective of the modality of task data, the task processing model includes but is not limited to: language processing model, image processing model, speech processing model, code processing model, etc. Taking the language processing model as an example, the language processing model can perform one or more natural language processing tasks, including but not limited to: machine translation task, speech recognition task, text analysis task, or text question-and-answer task.
[0183] The task processing result of the task to be processed is the output data of the task processing model executing the task to be processed, including the output data of natural language processing tasks, image processing tasks, etc. For example, for natural language processing tasks, there are text recognition results, response texts, inference results, classification results, generated images, etc. For example, for image processing tasks, there are classification results, detection results, segmentation results, enhanced images, denoised images, images in the target style, etc.
[0184] The model training of the task processing model has been described in detail in the above Figure 1 description of the embodiments of the specification and will not be elaborated here.
[0185] Exemplarily, the task to be processed is a text-to-image task, and the task processing model is a large language model trained based on the above model training method. Receive the prompt text input by the user: An orange cat is sunning itself on the green grass. Input the prompt text into the large language model, generate the corresponding picture, and feedback the picture to the front-end for rendering.
[0186] In the embodiments of this specification, task data of a task to be processed is obtained; the task data is input into a task processing model to obtain a task processing result of the task to be processed, where the task processing model is trained based on the above model training method. By directly receiving the status information sent by each training resource participating in the model training task and identifying abnormal training resources based on the status information of each training resource, fine-grained anomaly monitoring is achieved, abnormal training resources can be identified in a timely manner, and the monitoring cost is optimized; based on the task parameters of the model training task and the resource parameters of each target training resource, a target distribution strategy is determined when the task execution index reaches a preset index. The resource parameters of the target training resources other than the abnormal training resources are comprehensively considered and combined with the task parameters of the model training task to determine the target distribution strategy when the task execution index reaches the preset index, optimizing the sub-healthy cost and improving the global training efficiency; according to the target distribution strategy, the model training task is distributed to each target training resource for execution to obtain a target model that has completed training. According to the target distribution strategy, it quickly transitions from the current abnormal training state to a new training state, promptly resumes model training, optimizes the transition cost, realizes a comprehensive, efficient, and automatic recovery of the model training in case of anomalies, reduces the cost of model training, improves the reliability of model training, and processes the task data through the highly reliable task processing model obtained through training to obtain a task processing result, improving the accuracy and reliability of task processing.
[0187] The following combines the attached Figure 5 , taking the application of the model training method provided in this specification in a large language model as an example, further illustrates the model training method. Among them, Figure 5 FIG. shows a process flow chart of a model training method applied to a large language model provided by an embodiment of this specification, including the following specific steps:
[0188] Step 502: Receive the resource status information sent by each distributed node participating in the model training task for the large language model through a monitoring thread pre-deployed on the central processing unit of the distributed node.
[0189] Step 504: Identify abnormal nodes based on the resource status information of each distributed node and determine the abnormal types of the abnormal nodes.
[0190] Step 506: Determine the target impact level corresponding to the abnormal node based on the abnormal type, and determine the target operation information for the abnormal node based on the target impact level, where the target operation information includes re-training operation, restarting the training process operation, and reconfiguring the training resource operation.
[0191] Step 508: When the target operation information is to reconfigure distributed nodes, obtain the task backup information recorded during the execution of the model training task. Based on the task backup information, determine the task parameters of the model training task. Based on the task parameters of the model training task and the resource parameters of each target node, construct at least one distribution strategy. For at least one distribution strategy, calculate the weighted cumulative floating-point operations per second (FLOPS) under various distribution strategies, and determine the target distribution strategy whose weighted cumulative FLOPS reaches the weighted cumulative FLOPS threshold.
[0192] Step 510: Based on the target distribution strategy and the task backup information, determine the sub-tasks to be distributed corresponding to each target node, and distribute each sub-task to be distributed to the corresponding target node for execution to obtain the large language model that has completed training.
[0193] In the embodiments of this specification, through in-band error monitoring, it is used for real-time error identification without adding additional overhead; through the target distribution strategy generation mechanism, it is used to achieve optimized configuration reorganization; and through the fast transition strategy, it is used to reduce the downtime during state changes, respectively optimizing the monitoring cost, sub-healthy cost, and transition cost, realizing comprehensive, efficient, and automatic recovery of the model training with abnormal conditions, reducing the cost of model training, and improving the reliability of model training.
[0194] Corresponding to the above method embodiments, this specification also provides embodiments of a model training system. Figure 6 It shows a schematic structural diagram of a model training system provided by an embodiment of this specification. As Figure 6 shown, the system includes multiple training resources 602 participating in the model training task and a training scheduling end 604;
[0195] The multiple training resources 602 are used to send status information to the training scheduling end 604;
[0196] The training scheduling end 604 is used to receive the status information sent by each training resource 602, identify the abnormal training resources 602 based on the status information of each training resource 602, determine the target distribution strategy when the task execution metrics reach the preset metrics based on the task parameters of the model training task and the resource parameters of each target training resource 606, and distribute the model training task to each target training resource 606 according to the target distribution strategy, where the target training resource 606 is the training resource other than the abnormal training resource;
[0197] Each target training resource 606 is used to execute the model training task to obtain the target model that has completed training.
[0198] Optionally, the system further includes a front end;
[0199] The training scheduling terminal 604 is further configured to send the target distribution policy to the front end;
[0200] The front end is configured to receive the target distribution policy sent by the training scheduling terminal 604 and feedback a confirmation execution instruction for the target distribution policy to the training scheduling terminal 604;
[0201] The training scheduling terminal 604 is specifically configured to, when receiving the confirmation execution instruction fed back by the front end, distribute the model training task to each target training resource 606 for execution according to the target distribution policy, and obtain the target model that has completed training.
[0202] In the embodiments of this specification, a model training system is provided. The system includes a plurality of training resources participating in the model training task and a training scheduling terminal; the plurality of training resources are configured to send status information to the training scheduling terminal; the training scheduling terminal is configured to receive the status information sent by each training resource, identify abnormal training resources based on the status information of each training resource, determine a target distribution policy when the task execution index reaches a preset index based on the task parameters of the model training task and the resource parameters of each target training resource, and distribute the model training task to each target training resource according to the target distribution policy, where the target training resources are the training resources other than the abnormal training resources; each target training resource is configured to execute the model training task and obtain the target model that has completed training. By directly receiving the status information sent by each training resource participating in the model training task and identifying abnormal training resources based on the status information of each training resource, fine-grained anomaly monitoring is achieved, abnormal training resources can be identified in a timely manner, and the monitoring cost is optimized; based on the task parameters of the model training task and the resource parameters of each target training resource, a target distribution policy is determined when the task execution index reaches a preset index, comprehensively considering the resource parameters of the target training resources other than the abnormal training resources and combining them with the task parameters of the model training task, a target distribution policy is determined when the task execution index reaches a preset index, the sub-healthy cost is optimized, and the global training efficiency is improved; according to the target distribution policy, the model training task is distributed to each target training resource for execution, and the target model that has completed training is obtained. According to the target distribution policy, quickly transition from the current abnormal training state to a new training state, restore the model training in a timely manner, optimize the transition cost, achieve a comprehensive, efficient, and automatic recovery of the model training in case of anomalies, reduce the cost of model training, and improve the reliability of model training.
[0203] The above is a schematic solution of a model training system according to this embodiment. It should be noted that the technical solution of this model training system and the technical solution of the above model training method belong to the same concept. For the details not described in detail in the technical solution of the model training system, reference can be made to the description of the technical solution of the above model training method.
[0204] Corresponding to the above method embodiments, this specification also provides model training device embodiments. Figure 7 The structure diagram of a model training device provided by an embodiment of this specification is shown. As Figure 7 shown, the device includes:
[0205] A monitoring thread 702, configured to receive status information sent by each training resource participating in the model training task;
[0206] A status information monitoring module 704, configured to identify abnormal training resources based on the status information of each training resource;
[0207] A policy generation module 706, configured to determine a target distribution policy when the task execution index reaches a preset index based on the task parameters of the model training task and the resource parameters of each target training resource, where the target training resources are the training resources other than the abnormal training resources;
[0208] A policy execution module 708, configured to distribute the model training task to each target training resource for execution according to the target distribution policy, and obtain a target model that has completed training.
[0209] Optionally, the device further includes: an abnormal operation module, configured to determine the abnormal type of the abnormal training resource based on the status information; and determine the target operation information for the abnormal training resource based on the abnormal type;
[0210] Correspondingly, the policy generation module 706 is further configured to: when the target operation information is to reconfigure the training resource, determine a target distribution policy when the task execution index reaches a preset index based on the task parameters of the model training task and the resource parameters of each target training resource.
[0211] Optionally, the abnormal operation module is further configured to: determine the target impact level corresponding to the abnormal training resource based on the abnormal type; and determine the target operation information for the abnormal training resource based on the target impact level.
[0212] Optionally, the device further includes: a backup acquisition module, configured to acquire task backup information recorded during the execution of the model training task, where the task backup information includes task progress information and backup model information, and the backup model information is the model information of the model training task at a preset checkpoint; and determine the task parameters of the model training task based on the task backup information.
[0213] Optionally, the policy execution module 708 is further configured to: determine the sub-tasks to be distributed corresponding to each target training resource based on the target distribution policy and the task backup information; and distribute each sub-task to be distributed to the corresponding target training resource for execution, and obtain a target model that has completed training.
[0214] Optionally, the policy generation module 706 is further configured to: construct at least one distribution policy based on the task parameters of the model training task and the resource parameters of each target training resource; parse at least one distribution policy, and obtain a target distribution policy when the task execution metric reaches a preset metric.
[0215] Optionally, the policy generation module 706 is further configured to: calculate the task execution metric and the task distribution metric under various distribution policies for at least one distribution policy; determine the target distribution policy when the task execution metric reaches the preset metric under the constraint of the task distribution metric.
[0216] Optionally, the apparatus further includes: a first interaction module, configured to send the target distribution policy to the front end; correspondingly, the policy generation module 706 is further configured to: when receiving a confirmation execution instruction fed back by the front end, distribute the model training task to each target training resource for execution according to the target distribution policy, and obtain a target model that has completed training.
[0217] Optionally, the apparatus further includes: a second interaction module, configured to generate an exception prompt message based on the resource information of the abnormal training resource; send the exception prompt message to the front end.
[0218] In the embodiments of this specification, the status information sent by each training resource participating in the model training task is directly received, and based on the status information of each training resource, the abnormal training resources are identified, realizing fine-grained exception monitoring, and the abnormal training resources can be identified in time, optimizing the monitoring cost; based on the task parameters of the model training task and the resource parameters of each target training resource, the target distribution policy when the task execution metric reaches the preset metric is determined, comprehensively considering the resource parameters of the target training resources other than the abnormal training resources, combined with the task parameters of the model training task, and the target distribution policy when the task execution metric reaches the preset metric is determined, optimizing the sub-healthy cost and improving the global training efficiency; according to the target distribution policy, the model training task is distributed to each target training resource for execution, and a target model that has completed training is obtained. According to the target distribution policy, quickly transition from the current abnormal training state to a new training state, and restore the model training in time, optimizing the transition cost, realizing a comprehensive, efficient, and automatic restoration of the model training with abnormal conditions, reducing the cost of model training, and improving the reliability of model training.
[0219] The above is a schematic solution of a model training apparatus according to this embodiment. It should be noted that the technical solution of this model training apparatus and the technical solution of the above model training method belong to the same concept. For the details not described in detail in the technical solution of the model training apparatus, reference can be made to the description of the technical solution of the above model training method.
[0220] Corresponding to the above method embodiments, this specification also provides embodiments of a task processing apparatus. Figure 8 The following shows a schematic structural diagram of a task processing apparatus provided by an embodiment of this specification. As Figure 8 shown, the apparatus includes:
[0221] An acquisition model 802, which acquires task data of a task to be processed;
[0222] A processing module 804, configured to input the task data into a task processing model to obtain a task processing result of the task to be processed. Among them, the model training of the task processing model includes the following steps: receiving status information sent by each training resource participating in the model training task; identifying abnormal training resources based on the status information of each training resource; determining a target distribution strategy when the task execution index reaches a preset index based on the task parameters of the model training task and the resource parameters of each target training resource, where the target training resources are the training resources other than the abnormal training resources; distributing the model training task to each target training resource for execution according to the target distribution strategy to obtain a task processing model that has completed training.
[0223] In the embodiments of this specification, directly receiving the status information sent by each training resource participating in the model training task and identifying abnormal training resources based on the status information of each training resource realizes fine-grained anomaly monitoring, can timely identify abnormal training resources, and optimizes the monitoring cost; determining a target distribution strategy when the task execution index reaches a preset index based on the task parameters of the model training task and the resource parameters of each target training resource, comprehensively considering the resource parameters of the target training resources other than the abnormal training resources and combining them with the task parameters of the model training task to determine the target distribution strategy when the task execution index reaches a preset index, optimizes the sub-healthy cost, and improves the global training efficiency; distributing the model training task to each target training resource for execution according to the target distribution strategy to obtain a target model that has completed training, quickly transitioning from the current abnormal training state to a new training state according to the target distribution strategy, timely resuming model training, optimizing the transition cost, realizing comprehensive, efficient, and automatic recovery of the model training that has encountered abnormal conditions, reducing the cost of model training, enhancing the reliability of model training, and processing the task data through the highly reliable task processing model obtained through training to obtain a task processing result, improving the accuracy and reliability of task processing.
[0224] The above is a schematic solution of a task processing apparatus in this embodiment. It should be noted that the technical solution of this task processing apparatus and the technical solution of the above task processing method belong to the same concept. For the details not described in detail in the technical solution of the task processing apparatus, reference can be made to the description of the technical solution of the above task processing method.
[0225] Figure 9 The block diagram of a computing device provided by an embodiment of this specification is shown. The components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920. The processor 920 is connected to the memory 910 through a bus 930, and a database 950 is used to store data.
[0226] The computing device 900 further includes an access device 940, and the access device 940 enables the computing device 900 to communicate via one or more networks 960. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interfaces (for example, a Network Interface Controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).
[0227] In an embodiment of this specification, the above components of the computing device 900 and Figure 9 other components not shown may also be connected to each other, for example, through a bus. It should be understood that Figure 9 the shown block diagram of the computing device structure is only for example purposes and is not a limitation on the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0228] The computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 900 can also be a mobile or stationary server.
[0229] The processor 920 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-described model training method or task processing method.
[0230] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above-described model training method and task processing method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above-described model training method or task processing method.
[0231] This specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-described model training method or task processing method.
[0232] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above-described model training method and task processing method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above-described model training method or task processing method.
[0233] This specification also provides a computer program, which, when executed on a computer, causes the computer to execute the steps of the above-described model training method or task processing method.
[0234] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solutions of the above-described model training method and task processing method belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the descriptions of the technical solutions of the above-described model training method or task processing method.
[0235] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0236] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0237] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0238] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0239] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all details and do not limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A model training method, comprising: Receiving status information sent by each training resource participating in the model training task; Identifying abnormal training resources based on the status information of each training resource; Determining a target distribution strategy when the task execution index reaches a preset index based on the task parameters of the model training task and the resource parameters of each target training resource, wherein the target training resources are the training resources other than the abnormal training resources; Distributing the model training task to each target training resource for execution according to the target distribution strategy to obtain a target model that has completed training.
2. The method according to claim 1, before determining the target distribution strategy when the task execution index reaches a preset index based on the task parameters of the model training task and the resource parameters of each target training resource, further comprising: Determining the abnormal type of the abnormal training resource based on the status information; Determining target operation information for the abnormal training resource based on the abnormal type; The determining the target distribution strategy when the task execution index reaches a preset index based on the task parameters of the model training task and the resource parameters of each target training resource includes: When the target operation information is to reconfigure the training resources, determining the target distribution strategy when the task execution index reaches a preset index based on the task parameters of the model training task and the resource parameters of each target training resource.
3. The method according to claim 2, the determining the target operation information for the abnormal training resource based on the abnormal type includes: Determining the target impact level corresponding to the abnormal training resource based on the abnormal type; Determining the target operation information for the abnormal training resource based on the target impact level.
4. The method according to claim 1, before determining the target distribution strategy when the task execution index reaches a preset index based on the task parameters of the model training task and the resource parameters of each target training resource, further comprising: Obtaining task backup information recorded during the execution of the model training task, wherein the task backup information includes task progress information and backup model information, and the backup model information is the model information of the model training task at a preset checkpoint; Determining the task parameters of the model training task based on the task backup information.
5. The method according to claim 4, the distributing the model training task to each target training resource for execution according to the target distribution strategy to obtain a target model that has completed training includes: Determining the sub-tasks to be distributed corresponding to each target training resource based on the target distribution strategy and the task backup information; Distributing each sub-task to be distributed to the corresponding target training resource for execution to obtain a target model that has completed training.
6. The method according to any one of claims 1-5, the determining the target distribution strategy when the task execution index reaches a preset index based on the task parameters of the model training task and the resource parameters of each target training resource includes: Construct at least one distribution policy based on the task parameters of the model training task and the resource parameters of each target training resource; Parse the at least one distribution policy, and obtain a target distribution policy when the task execution metric reaches a preset metric.
7. The method according to claim 6, wherein the parsing the at least one distribution policy and obtaining a target distribution policy when the task execution metric reaches a preset metric includes: For the at least one distribution policy, calculate the task execution metric and the task distribution metric under various distribution policies; Under the constraint of the task distribution metric, determine the target distribution policy for which the task execution metric reaches the preset metric.
8. The method according to any one of claims 1-5, after determining the target distribution policy when the task execution metric reaches the preset metric based on the task parameters of the model training task and the resource parameters of each target training resource, further includes: Send the target distribution policy to the front end; The distributing the model training task to each target training resource for execution according to the target distribution policy to obtain a target model that has completed training includes: When receiving the confirmation execution instruction fed back by the front end, distribute the model training task to each target training resource for execution according to the target distribution policy to obtain a target model that has completed training.
9. The method according to any one of claims 1-5, after identifying an abnormal training resource based on the status information of each training resource, further includes: Generate an abnormal prompt message based on the resource information of the abnormal training resource; Send the abnormal prompt message to the front end.
10. A task processing method includes: Obtain the task data of the task to be processed; Input the task data into a task processing model to obtain the task processing result of the task to be processed, wherein the model training of the task processing model includes the following steps: Receive the status information sent by each training resource participating in the model training task; Identify abnormal training resources based on the status information of each training resource; Determine the target distribution policy when the task execution metric reaches the preset metric based on the task parameters of the model training task and the resource parameters of each target training resource, wherein the target training resource is the training resource other than the abnormal training resource; Distribute the model training task to each target training resource for execution according to the target distribution policy to obtain a task processing model that has completed training.
11. A model training system, the system includes a plurality of training resources participating in the model training task and a training scheduling terminal; The plurality of training resources are used to send status information to the training scheduling terminal; The training scheduling terminal is configured to receive the status information sent by each training resource, identify abnormal training resources based on the status information of each training resource, determine a target distribution strategy when the task execution metrics reach the preset metrics based on the task parameters of the model training task and the resource parameters of each target training resource, and distribute the model training task to each of the target training resources according to the target distribution strategy, where the target training resources are the training resources other than the abnormal training resources; Each of the target training resources is configured to execute the model training task to obtain a target model that has completed training.
12. The system according to claim 11, wherein the system further comprises a front end; The training scheduling terminal is further configured to send the target distribution strategy to the front end; The front end is configured to receive the target distribution strategy sent by the training scheduling terminal and feedback a confirmation execution instruction for the target distribution strategy to the training scheduling terminal; The training scheduling terminal is specifically configured to, when receiving the confirmation execution instruction fed back by the front end, distribute the model training task to each of the target training resources for execution according to the target distribution strategy to obtain a target model that has completed training.
13. A computing device, comprising: a memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. A computer-readable storage medium storing computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Cited By
Performance monitoring system and performance monitoring method
CN121501604A
Robot control method and system thereof, medium, equipment and program product
CN121515155A