A model training system, a model training task execution method, a device and a medium

By introducing management clusters into the model training system, the future state of the training cluster is predicted in real time, and through virtual training cluster simulation execution strategies, the problem of difficulty in capturing system state changes and optimizing resource utilization in the existing technology is solved, and efficient model training task execution is achieved.

CN119918624BActive Publication Date: 2025-06-06ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510404732.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-06
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

During the training of large-scale models, the existing technology is difficult to capture system state changes in real time, and cannot effectively predict equipment health status and performance bottlenecks, resulting in poor resource utilization and extended training cycles.

Method used

Provide a model training system, including training clusters and managing clusters. The management cluster predicts future states by obtaining real-time state data of the training cluster, and initializes the virtual training cluster. Generate execution strategies according to the predicted state, execute model training tasks through virtual training cluster simulation, determine the target strategy, and adjust the execution of the training cluster according to the target strategy.

Benefits of technology

Real-time prediction and optimization of training cluster states are realized, resource utilization and training efficiency are improved, and normal execution of model training tasks is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918624B_ABST
    Figure CN119918624B_ABST
Patent Text Reader

Abstract

The present application discloses a model training system, a method, an apparatus and a medium for executing a model training task. The management cluster in the model training system can obtain the real-time status data of the training cluster when executing the model training task. Through the real-time status data, the state of the training cluster when executing the model training task in the future set time period is predicted. The management cluster determines each simulator corresponding to each device included in the training cluster, and initializes the virtual training cluster corresponding to the training cluster through these simulators. The management cluster generates several execution strategies for the virtual training cluster according to the predicted state, and executes the model training task through the virtual training cluster simulation according to at least some of these execution strategies, obtains the performance indicators corresponding to at least some of the execution strategies, and determines the target strategy according to the obtained performance indicators, so as to execute the model training task according to the target strategy, thereby effectively improving the efficiency of the entire training cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of computer technology and artificial intelligence, and in particular to a model training system, a model training task execution method, a device and a medium. Background Art

[0002] With the continuous development of deep learning technology, the training of large models has gradually become an important engine to promote the advancement of artificial intelligence technology. Large model training usually consumes a lot of computing resources and relies on large high-performance computing clusters (including GPU / CPU nodes, network switching equipment, and storage systems) to work together. However, as the model scale and training cycle continue to grow, the operating status of cluster equipment becomes more complex, and system performance bottlenecks and equipment health issues become increasingly prominent. During weeks or even months of continuous training, if key nodes fail or performance degrades severely, it often leads to training interruptions, delays, or waste of resources.

[0003] In order to avoid the above problems, some scheduling strategies can be used to implement computing resource scheduling, so as to allocate the computing resources required in the large model training process. However, static or semi-dynamic scheduling strategies are currently used in large model training scenarios. These methods generally have the following problems: First, it is difficult to capture system status changes in real time, and there is a lack of flexible online prediction and rapid response methods for equipment health status and performance bottlenecks; second, the task and resource matching capabilities are limited, and the scheduling and allocation of training tasks cannot be quickly optimized according to the latest status information, resulting in poor resource utilization and extended training cycles; and there is a lack of low-overhead test and optimization sites. It is costly to try different strategies in the actual environment (such as pausing training, redeploying, etc.), and it is difficult to quickly evaluate the pros and cons of multiple strategies; finally, existing systems usually design task scheduling and status / fault management as independent modules, lacking a collaborative optimization mechanism between the two, which further limits the overall efficiency of the cluster. Summary of the invention

[0004] The embodiments of the present application provide a model training system, a model training task execution method, a device and a medium to partially solve the above-mentioned problems existing in the prior art.

[0005] This application adopts the following technical solutions:

[0006] The embodiment of the present application provides a model training system, which includes: a training cluster and a management cluster;

[0007] The training cluster is used to perform model training tasks;

[0008] The management cluster is used to obtain real-time status data of the training cluster when executing the model training task, and predict the status of the training cluster when executing the model training task within a set time period in the future based on the real-time status data, and determine each simulator corresponding to each device included in the training cluster as the predicted status, so as to initialize a virtual training cluster corresponding to the training cluster based on each simulator, and generate several execution strategies for the virtual training cluster based on the predicted status, and simulate the execution of the model training task through the virtual training cluster according to at least some of the several execution strategies to obtain performance indicators corresponding to at least some of the execution strategies, and determine a target strategy from the several execution strategies based on the performance indicators corresponding to at least some of the execution strategies, and execute the model training task through the training cluster according to the target strategy, wherein the virtual training cluster uses different virtual devices when simulating the execution of the model training task according to different execution strategies, and / or the virtual training cluster uses different data processing strategies when simulating the execution of the model training task according to different execution strategies.

[0009] Optionally, the model training system further includes:

[0010] A storage cluster, used to store real-time status data when the training cluster performs the model training task;

[0011] The management cluster is specifically used to obtain the real-time status data from the storage cluster.

[0012] Optionally, the management cluster is specifically used to obtain task parameters corresponding to the model training task, and predict the state of the training cluster when executing the model training task within a future set time based on the task parameters and the real-time status data.

[0013] Optionally, the management cluster is specifically used to select a candidate execution strategy from the several execution strategies, simulate and execute the model training task through the virtual training cluster according to the candidate execution strategy, obtain performance indicators corresponding to the candidate execution strategy, and reselect the candidate execution strategy from the remaining execution strategies according to the performance indicators corresponding to the candidate execution strategy until the target strategy is selected.

[0014] Optionally, the performance indicators include: simulation time spent on executing the model training task through the virtual training cluster simulation, simulation resource utilization when executing the model training task through the virtual training cluster simulation, and simulation load when executing the model training task through the virtual training cluster simulation;

[0015] The management cluster is specifically used to determine the function value of a preset objective function based on the simulation duration corresponding to the candidate execution strategy and the weight corresponding to the simulation duration, the simulation resource utilization corresponding to the candidate execution strategy and the weight corresponding to the simulation resource utilization, and the simulation load corresponding to the candidate execution strategy and the weight corresponding to the simulation load, wherein, the longer the simulation duration, the larger the function value, the larger the simulation resource utilization, the smaller the function value, and the larger the simulation load, the larger the function value; according to the function value, the posterior distribution of the objective function is fitted, and according to the posterior distribution, the candidate execution strategy is reselected from the remaining execution strategies.

[0016] The embodiment of the present application provides a model training execution method, which is applied to a management cluster in a model training system, wherein the model training system further includes a training cluster, including:

[0017] Acquire real-time status data of the training cluster when executing the model training task;

[0018] According to the real-time status data, predict the status of the training cluster when executing the model training task within a future set time period as the predicted status;

[0019] Determine each simulator corresponding to each device included in the training cluster, so as to initialize a virtual training cluster corresponding to the training cluster according to the each simulator;

[0020] According to the predicted state, several execution strategies for the virtual training cluster are generated, wherein the virtual training cluster uses different virtual devices when simulating and executing the model training task according to different execution strategies, and / or the virtual training cluster uses different data processing strategies when simulating and executing the model training task according to different execution strategies;

[0021] According to at least some of the execution strategies, the model training task is executed by the virtual training cluster simulation to obtain performance indicators corresponding to at least some of the execution strategies, and a target strategy is determined from the several execution strategies according to the performance indicators corresponding to at least some of the execution strategies;

[0022] According to the target strategy, the model training task is performed by the training cluster.

[0023] Optionally, according to at least some of the execution strategies, the model training task is executed by the virtual training cluster simulation to obtain performance indicators corresponding to at least some of the execution strategies, and according to the performance indicators corresponding to at least some of the execution strategies, a target strategy is determined from the several execution strategies, specifically including:

[0024] Select a candidate execution strategy from the several execution strategies, execute the model training task through the virtual training cluster simulation according to the candidate execution strategy, obtain the performance indicator corresponding to the candidate execution strategy, and reselect the candidate execution strategy from the remaining execution strategies according to the performance indicator corresponding to the candidate execution strategy until the target strategy is selected.

[0025] Optionally, the performance indicators include: simulation time spent on executing the model training task through the virtual training cluster simulation, simulation resource utilization when executing the model training task through the virtual training cluster simulation, and simulation load when executing the model training task through the virtual training cluster simulation;

[0026] Reselecting a candidate execution strategy from the remaining execution strategies according to the performance indicator corresponding to the candidate execution strategy, specifically including:

[0027] Determine the function value of the preset objective function according to the simulation duration corresponding to the candidate execution strategy and the weight corresponding to the simulation duration, the simulation resource utilization corresponding to the candidate execution strategy and the weight corresponding to the simulation resource utilization, and the simulation load corresponding to the candidate execution strategy and the weight corresponding to the simulation load, wherein the longer the simulation duration, the greater the function value, the greater the simulation resource utilization, the smaller the function value, and the greater the simulation load, the greater the function value;

[0028] According to the function value, a posterior distribution of the objective function is fitted, and according to the posterior distribution, a candidate execution strategy is reselected from the remaining execution strategies.

[0029] The present application embodiment provides a model training task execution device, including:

[0030] The acquisition module is used to obtain real-time status data when the training cluster executes model training tasks;

[0031] A prediction module, used to predict the state of the training cluster when executing the model training task within a future set time period according to the real-time state data, as a predicted state;

[0032] An initialization module, used to determine each simulator corresponding to each device included in the training cluster, so as to initialize a virtual training cluster corresponding to the training cluster according to the each simulator;

[0033] A generation module, used to generate a plurality of execution strategies for the virtual training cluster according to the predicted state, wherein the virtual training cluster uses different virtual devices when simulating and executing the model training task according to different execution strategies, and / or the virtual training cluster uses different data processing strategies when simulating and executing the model training task according to different execution strategies;

[0034] a determination module, configured to perform the model training task through the virtual training cluster simulation according to at least some of the execution strategies, obtain performance indicators corresponding to the at least some of the execution strategies, and determine a target strategy from the several execution strategies according to the performance indicators corresponding to the at least some of the execution strategies;

[0035] An execution module is used to execute the model training task through the training cluster according to the target strategy.

[0036] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned model training task execution method is implemented.

[0037] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:

[0038] The model training system and the model training task execution method provided by the embodiments of the present application are such that the management cluster in the model training system can obtain the real-time status data of the training cluster when executing the model training task, and then, through the real-time status data, the state of the training cluster when executing the model training task within the future set time period can be predicted, and the management cluster can determine each simulator corresponding to each device included in the training cluster, and initialize the virtual training cluster corresponding to the training cluster through these simulators. Afterwards, the management cluster can generate several execution strategies for the virtual training cluster according to the predicted state, and execute the model training task through the virtual training cluster simulation according to at least some of these execution strategies, and obtain the performance indicators corresponding to at least some of the execution strategies, so as to determine the target strategy from the several execution strategies according to the obtained performance indicators, and execute the model training task through the training cluster according to the target strategy.

[0039] It can be seen from the above method that the management cluster in the model training system can predict the state of the training cluster's future execution of the model training task based on the real-time state information of the training cluster executing the model training task, and the management cluster can create a virtual training cluster corresponding to the training cluster, and through the virtual training cluster, simulate several execution strategies determined based on the predicted state, thereby determining the optimal target strategy, and then adjust the training cluster according to the target strategy, or adjust the data processing method when the training cluster executes the model training task. Therefore, the model training system and method provided in the embodiment of the present application can establish an effective prediction mechanism, respond in advance to various conditions that may occur when the training cluster executes the model training task, so as to ensure the normal execution of the model training task, and through the created virtual training cluster, a better execution strategy can be simulated, thereby further improving the efficiency of the entire training cluster in executing the model training task. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The drawings described herein are used to provide a further understanding of the present specification and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present specification and do not constitute an improper limitation on the present application. In the drawings:

[0041] Figure 1 A schematic diagram of the architecture of a model training system provided in an embodiment of the present application;

[0042] Figure 2 A schematic diagram of a flow chart of a model training task execution provided in an embodiment of the present application;

[0043] Figure 3 A schematic diagram of the series connection of the entire model training task execution process provided in the embodiment of the present application;

[0044] Figure 4 A schematic diagram of a model training task execution device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0046] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0047] In order to solve the above problems, an embodiment of the present application provides a model training system, which includes a training cluster and a management cluster, wherein the training cluster is used to execute model training tasks, and the management cluster is used to execute task scheduling. The so-called task scheduling refers to adjusting the training cluster based on the status of the training cluster executing the model training task, or adjusting the data processing method when the training cluster executes the model training task.

[0048] The premise for the management cluster to perform task scheduling is that it needs to obtain data to determine the execution strategy corresponding to the task scheduling. For this reason, the above model training system also has a storage cluster, which is mainly used to manage various data. Among them, the data generated by the training cluster when performing model training tasks can be stored in the storage cluster, and the management cluster can obtain the required data from the storage cluster to determine the execution strategy corresponding to the task scheduling. The architecture of the above model training system is as follows: Figure 1 shown.

[0049] Figure 1 A schematic diagram of the architecture of a model training system provided in an embodiment of the present application.

[0050] from Figure 1 It can be seen that the model training system provided by the embodiment of the present application is mainly composed of three parts, a storage cluster, a management cluster and a training cluster. The management cluster and the training cluster can share the storage cluster, and the storage cluster maintains the various data required in the entire model training task execution process. Therefore, the data required by the training cluster when executing the model training task can be loaded from the storage cluster, and the state information in the form of snapshots generated by the training cluster during the execution of the model training task can be stored in the cluster.

[0051] The management cluster obtains the status information generated by the training cluster during the execution of the model training task from the storage cluster, and uses this as the basis to determine the execution strategy. After that, the management cluster performs task scheduling based on the determined execution strategy to ensure that the training cluster successfully executes the model training task.

[0052] In an embodiment of the present application, the training cluster is composed of multiple high-performance training servers, a computing network switch cluster, and an IB / Ethernet switch cluster.

[0053] Among them, each high-performance training server mainly includes multiple computing units (used for computing acceleration during model training, such as GPU), multiple storage units (responsible for storing relevant data required during the execution of model training tasks), multiple general computing units (used for task scheduling, special transcendental function calculations, etc., the general computing unit can refer to the CPU), multiple computing network cards (used to synchronize or transmit data between multiple computing units), and multiple switches (used to synchronize or transmit data between general computing units, storage units, computing units and network cards).

[0054] As for computing network switch clusters, they are mainly used for high-speed data transmission between various high-performance training servers, such as model parameter synchronization.

[0055] IB / Ethernet switch cluster is used to link the storage cluster and management cluster to achieve data access and task scheduling during model training execution.

[0056] In the embodiments of the present application, the training cluster can be a homogeneous cluster or a heterogeneous cluster. The so-called homogeneous cluster refers to a cluster composed of devices from the same manufacturer, and the so-called heterogeneous cluster refers to a cluster composed of devices from different manufacturers. The present application does not limit the specific form of the training cluster.

[0057] The management cluster in the model training system can be composed of high-performance CPU servers. Based on the status data of the training cluster when executing model training tasks, the management cluster responds in advance to possible future conditions of the training cluster, and explores the optimal execution strategy through the constructed virtual training cluster, thereby scheduling tasks for the training cluster to improve the overall model task execution efficiency of the training cluster.

[0058] The storage cluster in the model training system can be composed of storage servers, which are responsible for storing various types of information in the model training system and various data required during the execution of model training tasks, such as loading model parameters, training samples and other data to each high-performance training server in the training cluster through IB / Ethernet, or collecting real-time status information from each high-performance training server in the training cluster when executing model training tasks. The file system form based on which the storage cluster stores data is not specifically limited in the embodiments of this application, and can be various high-performance file systems, such as a shared file system (General Parallel File System, GPFS), a parallel distributed file system Lustre, etc.

[0059] The management cluster, storage cluster, and training cluster in the above model training system work together to complete the execution process of the model training task, and by formulating an execution strategy, the efficiency of the entire model training task execution process is ensured. The following is a detailed description of the process of the model training system executing the model training task. Figure 2 shown.

[0060] Figure 2 A flow chart of a method for executing a model training task provided in an embodiment of the present application includes the following steps:

[0061] S201: Acquire real-time status data of the training cluster when executing the model training task.

[0062] In the above-mentioned model training system, the management cluster is responsible for scheduling tasks for the training cluster. To this end, in an embodiment of the present application, the management cluster can obtain real-time status data when the training cluster performs model training tasks, wherein the status data mentioned in the embodiment of the present application may include data such as CPU / GPU utilization, storage I / O latency, network throughput, log records, and fault records when the training cluster performs model training tasks. Of course, the status data mentioned here includes not only the status of the training cluster on the device, but also the status of the training cluster when performing model training tasks, such as training accuracy, loss value, training time of the model to be trained, etc.

[0063] It should be pointed out that when the management cluster predicts the status of the training cluster when executing the model task, in addition to referring to the real-time status data obtained above, it can also combine the task parameters corresponding to the model training task. The task parameters mentioned here may include model structure data of the model to be trained, parameter scale of the model, size and sample distribution of the training sample set used for the training model, number of training rounds of the model to be trained, etc.

[0064] In addition, the management cluster can also obtain the cluster configuration parameters corresponding to the training cluster. The cluster configuration parameters mentioned here may include the number of devices in the training cluster (such as the high-performance training servers, general computing units, computing network cards, etc. mentioned above), the type of GPU / CPU, the network bandwidth and network topology of the training cluster, the file system type and storage capacity used by the storage cluster, etc.

[0065] Therefore, the management cluster can predict the future state of the training cluster in the subsequent process based on the acquired real-time status data, cluster configuration parameters, and task parameters.

[0066] In actual applications, the historical status of the training cluster can also reflect the future status of the training cluster in the process of executing the model training task to a certain extent. For this reason, in an embodiment of the present application, in addition to obtaining the above data, the management cluster can also obtain the historical status data of the training cluster. Among them, if the above real-time status data refers to the status data generated by the training cluster executing the model training task in the current period, then the historical status data mentioned here can refer to the historical status data generated by the training cluster executing the model training task before the current period.

[0067] Of course, the above historical status data may also include status data generated by the training cluster in the past when executing similar model training tasks. To this end, the management cluster may match historical tasks similar to the currently executed model training tasks in the saved model task execution records according to the above task parameters and cluster configuration parameters, thereby obtaining the matched historical status data of similar tasks as a basis for prediction. The model training task execution records may be saved in the above storage cluster, and the management cluster may obtain them from the storage cluster.

[0068] Furthermore, in an embodiment of the present application, the storage cluster can collect and store status data in real time or periodically from the logs generated by the training cluster executing model training tasks through preset commands, or collect status data through preset monitoring tools. The present application does not limit the commands or monitoring tools used, as long as it can ensure that data collection can be effectively completed.

[0069] S202: Predicting, based on the real-time status data, the status of the training cluster when executing the model training task within a future set time period as a predicted status.

[0070] After obtaining the above real-time status data, the management cluster can analyze the real-time status data (which can be analyzed together with the above task parameters, system configuration parameters and historical status data) to predict the status of the training cluster when performing model training tasks within a set time period in the future.

[0071] The reason why the management cluster needs to predict the status of the training cluster in advance is mainly to prevent the training cluster from encountering equipment failures such as GPU overheating, server failure, or performance bottlenecks such as network congestion and excessive storage latency during the subsequent execution of model training tasks, so as to prevent these problems from affecting the smooth execution of model training tasks.

[0072] To this end, the management cluster can use the real-time status data obtained to predict the status of the training cluster from different dimensions. For example, if the management cluster determines through the real-time status data obtained that the network throughput of the training cluster has decreased during the current model training task, it can be predicted that the network throughput of the training cluster will further decrease when performing model training tasks in the future set time period.

[0073] For another example, if the management cluster monitors through the real-time status data that the temperature of some GPUs in the training cluster has increased during the current model training task, it can be predicted that some GPUs will overheat when the training cluster performs model training tasks within a set time period in the future.

[0074] For another example, if the management cluster determines through the real-time status data obtained that the speed at which the data generated by the training cluster during the current model training task is written to the storage cluster has slowed down, it can be predicted that the data generated by the training cluster during the model training task executed within the future set time period may not be written normally to the storage cluster for storage.

[0075] Some of the examples listed above only use data from a single dimension to predict the future state of the training cluster. In actual applications, the management cluster can combine data from multiple dimensions to predict the future state of the training cluster.

[0076] For example, if the management cluster determines through the acquired real-time status data that the GPU temperature of a high-performance training server in the training cluster has increased significantly, and according to the GPU type of the high-performance training server, the computing power that can be increased in the high-performance training server is lower than that of other training servers, and the bandwidth based on which the high-performance training server performs data transmission when executing model training tasks is low, then it can be predicted that the high-performance training server may have an operating failure within a set time period in the future. Other examples are not given one by one here.

[0077] The above-mentioned management cluster predicts the future state of the training cluster based on the acquired data. It can be understood that the management cluster realizes the prediction through various preset rules. Of course, if an artificial intelligence model that can predict the future state of the training cluster is pre-deployed in the management cluster, the management cluster can also input the acquired data into the artificial intelligence model, and output the state of the training cluster executing the model training task within the future set time period through the artificial intelligence model. The artificial intelligence model mentioned here can be any model that can realize the prediction function, and this application does not limit the specific form of the artificial intelligence model.

[0078] In addition, after the management cluster predicts the status of the training cluster, the predicted status can be recorded so that in the subsequent process, the saved predicted status can be used to formulate an execution strategy for the training cluster to execute the model training task within a set time period in the future.

[0079] S203: Determine each simulator corresponding to each device included in the training cluster, so as to initialize a virtual training cluster corresponding to the training cluster according to the simulators.

[0080] In order to ensure that the training cluster can successfully perform the model training task within a set time period in the future, the management cluster needs to formulate an execution strategy for the training cluster based on the predicted state. In order to ensure the feasibility of the execution strategy and to meet the actual situation of the training cluster in the set time period in the future, in the embodiment of the present application, the management cluster can build a virtual training cluster that is basically consistent with the training cluster in architecture and function. The virtual training cluster is not composed of real cluster devices, but can be realized through a simulator.

[0081] Therefore, the management cluster can determine the simulators corresponding to the devices included in the training cluster according to the configuration information and architecture information corresponding to the training cluster, and then initialize the virtual training cluster corresponding to the training cluster according to these simulators.

[0082] Among them, the management cluster can use a computing simulator to simulate the computing units in the above-mentioned high-performance training server, and use a storage simulator to simulate the storage units in the above-mentioned high-performance training server. Through a computing network simulator, the computing units simulated by the computing simulator can be interconnected to form a data transmission network (such as a Scale-Up network). Through a network simulator, the computing units simulated by the computing simulator and the storage units simulated by the storage simulator can be interconnected to build a virtual high-performance training server, and a data transmission network (such as a Scale-Out network) can be formed through a network simulator to build a virtual training cluster.

[0083] It should be pointed out that in actual applications, the management cluster does not necessarily initialize the virtual training cluster corresponding to the training cluster after determining the above-mentioned predicted state. Since the configuration information and architecture information in the training cluster can be known in advance, the management cluster can actually initialize the virtual training cluster corresponding to the training cluster before predicting the state, or initialize the virtual training cluster corresponding to the training cluster before obtaining the above-mentioned real-time state data. In short, this application does not make clear restrictions on the timing of initializing the virtual training cluster.

[0084] S204: Generate several execution strategies for the virtual training cluster according to the predicted state.

[0085] After constructing the above-mentioned virtual training cluster, the management cluster can generate several execution strategies for the virtual training cluster by presetting the status and combining the actual situation of the training cluster. In fact, these execution strategies are also generated for the training cluster.

[0086] The execution strategy generated by the management cluster is mainly from two perspectives. One is for the training cluster equipment itself, that is, based on the actual situation of the training cluster when executing the model training task, it will determine whether there will be equipment with failure risks in the training cluster within a set time period in the future, or based on the actual situation of the training cluster when executing the model training task, it will determine whether it is necessary to increase or decrease the equipment in the training cluster, or modify the entire topology of the training cluster to adapt to the execution of the model task.

[0087] Among them, the aforementioned addition of equipment can deal with the situation of insufficient computing resources, that is, the idle equipment in the training cluster (such as computing unit GPU) can be enabled to improve the computing resources during the execution of model training tasks. Reducing equipment can deal with the situation of sufficient computing resources, that is, some of the already started equipment in the training cluster (such as computing unit GPU) can be disabled to save this part of computing resources for use in other model training tasks.

[0088] The above mentioned are mainly execution strategies for training cluster equipment, while the management cluster can also generate execution strategies from the perspective of data processing of model training tasks. The execution strategies generated from the perspective of data processing of model training tasks can be implemented from perspectives such as changing the data sharding method, checkpoint interval, parallelism equipment, load balancing strategy, etc.

[0089] Data sharding refers to the method of splitting a data set or model parameters into multiple computing nodes (such as GPUs, servers) according to a specific strategy, in order to improve computing efficiency, scalability, or resource utilization. The management cluster can determine the splitting strategy for allocating data sets or model parameters to each computing unit in the training cluster in the future according to the predicted status.

[0090] The checkpoint interval refers to the time interval between two consecutive checkpoint operations. Its core function is to achieve fault tolerance and fast recovery by regularly saving the state of the model or task, and to avoid data loss or retraining due to failures or interruptions. The state of the training cluster in a set time period in the future determines how long the checkpoint interval should be used to save the state of the model or task. For example, if the training cluster is prone to abnormal conditions such as failures in a set time period in the future, the checkpoint interval should be set relatively small, so as to ensure that when abnormal conditions such as failures occur, the state of the saved model or task can be as close to the time point of the failure as possible; if the training cluster is not prone to abnormal conditions such as failures in a set time period in the future, the checkpoint interval can be set larger to save resources as much as possible.

[0091] The parallelism setting refers to the parameters that control the number of computing tasks or threads executed simultaneously. The purpose is to improve task efficiency by rationally utilizing computing resources (such as CPU, GPU, and nodes). For example, if the task execution progress of the training cluster in a future set time period is blocked, the parallelism can be increased by increasing the number of threads to improve task efficiency. If the number of threads for task execution of the training cluster in a future set time period is sufficient, the number of threads used can be reduced to save resources without affecting the normal execution of the model training task.

[0092] Load balancing strategy refers to a series of methods that optimize resource utilization, improve system performance, and avoid single-point overload by reasonably allocating tasks or traffic to multiple GPUs, servers, and other devices. Its core goal is to ensure that the load of all devices is as balanced as possible, thereby improving task execution efficiency, reducing latency, and enhancing system reliability and scalability.

[0093] In this regard, the management cluster can adjust the allocation strategy of the training cluster when executing model training tasks within a set time period in the future based on the determined prediction status, and evenly distribute the data required to execute the task to each device to avoid single-point overload and ensure the smooth execution of the model training task.

[0094] Of course, in practical applications, the execution strategy determined from the perspective of data processing of model training tasks also includes other levels, which will not be explained one by one here.

[0095] Therefore, from the above content, we can know that since the management cluster initializes the virtual training cluster corresponding to the training cluster, the execution strategy generated by the management cluster is for the virtual training cluster. If from the perspective of equipment, the virtual training cluster actually uses different virtual devices to execute model training tasks according to different execution strategies, and if from the perspective of data processing, the virtual training cluster actually uses different data processing strategies when executing model training tasks according to different execution strategies.

[0096] In addition, in actual applications, the management cluster can also generate execution strategies from both the device perspective and the data processing perspective. At this time, the determined execution strategy not only includes the adjustment of the virtual devices or topology structure in the virtual training cluster, but also involves the adjustment of the data processing strategy used in the execution of the model training task. It appears specifically according to the actual situation, that is, the management cluster should generate an execution strategy for the virtual training cluster (that is, the training cluster) according to the determined prediction state.

[0097] S205: According to at least some of the execution strategies, the model training task is executed through the virtual training cluster simulation to obtain performance indicators corresponding to at least some of the execution strategies, and based on the performance indicators corresponding to at least some of the execution strategies, a target strategy is determined from the several execution strategies.

[0098] After generating various execution strategies for the virtual training cluster, the management cluster needs to follow these execution strategies and simulate the execution model training tasks through the virtual training cluster, so as to determine the target strategy to be executed from these execution strategies.

[0099] During this process, the management cluster needs to first synchronize the status of the virtual training cluster according to the status of the training cluster in currently executing the model training task. After that, the management cluster can execute the model training task in sequence through the virtual training cluster according to each execution strategy, thereby obtaining the performance indicators generated after executing the model training task according to each execution strategy.

[0100] The above performance indicators are used to measure the status of the virtual training cluster (equivalent to the training cluster) in executing model training tasks. Specifically, they can be measured from three aspects: the time spent on executing model training tasks, the resource utilization when executing model training tasks, and the load when executing model training tasks.

[0101] As for the duration, we hope that the duration is as short as possible, which indicates that the execution efficiency of the model training task is high. As for resource utilization, we hope that the resource utilization is as high as possible, which indicates that the model training task makes full use of the available computing resources, network resources, etc., while ensuring efficiency and avoiding resource waste as much as possible. As for the load, we hope that the load is as balanced as possible to prevent a single device from being overloaded.

[0102] In the embodiment of the present application, an objective function is provided in the management cluster, and the management cluster can calculate the function value of the objective function through the above three performance indicators, and thus determine the final target strategy from various execution strategies based on the calculated function value.

[0103] Specifically, for the above three performance indicators, each performance indicator has a corresponding weight. On this basis, for any execution strategy, the management cluster can perform weighted summation according to the simulation time corresponding to the execution strategy (the simulation time here refers to the time spent by the virtual training cluster when simulating the execution model training task according to the execution strategy) and the weight corresponding to the simulation time, the simulation resource utilization corresponding to the execution strategy (the simulation resource utilization here refers to the resource utilization when the virtual training cluster simulates the execution model training task according to the execution strategy) and the weight corresponding to the simulation resource utilization, the simulation load corresponding to the execution strategy (the simulation here refers to the load when the virtual training cluster simulates the execution model training task according to the execution strategy) and the weight corresponding to the simulation load, and obtain the function value of the preset objective function. For details, please refer to the following formula:

[0104]

[0105] in, Indicates the simulation duration, represents the weight corresponding to the simulation duration, represents the simulation resource utilization, represents the weight corresponding to the simulation resource utilization, Indicates the simulated load, Indicates the weight corresponding to the simulation load, Represented as the function value of the calculated objective function.

[0106] It can be seen from the above formula that the simulation time is positively correlated with the function value of the objective function, that is, the longer the simulation time is, the larger the function value is, and the shorter the simulation time is, the smaller the function value is; the simulation resource utilization is negatively correlated with the function value of the objective function, that is, the larger the simulation resource utilization is, the smaller the function value is, and the smaller the simulation resource utilization is, the larger the function value is; the simulation load is positively correlated with the function value of the objective function, that is, the larger the simulation load is, the larger the function value is, and the smaller the simulation load is, the smaller the function value is.

[0107] Therefore, after the management cluster calculates the function value corresponding to each execution strategy, it can determine the final target strategy based on the size of the function value. It can be understood that the function value corresponding to the target strategy is the smallest.

[0108] Of course, in practical applications, since the number of execution strategies generated by the management cluster may be large, if the simulation is performed blindly according to each execution strategy one by one, it may consume too much time and resources. Therefore, in the embodiment of the present application, the management cluster can explore the final target strategy by gradually exploring, and does not need to execute all execution strategies.

[0109] Specifically, the management cluster can first select a candidate execution strategy from the generated execution strategies, and then simulate the execution model training task through the virtual training cluster according to the candidate execution strategy, so as to obtain the performance indicators corresponding to the candidate execution strategy. Afterwards, the management cluster can reselect the candidate execution strategy from the remaining execution strategies according to the performance indicators corresponding to the candidate execution strategy, and repeat the previous process until the final target strategy is selected.

[0110] The candidate execution strategies selected by the management cluster from the above-mentioned several execution strategies may be multiple execution strategies, and these candidate execution strategies may be used as initial data of the proxy model. The proxy model mentioned here may be, for example, a Gaussian model.

[0111] The management cluster uses the proxy model to fit the posterior distribution of the objective function based on the above initial data. This posterior distribution can reflect the predicted value of each remaining execution strategy. The predicted value here mainly includes two aspects, one is the predicted mean and the other is the predicted variance.

[0112] Among them, the predicted mean reflects the difference between different execution strategies. If the difference between the means of two execution strategies is too large, it means that the difference between the two execution strategies is too large. Conversely, if the means of the two execution strategies are small, it means that the difference between the two execution strategies is relatively small.

[0113] The prediction variance is used to reflect the uncertainty of the execution strategy. The so-called uncertainty can be understood as the uncertainty of whether the execution strategy is the current optimal strategy.

[0114] Therefore, after fitting the above posterior distribution, the management cluster can first determine the current optimal execution strategy by calculating the function value of the objective function among the candidate execution strategies that have been selected. Then, through the fitted posterior distribution, one or more candidate execution strategies can be re-selected from the remaining execution strategies to determine the mean difference and variance difference between the re-selected candidate execution strategy and the current optimal execution strategy. For the re-selected candidate execution strategies, the management cluster can select the one with the smallest mean difference from the current optimal execution strategy, or the one with the largest variance difference from the current optimal execution strategy when the mean differences of the candidate execution strategies are similar. Selecting the one with the smallest mean difference is to select an execution strategy close to the current optimal execution strategy based on the current optimal execution strategy, and selecting the variance difference is to further explore the optimal solution when the mean differences are similar.

[0115] After reselecting the candidate execution strategy through the above-fitted posterior distribution, the reselected candidate execution strategy can be added to the initial data of the above-mentioned proxy model, so as to adjust or refit the posterior distribution of the objective function according to the updated initial data, and then reselect the candidate execution strategy from the remaining execution strategies based on the adjusted or refitted posterior distribution.

[0116] Through continuous iteration, the final target strategy can be determined from each execution strategy. In this way, the management cluster does not have to calculate the function values ​​of the objective functions of all execution strategies one by one. It is equivalent to fitting the posterior distribution used to explore the optimal target strategy by first determining the function values ​​of the objective functions of a part of the execution strategies. Through this posterior distribution, the final target strategy can be gradually explored, which greatly improves the efficiency of determining the target strategy and saves resources to a great extent.

[0117] In addition, the setting of the above weights can be determined according to actual needs. For example, if the training duration is the main goal, then the weight corresponding to the above simulation duration can be set relatively large, and the simulation resource utilization and simulation load are secondary and auxiliary goals.

[0118] It should also be emphasized that the above-mentioned virtual training cluster does not actually execute the model training task according to the generated execution strategy. This process is actually a simulation process, which requires simulating the process of executing the model training task according to the current actual state of the training cluster.

[0119] In addition, in actual applications, the management cluster can also decide whether to calculate the function value corresponding to each execution strategy one by one to determine the target strategy, or to explore the target strategy by calculating the function value of a part of the execution strategies according to the number of generated execution strategies. Specifically, a quantity threshold can be used. When the number of generated execution strategies exceeds the quantity threshold, the target strategy can be explored by calculating the function value of a part of the execution strategies. If the number of generated execution strategies does not exceed the quantity threshold, the final target strategy can be determined by calculating the function value corresponding to each execution strategy one by one.

[0120] S206: Execute the model training task through the training cluster according to the target strategy.

[0121] After determining the target strategy, the management cluster can perform task scheduling for the training cluster according to the target strategy, adjust the use of equipment in the training cluster (such as removing or replacing equipment that may fail within a set time period in the future, adding equipment, etc.), or adjust the data processing method used to perform model training tasks within a set time period in the future (such as adjusting the data sharding method, etc.). The training cluster can continue to perform model training tasks according to the target strategy.

[0122] In addition, after the training cluster continues to execute the model training task according to the target strategy for a period of time, the management cluster can continue to obtain the real-time status data of the training cluster, and thus continue to determine the target strategy suitable for the training cluster and the model training task being executed in the above manner, so that the training cluster can finally complete the model training task through the target strategy determined by the management cluster every once in a while.

[0123] It can be seen from the above method that the management cluster in the model training system can predict the state of the training cluster's future execution of the model training task based on the real-time state information of the training cluster executing the model training task, and the management cluster can create a virtual training cluster corresponding to the training cluster, and through the virtual training cluster, simulate several execution strategies determined based on the predicted state, thereby determining the optimal target strategy, and then adjust the training cluster according to the target strategy, or adjust the data processing method when the training cluster executes the model training task. Therefore, the model training system and method provided in the embodiment of the present application can establish an effective prediction mechanism, respond in advance to various conditions that may occur when the training cluster executes the model training task, so as to ensure the normal execution of the model training task, and through the created virtual training cluster, a better execution strategy can be simulated, thereby further improving the efficiency of the entire training cluster in executing the model training task.

[0124] Moreover, when determining the target strategy, the management cluster can explore the final target strategy through the above-mentioned exploration method without having to calculate the function value of the objective function of each execution strategy one by one, thereby further improving the efficiency of determining the execution strategy and improving the efficiency of the training cluster in executing model training tasks as a whole.

[0125] In order to further describe a model training task execution method provided in an embodiment of the present application, the various steps are connected in series below to further illustrate the entire process. Figure 3 shown.

[0126] Figure 3 A serial diagram of the entire model training task execution process provided in an embodiment of the present application.

[0127] Before executing the model training task, the management cluster can obtain detailed system configuration information and task parameters from the storage cluster. The task parameters have been explained in the above content and will not be repeated here. The detailed system configuration information is similar to the cluster configuration parameters mentioned above, and is mainly used to describe in detail the device scale, device type, topology, etc. of the entire model training system.

[0128] The management cluster can determine the initial execution strategy based on the obtained detailed system configuration information and task parameters, and execute the model training task through the training cluster according to this initial execution strategy.

[0129] After that, the management cluster can periodically obtain task snapshots of model training tasks and real-time status data of the training cluster from the storage cluster, where the task snapshots can reflect the execution status of the model training tasks, such as task execution progress, task duration, etc. In addition, it is necessary to further determine whether the obtained task snapshots or real-time status data are new data. Since the storage cluster needs to collect the actual status of the training cluster and task snapshots according to a certain collection cycle, the management cluster needs to determine whether the data obtained from the storage cluster is new. If not, it means that it has been obtained before, and you can wait for a while before obtaining the task snapshot and status data from the storage cluster.

[0130] After determining that the acquired task snapshot and status data are new data, the management cluster can predict the status of the training cluster executing the model training task within a set time period in the future based on the acquired data, that is, obtain the predicted status, and record the obtained predicted status in the status data table.

[0131] Afterwards, the management cluster can initialize a virtual training cluster corresponding to the training cluster and synchronize the state of the virtual training cluster to the current state of the training cluster. The management cluster can generate multiple execution strategies based on the determined predicted state and determine the final target strategy from these execution strategies through simulation of the virtual training cluster.

[0132] When the target strategy determines that the currently executed model training task does not match the resources currently provided by the training cluster (this current mismatch is reflected in the fact that the training cluster may not be able to perform the model training task normally within the future set time period), a new execution strategy, that is, the target strategy, can be executed to schedule tasks for the training cluster.

[0133] By repeating the above process, we can eventually adjust the equipment configuration and data processing method of the training cluster to adapt to the model training task and ensure the smooth execution of the model training task.

[0134] The above is a model training task execution method provided by one or more embodiments of the present application. Based on the same idea, the present application also provides a corresponding model training task execution device, such as Figure 4 shown.

[0135] Figure 4 A schematic diagram of a model training task execution device provided in an embodiment of the present application specifically includes:

[0136] The acquisition module 401 is used to obtain real-time status data when the training cluster executes the model training task;

[0137] A prediction module 402 is used to predict the state of the training cluster when executing the model training task within a future set time period according to the real-time state data as a predicted state;

[0138] An initialization module 403 is used to determine each simulator corresponding to each device included in the training cluster, so as to initialize a virtual training cluster corresponding to the training cluster according to each simulator;

[0139] A generation module 404 is used to generate a plurality of execution strategies for the virtual training cluster according to the predicted state, wherein the virtual training cluster uses different virtual devices when simulating and executing the model training task according to different execution strategies, and / or the virtual training cluster uses different data processing strategies when simulating and executing the model training task according to different execution strategies;

[0140] A determination module 405 is used to perform the model training task through the virtual training cluster simulation according to at least some of the execution strategies, obtain performance indicators corresponding to the at least some of the execution strategies, and determine a target strategy from the several execution strategies according to the performance indicators corresponding to the at least some of the execution strategies;

[0141] The execution module 406 is used to execute the model training task through the training cluster according to the target strategy.

[0142] Optionally, the prediction module 402 is specifically used to obtain task parameters corresponding to the model training task; and predict the state of the training cluster when executing the model training task within a set future time based on the task parameters and the real-time status data.

[0143] Optionally, the determination module 405 is specifically used to select a candidate execution strategy from the several execution strategies, execute the model training task through the virtual training cluster simulation according to the candidate execution strategy, obtain performance indicators corresponding to the candidate execution strategy, and reselect the candidate execution strategy from the remaining execution strategies according to the performance indicators corresponding to the candidate execution strategy until the target strategy is selected.

[0144] Optionally, the performance indicators include: simulation time spent on executing the model training task through the virtual training cluster simulation, simulation resource utilization when executing the model training task through the virtual training cluster simulation, and simulation load when executing the model training task through the virtual training cluster simulation;

[0145] The determination module 405 is specifically used to determine the function value of the preset objective function according to the simulation duration corresponding to the candidate execution strategy and the weight corresponding to the simulation duration, the simulation resource utilization corresponding to the candidate execution strategy and the weight corresponding to the simulation resource utilization, and the simulation load corresponding to the candidate execution strategy and the weight corresponding to the simulation load, wherein the longer the simulation duration, the larger the function value, the larger the simulation resource utilization, the smaller the function value, and the larger the simulation load, the larger the function value; according to the function value, fit the posterior distribution of the objective function, and reselect the candidate execution strategy from the remaining execution strategies according to the posterior distribution.

[0146] The present application also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 A model training task execution method is provided.

[0147] Of course, in addition to software implementation, this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0148] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with hardware entity modules. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0149] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.

[0150] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0151] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0152] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0153] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0154] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0156] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0157] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0158] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0159] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0160] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0161] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0162] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0163] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.

Claims

1. A model training system, characterized in that: The model training system includes: a training cluster and a management cluster; The training cluster is used to perform model training tasks; The management cluster is used to obtain real-time status data of the training cluster when executing the model training task, and predict the status of the training cluster when executing the model training task within a set time period in the future based on the real-time status data, and determine each simulator corresponding to each device included in the training cluster as the predicted status, so as to initialize a virtual training cluster corresponding to the training cluster based on each simulator, and generate several execution strategies for the virtual training cluster based on the predicted status, and simulate the execution of the model training task through the virtual training cluster according to at least some of the several execution strategies to obtain performance indicators corresponding to at least some of the execution strategies, and determine a target strategy from the several execution strategies based on the performance indicators corresponding to at least some of the execution strategies, and execute the model training task through the training cluster according to the target strategy, wherein the virtual training cluster uses different virtual devices when simulating the execution of the model training task according to different execution strategies, and / or the virtual training cluster uses different data processing strategies when simulating the execution of the model training task according to different execution strategies.

2. The model training system according to claim 1, characterized in that: The model training system also includes: A storage cluster, used to store real-time status data when the training cluster performs the model training task; The management cluster is specifically used to obtain the real-time status data from the storage cluster.

3. The model training system according to claim 1 or 2, characterized in that: The management cluster is specifically used to obtain task parameters corresponding to the model training task, and predict the state of the training cluster when executing the model training task within a future set time based on the task parameters and the real-time status data.

4. The model training system according to claim 1, characterized in that: The management cluster is specifically used to select a candidate execution strategy from the several execution strategies, simulate and execute the model training task through the virtual training cluster according to the candidate execution strategy, obtain performance indicators corresponding to the candidate execution strategy, and reselect the candidate execution strategy from the remaining execution strategies according to the performance indicators corresponding to the candidate execution strategy until the target strategy is selected.

5. The model training system according to claim 4, characterized in that: The performance indicators include: the simulation time spent on executing the model training task through the virtual training cluster simulation, the simulation resource utilization rate when executing the model training task through the virtual training cluster simulation, and the simulation load when executing the model training task through the virtual training cluster simulation; The management cluster is specifically used to determine the function value of a preset objective function based on the simulation duration corresponding to the candidate execution strategy and the weight corresponding to the simulation duration, the simulation resource utilization corresponding to the candidate execution strategy and the weight corresponding to the simulation resource utilization, and the simulation load corresponding to the candidate execution strategy and the weight corresponding to the simulation load, wherein, the longer the simulation duration, the larger the function value, the larger the simulation resource utilization, the smaller the function value, and the larger the simulation load, the larger the function value; according to the function value, the posterior distribution of the objective function is fitted, and according to the posterior distribution, the candidate execution strategy is reselected from the remaining execution strategies.

6. A method for executing a model training task, characterized in that: The method is applied to a management cluster in a model training system, and the model training system also includes a training cluster, including: Acquire real-time status data of the training cluster when executing the model training task; According to the real-time status data, predict the status of the training cluster when executing the model training task within a future set time period as the predicted status; Determine each simulator corresponding to each device included in the training cluster, so as to initialize a virtual training cluster corresponding to the training cluster according to each simulator; According to the predicted state, several execution strategies for the virtual training cluster are generated, wherein the virtual training cluster uses different virtual devices when simulating and executing the model training task according to different execution strategies, and / or the virtual training cluster uses different data processing strategies when simulating and executing the model training task according to different execution strategies; According to at least some of the execution strategies, the model training task is executed by the virtual training cluster simulation to obtain performance indicators corresponding to at least some of the execution strategies, and a target strategy is determined from the several execution strategies according to the performance indicators corresponding to at least some of the execution strategies; According to the target strategy, the model training task is performed by the training cluster.

7. The method according to claim 6, characterized in that According to at least some of the execution strategies, the model training task is executed by the virtual training cluster simulation to obtain performance indicators corresponding to the at least some of the execution strategies, and according to the performance indicators corresponding to the at least some of the execution strategies, a target strategy is determined from the several execution strategies, specifically including: Select a candidate execution strategy from the several execution strategies, execute the model training task through the virtual training cluster simulation according to the candidate execution strategy, obtain the performance indicator corresponding to the candidate execution strategy, and reselect the candidate execution strategy from the remaining execution strategies according to the performance indicator corresponding to the candidate execution strategy until the target strategy is selected.

8. The method according to claim 7, characterized in that The performance indicators include: the simulation time spent on executing the model training task through the virtual training cluster simulation, the simulation resource utilization rate when executing the model training task through the virtual training cluster simulation, and the simulation load when executing the model training task through the virtual training cluster simulation; Reselecting a candidate execution strategy from the remaining execution strategies according to the performance indicator corresponding to the candidate execution strategy, specifically including: Determine the function value of the preset objective function according to the simulation duration corresponding to the candidate execution strategy and the weight corresponding to the simulation duration, the simulation resource utilization corresponding to the candidate execution strategy and the weight corresponding to the simulation resource utilization, and the simulation load corresponding to the candidate execution strategy and the weight corresponding to the simulation load, wherein the longer the simulation duration, the greater the function value, the greater the simulation resource utilization, the smaller the function value, and the greater the simulation load, the greater the function value; According to the function value, a posterior distribution of the objective function is fitted, and according to the posterior distribution, a candidate execution strategy is reselected from the remaining execution strategies.

9. A model training task execution device, characterized in that: include: The acquisition module is used to obtain real-time status data when the training cluster executes model training tasks; A prediction module, used to predict the state of the training cluster when executing the model training task within a future set time period according to the real-time state data, as a predicted state; An initialization module, used to determine each simulator corresponding to each device included in the training cluster, so as to initialize a virtual training cluster corresponding to the training cluster according to the each simulator; A generation module, used to generate a plurality of execution strategies for the virtual training cluster according to the predicted state, wherein the virtual training cluster uses different virtual devices when simulating and executing the model training task according to different execution strategies, and / or the virtual training cluster uses different data processing strategies when simulating and executing the model training task according to different execution strategies; a determination module, configured to perform the model training task through the virtual training cluster simulation according to at least some of the execution strategies, obtain performance indicators corresponding to the at least some of the execution strategies, and determine a target strategy from the several execution strategies according to the performance indicators corresponding to the at least some of the execution strategies; An execution module is used to execute the model training task through the training cluster according to the target strategy.

10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 6 to 8 is implemented.

Citation Information

Patent Citations

  • Distributed training method and device for deep learning model

    CN112000473A

  • Cluster configuration automatic optimization method and system based on machine learning

    CN112445746A