Expert parallel training time-consuming prediction method, device, equipment, medium and product

By sampling multiple times the training sample set of the hybrid expert model to monitor the activation status of the expert network, combined with the computing node resource status information, the problem of inaccurate prediction of parallel training of the hybrid expert model is solved, and the training efficiency and resource utilization rate are improved.

CN120373423BActive Publication Date: 2025-09-02SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510874226.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-02
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The existing model training time-consuming prediction scheme cannot accurately predict expert parallel training of mixed expert models, resulting in low training efficiency and low resource utilization.

Method used

By sampling the training sample set of the mixed expert model multiple times, monitoring the activation state parameters of the expert network, and combining the computing node resource status information of the heterogeneous computing system, the time-consuming prediction results of iterative training are calculated.

Benefits of technology

It realizes accurate prediction of the parallel training time of hybrid expert models, optimizes the expert paralleling method of heterogeneous computing systems, and improves training efficiency and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373423B_ABST
    Figure CN120373423B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment, medium and product for predicting the time consumption of expert parallel training, which relates to the field of artificial intelligence technology. The method comprises the following steps: the training sample set used for hybrid expert model training is sampled multiple times to obtain training samples, and the training samples are respectively input into the hybrid expert model to monitor the activation state parameters of the expert network of the hybrid expert model, thereby accurately predicting the activation status of the expert network in the expert parallel training; based on the activation state parameters of the expert network and the resource state information of the computing nodes of the heterogeneous computing system, the time consumption prediction results of the computing nodes performing iterative training are calculated, and based on the time consumption prediction results of the computing nodes, the time consumption prediction results of the heterogeneous computing systems performing iterative training are determined, thereby accurately predicting the time consumption of the expert parallel training of the hybrid expert model, and solving the problem in the related art that the model training time consumption prediction scheme cannot accurately predict the time consumption of the expert parallel training of the hybrid expert model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, medium and product for predicting the time consumption of expert parallel training. Background Art

[0002] With the advancement of artificial intelligence (AI) technology, the scale of neural network models used continues to increase, leading to an increase in the hardware cost and time required for model training. The Mixture of Experts (MoE) model is a machine learning method. For MoEs, using parallel expert training, deploying different expert networks on different computing nodes, can effectively improve the training efficiency of MoEs. However, because MoEs use sparse activation computation, traditional prediction methods that require time-consuming model training are inapplicable to MoEs. Summary of the Invention

[0003] The present invention provides a method, apparatus, device, medium and product for predicting the time consumption of expert parallel training, so as to at least solve the problem that the model training time consumption prediction scheme in the related art cannot accurately predict the time consumption of expert parallel training of the hybrid expert model.

[0004] The present invention provides a method for predicting the time consumption of expert parallel training, comprising:

[0005] Determine an expert network of hybrid expert models deployed on computing nodes of a heterogeneous computing system;

[0006] Obtaining a training sample set of the hybrid expert model, inputting training samples obtained by multiple sampling from the training sample set into the hybrid expert model respectively, and monitoring activation state parameters of the expert network;

[0007] Calculating a prediction result of the time consumption of the computing node to perform iterative training based on the activation state parameters of the expert network and the resource state information of the computing node;

[0008] The time consumption prediction result of the heterogeneous computing system performing iterative training is determined according to the time consumption prediction result of the computing node.

[0009] The present invention also provides a device for predicting the time consumption of expert parallel training, comprising:

[0010] A task information collection module, used to determine the expert network of the hybrid expert model deployed by the computing nodes of the heterogeneous computing system;

[0011] A system information collection module, configured to obtain resource status information of the computing node;

[0012] An expert state monitoring module is used to obtain a training sample set of the hybrid expert model, input training samples obtained by multiple sampling from the training sample set into the hybrid expert model, and monitor the activation state parameters of the expert network;

[0013] A calculation module is used to calculate the time consumption prediction result of the computing node performing iterative training based on the activation state parameters of the expert network and the resource state information of the computing node; and determine the time consumption prediction result of the heterogeneous computing system performing iterative training based on the time consumption prediction result of the computing node.

[0014] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned expert parallel training time consumption prediction methods when executing the computer program.

[0015] The present invention also provides a non-volatile storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned expert parallel training time consumption prediction methods are implemented.

[0016] The present invention also provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned expert parallel training time consumption prediction methods when executed by a processor.

[0017] Through the present invention, the training samples obtained by sampling the training sample set used for hybrid expert model training multiple times are respectively input into the hybrid expert model to monitor the activation state parameters of the expert network of the hybrid expert model, thereby accurately predicting the activation status of the expert network in the expert parallel training; according to the activation state parameters of the expert network and the resource status information of the computing nodes of the heterogeneous computing system, the time consumption prediction results of the computing nodes performing iterative training are calculated, and according to the time consumption prediction results of the computing nodes, the time consumption prediction results of the heterogeneous computing system performing iterative training are determined, thereby achieving accurate prediction of the time consumption of the expert parallel training of the hybrid expert model, solving the problem in the related art that the model training time consumption prediction scheme cannot accurately predict the time consumption of the expert parallel training of the hybrid expert model, and utilizing the obtained accurate expert parallel time consumption prediction results can help to pre-optimize the expert parallel mode of the heterogeneous computing system, thereby improving the efficiency of the expert parallel training and improving the resource utilization of the heterogeneous computing system. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 A flowchart of a first expert parallel training time-consuming prediction method provided by an embodiment of the present invention;

[0020] Figure 2 A schematic diagram of the deployment architecture of the first hybrid expert model provided by an embodiment of the present invention;

[0021] Figure 3 A schematic diagram of the deployment architecture of the second hybrid expert model provided by an embodiment of the present invention;

[0022] Figure 4 This is a flowchart of a second expert parallel training time-consuming prediction method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0024] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0025] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0026] To address the problem in related art that model training time consumption prediction schemes cannot accurately predict the time consumption of expert parallel training of a hybrid expert model, an embodiment of the present invention provides an expert parallel training time consumption prediction scheme. The scheme comprises: inputting training samples obtained by multiple sampling of a training sample set used for hybrid expert model training into the hybrid expert model to monitor the activation state parameters of the expert network of the hybrid expert model, thereby accurately predicting the activation status of the expert network during expert parallel training; calculating the time consumption prediction results of the computing nodes performing iterative training based on the activation state parameters of the expert network and the resource state information of the computing nodes of the heterogeneous computing system; and determining the time consumption prediction results of the heterogeneous computing system performing iterative training based on the time consumption prediction results of the computing nodes, thereby accurately predicting the time consumption of the expert parallel training of the hybrid expert model, thereby addressing the problem in related art that model training time consumption prediction schemes cannot accurately predict the time consumption of expert parallel training of the hybrid expert model. The obtained accurate expert parallel time consumption prediction results can help to pre-optimize the expert parallel mode of the heterogeneous computing system, thereby improving the efficiency of expert parallel training and improving the resource utilization of the heterogeneous computing system.

[0027] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the expert parallel training time-consuming prediction solution depends, the specific application environment architecture or specific hardware architecture is described here.

[0028] The expert parallel training time prediction solution provided by this invention can be deployed on heterogeneous computing systems. Heterogeneous computing systems include multiple computing devices (computing nodes), and different computing nodes may have different parameters such as computing core type, computing power, memory size, and communication bandwidth.

[0029] In the heterogeneous computing system targeted by the present invention, the computing core type of the computing node may include, but is not limited to, one or more of a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a neural network processor (NPU), a microcontroller unit (MCU), and an application-specific integrated circuit (ASIC).

[0030] In the heterogeneous computing system targeted by the present invention, different computing nodes can be interconnected through a network, such as Ethernet; or they can be interconnected through a bus, such as a high-speed serial computer expansion bus (Peripheral Component Interconnect Express, PCIe) or an NVIDIA high-speed interconnect bus.

[0031] The expert parallel training time consumption prediction method provided by the embodiment of the present invention can be applied to one or more computing nodes in a heterogeneous computing system, and can also be applied to a control node outside the heterogeneous computing system.

[0032] An embodiment of the present invention provides a method for predicting the time consumption of expert parallel training. The method is described in detail below in conjunction with the execution flow of the method.

[0033] Figure 1 A flowchart of a first expert parallel training time-consuming prediction method provided by an embodiment of the present invention; Figure 2 A schematic diagram of the deployment architecture of the first hybrid expert model provided by an embodiment of the present invention; Figure 3 A schematic diagram of the deployment architecture of the second hybrid expert model provided in an embodiment of the present invention.

[0034] like Figure 1 As shown, the method for predicting the time consumption of expert parallel training provided by the embodiment of the present invention may include: S101: determining an expert network of a hybrid expert model deployed on computing nodes of a heterogeneous computing system.

[0035] S102: Obtain a training sample set of the hybrid expert model, input the training samples obtained by multiple sampling from the training sample set into the hybrid expert model respectively, and monitor the activation state parameters of the expert network.

[0036] S103: Calculate and obtain a time consumption prediction result of the computing node executing iterative training based on the activation state parameters of the expert network and the resource state information of the computing node.

[0037] S104: Determine a time consumption prediction result of the heterogeneous computing system executing iterative training according to the time consumption prediction result of the computing node.

[0038] In some optional implementations of the present invention, an expert network can be deployed on a computing node. Figure 2As shown, compute node 1 deploys the gating network, expert network 1, and expert network 2, while compute node 2 deploys the gating network, expert network 3, and expert network 4. When input data 1 enters compute node 1, the gating network decides to activate expert network 1 and expert network 3 of the hybrid expert model. The expert outputs of expert networks 1 and 3 are then aggregated to produce output 1. When input data 2 enters compute node 2, the gating network decides to activate expert network 2 and expert network 4 of the hybrid expert model. The expert outputs of expert networks 2 and 4 are then aggregated to produce output 2. In a single training iteration, if the input is input data 1, expert network 1 and expert network 3 are activated. The time spent on compute node 1 is the computation time of expert network 1, the memory access time of compute node 1, and the expert input and output time of expert network 1. The time spent on compute node 2 is the computation time of expert network 3, the memory access time of compute node 2, and the expert input and output time of expert network 3.

[0039] In some other optional implementations of the present invention, one or more expert networks may be deployed on a computing node. Figure 3 As shown, expert network 1 is deployed on computing node 1, expert network 2 is deployed on computing node 2, and so on, expert network n is deployed on computing node n. After the input data passes through the gating network decision, expert network 1 and expert network 2 are activated, and the input data is sent to computing node 1 and computing node 2, and then the expert outputs of expert network 1 and expert network 2 are summarized to obtain the output result.

[0040] It should be noted that the above Figure 2 、 Figure 3 The deployment architecture shown here only applies to situations where only one expert network or one or more expert networks are deployed on a computing node. In actual applications, other specific deployment methods can be used. In embodiments of the present invention, one or more expert networks can be deployed on a computing node, and the same expert network can be deployed on different computing nodes.

[0041] The embodiments of the present invention can be applied before deploying a hybrid expert model in a heterogeneous computing system. By predicting the time consumption of the heterogeneous computing system to perform expert parallel training based on the expert activation status, the expert parallel mode of the heterogeneous computing system can be optimized, the efficiency of expert parallel training can be improved, and the resource utilization of the heterogeneous computing system can be improved.

[0042] For S101 , information of expert parallel training tasks required for time-consuming prediction is determined, that is, the expert network deployed on each computing node of the heterogeneous computing system is determined.

[0043] For S102, the training sample set is a collection of samples used to train the hybrid expert model. The gating network is used to decide which expert network to activate based on the type of input data. Therefore, the activation of the expert network can be monitored and the activation state parameters can be calculated by sampling the training samples and inputting them into the hybrid expert model in the same way as training the hybrid expert model.

[0044] In an embodiment of the present invention, monitoring the activation state parameters of the expert network in S102 may include: monitoring the activation state of the expert network during historical iterative training; determining the activation probability of the expert network during a single iterative training session based on the activation state of the expert network during multiple rounds of historical iterative training, and using the activation probability as the activation state parameter of the expert network. The activation state of the expert network during multiple rounds of historical iterative training can be monitored by pre-deploying the hybrid expert model or obtaining historical iterative training data when the hybrid expert model is trained based on the same training sample. For each expert network, the activation probability of the expert network can be obtained by dividing the number of activations during multiple rounds of historical iterative training by the number of historical iterative training sessions.

[0045] In an embodiment of the present invention, monitoring the activation state of an expert network during historical iterative training may include: obtaining a pre-trained gating network of a hybrid expert model; sampling training samples from a training sample set and inputting them into the gating network, and monitoring the weights of the expert network output by the gating network; and determining the activation state of the expert network based on the weights of the expert network. Because deploying a hybrid expert network requires high hardware costs, a pre-trained gating network can be deployed to determine the activation probability of the expert network under the training sample set.

[0046] Since the sparse activation strategy of the hybrid expert network is to select the top k expert networks for activation based on the output of the gating network, the activation state of the expert network is determined according to the weight of the expert network, which can include: obtaining the number of expert activations in a calculation of the hybrid expert model; sorting the weights of the expert networks from large to small, and determining the corresponding number of expert networks as activated expert networks based on the number of expert activations, and the other expert networks as inactivated expert networks.

[0047] For S103, the time consumption prediction result of the computing node in one iterative training can be calculated based on the activation probability of the expert network when the hybrid expert model is trained using the training sample set, combined with the resource status information of the computing node of the heterogeneous computing system and the expert network deployed by the computing node.

[0048] Regarding S104 , considering that the time consumption of the heterogeneous computing system for executing one iterative training is affected by the computing node with the longest time consumption, the time consumption prediction result of the heterogeneous computing system for executing the iterative training can be determined based on the time consumption prediction results of each computing node.

[0049] The embodiment of the present invention provides a method for predicting the time consumption of expert parallel training. The method inputs the training samples obtained by multiple sampling of the training sample set used for hybrid expert model training into the hybrid expert model to monitor the activation state parameters of the expert network of the hybrid expert model, thereby accurately predicting the activation status of the expert network in the expert parallel training; according to the activation state parameters of the expert network and the resource status information of the computing nodes of the heterogeneous computing system, the time consumption prediction results of the computing nodes performing iterative training are calculated, and according to the time consumption prediction results of the computing nodes, the time consumption prediction results of the heterogeneous computing system performing iterative training are determined, thereby accurately predicting the time consumption of the expert parallel training of the hybrid expert model, solving the problem in the related art that the model training time consumption prediction scheme cannot accurately predict the time consumption of the expert parallel training of the hybrid expert model, and using the obtained accurate expert parallel time consumption prediction results can help to pre-optimize the expert parallel mode of the heterogeneous computing system, thereby improving the efficiency of the expert parallel training and improving the resource utilization of the heterogeneous computing system.

[0050] Figure 4 This is a flowchart of a second expert parallel training time-consuming prediction method provided by an embodiment of the present invention.

[0051] Based on the above embodiment, the embodiment of the present invention continues to describe the step of predicting the time consumption of executing the expert parallel training task in the heterogeneous computing system according to the activation state parameters of the expert network.

[0052] In an embodiment of the present invention, in S103, the time consumption prediction result of the computing node performing iterative training is calculated based on the activation state parameters of the expert network and the resource status information of the computing node, which may include: calculating the forward propagation computing time prediction result of the computing node and the backward propagation computing time prediction result of the computing node based on the activation probability of the expert network, the computing power resource parameters of the computing node and the memory bandwidth of the computing node; calculating the forward propagation communication time prediction result of the computing node and the backward propagation communication time prediction result of the computing node based on the activation probability of the expert network and the network resource parameters of the computing node.

[0053] like Figure 4 As shown, the expert parallel training time-consuming prediction method provided by the embodiment of the present invention can be implemented based on four modules: a task information collection module, a system information collection module, an expert status monitoring module and a calculation module.

[0054] The task information collection module is used to parse or calculate the necessary input information based on the user-entered expert parallel training task information. This information may include the forward propagation and backpropagation computational cost of the expert network, expressed in floating-point operations (FLOPs). The forward propagation computational cost can be estimated based on the neural layer composition of the expert network, where it is equal to the sum of the computational cost of all neural layers. The computational cost of a neural layer can be estimated using mathematical methods. For example, for the forward computation of a fully connected layer, assuming the input data dimensions are (N, D), the hidden layer weight dimensions are (D, out), and the output is (N, out), the computational cost is FLOPs = N x (2 * D - 1) * out. Different neural layers have different calculation methods, which are not listed here. The backpropagation computational cost can be estimated by multiplying the forward propagation computational cost by 2. Forward and backpropagation computational costs can also be calculated using existing open-source tools such as torchstat. Generally, the network shape of the expert networks is the same, so the computational cost of only one expert network needs to be calculated.

[0055] The input and output data volumes of the expert neural network. The input data volume can be estimated by multiplying the batch size by the size of the tensor input to the expert neural network. For example, if the batch size is 10 and the input size of the neural network layer is [10, 10], the input data volume is 10*10*10*data precision (for example, 32-bit floating point numbers fp32 is 4 bytes). This information can also be calculated using existing open-source tools such as torchstat. The output data volume can also be estimated by multiplying the output tensor size of the expert network by the batch size. This information can all be calculated using existing open-source tools such as torchstat.

[0056] The amount of memory accessed during the forward and backward propagation of the expert network. For the forward propagation of the expert network, the memory data size can be calculated as the expert network's input data size plus the number of parameters. Given the input data size, the parameter size can be counted using existing open-source tools such as torchstat, or mathematically calculated based on the expert network's structure. The amount of memory accessed during the backward propagation of the expert network can be calculated using data statistics. For example, given a model parameter size of X, in PyTorch automatic mixed-precision training, the data size includes parameters, gradients, optimizer state, and activation values. The data size for parameters, gradients, and optimizer state is 16X bytes, and the activation value is equal to the output of each layer of the expert neural network plus the input of the first layer (known from the calculation of the input and output data sizes of the expert neural network). This information can also be counted using existing open-source tools such as torchstat.

[0057] During parallel expert training, the expert network is assigned to each compute node. This data can be specified manually, typically through programmatic programming. Each unique compute node needs to collect information about the expert network running on it.

[0058] The total number of compute nodes participating in the task.

[0059] The total number of expert networks.

[0060] If the above information cannot be collected, the task information collection module will inform the user to make the corresponding input. The collected information will be sent to the calculation module to predict the time required for iterative training.

[0061] The system information collection module is used to collect resource status information for compute nodes in a heterogeneous computing system. This information may include the computing power resource parameters for each compute node. If the compute node is used only for inference computing tasks of a hybrid expert model, the peak computing power parameter (in floating-point operations per second) of the compute node can be used. This parameter can be obtained from the compute node's product manual or actual testing. The memory bandwidth of the compute node. If the compute node is used only for inference computing tasks of a hybrid expert model, the peak memory bandwidth of the compute node can be used. This parameter can be obtained from the compute node's product manual or actual testing. The network resource parameters of the compute node may include the bandwidth and latency of the compute node's external links. These parameters can be obtained using standard benchmark tools such as iperf.

[0062] The system status monitoring module is used to monitor the activation probability of the expert network when training based on the training sample set. It can intercept multiple adjacent rounds of historical iterative training and calculate the probability of the expert network being activated in each round of iterative training.

[0063] Based on the information collected by the task information collection module, the system information collection module and the system status monitoring module, the computing module performs time consumption prediction of each computing node in the expert parallel training according to the deployment method of the hybrid expert model in the heterogeneous computing system, and predicts the time consumption of the heterogeneous computing system to perform iterative training based on the time consumption prediction results of the computing nodes.

[0064] The calculation steps of the calculation module are further introduced below.

[0065] In an embodiment of the present invention, the forward propagation computation time prediction result of the computing node is calculated based on the activation probability of the expert network, the computing power resource parameters of the computing node, and the memory bandwidth of the computing node, which may include: calculating the first time prediction result of the computing node based on the activation probability of the expert network on the computing node, the forward propagation computation amount of the expert network, and the computing power resource parameters of the computing node; calculating the second time prediction result of the computing node based on the activation probability of the expert network on the computing node, the forward propagation memory access data amount of the expert network, and the memory bandwidth of the computing node; and determining the larger value of the first time prediction result and the second time prediction result as the forward propagation computation time prediction result of the computing node.

[0066] In an embodiment of the present invention, the back propagation computation time prediction result of the computing node is calculated based on the activation probability of the expert network, the computing power resource parameters of the computing node, and the memory bandwidth of the computing node, which may include: calculating the third time prediction result of the computing node based on the activation probability of the expert network on the computing node, the back propagation computation amount of the expert network, and the computing power resource parameters of the computing node; calculating the fourth time prediction result of the computing node based on the activation probability of the expert network on the computing node, the back propagation memory access data amount of the expert network, and the memory bandwidth of the computing node; and determining the larger value of the third time prediction result and the fourth time prediction result as the back propagation computation time prediction result of the computing node.

[0067] In an embodiment of the present invention, the forward propagation communication time prediction result of the computing node is calculated based on the activation probability of the expert network and the network resource parameters of the computing node, which may include: determining the first expert input data amount output by the computing node to other computing nodes in the forward propagation based on the activation probability of the expert network on the computing node and the amount of expert input data sent to the computing nodes where other activated expert networks are located when the expert network is activated; determining the first expert output data amount received by the computing node from other computing nodes in the forward propagation based on the activation probability of the expert network on the computing node and the amount of expert output data output by the computing nodes where other activated expert networks are located when the expert network is activated; and calculating the forward propagation communication time prediction result of the computing node based on the sum of the first expert input data amount and the first expert output data amount divided by the bandwidth of the computing node's external communication link plus the communication delay of the computing node's external communication link.

[0068] In an embodiment of the present invention, the backward propagation communication time prediction result of the computing node is calculated based on the activation probability of the expert network and the network resource parameters of the computing node, which may include: determining the second expert output data volume output by the computing node from other computing nodes during backward propagation based on the activation probability of the expert network on the computing node and the amount of expert output data output by the computing nodes where other activated expert networks are located when the expert network is activated; determining the second expert input data volume output by the computing node to other computing nodes during backward propagation based on the activation probability of the expert network on the computing node and the amount of expert input data sent to the computing nodes where other activated expert networks are located when the expert network is activated; and calculating the backward propagation communication time prediction result of the computing node based on the sum of the second expert input data volume and the second expert output data volume divided by the bandwidth of the computing node's external communication link plus the communication delay of the computing node's external communication link.

[0069] Based on the information collected by the task information collection module, system information collection module and system status monitoring module, the input parameters of the calculation module may include: the total number of computing nodes participating in the task , the total number of expert networks (When deploying an expert network on a computing node, = ), compute nodes The bandwidth of the external communication link , compute nodes The delay of the external communication link , compute nodes Computing resource parameters (unit: floating point operations per second FLOPS), computing nodes Memory bandwidth , the forward propagation computation of the expert network , the amount of back propagation calculation of the expert network , expert network in computing nodes The amount of data accessed in the forward propagation , expert network in computing nodes The amount of data accessed by back propagation on , the amount of input data of the expert network , the output data volume of the expert network .

[0070] The time consumption prediction principle introduced in the embodiment of the present invention will be introduced below in combination with actual application scenarios.

[0071] For the scenario where one expert network is deployed on one computing node, determining the time consumption prediction result of the heterogeneous computing system performing iterative training based on the time consumption prediction result of the computing node in S104 may include: calculating the forward propagation computing time prediction result of the heterogeneous computing system based on the forward propagation computing time prediction results of multiple computing nodes and the activation probability of the expert network; calculating the back propagation computing time prediction result of the heterogeneous computing system based on the back propagation computing time prediction results of multiple computing nodes and the activation probability of the expert network; calculating the forward propagation communication time prediction result of the heterogeneous computing system based on the forward propagation communication time prediction results of multiple computing nodes and the activation probability of the expert network; calculating the back propagation communication time prediction result of the heterogeneous computing system based on the back propagation communication time prediction results of multiple computing nodes and the activation probability of the expert network; and using the sum of the forward propagation computing time prediction result of the heterogeneous computing system, the back propagation computing time prediction result of the heterogeneous computing system, the forward propagation communication time prediction result of the heterogeneous computing system, and the back propagation communication time prediction result of the heterogeneous computing system as the time consumption prediction result of the heterogeneous computing system performing iterative training.

[0072] That is to say, the time consumption prediction results of the computing system in the four stages of forward propagation calculation, forward propagation communication, backpropagation calculation, and backpropagation communication are calculated respectively, and the sum of the results is used to obtain the time consumption prediction results of the heterogeneous computing system in the first iterative training.

[0073] In an embodiment of the present invention, the forward propagation calculation time consumption prediction result of the heterogeneous computing system is calculated based on the forward propagation calculation time consumption prediction results of multiple computing nodes and the activation probability of the expert network, which can include: arranging the forward propagation calculation time consumption prediction results of the computing nodes in order from large to small and substituting them into each element of the expected time consumption calculation formula respectively, substituting the corresponding activation probability of the expert network into each element of the expected time consumption calculation formula, and calculating the forward propagation calculation time consumption prediction result of the heterogeneous computing system; wherein, the expected time consumption calculation formula is to calculate the expected time consumption based on the sum of multiple elements, and the elements of the expected time consumption calculation formula are the product of the input time consumption prediction result, the activation probability of the expert network corresponding to the input time consumption prediction result, and the probability that the expert network is not activated is greater than the probability corresponding to the input time consumption prediction result.

[0074] In an embodiment of the present invention, the back propagation calculation time consumption prediction result of the heterogeneous computing system is calculated based on the back propagation calculation time consumption prediction results of multiple computing nodes and the activation probability of the expert network, which can include: arranging the back propagation calculation time consumption prediction results of the computing nodes in order from large to small and substituting them into each element of the expected time consumption calculation formula respectively, substituting the corresponding activation probability of the expert network into each element of the expected time consumption calculation formula, and calculating the back propagation calculation time consumption prediction result of the heterogeneous computing system; wherein, the expected time consumption calculation formula is to calculate the expected time consumption based on the sum of multiple elements, and the elements of the expected time consumption calculation formula are the product of the input time consumption prediction result, the activation probability of the expert network corresponding to the input time consumption prediction result, and the probability that the expert network is not activated is greater than the probability corresponding to the input time consumption prediction result.

[0075] In an embodiment of the present invention, the forward propagation communication time consumption prediction result of the heterogeneous computing system is calculated based on the forward propagation communication time consumption prediction results of multiple computing nodes and the activation probability of the expert network, which can include: arranging the forward propagation communication time consumption prediction results of the computing nodes in order from large to small and substituting them into each element of the expected time consumption calculation formula respectively, substituting the corresponding activation probability of the expert network into each element of the expected time consumption calculation formula, and calculating the forward propagation communication time consumption prediction result of the heterogeneous computing system; wherein the expected time consumption calculation formula is to calculate the expected time consumption based on the sum of multiple elements, and the elements of the expected time consumption calculation formula are the product of the input time consumption prediction result, the activation probability of the expert network corresponding to the input time consumption prediction result, and the probability that the expert network is not activated is greater than the probability corresponding to the input time consumption prediction result.

[0076] In an embodiment of the present invention, the back propagation communication time prediction result of the heterogeneous computing system is calculated based on the back propagation communication time prediction results of multiple computing nodes and the activation probability of the expert network, which can include: arranging the back propagation communication time prediction results of the computing nodes in order from large to small and substituting them into each element of the expected time calculation formula respectively, substituting the corresponding activation probability of the expert network into each element of the expected time calculation formula, and calculating the back propagation communication time prediction result of the heterogeneous computing system; wherein the expected time calculation formula is to calculate the expected time according to the sum of multiple elements, and the elements of the expected time calculation formula are the product of the input time prediction result, the activation probability of the expert network corresponding to the input time prediction result, and the probability that the expert network is not activated is greater than the probability corresponding to the input time prediction result.

[0077] The expected time calculation formula provided in the embodiment of the present invention can be expressed as:

[0078] ;

[0079] in, Indicates time-consuming expectations, 、 、 … Indicates that the single time consumption prediction results of each computing node are arranged from large to small. 、 、 … 、 Represents the activation probability of the expert network on the computing node after the corresponding single-item time consumption prediction results are arranged from largest to smallest. The types of single-item time consumption prediction results include forward propagation computation time prediction results, backward propagation computation time prediction results, forward propagation communication time prediction results, and backward propagation communication time prediction results.

[0080] The meaning of this expected time calculation formula is that since the single time consumption of a heterogeneous computing system in a stage is affected by the computing node with the longest time consumption, if the expert network on the computing node with the longest time consumption is activated, there is no need to consider the time consumption of other computing nodes.

[0081] In the embodiment of the present invention, the computing node The forward propagation calculation time prediction results It can be calculated by the following formula:

[0082] ;

[0083] in, represents the forward propagation computational amount of the expert network, Represents a compute node The computing power resource parameters, Indicates that the expert network is in the computing node The amount of data accessed by the forward propagation on Represents a compute node Memory bandwidth, Indicates maximum value calculation. 、 Represents computing nodes The forward propagation computation time and memory access time are determined by the forward propagation computation time and memory access time. The final forward propagation computation time depends on one of the computation bottleneck and the memory access bottleneck.

[0084] Each computing node The forward propagation calculation time prediction results Arrange them in descending order and substitute them into the expected time calculation formula. 、 、 … , and substitute the activation probability of the corresponding expert network into the expected time calculation formula 、 、 … 、 , thereby calculating the forward propagation calculation time prediction result of the heterogeneous computing system .

[0085] Compute nodes Back propagation calculation time-consuming prediction results It can be calculated by the following formula:

[0086] ;

[0087] in, represents the forward propagation computational amount of the expert network, Represents a compute node The computing power resource parameters, Indicates that the expert network is in the computing node The amount of data accessed by the forward propagation on Represents a compute node Memory bandwidth, Indicates maximum value calculation. 、 Represents computing nodes The back propagation computation time and memory access time of the algorithm are calculated. The final back propagation computation time depends on one of the computation bottleneck and the memory access bottleneck.

[0088] Each computing node The forward propagation calculation time prediction results Arrange them in descending order and substitute them into the expected time calculation formula. 、 、 … , and substitute the activation probability of the corresponding expert network into the expected time calculation formula 、 、 … 、 , thereby calculating the forward propagation calculation time prediction result of the heterogeneous computing system .

[0089] Compute nodes The forward propagation communication time prediction results It can be calculated by the following formula:

[0090] ;

[0091] in, represents the amount of input data of the expert network, represents the output data volume of the expert network, Represents a compute node The bandwidth of the external communication link, Represents a compute node The delay of the external communication link, Indicates the number of experts activated in the hybrid expert model setting. Compute node The forward propagation communication includes two communications. The first is to receive the expert input data from each computing node except itself, and the second is to receive the expert output data from each expert network except its own expert network.

[0092] Each computing node The forward propagation communication time prediction results Arrange them in descending order and substitute them into the expected time calculation formula. 、 、 … , and substitute the activation probability of the corresponding expert network into the expected time calculation formula 、 、 … 、 , thereby calculating the forward propagation communication time prediction result of the heterogeneous computing system .

[0093] Compute nodes The back propagation communication time prediction results It can be calculated by the following formula:

[0094] ;

[0095] in, represents the amount of input data of the expert network, represents the output data volume of the expert network, Represents a compute node The bandwidth of the external communication link, Represents a compute node The delay of the external communication link, Indicates the number of experts activated in the hybrid expert model setting. Compute node The backpropagation communication includes two communications, which is opposite to the forward propagation communication. The first is that each expert network except its own expert network receives the expert output data once, and the second is that each computing node except its own receives the expert input data once.

[0096] Each computing node The back propagation communication time prediction results Arrange them in descending order and substitute them into the expected time calculation formula. 、 、 … , and substitute the activation probability of the corresponding expert network into the expected time calculation formula 、 、 … 、 , thereby calculating the back propagation communication time prediction result of the heterogeneous computing system .

[0097] The time-consuming prediction result of the heterogeneous computing system in one iteration training is for:

[0098] .

[0099] If the expert parallel mode of the heterogeneous computing system deploys one or more expert networks for one computing node, that is, there is a situation where multiple expert networks are deployed on one computing node, based on the calculation principle of the calculation module provided in the above embodiment, in order to simplify the solution, S103 determines the time consumption prediction result of the heterogeneous computing system performing iterative training according to the time consumption prediction result of the computing node, which may include: taking the maximum value of the forward propagation calculation time consumption prediction results of multiple computing nodes as the forward propagation calculation time consumption prediction result of the heterogeneous computing system; taking the maximum value of the backward propagation calculation time consumption prediction results of multiple computing nodes as the backward propagation calculation time consumption prediction result of the heterogeneous computing system. The forward propagation communication time consumption prediction result of the heterogeneous computing system is used; the maximum value among the forward propagation communication time consumption prediction results of multiple computing nodes is used as the forward propagation communication time consumption prediction result of the heterogeneous computing system; the maximum value among the backward propagation communication time consumption prediction results of multiple computing nodes is used as the backward propagation communication time consumption prediction result of the heterogeneous computing system; the sum of the forward propagation computing time consumption prediction result of the heterogeneous computing system, the backward propagation computing time consumption prediction result of the heterogeneous computing system, the forward propagation communication time consumption prediction result of the heterogeneous computing system and the backward propagation communication time consumption prediction result of the heterogeneous computing system is used as the time consumption prediction result for the heterogeneous computing system to perform iterative training.

[0100] Specifically, for the four phases of forward propagation computation, forward propagation communication, backpropagation computation, and backpropagation communication in a single iterative training iteration, we can first calculate an estimated value for each compute node's timing prediction for that phase based on the activation probability of the expert network on that compute node, the task parameters of the expert network, and the resource status information of the compute node. Because each phase is executed in parallel by multiple compute nodes, the estimated timing of the compute node with the largest timing result can be used as the system timing prediction for that phase.

[0101] In an embodiment of the present invention, the calculation steps of the forward propagation computing time prediction result of the computing node may include: performing weighted summation of the forward propagation computing amount of the expert network on the computing node with the activation probability of the expert network as the weight to obtain the total forward propagation computing amount of the computing node, and obtaining the first time prediction result of the computing node by the ratio of the total forward propagation computing amount of the computing node to the computing power resource parameter of the computing node; performing weighted summation of the forward propagation memory access data amount of the expert network on the computing node with the activation probability of the expert network as the weight to obtain the memory access data amount of the computing node, and obtaining the second time prediction result of the computing node by the ratio of the memory access data amount of the computing node to the memory bandwidth of the computing node; and determining the larger value of the first time prediction result and the second time prediction result as the forward propagation computing time prediction result of the computing node.

[0102] The forward propagation calculation time prediction result of the heterogeneous computing system is It can be calculated by the following formula:

[0103] ;

[0104] in, Represents a compute node Expert Network The activation probability, Represents a compute node The total number of expert networks on Indicates the total number of computing nodes, represents the forward propagation computational amount of the expert network, Represents a compute node The computing power resource parameters, Indicates that the expert network is in the computing node The amount of data accessed by the forward propagation on Represents a compute node Memory bandwidth, Indicates maximum value calculation. 、 Represents computing nodes The forward propagation computation time and memory access time are determined by the forward propagation computation time and memory access time. The final forward propagation computation time depends on one of the computation bottleneck and the memory access bottleneck.

[0105] In an embodiment of the present invention, the calculation steps of the back propagation calculation time prediction result of the computing node may include: performing weighted summation of the back propagation calculation amount of the expert network on the computing node with the activation probability of the expert network as the weight to obtain the total back propagation calculation amount of the computing node, and obtaining the third time prediction result of the computing node by the ratio of the total back propagation calculation amount of the computing node to the computing power resource parameter of the computing node; performing weighted summation of the back propagation memory access data amount of the expert network on the computing node with the activation probability of the expert network as the weight to obtain the memory access data amount of the computing node, and obtaining the fourth time prediction result of the computing node by the ratio of the memory access data amount of the computing node to the memory bandwidth of the computing node; and determining the larger value of the third time prediction result and the fourth time prediction result as the back propagation calculation time prediction result of the computing node.

[0106] The back propagation calculation time prediction result of the heterogeneous computing system is It can be calculated by the following formula:

[0107] ;

[0108] in, Represents a compute node Expert Network The activation probability, Represents a compute node The total number of expert networks on Indicates the total number of computing nodes, represents the amount of back propagation computation of the expert network, Represents a compute node The computing power resource parameters, Indicates that the expert network is in the computing node The amount of back-propagation memory access data on , Represents a compute node Memory bandwidth, Indicates maximum value calculation. 、 Represents computing nodes The back propagation computation time and memory access time of the algorithm are calculated. The final back propagation computation time depends on one of the computation bottleneck and the memory access bottleneck.

[0109] In an embodiment of the present invention, the steps for calculating the forward propagation communication time prediction result of a computing node may include: determining the amount of first expert input data output by the computing node to other computing nodes during forward propagation based on the activation probability of the expert network on the computing node and the amount of expert input data sent to computing nodes where other activated expert networks are located when the expert network is activated; determining the amount of first expert output data received by the computing node from other computing nodes during forward propagation based on the activation probability of the expert network on the computing node and the amount of expert output data output by computing nodes where other activated expert networks are located when the expert network is activated; and calculating the forward propagation communication time prediction result of the computing node based on the sum of the first expert input data amount and the first expert output data amount divided by the bandwidth of the computing node's external communication link plus the communication delay of the computing node's external communication link.

[0110] The forward propagation communication time prediction result of the heterogeneous computing system is It can be obtained through the following communication:

[0111] ;

[0112] in, Represents a compute node Expert Network The activation probability, Represents a compute node The total number of expert networks on Indicates the total number of computing nodes, represents the amount of back propagation computation of the expert network, Represents a compute node The computing power resource parameters, Indicates that the expert network is at the communication node The amount of data accessed by the forward propagation on Represents a communication node Memory bandwidth, Indicates maximum value communication. Computing node The forward propagation communication includes two communications. The first is to receive the expert input data from each computing node except itself, and the second is to receive the expert output data from each expert network except its own expert network.

[0113] In an embodiment of the present invention, the steps for calculating the predicted result of the backpropagation communication time of a computing node may include: determining the amount of second expert input data output by the computing node from other computing nodes during backpropagation based on the activation probability of the expert network on the computing node and the amount of expert output data output by the computing nodes where other activated expert networks are located when the expert network is activated; determining the amount of second expert input data output by the computing node to other computing nodes during backpropagation based on the activation probability of the expert network on the computing node and the amount of expert input data sent to the computing nodes where other activated expert networks are located when the expert network is activated; and calculating the predicted result of the backpropagation communication time of the computing node based on the sum of the second expert input data amount and the second expert output data amount divided by the bandwidth of the computing node's external communication link and adding the communication delay of the computing node's external communication link.

[0114] The back propagation communication time prediction result of heterogeneous computing system is It can be obtained through the following communication:

[0115] ;

[0116] in, Represents a compute node Expert Network The activation probability, Represents a compute node The total number of expert networks on Indicates the total number of computing nodes, represents the back-propagation communication volume of the expert network, Represents a communication node The computing power resource parameters, Indicates that the expert network is at the communication node The amount of back-propagation memory access data on , Represents a communication node Memory bandwidth, Indicates maximum value communication. Computing node The backpropagation communication includes two communications, which is opposite to the forward propagation communication. The first is that each expert network except its own expert network receives the expert output data once, and the second is that each computing node except its own receives the expert input data once.

[0117] The time-consuming prediction result of the heterogeneous computing system in one iteration training is for:

[0118] .

[0119] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0120] An embodiment of the present invention also provides an expert parallel training time consumption prediction device, comprising: a task information collection module, used to determine the expert network of the hybrid expert model deployed by the computing nodes of the heterogeneous computing system; a system information collection module, used to obtain the resource status information of the computing nodes; an expert status monitoring module, used to obtain the training sample set of the hybrid expert model, input the training samples obtained by multiple sampling from the training sample set into the hybrid expert model respectively, and monitor the activation state parameters of the expert network; a calculation module, used to calculate the time consumption prediction result of the computing node performing iterative training based on the activation state parameters of the expert network and the resource status information of the computing node; and determine the time consumption prediction result of the heterogeneous computing system performing iterative training based on the time consumption prediction result of the computing node.

[0121] The description of the features in the embodiment corresponding to the expert parallel training time consumption prediction device can be found in the relevant description of the embodiment corresponding to the expert parallel training time consumption prediction method, which will not be repeated here.

[0122] An embodiment of the present invention further provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned expert parallel training time consumption prediction method embodiments.

[0123] An embodiment of the present invention further provides a non-volatile storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned expert parallel training time consumption prediction method embodiments when running.

[0124] In an exemplary embodiment, the non-volatile storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0125] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned expert parallel training time consumption prediction method embodiments are implemented.

[0126] An embodiment of the present invention also provides another computer program product, including a non-volatile storage medium, the non-volatile storage medium storing a computer program, and when the computer program is executed by a processor, implementing the steps of any of the above-mentioned expert parallel training time consumption prediction method embodiments.

[0127] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0128] The above is a detailed introduction to the expert parallel training time prediction method, device, equipment, medium and product provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A method for predicting the time consumption of expert parallel training, characterized in that: include: Determine an expert network of hybrid expert models deployed on computing nodes of a heterogeneous computing system; Obtaining a training sample set of the hybrid expert model, inputting training samples obtained by multiple sampling from the training sample set into the hybrid expert model respectively, and monitoring activation state parameters of the expert network; Calculating a prediction result of the time consumption of the computing node to perform iterative training based on the activation state parameters of the expert network and the resource state information of the computing node; Determining a time consumption prediction result of the heterogeneous computing system performing iterative training according to the time consumption prediction result of the computing node; The step of monitoring the activation state parameters of the expert network includes: Monitoring the activation state of the expert network during historical iterative training; Determining the activation probability of the expert network in one iterative training according to the activation state of the expert network in multiple rounds of historical iterative training, and using the activation probability as the activation state parameter of the expert network; The method further comprises calculating, based on the activation state parameters of the expert network and the resource state information of the computing node, a time consumption prediction result of the computing node performing iterative training, including: Calculate a forward propagation computation time prediction result of the computing node and a backward propagation computation time prediction result of the computing node based on the activation probability of the expert network, the computing resource parameters of the computing node, and the memory bandwidth of the computing node; According to the activation probability of the expert network and the network resource parameters of the computing node, a forward propagation communication time prediction result of the computing node and a backward propagation communication time prediction result of the computing node are calculated.

2. The expert parallel training time-consuming prediction method according to claim 1, characterized in that: Monitor the activation status of the expert network in historical iterative training, including: Obtaining a pre-trained gating network of the hybrid expert model; Sampling training samples from the training sample set and inputting them into the gating network, and monitoring the weights of the expert network output by the gating network; An activation state of the expert network is determined according to the weight of the expert network.

3. The expert parallel training time-consuming prediction method according to claim 2, characterized in that: Determining the activation state of the expert network according to the weight of the expert network includes: Obtaining the number of experts activated in one calculation of the hybrid expert model; The weights of the expert networks are sorted from large to small, and the expert networks corresponding to the number of experts activated are determined as activated expert networks, and the other expert networks are determined as inactivated expert networks.

4. The expert parallel training time-consuming prediction method according to claim 1, characterized in that: The forward propagation computation time prediction result of the computing node is calculated based on the activation probability of the expert network, the computing resource parameters of the computing node, and the memory bandwidth of the computing node, including: Calculate a first time consumption prediction result of the computing node according to the activation probability of the expert network on the computing node, the forward propagation calculation amount of the expert network, and the computing resource parameter of the computing node; Calculate a second time consumption prediction result of the computing node according to the activation probability of the expert network on the computing node, the amount of memory accessed by the forward propagation of the expert network, and the memory bandwidth of the computing node; A larger value between the first time consumption prediction result and the second time consumption prediction result is determined as the forward propagation calculation time consumption prediction result of the computing node.

5. The expert parallel training time-consuming prediction method according to claim 1, characterized in that: The back propagation calculation time consumption prediction result of the computing node is calculated based on the activation probability of the expert network, the computing resource parameters of the computing node, and the memory bandwidth of the computing node, including: Calculate a third time consumption prediction result of the computing node according to the activation probability of the expert network on the computing node, the back propagation calculation amount of the expert network, and the computing resource parameters of the computing node; Calculate a fourth time consumption prediction result of the computing node according to the activation probability of the expert network on the computing node, the amount of back-propagation memory access data of the expert network, and the memory bandwidth of the computing node; A larger value between the third time consumption prediction result and the fourth time consumption prediction result is determined as the back propagation calculation time consumption prediction result of the computing node.

6. The expert parallel training time-consuming prediction method according to claim 1, characterized in that: Calculating a forward propagation communication time prediction result of the computing node according to the activation probability of the expert network and the network resource parameters of the computing node includes: Determining the amount of first expert input data output by the computing node to the other computing nodes in forward propagation according to the activation probability of the expert network on the computing node and the amount of expert input data sent by the expert network to the computing nodes where the other activated expert networks are located when the expert network is activated; Determining the amount of first expert output data input by the computing node from other computing nodes in the forward propagation, based on the activation probability of the expert network on the computing node and the amount of expert output data output by the computing nodes where other activated expert networks are located when the expert network is activated; The forward propagation communication time prediction result of the computing node is calculated based on the sum of the first expert input data volume and the first expert output data volume divided by the bandwidth of the computing node's external communication link and added to the communication delay of the computing node's external communication link.

7. The expert parallel training time-consuming prediction method according to claim 1, characterized in that: The back propagation communication time prediction result of the computing node is calculated based on the activation probability of the expert network and the network resource parameters of the computing node, including: Determining the amount of second expert output data output by the computing node from other computing nodes that are located at the same time as the expert network is activated, based on the activation probability of the expert network on the computing node and the amount of expert output data output by the computing node where the other activated expert networks are located when the expert network is activated; Determining the amount of second expert input data output by the computing node to the other computing nodes in back propagation according to the activation probability of the expert network on the computing node and the amount of expert input data sent by the expert network to the computing nodes where the other activated expert networks are located when the expert network is activated; The predicted result of the back propagation communication time consumption of the computing node is calculated based on the sum of the second expert input data volume and the second expert output data volume divided by the bandwidth of the external communication link of the computing node and added to the communication delay of the external communication link of the computing node.

8. The expert parallel training time-consuming prediction method according to claim 1, characterized in that: One computing node deploys one expert network; Determining a predicted result of time consumption for executing iterative training by the heterogeneous computing system according to the predicted result of time consumption of the computing node includes: Calculating a forward propagation computation time prediction result of the heterogeneous computing system based on the forward propagation computation time prediction results of the plurality of computing nodes and the activation probability of the expert network; Calculating a backpropagation calculation time prediction result of the heterogeneous computing system based on the backpropagation calculation time prediction results of the plurality of computing nodes and the activation probability of the expert network; Calculating a forward propagation communication time prediction result of the heterogeneous computing system based on the forward propagation communication time prediction results of the plurality of computing nodes and the activation probability of the expert network; Calculating a backpropagation communication time prediction result of the heterogeneous computing system based on the backpropagation communication time prediction results of the plurality of computing nodes and the activation probability of the expert network; The sum of the forward propagation computing time prediction result of the heterogeneous computing system, the backward propagation computing time prediction result of the heterogeneous computing system, the forward propagation communication time prediction result of the heterogeneous computing system and the backward propagation communication time prediction result of the heterogeneous computing system is used as the time consumption prediction result of the heterogeneous computing system performing iterative training.

9. The expert parallel training time-consuming prediction method according to claim 8, characterized in that: The forward propagation computation time prediction result of the heterogeneous computing system is calculated based on the forward propagation computation time prediction results of the plurality of computing nodes and the activation probability of the expert network, including: Arrange the forward propagation computation time prediction results of the computing nodes in descending order and substitute them into each element of the expected computation time calculation formula respectively; substitute the corresponding activation probability of the expert network into each element of the expected computation time calculation formula to calculate and obtain the forward propagation computation time prediction result of the heterogeneous computing system; Among them, the time expectation calculation formula is to calculate the time expectation based on the sum of multiple elements, corresponding to the order of the input time prediction results from large to small. The elements of the time expectation calculation formula are the input time prediction result, the activation probability of the expert network corresponding to the input time prediction result, and the product of the probability that the expert network is not activated, which is greater than the probability corresponding to the input time prediction result.

10. The expert parallel training time-consuming prediction method according to claim 8, characterized in that: The method further comprises calculating a back propagation calculation time prediction result of the heterogeneous computing system based on the back propagation calculation time prediction results of the plurality of computing nodes and the activation probability of the expert network, including: Arrange the back propagation calculation time consumption prediction results of the computing nodes in descending order and substitute them into each element of the expected time consumption calculation formula respectively, substitute the corresponding activation probability of the expert network into each element of the expected time consumption calculation formula, and calculate the back propagation calculation time consumption prediction result of the heterogeneous computing system; Among them, the time expectation calculation formula is to calculate the time expectation based on the sum of multiple elements, corresponding to the order of the input time prediction results from large to small. The elements of the time expectation calculation formula are the input time prediction result, the activation probability of the expert network corresponding to the input time prediction result, and the product of the probability that the expert network is not activated, which is greater than the probability corresponding to the input time prediction result.

11. The method for predicting the time consumption of expert parallel training according to claim 8, characterized in that: The forward propagation communication time prediction result of the heterogeneous computing system is calculated based on the forward propagation communication time prediction results of the plurality of computing nodes and the activation probability of the expert network, including: Arrange the forward propagation communication time consumption prediction results of the computing nodes in descending order and substitute them into each element of the expected time consumption calculation formula respectively, substitute the corresponding activation probability of the expert network into each element of the expected time consumption calculation formula, and calculate the forward propagation communication time consumption prediction result of the heterogeneous computing system; Among them, the time expectation calculation formula is to calculate the time expectation based on the sum of multiple elements, corresponding to the order of the input time prediction results from large to small. The elements of the time expectation calculation formula are the input time prediction result, the activation probability of the expert network corresponding to the input time prediction result, and the product of the probability that the expert network is not activated, which is greater than the probability corresponding to the input time prediction result.

12. The method for predicting the time consumption of expert parallel training according to claim 8, characterized in that: The method further comprises calculating a back propagation communication time prediction result of the heterogeneous computing system based on the back propagation communication time prediction results of the plurality of computing nodes and the activation probability of the expert network, including: Arrange the backpropagation communication time prediction results of the computing nodes in descending order and substitute them into each element of the expected time calculation formula respectively, substitute the corresponding activation probability of the expert network into each element of the expected time calculation formula, and calculate the backpropagation communication time prediction result of the heterogeneous computing system; Among them, the time expectation calculation formula is to calculate the time expectation based on the sum of multiple elements, corresponding to the order of the input time prediction results from large to small. The elements of the time expectation calculation formula are the input time prediction result, the activation probability of the expert network corresponding to the input time prediction result, and the product of the probability that the expert network is not activated, which is greater than the probability corresponding to the input time prediction result.

13. The expert parallel training time-consuming prediction method according to claim 1, characterized in that: One computing node deploys one or more expert networks.

14. The method for predicting the time consumption of expert parallel training according to claim 13, characterized in that: Determining a predicted result of time consumption for executing iterative training by the heterogeneous computing system according to the predicted result of time consumption of the computing node includes: Taking the maximum value among the forward propagation computation time consumption prediction results of the plurality of computing nodes as the forward propagation computation time consumption prediction result of the heterogeneous computing system; Taking the maximum value among the back propagation calculation time consumption prediction results of the plurality of computing nodes as the back propagation calculation time consumption prediction result of the heterogeneous computing system; Taking the maximum value among the forward propagation communication time consumption prediction results of the plurality of computing nodes as the forward propagation communication time consumption prediction result of the heterogeneous computing system; Taking the maximum value among the reverse propagation communication time consumption prediction results of the plurality of computing nodes as the reverse propagation communication time consumption prediction result of the heterogeneous computing system; The sum of the forward propagation computing time prediction result of the heterogeneous computing system, the backward propagation computing time prediction result of the heterogeneous computing system, the forward propagation communication time prediction result of the heterogeneous computing system and the backward propagation communication time prediction result of the heterogeneous computing system is used as the time consumption prediction result of the heterogeneous computing system performing iterative training.

15. An expert parallel training time-consuming prediction device, characterized in that: include: A task information collection module, used to determine the expert network of the hybrid expert model deployed by the computing nodes of the heterogeneous computing system; A system information collection module, configured to obtain resource status information of the computing node; An expert state monitoring module is used to obtain a training sample set of the hybrid expert model, input training samples obtained by multiple sampling from the training sample set into the hybrid expert model, and monitor the activation state parameters of the expert network; A calculation module, configured to calculate a time-consuming prediction result of the iterative training performed by the computing node based on the activation state parameters of the expert network and the resource state information of the computing node; Determining a time consumption prediction result of the heterogeneous computing system performing iterative training according to the time consumption prediction result of the computing node; The step of monitoring the activation state parameters of the expert network includes: Monitoring the activation state of the expert network during historical iterative training; Determining the activation probability of the expert network in one iterative training according to the activation state of the expert network in multiple rounds of historical iterative training, and using the activation probability as the activation state parameter of the expert network; The method further comprises calculating, based on the activation state parameters of the expert network and the resource state information of the computing node, a time consumption prediction result of the computing node performing iterative training, including: Calculate a forward propagation computation time prediction result of the computing node and a backward propagation computation time prediction result of the computing node based on the activation probability of the expert network, the computing resource parameters of the computing node, and the memory bandwidth of the computing node; According to the activation probability of the expert network and the network resource parameters of the computing node, a forward propagation communication time prediction result of the computing node and a backward propagation communication time prediction result of the computing node are calculated.

16. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of the expert parallel training time consumption prediction method as described in any one of claims 1 to 14 when executing the computer program.

17. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the expert parallel training time consumption prediction method according to any one of claims 1 to 14 are implemented.

18. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the expert parallel training time consumption prediction method according to any one of claims 1 to 14 are implemented.

Citation Information

Patent Citations

  • Heterogeneous computing system and training time consumption prediction method, equipment, medium and product thereof

    CN119204360A

  • Hybrid expert model training method and device, computer equipment, readable storage medium and program product

    CN119830952A