Expert parallel training time consumption prediction method, device, equipment, medium and product

By sampling the training sample set of the hybrid expert model multiple times to monitor the activation state parameters of the expert network, combined with the resource state information of the heterogeneous computing system, the problem of inaccurate prediction of parallel training of hybrid expert model experts is solved, and training efficiency and resource utilization are improved.

CN120373423AActive Publication Date: 2025-07-25SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510874226.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The existing model training time-consuming prediction scheme cannot accurately predict expert parallel training of mixed expert models, resulting in low training efficiency and low resource utilization.

Method used

By sampling the training sample set of the mixed expert model multiple times, monitoring the activation state parameters of the expert network, and combining the computing node resource status information of the heterogeneous computing system, the time-consuming prediction results of iterative training are calculated to achieve accurate prediction of the time-consuming time-consuming expert training.

Benefits of technology

The efficiency of expert parallel training is improved, the resource utilization rate of heterogeneous computing systems is optimized, and the accuracy and efficiency of the training process is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373423A_ABST
    Figure CN120373423A_ABST
Patent Text Reader

Abstract

The invention discloses an expert parallel training time consumption prediction method and device, equipment, a medium and a product, and relates to the technical field of artificial intelligence. The method comprises the following steps: inputting training samples obtained by sampling a training sample set adopted by hybrid expert model training for multiple times into a hybrid expert model to monitor activation state parameters of an expert network of the hybrid expert model, so as to accurately predict activation conditions of the expert network in expert parallel training; according to the activation state parameters of the expert network and the resource state information of the computing nodes of the heterogeneous computing system, computing a time consumption prediction result of the computing nodes for executing iterative training, and determining a time consumption prediction result of the heterogeneous computing system for executing the iterative training according to the time consumption prediction result of the computing nodes, accurate prediction is carried out on the time consumption of expert parallel training of the hybrid expert model, and the problem that the time consumption of expert parallel training of the hybrid expert model cannot be accurately predicted by a model training time consumption prediction scheme in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment, medium and product for predicting the time consumption of parallel expert training. Background Art

[0002] With the development of artificial intelligence (AI) technology, the scale of the neural network models adopted has been continuously increasing, resulting in an increasing hardware cost and model training time consumption required for model training. The Mixture of experts (MoE) is a machine learning method. For the Mixture of experts model, a parallel expert training method is adopted, and different expert networks are deployed on different computing nodes, which can effectively improve the training efficiency of the Mixture of experts model. Due to the sparse activation calculation method adopted by the Mixture of experts model, traditional model training time consumption prediction schemes cannot be applied to the Mixture of experts model. Summary of the Invention

[0003] The present invention provides a method, device, equipment, medium and product for predicting the time consumption of parallel expert training, so as to at least solve the problem that the model training time consumption prediction scheme in the related art cannot accurately predict the time consumption of parallel expert training of the Mixture of experts model.

[0004] The present invention provides a method for predicting the time consumption of parallel expert training, including: Determining the expert networks of the Mixture of experts model deployed on the computing nodes of the heterogeneous computing system; Obtaining the training sample set of the Mixture of experts model, respectively inputting the training samples obtained by sampling multiple times from the training sample set into the Mixture of experts model, and monitoring the activation state parameters of the expert networks; Calculating the time consumption prediction result of the computing node for performing iterative training according to the activation state parameters of the expert networks and the resource state information of the computing node; Determining the time consumption prediction result of the heterogeneous computing system for performing iterative training according to the time consumption prediction result of the computing node.

[0005] The present invention also provides a device for predicting the time consumption of parallel expert training, including: A task information collection module, configured to determine the expert networks of the Mixture of experts model deployed on the computing nodes of the heterogeneous computing system; A system information collection module, configured to obtain the resource state information of the computing node; An expert state monitoring module, configured to obtain the training sample set of the Mixture of experts model, respectively input the training samples obtained by sampling multiple times from the training sample set into the Mixture of experts model, and monitor the activation state parameters of the expert networks; A calculation module, configured to calculate a time-consuming prediction result of the computing node for performing iterative training according to the activation state parameter of the expert network and the resource state information of the computing node; and determine a time-consuming prediction result of the heterogeneous computing system for performing iterative training according to the time-consuming prediction result of the computing node.

[0006] The present invention further provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any one of the above-mentioned expert parallel training time-consuming prediction methods when executing the computer program.

[0007] The present invention further provides a non-volatile storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above-mentioned expert parallel training time-consuming prediction methods are implemented.

[0008] The present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any one of the above-mentioned expert parallel training time-consuming prediction methods are implemented.

[0009] Through the present invention, since the training samples obtained by sampling the training sample set used for training the mixture of experts model multiple times are respectively input into the mixture of experts model to monitor the activation state parameters of the expert network of the mixture of experts model, the activation situation of the expert network in expert parallel training can be accurately predicted; according to the activation state parameter of the expert network and the resource state information of the computing nodes of the heterogeneous computing system, calculate the time-consuming prediction result of the computing node for performing iterative training, and determine the time-consuming prediction result of the heterogeneous computing system for performing iterative training according to the time-consuming prediction result of the computing node, so as to accurately predict the time consumption of the expert parallel training of the mixture of experts model, solve the problem that the model training time-consuming prediction scheme in the related art cannot accurately predict the time consumption of the expert parallel training of the mixture of experts model, and the accurate expert parallel time-consuming prediction result obtained can help to pre-optimize the expert parallel mode of the heterogeneous computing system, thereby improving the efficiency of expert parallel training and the resource utilization rate of the heterogeneous computing system. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0011] Figure 1 It is a flowchart of the first expert parallel training time-consuming prediction method provided by the embodiment of the present invention; Figure 2Schematic diagram of the deployment architecture of the first hybrid expert model provided by an embodiment of the present invention; Figure 3 Schematic diagram of the deployment architecture of the second hybrid expert model provided by an embodiment of the present invention; Figure 4 Flowchart of the second expert parallel training time-consuming prediction method provided by an embodiment of the present invention. Detailed implementation manners

[0012] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0013] It should be noted that in the description of the present invention, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0014] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0015] To solve the problem that the model training time-consuming prediction scheme in the related art cannot accurately predict the time-consuming of the expert parallel training of the hybrid expert model, an embodiment of the present invention provides an expert parallel training time-consuming prediction scheme. By inputting the training samples obtained by sampling the training sample set used in the hybrid expert model training into the hybrid expert model multiple times to monitor the activation state parameters of the expert network of the hybrid expert model, the activation situation of the expert network in the expert parallel training can be accurately predicted; according to the activation state parameters of the expert network and the resource state information of the computing nodes of the heterogeneous computing system, calculate the time-consuming prediction result of the computing node to execute iterative training, and determine the time-consuming prediction result of the heterogeneous computing system to execute iterative training according to the time-consuming prediction result of the computing node, so as to accurately predict the time-consuming of the expert parallel training of the hybrid expert model, solve the problem that the model training time-consuming prediction scheme in the related art cannot accurately predict the time-consuming of the expert parallel training of the hybrid expert model, and the accurate expert parallel time-consuming prediction result obtained can help to pre-optimize the expert parallel manner of the heterogeneous computing system, thereby improving the efficiency of expert parallel training and the resource utilization rate of the heterogeneous computing system.

[0016] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the expert parallel training time-consuming prediction scheme depends, the specific application environment architecture or specific hardware architecture is described herein.

[0017] The expert parallel training time-consuming prediction scheme provided by the present invention can be deployed based on a heterogeneous computing system. The heterogeneous computing system includes multiple computing devices (computing nodes), and parameters such as the computing core type, computing power, memory size, and communication bandwidth of different computing nodes may be different.

[0018] In the heterogeneous computing system targeted by the present invention, the computing core type of the computing node may include, but is not limited to, one or more of a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), a neural network processor (NPU), a microcontroller unit (MCU), and an application-specific integrated circuit (ASIC).

[0019] In the heterogeneous computing system targeted by the present invention, different computing nodes can be interconnected through a network, such as through an Ethernet; or can be interconnected through a bus, such as through a peripheral component interconnect express (PCIe) or an NVIDIA high-speed interconnect bus.

[0020] The expert parallel training time-consuming prediction method provided by the embodiments of the present invention can be applied to one or more computing nodes in the heterogeneous computing system, or can be applied to a control node outside the heterogeneous computing system.

[0021] The embodiments of the present invention provide an expert parallel training time-consuming prediction method. In combination with the execution process of the expert parallel training time-consuming prediction method, the method is described in detail below.

[0022] Figure 1 It is a flowchart of the first expert parallel training time-consuming prediction method provided by the embodiments of the present invention; Figure 2 It is a schematic diagram of the deployment architecture of the first hybrid expert model provided by the embodiments of the present invention; Figure 3Schematic diagram of the deployment architecture of the second hybrid expert model provided by the embodiments of the present invention.

[0023] As Figure 1 shown, the method for predicting the time-consuming of parallel training of experts provided by the embodiments of the present invention may include: S101: Determine the expert network of the hybrid expert model deployed on the computing nodes of the heterogeneous computing system.

[0024] S102: Obtain the training sample set of the hybrid expert model, input the training samples obtained by sampling multiple times from the training sample set into the hybrid expert model respectively, and monitor the activation state parameters of the expert network.

[0025] S103: Calculate the time-consuming prediction result of the computing node for performing iterative training according to the activation state parameters of the expert network and the resource state information of the computing node.

[0026] S104: Determine the time-consuming prediction result of the heterogeneous computing system for performing iterative training according to the time-consuming prediction result of the computing node.

[0027] In some alternative embodiments of the embodiments of the present invention, one expert network may be deployed on one computing node. As Figure 2 shown, it is recorded that the gating network, expert network 1, and expert network 2 are deployed on computing node 1, and the gating network, expert network 3, and expert network 4 are deployed on computing node 2. After the input data 1 enters computing node 1, the expert network 1 and expert network 3 of the hybrid expert model are activated through the decision of the gating network, and then the expert outputs of expert network 1 and expert network 3 are aggregated to obtain the output result 1. After the input data 2 enters computing node 2, the expert network 2 and expert network 4 of the hybrid expert model are activated through the decision of the gating network, and then the expert outputs of expert network 2 and expert network 4 are aggregated to obtain the output result 2. In one iterative training, if the input is input data 1, then expert network 1 and expert network 3 are activated, and the time-consuming on computing node 1 is the computing time-consuming of expert network 1, the memory access time-consuming of computing node 1, and the expert input / output time-consuming of expert network 1, and the time-consuming on computing node 2 is the computing time-consuming of expert network 3, the access time-consuming of computing node 2, and the expert input / output time-consuming of expert network 3.

[0028] In some other alternative embodiments of the embodiments of the present invention, one or more expert networks may also be deployed on one computing node. As Figure 3 shown, it is recorded that expert network 1 is deployed on computing node 1, expert network 2 is deployed on computing node 2,..., expert network n is deployed on computing node n. After the input data passes through the decision of the gating network and activates expert network 1 and expert network 2, the input data is sent to computing node 1 and computing node 2, and then the expert outputs of expert network 1 and expert network 2 are aggregated to obtain the output result.

[0029] It should be noted that the above Figure 2 、 Figure 3 The deployment architectures shown are only applicable to the cases where only one expert network is deployed on one computing node or one or more expert networks are deployed on one computing node. In actual applications, other specific deployment methods can be adopted. In the embodiments of the present invention, one or more expert networks can be deployed on one computing node, and the same expert network can be deployed on different computing nodes.

[0030] The embodiments of the present invention can be applied before deploying a mixture of experts model to a heterogeneous computing system. By predicting the time consumption of the heterogeneous computing system for executing expert parallel training according to the expert activation situation, the expert parallel mode of the heterogeneous computing system can be optimized, the efficiency of expert parallel training can be improved, and the resource utilization rate of the heterogeneous computing system can be increased.

[0031] For S101, determine the information of the expert parallel training task for which time consumption prediction is required, that is, determine the expert networks deployed on each computing node of the heterogeneous computing system.

[0032] For S102, the training sample set is a set of samples used to train the mixture of experts model. The gating network is used to make a decision on which expert network to activate according to the type of input data. Therefore, the activation situation of the expert network can be monitored by sampling the training samples and inputting them into the mixture of experts model in the same way as training the mixture of experts model, and the activation state parameter can be measured.

[0033] In the embodiments of the present invention, monitoring the activation state parameter of the expert network in S102 may include: monitoring the activation state of the expert network in historical iterative training; determining the activation probability of the expert network in one iterative training according to the activation states of the expert network in multiple rounds of historical iterative training, and using the activation probability as the activation state parameter of the expert network. The activation state of the expert network in multiple rounds of historical iterative training can be monitored by pre-deploying the mixture of experts model or obtaining the historical iterative training data when the mixture of experts model is trained based on the same training samples. For each expert network, dividing the number of activation times in multiple rounds of historical iterative training by the number of historical iterative training rounds can obtain the activation probability of the expert network.

[0034] In the embodiments of the present invention, monitoring the activation state of the expert network in historical iterative training may include: obtaining the gating network of the pre-trained mixture of experts model; sampling training samples from the training sample set and inputting them into the gating network, and monitoring the weights of the expert networks output by the gating network; determining the activation state of the expert network according to the weights of the expert networks. Since the deployment of the mixture of experts network requires high hardware costs, the activation probability of the expert network under the training sample set can be determined by deploying the pre-trained gating network.

[0035] Since the sparse activation strategy of the mixture-of-experts network is to select the top k expert networks for activation according to the output of the gating network, determining the activation status of the expert networks according to the weights of the expert networks may include: obtaining the number of expert activations in one calculation of the mixture-of-experts model; sorting the weights of the expert networks from largest to smallest, and determining the corresponding number of expert networks with the largest weights as the activated expert networks according to the number of expert activations, and other expert networks as the unactivated expert networks.

[0036] For S103, the predicted result of the time consumption of the computing node in one iterative training can be calculated according to the activation probability of the expert network when the mixture-of-experts model is trained using the training sample set, in combination with the resource status information of the computing nodes of the heterogeneous computing system and the expert networks deployed on the computing nodes.

[0037] For S104, considering that the time consumption of one iterative training of the heterogeneous computing system is affected by the computing node with the largest time consumption, the predicted result of the time consumption of the heterogeneous computing system for iterative training can be determined according to the predicted results of the time consumption of each computing node.

[0038] The method for predicting the time consumption of expert parallel training provided by the embodiments of the present invention accurately predicts the activation situation of the expert networks in expert parallel training by respectively inputting the training samples obtained by sampling the training sample set used for training the mixture-of-experts model multiple times into the mixture-of-experts model to monitor the activation status parameters of the expert networks of the mixture-of-experts model; calculates the predicted result of the time consumption of the computing node for iterative training according to the activation status parameters of the expert networks and the resource status information of the computing nodes of the heterogeneous computing system, and determines the predicted result of the time consumption of the heterogeneous computing system for iterative training according to the predicted results of the time consumption of the computing nodes, so as to accurately predict the time consumption of the expert parallel training of the mixture-of-experts model, solve the problem that the model training time consumption prediction scheme in the related art cannot accurately predict the time consumption of the expert parallel training of the mixture-of-experts model, and the obtained accurate predicted result of the expert parallel time consumption can help to pre-optimize the expert parallel mode of the heterogeneous computing system, thereby improving the efficiency of expert parallel training and the resource utilization rate of the heterogeneous computing system.

[0039] Figure 4 It is a flowchart of the second method for predicting the time consumption of expert parallel training provided by the embodiments of the present invention.

[0040] On the basis of the above embodiments, the embodiments of the present invention continue to describe the steps of predicting the time consumption of the heterogeneous computing system for executing the expert parallel training task according to the activation status parameters of the expert networks.

[0041] In the embodiment of the present invention, in S103, according to the activation state parameter of the expert network and the resource state information of the computing node, the predicted result of the time consumption for the computing node to perform iterative training can be calculated, including: according to the activation probability of the expert network, the computing power resource parameter of the computing node, and the memory bandwidth of the computing node, calculating the predicted result of the forward propagation calculation time consumption of the computing node and the predicted result of the backward propagation calculation time consumption of the computing node; according to the activation probability of the expert network and the network resource parameter of the computing node, calculating the predicted result of the forward propagation communication time consumption of the computing node and the predicted result of the backward propagation communication time consumption of the computing node.

[0042] As Figure 4 shown, the method for predicting the time consumption of expert parallel training provided by the embodiment of the present invention can be implemented based on four modules: a task information collection module, a system information collection module, an expert state monitoring module, and a calculation module.

[0043] Among them, the task information collection module is used to parse or calculate the necessary input information according to the information of the expert parallel training task input by the user, including: the forward propagation calculation amount and the backward propagation calculation amount of the expert network, which can be represented by the number of floating point operations (FLOPs). The forward propagation calculation amount can be estimated according to the neural layer composition of the expert network, and the calculation amount is equal to the sum of the calculation amounts of all neural layers. The calculation amount of the neural layer can be estimated using mathematical methods. For example, for the forward calculation of the fully connected layer, assuming the input data dimension (N, D), the hidden layer weight dimension (D, out), and the output (N, out), then the calculation amount is FLOPs = N x (2 * D - 1) * out. Different neural layers have different calculation methods, which will not be listed one by one here. The backward propagation calculation amount can be estimated by multiplying the forward propagation calculation amount by 2. The forward propagation calculation amount and the backward propagation calculation amount can also be statistically calculated using existing open source tools such as torchstat. Usually, the network shapes of the expert networks are the same, so only the calculation amount of one expert network needs to be calculated.

[0044] The input data volume and output data volume of the expert neural network. The input data volume can be estimated by multiplying the batch size by the tensor size of the input to the expert neural network. For example, if the batch size is 10 and the input size of the neural network layer is [10, 10], then the input data volume is 10 * 10 * 10 * data precision (e.g., for 32-bit floating-point number fp32, it is 4 bytes). The above information can also be statistically analyzed using existing open-source tools such as torchstat. The output data volume can also be estimated by multiplying the output tensor size of the expert network by the batch size. The above information can also be statistically analyzed using existing open-source tools such as torchstat.

[0045] The memory access data volume during the forward propagation and backward propagation of the expert network. For the memory access data volume during the forward propagation of the expert network, the size of this memory data can be the input data volume of the expert network + the number of parameters. Among them, the input data volume is known, and the number of parameters can be statistically analyzed using existing open-source tools such as torchstat or calculated mathematically according to the network structure of the expert network. For the memory access data volume during the backward propagation of the expert network, it can be calculated through data statistics. For example, given that the number of model parameters is X, in PyTorch automatic mixed-precision training, the size of this data volume includes parameters, gradients, optimizer states, and activation values. Among them, the data volume of parameters, gradients, and optimizer states is 16X bytes, and the activation values are equal to the output of each layer of the expert neural network plus the input of the first layer (known through the calculation of the input data volume and output data volume of the above expert neural network). The above information can also be statistically analyzed using existing open-source tools such as torchstat.

[0046] The expert networks allocated on each computing node during expert parallel training. This data can be specified according to manual input, generally by writing a program. Each unique computing node needs to collect information about the expert networks running on it.

[0047] The total number of computing nodes participating in the task.

[0048] The total number of expert networks.

[0049] If the above information cannot be collected completely, the task information collection module notifies the user to make corresponding inputs. The above collected information will be sent to the computing module for predicting the time consumption of iterative training.

[0050] The system information collection module is used to collect the resource status information of the computing nodes in the heterogeneous computing system. Specifically, it can include the computing power resource parameters of each computing node. If the computing node is only used to execute the inference calculation task of the mixture-of-experts model, the peak computing power parameter of the computing node (in floating-point operations per second, FLOPS) can be adopted, and this parameter can be obtained through the product manual of the computing node or actual tests. The memory bandwidth of the computing node. If the computing node is only used to execute the inference calculation task of the mixture-of-experts model, the peak memory bandwidth of the computing node can be adopted, and this parameter can be obtained through the product manual of the computing node or actual tests. The network resource parameters of the computing node, specifically including the bandwidth and delay information of the external links of the computing node. These parameters can be obtained by testing with some standard (benchmark) tools such as iperf.

[0051] The system status monitoring module is used to monitor the activation probability of the expert network during training based on the training sample set. It can intercept adjacent multiple rounds of historical iterative training and calculate the probability that the expert network is activated in each round of iterative training.

[0052] According to the information collected by the three modules of the task information collection module, the system information collection module, and the system status monitoring module, the calculation module predicts the time consumption of each computing node in the parallel training of experts according to the deployment method of the mixture-of-experts model in the heterogeneous computing system, and predicts the time consumption of the heterogeneous computing system for iterative training according to the predicted results of the time consumption of the computing nodes.

[0053] The following further introduces the calculation steps of the calculation module.

[0054] In the embodiment of the present invention, according to the activation probability of the expert network, the computing power resource parameters of the computing node, and the memory bandwidth of the computing node, the predicted result of the forward propagation calculation time consumption of the computing node can be calculated, which can include: calculating the first predicted result of the time consumption of the computing node according to the activation probability of the expert network on the computing node, the forward propagation calculation amount of the expert network, and the computing power resource parameters of the computing node; calculating the second predicted result of the time consumption of the computing node according to the activation probability of the expert network on the computing node, the forward propagation memory access data volume of the expert network, and the memory bandwidth of the computing node; determining the larger value of the first predicted result and the second predicted result as the predicted result of the forward propagation calculation time consumption of the computing node.

[0055] In an embodiment of the present invention, according to the activation probability of the expert network, the computing power resource parameters of the computing node, and the memory bandwidth of the computing node, the predicted result of the backpropagation computing time of the computing node can be calculated, which may include: calculating a third predicted time result of the computing node according to the activation probability of the expert network on the computing node, the backpropagation computation amount of the expert network, and the computing power resource parameters of the computing node; calculating a fourth predicted time result of the computing node according to the activation probability of the expert network on the computing node, the memory access data amount of the backpropagation of the expert network, and the memory bandwidth of the computing node; and determining the larger value of the third predicted time result and the fourth predicted time result as the predicted result of the backpropagation computing time of the computing node.

[0056] In an embodiment of the present invention, according to the activation probability of the expert network and the network resource parameters of the computing node, the predicted result of the forward propagation communication time of the computing node can be calculated, which may include: determining a first expert input data amount output by the computing node to other computing nodes in the forward propagation according to the activation probability of the expert network on the computing node and the amount of expert input data sent by the computing node where the expert network is activated to the computing nodes where other activated expert networks are located when the expert network is activated; determining a first expert output data amount input by the computing node from other computing nodes in the forward propagation according to the activation probability of the expert network on the computing node and the amount of expert output data output by the computing nodes where other activated expert networks are located when the expert network is activated; and calculating the predicted result of the forward propagation communication time of the computing node by dividing the sum of the first expert input data amount and the first expert output data amount by the bandwidth of the external communication link of the computing node and then adding the communication delay of the external communication link of the computing node.

[0057] In an embodiment of the present invention, according to the activation probability of the expert network and the network resource parameters of the computing node, the predicted result of the backpropagation communication time of the computing node can be calculated, which may include: determining a second expert output data amount received by the computing node from other computing nodes in the backpropagation according to the activation probability of the expert network on the computing node and the amount of expert output data output by the computing nodes where other activated expert networks are located when the expert network is activated; determining a second expert input data amount output by the computing node to other computing nodes in the backpropagation according to the activation probability of the expert network on the computing node and the amount of expert input data sent by the computing node where the expert network is activated to the computing nodes where other activated expert networks are located when the expert network is activated; and calculating the predicted result of the backpropagation communication time of the computing node by dividing the sum of the second expert input data amount and the second expert output data amount by the bandwidth of the external communication link of the computing node and then adding the communication delay of the external communication link of the computing node.

[0058] Based on the information collected by the task information collection module, the system information collection module, and the system status monitoring module, the input parameters of the calculation module may include: the total number of computing nodes participating in the task , the total number of expert networks (when one expert network is deployed on one computing node, = ), the bandwidth of the external communication link of the computing node , the latency of the external communication link of the computing node , the computing power resource parameter of the computing node (in floating-point operations per second FLOPS), the memory bandwidth of the computing node , the forward propagation computation volume of the expert network , the backward propagation computation volume of the expert network (in floating-point operations per second FLOPS), the forward propagation memory access data volume of the expert network on the computing node , the backward propagation memory access data volume of the expert network on the computing node , the input data volume of the expert network , the output data volume of the expert network , the forward propagation memory access data volume of the expert network on the computing node , the backward propagation memory access data volume of the expert network on the computing node , the forward propagation memory access data volume of the expert network on the computing node , the backward propagation memory access data volume of the expert network on the computing node , the input data volume of the expert network , the output data volume of the expert network .

[0059] Combined with the time-consuming prediction principle introduced in the embodiments of the present invention, the following is introduced in combination with an actual application scenario.

[0060] For the scenario where one expert network is deployed on one computing node, in S104, determining the time-consuming prediction result of the heterogeneous computing system for iterative training according to the time-consuming prediction result of the computing node may include: calculating the forward propagation computing time-consuming prediction result of the heterogeneous computing system based on the forward propagation computing time-consuming prediction results of multiple computing nodes and the activation probability of the expert network; calculating the backward propagation computing time-consuming prediction result of the heterogeneous computing system based on the backward propagation computing time-consuming prediction results of multiple computing nodes and the activation probability of the expert network; calculating the forward propagation communication time-consuming prediction result of the heterogeneous computing system based on the forward propagation communication time-consuming prediction results of multiple computing nodes and the activation probability of the expert network; calculating the backward propagation communication time-consuming prediction result of the heterogeneous computing system based on the backward propagation communication time-consuming prediction results of multiple computing nodes and the activation probability of the expert network; using the sum of the forward propagation computing time-consuming prediction result of the heterogeneous computing system, the backward propagation computing time-consuming prediction result of the heterogeneous computing system, the forward propagation communication time-consuming prediction result of the heterogeneous computing system, and the backward propagation communication time-consuming prediction result of the heterogeneous computing system as the time-consuming prediction result of the heterogeneous computing system for iterative training.

[0061] That is to say, the time-consuming prediction results of the system are calculated for the four stages of forward propagation calculation, forward propagation communication, backward propagation calculation, and backward propagation communication respectively, and the sum is obtained as the time-consuming prediction result of the heterogeneous computing system in the first iteration training.

[0062] In the embodiment of the present invention, according to the forward propagation calculation time-consuming prediction results of multiple computing nodes and the activation probability of the expert network, calculating the forward propagation calculation time-consuming prediction result of the heterogeneous computing system may include: arranging the forward propagation calculation time-consuming prediction results of the computing nodes in descending order and substituting them into the respective elements of the expected time-consuming formula, substituting the corresponding activation probabilities of the expert network into the respective elements of the expected time-consuming formula, and calculating the forward propagation calculation time-consuming prediction result of the heterogeneous computing system; wherein, the expected time-consuming formula calculates the expected time-consuming according to the sum value of multiple elements, corresponding to the descending order of the input time-consuming prediction results, and the elements of the expected time-consuming formula are the input time-consuming prediction results, the product of the activation probability of the expert network corresponding to the input time-consuming prediction results, and the probability that the expert network corresponding to the input time-consuming prediction result greater than the input time-consuming prediction result is not activated.

[0063] In the embodiment of the present invention, according to the backward propagation calculation time-consuming prediction results of multiple computing nodes and the activation probability of the expert network, calculating the backward propagation calculation time-consuming prediction result of the heterogeneous computing system may include: arranging the backward propagation calculation time-consuming prediction results of the computing nodes in descending order and substituting them into the respective elements of the expected time-consuming formula, substituting the corresponding activation probabilities of the expert network into the respective elements of the expected time-consuming formula, and calculating the backward propagation calculation time-consuming prediction result of the heterogeneous computing system; wherein, the expected time-consuming formula calculates the expected time-consuming according to the sum value of multiple elements, corresponding to the descending order of the input time-consuming prediction results, and the elements of the expected time-consuming formula are the input time-consuming prediction results, the product of the activation probability of the expert network corresponding to the input time-consuming prediction results, and the probability that the expert network corresponding to the input time-consuming prediction result greater than the input time-consuming prediction result is not activated.

[0064] In an embodiment of the present invention, according to the forward propagation communication time consumption prediction results of multiple computing nodes and the activation probabilities of the expert networks, calculating the forward propagation communication time consumption prediction result of the heterogeneous computing system may include: arranging the forward propagation communication time consumption prediction results of the computing nodes in descending order and substituting them into the respective elements of the expected time consumption calculation formula, substituting the corresponding activation probabilities of the expert networks into the respective elements of the expected time consumption calculation formula, and calculating the forward propagation communication time consumption prediction result of the heterogeneous computing system; wherein, the expected time consumption calculation formula calculates the expected time consumption based on the sum of multiple elements. Corresponding to the order of the input time consumption prediction results from large to small, the elements of the expected time consumption calculation formula are the input time consumption prediction result, the product of the activation probability of the expert network corresponding to the input time consumption prediction result, and the probability that the expert network corresponding to the input time consumption prediction result is not activated and is greater than the input time consumption prediction result.

[0065] In an embodiment of the present invention, according to the backward propagation communication time consumption prediction results of multiple computing nodes and the activation probabilities of the expert networks, calculating the backward propagation communication time consumption prediction result of the heterogeneous computing system may include: arranging the backward propagation communication time consumption prediction results of the computing nodes in descending order and substituting them into the respective elements of the expected time consumption calculation formula, substituting the corresponding activation probabilities of the expert networks into the respective elements of the expected time consumption calculation formula, and calculating the backward propagation communication time consumption prediction result of the heterogeneous computing system; wherein, the expected time consumption calculation formula calculates the expected time consumption based on the sum of multiple elements. Corresponding to the order of the input time consumption prediction results from large to small, the elements of the expected time consumption calculation formula are the input time consumption prediction result, the product of the activation probability of the expert network corresponding to the input time consumption prediction result, and the probability that the expert network corresponding to the input time consumption prediction result is not activated and is greater than the input time consumption prediction result.

[0066] The expected time consumption calculation formula provided by the embodiment of the present invention can be expressed as: ; Wherein, represents the expected time consumption, , , …… represent the single-item time consumption prediction results of each computing node arranged in descending order, , , …… , represent the activation probabilities of the expert networks on the computing nodes corresponding to the single-item time consumption prediction results arranged in descending order. Among them, the types of single-item time consumption prediction results include forward propagation calculation time consumption prediction results, backward propagation calculation time consumption prediction results, forward propagation communication time consumption prediction results, and backward propagation communication time consumption prediction results.

[0067] The meaning of the formula for calculating the expected time consumption is that since the single-item time consumption of the heterogeneous computing system in a stage is affected by the computing node with the largest time consumption, if the expert network on the computing node with the largest time consumption is activated, the time consumption of other computing nodes does not need to be considered.

[0068] Then in the embodiment of the present invention, for the computing node the predicted result of the forward propagation computing time consumption can be calculated by the following formula: ; wherein, represents the forward propagation computing volume of the expert network, represents the computing power resource parameter of the computing node , represents the forward propagation memory access data volume of the expert network on the computing node , represents the memory bandwidth of the computing node , represents the maximum value calculation. , respectively represent the forward propagation computing time consumption and memory access time consumption of the computing node . The final forward propagation computing time consumption depends on one of the computing bottleneck and the memory access bottleneck.

[0069] After arranging the predicted results of the forward propagation computing time consumption of each computing node in descending order, substitute them into in the formula for calculating the expected time consumption respectively, , , ... , and substitute the activation probability of the corresponding expert network into , , ... , , so as to calculate the predicted result of the forward propagation computing time consumption of the heterogeneous computing system .

[0070] For the computing node the predicted result of the backward propagation computing time consumption can be calculated by the following formula: ; wherein, represents the forward propagation computing volume of the expert network, represents the computing power resource parameter of the computing node , represents the expert network on the computing node The amount of data fetched during forward propagation on represents the memory bandwidth of the computing node and represents the maximum value calculation. and respectively represent the time consumption and data fetching time consumption of the reverse propagation calculation of the computing node . The final reverse propagation calculation time consumption depends on either the calculation bottleneck or the data fetching bottleneck.

[0071] After arranging the predicted results of the forward propagation calculation time consumption of each computing node in descending order, substitute them into in the expected time consumption calculation formula respectively, and and ... , and substitute the activation probabilities of the corresponding expert networks into and and ... and in the expected time consumption calculation formula, so as to calculate and obtain the predicted result of the forward propagation calculation time consumption of the heterogeneous computing system .

[0072] The predicted result of the forward propagation communication time consumption of the computing node can be calculated by the following formula: ; where represents the amount of input data of the expert network, represents the amount of output data of the expert network, represents the bandwidth of the external communication link of the computing node , represents the delay of the external communication link of the computing node

[0073]

[0073] After arranging the predicted results of the forward propagation communication time consumption of each computing node in descending order, substitute them into in the expected time consumption calculation formula respectively, and and ... and substitute the activation probability of the corresponding expert network into the expected time-consuming calculation formula for , , …… , , so as to calculate the predicted result of the forward propagation communication time consumption of the heterogeneous computing system .

[0074] Computing node The predicted result of the backward propagation communication time consumption can be calculated by the following formula: ; where represents the amount of input data of the expert network, represents the amount of output data of the expert network, represents the computing node the bandwidth of the external communication link, represents the computing node the latency of the external communication link, represents the number of expert activations set by the mixture of experts model. The backward propagation communication of the computing node includes two communications. Contrary to the forward propagation communication, the first is that each expert network except its own receives the amount of expert output data once, and the second is that each computing node except itself receives the amount of expert input data once.

[0075] Arrange the predicted results of the backward propagation communication time consumption of each computing node in descending order, and substitute them into the in the expected time-consuming calculation formula respectively , , …… , and substitute the activation probability of the corresponding expert network into the , , …… , , so as to calculate the predicted result of the backward propagation communication time consumption of the heterogeneous computing system .

[0076] Then the predicted result of the time consumption of the heterogeneous computing system in one iteration training is: .

[0077] If the expert parallel mode of the heterogeneous computing system is to deploy one or more expert networks on one computing node, that is, when there is a situation where multiple expert networks are deployed on one computing node, based on the computing principle of the computing module provided in the above embodiment, to simplify the solution, determining the time-consuming prediction result of the heterogeneous computing system for performing iterative training in S103 may include: using the maximum value among the forward propagation computing time-consuming prediction results of multiple computing nodes as the forward propagation computing time-consuming prediction result of the heterogeneous computing system; using the maximum value among the backward propagation computing time-consuming prediction results of multiple computing nodes as the backward propagation computing time-consuming prediction result of the heterogeneous computing system; using the maximum value among the forward propagation communication time-consuming prediction results of multiple computing nodes as the forward propagation communication time-consuming prediction result of the heterogeneous computing system; using the maximum value among the backward propagation communication time-consuming prediction results of multiple computing nodes as the backward propagation communication time-consuming prediction result of the heterogeneous computing system; using the sum of the forward propagation computing time-consuming prediction result of the heterogeneous computing system, the backward propagation computing time-consuming prediction result of the heterogeneous computing system, the forward propagation communication time-consuming prediction result of the heterogeneous computing system, and the backward propagation communication time-consuming prediction result of the heterogeneous computing system as the time-consuming prediction result of the heterogeneous computing system for performing iterative training.

[0078] That is to say, for the four stages of forward propagation computing, forward propagation communication, backward propagation computing, and backward propagation communication in one iterative training, the estimated values of the time-consuming prediction results of each computing node in the corresponding stage can be calculated respectively based on the activation probability of the expert network on the computing node, the task parameters of the expert network, and the resource status information of the computing node. Since each stage is executed in parallel by multiple computing nodes, the estimated value of the time-consuming result of the computing node with the largest time-consuming result can be used as the system time-consuming prediction result of this stage.

[0079] In the embodiment of the present invention, the calculation steps of the forward propagation computing time-consuming prediction result of the computing node may include: performing weighted summation on the forward propagation computing amount of the expert network on the computing node with the activation probability of the expert network as the weight to obtain the total forward propagation computing amount of the computing node, and obtaining the first time-consuming prediction result of the computing node by taking the ratio of the total forward propagation computing amount of the computing node to the computing power resource parameter of the computing node; performing weighted summation on the forward propagation memory access data amount of the expert network on the computing node with the activation probability of the expert network as the weight to obtain the memory access data amount of the computing node, and obtaining the second time-consuming prediction result of the computing node by taking the ratio of the memory access data amount of the computing node to the memory bandwidth of the computing node; determining the larger value between the first time-consuming prediction result and the second time-consuming prediction result as the forward propagation computing time-consuming prediction result of the computing node.

[0080] Then the forward propagation computing time-consuming prediction result of the heterogeneous computing system can be calculated by the following formula: ; wherein, represents the activation probability of the expert network on the computing node ; represents the total number of expert networks on the computing node represents the total number of computing nodes, represents the forward propagation computation amount of the expert network, represents the computing power resource parameter of the computing node represents the forward propagation memory access data amount of the expert network on the computing node represents the memory bandwidth of the computing node represents the maximum value calculation. , respectively represent the forward propagation computation time consumption and memory access time consumption of the computing node, and the final forward propagation computation time consumption depends on one of the computing bottleneck and the memory access bottleneck.

[0081] In the embodiment of the present invention, the calculation steps of the reverse propagation computation time consumption prediction result of the computing node may include: weighted summing the reverse propagation computation amounts of the expert networks on the computing node with the activation probability of the expert network as the weight to obtain the total reverse propagation computation amount of the computing node, and obtaining the third time consumption prediction result of the computing node by taking the ratio of the total reverse propagation computation amount of the computing node to the computing power resource parameter of the computing node; weighted summing the reverse propagation memory access data amounts of the expert networks on the computing node with the activation probability of the expert network as the weight to obtain the memory access data amount of the computing node, and obtaining the fourth time consumption prediction result of the computing node by taking the ratio of the memory access data amount of the computing node to the memory bandwidth of the computing node; determining the larger value of the third time consumption prediction result and the fourth time consumption prediction result as the reverse propagation computation time consumption prediction result of the computing node.

[0082] Then, the reverse propagation computation time consumption prediction result of the heterogeneous computing system ; wherein, represents the activation probability of the expert network on the computing node ; represents the total number of expert networks on the computing node represents the total number of computing nodes, represents the reverse propagation computation amount of the expert network Represents the computing power resource parameters of the computing node ; Represents the amount of memory access data for backpropagation of the expert network on the computing node ; Represents the memory bandwidth of the computing node ; Represents the maximum value calculation. 、 Respectively represent the backpropagation calculation time and memory access time of the computing node . The final backpropagation calculation time depends on one of the calculation bottleneck and the memory access bottleneck.

[0083] In the embodiments of the present invention, the calculation steps of the prediction result of the forward propagation communication time of the computing node may include: determining the first amount of expert input data output by the computing node in the forward propagation according to the activation probability of the expert network on the computing node and the amount of expert input data sent to the computing nodes where other activated expert networks are located when the expert network is activated; determining the first amount of expert output data received by the computing node from other computing nodes in the forward propagation according to the activation probability of the expert network on the computing node and the amount of expert output data output by the computing nodes where other activated expert networks are located when the expert network is activated; calculating the prediction result of the forward propagation communication time of the computing node by dividing the sum of the first amount of expert input data and the first amount of expert output data by the bandwidth of the external communication link of the computing node and then adding the communication delay of the external communication link of the computing node.

[0084] Then the prediction result of the forward propagation communication time of the heterogeneous computing system Can be obtained through the following formula communication: ; Wherein, Represents the activation probability of the expert network On the computing node ; Represents the total number of expert networks on the computing node ; Represents the total number of computing nodes Represents the backpropagation calculation amount of the expert network Represents the computing power resource parameters of the computing node ; Represents the amount of memory access data for forward propagation of the expert network on the communication node ; Represents the memory bandwidth of the communication node ; Represents the maximum value communication. Computing node The forward propagation communication includes two communications. The first is to receive the amount of expert input data once from each computing node except itself, and the second is to receive the amount of expert output data once from each expert network except its own expert network.

[0085] In the embodiment of the present invention, the calculation steps of the reverse propagation communication time-consuming prediction result of the computing node may include: determining the second expert input data amount received by the computing node from other computing nodes in the reverse propagation according to the activation probability of the expert network on the computing node and the amount of expert output data output by the computing nodes where other activated expert networks are located when the expert network is activated; determining the second expert input data amount output by the computing node to other computing nodes in the reverse propagation according to the activation probability of the expert network on the computing node and the amount of expert input data sent to the computing nodes where other activated expert networks are located when the expert network is activated; calculating the reverse propagation communication time-consuming prediction result of the computing node by dividing the sum of the second expert input data amount and the second expert output data amount by the bandwidth of the external communication link of the computing node and then adding the communication delay of the external communication link of the computing node.

[0086] Then the reverse propagation communication time-consuming prediction result of the heterogeneous computing system can be obtained through the following formula for communication: ; where represents the activation probability of the expert network on the computing node , represents the total number of expert networks on the computing node , represents the total number of computing nodes, represents the reverse propagation communication volume of the expert network, represents the computing power resource parameter of the communication node , represents the reverse propagation memory access data volume of the expert network on the communication node , represents the memory bandwidth of the communication node , represents the maximum value seeking communication. The reverse propagation communication of the computing node includes two communications. Contrary to the forward propagation communication, the first is to receive the amount of expert output data once from each expert network except its own expert network, and the second is to receive the amount of expert input data once from each computing node except itself.

[0087] Then the time-consuming prediction result of the heterogeneous computing system in one iteration training is: 。

[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0089] The embodiment of the present invention also provides an expert parallel training time-consuming prediction device, including: a task information collection module for determining the expert network of the mixture of experts model deployed on the computing nodes of the heterogeneous computing system; a system information collection module for obtaining the resource status information of the computing nodes; an expert status monitoring module for obtaining the training sample set of the mixture of experts model, inputting the training samples obtained by sampling multiple times from the training sample set into the mixture of experts model respectively, and monitoring the activation state parameters of the expert network; a calculation module for calculating the time-consuming prediction result of the computing node to perform iterative training according to the activation state parameters of the expert network and the resource status information of the computing node; and determining the time-consuming prediction result of the heterogeneous computing system to perform iterative training according to the time-consuming prediction result of the computing node.

[0090] For the description of the features in the corresponding embodiment of the expert parallel training time-consuming prediction device, reference can be made to the relevant description of the corresponding embodiment of the expert parallel training time-consuming prediction method, which will not be elaborated here.

[0091] The embodiment of the present invention also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the expert parallel training time-consuming prediction method.

[0092] The embodiment of the present invention also provides a non-volatile storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the expert parallel training time-consuming prediction method when running.

[0093] In an exemplary embodiment, the above non-volatile storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc and other media that can store computer programs.

[0094] The embodiment of the present invention also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the expert parallel training time-consuming prediction method.

[0095] An embodiment of the present invention also provides another computer program product, including a non-volatile storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the expert parallel training time-consuming prediction method are implemented.

[0096] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0097] The above has introduced in detail the expert parallel training time-consuming prediction method, device, equipment, medium, and product provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A method for predicting the time-consuming of parallel training of experts, characterized in that, Including: Determine the expert network of the mixture-of-experts model deployed on the computing nodes of the heterogeneous computing system; Obtain the training sample set of the mixture-of-experts model, input the training samples obtained by sampling from the training sample set multiple times into the mixture-of-experts model respectively, and monitor the activation state parameters of the expert network; Calculate the predicted result of the time consumption for the computing node to perform iterative training according to the activation state parameters of the expert network and the resource state information of the computing node; Determine the predicted result of the time consumption for the heterogeneous computing system to perform iterative training according to the predicted result of the time consumption of the computing node.

2. The expert parallel training time-consuming prediction method according to claim 1, characterized in that, Monitoring the activation state parameters of the expert network includes: Monitoring the activation state of the expert network in historical iterative training; Determine the activation probability of the expert network in one iterative training according to the activation states of the expert network in multiple rounds of historical iterative training, and use the activation probability as the activation state parameter of the expert network.

3. The expert parallel training time-consuming prediction method according to claim 2, wherein Monitoring the activation state of the expert network in historical iterative training includes: Obtain the gating network of the pre-trained mixture-of-experts model; Sample a training sample from the training sample set and input it into the gating network, and monitor the weights of the expert network output by the gating network; Determine the activation state of the expert network according to the weights of the expert network.

4. The expert parallel training time-consuming prediction method according to claim 3, characterized in that Determining the activation state of the expert network according to the weights of the expert network includes: Obtain the number of activated experts in one calculation of the mixture-of-experts model; Sort the weights of the expert network from large to small, and determine the corresponding number of the expert networks with the largest weights as the activated expert networks according to the number of activated experts, and the other expert networks as the unactivated expert networks.

5. The expert parallel training time consumption prediction method according to claim 2, characterized in that Calculating the predicted result of the time consumption for the computing node to perform iterative training according to the activation state parameters of the expert network and the resource state information of the computing node includes: Calculate the predicted result of the forward propagation calculation time consumption of the computing node and the predicted result of the backward propagation calculation time consumption of the computing node according to the activation probability of the expert network, the computing power resource parameters of the computing node, and the memory bandwidth of the computing node; Calculate the predicted result of the forward propagation communication time consumption of the computing node and the predicted result of the backward propagation communication time consumption of the computing node according to the activation probability of the expert network and the network resource parameters of the computing node.

6. The method for predicting the time consumption of parallel training of experts according to claim 5, wherein Calculating the predicted result of the forward propagation calculation time consumption of the computing node according to the activation probability of the expert network, the computing power resource parameters of the computing node, and the memory bandwidth of the computing node includes: Calculate the first predicted result of the time consumption of the computing node according to the activation probability of the expert network on the computing node, the forward propagation calculation amount of the expert network, and the computing power resource parameters of the computing node; Calculate the second predicted result of the time consumption of the computing node according to the activation probability of the expert network on the computing node, the forward propagation memory access data volume of the expert network, and the memory bandwidth of the computing node. Determine the larger value between the first time-consuming prediction result and the second time-consuming prediction result as the forward propagation calculation time-consuming prediction result of the computing node.

7. The method for predicting the time consumption of expert parallel training according to claim 5, wherein Calculate the backward propagation calculation time-consuming prediction result of the computing node based on the activation probability of the expert network, the computing power resource parameters of the computing node, and the memory bandwidth of the computing node, including: Calculate the third time-consuming prediction result of the computing node according to the activation probability of the expert network on the computing node, the backward propagation calculation amount of the expert network, and the computing power resource parameters of the computing node; Calculate the fourth time-consuming prediction result of the computing node according to the activation probability of the expert network on the computing node, the backward propagation memory access data amount of the expert network, and the memory bandwidth of the computing node; Determine the larger value between the third time-consuming prediction result and the fourth time-consuming prediction result as the backward propagation calculation time-consuming prediction result of the computing node.

8. The expert parallel training time-consuming prediction method according to claim 5, wherein Calculate the forward propagation communication time-consuming prediction result of the computing node according to the activation probability of the expert network and the network resource parameters of the computing node, including: Determine the first expert input data amount output by the computing node to other computing nodes in forward propagation according to the activation probability of the expert network on the computing node and the amount of expert input data sent by the computing node where the expert network is activated to the computing nodes where other activated expert networks are located; Determine the first expert output data amount input by the computing node from other computing nodes in forward propagation according to the activation probability of the expert network on the computing node and the amount of expert output data output by the computing nodes where other activated expert networks are located when the expert network is activated; Calculate the forward propagation communication time-consuming prediction result of the computing node by dividing the sum of the first expert input data amount and the first expert output data amount by the bandwidth of the external communication link of the computing node and then adding the communication delay of the external communication link of the computing node.

9. The method for predicting the time consumption of expert parallel training according to claim 5, wherein Calculate the backward propagation communication time-consuming prediction result of the computing node according to the activation probability of the expert network and the network resource parameters of the computing node, including: Determine the second expert output data amount received by the computing node from other computing nodes in backward propagation according to the activation probability of the expert network on the computing node and the amount of expert output data output by the computing nodes where other activated expert networks are located when the expert network is activated; Determine the second expert input data amount output by the computing node to other computing nodes in backward propagation according to the activation probability of the expert network on the computing node and the amount of expert input data sent by the computing node where the expert network is activated to the computing nodes where other activated expert networks are located; Calculate the predicted result of the backpropagation communication time consumption of the computing node by dividing the sum of the second expert input data volume and the second expert output data volume by the bandwidth of the external communication link of the computing node and then adding the communication delay of the external communication link of the computing node.

10. The expert parallel training time-consuming prediction method according to claim 5, wherein Deploy one of the expert networks on one of the computing nodes; Determine the predicted result of the time consumption for the heterogeneous computing system to perform iterative training according to the predicted result of the time consumption of the computing node, including: Calculate the predicted result of the forward propagation calculation time consumption of the heterogeneous computing system according to the predicted results of the forward propagation calculation time consumption of multiple computing nodes and the activation probability of the expert network; Calculate the predicted result of the backpropagation calculation time consumption of the heterogeneous computing system according to the predicted results of the backpropagation calculation time consumption of multiple computing nodes and the activation probability of the expert network; Calculate the predicted result of the forward propagation communication time consumption of the heterogeneous computing system according to the predicted results of the forward propagation communication time consumption of multiple computing nodes and the activation probability of the expert network; Calculate the predicted result of the backpropagation communication time consumption of the heterogeneous computing system according to the predicted results of the backpropagation communication time consumption of multiple computing nodes and the activation probability of the expert network; Use the sum of the predicted result of the forward propagation calculation time consumption of the heterogeneous computing system, the predicted result of the backpropagation calculation time consumption of the heterogeneous computing system, the predicted result of the forward propagation communication time consumption of the heterogeneous computing system, and the predicted result of the backpropagation communication time consumption of the heterogeneous computing system as the predicted result of the time consumption for the heterogeneous computing system to perform iterative training.

11. The method for predicting the time consumption of expert parallel training according to claim 10, wherein Calculate the predicted result of the forward propagation calculation time consumption of the heterogeneous computing system according to the predicted results of the forward propagation calculation time consumption of multiple computing nodes and the activation probability of the expert network, including: Arrange the predicted results of the forward propagation calculation time consumption of the computing nodes in descending order and substitute them into the respective elements of the expected time consumption formula, substitute the corresponding activation probabilities of the expert network into the respective elements of the expected time consumption formula, and calculate the predicted result of the forward propagation calculation time consumption of the heterogeneous computing system; Among them, the expected time consumption formula calculates the expected time consumption based on the sum of multiple elements. Corresponding to the descending order of the input predicted results of the time consumption, the elements of the expected time consumption formula are the product of the input predicted result of the time consumption, the activation probability of the expert network corresponding to the input predicted result of the time consumption, and the probability that the expert network corresponding to the input predicted result of the time consumption is not activated.

12. The expert parallel training time-consuming prediction method according to claim 10, characterized in that Calculate the predicted result of the backpropagation calculation time consumption of the heterogeneous computing system according to the predicted results of the backpropagation calculation time consumption of multiple computing nodes and the activation probability of the expert network, including: Arrange the predicted results of the backpropagation calculation time consumption of the computing nodes in descending order and substitute them into the respective elements of the expected time consumption formula, substitute the corresponding activation probabilities of the expert network into the respective elements of the expected time consumption formula, and calculate the predicted result of the backpropagation calculation time consumption of the heterogeneous computing system; Among them, the formula for calculating the expected time consumption is used to calculate the expected time consumption based on the sum of multiple elements. For the order of the input time consumption prediction results from large to small, the elements of the formula for calculating the expected time consumption are the input time consumption prediction results, the activation probability of the expert network corresponding to the input time consumption prediction results, and the product of the probability that the expert network corresponding to the input time consumption prediction results is not activated and is greater than the input time consumption prediction results.

13. The expert parallel training time-consuming prediction method according to claim 10, characterized in that Based on the forward propagation communication time consumption prediction results of multiple calculation nodes and the activation probability of the expert network, calculating the forward propagation communication time consumption prediction result of the heterogeneous computing system includes: After arranging the forward propagation communication time consumption prediction results of the calculation nodes in descending order, substituting them into the respective elements of the formula for calculating the expected time consumption, and substituting the corresponding activation probability of the expert network into the respective elements of the formula for calculating the expected time consumption, to calculate the forward propagation communication time consumption prediction result of the heterogeneous computing system; Among them, the formula for calculating the expected time consumption is used to calculate the expected time consumption based on the sum of multiple elements. For the order of the input time consumption prediction results from large to small, the elements of the formula for calculating the expected time consumption are the input time consumption prediction results, the activation probability of the expert network corresponding to the input time consumption prediction results, and the product of the probability that the expert network corresponding to the input time consumption prediction results is not activated and is greater than the input time consumption prediction results.

14. The expert parallel training time-consuming prediction method according to claim 10, wherein Based on the backward propagation communication time consumption prediction results of multiple calculation nodes and the activation probability of the expert network, calculating the backward propagation communication time consumption prediction result of the heterogeneous computing system includes: After arranging the backward propagation communication time consumption prediction results of the calculation nodes in descending order, substituting them into the respective elements of the formula for calculating the expected time consumption, and substituting the corresponding activation probability of the expert network into the respective elements of the formula for calculating the expected time consumption, to calculate the backward propagation communication time consumption prediction result of the heterogeneous computing system; Among them, the formula for calculating the expected time consumption is used to calculate the expected time consumption based on the sum of multiple elements. For the order of the input time consumption prediction results from large to small, the elements of the formula for calculating the expected time consumption are the input time consumption prediction results, the activation probability of the expert network corresponding to the input time consumption prediction results, and the product of the probability that the expert network corresponding to the input time consumption prediction results is not activated and is greater than the input time consumption prediction results.

15. The method for predicting the time consumption of parallel training of experts according to claim 5, characterized in that, One calculation node deploys one or more of the expert networks.

16. The expert parallel training time-consuming prediction method according to claim 15, wherein Based on the time consumption prediction result of the calculation node to determine the time consumption prediction result of the heterogeneous computing system for performing iterative training, including: Taking the maximum value among the forward propagation calculation time consumption prediction results of multiple calculation nodes as the forward propagation calculation time consumption prediction result of the heterogeneous computing system; Taking the maximum value among the backward propagation calculation time consumption prediction results of multiple calculation nodes as the backward propagation calculation time consumption prediction result of the heterogeneous computing system; Taking the maximum value among the forward propagation communication time consumption prediction results of multiple calculation nodes as the forward propagation communication time consumption prediction result of the heterogeneous computing system; Use the maximum value among the backpropagation communication time prediction results of multiple said computing nodes as the backpropagation communication time prediction result of the heterogeneous computing system; Use the sum of the forward propagation calculation time prediction result of the heterogeneous computing system, the backpropagation calculation time prediction result of the heterogeneous computing system, the forward propagation communication time prediction result of the heterogeneous computing system, and the backpropagation communication time prediction result of the heterogeneous computing system as the time prediction result for the heterogeneous computing system to perform iterative training.

17. An expert parallel training time-consuming prediction device, characterized in that, Comprising: A task information collection module, configured to determine the expert network of the mixture-of-experts model deployed on the computing nodes of the heterogeneous computing system; A system information collection module, configured to obtain the resource status information of the computing nodes; An expert status monitoring module, configured to obtain the training sample set of the mixture-of-experts model, input the training samples obtained by sampling multiple times from the training sample set into the mixture-of-experts model respectively, and monitor the activation state parameters of the expert network; A calculation module, configured to calculate the time prediction result for the computing nodes to perform iterative training according to the activation state parameters of the expert network and the resource status information of the computing nodes; Determine the time prediction result for the heterogeneous computing system to perform iterative training according to the time prediction results of the computing nodes.

18. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the expert parallel training time prediction method according to any one of claims 1 to 16 when executing the computer program.

19. A non-volatile storage medium, characterized in that, A computer program is stored in the non-volatile storage medium, wherein the computer program implements the steps of the expert parallel training time prediction method according to any one of claims 1 to 16 when executed by a processor.

20. A computer program product comprising a computer program, characterized in that, The computer program implements the steps of the expert parallel training time prediction method according to any one of claims 1 to 16 when executed by a processor.

Citation Information

Patent Citations

  • Training time length prediction method and device, multivariate heterogeneous computing equipment and medium

    CN116244159A

  • Heterogeneous computing platform and task simulation and time consumption prediction method, device and equipment thereof

    CN117971630A

  • Heterogeneous computing system and training time consumption prediction method, equipment, medium and product thereof

    CN119204360A

  • Hybrid expert model training method and device, computer equipment, readable storage medium and program product

    CN119830952A

  • Reliable reasoning scheduling method based on edge hybrid expert large model

    CN119903923A