Model deployment method and device, equipment, storage medium and product

By acquiring and analyzing the performance and task information of heterogeneous computing nodes in a multi-heterogeneous computing system, and performing model compression and deployment, the problem of low hardware resource utilization is solved, and time consumption balance and inference task acceleration are achieved.

CN120745844AActive Publication Date: 2025-10-03SHANDONG HAILIANG INFORMATION TECH RES INST
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511224090.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-10-03
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

In a multi-heterogeneous computing system, the performance and communication capability differences of heterogeneous computing nodes lead to low hardware resource utilization, and the parallel computing of expert models needs to wait for the slowest model to be calculated.

Method used

By obtaining the distributed reasoning task information of the hybrid expert model and the performance information of the heterogeneous computing system, the total time consumed by each expert model on the heterogeneous computing node is determined. Based on the load balancing principle, the total time consumption, communication time consumption and computing time consumption are analyzed, and the model is compressed to adjust the deployment to ensure the time consumption balance of each expert model.

Benefits of technology

This improves the utilization of hardware resources, avoids long waits for the slowest model to complete calculation, and accelerates expert neural network inference tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745844A_ABST
    Figure CN120745844A_ABST
Patent Text Reader

Abstract

The invention discloses a model deployment method and device, equipment, a storage medium and a product, and relates to the technical field of multivariate heterogeneous computing systems. And determining the total time consumption of each expert model for executing the reasoning task on the corresponding heterogeneous computing node. Based on a load balancing principle, analyzing total time consumption, communication time consumption and calculation time consumption of all expert models to determine a compression ratio; and performing iterative compression on each expert model according to a model compression strategy to obtain each compressed expert model meeting an error requirement and a compression rate requirement. And deploying the compressed expert models at the corresponding heterogeneous computing nodes. The expert model is compressed, and the compression degree of the expert model is determined based on the compression rate, so that the calculation time consumption of different heterogeneous computing power in the expert operation layer is balanced as much as possible, and the utilization rate of hardware resources is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of multi-heterogeneous computing systems, and in particular to a model deployment method, apparatus, device, storage medium, and product. Background Art

[0002] In recent years, the concept of multi-heterogeneous computing systems has gradually emerged. In such systems, heterogeneous computing power (i.e., heterogeneous compute nodes) with varying computational performance are integrated into a single distributed computing environment and collaborate to complete distributed inference tasks using a mixture of experts (MEXs). Mixture of Experts (MoE) is a model parallelization method that improves training and inference efficiency by partitioning large models into multiple expert models and deploying each expert model on different heterogeneous compute nodes.

[0003] Due to differences in performance and communication capabilities among heterogeneous computing nodes in a heterogeneous computing system, expert models on different nodes may end their computation tasks at different times. Parallel computation of expert models requires waiting for the slowest expert model to complete, resulting in low hardware resource utilization.

[0004] It can be seen that how to improve the utilization of hardware resources is a problem that those skilled in the art need to solve. Summary of the Invention

[0005] The present application provides a model deployment method, apparatus, device, storage medium and product to at least solve the problem of low utilization of hardware resources in related technologies.

[0006] This application provides a model deployment method, including: Obtaining distributed reasoning task information of a hybrid expert model and performance information of a heterogeneous computing system; wherein the heterogeneous computing system includes multiple heterogeneous computing nodes; and the hybrid expert model includes multiple expert models; Based on the performance information of the heterogeneous computing system and the distributed reasoning task information, the total time consumed by each expert model to execute the reasoning task on its corresponding heterogeneous computing node is determined; the total time consumed includes communication time and computing time; Based on the load balancing principle, the total time consumption, communication time consumption and calculation time consumption of all expert models are analyzed to determine the compression ratio; Iteratively compress each expert model according to the model compression strategy to obtain each compressed expert model that meets the error requirements and the compression ratio requirements; According to the corresponding relationship between expert models and heterogeneous computing nodes, each compressed expert model is deployed on the corresponding heterogeneous computing node.

[0007] The present application also provides a model deployment device, comprising an acquisition unit, a time consumption determination unit, a multiplication rate determination unit, a compression unit, and a deployment unit; An acquisition unit, configured to acquire distributed reasoning task information of a hybrid expert model and performance information of a heterogeneous computing system; wherein the heterogeneous computing system includes a plurality of heterogeneous computing nodes; and the hybrid expert model includes a plurality of expert models; A time consumption determination unit is used to determine the total time consumption of each expert model for executing the reasoning task on its corresponding heterogeneous computing node based on the performance information of the heterogeneous computing system and the distributed reasoning task information; wherein the total time consumption includes communication time consumption and computing time consumption; A compression ratio determination unit is used to analyze the total time consumption, communication time consumption, and calculation time consumption of all expert models based on the load balancing principle to determine the compression ratio; A compression unit, configured to iteratively compress each expert model according to a model compression strategy to obtain compressed expert models that meet error requirements and compression ratio requirements; The deployment unit is used to deploy each compressed expert model on the corresponding heterogeneous computing node according to the corresponding relationship between the expert model and the heterogeneous computing node.

[0008] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned model deployment methods when executing the computer program.

[0009] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned model deployment methods are implemented.

[0010] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned model deployment methods when executed by a processor.

[0011] The present application obtains the distributed reasoning task information of the hybrid expert model and the performance information of the heterogeneous computing system; wherein, the heterogeneous computing system includes multiple heterogeneous computing nodes; and the hybrid expert model includes multiple expert models. The distributed reasoning task information includes the architecture information of the expert model. Different architectures will result in different time consumption for executing reasoning tasks. The performance of the heterogeneous computing nodes will also affect the time consumption of reasoning tasks. Therefore, based on the performance information of the heterogeneous computing system and the distributed reasoning task information, the total time consumption of each expert model to execute the reasoning task on its corresponding heterogeneous computing node can be determined; wherein the total time consumption includes communication time consumption and computation time consumption. In order to make the time consumption of different expert models to execute reasoning tasks as balanced as possible, the total time consumption, communication time consumption, and computation time consumption of all expert models can be analyzed based on the load balancing principle to determine the compression ratio; each expert model is iteratively compressed according to the model compression strategy to obtain each compressed expert model that meets the error requirements and the compression ratio requirements. According to the correspondence between the expert model and the heterogeneous computing node, each compressed expert model is deployed on the corresponding heterogeneous computing node. In this application, the expert models deployed on different heterogeneous computing nodes are adjusted by model compression, and the compression degree of the expert model is determined based on the compression ratio, so that the calculation time of different heterogeneous computing powers in the expert operation layer is as balanced as possible, which accelerates the expert neural network reasoning task and effectively avoids the need to wait for a long time for the slowest expert model to complete calculation, thereby improving the utilization rate of hardware resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 A schematic diagram of distributed reasoning performed on a hybrid expert model on a heterogeneous computing system; Figure 2 A flowchart of a model deployment method provided in an embodiment of the present application; Figure 3 A schematic diagram of the time consumption of executing a computing task by multiple compressed expert models provided in an embodiment of the present application; Figure 4 A flowchart of a method for compressing an expert model provided in an embodiment of the present application; Figure 5 A schematic diagram of the structure of a model deployment device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0014] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0015] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0016] In the MoE architecture, the hybrid expert model is split into multiple expert models, each responsible for a portion of the task. The expert model is typically a small feedforward neural network (FNN). A heterogeneous computing system consists of multiple heterogeneous compute nodes, which can be accelerator cards from different manufacturers and with varying computing performance. Communication between heterogeneous compute nodes can be intra-server or inter-server.

[0017] Figure 1 A schematic diagram of distributed reasoning performed by a hybrid expert model on a heterogeneous computing system. The hybrid expert model contains multiple expert models. Figure 1 In this paper, four expert models are taken as examples, namely expert model 1, expert model 2, expert model 3 and expert model 4. Each expert model has its corresponding input and output. Figure 1 Different labels are used to distinguish the input and output of different expert models. Each expert model is deployed on a corresponding heterogeneous computing node. A heterogeneous computing node is a heterogeneous computing power. Figure 2 In this example, four heterogeneous computing forces are used: Heterogeneous Computing Force 1, Heterogeneous Computing Force 2, Heterogeneous Computing Force 3, and Heterogeneous Computing Force 4. Multiple expert models share a large model space, but each time an inference task is executed, the input data only activates a subset of the expert models for calculation. The expert selection gate function assigns the corresponding task to the activated expert model. Figure 1 The example above uses activated expert models 1, 2, and 3, while inactivating expert model 4. By distributing different expert models to heterogeneous computing nodes, the MoE architecture enables parallel computing across these expert models. This parallelization significantly improves inference speed and model scalability.

[0018] Due to the differences in the performance of heterogeneous computing power and communication capabilities in heterogeneous computing systems, the expert models under different heterogeneous computing nodes have different completion times for executing inference tasks. The expert parallel computing process needs to wait for the slowest expert neural network to complete the calculation, resulting in low utilization of hardware resources.

[0019] This application considers that if the time consumption of all expert models can be made consistent or similar, the load of each heterogeneous computing power can be guaranteed to be as balanced as possible, thereby improving the utilization rate of hardware resources. Therefore, the embodiments of this application provide a model deployment method, device, equipment, storage medium and product, which adjust the expert models deployed on different heterogeneous computing nodes by model compression, and determine the compression level of the expert model based on the compression ratio, so that the calculation time consumption of different heterogeneous computing powers at the expert operation layer is as balanced as possible, accelerating the expert neural network reasoning task, effectively avoiding the need to wait for a long time for the slowest expert model to complete calculation, and thus improving the utilization rate of hardware resources.

[0020] When performing model compression, the target expert model after each compression can be fine-tuned. Fine-tuning refers to performing small-scale training on the pre-trained model for specific task objectives (downstream tasks) and task data (downstream data), making minor adjustments to the pre-trained model parameters, and ultimately obtaining a model that is adapted to the specific task and data.

[0021] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0022] Figure 2 A flowchart of a model deployment method provided in an embodiment of the present application includes: S201: Obtaining distributed reasoning task information of the hybrid expert model and performance information of the heterogeneous computing system.

[0023] The heterogeneous computing system includes multiple heterogeneous computing nodes, and a heterogeneous computing node can be regarded as a heterogeneous computing power. The hybrid expert model includes multiple expert models.

[0024] Distributed reasoning task information includes the architecture information of the expert models, such as the network structure, input and output sizes of each expert model, the number of expert models activated in each round of reasoning, and the correspondence between the expert models and heterogeneous computing nodes.

[0025] The network structure of the expert model may include the number of neural layers contained in the expert model, the type of each layer, the input and output structure of each layer, etc. Currently, most expert neural networks are fully connected layers.

[0026] Based on the correspondence between expert models and heterogeneous computing nodes, it is possible to determine on which heterogeneous computing node each expert model is deployed. This correspondence can be directly obtained through the task requirements input by the user.

[0027] The performance information of a heterogeneous computing system characterizes the network communication capabilities of the heterogeneous computing power within the system. This information may include uplink and downlink bandwidth and latency information for each heterogeneous computing power's external links. Uplink and downlink bandwidth include both uplink and downlink bandwidths. Latency information includes both uplink and downlink latency.

[0028] In actual applications, server performance testing tools (benchmark) can be used to obtain performance information of each heterogeneous computing node.

[0029] S202: Determine the total time consumed by each expert model to execute the reasoning task on its corresponding heterogeneous computing node based on the performance information of the heterogeneous computing system and the distributed reasoning task information.

[0030] The total time consumption may include communication time consumption and calculation time consumption.

[0031] In this embodiment of the present application, the computational time required for each expert model to execute an inference task on its corresponding heterogeneous computing node can be obtained. The communication time is determined based on the input and output sizes included in the distributed inference task information, the uplink bandwidth, downlink bandwidth, and latency information included in the performance information, and the number of expert models activated in each inference round. The sum of the computational time and communication time of each expert model is used as the total time required for each expert model.

[0032] To determine communication time, the downlink data volume can be determined based on the input size and the number of expert models activated in each inference round. The ratio of the downlink data volume to the downlink bandwidth is used as the downlink communication time. The uplink data volume can be determined based on the output size and the number of expert models activated in each inference round. The ratio of the uplink data volume to the uplink bandwidth is used as the uplink communication time. The downlink communication time, the uplink communication time, and the uplink and downlink delays included in the performance information are summed to obtain the communication time.

[0033] In the specific implementation, the total time consumption can be calculated according to the following formula: ; in, represents the total time consumed by the i-th expert model, represents the computation time of the i-th expert model, , represents the communication time of the i-th expert model, , M irepresents the i-th expert model, A represents the number of expert models activated in each round of reasoning, and D in Indicates the input size, D out Indicates the output size, Indicates the uplink bandwidth. Indicates the downlink bandwidth, L i Indicates the delay. Since the uplink delay and downlink delay are the same, the delay in the communication time is 2L i .

[0034] S203: Based on the load balancing principle, the total time consumption, communication time consumption and calculation time consumption of all expert models are analyzed to determine the compression ratio.

[0035] Because the correspondence between expert models and heterogeneous computing nodes is fixed, the communication time of each expert model is difficult to change when heterogeneous computing power is specified. Therefore, in the embodiments of this application, model compression is used to change the computation time, so that each expert model can complete the inference task as simultaneously as possible. By obtaining the computation time and communication time of each expert model, the sum of the computation time and communication time is used as the total time, providing a reference for determining the compression ratio.

[0036] To evaluate the degree of model compression, we can first determine the compression ratio. The compression ratio can be regarded as the ratio of the computational effort of the expert model after compression to the computational effort of the expert model before compression.

[0037] In an embodiment of the present application, the total time consumption of all expert models can be compared with the communication time consumption to determine the optimal load balancing time consumption; based on the total time consumption of each expert model, the target expert model with the largest total time consumption is selected. If the total time consumption of the target expert model is less than or equal to the optimal load balancing time consumption, the target expert model can be directly output. If the total time consumption of the target expert model is greater than the optimal load balancing time consumption, the compression ratio is determined based on the optimal load balancing time consumption and the calculation time and total time consumption of the target expert model.

[0038] To achieve a balanced total time consumption across all expert models, the optimal load balancing time can be determined by minimizing the total time consumption of all expert models. Considering that some expert models may have long communication times, the minimum communication time consumption of all expert models can be considered when determining the optimal load balancing time.

[0039] Therefore, in an embodiment of the present application, the minimum total time consumption can be screened out based on the total time consumption of all expert models; the maximum communication time consumption can be screened out based on the communication time consumption of all expert models; and the maximum value between the minimum total time consumption and the maximum communication time consumption can be selected as the optimal load balancing time consumption.

[0040] In specific implementation, the optimal load balancing time can be determined according to the following formula: ; Among them, O represents the optimal load balancing time.

[0041] After determining the optimal load balancing time, the difference between the total time of the target expert model and the optimal load balancing time can be divided by the calculation time of the target expert model to obtain the calculation amount compression ratio; based on the sum of the calculation amount compression ratio and the compression ratio being one, the compression ratio is determined.

[0042] In a specific implementation, the compression ratio can be determined according to the following formula: ; Among them, q represents the compression ratio, C i Indicates the computation time of the i-th expert model.

[0043] S204: Iteratively compressing each expert model according to the model compression strategy to obtain compressed expert models that meet the error requirements and the compression ratio requirements.

[0044] The model compression strategy is used to indicate the way to compress the model, which may include the compression step size and which neurons to compress. In the embodiment of the present application, model compression can be achieved by using a model pruning method.

[0045] In practical applications, we can evaluate the importance of each neuron in the expert model to the overall output of the model, and prune the neurons with low importance according to the compression step size to obtain the expert model after a single compression.

[0046] In order to avoid the impact of excessive compression on the accuracy of the model, the compression step size can be set to a smaller value. Through multiple compressions, the expert model after compression that meets the error requirements and the compression ratio requirements is finally determined.

[0047] Figure 3 A schematic diagram of the time consumption of multiple compressed expert models performing computing tasks provided in an embodiment of the present application is provided. Figure 3 This example uses three expert models: Expert Model 1, Expert Model 2, and Expert Model 3. These models are deployed on different heterogeneous computing nodes. The time it takes for an expert model to perform a computational task primarily consists of communication time and computation time. Communication time can include both the time spent receiving data and the time spent sending data. The time spent receiving data can be referred to as the communication receiving time, and the time spent sending data can be referred to as the communication sending time. By compressing the expert models, the computation time of each expert model can be adjusted, making the total time of the three expert models nearly consistent.

[0048] S205: deploying each compressed expert model on a corresponding heterogeneous computing node according to the corresponding relationship between the expert model and the heterogeneous computing node.

[0049] Based on the correspondence between expert models and heterogeneous computing nodes, the deployment location of each expert model can be determined, that is, on which heterogeneous computing node each expert model is deployed. After the expert model is compressed, each compressed expert model can be deployed on the corresponding heterogeneous computing node based on the correspondence.

[0050] In the embodiment of the present application, by comprehensively considering the performance of the heterogeneous computing system, the architecture of the hybrid expert model, and the total time taken by each expert model to execute the reasoning task on its corresponding heterogeneous computing node, the expert model is compressed, so that the computing time of different heterogeneous computing powers at the expert operation layer is balanced as much as possible, thereby accelerating the expert neural network reasoning task, achieving the effect of completing the reasoning request faster with the same computing cost, and improving the utilization rate of the hardware resources of the heterogeneous computing system.

[0051] After deploying the compressed expert models on the corresponding heterogeneous computing nodes, they can be used to perform text answering tasks. For example, a question can be input to the expert model, and the expert model can output the corresponding answer.

[0052] As can be seen from the above technical solution, distributed inference task information of a hybrid expert model and performance information of a heterogeneous computing system are obtained. The heterogeneous computing system includes multiple heterogeneous computing nodes, and the hybrid expert model includes multiple expert models. The distributed inference task information includes the architecture information of the expert model. Different architectures will result in different execution times for inference tasks. The performance of heterogeneous computing nodes will also affect the inference task time. Therefore, based on the performance information of the heterogeneous computing system and the distributed inference task information, the total execution time of each expert model on its corresponding heterogeneous computing node can be determined. The total execution time includes communication time and computation time. To ensure that the execution time of different expert models is as balanced as possible, the total execution time, communication time, and computation time of all expert models can be analyzed based on the principle of load balancing to determine a compression ratio. Each expert model is then iteratively compressed according to the model compression strategy to obtain compressed expert models that meet error and compression ratio requirements. Based on the correspondence between the expert model and the heterogeneous computing node, each compressed expert model is deployed on the corresponding heterogeneous computing node. In this application, the expert models deployed on different heterogeneous computing nodes are adjusted by model compression, and the compression degree of the expert model is determined based on the compression ratio, so that the calculation time of different heterogeneous computing powers in the expert operation layer is as balanced as possible, which accelerates the expert neural network reasoning task and effectively avoids the need to wait for a long time for the slowest expert model to complete calculation, thereby improving the utilization rate of hardware resources.

[0053] Figure 4 A flowchart of a method for compressing an expert model provided in an embodiment of the present application, the method comprising: S401: Pruning multiple neurons with the lowest importance scores in the target expert model based on the compression step size to obtain a compressed target expert model.

[0054] Each time model compression is performed, the expert model with the largest total time consumption can be selected from all uncompressed expert models as the target expert model.

[0055] When performing model compression, in order to prevent the shape of the input and output from being changed and affecting the performance of the expert model, the importance of the output neurons in the fully connected layers of the expert model except the first and last layers is evaluated.

[0056] In an embodiment of the present application, the L1 norm of the weight column can be used as an importance score. The higher the importance score, the more important the neuron. The L1 norm of the weight of each neuron in the target expert model is used as the importance score of each neuron; in ascending order of the importance score of each neuron, multiple neurons to be pruned with the lowest importance score consistent with the number of compression steps are determined. Each neuron to be pruned is pruned from the target expert model to obtain a compressed target expert model.

[0057] For example, for the neural network M i Except for the output neurons in the first and last fully connected layers (to prevent changing the shape of input and output and affecting model performance), the L1 norm of their weight columns is calculated as the importance score: ; in, represents the L1 norm, W represents the weight matrix of the fully connected layer of the target expert model, W j represents the weight of the j-th output neuron, represents the importance score of the j-th output neuron.

[0058] In the embodiment of the present application, the product of the set search step length and the number of pruning times can be used as the compression step length. The search step length can be represented by K, the number of pruning times is represented by u, and the compression step length is u*K.

[0059] After determining the importance score of each neuron, u*K neurons can be pruned to obtain a single compressed expert model. i Represents the expert model before compression, using Represents the expert model after single compression, using m i It represents the compressed expert model that meets the error requirements and the compression ratio requirements.

[0060] S402: Determine the output error of the compressed target expert model according to the deviation between the output vector of the target expert model and the output vector of the compressed target expert model.

[0061] In the embodiment of the present application, the distillation loss can be used to determine the Perform fine-tuning training, where the distillation loss is: ; in, represents the distillation loss value, represents the L2 norm, represents the distillation loss function, represents the output vector of the expert model after single compression, represents the output vector of the expert model before compression, In order to prevent the denominator from being zero, Can be set to 0.000001.

[0062] In order to evaluate the output error of the compressed target expert model, in practical applications, the distillation loss value can be used as the output error.

[0063] S403: The ratio of the calculation amount of the compressed target expert model to the calculation amount of the target expert model is used as the calculation amount change value.

[0064] For ease of description, we can use Indicates the computational cost of the compressed target expert model, using Represents the computational effort of the target expert model before compression.

[0065] S404: When the output error of the compressed target expert model is less than the error threshold and the computational change of the compressed target expert model is greater than the compression ratio, adjust the compression step size and use the compressed target expert model as the latest target expert model.

[0066] The error threshold can be flexibly set based on actual needs. When the accuracy of the expert model is required to be high, the error threshold can be set smaller; when the accuracy of the expert model is not required to be very high, the error threshold can be set larger. In actual applications, the default error threshold can be set to 0.1%.

[0067] In the specific implementation, when the output error of the compressed target expert model is less than the error threshold, if , indicating that the compressed target expert model does not meet the compression requirements. At this time, the compressed target expert model can be further compressed.

[0068] Each time the pruning operation of the expert model is performed, the number of pruning times is increased by one; in the initial state, the value of the pruning number is one; the product of the set search step size and the current value of the pruning number is used as the adjusted compression step size.

[0069] When the compressed target expert model needs to be further compressed, the compression step size can be adjusted, and the compressed target expert model can be used as the latest target expert model. Then, the operation steps of pruning multiple neurons with the lowest importance scores in the target expert model based on the compression step size can be returned to S501 to obtain the compressed target expert model.

[0070] S405: When the output error of the compressed target expert model is less than the error threshold and the calculation amount change value of the compressed target expert model is less than or equal to the compression ratio, output the compressed target expert model.

[0071] In the specific implementation, when the output error of the compressed target expert model is less than the error threshold, if , indicating that the compressed target expert model meets the compression requirements. At this time, the compression of the target expert model can be ended and the compressed target expert model can be directly output.

[0072] Considering that in practical applications, there may be cases where the output error of the compressed target expert model is greater than or equal to the error threshold, the following will describe the model compression process in this case.

[0073] When the output error of the compressed target expert model is greater than or equal to the error threshold, it means that the current compressed target expert model is unavailable. At this time, the total time consumption of the target expert model can be adjusted according to the computational amount of the target expert model and the computational amount of the compressed target expert model.

[0074] In an embodiment of the present application, the ratio of the computational amount of the compressed target expert model to the computational amount of the target expert model can be multiplied by the computational time of the target expert model to obtain the adjusted computational time; the sum of the adjusted computational time and the communication time of the target expert model is used as the adjusted total time of the target expert model.

[0075] For example, the ratio of the computational cost of the target expert model after compression to the computational cost of the target expert model before compression is multiplied by C i Estimate the computation time, and the adjusted computation time is recorded as c i Using c i Modify and update the current total time: .

[0076] When the change in the computational complexity of the compressed target expert model is less than or equal to the compression ratio, the operation step of selecting the target expert model with the largest total time consumption can be returned to perform the compression operation on the newly selected target expert model.

[0077] When the change in the computational amount of the compressed target expert model is greater than the compression ratio, it indicates that the optimal load balancing time setting is unreasonable. At this time, the adjusted total time of the target expert model can be used as the optimal load balancing time, and the operation steps of selecting the target expert model with the largest total time can be returned to perform compression operations on the newly selected target expert model.

[0078] In an embodiment of the present application, when the output error of the compressed target expert model is large and the change in the amount of calculation is greater than the compression ratio, the optimal load balancing time is adjusted to ensure the smooth execution of the expert model compression process, thereby improving the inference speed of the expert model while ensuring the accuracy of the expert model.

[0079] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0080] Figure 5 A schematic diagram of the structure of a model deployment device provided in an embodiment of the present application, comprising an acquisition unit 51, a time consumption determination unit 52, a magnification determination unit 53, a compression unit 54, and a deployment unit 55; An acquisition unit 51 is configured to acquire distributed reasoning task information of a hybrid expert model and performance information of a heterogeneous computing system; wherein the heterogeneous computing system includes a plurality of heterogeneous computing nodes; and the hybrid expert model includes a plurality of expert models; The time consumption determination unit 52 is used to determine the total time consumption of each expert model to execute the reasoning task on its corresponding heterogeneous computing node based on the performance information of the heterogeneous computing system and the distributed reasoning task information; wherein the total time consumption includes communication time consumption and computing time consumption; A compression ratio determination unit 53 is configured to analyze the total time consumption, communication time consumption, and computation time consumption of all expert models based on a load balancing principle to determine a compression ratio; The compression unit 54 is used to iteratively compress each expert model according to the model compression strategy to obtain each compressed expert model that meets the error requirement and the compression ratio requirement; The deployment unit 55 is configured to deploy each compressed expert model on a corresponding heterogeneous computing node according to the corresponding relationship between the expert model and the heterogeneous computing node.

[0081] In some embodiments, the distributed reasoning task information includes the network structure of each expert model, input and output sizes, the number of expert models activated in each round of reasoning, and the correspondence between the expert models and heterogeneous computing nodes; The time consumption determination unit includes a calculation time consumption acquisition subunit, a communication time consumption determination subunit and a total time consumption determination subunit; The computation time acquisition subunit is used to obtain the computation time of each expert model executing the inference task on its corresponding heterogeneous computing node; A communication time determination subunit is configured to determine the communication time based on the input size and output size included in the distributed reasoning task information, the uplink bandwidth, downlink bandwidth, and latency information included in the performance information, and the number of expert models activated in each round of reasoning; The total time consumption determining subunit is used to take the sum of the computation time consumption and the communication time consumption of each expert model as the total time consumption of each expert model.

[0082] In some embodiments, the communication time determination subunit is used to determine the amount of downlink data based on the input size and the number of expert models activated in each round of reasoning; The ratio of the downlink data volume to the downlink bandwidth is used as the downlink communication time. The amount of uplink data is determined based on the output size and the number of expert models activated in each round of inference. The ratio of the uplink data volume to the uplink bandwidth is used as the uplink communication time consumption; The downlink communication time, the uplink communication time, and the uplink delay and the downlink delay included in the performance information are summed to obtain the communication time.

[0083] In some embodiments, the ratio determination unit includes a comparison subunit, a selection subunit, an output subunit, and a compression ratio determination subunit; A comparison subunit is used to compare the total time consumption of all expert models with the communication time consumption to determine the optimal load balancing time consumption; A selection subunit is used to select the target expert model with the largest total time consumption according to the total time consumption of each expert model; An output subunit, configured to output a target expert model when the total time consumed by the target expert model is less than or equal to the optimal load balancing time consumed; The compression ratio determination subunit is used to determine the compression ratio according to the optimal load balancing time and the calculation time and total time of the target expert model when the total time consumption of the target expert model is greater than the optimal load balancing time consumption.

[0084] In some embodiments, the comparison subunit is used to filter out the minimum total time consumption based on the total time consumption of all expert models; Based on the communication time of all expert models, the maximum communication time is selected; The maximum value is selected from the minimum total time and the maximum communication time as the optimal load balancing time.

[0085] In some embodiments, the compression ratio determination subunit is used to divide the difference between the total time consumption of the target expert model and the optimal load balancing time consumption by the calculation time consumption of the target expert model to obtain the calculation amount compression ratio; and determine the compression ratio based on the sum of the calculation amount compression ratio and the compression ratio being one.

[0086] In some embodiments, the compression unit includes a pruning subunit, an error determination subunit, a change value determination subunit, an adjustment subunit, and a model output subunit; A pruning subunit, used to prune multiple neurons with the lowest importance scores in the target expert model based on the compression step size to obtain a compressed target expert model; an error determination subunit, configured to determine an output error of the compressed target expert model based on a deviation between an output vector of the target expert model and an output vector of the compressed target expert model; a change value determining subunit, configured to use the ratio of the computational amount of the compressed target expert model to the computational amount of the target expert model as a computational amount change value; an adjusting subunit, configured to adjust the compression step size when the output error of the compressed target expert model is less than the error threshold and the change in the computational complexity of the compressed target expert model is greater than the compression ratio, use the compressed target expert model as the latest target expert model, and return an operation step of pruning multiple neurons with the lowest importance scores in the target expert model based on the compression step size to obtain the compressed target expert model; The model output subunit is used to output the compressed target expert model when the output error of the compressed target expert model is less than the error threshold and the calculation amount change value of the compressed target expert model is less than or equal to the compression ratio.

[0087] In some embodiments, the pruning subunit is used to use the L1 norm of the weight of each neuron in the target expert model as the importance score of each neuron; determine a plurality of neurons to be pruned with the lowest importance scores consistent with the number of compression steps in ascending order of the importance scores of each neuron; and prune each neuron to be pruned from the target expert model to obtain a compressed target expert model.

[0088] In some embodiments, the adjustment unit is used to increase the number of pruning times by one each time a pruning operation of the expert model is performed; wherein, in the initial state, the value of the number of pruning times is one; and the product of the set search step size and the current value of the number of pruning times is used as the adjusted compression step size.

[0089] In some embodiments, it further includes a time-consuming adjustment unit and an as unit; The time consumption adjustment unit is used to adjust the total time consumption of the target expert model according to the calculation amount of the target expert model and the calculation amount of the compressed target expert model when the output error of the compressed target expert model is greater than or equal to the error threshold; and trigger the selection subunit to execute the operation step of selecting the target expert model with the largest total time consumption when the change value of the calculation amount of the compressed target expert model is less than or equal to the compression ratio; As a unit, it is used to use the adjusted total time consumption of the target expert model as the optimal load balancing time consumption when the change in the calculation amount of the compressed target expert model is greater than the compression ratio, and trigger the selection sub-unit to execute the operation steps of selecting the target expert model with the largest total time consumption.

[0090] In some embodiments, the time consumption adjustment unit is used to perform a multiplication operation on the ratio of the computational amount of the compressed target expert model to the computational amount of the target expert model and the computational consumption of the target expert model to obtain the adjusted computational consumption; and the sum of the adjusted computational consumption and the communication consumption of the target expert model is used as the adjusted total time consumption of the target expert model.

[0091] As can be seen from the above technical solution, distributed inference task information of a hybrid expert model and performance information of a heterogeneous computing system are obtained. The heterogeneous computing system includes multiple heterogeneous computing nodes, and the hybrid expert model includes multiple expert models. The distributed inference task information includes the architecture information of the expert model. Different architectures will result in different execution times for inference tasks. The performance of heterogeneous computing nodes will also affect the inference task time. Therefore, based on the performance information of the heterogeneous computing system and the distributed inference task information, the total execution time of each expert model on its corresponding heterogeneous computing node can be determined. The total execution time includes communication time and computation time. To ensure that the execution time of different expert models is as balanced as possible, the total execution time, communication time, and computation time of all expert models can be analyzed based on the principle of load balancing to determine a compression ratio. Each expert model is then iteratively compressed according to the model compression strategy to obtain compressed expert models that meet error and compression ratio requirements. Based on the correspondence between the expert model and the heterogeneous computing node, each compressed expert model is deployed on the corresponding heterogeneous computing node. In this application, the expert models deployed on different heterogeneous computing nodes are adjusted by model compression, and the compression degree of the expert model is determined based on the compression ratio, so that the calculation time of different heterogeneous computing powers in the expert operation layer is as balanced as possible, which accelerates the expert neural network reasoning task and effectively avoids the need to wait for a long time for the slowest expert model to complete calculation, thereby improving the utilization rate of hardware resources.

[0092] For the description of the features in the embodiment corresponding to the model deployment device, please refer to the relevant description of the embodiment corresponding to the model deployment method, and no further details will be given here.

[0093] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned model deployment method embodiments.

[0094] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned model deployment method embodiments when running.

[0095] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0096] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned model deployment method embodiments are implemented.

[0097] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned model deployment method embodiments are implemented.

[0098] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0099] The above is a detailed introduction to a model deployment method, device, equipment, storage medium and product provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only applicable to help understand the method of this application and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of this application.

Claims

1. A model deployment method, characterized in that: include: Obtaining distributed reasoning task information of a hybrid expert model and performance information of a heterogeneous computing system; wherein the heterogeneous computing system includes multiple heterogeneous computing nodes; and the hybrid expert model includes multiple expert models; Determine, based on the performance information of the heterogeneous computing system and the distributed reasoning task information, the total time consumed by each expert model to execute the reasoning task on its corresponding heterogeneous computing node; wherein the total time consumed includes communication time and computing time; Based on the load balancing principle, the total time consumption, communication time consumption and calculation time consumption of all the expert models are analyzed to determine the compression ratio; Iteratively compressing each of the expert models according to the model compression strategy to obtain compressed expert models that meet error requirements and compression ratio requirements; According to the corresponding relationship between expert models and heterogeneous computing nodes, each compressed expert model is deployed on the corresponding heterogeneous computing node.

2. The model deployment method according to claim 1, characterized in that: The distributed reasoning task information includes the network structure, input and output size of each expert model, the number of expert models activated in each round of reasoning, and the correspondence between the expert models and heterogeneous computing nodes; Determining the total time consumed by each expert model to execute the reasoning task on its corresponding heterogeneous computing node based on the performance information of the heterogeneous computing system and the distributed reasoning task information, including: Obtain the computational time required for each expert model to perform reasoning tasks on its corresponding heterogeneous computing nodes; Determining communication time based on the input size and output size included in the distributed reasoning task information, the uplink bandwidth, downlink bandwidth and latency information included in the performance information, and the number of expert models activated in each round of reasoning; The sum of the computation time and communication time of each expert model is taken as the total time of each expert model.

3. The model deployment method according to claim 2, characterized in that: Determining communication time based on the input size and output size included in the distributed reasoning task information, the uplink bandwidth, downlink bandwidth, and latency information included in the performance information, and the number of expert models activated in each round of reasoning, including: Determining the amount of downlink data based on the input size and the number of expert models activated in each round of reasoning; Taking the ratio of the downlink data volume to the downlink bandwidth as the downlink communication time consumption; Determining the amount of uplink data based on the output size and the number of expert models activated in each round of reasoning; Taking the ratio of the uplink data volume to the uplink bandwidth as the uplink communication time consumption; The downlink communication time, the uplink communication time, and the uplink delay and downlink delay included in the performance information are summed to obtain the communication time.

4. The model deployment method according to claim 1, characterized in that: Based on the load balancing principle, the total time consumption, communication time consumption, and computation time consumption of all the expert models are analyzed to determine the compression ratio, including: Compare the total time of all expert models with the communication time to determine the optimal load balancing time; According to the total time consumption of each expert model, a target expert model with the largest total time consumption is selected; When the total time consumption of the target expert model is less than or equal to the optimal load balancing time consumption, outputting the target expert model; When the total time consumption of the target expert model is greater than the optimal load balancing time consumption, a compression ratio is determined according to the optimal load balancing time consumption and the calculation time consumption and the total time consumption of the target expert model.

5. The model deployment method according to claim 4, characterized in that: Compare the total time of all expert models with the communication time to determine the optimal load balancing time, including: Based on the total time consumption of all expert models, filter out the minimum total time consumption; Based on the communication time of all expert models, the maximum communication time is selected; A maximum value is selected from the minimum total time consumption and the maximum communication time consumption to serve as the optimal load balancing time consumption.

6. The model deployment method according to claim 4, characterized in that: Determining the compression ratio according to the optimal load balancing time and the calculation time and total time of the target expert model includes: Dividing the difference between the total time consumption of the target expert model and the optimal load balancing time consumption by the calculation time consumption of the target expert model to obtain a calculation amount compression ratio; The compression ratio is determined according to the calculated amount compression ratio value and the compression ratio being one.

7. The model deployment method according to claim 4, characterized in that: Iteratively compressing each of the expert models according to the model compression strategy to obtain compressed expert models that meet the error requirements and the compression ratio requirements, including: Pruning a plurality of neurons with the lowest importance scores in the target expert model based on the compression step size to obtain a compressed target expert model; determining an output error of the compressed target expert model according to a deviation between an output vector of the target expert model and an output vector of the compressed target expert model; taking the ratio of the computation amount of the compressed target expert model to the computation amount of the target expert model as a computation amount change value; When the output error of the compressed target expert model is less than the error threshold and the calculation amount change value of the compressed target expert model is greater than the compression ratio, the compression step size is adjusted, the compressed target expert model is used as the latest target expert model, and the operation steps of pruning the multiple neurons with the lowest importance scores in the target expert model based on the compression step size are returned to obtain the compressed target expert model; When the output error of the compressed target expert model is less than the error threshold and the calculation amount change value of the compressed target expert model is less than or equal to the compression ratio, the compressed target expert model is output.

8. The model deployment method according to claim 7, characterized in that: Pruning multiple neurons with the lowest importance scores in the target expert model based on the compression step size to obtain a compressed target expert model, including: The L1 norm of the weight of each neuron in the target expert model is used as the importance score of each neuron; Determining, in ascending order of the importance scores of the neurons, a plurality of neurons to be pruned having the lowest importance scores that are consistent with the number of compression steps; Each of the neurons to be pruned is pruned from the target expert model to obtain a compressed target expert model.

9. The model deployment method according to claim 7, characterized in that: Adjusting the compression step size includes: Each time the pruning operation of the expert model is performed, the number of pruning times is increased by one; wherein, in the initial state, the value of the number of pruning times is one; The product of the set search step size and the current value of the pruning times is used as the adjusted compression step size.

10. The model deployment method according to claim 7, characterized in that: Also includes: When the output error of the compressed target expert model is greater than or equal to the error threshold, adjusting the total time consumption of the target expert model according to the computational amount of the target expert model and the computational amount of the compressed target expert model; When the calculation amount change value of the compressed target expert model is less than or equal to the compression ratio, returning to the operation step of selecting the target expert model with the largest total time consumption; When the change in the computational complexity of the compressed target expert model is greater than the compression ratio, the adjusted total time consumption of the target expert model is used as the optimal load balancing time consumption, and the operation step of selecting the target expert model with the largest total time consumption is returned.

11. The model deployment method according to claim 10, characterized in that: Adjusting the total time consumption of the target expert model according to the computational amount of the target expert model and the computational amount of the compressed target expert model includes: Performing a multiplication operation on the ratio of the calculation amount of the compressed target expert model to the calculation amount of the target expert model and the calculation time of the target expert model to obtain an adjusted calculation time; The sum of the adjusted calculation time consumption and the communication time consumption of the target expert model is used as the adjusted total time consumption of the target expert model.

12. A model deployment device, characterized in that: It includes an acquisition unit, a time consumption determination unit, a compression ratio determination unit, a compression unit and a deployment unit; An acquisition unit, configured to acquire distributed reasoning task information of a hybrid expert model and performance information of a heterogeneous computing system; wherein the heterogeneous computing system includes a plurality of heterogeneous computing nodes; and the hybrid expert model includes a plurality of expert models; A time consumption determination unit is used to determine the total time consumed by each expert model to execute the reasoning task on its corresponding heterogeneous computing node based on the performance information of the heterogeneous computing system and the distributed reasoning task information; wherein the total time consumption includes communication time consumption and computing time consumption; A compression ratio determination unit, configured to analyze the total time consumption, communication time consumption, and calculation time consumption of all the expert models based on a load balancing principle to determine a compression ratio; A compression unit, configured to iteratively compress each of the expert models according to a model compression strategy to obtain compressed expert models that meet error requirements and compression ratio requirements; The deployment unit is used to deploy each compressed expert model on the corresponding heterogeneous computing node according to the corresponding relationship between the expert model and the heterogeneous computing node.

13. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the model deployment method according to any one of claims 1 to 11 when executing the computer program.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the model deployment method according to any one of claims 1 to 11 are implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the model deployment method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Quantization method, device and equipment of image segmentation model and computer storage medium

    CN116630632A

  • Model compression deployment method and device, server and storage medium

    CN116776953A

  • Model compression method and device, electronic equipment and storage medium

    CN117973481A

  • Model pruning method and device for heterogeneous computing cluster and storage medium

    CN119227769A

  • Decentralized artificial intelligence (AI) / machine learning training system

    US20220344049A1