Distributed CNN reasoning method on heterogeneous edge device cluster
By parallel training and slicing the CNN model and optimizing the pipeline inference mode, the computing and memory resource limitation problems of high-performance CNN models on edge devices are solved, and efficient inference task execution and energy consumption management are achieved.
Patent Information
- Application Number
- CN202510265260.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
The high-performance CNN model puts higher requirements on the computing power, memory resources and power consumption of edge devices, making it difficult to directly deploy to edge devices to perform inference tasks, and there is a lot of room for exploration on how to find high-quality model cutting points and computing resource allocation solutions.
The CNN model is trained through parallel characteristics and divided into multiple parts. The throughput, memory load and energy consumption of different pipeline schemes are evaluated according to available resources, and the pipeline reasoning model is optimized to reduce latency, improve system throughput, and consider energy consumption factors when allocating resources.
It realizes efficient execution of CNN inference tasks on edge devices, improves system throughput, reduces overall energy consumption, and effectively utilizes edge device resources and reduces resource waste.
Smart Images

Figure CN120218131A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a distributed CNN inference method on a heterogeneous edge device cluster. Background Art
[0002] In recent years, convolutional neural networks have made great progress in the field of deep learning and have been widely applied to intelligent fields such as computer vision, natural language processing, and automation. Specifically, CNN models are currently widely used in many fields such as intelligent healthcare, autonomous driving, intelligent transportation, fintech, text classification, and speech recognition. In the future, with the continuous improvement of CNN models and technologies, they will continue to be widely applied in various industries and explore their roles in more fields, promoting the intelligent development of various fields in modern society. Unfortunately, both the depth and complexity of CNN models are increasing continuously. From the previous LeNet-5 with only a few layers to the current ResNet-101 with more than a hundred layers, we have witnessed the continuous improvement of the inference performance of CNN models. However, it must also be noted that high-performance CNN models pose higher requirements for the computing power, memory resources, and power consumption of devices. Edge devices usually have limited computing and memory resources and low power, which seriously hinders the direct deployment of high-performance network models to edge devices to perform inference tasks.
[0003] An automated distributed inference framework can automatically cut the model and generate corresponding communication and configuration files, and at the same time deploy each sub-model to each device to achieve distributed inference. Given available computing devices and resources and a pre-trained model, since modern high-performance CNN models usually have many layers, how to determine the model cutting points and allocate computing resources for each sub-model to form a specific pipeline scheme usually has a large exploration space. Among numerous pipeline schemes, the throughput, energy consumption, and memory occupancy of model inference under different pipeline schemes are all different. How to find a high-quality pipeline scheme is a problem worthy of exploration. Summary of the Invention
[0004] To solve the above technical problems, the present invention is realized through the following technical solutions:
[0005] The present invention provides a distributed CNN inference method on a heterogeneous edge device cluster, including the following steps:
[0006] Step S1: Train a CNN model using parallel characteristics;
[0007] Step S2: Split and allocate the trained CNN model to form pipeline inference;
[0008] Step S3: Evaluate the throughput, memory load, and energy consumption of different pipelines based on the available resources of the allocated CNN model;
[0009] Step S4: performing a segmentation operation on the pre-trained CNN model to form a pipeline solution;
[0010] The step S4 includes the following sub-steps:
[0011] Step S41, obtaining summary information of each level model by analyzing the trained CNN model;
[0012] Step S42, dividing the CNN model into multiple parts, calculating the resource allocation scheme by dividing the parts, and calculating the number of resource allocation schemes according to the number of divisions;
[0013] Step S43: Divide the CNN model into j parts, distribute them to j devices, select computing resources on each device for distribution, and achieve memory load balancing;
[0014] Step S44, evaluating each pipeline solution by counting the system throughput, the memory load deviation of each device, and the maximum energy consumption of a single device when each pipeline solution is executed.
[0015] Furthermore, in step S1, GPipe is used to divide the model into multiple segments, each segment is assigned to a different GPU, and the input data is divided into small micro-batches so that each GPU can process different batches of training data at the same time, thereby utilizing pipeline parallelism to distribute the training of large models across multiple devices.
[0016] Furthermore, the step S2 includes the following steps:
[0017] Step S21, model segmentation: vertically segment the trained CNN model into multiple sub-models, and deploy each sub-model on the computing resources of different edge devices. When receiving multiple requests, each edge device executes each sub-model in turn and synchronizes and communicates the intermediate inference results, thereby forming a pipeline inference mode;
[0018] Step S22, computing resource allocation: When a resource has a strong computing capability, a sub-model with a larger computing amount is allocated to it. Similarly, when a resource has a weaker computing capability, a sub-model with a smaller computing amount is allocated to it. Ultimately, the computing delays between the various stages of the pipeline are guaranteed to be close, thereby reducing the waiting time for each pipeline stage to transmit intermediate results, avoiding a certain stage of the pipeline from becoming the performance bottleneck of the entire pipeline, accelerating the efficiency of the pipeline inference, thereby improving throughput and reducing the energy consumption of related equipment.
[0019] Furthermore, step S3 includes the following steps:
[0020] Step S31. Resource Allocation for the CNN Model: For a CNN model, there are n available computing resources for allocation, denoted as R = [r1, r2..., rn]. The computing power corresponding to each resource is denoted as c = [c1, c2..., cn]. The model is vertically divided into n parts to form a pipeline with n stages, denoted as s = [s1, s2..., sn]. If the average time for each stage to process one picture is corresponding to T = [t1, t2..., tn], there are a total of j available devices denoted as D = [d1, d2..., dj], and the total memory of each device is denoted as M = [M1, M2..., Mj]. Each device corresponds to multiple resources and can be allocated to different stages of the pipeline simultaneously;
[0021] Step S32: Select the pipeline stage with the highest latency for processing one picture, and obtain the number of images processed per second in this stage as the system throughput Tsystem;
[0022]
[0023] The processing of one picture by each pipeline stage should include computing latency and communication latency. The computing latency is to perform inference through the relevant CNN layers of the input. Except that the first stage receives the input of the system and the last stage directly outputs the final result, the communication latency includes receiving the intermediate result from the previous stage and sending the intermediate inference result of this stage to the next stage. The computing latency depends on the number of multiply-accumulate operations of the sub-model in this stage and the computing power of the computing resources allocated to this stage. The communication latency involves intra-node communication between the GPU and CPU within the same device and network communication between different devices;
[0024] Step S33. Memory Load: The memory occupancy Mi of the pipeline stage comes from three parts. One is the memory occupancy for storing the weights, parameters, and biases of the sub-model in this stage during inference The second is the memory occupancy for storing intermediate results during the inference of the CNN layer The third is the memory occupancy for receiving intermediate inference results from other pipeline stages
[0025]
[0026] If the computing resources used in different pipeline stages are the same device, and if device dj has k computing resources allocated to the pipeline stage set The total memory occupancy of this device The formula is as follows:
[0027]
[0028] Statistically calculate the total memory occupancy of all devices participating in pipeline inference under the corresponding pipeline scheme, and calculate the memory utilization rate of a single device. The formula for the memory utilization rate Uj of a single device j is as follows:
[0029]
[0030] After obtaining the memory utilization rate of each device, calculate the average memory utilization rate of all devices The calculation process is as follows:
[0031]
[0032] Then, the memory load deviation σU under the corresponding pipeline scheme can be calculated as follows:
[0033]
[0034] The smaller the value of σU, the closer the memory usage rates of each device are, the more balanced the memory loads of each device are, and the higher the quality of the pipeline scheme;
[0035] Step S34, Energy consumption: The energy consumed in each pipeline stage is concentrated in the process of executing the CNN sub-model inference and the process of intermediate data communication with other pipeline stages. If the energy consumption for executing inference in one of the pipeline stages is The energy consumption for data communication with other stages is The formula for the total energy consumption Ei of each pipeline stage is as follows:
[0036]
[0037] If k computing resources of one device dj are allocated to the pipeline stage set Sj = {sj1, sj2..., sjk}, the total energy consumption of this device Is expressed as:
[0038]
[0039] Statistically calculate the energy consumed when all devices execute pipeline inference, and select the maximum energy consumption Emax of a single device as the evaluation index. The lower the maximum energy consumption of a single device, the higher the quality of the corresponding pipeline scheme;
[0040] Different pipeline schemes have different performance in various aspects, which depends on which application goal the current application scenario mainly focuses on. For some scenarios with high real-time requirements, the main concern is the system throughput under the current pipeline scheme. In this case, a pipeline scheme that minimizes the inference latency should be found to improve the system throughput. At the same time, when the device has a low inference latency, the device can also reduce the energy consumption during inference. For some scenarios, there may be no strict requirement for throughput, but there is a strict requirement for the memory occupancy of the device. Since the memory resources of edge devices are usually very scarce, this requires us to reasonably allocate sub-models to each device when designing the pipeline scheme to balance the memory load of each device. At the same time, in some scenarios, it may be necessary to comprehensively consider latency, energy consumption, and device memory occupancy. In this case, it is also crucial to find a pipeline scheme with relatively balanced performance in all aspects.
[0041] Further, in step S41, a pre-trained CNN model is given. By analyzing the model, relevant information of each layer and the overall summary information of the model are obtained. The overall summary information of the model includes the layer name, layer type, number of multiply-accumulate operations during forward propagation, memory occupancy of the layer, number of layer parameters, input and output shapes, total number of multiply-accumulate operations corresponding to all layers of the model, total memory occupancy during model layer execution, and total number of parameters of the model.
[0042] Further, in step S42, a CNN model M and n available resources are given. The model is divided into multiple parts and a different resource is assigned to each part. It can be divided into at most n parts to form a pipeline with n stages. The model can be divided into two parts to form a two-stage pipeline, or it can be divided into three parts to form a three-stage pipeline, and so on. It can be divided into at most n parts to form a pipeline with n stages. If the model is divided into two parts, two resources need to be selected from n resources and assigned to these two parts, then there are n*(n - 1) ways of allocation. Similarly, if the model is divided into three parts, then there are n*(n - 1)*(n - 2) ways of allocation. For dividing the model into n parts, there are n! ways of allocation in total. If the model is divided into p parts, then the total number of resource allocation methods solutions(n, p) can be calculated by the formula:
[0043]
[0044] For different numbers of divisions, the formula for the total number of resource allocation schemes solutions(n, p) of the model is as follows:
[0045]
[0046] After splitting into resource allocation schemes corresponding to different parts, let the total computational amount of the model be X, which is obtained by statistical analysis using relevant model analysis tools. When the model is split into p parts, the computational amounts of each part are denoted as X1, X2,..., XP, and the computational capabilities corresponding to the resources allocated to each part are C1, C2,..., CP, respectively. Then we have:
[0047] X1 + X2 +... + XP = X
[0048] To ensure the inference of the pipeline scheme, the delays of executing CNN inference in each stage are equal and denoted as L. Then we have:
[0049]
[0050] Through the above two formulas, the computational amounts allocated to each part are obtained to find the relevant cut points. Combining the above formulas, the expression of L can be simplified as:
[0051]
[0052] To make the pipeline scheme have lower latency and energy consumption, the optimization goal is to minimize L under the condition of satisfying formula 2. The problem is then defined as finding the top p resources with the highest computational capabilities from n resources and allocating them to the corresponding p stages. There are a total of p! corresponding pipeline schemes that minimize the latency L. Among all these pipeline schemes, by statistically calculating the total size of the outputs of all cut points corresponding to each scheme and selecting the fifty pipeline schemes with the smallest total size of the outputs of all cut points for testing. If the number of all schemes is less than fifty, then all of them are selected for subsequent on-machine testing. This is because when the communication bandwidth is stable and the intermediate results transmitted by each cut point are relatively small, the communication latency between the stages of the pipeline can be effectively reduced, further improving the system throughput and reducing the latency. At the same time, it can also reduce the number of schemes that need to be executed on the machine and speed up the search.
[0055] By finding the optimal pipeline schemes for splitting the model into two parts, three parts, up to n parts; successively evaluating the performance of the found schemes using the AutoD i CE framework during execution, including system throughput, the maximum energy consumption of a single device, and the standard deviation of the memory load of each device to evaluate the effectiveness of the pipeline scheme.
[0056] Further, in step S43, for a device, it can be allocated to different stages of the pipeline simultaneously. When the pipeline is executing normally, two resources work simultaneously. In resource-constrained edge devices, the GPU and CPU usually share memory. When the two resources execute CNN inference simultaneously, this usually occupies a lot of memory of a single device. Each device is only responsible for one stage of the pipeline, reducing the memory load pressure on the device. The memory sizes of different devices are different. For devices with rich memory resources, sub-models with larger memory occupancy are allocated to them, while for devices with scarce memory resources, sub-models with smaller memory occupancy are allocated.
[0057] When there are j available devices, the model is divided into j parts, and each device is responsible for one stage of the pipeline. At the same time, in order to make the throughput of the pipeline scheme relatively high and the energy consumption of each device relatively low, the resource with the strongest computing power of each device is selected for allocation. Then, the specific splitting points need to be determined. If the memory load of each device needs to be balanced, only the memory consumption of the sub-model assigned to each device needs to match its own memory resources. At this time, the model splitting points can be determined according to the total memory ratio of each device. The memory occupancy of each pipeline stage mainly consists of three parts. Compared with and the memory occupancy of each pipeline stage for receiving intermediate results from other stages is relatively small and changes with the change of the cutting point. The main consideration is the and of each pipeline stage to approximate the total memory occupancy of each stage.
[0058] Let the total memory occupancy of the CNN model be Mtotal, and there are j available devices in total. After the model is divided into j sub-models and each sub-model is allocated to j devices respectively, the memory occupancy on each device is Then there is:
[0059]
[0060] To make the load of each device balanced, the following formula should be ensured to hold:
[0061]
[0062] Therefore, the optimization goal is to divide the model into j parts and make the above formula (1) and formula (2) hold simultaneously. At the same time, to achieve memory load balance and ensure that the pipeline inference has relatively high throughput and low energy consumption, the resource with the strongest computing power of each device is selected for allocation.
[0063] Furthermore, in step S44, after finding the pipeline scheme by the above method, the AutoDiCE framework is used to segment the model according to the corresponding pipeline scheme and deploy it to each device for actual testing, and each pipeline scheme is evaluated by statistically analyzing the system throughput, the memory load deviation of each device, and the maximum energy consumption of a single device when each pipeline scheme is executed;
[0064] After the test is completed, each pipeline solution corresponds to three indicators: Tsystem, σU, and Emax. The three indicators are normalized to [0, 1], and then the weighted comprehensive score of each solution is calculated as follows:
[0065] score=αTsystem+βσU+γEmax,(α+β+γ=1)
[0066] At the same time, the standard deviation of each pipeline scheme on the three indicators is calculated to measure the fluctuations between these indicators. When system throughput is the main indicator to be considered, a larger weight is set for α. Finally, each pipeline scheme is arranged in descending order by weighted comprehensive score. When the weighted comprehensive scores are the same, auxiliary sorting is performed according to the standard deviation. The scheme with a small standard deviation will be given priority. The top-ranked scheme with a higher weighted comprehensive score represents a better system throughput performance. At the same time, a higher weight can be set for β and the schemes are arranged in ascending order according to the weighted comprehensive score to select a pipeline scheme that balances the memory load of each device, and a higher weight can be set for γ and the schemes are arranged in ascending order according to the weighted comprehensive score to find a pipeline scheme that makes each device have lower energy consumption. When close values are set for the three indicators, a scheme with relatively balanced performance in all aspects can be selected, which depends on the actual application requirements.
[0067] The present invention has the following beneficial effects:
[0068] 1. In the process of model training and reasoning, the present invention utilizes parallel characteristics to train the model, and optimizes the pipeline reasoning mode to reduce the waiting time of each pipeline stage, thereby improving the system throughput. At the same time, energy consumption factors are considered in resource allocation to reduce the overall energy consumption. The pipeline reasoning method in the embodiment performs similarly or even better than NSGA2 in terms of system throughput and energy consumption indicators. In some experiments, the best energy consumption solution is only 0.8% different from NSGA2, and the system throughput gap is also small, and the search time is greatly shortened, reducing the workload pressure of the equipment.
[0069] 2. The present invention reasonably divides and distributes the model to match the computing amount with the computing power of the device, achieves memory load balancing, and can effectively utilize edge device resources for CNN reasoning tasks, reducing resource waste. In experiments under different device configurations, the model part can be reasonably allocated according to the device situation to ensure the smooth progress of the reasoning process.
[0070] 3. The present invention evaluates the pipeline scheme by comprehensively considering multiple indicators such as system throughput, memory load deviation, and maximum energy consumption of a single device, calculates the weighted comprehensive score after normalizing the indicators for ranking, and can also adjust the indicator weights according to the requirements of different application scenarios, so as to be able to select a high-quality pipeline scheme that meets the actual application requirements, ensure the relative balance of the overall performance in different scenarios, increase the relevant weight screening scheme in scenarios with strict memory load requirements, and focus on the throughput indicator selection scheme in scenarios that emphasize real-time performance.
[0071] Of course, any product implementing the present invention does not necessarily need to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0073] Figure 1 It is a schematic flow diagram of a distributed CNN inference method on a heterogeneous edge device cluster according to the present invention;
[0074] Figure 2 They are the ordinary inference mode and the pipeline inference mode;
[0075] Figure 3 It is a schematic diagram of the overall system architecture;
[0076] Figure 4 It is the comparison of system throughput in the first group of experiments;
[0077] Figure 5 It is the comparison of device energy consumption in the first group of experiments;
[0078] Figure 6 It is the comparison of the memory load of each device in the first group of experiments;
[0079] Figure 7 It is the comparison of the overall performance of the schemes with system throughput higher than 25 in the first group of experiments;
[0080] Figure 8 It is the comparison of the overall performance of the schemes with a standard deviation of memory load less than 0.01 in the first group of experiments;
[0081] Figure 9 It is the comparison of system throughput in the second group of experiments;
[0082] Figure 10For the comparison of device energy consumption in the second group of experiments;
[0083] Figure 11 For the comparison of device memory load in the second group of experiments;
[0084] Figure 12 For the overall comparison of the performance of the solutions with system throughput greater than 20 in the second group of experiments;
[0085] Figure 13 For the overall comparison of the performance of the solutions with standard deviation of memory load less than 0.01 in the second group of experiments. Specific implementation manners
[0086] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0087] Please refer to Figure 1 As shown, the present invention is a distributed CNN inference method on a heterogeneous edge device cluster, including the following steps:
[0088] Step S1: Train a CNN model using parallel characteristics;
[0089] Step S2: Split and allocate the trained CNN model to form pipelined inference;
[0090] Step S3: Evaluate the throughput, memory load, and energy consumption of different pipelines according to the available resources of the allocated CNN model;
[0091] Step S4: Perform a segmentation operation on the pre-trained CNN model to form a pipeline solution;
[0092] In step S4, the following sub-steps are included:
[0093] Step S41: Obtain the summary information of each hierarchical model by analyzing the trained CNN model;
[0094] Step S42: Split the CNN model into multiple parts, calculate the resource allocation scheme through the split parts, and calculate the number of resource allocation schemes according to the number of splits;
[0095] Step S43: Split the CNN model into j parts, allocate them to j devices, and select the computing resources on each device for allocation to achieve memory load balancing;
[0096] Step S44: Evaluate each pipeline scheme by statistically analyzing the system throughput, the memory load deviation of each device, and the maximum energy consumption of a single device during the execution of each pipeline scheme.
[0097] In step S1, GPipe is used to divide the model into multiple segments, each segment is assigned to a different GPU, and the input data is sliced into small batches of micro-batches to enable each GPU to process different batches of training data simultaneously, thereby using the pipeline parallelism method to perform distributed training of large models across multiple devices.
[0098] Step S2 includes the following steps:
[0099] Step S21: Model segmentation: Vertically divide the trained CNN model into multiple sub-models, each sub-model is deployed to the computing resources of different edge devices. When receiving multiple requests, each edge device sequentially executes each sub-model and synchronizes and communicates the intermediate inference results, thereby forming a pipeline inference mode.
[0100] Step S22: Computing resource allocation: When a resource has strong computing power, assign a sub-model with a larger amount of computation to it. Similarly, when a resource has weak computing power, assign a sub-model with a smaller amount of computation to it. Ultimately, ensure that the computing delays between the various stages of the pipeline are close, thereby reducing the time for each stage of the pipeline to wait for the transmission of intermediate results from each other, avoiding a certain stage of the pipeline from becoming the performance bottleneck of the entire pipeline, accelerating the pipeline inference efficiency, thereby increasing the throughput and reducing the energy consumption of related devices.
[0101] Step S3 includes the following steps:
[0102] Step S31: Resource allocation of the CNN model: For a CNN model, there are n available computing resources for allocation, denoted as R = [r1, r2..., rn]. The computing power corresponding to each resource is denoted as c = [c1, c2..., cn]. Vertically divide the model into n parts and form a pipeline with n stages, denoted as s = [s1, s2..., sn]. If the average time for each stage to process one picture corresponds to T = [t1, t2..., tn], there are a total of j available devices denoted as D = [d1, d2..., dj], and the total memory of each device is denoted as M = [M1, M2..., Mj]. Each device corresponds to multiple resources and can be simultaneously allocated to different stages of the pipeline.
[0103] Step S32: Select the pipeline stage with the highest processing delay for one picture, and obtain the number of images processed per second at this stage as the system throughput Tsystem.
[0104]
[0105] When each pipeline stage finishes processing an image, it should include computing latency and communication latency. The computing latency is obtained by performing inference on the relevant CNN layers of the input execution. Except for the first stage that receives the input of the system and the last stage that directly outputs the final result, the communication latency includes receiving the intermediate results from the previous stage and sending the intermediate inference results of this stage to the next stage. The computing latency depends on the number of multiply-accumulate operations of the sub-model in this stage and the computing power of the computing resources allocated to this stage. The communication latency involves intra-node communication between the GPU and CPU within the same device and network communication between different devices;
[0106] Step S33, Memory Load: The memory occupancy Mi of the pipeline stage comes from three parts. One is the memory occupancy for storing the weights, parameters, and biases of the sub-model in this stage during inference The second is the memory occupancy for storing intermediate results during the inference of the CNN layer The third is the memory occupancy for receiving intermediate inference results from other pipeline stages
[0107]
[0108] If the computing resources used in different pipeline stages are the same device, and if device dj has k computing resources allocated to the pipeline stage set The total memory occupancy of this device The formula is as follows:
[0109]
[0110] Statistically calculate the total memory occupancy of all devices participating in the pipeline inference under the corresponding pipeline scheme, and calculate the memory utilization rate of a single device. The formula for the memory utilization rate Uj of a single device j is as follows:
[0111]
[0112] After obtaining the memory utilization rate of each device, calculate the average memory utilization rate of all devices The calculation process is as follows:
[0113]
[0114] Then, the memory load deviation σU under the corresponding pipeline scheme can be calculated as follows:
[0115]
[0116] The smaller the value of σU, the closer the memory usage rates of each device are, the more balanced the memory loads of each device are, and the higher the quality of the pipeline scheme;
[0117] Step S34, Energy Consumption: The energy consumed in each pipeline stage is concentrated in the process of executing the CNN sub-model inference and the process of intermediate data communication with other pipeline stages. If the energy consumption corresponding to the execution of inference in one pipeline stage is and the energy consumption for data communication with other stages is The formula for the total energy consumption Ei of each pipeline stage is as follows:
[0118]
[0119] If k computing resources of one device dj are allocated to the pipeline stage set Sj = {sj1, sj2..., sjk}, the total energy consumption of this device is expressed as:
[0120]
[0121] Statistically analyze the energy consumed when all devices execute pipeline inference, and select the maximum energy consumption Emax of a single device as the evaluation index. The lower the maximum energy consumption of a single device, the higher the quality of the corresponding pipeline scheme.
[0122] Step S42, Given a CNN model M and n available resources, divide the model into multiple parts and allocate a different resource to each part; it can be divided into at most n parts to form a pipeline with n stages; the model can be divided into two parts to form a two-stage pipeline, or it can be divided into three parts to form a three-stage pipeline, and so on, and it can be divided into at most n parts to form a pipeline with n stages; if the model is divided into two parts, two resources need to be selected from n resources and allocated to these two parts, then there are n*(n - 1) allocation methods. Similarly, if the model is divided into three parts, then there are n*(n - 1)*(n - 2) allocation methods. For dividing the model into n parts, there are a total of n! allocation methods; if the model is divided into p parts, then the total number of resource allocation methods solutions(n, p) at this time can be calculated by the formula:
[0123]
[0124] For different numbers of divisions, the formula for the total number of resource allocation solutions solutions(n, p) of the model is as follows:
[0125]
[0126] After splitting into resource allocation schemes corresponding to different parts, let the total computational amount of the model be X, which is obtained by statistical analysis using relevant model analysis tools; when the model is split into p parts, the computational amounts of each part are denoted as X1, X2,..., XP, and the computational capabilities corresponding to the resources allocated to each part are C1, C2,..., CP, respectively. Then we have:
[0127] X1 + X2 +... + XP = X
[0128] To ensure the inference of the pipeline scheme, the delays of executing CNN inference in each stage are equal and denoted as L. Then we have:
[0129]
[0130] Through the above two formulas, the computational amounts allocated to each part are obtained, and thus the relevant cutting points are found; combining the above formulas, the expression of L can be simplified as:
[0131]
[0132] To make the pipeline scheme have lower latency and energy consumption, the optimization goal is to minimize L under the condition of satisfying Formula 2. The problem is then defined as finding the top p resources with the highest computational capabilities from n resources and allocating them to the corresponding p stages. There are a total of p! corresponding pipeline schemes that minimize the latency L. Among all these pipeline schemes, by statistically calculating the total size of the outputs of all cutting points corresponding to each scheme and selecting the fifty pipeline schemes with the smallest total size of the outputs of all cutting points for testing. If the number of all schemes is less than fifty, then all of them are selected for subsequent on-machine testing. This is because when the communication bandwidth is stable, when the intermediate results transmitted by each cutting point are relatively small, the communication latency between the stages of the pipeline can be effectively reduced, further improving the system throughput and reducing the latency; at the same time, it can also reduce the number of schemes that need to be executed on the machine and speed up the search;
[0135] By finding the optimal pipeline schemes for splitting the model into two parts, three parts, up to n parts; successively evaluating the performance during execution of the found schemes using the AutoDiCE framework, including system throughput, the maximum energy consumption of a single device, and the standard deviation of the memory load of each device to evaluate the effectiveness of the pipeline scheme;
[0136] Step S43: For a device, it can be allocated to different stages of the pipeline simultaneously. When the pipeline is executing normally, two resources work simultaneously. In resource-constrained edge devices, the GPU and CPU usually share memory. When the two resources execute CNN inference simultaneously, this usually occupies a lot of memory of a single device. Each device is only responsible for one stage of the pipeline, reducing the memory load pressure on the device. The memory sizes of different devices are different. For devices with rich memory resources, sub-models with larger memory occupancy are allocated to them, while for devices with scarce memory resources, sub-models with smaller memory occupancy are allocated.
[0137] When there are j available devices, the model is divided into j parts, and each device is responsible for one stage of the pipeline stage. At the same time, in order to make the throughput of the pipeline scheme relatively high and the energy consumption of each device relatively low, the resource with the strongest computing power of each device is selected for allocation. After that, the specific splitting point needs to be determined. If the memory load of each device is to be balanced, only the memory consumption of the sub-model allocated to each device needs to match its own memory resources. At this time, the model splitting point can be determined according to the total memory ratio of each device. The memory occupancy of each pipeline stage mainly consists of three parts. Compared with and the memory occupancy of each pipeline stage for receiving intermediate results from other stages is relatively small and changes with the change of the cutting point. The main consideration is the and of each pipeline stage to approximate the total memory occupancy of each stage.
[0138] Let the total memory occupancy of the CNN model be Mtotal, and there are j available devices in total. After the model is divided into j sub-models and each sub-model is allocated to j devices respectively, the memory occupancies on each device are Then there is:
[0139]
[0140] To make the load of each device balanced, the following formula should be ensured to hold:
[0141]
[0142] Therefore, the optimization goal is to divide the model into j parts and make the above formula one and formula two hold simultaneously. At the same time, to achieve memory load balance and ensure that the pipeline inference has a relatively high throughput and low energy consumption, the resource with the strongest computing power of each device is selected for allocation;
[0143] Step S44, after finding the pipeline scheme through the above method, use the AutoDiCE framework to segment the model according to the corresponding pipeline scheme and deploy it to each device for actual testing, and evaluate each pipeline scheme by counting the system throughput, memory load deviation of each device, and maximum energy consumption of a single device when executing each pipeline scheme;
[0144] After the test is completed, each pipeline solution corresponds to three indicators: Tsystem, σU, and Emax. The three indicators are normalized to [0, 1], and then the weighted comprehensive score of each solution is calculated as follows:
[0145] score=αTsystem+βσU+γEmax,(α+β+γ=1)
[0146] At the same time, the standard deviation of each pipeline scheme on the three indicators is calculated to measure the fluctuations between these indicators. When system throughput is the main indicator to be considered, a larger weight is set for α. Finally, each pipeline scheme is arranged in descending order by weighted comprehensive score. When the weighted comprehensive scores are the same, auxiliary sorting is performed according to the standard deviation. The scheme with a small standard deviation will be given priority. The top-ranked scheme with a higher weighted comprehensive score represents a better system throughput performance. At the same time, a higher weight can be set for β and the schemes are arranged in ascending order according to the weighted comprehensive score to select a pipeline scheme that balances the memory load of each device, and a higher weight can be set for γ and the schemes are arranged in ascending order according to the weighted comprehensive score to find a pipeline scheme that makes each device have lower energy consumption. When close values are set for the three indicators, a scheme with relatively balanced performance in all aspects can be selected, which depends on the actual application requirements.
[0147] Embodiment 1
[0148] The equipment used in this embodiment is as follows:
[0149] In the experiment, high-quality pipeline schemes were explored in heterogeneous device clusters with different numbers. The devices specifically used were NVIDIA Jetson Nano and NVIDIA Jetson AGX Xavier, which differed in terms of computing performance, memory resources, and power consumption. NVIDIA Jetson Nano was equipped with 4GB of memory, had a 128-core GPU with a Maxwell architecture and a quad-core ARM Cortex-A57 CPU. NVIDIA Jetson AGX Xavier was equipped with 32GB of memory, had a 512-core GPU with a Volta architecture and an 8-core ARM architecture CPU. In the experiment, for Nano, all 4 CPU cores and the GPU were used, and for Xavier, 6 CPU cores and the GPU were used. In the first group of experiments, one NVIDIA Jetson Nano and one NVIDIA Jetson AGX Xavier were used, so there were a total of 4 available resources. In the second group of experiments, three NVIDIA Jetson Nano and one NVIDIA Jetson AGX Xavier were used, and in such a configuration, there were a total of eight available computing resources. The devices were connected through a gigabit network switch. In terms of model selection, the mainstream CNN model ResNet-101 was chosen because it was widely used and had many layers. Therefore, there was a huge search space for searching high-quality pipeline schemes for it, so it was representative. In the actual on-machine evaluation, the system throughput, maximum energy consumption per device, and standard deviation of memory load of each device corresponding to the pipeline inference of the CNN model on 20 images in sequence were collected for each scheme. The images used were 20 classes of images selected from ImageNet-100.
[0150] In NSGA2, the population size of the NSGA-II algorithm used was 100 individuals, the mutation probability was 0.2, and the crossover probability was 0.5. For the first group of experiments, the population generation was set to 50, and in the second group of experiments, the population generation was set to 100. All the pipeline schemes found by the method were also tested on the machine. By comparing the performance of the pipeline schemes found by this method in terms of the three evaluation indicators, the effectiveness and efficiency of the method were demonstrated.
[0151] Such as Figure 4 、 5As shown in Figures 5 and 6, in the first group of experiments, it can be seen that the pipeline inference method took 1.3 hours, which was 74.23 times shorter than the method used by NSGA2. This greatly shortened the search time, improved the efficiency, reduced the workload pressure on the device, and saved resources. In terms of the performance of each indicator, it can be seen that compared with the best system throughput solution found by NSGA2, we only differed by 1.27%. It can be said that our performance was consistent with the best solution. At the same time, in terms of device energy consumption, the best energy consumption solution found by the pipeline only differed by 0.8% from the best of NSGA2. In terms of memory load, the gap between the pipeline inference method and NSGA2 was relatively large. By statistically analyzing the performance of each point in the figure in other aspects for systematic comparison, such as Figure 8 shown, the pipeline solution was far superior to the NSGA2 solution in terms of system throughput and energy consumption. The reason is that NSGA2 only focused on finding the best solution for a certain indicator and did not consider other indicators. Therefore, it was easy to find some extreme model partitioning and resource allocation solutions, resulting in a significant decline in performance in other indicators. Such a solution is usually not advisable in practice. In the pipeline inference method, while considering memory load balancing, the performance of other indicators is also considered. Therefore, the pipeline solution can ensure relatively balanced memory load while also ensuring high system throughput and low energy consumption of the system. The solutions with a system throughput higher than 25 were selected for comprehensive comparison, as Figure 7 shown, it can be seen that the solutions with higher system throughput were also more balanced in other aspects.
[0152] Example Two
[0153] In the second experiment, as follows Figure 9 、 10 and 11, the total running time of the pipeline inference method was 12.16 hours, while NSGA2 took 285 hours. This indicates that the pipeline inference method improved the search efficiency by 23.44 times. Specifically, in terms of various indicators, the best system throughput of the pipeline solution was only 9.8% lower than that of NSGA2, and the difference in the minimum device energy consumption was 9.6% lower than that of NSGA2. Even though the search time was significantly reduced, the pipeline solutions generated by the pipeline inference method still maintained relatively good performance, which was acceptable in general. In addition, in terms of the memory load between devices, the pipeline inference method was also able to achieve a relatively balanced memory load pipeline solution;
[0154] Similarly, to systematically compare the overall performance of the pipeline inference method in various indicators, all pipeline solutions with a system throughput greater than 20 were selected from Figure 9 for comprehensive evaluation. The results are as Figure 12As shown, although A1 lags behind B5 and B6 in terms of system throughput and energy consumption, A1 performs better in terms of memory load balancing. Although B1 and B2 perform excellently in terms of memory load balancing, they lag behind A1 in terms of system throughput and device energy consumption. Generally speaking, the pipeline inference method provides a more balanced performance in all metrics because sub-models with higher computational requirements are usually assigned to devices with stronger computational capabilities, and devices with stronger computational capabilities usually have richer memory resources. Therefore, while maintaining a good inter-device memory load balance, A1 can also achieve a high throughput and low energy consumption;
[0155] Selected from Figure 11 pipeline solutions with a standard deviation of memory load less than 0.01 for comprehensive comparison. As Figure 13 shown, it can be seen that in terms of system throughput and device energy consumption, the pipeline inference method is superior to most NSGA2 solutions. This advantage stems from considering memory load balancing, system throughput, and device energy consumption. Those solutions that are similar to or better than A1 and A2 in terms of system throughput and energy consumption also exhibit similar standard deviations of memory load. Therefore, the pipeline inference method can achieve a high system throughput and low device energy consumption while maintaining an inter-device memory load balance.
[0156] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0157] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not elaborate on all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A distributed CNN inference method on a heterogeneous edge device cluster, characterized in that: The following steps are involved: Step S1: train the CNN model using parallel features; Step S2: Split and distribute the trained CNN model to form pipeline reasoning; Step S3: Evaluate the throughput, memory load, and energy consumption of different pipelines based on the available resources of the allocated CNN model; Step S4: performing a segmentation operation on the pre-trained CNN model to form a pipeline solution; The step S4 includes the following sub-steps: Step S41, obtaining summary information of each level model by analyzing the trained CNN model; Step S42, dividing the CNN model into multiple parts, calculating the resource allocation scheme by dividing the parts, and calculating the number of resource allocation schemes according to the number of divisions; Step S43: Divide the CNN model into j parts, distribute them to j devices, select computing resources on each device for distribution, and achieve memory load balancing; Step S44, evaluating each pipeline solution by counting the system throughput, the memory load deviation of each device, and the maximum energy consumption of a single device when each pipeline solution is executed.
2. According to the method of distributed CNN inference on a heterogeneous edge device cluster in claim 1, it is characterized in that: In step S1, GPipe is used to divide the model into multiple segments, each segment is assigned to a different GPU, and the input data is divided into small micro-batches so that each GPU can process different batches of training data at the same time, and the large model is distributedly trained across multiple devices in a pipeline parallel manner.
3. According to the method of distributed CNN inference on a heterogeneous edge device cluster in claim 1, it is characterized in that: The step S2 includes the following steps: Step S21, model segmentation: vertically segment the trained CNN model into multiple sub-models, and deploy each sub-model on the computing resources of different edge devices. When receiving multiple requests, each edge device executes each sub-model in turn and synchronizes and communicates the intermediate inference results, thereby forming a pipeline inference mode; Step S22, computing resource allocation: according to the computing power of a resource, a sub-model with matching computing capacity is allocated to it, so that the computing delay between each stage of the pipeline is consistent.
4. The distributed CNN inference method on a heterogeneous edge device cluster according to claim 1, characterized in that: The step S3 comprises the following steps: Step S31, resource allocation of CNN model: suppose a CNN model has n available computing resources that can be allocated, the computing resource set is represented by R = [r1, r2 ..., rn], rn represents the nth available computing resource, the computing power set corresponding to each resource is represented by c = [c1, c2 ..., cn], cn represents the computing power corresponding to the nth resource, the CNN model is vertically divided into n parts and forms an n-stage pipeline, the pipeline set is represented by s = [s1, s2 ..., sn], if the average time for processing a picture in each stage corresponds to T = [t1, t2 ..., tn], there are a total of j available devices represented by D = [d1, d2 ..., dj], and the total memory capacity of each device is represented by M = [M1, M2 ..., Mj], each device corresponds to multiple resources, which can be allocated to different stages of the pipeline at the same time; Step S32: select a picture to complete the pipeline stage with the highest processing delay, and obtain the number of images processed per second in this stage as the system throughput Tsystem; Step S33: The memory usage Mi of the pipeline stage is divided into three parts, among which the first part is the memory usage used to store the weights, parameters, and biases of the sub-models in this stage when performing inference. The second part is the memory usage used to store intermediate results during CNN layer inference. The third part is the memory usage for receiving intermediate inference results from other pipeline stages So the memory usage Mi of the pipeline stage is expressed as: The computing resources used by different pipeline stages are the same device. If device dj has k computing resources allocated to the pipeline stage, the set of allocated computing resources is The total memory usage of this device The formula is as follows: Statistics are given for the total memory usage of all devices participating in pipeline reasoning under the corresponding pipeline scheme, and the memory utilization of a single device is calculated. The calculation formula for the memory utilization Uj of a single device j is as follows: According to the memory utilization Uj of each device, calculate the average memory utilization of all devices The calculation formula is as follows: Based on average memory utilization Calculate the memory load deviation σU under the corresponding pipeline solution. The specific calculation formula is as follows: Step S34: The energy consumed by each pipeline stage is concentrated in the process of executing the CNN sub-model inference and the process of communicating intermediate data with other pipeline stages. If the energy consumption of executing inference in one pipeline stage is The energy consumption for data communication with other stages is The total energy consumption Ei of each pipeline stage is as follows: If one of the devices dj has k computing resources allocated to the pipeline stage set Sj = {sj1, sj2..., sjk}, the total energy consumption of this device is It is expressed as: The energy consumed by all devices when executing pipeline inference is counted, and the maximum energy consumption Emax of a single device is selected as the evaluation indicator.
5. The distributed CNN inference method on a heterogeneous edge device cluster according to claim 1, characterized in that: The step S41, given a pre-trained CNN model, obtains relevant information of each layer and the overall summary information of the model by analyzing the model, the overall summary information of the model includes the layer name, layer type, number of multiplication and accumulation operations during forward propagation, memory usage of the layer, number of layer parameters, input and output shapes, total number of multiplication and accumulation operations corresponding to all layers of the model, total memory usage when the model layer is executed and the total number of parameters of the model.
6. The distributed CNN inference method on a heterogeneous edge device cluster according to claim 1, characterized in that: The step S42, given a CNN model M and n available resources, divide the model into multiple parts and allocate a different resource to each part; It is divided into n parts to form an n-stage pipeline; The model is divided into n parts, with n allocation methods in total; if the model is divided into p parts, the total number of resource allocation methods solutions(n,p) is calculated as follows: For different numbers of splits, the total number of resource allocation solutions (n, p) that the model has, the total number of pipeline solutions is calculated as follows:
7. The distributed CNN inference method on a heterogeneous edge device cluster according to claim 1, characterized in that: In step S43, when the model is segmented based on the memory amount, the total memory usage of the CNN model is assumed to be Mtotal, and there are j available devices in total. After the model is segmented into j sub-models, each sub-model is assigned to j devices respectively, and the memory usage on each device is respectively Then we have: In order to balance the load of each device, the following formula should be ensured: Therefore, the optimization goal is to divide the model into j parts and make the above formula 1 and formula 2 valid at the same time.
8. The distributed CNN inference method on a heterogeneous edge device cluster according to claim 1, characterized in that: In step S44, after finding the pipeline scheme by the above method, the AutoDiCE framework is used to segment the model according to the corresponding pipeline scheme and deploy it to each device for actual testing, and each pipeline scheme is evaluated by counting the system throughput Tsystem, the memory load deviation σU of each device, and the maximum energy consumption Emax of a single device when each pipeline scheme is executed; After the test is completed, each pipeline solution corresponds to three indicators: Tsystem, σU, and Emax. The three indicators are normalized to [0, 1], and then the weighted comprehensive score of each solution is calculated. The calculation formula is as follows: score=αTsystem+βσU+γEmax, (α+β+γ=1).