A Method for Accelerating Inference of Heterogeneous Processors in Terminal Devices under Temperature Constraints

By building a dynamic frequency setting model under temperature constraints and a single-layer particle size parallel method in smart terminal devices, optimizing the working frequency and computing task allocation of heterogeneous processors, the problem that smart terminal devices are difficult to meet the delay and accuracy requirements of deep neural network inference tasks in industrial production environments is solved, and the inference speed is improved and the equipment is stable operation.

CN114117918BActive Publication Date: 2025-06-13SOUTHEAST UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111426929.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-28
Publication Date
2025-06-13
Estimated Expiration
2041-11-28

AI Technical Summary

Technical Problem

Due to insufficient computing performance, existing smart terminal devices are difficult to meet the latency and accuracy requirements of deep neural network inference tasks in industrial production environments, and the ambient temperature has a significant impact on equipment performance.

Method used

A method for inferring the inference acceleration of the heterogeneous processor of terminal devices under temperature constraints is proposed. By constructing a dynamic frequency setting model, the working frequency of the heterogeneous processor is adjusted, and a single-layer particle size parallel method and a temperature-aware dynamic frequency algorithm are used to optimize the calculation task allocation of deep neural networks to achieve the improvement of inference speed and stable equipment operation.

Benefits of technology

It effectively improves the inference speed of the deep neural network of terminal devices, reduces the inference delay, and improves its application capabilities in industrial production environments while ensuring the stable operation of the equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114117918B_ABST
    Figure CN114117918B_ABST
Patent Text Reader

Abstract

The present invention provides a method for accelerating inference of heterogeneous processors in a terminal device under temperature constraints. Aiming at intelligent terminal devices equipped with multiple heterogeneous processors in an industrial production environment, it solves the problem of low inference efficiency of terminal devices caused by inter-layer heterogeneity, processor heterogeneity, and environmental temperature in deep neural networks. The present invention first considers the environmental temperature of industrial production and the power of the terminal device processor, establishes a dynamic frequency model of the terminal device under temperature constraints, and uses a temperature-aware dynamic frequency algorithm to set the device frequency; then, according to the calculation methods and structural characteristics of different layers in the deep neural network, a single-layer parallel method for the deep neural network is designed; finally, using the heterogeneous processors in the terminal device, a single-layer computing task allocation method for the deep neural network oriented to heterogeneous processors is designed, ensuring low latency and robustness in the collaborative inference of heterogeneous processors in the terminal device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of edge computing and industrial Internet, and particularly relates to a method for accelerating heterogeneous processor inference of terminal devices under temperature constraints. Background Art

[0002] With the rapid development of hardware technology, currently, a variety of intelligent terminal devices such as smart phones, smart cameras, wearable devices, and drones have been widely used in various fields of people's production and life. These intelligent devices are often equipped with sensors such as cameras and radars, as well as various processors such as central processors and graphics processors. Some devices even have processors for specific field calculations, such as the acceleration processor developed by AMD for accelerating image processing and the neural network processor developed by Google for accelerating artificial intelligence algorithms. These processors have strong computing capabilities and can analyze and calculate the collected data after the sensors collect information such as images and sounds. However, on the one hand, the number of layers of deep neural networks is continuously deepening, and the amount of data and calculations is continuously increasing; on the other hand, the vast majority of widely used intelligent terminal devices are limited by power consumption, processor performance, etc., and do not have the powerful computing capabilities of traditional computers and servers used for deep neural network training and inference. Therefore, the training and inference delays of intelligent terminal devices are continuously increasing. For some latency-sensitive tasks, the traditional computing methods of intelligent terminal devices can no longer meet the latency and accuracy requirements of the tasks.

[0003] A deep neural network includes an input layer, an output layer, and several hidden layers. After the input layer converts inputs such as pictures and texts into data forms such as vectors, the intermediate hidden layers perform different types of calculations, and finally, results such as picture classification and text recognition are output through the output layer. Deep neural network inference has the characteristics of a large number of parameters, a large amount of calculations, and high requirements for response speed. Due to performance limitations, existing intelligent terminal devices often have difficulty meeting the latency requirements of deep neural network inference tasks. To improve the inference speed of intelligent terminal devices, the traditional method is to use the cloud computing mode to execute tasks, that is, to send data such as images collected by the terminal device to the cloud computing center, use the powerful computing capabilities of the cloud computing center to complete the inference task, and send the result back to the terminal device. Edge computing is a hierarchical distributed computing architecture. By deploying a certain number of edge devices with stronger computing and storage capabilities than intelligent terminal devices at the network edge, it can provide communication, storage, and computing resources for intelligent terminal devices in the network to reduce the deep neural network inference delay of intelligent terminal devices and make better use of network resources.

[0004] In recent years, with the development of the industrial Internet, deep neural networks have gradually been used in industrial production environments to assist production. For example, during the product production process, sensors such as cameras are used to collect product information (images, vibration signals, etc.) on the production line, and these information are analyzed through neural network inference, enabling functions such as product surface detection and system fault diagnosis. Since industrial production environments often have the characteristics of high temperature and high risk, using deep neural network applications to replace manual detection can not only improve industrial production efficiency but also protect the safety of workers, which is of great significance. With the rapid development of intelligent terminal devices such as intelligent cameras, deep neural network applications in industrial production environments are usually deployed in these intelligent terminal devices. However, on the one hand, existing intelligent terminal devices often have insufficient overall computing performance and are often difficult to meet the requirements of deep neural network applications for inference accuracy and latency; on the other hand, the complexity of industrial production environments also places many restrictions on the performance of intelligent terminal devices. Therefore, the relatively high requirements for accuracy and inference latency and the relatively weak computing power of terminal devices restricted by industrial production environments pose new challenges to traditional deep neural network inference methods based on intelligent terminal devices. Summary of the Invention

[0005] Aiming at the three deficiencies of traditional terminal inference methods that do not consider that different types of layers have different computing methods and data structures, ignore the different performances and computing characteristics of heterogeneous processors, and the influence of the environment where intelligent terminal devices are located on the processors in the devices, the present invention proposes a method for accelerating inference of heterogeneous processors in terminal devices for single-layer granularity neural network computing. Under the condition of considering the temperature limit of intelligent terminal devices, the working frequency of heterogeneous processors in the device is adjusted, and a method of using heterogeneous processors for fine-grained cooperative acceleration of single-layer deep neural networks is used to improve the inference speed of terminal devices and ensure the stable operation of terminal devices.

[0006] To achieve the above object, the technical solution of the present invention is as follows:

[0007] A method for accelerating inference of heterogeneous processors in terminal devices under temperature constraints, comprising the following steps:

[0008] Step 1: Construct a dynamic frequency setting model for terminal devices under temperature constraints, analyze the relationship between power consumption control and clock frequency constraints of terminal devices in industrial production environments, and model by actually measuring the environmental temperature and device power consumption;

[0009] Step 2: Select the parallel method for single-layer granularity of the neural network, characterize the computational amount of each layer of the deep neural network, analyze the data structures and computational amounts of three common layers, namely the convolutional layer, pooling layer, and fully connected layer, and combine the computing methods and structural characteristics of heterogeneous processors to estimate the computing delay of each layer on each processor, so as to determine the single-layer parallel method of the deep neural network;

[0010] Step 3: Based on Steps 1 and 2, provide a single-layer granularity calculation load division for the deep neural network inference process, specifically including:

[0011] First, considering the high-temperature environment in industrial production, according to the dynamic frequency model of the terminal device under temperature constraints established in Step 1, set the device processor frequency, so as to limit the device power consumption and keep the device temperature within a reasonable operating range.

[0012] After that, according to the single-layer parallel method of the deep neural network designed in Step 2, select the single-layer granularity parallel modes of different layers and their combinations. The optional modes are data parallelism and model parallelism. Further consider the calculation time caused by merging the output results of two processors for each layer, that is, the additional delay after parallelization.

[0013] Finally, implement the single-layer calculation task allocation of the heterogeneous processor for the deep neural network. The goal of the task allocation is to minimize the total inference delay of the terminal device; transform the problem of accelerating the inference of the heterogeneous processor of the terminal device under temperature constraints into an optimization problem subject to certain constraints, and use the temperature-aware dynamic frequency algorithm TADF and the single-layer heterogeneous processor load allocation algorithm HSWD algorithm to perform load allocation for the calculation tasks of each layer, so that the inference delay of each layer is the lowest.

[0014] Furthermore, when constructing the dynamic frequency setting model of the terminal device under temperature constraints in Step 1, based on the modeling key parameters, the frequency f of the heterogeneous processor in the terminal device processor , the power consumption P of the heterogeneous processor processor , obtain the total power consumption P of the terminal device; based on the modeling key parameters, the ambient temperature T eno (t) at time t and the device temperature T(t), obtain the device steady-state operating temperature T(∞); the floating-point operation speed of the heterogeneous processor and the device steady-state operating temperature follow certain constraints.

[0015] Furthermore, Step 1 specifically includes the following process:

[0016] First, model the characteristics of the intelligent terminal device. For an intelligent terminal device D equipped with a CPU and a GPU, the frequency of the heterogeneous processor in the device is determined by the processor clock frequency f clock and the number of floating-point operations per clock cycle n processor , that is and The processor power consumption is related to the clock frequency of the processor, where P processor = Ψ(f clock ) 3 , Ψ (W / ((cycle / s)) 3 ) is a coefficient determined by the processor architecture. Therefore, the processor power consumption is expressed as follows:

[0017]

[0018] Among them, γ C = Ψ C / (n C ) 3 , γ G = Ψ G / (n G ) 3 ;

[0019] In addition, the standby power consumption of the device estimates the relationship between the standby power consumption of the device, the environment, and the device voltage with high precision through a linear model, that is, P idle = V(β 1 T eno + β 0 ), the coefficients β 1 and β 0 are related to the performance of the device. Therefore, the total power consumption of the terminal device is:

[0020] P = P idle + P C + P G

[0021] = V(β 1 T eno + β 0 ) + γ C (f C ) 3 + γ G (f G ) 3

[0022] Due to the influence of the environmental temperature T eno (t) and the self-heat power consumption factor of the device, when the processor working frequency and the environmental temperature of the device remain stable, the stable temperature model T(t→∞) will be reached after the device works continuously for a long time; according to the thermal circuit model, the temperature of the device is expressed as a function related to the power consumption of the device. When the device D runs at power P, the temperature of the device at time t is expressed as:

[0023]

[0024] Among them, R (℃ / W) and C (J / K) represent thermal resistance and heat capacity respectively;

[0025] From this, when t→∞, the stable operating temperature of the device is:

[0026] T(∞) = T eno (∞) + P·R

[0027] = T eno(∞)+(P idle +P C +P G )·R

[0028] =(1+VRβ 1 )·T eno (∞)+Rγ C ·(f C ) 3 +Rγ G ·(f G ) 3 +VRβ 0

[0029] =α 1 ·T eno (∞)+α 2 ·(f C ) 3 +α 3 (f G ) 3 +α 0

[0030] Keep the temperature of the device always lower than its maximum stable operating temperature T max ; Correspondingly, the floating-point operation speeds of the CPU and GPU in device D should comply with the constraint:

[0031] α 2 ·(f C ) 3 +α 3 ·(f G ) 3 ≤T max -α 1 T eno (∞)-α 0 。

[0032] Furthermore, in the second step, the computational workload W of common layers in the deep neural network is characterized, and the computational workload of each layer on each heterogeneous processor is analyzed separately Combined with the data structure and computational characteristics of each layer, data parallelism or model parallelism at the single-layer granularity is selected.

[0033] Furthermore, the second step specifically includes the following process:

[0034] First, model the deep neural network, and calculate the floating-point computational workloads W of the convolutional layer, pooling layer, and fully connected layer respectively;

[0035] Secondly, consider the single-layer granularity parallel modes of the above three network layers, namely data parallelism and model parallelism. For convolutional layers and fully connected layers, use the model parallelism method to divide the convolutional kernels; for the calculation of pooling layers, use the data parallelism method, and divide the input matrix by channels to achieve parallelism.

[0036] Further, when implementing the allocation of the deep neural network single-layer calculation tasks for heterogeneous processors in step three, analyze the performance of the heterogeneous processors in the terminal device and the structural characteristics of the single layer in the deep neural network, divide the calculation tasks of the single layer in the deep neural network, and merge the results after parallel calculation; First, characterize the computational workload W C , W G of each layer of the neural network on the CPU and GPU, and then analyze the inference latency on the CPU and GPU Thus, the latency of using the parallel method to execute layer L i can be expressed as The total latency consists of two parts: the maximum running time during the parallel process and the calculation result merging time. Finally, calculate the shortest execution latency t i of layer L i :

[0037]

[0038] Among them, is the additional latency caused by merging the output results of the two processors when using the parallel method for inference, and are the times required to execute layer l i using only the CPU and GPU respectively.

[0039] Further, the solution objective of the optimization problem is the minimum latency that conforms to the constraints of the mathematical model set by the problem:

[0040]

[0041] where t i represents the execution time of the i-th layer in the neural network, W C , W G represent the total computational workloads on the CPU and GPU respectively, f C , f G represent the stable working frequencies of the CPU and GPU respectively, and represent the maximum floating-point operation speeds of the CPU and GPU in D respectively, and P max represents the maximum power consumption of D.

[0042] The beneficial effects of the present invention are:

[0043] The method for accelerating inference of heterogeneous processors in a terminal device under temperature constraints provided by the present invention utilizes fine-grained cooperation of heterogeneous processors to accelerate a single-layer deep neural network, improving the inference speed of the terminal device and ensuring the stable operation of the terminal device. The method of the present invention helps to construct an optimization system for deep neural network inference based on edge-cloud collaboration, and can achieve deep neural network inference with low latency and high robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a schematic flowchart of the method for accelerating inference of heterogeneous processors in a terminal device under temperature constraints provided by the present invention.

[0045] Figure 2 For the temperature of Jetson Nano and the processor clock frequency in the MaxN mode at room temperature of 40°C;

[0046] Figure 3 It is the single-layer granularity parallel mode described in the present invention;

[0047] Figure 4 It is the system working flowchart. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The following will describe in detail the technical solutions provided by the present invention in combination with specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0049] The applicable scope of the present invention includes, but is not limited to, intelligent terminal devices equipped with multiple heterogeneous processors in an industrial production environment, and can also be implemented in other scenarios with temperature constraints.

[0050] The present invention provides a method for accelerating inference of heterogeneous processors in a terminal device under temperature constraints, mainly aiming at intelligent terminal devices equipped with multiple heterogeneous processors in an industrial production environment, and solving the problem of low inference efficiency of terminal devices caused by layer heterogeneity, processor heterogeneity and environmental temperature in deep neural networks. The core mechanism of the present invention mainly includes three parts: dynamic frequency setting of terminal devices under temperature constraints, single-layer granularity parallel method for deep neural networks, and single-layer granularity calculation load division in the deep neural network inference process. The present invention first considers the environmental temperature of industrial production and the power of the terminal device processor, establishes a dynamic frequency model of the terminal device under temperature constraints, and uses a temperature-aware dynamic frequency algorithm to set the device frequency. Then, according to the calculation methods and structural characteristics of different layers in the deep neural network, a single-layer parallel method for the deep neural network is designed. Finally, using the heterogeneous processors in the terminal device, a single-layer calculation task allocation method for the deep neural network facing heterogeneous processors is designed, ensuring low latency and robustness in the collaborative inference of heterogeneous processors in the terminal device.

[0051] The heterogeneous processor inference acceleration method for terminal devices under temperature constraints provided by the present invention has a process as Figure 1 shown, including the following steps:

[0052] Step 1, dynamic frequency setting of terminal devices under temperature constraints. According to the ambient temperature and processor power, establish a dynamic frequency model for terminal devices under temperature constraints, analyze the relationship between power consumption control and clock frequency constraints of terminal devices in industrial production environments, and provide a basis for the selection of the neural network single-layer granularity parallel mode in the subsequent step 2 and the calculation load division in step 3.

[0053] First, model the intelligent terminal device. For an intelligent terminal device D equipped with a CPU and a GPU, the relevant attributes in this device are shown in Table 1.

[0054] Table 1 Device attribute symbols

[0055]

[0056] Let the clock frequencies (cycles / second) of the CPU and GPU in D be and respectively. The number of floating-point operations per cycle is n C and n G . respectively. Then, the floating-point operation speeds (times / second) of the CPU and GPU can be expressed as:

[0057]

[0058] a. Power consumption calculation

[0059] The power consumption of the intelligent terminal device consists of the standby power consumption of the device itself and the power consumption generated when the processor in the device runs. The processor power consumption is related to the clock frequency of the processor, and their corresponding relationship can be expressed as: P processor = Ψ(f clock ) 3 , where P processor represents the power consumption of the processor, Ψ (W / ((cycle / s)) 3 ) is a coefficient determined by the processor architecture, and f clock is the clock frequency of the processor. Then, the power consumptions of the CPU and GPU in D are:

[0060]

[0061] where γ C = Ψ C / (n C ), Ψ 3 = Ψ G / (n G / (n G )3 . It is not difficult to conclude that the power consumption of the processor is positively correlated with the floating-point operation speed of the processor. In addition, the standby power consumption of the device can accurately estimate the relationship between the standby power consumption of the device, the environment, and the device voltage through a linear model, that is, P idle = V(β 1 T eno + β 0 ), and the coefficients β 1 and β 0 are related to the performance of the device. The total power consumption of the device including standby power consumption is:

[0062] P = P idle + P C + P G = V(β 1 T eno + β 0 ) + γ C (f C ) 3 + ΥG(f G ) 3 .

[0063] Due to the influence of the environmental temperature T eno (t) and the self-heat power consumption factor of the device, it is necessary to describe the stable temperature model T(t→∞) that the device will reach after long-term continuous operation when the processor working frequency and environmental temperature of the device remain stable.

[0064] b. Temperature-Frequency Constraint Analysis

[0065] For device D, its power consumption is affected by the floating-point operation speeds of the CPU and GPU in the device and the environmental temperature. According to the thermal circuit model, the temperature of the device can be expressed as a function related to the device power consumption. When device D operates at power P, the temperature of the device at time t can be expressed as:

[0066]

[0067] where R (℃ / W) and C (J / K) represent thermal resistance and heat capacity respectively.

[0068] Then, as the device continues to operate, when t→∞, the temperature of the device can be expressed as:

[0069]

[0070] The constants in the formula are all abbreviated with the symbol α. That is, when the processor working frequency and environmental temperature of the device remain stable, the device will reach a stable temperature T(∞) after long-term continuous operation. Therefore, to enable the device to operate stably for a long time, the temperature of the device should always be lower than its maximum stable operating temperature T maxAccordingly, the floating-point operation speeds of the CPU and GPU in device D should comply with the constraints:

[0071] α 2 ·(f C ) 3 +α 3 ·(f G ) 3 ≤T max -α 1 ·T eno (∞)-α 0 .

[0072] like Figure 2 As shown in the figure, the relationship between device temperature and processor clock frequency in MaxN mode (full power mode W=10) at room temperature of 40℃. This experiment proves that ambient temperature has a very significant impact on smart terminal devices. When the terminal device runs at a high ambient temperature for a period of time, the device temperature may exceed the maximum normal operating temperature of the device. At this excessively high device temperature, the processor performance in the device will also be significantly affected.

[0073] Step 2: Single-layer granularity parallel method for deep neural networks.

[0074] c. Analysis and selection of single-layer granularity parallel methods for neural networks

[0075] Since trained models are often used to perform inference tasks in smart terminal devices, the present invention mainly focuses on the inference process of deep neural networks in smart terminal devices. Among various types of deep neural networks, the present invention mainly focuses on convolutional neural networks because they are the most widely used in smart terminal devices. It should be noted that this step c is also applicable to other neural network layers that can use data parallelism or model parallelism.

[0076] The input of each layer in the convolutional neural network is an H i *W i *D i The matrix, where H i and W i Respectively represent the length and width of the matrix, D i Represents the depth of the matrix. Similarly, the output of each layer is H o *W o *D o The matrix, where H o and W o Respectively represent the length and width of the matrix, D o Indicates the depth of the matrix. In addition to the input layer and the output layer, there are several convolutional layers, pooling layers, and fully connected layers in the hidden layer.

[0077] First, model the deep neural network and calculate the floating-point computation amounts \(W\) of the convolutional layer, pooling layer, and fully connected layer respectively. Analyze the computation amounts of each layer on each heterogeneous processor separately.

[0078] Secondly, consider the single-layer granularity parallel mode of the above three network layers. Combine the data structure and computation characteristics of each layer to estimate the computation latency of each layer on each processor, and select data parallelism or model parallelism at the single-layer granularity. The purpose of this step is to solve the minimum latency for the single-layer granularity of the neural network to execute on the terminal device, providing a basis for the task partitioning and allocation in step three.

[0079] For the convolutional layer and the fully connected layer, since the computation method of these two types of layers is to perform a convolution operation on the input data through a convolution kernel, the convolution kernel can be partitioned using the model parallelism method. As Figure 3 (a) shows, for a convolutional layer or a fully connected layer, first partition the convolution kernels within the layer according to the performance of the heterogeneous processor by the number of channels. Then, input the input data of this layer into the CPU and GPU completely respectively. The CPU and GPU respectively use the allocated convolution kernels to perform convolution operations on all input data and output the results. After the outputs of the two processors are combined, they become the output of this layer and the input of the next layer. Since the intelligent terminal device uses an integrated architecture with shared memory, the data transmission during the computation does not cause additional transmission latency. And during the process of partitioning the convolution kernels by channels, there is no overlapping part between the convolution kernels allocated to the two processors. Therefore, the additional latency of this layer only comes from the computation latency caused by combining the output results of the two processors.

[0080] For the pooling layer, since the pooling layer only performs average pooling or max pooling operations on the input data according to a preset window size, and the pooling layer itself has no data such as convolution kernels and is not suitable for the model parallelism method, the data parallelism method is adopted for the computation of the pooling layer, and parallelism is achieved by partitioning the input matrix by channels. As Figure 3 (b) shows, after the input data is partitioned by channels, it is respectively transmitted to the CPU and GPU. Then, the CPU and GPU respectively perform pooling operations on the input data and obtain the corresponding output results. After the data in the two processors are combined, the output data of this pooling layer is obtained. Similar to the convolutional layer and the fully connected layer, the additional latency only comes from the computation latency caused by combining the output data of the two processors.

[0081] In addition, there are also layers such as activation functions and SoftMax in the convolutional neural network. However, since the convolutional layer, pooling layer, and fully connected layer mainly occupy the computation amount and data amount in the convolutional neural network, in the subsequent implementation step d of the present invention, the design of the computation load allocation method for these three layers is mainly considered.

[0082] Step 3: Based on Steps 1 and 2, provide a single-layer granularity calculation load division for the deep neural network inference process. As Figure 4 shown, first, considering the high-temperature environment in industrial production, sense the environmental temperature, and according to the dynamic frequency model of the terminal device under temperature constraints established in Step 1, set the power consumption of the device processor to keep the device temperature within a reasonable operating range; then, according to the single-layer parallel method designed in Step 2, select the single-layer granularity parallel mode for different layers and their combinations. The optional modes are data parallelism and model parallelism. It is necessary to further consider the calculation time caused by merging the output results of the two processors for each layer, that is, the additional delay after parallelization; finally, implement the single-layer calculation task allocation of heterogeneous processors. The goal of the task allocation is to minimize the total inference delay of the terminal device.

[0083] d. Single-layer granularity calculation load division for the deep neural network inference process

[0084] According to the single-layer granularity parallel method proposed in Implementation Step c, in this step, the calculation tasks of the single layers in the deep neural network will be divided according to the performance of the heterogeneous processors in the terminal device and the structural characteristics of the single layers in the deep neural network. First, model the deep neural network. For a deep neural network G, the parameters of the device are shown in Table 2.

[0085] Table 2 Deep neural network related attributes

[0086]

[0087] For a deep neural network G with a total of n layers, L = {l 1 , l 2 , l n} is the set of each layer in G, where l i ∈ L is the i-th layer in G. W = {W 1 , W 2 , …, W n} is the amount of computation for each layer in G, where W i ∈ W is the amount of computation for the i-th layer in G. and respectively represent the amount of computation for each layer of the CPU and GPU when performing the inference task of the deep neural network G:

[0088]

[0089] Obviously, when only using the CPU to perform inference on layer l i , Similarly, when only using the GPU to perform inference on layer l i , Therefore, when the floating-point operation speeds of the CPU and GPU are f C and f G respectively, when only a single processor is used to perform inference on layer l i , the latencies of using the CPU and GPU are and respectively. For layer l i , when the CPU and GPU are used simultaneously to perform calculations on this layer, if this layer is a convolutional layer or a fully connected layer with the number of convolutional kernel channels being N, the CPU and GPU will respectively use N C and N G convolutional kernels of channels to perform convolution calculations on the input data, where N = N C + N G . Then, the computational amounts of the CPU and GPU on this layer can be respectively expressed as:

[0090]

[0091] If this layer is a pooling layer with the number of input data channels being M, when the CPU and GPU are used simultaneously to perform calculations on this layer, the CPU and GPU will respectively perform pooling operations on the input data of M C and M G , where M = M C + M G . At this time, the computational amounts of the CPU and GPU on this layer are respectively:

[0092]

[0093] Then, the latency of using the parallel method to execute layer l i can be expressed as:

[0094]

[0095] The total latency consists of two parts: the maximum running time during the parallel process and the calculation result merging time. And the shortest execution time t i of layer l i in G can be expressed as:

[0096]

[0097] Among them, is the additional latency caused by merging the output results of the two processors during parallel inference, and are the times required to execute layer l i using only the CPU and GPU respectively.

[0098] a. Solving the problem of accelerating heterogeneous processor inference in a terminal device under temperature constraints

[0099] According to the method proposed in implementation steps a - d, the problem of accelerating heterogeneous processor inference of a terminal device under temperature constraints can be transformed into an optimization problem subject to certain constraints:

[0100]

[0101] Among them, and respectively represent the maximum floating - point operation speeds of the CPU and GPU in D, while P max represents the maximum power consumption of D. That is, under the power consumption and temperature constraints of the device, by setting the floating - point operation speeds of the CPU and GPU, the computing power of the device is maximized, and the computing load of each layer in the neural network is reasonably allocated to minimize the single - layer inference latency of each layer, so as to obtain the lowest inference latency of the deep neural network of the terminal device and ensure the long - term stable operation of the device.

[0102] To achieve the optimization goal in the above - mentioned optimization problem, first, based on the maximum operating temperature limit of the device and the maximum floating - point operation speed limits of the CPU and GPU, the floating - point operation speeds of the CPU and GPU are set to maximize the computing power of the device. The specific algorithm is as follows.

[0103]

[0104]

[0105] In Algorithm 1, the complexity of determining whether the maximum floating - point operation speed meets the constraints is The complexity of adjusting the floating - point operation speed of the processor is Therefore, the complexity of Algorithm 1 is

[0106] After determining the floating - point operation speeds of the CPU and GPU, it is necessary to allocate the computing load according to the computing performance of the CPU and GPU and the computing volume and structural characteristics of each layer in the deep neural network. The HSWD algorithm proposed in the present invention sets the computing load allocation of each layer according to the floating - point operation speeds of the CPU and GPU and the computing volume of each layer of the deep neural network D and The specific process of this algorithm is shown in Algorithm 2.

[0107]

[0108]

[0109] By traversing the entire model, Algorithm 2 can obtain the optimal computing load allocation method for each layer in G, thereby obtaining the lowest inference latency of this neural network. Since Algorithm 2 traverses the entire deep neural network layer by layer, its algorithm complexity is

[0110] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.

Claims

1. An inference acceleration method for heterogeneous processors of terminal devices under temperature constraints, characterized in that, it includes the following steps: Step 1: Build a dynamic frequency setting model for the terminal device under temperature constraints, analyze the relationship between power consumption control and clock frequency constraints of the terminal device in the industrial production environment, and model by actually measuring the environmental temperature and device power consumption; when building a dynamic frequency setting model for the terminal device under temperature constraints, based on the key parameters of the model, the frequency f of the heterogeneous processor in the terminal device processor and the power consumption P of the heterogeneous processor processor , the total power consumption P of the terminal device is obtained; based on the key parameters of the model, the environmental temperature T eno (t) at time t and the device temperature T(t), the steady-state operating temperature T(∞) of the device is obtained; the floating-point operation speed of the heterogeneous processor and the steady-state operating temperature of the device follow certain constraints; Step 2: Selection of the parallel mode at the single-layer granularity of the neural network, characterizing the computational workload of each layer of the deep neural network, analyzing the data structures and computational workloads of three common layers, namely the convolutional layer, pooling layer, and fully connected layer, and combining the computational methods and structural characteristics of the heterogeneous processors to estimate the computational latency of each layer on each processor, so as to determine the single-layer parallel method of the deep neural network; Step 3: Based on Steps 1 and 2, provide the division of the single-layer granularity computational load in the inference process of the deep neural network, specifically including: First, considering the high-temperature environment of industrial production, according to the dynamic frequency model of the terminal device under temperature constraints established in Step 1, set the frequency of the device processor, so as to limit the power consumption of the device to keep the temperature of the device within a reasonable operating range; After that, according to the single-layer parallel method of the deep neural network designed in Step 2, select the single-layer granularity parallel modes of different layers and their combinations. The optional modes are data parallelism and model parallelism. Further consider the computational time caused by merging the output results from two processors for each layer, that is, the additional latency after parallelization; Finally, realize the allocation of the single-layer computational tasks of the heterogeneous processors for the deep neural network. The goal of the task allocation is to minimize the total inference latency of the terminal device; transform the problem of accelerating the inference of heterogeneous processors of terminal devices under temperature constraints into an optimization problem subject to certain constraints, and use the temperature-aware dynamic frequency algorithm TADF and the single-layer heterogeneous processor workload distribution algorithm HSWD algorithm to allocate the computational tasks of each layer, so that the inference latency of each layer is the lowest; The temperature-aware dynamic frequency algorithm TADF sets the floating-point operation speeds of the CPU and GPU based on the maximum operating temperature limit of the device and the maximum floating-point operation speeds of the CPU and GPU; the TADF algorithm is as follows: The single-layer heterogeneous processor workload distribution algorithm HSWD algorithm sets the computational workload distribution of each layer according to the floating-point operation speeds of the CPU and GPU and the computational workload of each layer of the deep neural network; the HSWD algorithm is as follows: Among them, f C is the floating-point operation speed of the CPU, and f G is the floating-point operation speed of the GPU. and respectively represent the maximum floating-point operation speeds of the CPU and the GPU. T max is the maximum stable operating temperature, and T eno (∞) is the stable ambient temperature. W i is the computational workload of the i-th layer in the model, and W C is the set of CPU computational workloads for each layer of the model, and W G is the set of GPU computational workloads for each layer of the model. is the CPU computational workload of the i-th layer in the model. is the GPU computational workload of the i-th layer in the model. is the latency of executing layer L i using the parallel method. and are the times required to execute layer L i using only the CPU and the GPU respectively. is the additional time delay caused by merging the output results of the two processors when performing inference using the parallel method. n is the number of floating-point operations, and α 0 , α 1 , α 2 , α 3 are constants in the formula.

2. The inference acceleration method for heterogeneous processors of terminal devices under temperature constraints according to claim 1, characterized in that, the specific process of Step 1 is as follows: First, model the characteristics of the intelligent terminal device. For an intelligent terminal device D equipped with a CPU and a GPU, the frequency of the heterogeneous processors in this device is represented by the processor clock frequency f clock and the number of floating-point operations per clock cycle n processor , that is and is the CPU clock frequency, is the GPU clock frequency, n C is the number of floating-point operations per cycle of the CPU, n G is the number of floating-point operations per cycle of the GPU; the processor power consumption is related to the clock frequency of the processor, where P processor = Ψ(f clock ) 3 , Ψ is a coefficient determined by the processor architecture, so the processor power consumption is expressed as follows: where, Υ C = Ψ C / (n C ) 3 , Υ G = Ψ G / (n G ) 3 ; In addition, the standby power consumption of the device estimates the relationship between the standby power consumption of the device, the environment, and the device voltage with high precision through a linear model, that is, P idle = V(β 1 T eno + β 0 ), the coefficients β 1 and β 0 are related to the performance of the device, and V is the device voltage; therefore, the total power consumption of the terminal device is: P = P idle + P C + P G = V(β 1 T eno + β 0 ) + Υ C (f C ) 3 + Υ G (f G ) 3 Due to the environmental temperature T eno (t) and the influence of the device's own thermal power consumption factor, when the processor working frequency and environmental temperature of the device remain stable, the stable temperature model will be reached after the device works continuously for a long time; according to the thermal circuit model, the temperature of the device is expressed as a function related to the device power consumption. When the device D operates at power P, the temperature of the device at time t is expressed as: Where R and C respectively represent thermal resistance and heat capacity, and T(0) is the initial temperature; From this, it can be obtained that when t→∞, the stable operating temperature of the device is: T(∞) = T eno (∞) + P·R = T eno (∞) + (P idle + P C + P G )·R =(1 + VRβ 1 )·T eno (∞)+RΥ C ·(f C ) 3 +RΥ G ·(f G ) 3 +VRβ 0 = α 1 · T eno (∞)+ α 2 ·(f C ) 3 + α 3 ·(f G ) 3 + α 0 Keep the temperature of the device always lower than its maximum stable operating temperature T max ; correspondingly, the floating-point operation speeds of the CPU and GPU in device D should comply with the constraint: α 2 ·(f C ) 3 +α 3 ·(f G ) 3 ≤T max -α 1 ·T eno (∞)-α 0 。 3. The inference acceleration method for heterogeneous processors of terminal devices under temperature constraints according to claim 1, characterized in that, In step 2, calculate the computational workload W of common layers in the deep neural network, and separately analyze the computational workload of each layer on each heterogeneous processor. Combined with the data structure and computational characteristics of each layer, select data parallelism or model parallelism at the single-layer granularity.

4. The inference acceleration method for heterogeneous processors of terminal devices under temperature constraints according to claim 1, characterized in that, the specific process of Step 2 is as follows: First, model the deep neural network, and calculate the floating-point computational workload W of the convolutional layer, pooling layer, and fully connected layer respectively; Secondly, consider the single-layer granularity parallel modes of the above three network layers, namely data parallelism and model parallelism; for convolutional layers and fully connected layers, use the model parallelism method to divide the convolutional kernels; for the calculation of pooling layers, use the data parallelism method, and achieve parallelism by dividing the input matrix according to channels.

5. The method for accelerating inference of a heterogeneous processor of a terminal device under temperature constraints according to claim 1, characterized in that When realizing the assignment of the deep neural network single-layer computing task for heterogeneous processors in the third step, the performance of the heterogeneous processors in the terminal device and the structural characteristics of a single layer in the deep neural network will be analyzed, the computing tasks of a single layer in the deep neural network will be divided, and the results will be merged after parallel computing; First, characterize the computing amount W of each layer of the neural network on the CPU and GPU C , W G , and then analyze the inference latency on the CPU and GPU So as to execute layer L in parallel i The latency can be expressed as The total latency consists of two parts: the maximum running time and the computing result merging time during the parallel process, and finally calculate the shortest execution latency t of layer L i : i ​ Among them, is the additional latency caused by merging the output results of two processors when performing inference in parallel, and n is the number of floating-point operations.

6. The method for accelerating inference of a heterogeneous processor of a terminal device under temperature constraints according to claim 1 or 5, characterized in that the solution objective of the optimization problem is the minimum latency that conforms to the constraints of the mathematical model set by the problem: α 2 ·(f C ) 3 +α 3 ·(f G ) 3 ≤, F = T max -α 1 ·T eno (∞)-α 0 where W is the set of computational amounts of each layer of the model.