Self-adaptive end-edge cooperation method based on DNN division and feature compression
By adopting an adaptive edge-end collaborative inference framework based on mobile devices and edge devices, dynamically adjusting the segmentation points and feature compression rates of the DNN model, the problems of high DNN inference delay and energy consumption in the prior art are solved, and efficient and real-time inference performance optimization is achieved.
Patent Information
- Application Number
- CN202510154350.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-24
AI Technical Summary
When running deep neural networks (DNNs) on mobile devices and edge devices, it is difficult for the prior art to effectively reduce inference latency and energy consumption while meeting the needs of real-time and efficient computing resource allocation.
Adaptive edge-end collaborative inference framework based on deep reinforcement learning (DRL) is adopted to dynamically adjust the model segmentation points and feature compression ratio through DNN division and feature compression, and optimize the inference performance of DNN models on mobile devices and edge devices.
It realizes efficient inference performance optimization in dynamic network environment, reduces inference delay and energy consumption, and ensures the accuracy of inference results and enhances user privacy protection.
Smart Images

Figure CN120197668A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an adaptive edge - end collaborative inference framework based on deep reinforcement learning (DRL) for optimizing the inference performance of deep neural networks (DNNs) on mobile devices and edge devices. This framework is applicable to application scenarios that require real - time inference and efficient computing resource allocation, such as autonomous driving and augmented reality, etc. Background Art
[0002] With the development of artificial intelligence technology, deep neural networks have been widely used in mobile applications. However, due to the limited computing power and storage resources of mobile devices, the solution of directly deploying DNN models on mobile devices and completing the computing process is difficult to meet the real - time requirements of applications. The traditional solution to solve this problem is to upload tasks to the cloud for processing, but this method has problems of data transmission delay and privacy security. The emergence of edge computing technology provides a new idea to solve these problems, that is, by offloading part of the computing tasks to edge devices to reduce transmission delay and overall latency. However, due to the dynamics and complexity of the real network environment, how to dynamically select a suitable partial offloading scheme in a dynamically changing network environment remains a challenging problem. Summary of the Invention
[0003] In order to overcome the deficiencies of the prior art and achieve efficient adaptive split - computing end - edge collaborative computing, the present invention proposes an adaptive end - edge collaborative method based on DNN partitioning and feature compression, aiming to optimize the inference performance of DNN models on mobile devices and edge devices. This framework combines deep reinforcement learning, model split - computing, and feature compression to ensure that while meeting the accuracy requirements, the inference delay and energy consumption are effectively reduced.
[0004] To achieve the above - mentioned technical tasks, the present invention adopts the following technical solutions:
[0005] An adaptive end - edge collaborative method based on DNN partitioning and feature compression, comprising:
[0006] Step 1: Collect network environment state information s: In an edge computing system composed of a wireless device (WD) and an edge computing server (ECS), the wireless device collects the current network environment state information s, which is the input of the policy model and includes signal quality Q, transmission rate V, channel gain S, and task volume A;
[0007] s = [Q, V, S, A]
[0008] where the task model to be executed is M, and the model contains N main weight layers, including N convOne convolutional layer and N fc fully connected layers;
[0009] Step 2: Input the environmental state information into the policy model π;
[0010] Step 3: Output the probability distribution of the joint actions of the model segmentation point and the compression rate;
[0011] Step 4: The mobile device performs the calculation of the first half of the DNN model;
[0012] Step 5: After several layers of the head model M head are calculated, M head is the model from the first layer to the d-th layer, and the intermediate layer feature data X ∈ R b×c×h×w is obtained. Then, the intermediate layer feature data X is calculated through the channel attention module to obtain the channel attention value A ∈ R b×c×1×1 . The normalized result A norm ∈ R b×c×1×1 after the sigmoid operation of A is used as the channel importance evaluation result I ∈ R b×c×1×1 , where b is the batch number, c is the number of channels, and h and w are the height and width of the feature map;
[0013] Step 6: According to the channel importance result I in Step 5, perform channel reduction on the intermediate layer feature data X; first, sort the channel importance values I to obtain the result I sort = [I1, I2…, I b×c ; then, perform channel reduction on the intermediate feature data according to the decision a obtained from the DRL decision model, where the number of reserved channels is b×c×comp, corresponding to the channel data with higher importance, that is, I selected = [I1, I2…, I b×c×comp ; according to the selected channel index list I selected , extract the corresponding important channel features from the intermediate layer feature data X to obtain the important channel feature tensor, that is, the result X selected after channel reduction;
[0014] Step 7: Perform feature quantization on the result X selected after channel reduction obtained in Step 6 to obtain the final data D comp for communication transmission; perform quantization processing on the reduced feature data X selected , and quantize the data type from float32 to float16 to further reduce the data volume;
[0015] D comp = f comp (X selected );
[0016] Step Eight: Send the compressed intermediate feature data D comp to the edge computing server ECS. The wireless device WD sends the quantized intermediate feature data D comp to the edge computing server ECS via the wireless network;
[0017] D decomp = f decomp (D comp );
[0018] Step Nine: Edge device de - quantization and continued calculation. After receiving the intermediate feature data, the edge device performs de - quantization operations to restore the features to their original size and continues to execute the second half of the DNN model's calculation.
[0019] Result = M tail (D decomp );
[0020] Step Ten: Collect inference performance data: After the inference task is completed, the system collects the inference task latency L and the total energy consumption E during the inference process. The inference task latency L consists of the computing latency on the WD, the transmission latency in the wireless network, and the computing latency on the ECS, which is obtained by detecting the end - to - end latency on the WD. The total energy consumption E is obtained through the WD using the command line;
[0021] Step Eleven: Calculate the reward function value: According to the inference performance data collected in Step Ten, calculate the reward function value according to the reward function formula,
[0022] R = ω L ·R L + ω E ·R E
[0023] where ω L and ω E are the trade - off coefficients of the task latency L and the total energy consumption E respectively, and the calculation methods of R L and R E are as follows:
[0024] R L = -tanh[A(L - B)] + C
[0025] R E = -tanh[A(E - B)] + C
[0026] where tanh represents the hyperbolic tangent function, and A, B, C are empirical parameter settings after a large number of experimental evaluations;
[0027] Step Twelve: Calculate the gradient value: Perform gradient backpropagation according to the reward function value,
[0028]
[0029] where p i is the probability value corresponding to the selected action i, and R bl is the baseline for optimizing the reward value, which refers to the reference value used to evaluate and compare the performance of the algorithm during a given iteration. R bl is determined by statistically averaging the reward values of the 50 iterations forward from the current round t;
[0030] Step Thirteen: Backpropagate the calculated gradient values to implement gradient backpropagation and optimize the DRL model parameters; after each complete inference, adjust the parameters of the DRL model according to the performance data during the inference process collected in Step Ten, and optimize the policy model parameters for the offloading strategy selection to continuously adapt to the changing network environment, improve the system inference performance and environmental adaptability. Subsequently, the model will be continuously iteratively updated in this update manner to optimize the performance.
[0031] Furthermore, in Step Two, the policy model π used is a DRL model. First, the collected environmental state information s is input into the deep reinforcement learning DRL model based on policy gradients. This model adopts a fully connected network structure, including an input layer, an output layer, and two hidden layers. Among them, the input dimension is 4, the output dimension is 1, and each hidden layer has 128 neurons, and their weights are initialized using a normal distribution with a mean of 0 and a standard deviation of 0.1;
[0032] p = π θ (s)
[0033] where π is the policy model, θ is the parameter of the policy model π, and p = [p1, p2,..., p N is the probability distribution of candidate actions in state s, that is, the output of the policy model π, and satisfies
[0034] Still further, in Step Three, the policy model outputs the probability distribution of the model split points according to the input environmental state information. The sampler performs probability sampling of the execution actions according to the action probability distribution, thereby dynamically determining the action policy, that is, the best split layer and compression policy of the DNN model;
[0035] a ~ p
[0036] where a is the action sampled by the sampler according to the probability distribution p of the candidate actions. Subsequently, the splitting calculation and compression process corresponding to a will be executed. The sampled action a includes the split point decision d and compression decision c of the splitting calculation;
[0037] a = [d, c]
[0038] The compression operation can be decomposed into four processes: calculating the channel importance of features, sorting the channel importance, reducing the channels of the feature map, and feature quantization. Among them, the compression decision is defined as the compression rate, and the calculation method of the compression rate is as follows:
[0039]
[0040] According to the splitting point decision d calculated by the model, the model M can be divided into the head model M head and the tail model M tail ;
[0041] M head , M tail = f divide (M, d)
[0042] where M head is the model from the first layer to the d-th layer, and M tail is the model of the remaining part.
[0043] Furthermore, in the fourth step, the mobile device performs the calculation of the first half of the DNN model according to the determined model splitting point to obtain the intermediate layer feature data X;
[0044] X = M head (D in )
[0045] Next, before the data transmission starts, a compression process of the intermediate feature data based on the channel attention mechanism will be performed to reduce the amount of data transmitted, thereby improving the overall performance.
[0046] The operation steps of the channel importance evaluation based on the channel attention module in the fifth step are as follows:
[0047] 5.1) By performing max pooling on the intermediate layer feature data X, P max ;
[0048]
[0049] 5.2) By performing average pooling on the intermediate layer feature data X, P avg ;
[0050]
[0051] 5.3) Respectively pass the average pooling feature P avg and the max pooling feature P max through a series of operations including a fully connected layer, ReLU activation operation, and sigmoid operation, and finally obtain two intermediate features F avg ∈ R b×c and F max∈R b×c ;
[0052] F max = f2(f1(P max ))
[0053] F avg = f2(f1(P avg ))
[0054] where f1 represents the set of the first fully connected layer and ReLU operation, and f2 represents the set of the second fully connected layer and sigmoid operation;
[0055] 5.4) Add the two intermediate features F avg and F max to obtain the channel attention weight value A ∈ R b×c×1×1 , then perform sigmoid activation operation on A to obtain the normalized attention weight A norm ∈ R b×c×1×1 , that is, the channel importance value I;
[0056] A norm = σ(A) = σ(F max + F avg ) = I.
[0057] The beneficial effects of the present invention are as follows: The present invention adopts a dynamic optimization strategy. Compared with the traditional static model segmentation method, the present invention can dynamically adjust the model segmentation point and feature compression rate according to the real-time network state and task requirements, improving the adaptability and flexibility of the inference performance to the dynamic network environment. Through edge-cloud collaborative inference, the present invention efficiently and fully utilizes the computing resources of mobile devices and edge devices, reducing the inference latency and energy consumption, while ensuring the accuracy of the inference results. Compared with uploading the task completely to the cloud, the present invention reduces the data transmission volume, reduces the risk of data leakage, and enhances user privacy protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a flowchart of an adaptive edge-cloud collaboration method based on DNN partitioning and feature compression.
[0059] Figure 2 is a structural diagram of the DRL algorithm.
[0060] Figure 3 is a flowchart of the feature compression module. DETAILED DESCRIPTION OF THE INVENTION
[0061] The present invention will be further described below with reference to the accompanying drawings.
[0062] Refer to Figures 1 to 3, an adaptive edge-cloud collaboration method based on DNN partitioning and feature compression, comprising the following steps:
[0063] Step 1, collect network environment status information s: In an edge computing system composed of a wireless device (WD) and an edge computing server (ECS), the mobile device collects the current network environment status information s, which is the input of the policy model and includes signal quality Q, transmission rate V, channel gain S, and task volume A;
[0064] s = [Q, V, S, A]
[0065] Where the task model to be executed is M, which contains N main weight layers, including N conv convolutional layers and N fc fully connected layers. In specific implementation, the VGG16 model is used as the task model for image classification tasks. This model contains 16 main weight layers, including 13 convolutional layers and 3 fully connected layers;
[0066] Step 2, input the environment status information into the policy model π. The policy model π used is a DRL model. First, input the collected environment status information s into the deep reinforcement learning DRL model based on policy gradient. This model adopts a fully connected network structure, including an input layer, an output layer, and two hidden layers. Among them, the input dimension is 4, the output dimension is 1, each hidden layer has 128 neurons, and their weights are initialized using a normal distribution with a mean of 0 and a standard deviation of 0.1;
[0067] p = π θ (s)
[0068] Where π is the policy model, θ is the parameter of the policy model π, p = [p1, p2, …, p N is the probability distribution of candidate actions in state s, that is, the output of the policy model π, and satisfies
[0069] Step 3, output the probability distribution of the joint action of the model segmentation point and the compression rate: The policy model outputs the probability distribution of the model segmentation point according to the input environment status information. The sampler performs probability sampling of the execution action according to the action probability distribution, so as to dynamically determine the action strategy, that is, the best segmentation layer and compression strategy of the DNN model;
[0070] a ∼ p
[0071] Among them, a is the action sampled by the sampler according to the probability distribution p of the candidate action. Subsequently, the splitting calculation and compression process corresponding to a will be executed. The sampled action a includes the splitting point decision d and the compression decision c of the splitting calculation;
[0072] a = [d, c]
[0073] The compression operation can be decomposed into four processes: channel importance calculation of features, channel importance ranking, channel reduction of feature maps, and feature quantization. For the specific implementation, please refer to the description in Steps Five to Seven; among them, the compression decision is defined as the compression ratio, and the calculation method of the compression ratio is as follows:
[0074]
[0075] According to the splitting point decision d calculated by the model, the model M can be divided into a head model M head and a tail model M tail ;
[0076] M head , M tail = f divide (M, d)
[0077] Among them, M head is the model from the first layer to the d-th layer, and M tail is the model of the remaining part;
[0078] Step Four: The mobile device executes the first half of the DNN model calculation. The mobile device executes the first half of the DNN model calculation according to the determined model splitting point to obtain the intermediate layer feature data X;
[0079] X = M head (D in )
[0080] Next, before the data transmission starts, a compression process of the intermediate feature data based on the channel attention mechanism will be performed to reduce the amount of data transmitted, thereby improving the overall performance;
[0081] Step Five: After several layers of calculation of the head model M head , the intermediate layer feature data X ∈ R b×c×h×w of the model inference is obtained. Then, the channel attention value A ∈ R b ×c×1×1 is calculated for the intermediate layer feature data X through the channel attention module. In the present invention, the normalized result A norm ∈ R b×c×1×1 after the sigmoid operation of A is used as the channel importance evaluation result I ∈ R b×c×1×1 , where b is the number of batches, c is the number of channels, and h and w are the height and width of the feature map;
[0082] The operating steps for channel importance evaluation based on the channel attention module are as follows:
[0083] 5.1) Max-pool the intermediate layer feature data X to obtain P max ;
[0084]
[0085] 5.2) Average-pool the intermediate layer feature data X to obtain P avg ;
[0086]
[0087] 5.3) Respectively pass the average-pooled feature P avg and the max-pooled feature P max through a series of operations including a fully connected layer, ReLU activation operation, and sigmoid operation, and finally obtain two intermediate features F avg ∈ R b×c and F max ∈ R b×c ;
[0088] F max = f2(f1(P max ))
[0089] F avg = f2(f1(P avg ))
[0090] where f1 represents the set of the first fully connected layer and ReLU operation, and f2 represents the set of the second fully connected layer and sigmoid operation.
[0091] 5.4) Add the two intermediate features F avg and F max to obtain the channel attention weight value A ∈ R b×c×1×1 , and then perform sigmoid activation on A to obtain the normalized attention weight A norm ∈ R b×c×1×1 , that is, the channel importance value I;
[0092] A norm = σ(A) = σ(F max + F avg ) = I;
[0093] Step Six: According to the channel importance result I in Step Five, perform channel reduction on the intermediate layer feature data X: First, sort the channel importance value I to obtain the result I sort = [I1, I2…, I b×c; Then, according to the decision a obtained from the DRL decision model, channel reduction is performed on the intermediate feature data, where the number of retained channels is b×c×comp, corresponding to the channel data with relatively high importance, i.e., I selected = [I1, I2…, I b×c×comp ; According to the selected channel index list I selected , the corresponding important channel features are extracted from the intermediate layer feature data X to obtain the important channel feature tensor, which is the result X selected after channel reduction;
[0094] Step Seven: Feature quantization is performed on the result X seleted after channel reduction obtained in Step Six to obtain the final data D comp for communication transmission: The reduced feature data X seleted is quantized, and the data type is quantized from float32 to float16 to further reduce the data volume;
[0095] D comp = f comp (X selected )
[0096] Steps Five to Seven complete the entire feature compression process. The next steps transmit the processed data to the edge device, perform decompression operations, and then continue the remaining calculation process:
[0097] Step Eight: Transmit the compressed intermediate feature data D comp to the edge computing server ECS. The wireless device WD transmits the quantized intermediate feature data D comp to the edge computing server ECS through the wireless network;
[0098] D decomp = f decomp (D comp )
[0099] Step Nine: Edge device dequantization and continued calculation: After receiving the intermediate feature data, the edge device performs dequantization operations to restore the features to their original size and continues to execute the second half of the DNN model calculation;
[0100] Result = M tail (D decomp )
[0101] Step Ten: Collect inference performance data: After the inference task is completed, the system collects the inference task latency L and total energy consumption E during the inference process. The inference task latency L consists of the calculation latency on the WD, the transmission latency in the wireless network, and the calculation latency on the ECS, and is obtained by detecting the end-to-end latency on the WD. The total energy consumption E is obtained through the WD using the command line;
[0102] Step Eleven: Calculate the reward function value: Calculate the reward function value according to the reward function formula based on the inference performance data collected in Step Ten;
[0103] R = ω L ·R L + ω E ·R E
[0104] where ω L and ω E are the trade-off coefficients of task latency L and total energy consumption E respectively, and R L and R E are calculated as follows:
[0105] R L = -tanh[A(L - B)] + C
[0106] R E = -tanh[A(E - B)] + C
[0107] where tanh represents the hyperbolic tangent function, and A, B, and C are empirical parameter settings after a large number of experimental evaluations;
[0108] Step Twelve: Calculate the gradient value and perform gradient backpropagation according to the reward function value,
[0109]
[0110] where p i is the probability value corresponding to the selected action i, and R bl is the baseline for optimizing the reward value, which refers to the reference value used to evaluate and compare the algorithm performance during a given iteration. Specifically, R bl is determined by calculating the average of the reward values in the 50 iteration rounds forward from the current round t;
[0111] Step Thirteen: Backpropagate the calculated gradient value to implement gradient backpropagation, optimize the DRL model parameters. After each complete inference, adjust the parameters of the DRL model according to the performance data in the inference process collected in Step Ten, and optimize the policy model parameters for the offloading strategy selection to continuously adapt to the changing network environment, improve the system inference performance and environmental adaptability. Subsequently, the model will be continuously iteratively updated in this update manner to optimize the performance.
[0112] This embodiment elaborates in detail the specific implementation steps of an adaptive edge-cloud collaboration method based on DNN partitioning and feature compression. This method achieves efficient joint inference performance on mobile devices and edge devices through deep reinforcement learning (DRL) technology. The technical solutions adopted in this embodiment, including but not limited to model segmentation, feature compression, channel attention mechanism, and quantization processing, are all for achieving adaptive optimization in different network environments. Through the detailed description of this embodiment, the specific application and implementation manner of the present invention can be clearly demonstrated, thereby providing operable technical guidance for technicians.
[0113] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept and is only for illustrative purposes. The protection scope of the present invention should not be regarded as limited to the specific forms stated in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be conceived by those of ordinary skill in the art based on the inventive concept.
Claims
1. An adaptive end-edge collaboration method based on DNN partitioning and feature compression, characterized in that: The method comprises the following steps: Step 1. Collect network environment status information s: In the edge computing system composed of wireless devices WD and edge computing servers ECS, the wireless devices collect the current network environment status information s, which is the input of the strategy model including signal quality Q, transmission rate V, channel gain S and task amount A; s=[Q,V,S,A] The task model to be executed is M, which contains N main weight layers, including N conv convolutional layers and N fc Fully connected layers; Step 2: Input the environment status information into the strategy model π; Step 3: Output the probability distribution of the joint action of the model segmentation point and compression rate; Step 4: The mobile device performs the calculation of the first half of the DNN model; Step 5. On the head model M head After several layers of calculations are completed, M head It is the model from the first layer to the dth layer, and the intermediate layer feature data X∈R for model reasoning is obtained b×c×h×w , and then the channel attention module is used to calculate the intermediate layer feature data X to obtain the channel attention value A∈R b×c×1×1 , the normalized result A after the sigmoid operation norm ∈R b×c×1×1 As the importance evaluation result of the channel I∈R b×c×1×1 , where b is the number of batches, c is the number of channels, h and w are the height and width of the feature map; Step 6: According to the channel importance result I of step 5, the channel reduction is performed on the intermediate layer feature data X; first, the channel importance value I is sorted to obtain the result I after I is sorted in descending order sort =[I1, I2…, I b×c ]; Then, according to the decision a obtained by the DRL decision model, the intermediate feature data is channel-reduced, where the number of channels retained is b×c×comp, corresponding to the channel data with higher importance, i.e., I selected =[I1, I2…, I b×c×comp ]; According to the selected channel index list I selected , extract the corresponding important channel features from the intermediate layer feature data X, and obtain the important channel feature tensor, that is, the result after channel reduction X selected ; Step 7: The result X after channel reduction obtained in step 6 selected Perform feature quantization to obtain the final data D for communication transmission comp ; For the reduced feature data X selected Perform quantization processing to quantize the data type from float32 to float16 to further reduce the amount of data; D comp =f comp (X selected ); Step 8: Send the compressed intermediate feature data D comp To the edge computing server ECS, the wireless device WD, the quantized intermediate feature data D comp Send to the edge computing server ECS via wireless network; D decomp =F decomp (D comp ); Step 9: Dequantization and continued calculation of the edge device. After receiving the intermediate feature data, the edge device performs a dequantization operation to restore the feature to its original size and continues to perform the second half of the calculation of the DNN model. Result=M tail (D decomp ); Step 10: Collect inference performance data: After the inference task is completed, the system collects the inference task delay L and total energy consumption E during the inference process. The inference task delay L is composed of the calculation delay on the WD, the transmission delay in the wireless network, and the calculation delay on the ECS. It is obtained by detecting the end-to-end delay on the WD. The total energy consumption E is obtained by using the command line of the WD. Step 11: Calculate the reward function value: Based on the inference performance data collected in step 10, calculate the reward function value according to the reward function formula. R=ω L ·R L +oh E ·R E where ω L and ω E are the trade-off coefficients between task delay L and total energy consumption E, R L and R E The calculation method is as follows: R L =-tan h[A(LB)]+C R E =-tan h[A(EB)]+C Among them, tanh represents the hyperbolic tangent function, and A, B, and C are empirical parameter settings after a large number of experimental evaluations; Step 12: Calculate the gradient value: Return the gradient according to the reward function value. where p i is the probability value corresponding to the selected action i, R bl It is the baseline for reward value optimization, which refers to the reference value used to evaluate and compare algorithm performance in a given iteration process. bl It is determined by counting the average reward values of the 50 iteration rounds forward from the current round t; Step 13: Send back the calculated gradient value to realize gradient back propagation and optimize the DRL model parameters. After each complete reasoning, adjust the parameters of the DRL model according to the performance data of the reasoning process collected in step 10, and optimize the policy model parameters of the offloading policy selection to continuously adapt to the changing network environment and improve the system reasoning performance and environmental adaptability. In the future, this update method will be used to continuously iterate and update the model and optimize the performance.
2. The adaptive end-edge collaboration method based on DNN partitioning and feature compression according to claim 1, characterized in that: In the step 2, the policy model π used is a DRL model. First, the collected environment state information s is input into a deep reinforcement learning DRL model based on policy gradient. The model adopts a fully connected network structure, including an input layer, an output layer and two hidden layers, wherein the input dimension is 4, the output dimension is 1, each hidden layer has 128 neurons, and their weights are initialized using a normal distribution with a mean of 0 and a standard deviation of 0.1; p=π θ (s) Where π is the policy model, θ is the parameter of the policy model π, and p = [p1, p2, ..., p N ] is the probability distribution of candidate actions in state s, that is, the output of the policy model π, and satisfies 3. The adaptive end-edge collaboration method based on DNN partitioning and feature compression according to claim 1 or 2, characterized in that: In the step 3, the strategy model outputs the probability distribution of the model segmentation points according to the input environment state information, and the sampler performs probability sampling of the execution action according to the action probability distribution, thereby dynamically determining the action strategy, that is, the optimal segmentation layer and compression strategy of the DNN model; a~p Where a is the action sampled by the sampler after performing probability sampling according to the probability distribution p of the candidate action. The split calculation and compression process corresponding to a will be performed later. The sampled action a includes the split point decision d and the compression decision c of the split calculation; a=[d,c] The compression operation is decomposed into four processes: channel importance calculation of features, channel importance sorting, channel reduction of feature maps, and feature quantization. The compression decision is defined as the compression rate, and the compression rate calculation method is as follows: According to the split point decision d calculated by the model, the model M is divided into the head model M head and tail model M tail ; M head ,M tail =f divide (M,d) Among them, M head is the model from the first layer to the dth layer, M tail is the model for the rest.
4. The adaptive end-edge collaboration method based on DNN partitioning and feature compression according to claim 1 or 2, characterized in that: In the step 4, the mobile device performs the calculation of the first half of the DNN model according to the determined model segmentation point to obtain the intermediate layer feature data X; X=M head (D in ) Next, before data transmission begins, an intermediate feature data compression process based on the channel attention mechanism will be performed to reduce the amount of transmitted data and thus improve the overall performance.
5. The adaptive end-edge collaboration method based on DNN partitioning and feature compression according to claim 1 or 2, characterized in that: In step 5, the operation steps of channel importance evaluation based on the channel attention module are as follows: 5.1) By performing maximum pooling on the intermediate layer feature data X, we get P max ; 5.2) By performing average pooling on the intermediate layer feature data X, we get P avg ; 5.3) Average pooling feature P avg And the maximum pooling feature P max Through a series of operations including the fully connected layer, ReLU activation operation, and sigmoid operation, two intermediate features F are finally obtained. avg ∈R b×c and F max ∈R b×c ; F max =f2(f1(P max )) F avg =f2(f1(P avg )) Where f1 represents the set of the first fully connected layer and ReLU operations, and f2 represents the set of the second fully connected layer and sigmoid operations; 5.4) The two intermediate features F avg and F max Add together to get the channel attention weight value A∈R b×c×1×1 , then perform sigmoid activation on A to obtain the normalized attention weight A norm ∈R b×c×1×1 , that is, the channel importance value I; A norm =σ(A)=σ(F max +F avg )=I.
Citation Information
Cited By
Image detection method based on edge calculation and attention perception compression
CN121937846A
Task processing method and device, storage medium and electronic equipment
CN122019061A
Task processing methods and devices, storage media, electronic devices
CN122019061B