Collaborative intelligent feature compression algorithm based on channel attention
By adopting a collaborative intelligent feature compression algorithm with channel attention in video analysis, the problems of large computing volume and high communication cost of light mobile edge terminals are solved, and efficient feature compression and object detection are achieved.
Patent Information
- Application Number
- CN202410096798.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-24
- Publication Date
- 2025-07-25
AI Technical Summary
Existing deep learning networks are computationally expensive, communication costs are high, and easily congested and difficult to operate efficiently when performing video analysis on lightweight mobile edge terminals.
A collaborative intelligent feature compression algorithm based on channel attention is adopted to input the video frames into the YOLOX network for feature extraction, select split points with a feature channel dimension of 256, compress features through the channel attention network, and perform subsequent calculations in the cloud, and reconstruct using a generalized division normalization layer.
It effectively reduces the communication bandwidth pressure, ensures the performance of machine vision tasks, and improves the accuracy and efficiency of object detection.
Smart Images

Figure BDA0004678705590000022 
Figure BDA0004678705590000023 
Figure BDA0004678705590000031
Abstract
Description
Technical Field
[0001] The present invention relates to video feature compression technology, and specifically to a collaborative intelligent feature compression algorithm based on channel attention, belonging to the field of image communication. Background Art
[0002] With the rapid development of technology and the advancement of the modernization process, more and more surveillance videos, online videos, etc. begin to use machine vision to perform intelligent analysis on video content, greatly changing our lives. Benefiting from the development of deep learning, deep neural network DNN is now mostly used for intelligent analysis tasks such as object detection, image segmentation, and video tracking. Since DNN often has a complex structure and a huge amount of computation, it has high requirements for the performance of the deployed device and is difficult to operate efficiently on lightweight mobile edge terminals such as surveillance cameras. Therefore, mobile devices are often used as sensors to capture data, and then the data is transmitted to the cloud to perform heavy neural network tasks. However, this method faces challenges of long latency and high communication cost. In addition, multiple mobile devices may cause congestion, thus limiting the processing throughput of the cloud.
[0003] An effective solution is to use a collaborative intelligent framework, use a neural network to extract and compress video features, and transmit the compressed features to the cloud for intelligent analysis. First, carefully select the split point to divide the intelligent analysis model into two parts. One part is deployed on the mobile terminal to extract feature maps from the video data captured on the terminal, and the other part is deployed in the cloud for subsequent complex calculations. The extracted feature maps are intended to balance the workloads of mobile devices, the cloud, and network communication. Among them, the effect of the feature compression network is crucial for the collaborative intelligent framework. Summary of the Invention
[0004] The purpose of the present invention is to reduce the communication bandwidth pressure by feature extraction and compression of video streams, while ensuring the performance of machine vision tasks.
[0005] A collaborative intelligent feature compression algorithm based on channel attention proposed by the present invention mainly includes the following operating steps:
[0006] (1) Input the video frame into the basic intelligent analysis network YOLOX.
[0007] (2) Select the split point from the feature extractor part of YOLOX to determine the intermediate features to be compressed.
[0008] (3) Input the intermediate features selected in (2) into our collaborative intelligent feature compression network based on channel attention to obtain the compressed features.
[0009] (4) Upload the compressed features to the cloud and input them into the second half of the YOLOX network deployed on the cloud to obtain the video object detection results.
[0010] Specifically, assume that the original video to be analyzed is F, which consists of n frames, i.e., {F1, F2, F3,..., F n} ∈ F. Process it using the collaborative intelligent feature compression algorithm based on channel attention. The overall diagram of the collaborative intelligent feature compression algorithm based on channel attention is as Figure 1 shown. We select YOLOX as the intelligent analysis network therein, and the focus of the network lies in the implementation of the feature compression module. When the video frame input into the YOLOX feature extractor is F t , we select the feature with a feature channel dimension of 256 as the feature at the splitting point, i.e., the feature to be compressed. Input the feature to be compressed into the compression network based on channel attention, and the channel attention network structure is as Figure 1 shown. First, determine the number of channels after feature compression according to the selection of the compression ratio, denoted as 2n. First, use a channel scoring network that combines convolution and fully connected layers to score different channels of the intermediate features according to their different importance for the intelligent analysis task. Select the top n scores from high to low, completely retain the corresponding channels, and weight the remaining channels according to the score size to fuse them into an n-channel feature. Finally, concatenate the total 2n channels obtained before and after to obtain the preliminarily compressed feature.
[0011] S = SN(F sp )
[0012] S′ = Normalization(S)
[0013] where F sp is a splitting point feature with 256 channels, S is its corresponding 256 scores, and S′ is the set of scores after its weight normalization.
[0014] TopIndices(S′, n) = argsort(S, descending)[:n]
[0015]
[0016] where i1, i2,..., i n are the indices of the n highest-scoring feature channels in TopIndices(S, n), and n depends on the set compression ratio.
[0017]
[0018] F′ sp = S′ * F sp
[0019]
[0020] Among them, j1, j2, …, j 256-n are the indices of the remaining channels after removing the n highest-scoring feature channels in SelectedFeatures.
[0021]
[0022] F′2 = GDN(F2)
[0023] F compress = Concat(F1, F′2)
[0024] F2 is the feature obtained by convolving and fusing the remaining features, with a total of n channels. F′2 is the feature obtained by performing high-precision generalized divisive normalization on F2. GDN is the generalized divisive normalization operation. When bandwidth permits, if the high-precision mode is selected, that is, without performing spatial redundancy removal operation, then F at this time compress is the compressed feature.
[0025] If a smaller transmission bitrate is required, then on the basis of the above compressed features, a convolution operation with a convolution kernel size of 2 and a stride of 2 is used to reduce the feature size and remove the spatial redundancy of the features. The complete structure diagram of the channel aggregation network based on channel attention is as Figure 3 shown. Brief Description of the Drawings
[0026] Figure 1 is the overall diagram of the collaborative intelligent feature compression algorithm based on channel attention in the present invention.
[0027] Figure 2 is the structure diagram of the channel aggregation network NCAnet based on channel attention in the present invention.
[0028] Figure 3 is the complete structure diagram of the feature compression method based on channel attention in the present invention.
[0029] Figure 4 is the comparison diagram of the reconstructed features obtained by different algorithms for a certain video frame. Among them, Figure (a) is the original feature map, Figure (b) is the feature map at the corresponding network position after the video is compressed by comparison method 1, Figure (c) is the reconstructed feature by the algorithm of the present invention, and Figure (d) is the reconstructed feature by comparison method 2.
[0030] Figure 5 is the bitrate-accuracy performance comparison diagram of different compression algorithms. Detailed Embodiment
[0031] The present invention will be further described in detail below in conjunction with embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and should not be construed as limiting the protection scope of the present invention. Those skilled in the art can make some non-essential improvements and adjustments to the present invention based on the above-mentioned invention content and implement it specifically, which should still fall within the protection scope of the present invention.
[0032] Figure 1 Specifically, it relates to a collaborative intelligent feature compression algorithm based on channel attention, which can be specifically divided into the following steps:
[0033] (1) Input the video frame into the basic intelligent analysis network YOLOX.
[0034] (2) Select a split point from the feature extractor part of YOLOX to determine the intermediate feature to be compressed.
[0035] (3) Input the intermediate feature selected in (2) into our collaborative intelligent feature compression network based on channel attention to obtain the compressed feature.
[0036] (4) Upload the compressed feature to the cloud and input it into the second half of the YOLOX network deployed on the cloud to obtain the video object detection result.
[0037] As described in the above steps (1), (2), (3), and (4), through experimental verification, the information retained by this algorithm during channel fusion is the most complete, and the loss of the reconstructed intermediate feature on the object detection effect is the smallest. It can effectively relieve the communication bandwidth pressure and ensure the performance of machine vision tasks.
[0038] To better illustrate the effectiveness of the present invention, we show the impact of the patent algorithm and the comparison method on the performance of the object detection task under the transmission bit rate of 0.2 bpp, as shown in Table 1, where Raw is the original video without compression. On the other hand, to more intuitively show the effectiveness of the algorithm of the present invention, in Figure 4 we show various reconstructed feature maps input into the YOLOX detection end. Finally, to illustrate that the algorithm of the present invention has good performance under various bit rates, we conduct experiments on the algorithm of the present invention and the comparison algorithm under various bit rates, and the results are as Figure 5 shown.
[0039] Table 1 Impact of different compression algorithms on task performance (mAP)
[0040] RAW Comparison Method 1 Comparison Method 2 Method of the Present Invention Bus 97.01% 96.56% 96.04% 96.52% Car 96.53% 95.16% 95.28% 95.21% Cyclist 87.88% 85.18% 85.66% 86.31% Motorcyclist 75.79% 72.48% 74.11% 73.78% Pedestrian 81.80% 78.27% 76.09% 75.75% Traffic cone 90.25% 86.21% 87.60% 88.16% Truck 78.89% 76.44% 77.89% 77.78% Van 80.74% 78.79% 77.61% 79.44% Average 86.11% 83.64% 83.79% 84.12%
[0041] The comparison methods are:
[0042] Method 1: Video coding standard HEVC, encoder is HM16.20.
[0043] Method 2: The method proposed by Shao J, Zhang J, etc., reference "Bottlenet++: An end-to-end approach for feature compression in device-edge co-inference systems[C] / / 2020 IEEE International Conference on Communications Workshops(ICC Workshops). IEEE, 2020: 1-6."
Claims
1. A channel aggregation algorithm combined with channel attention, which has good effect in feature compression, is characterized in that Based on the different task correlations of different channels according to their characteristics, different processing methods are adopted to remove redundancy in the channels while retaining the most relevant information for intelligent analysis tasks. The steps are as follows: (1) Input the features at the splitting point into the channel scoring network based on fully connected layers; The specific method is to first input the features into the global pooling layer to make the width and height of the output features both 1. Subsequently, the obtained features are input into the fully connected layer for channel dimensionality reduction to extract the most important information from the input channels. Then, the RELU activation function is used to introduce a non-linear transformation. Next, a fully connected network is used to restore the number of feature channels to the original number, adjust the channel weights of the model, and finally, the output is restricted to the range of (0,1) through the Sigmoid activation function to generate the task correlation scores on each channel and perform normalization; (2) Perform different processing on different channels by combining the channel scores obtained in (1); The specific method is to extract and aggregate the channels with the top n scores according to the scores, where n is half of the number of feature channels after compression. The top n channels are completely retained, and the remaining channels with lower scores are weighted and aggregated with the normalized scores obtained in (1), and the target number of aggregated channels is n; (3) Aggregate the two types of features obtained in (2) to obtain the preliminarily compressed features; For the two types of features obtained in (2), a concatenation operation is performed in the channel dimension to obtain features with a preliminarily compressed channel number of 2n.
2. A method for applying generalized division normalization to feature compression, characterized in that, Define a new natural image probability model through a reversible non-linear transformation, and determine the parameters by minimizing the KL divergence between the transformed data distribution and the Gaussian target. It is innovatively used in the feature compression task to introduce less noise during the feature compression and feature reconstruction processes and more accurately restore the compressed features.