Efficient target detection method and system based on bottleneck design and double-line distillation
By combining bottleneck design and double-line distillation technology, the problem of excessive use of object detection algorithm computing resources on the edge computing platform is solved, efficient object detection is achieved, and the requirements of real-time and accuracy are met.
Patent Information
- Application Number
- CN202510075140.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
When existing object detection algorithms are deployed on edge computing platforms, the computing resources are used too much, resulting in inference delay and real-time problems, which are difficult to meet the needs of complex tasks.
The efficient object detection method based on bottleneck design and double-line distillation is adopted. Through the combination of lightweight bottleneck design and double-line distillation module, the parameter quantity and calculation complexity are compressed, while improving detection accuracy and inference speed.
While maintaining the accuracy and accuracy of object detection, it significantly reduces the computational complexity and parameter quantity, accelerates the inference speed on the edge computing platform, and meets the real-time requirements.
Smart Images

Figure CN119992053A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to an efficient target detection method and system based on bottleneck design and double-line distillation, and belongs to the technical field of image processing. Background Art
[0002] At present, object detection technology has become an important branch of research in the field of computer vision and artificial intelligence. It is mainly used in complex related tasks such as augmented reality (AR), metaverse, smart transportation, and autonomous driving. These tasks often need to be deployed on platforms with limited computing resources such as edge computing.
[0003] In recent years, deep neural networks have been used for their high reasoning speed, robustness, and ability to learn low-level and high-level semantic information. However, they require a large amount of computing resources. If only one target detection algorithm is deployed on a distributed device, once there are too many image resources to be processed, there will often be problems such as high reasoning latency, inability to process in real time, and excessive CPU and video memory resources. It is even more impossible to cooperate with other algorithms on edge computing devices with limited computing resources to complete more complex tasks that are closer to practical applications.
[0004] In recent years, methods such as model pruning, network quantization, knowledge distillation, neural network search, lightweight network design, and bottleneck design have been proposed and applied to the lightweight of target detection networks. Many methods have been proposed and have shown excellent results. Among them, knowledge distillation and bottleneck design have become the first choice for efficient construction of target detection models due to their adaptability to various platforms. Bottleneck design starts with module design and directly reduces the computational complexity of the model. Knowledge distillation is used to transfer the effective features and gradient information of the larger target detection network teacher model to the method as an efficient target detection network student model of the student network, so that the student model can achieve detection accuracy and precision similar to or even exceeding that of the teacher model.
[0005] When designing a lightweight bottleneck module for the target detection network for the edge computing platform, this method often needs to take into account its limited computing resources, the platform's support for operator acceleration, the platform's optimization between different operators and other special characteristics, in order to design a relatively good bottleneck that can balance accuracy and speed for the edge computing platform, so as to further build an efficient target detection network. Summary of the invention
[0006] Purpose of the invention: In order to overcome the shortcomings of the prior art, the present invention provides an efficient target detection method based on bottleneck design and two-line distillation. Through the combination of lightweight bottleneck design and two-line distillation modules, while maintaining the target detection precision and accuracy, a large amount of parameters and computational complexity are compressed, thereby accelerating the inference speed when deployed on the edge computing platform and meeting the real-time requirements during application.
[0007] Technical solution: To achieve the above purpose, the technical solution adopted by the present invention is:
[0008] An efficient target detection method based on bottleneck design and double-line distillation includes the following steps:
[0009] Step 1: Collect image data of track scene obstacles and mark the location information of the targets to be detected to obtain a training set.
[0010] Step 2: construct an efficient target detection neural network based on lightweight bottleneck design and dual-line distillation module. The efficient target detection neural network includes an efficient target detection neural network constructed using lightweight bottleneck, a dual-line distillation module in which a feature distillation module and a logic distillation module are mixed in proportion, and the efficient target detection neural network constructed using lightweight bottleneck includes a teacher network and a student network. The distillation weight of the student network is obtained by offline distillation of the student network using the training weight of the teacher network through the dual-line distillation module.
[0011] Step 3: Use the training set to train the efficient target detection neural network. At this time, the teacher network is trained, and the training weights of the teacher network are finally obtained. Then, the training weights of the teacher network are used to perform offline distillation on the student network through a two-line distillation module, and finally the distillation weights of the student network are obtained.
[0012] Step 4: Input the target image to be detected into the student network for recognition to obtain detection information.
[0013] Preferably: the loss calculation formula of the dual-line distillation module is as follows:
[0014]
[0015] Among them, L total represents the bilinear distillation loss function, β1 represents the coefficient of the total logistic distillation loss function, β2 represents the coefficient of the feature distillation loss function, n represents the number of feature maps in each training batch, K represents the total number of categories to be detected, ω i,j Represents the absolute difference between the gradients of the teacher model and the student model for the jth category at the ith position after Sigmoid function conversion, ω .,j Represents the absolute difference between the teacher gradient and the student gradient in the jth category after the Sigmoid function conversion, Represents the conversion value of the gradient of the teacher model for the jth category at the ith position after the Sigmod function, Represents the conversion value of the student model gradient for the jth category at the ith position after the Sigmod function, μ i ′ represents the intersection-over-union ratio between the bounding boxes predicted by the teacher model and the student model at the i-th anchor position, α1 represents the optimized value obtained during training, α2 represents the optimized value obtained during training, N represents the batch size, H represents the feature map height, and W represents the feature map width. represents the eigenvalue of the student model at the position (h,w) in the feature map of the i-th sample, Represents the feature value of the teacher model at the (h, w) position in the feature map of the i-th sample.
[0016] Preferably: the feature distillation module is used to compensate for the loss of information between channels in deep convolution, select more important channels, and compensate for the feature information lost in the dimensionality reduction operation of the point multiplication convolution channel.
[0017] Preferably: the feature distillation module performs softmax normalization on each channel to obtain a probability distribution representing the relative importance of each position in the channel, and then calculates the asymmetric KL divergence between the probability distributions of the teacher network and the student network response channels as a loss, so that the student network imitates the teacher network in the foreground salient area.
[0018] Preferably, the logic distillation module is used to compensate for the reduced accuracy of local spatial feature information interaction caused by replacing standard convolution with point multiplication convolution, while improving the interaction of the entire network with global feature information.
[0019] Preferably: the logistic distillation module maps the classification gradients into a classification map in the binary system of the number of categories, then uses the Sigmod protocol to obtain the scores, and applies a binary cross entropy loss to extract each binary classification map in the teacher-student network.
[0020] Preferably: the teacher network and the student network include a feature extraction module, a multi-scale feature fusion module, and a prediction head module which are connected in sequence.
[0021] Another object of the present invention is to provide an efficient target detection system based on bottleneck design and two-line distillation, which adopts the efficient target detection method based on bottleneck design and two-line distillation, including an input unit, an efficient target detection neural network unit, and an output unit, wherein:
[0022] The input unit is used to input a target image to be detected.
[0023] The efficient target detection neural network unit is provided with an efficient target detection neural network, which includes an efficient target detection neural network constructed using a lightweight bottleneck, a dual-line distillation module in which a feature distillation module and a logic distillation module are mixed in proportion, and the efficient target detection neural network constructed using a lightweight bottleneck includes a teacher network and a student network. The network structures of the two networks are generally similar, and the number of feature channels of the teacher network is 1 / 2 more than that of the student network, so the teacher network has a stronger feature extraction capability and detection accuracy. The distillation weight of the student network is obtained by offline distillation of the student network using the training weight of the teacher network through the dual-line distillation module. The efficient target detection neural network unit is used to identify the target image to be detected through the efficient target detection neural network to obtain detection information.
[0024] The output unit is used to output detection information.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] The present invention effectively utilizes the complementarity of lightweight bottleneck design to compress computational complexity and dual-line distillation module to improve detection accuracy, uses point multiplication convolution to reduce the channel dimension of features, and then uses deep convolution to further compress the computational amount, and then uses splicing operation to complete the reuse of the previous step features, so as to maximize the detection accuracy while accelerating the reasoning speed, so as to obtain the lightweight bottleneck mentioned in the steps of the method, and then use the bottleneck to construct feature information extraction module, feature fusion module, prediction head module, and finally obtain an efficient target detection network, and the operators of point multiplication convolution and deep convolution are used because of the existence of various restricted conditions of the edge computing platform. Among them, feature distillation is used to make up for the lack of information interaction on channel features brought by deep convolution, and logical distillation is used to make up for the lack of local spatial information interaction brought by point multiplication convolution, and then the two distillation methods are combined in proportion to form the dual-line distillation module of the method. In summary, the bottleneck design and dual-line distillation strategy provided by the present invention compress a large number of parameters and computational complexity while maintaining the accuracy and precision of target detection, speed up the reasoning speed when deployed on the edge computing platform, and meet the real-time requirements of application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a schematic diagram of the process of the efficient target detection method based on bottleneck design and double-line distillation of the present invention.
[0028] Figure 2 It is a lightweight bottleneck module.
[0029] Figure 3 It is a lightweight basic module.
[0030] Figure 4The overall structure of the neural network. DETAILED DESCRIPTION
[0031] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0032] An efficient target detection method based on bottleneck design and two-line distillation, such as Figure 1-4 As shown, the following steps are included:
[0033] Step 1: Collect image data of track scene obstacles and mark the location information of the targets to be detected to obtain a training set.
[0034] First, the target detection image data is collected and marked to obtain a labeling file, and the target detection image data and the labeling file are combined to obtain a training set.
[0035] Step 2: construct an efficient target detection neural network based on lightweight bottleneck design and dual-line distillation module. The efficient target detection neural network includes an efficient target detection neural network constructed using lightweight bottleneck, a dual-line distillation module in which a feature distillation module and a logic distillation module are mixed in proportion, and the efficient target detection neural network constructed using lightweight bottleneck includes a teacher network and a student network. The distillation weight of the student network is obtained by offline distillation of the student network using the training weight of the teacher network through the dual-line distillation module.
[0036] The lightweight bottleneck first uses point multiplication convolution to reduce the number of channels of the feature map to half of the original feature, and then uses the batch normalization layer and the activation function layer to enhance the nonlinear ability of feature extraction to obtain a set of basic features. Then, deep convolution is used to complete lightweight information extraction on the basic features to obtain a set of lightweight features. Then, the splicing operation is used to splice the basic features with the lightweight features in the channel dimension to obtain the spliced features, and then the spliced features are added to the lightweight features using the short-circuit operation.
[0037] The key point of the lightweight bottleneck design is that we found in the experiment that among the feature images obtained after a set of feature images have passed through two layers of standard convolution, many feature images have a high degree of overlap. Therefore, the concept of constructing a lightweight bottleneck is to use less computation and parameters to generate features similar to those obtained by two layers of standard convolution. Therefore, we first use point multiplication convolution to complete channel dimensionality reduction, and reuse these features through splicing operations after the cheap spatial feature extraction of deep convolution, so as to achieve the effect of feature extraction capability brought by two layers of standard convolution as much as possible.
[0038] Then, the lightweight bottleneck is used to build a lightweight basic module, and then the lightweight basic module is used to build a feature extraction module, a multi-scale feature fusion module and a prediction head module. The structures of the four modules are as follows: Figure 3 and Figure 4 shown.
[0039] The feature distillation module is used to compensate for the loss of information between channels in deep convolution, and pays more attention to how to transfer knowledge at the channel level, and compensates for the feature information lost by the dimension reduction operation of the point multiplication convolution channel. The logical distillation module is used to compensate for the reduced accuracy of the interaction of local spatial feature information caused by replacing the standard convolution with the point multiplication convolution, while improving the interaction of the entire network with global feature information. The feature distillation module performs softmax normalization on each channel to obtain a probability distribution that represents the relative importance of each position in the channel, and then calculates the asymmetric KL divergence between the probability distributions of the teacher network and the student network response channel as a loss, so that the student network imitates the teacher network in the foreground salient area. The logical distillation module maps the classification gradient to a classification mapping of the number of categories, then uses the Sigmod protocol to obtain the score, and applies the binary cross entropy loss to extract each binary classification map from the teacher to the student network.
[0040]
[0041] L logits (x) = α1·L cls (x)+α2·L loc (x)
[0042] Right now:
[0043]
[0044] L total (x) = β1·L logits (x)+β2·L feature (x)
[0045] Right now:
[0046]
[0047] Among them, L cls represents the classification loss function of logistic distillation, n represents the number of feature maps in each training batch, K represents the total number of categories to be detected, ω i,j It represents the absolute difference between the gradients of the teacher model and the student model for the jth category at the ith position after the Sigmoid function conversion. Represents the conversion value of the gradient of the teacher model for the jth category at the ith position after the Sigmod function, Represents the conversion value of the student model gradient for the jth category at the ith position after the Sigmod function, L loc represents the localization loss function of logical distillation, ω .,j represents the absolute difference between the teacher gradient and the student gradient in the jth category after the Sigmoid function conversion, μ i ′ represents the intersection-over-union ratio between the bounding boxes predicted by the teacher model and the student model at the i-th anchor position, L logits is the total logistic distillation loss function, α1 represents the optimized value obtained during training, α2 represents the optimized value obtained during training, and L feature is the feature distillation loss function, N represents the batch size, H represents the feature map height, and W represents the feature map width. represents the eigenvalue of the student model at the position (h,w) in the feature map of the i-th sample, represents the eigenvalue of the teacher model at the position (h, w) in the feature map of the i-th sample, L total It represents the bilinear distillation loss function, and also represents the relationship between logical distillation and feature distillation. β1 represents the coefficient of the total logical distillation loss function, and β2 represents the coefficient of the feature distillation loss function.
[0048] A two-line distillation module is used to compensate for the loss of precision and accuracy caused by the lightweight bottleneck design. Feature distillation is used to compensate for the loss of information between channels in deep convolution, improve channel information interaction, select more important channels, and compensate for the feature information lost in the channel dimensionality reduction operation of the point multiplication convolution. Another logical distillation method compensates for the reduced accuracy of local spatial feature information interaction caused by replacing standard convolution with point multiplication convolution, while improving the interaction of the entire network with global feature information. The use of this two-line distillation module can well compensate for the loss of precision and accuracy caused by the lightweight bottleneck design.
[0049] The teacher network and the student network include a feature extraction module, a multi-scale feature fusion module, and a prediction head module connected in sequence, and have the same overall structure. The number of feature channels of the teacher network is 1 / 2 more than that of the student network, so it has stronger feature extraction capability and higher target detection accuracy.
[0050] Step 3: Use the training set to train the efficient target detection neural network. At this time, the teacher network is trained, and the training weights of the teacher network are finally obtained. Then, the training weights of the teacher network are used to perform offline distillation on the student network through a two-line distillation module, and finally the distillation weights of the student network are obtained.
[0051] The present invention first inputs the training set into the teacher network for training to obtain the training weights of the teacher network, then uses the teacher network weights to perform offline distillation on the student model through a two-line distillation module, inputs the detection category information and bounding box information into the student network, and finally obtains the distilled student network weights.
[0052] Step 4: Input the target image to be detected into the student network for recognition to obtain target confidence and model inference speed information.
[0053] Another object of the present invention is to provide an efficient target detection system based on bottleneck design and two-line distillation, which adopts the efficient target detection method based on bottleneck design and two-line distillation, including an input unit, an efficient target detection neural network unit, and an output unit, wherein:
[0054] The input unit is used to input a target image to be detected.
[0055] The efficient target detection neural network unit is provided with an efficient target detection neural network, which includes an efficient target detection neural network constructed using a lightweight bottleneck, a dual-line distillation module in which a feature distillation module and a logic distillation module are mixed in proportion, and the efficient target detection neural network constructed using a lightweight bottleneck includes a teacher network and a student network, and the distillation weight of the student network is obtained by offline distillation of the student network using the training weight of the teacher network through the dual-line distillation module. The efficient target detection neural network unit is used to identify the target image to be detected through the efficient target detection neural network to obtain detection information.
[0056] The output unit is used to output detection information.
[0057] The focus of the present invention is on lightweight bottleneck design, using the property of a feature image having a large number of conflicting feature maps after two layers of convolution as the construction concept. First of all, point multiplication convolution and depth convolution are common operations in lightweight bottleneck design. The difference of the lightweight bottleneck of the present invention is that we will reuse the features after the dimension reduction of point multiplication convolution through splicing operations, so that a lightweight bottleneck with feature extraction capabilities closer to the bottleneck of two layers of standard convolution can be constructed with fewer parameters and calculations. The two-line distillation module in step 2 uses the characteristics of two different distillations to mix them in proportion, making full use of the accuracy improvement effect of the two distillations to make up for the loss of local spatial information brought by the lightweight bottleneck and the loss of channel feature interaction information.
[0058] The present invention effectively utilizes the compression parameters and calculation amount brought by point multiplication convolution and depth convolution as well as the ability to accelerate network reasoning speed to construct a lightweight bottleneck module, and uses the lightweight bottleneck module to construct a lightweight basic module, and then constructs a feature extraction module, a multi-scale feature fusion module, and a prediction head module, and combines the use of feature distillation to enhance the interaction between feature information channels and the improvement of feature space information interaction by logic distillation to compensate for the loss of detection accuracy caused by the lightweight bottleneck, and combines the characteristics of the edge computing platform of practical application to finally construct an efficient target detection network, which greatly accelerates the reasoning speed while maintaining the detection accuracy at a high level. Therefore. The present invention is deployed on the edge computing platform through lightweight bottleneck design and dual-line distillation, while maintaining the network detection accuracy, reducing the network calculation complexity and accelerating the network reasoning speed.
[0059] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. An efficient target detection method based on bottleneck design and double-line distillation, characterized in that: The steps include: Step 1: Collect target detection image data and mark them to obtain a training set; Step 2, constructing an efficient target detection neural network based on a lightweight bottleneck design and a two-line distillation module; the efficient target detection neural network includes an efficient target detection neural network constructed using a lightweight bottleneck, and a two-line distillation module in which a feature distillation module and a logic distillation module are mixed in proportion, the efficient target detection neural network constructed using a lightweight bottleneck includes a teacher network and a student network, and the distillation weight of the student network is obtained by offline distilling the student network through the two-line distillation module using the training weight of the teacher network; Step 3: Use the training set to train the efficient target detection neural network. At this time, the teacher network is trained, and the training weights of the teacher network are finally obtained. Then, the training weights of the teacher network are used to perform offline distillation on the student network through a two-line distillation module, and finally the distillation weights of the student network are obtained. Step 4: Input the target image to be detected into the student network for recognition to obtain detection information.
2. According to claim 1, the efficient target detection method based on bottleneck design and double-line distillation is characterized in that: The loss calculation formula of the dual-line distillation module is as follows: Among them, L total represents the bilinear distillation loss function, β1 represents the coefficient of the total logistic distillation loss function, β2 represents the coefficient of the feature distillation loss function, n represents the number of feature maps in each training batch, K represents the total number of categories to be detected, ω i,j Represents the absolute difference between the gradients of the teacher model and the student model for the jth category at the ith position after Sigmoid function conversion, ω .,j Represents the absolute difference between the teacher gradient and the student gradient in the jth category after the Sigmoid function conversion, Represents the conversion value of the gradient of the teacher model for the jth category at the ith position after the Sigmod function, Represents the conversion value of the student model gradient for the jth category at the ith position after the Sigmod function, μ i ′ represents the intersection-over-union ratio between the bounding boxes predicted by the teacher model and the student model at the i-th anchor position, α1 represents the optimized value obtained during training, α2 represents the optimized value obtained during training, N represents the batch size, H represents the feature map height, and W represents the feature map width. represents the eigenvalue of the student model at the position (h,w) in the feature map of the i-th sample, Represents the feature value of the teacher model at the (h, w) position in the feature map of the i-th sample.
3. The efficient target detection method based on bottleneck design and double-line distillation according to claim 2, characterized in that: The feature distillation module is used to compensate for the loss of information between channels in deep convolution, select more important channels, and compensate for the feature information lost in the dimensionality reduction operation of the point multiplication convolution channel.
4. The efficient target detection method based on bottleneck design and double-line distillation according to claim 3, characterized in that: The feature distillation module performs softmax normalization on each channel to obtain a probability distribution that represents the relative importance of each position in the channel, and then calculates the asymmetric KL divergence between the probability distributions of the teacher network and the student network response channels as the loss, so that the student network imitates the teacher network in the foreground salient area.
5. The efficient target detection method based on bottleneck design and double-line distillation according to claim 4, characterized in that: The logic distillation module is used to compensate for the reduced accuracy of local spatial feature information interaction caused by replacing standard convolution with point multiplication convolution, while improving the interaction of the entire network with global feature information.
6. The efficient target detection method based on bottleneck design and double-line distillation according to claim 5, characterized in that: The logistic distillation module maps the classification gradients into classification maps in the binary system of the number of categories, then uses the Sigmod protocol to obtain the scores and applies the binary cross entropy loss to extract each binary classification map in the teacher-to-student network.
7. The efficient target detection method based on bottleneck design and double-line distillation according to claim 6, characterized in that: The student network includes a feature extraction module, a multi-scale feature fusion module, and a prediction head module which are connected in sequence.
8. An efficient target detection system based on bottleneck design and double-line distillation, characterized in that: The efficient target detection method based on bottleneck design and two-line distillation as described in claim 1 comprises an input unit, an efficient target detection neural network unit, and an output unit, wherein: The input unit is used to input the target image to be detected; The efficient target detection neural network unit is provided with an efficient target detection neural network, the efficient target detection neural network includes an efficient target detection neural network constructed using a lightweight bottleneck, a dual-line distillation module in which a feature distillation module and a logic distillation module are mixed in proportion, the efficient target detection neural network constructed using a lightweight bottleneck includes a teacher network and a student network, and the distillation weight of the student network is obtained by performing offline distillation on the student network through the dual-line distillation module using the training weight of the teacher network; the efficient target detection neural network unit is used to identify the target image to be detected through the efficient target detection neural network to obtain detection information; The output unit is used to output detection information.
Citation Information
Cited By
Fruit target detection method based on lightweight multi-scale attention mechanism
CN120259793A