Lightweight Object Detection Method, Device, Equipment and Medium for Unmanned Aerial Vehicle
By building a lightweight object detection model on the drone, using feature pyramids and expansion convolutions to optimize the YOLOV11 model, and combining sparse training and channel pruning, the detection problem of the drone's computing resources is solved, and efficient and accurate object detection is achieved.
Patent Information
- Application Number
- CN202510518881.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing deep learning object detection model has large parameters and high computing complexity, so it is impossible to achieve real-time and accurate target detection on drones with limited computing resources, resulting in increased power consumption of drones, slow detection speed, and even lag.
Complex models are constructed based on YOLOV11, and the feature pyramid and expansion convolution are optimized. Combined with residual structure and self-attention mechanism, a lightweight model is constructed. Through sparse training and channel pruning, the lightweight model is optimized using the knowledge distillation framework to achieve object detection.
It realizes high-precision and low-complexity target detection on drones, meets real-time detection needs, reduces power consumption and improves detection efficiency.
Smart Images

Figure CN120047862B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer vision object detection, and particularly to a lightweight object detection method, device, equipment, and medium for unmanned aerial vehicles (UAVs). Background Art
[0002] With the rapid development of UAV technology, it has been widely used in many fields such as security monitoring, logistics distribution, agricultural plant protection, and environmental monitoring. In these application scenarios, UAVs need to detect and identify various target objects in real time and accurately, such as people and vehicles in security monitoring, pest areas and crop growth conditions in agricultural plant protection, etc.
[0003] Currently, common object detection technologies are mainly based on deep learning algorithms, such as object detection models based on convolutional neural networks (CNNs), like Faster R-CNN, YOLO series, etc. These algorithms can achieve relatively ideal detection accuracy in an environment with sufficient computing resources.
[0004] However, due to its own hardware limitations, such as limited computing power and insufficient battery life, UAVs have put forward lightweight requirements for the carried object detection methods. Existing deep learning object detection models have a large number of parameters and high computational complexity. When directly applied to UAVs, it will cause a significant increase in the power consumption of UAVs, seriously shortening the flight time. At the same time, the detection speed is also difficult to meet the real-time requirements, and even situations such as freezing and detection interruption may occur due to the exhaustion of computing resources, making it impossible to achieve efficient and stable object detection, which greatly limits the further application and development of UAVs in related fields. Summary of the Invention
[0005] Based on this, it is necessary to provide a lightweight object detection method, device, equipment, and medium for UAVs that can achieve accurate object detection for the above technical problems.
[0006] A lightweight object detection method for UAVs, the method includes:
[0007] Construct a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to construct a first optimization unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network to a second optimization unit;
[0008] Construct a lightweight model based on the MobileNetV4 network;
[0009] Sparsely train the complex model using the training dataset, discriminate the importance of each channel in the complex model, and prune the secondary channels using channel pruning, and obtain the pruned and optimized complex model through fine-tuning training;
[0010] Use the pruned and optimized complex model as the teacher network, use the lightweight model as the student network, construct a knowledge distillation framework, and adjust the parameters in the lightweight model using the training dataset under the knowledge distillation framework to obtain the trained lightweight model, and use the trained lightweight model as the lightweight object detection network;
[0011] Obtain a real-time image to be subjected to object detection, and perform object detection on the real-time image using the lightweight object detection network.
[0012] In one embodiment, in the first optimization unit:
[0013] Use three parallel dilated convolutional sub-branches with different dilation rates to perform convolution, batch normalization, and activation operations on the input feature map respectively to obtain feature maps of different scales;
[0014] After performing a fusion operation on the feature maps of different scales, obtain a fused feature map, and use the fused feature map as the output data of the first optimization unit.
[0015] In one embodiment, in the second optimization unit:
[0016] Use the first optimization unit to perform preliminary feature extraction on the input feature map to obtain a preliminary feature image;
[0017] After sequentially processing the preliminary feature image through the first C3K_2 and the second C3K_2, splice it with the scale-fused feature image and the data processed by the first C3K_2 in the channel dimension to obtain a spliced image;
[0018] After processing the spliced image through the first optimization unit again, and then processing it using the Transformer unit, obtain the output data of the second optimization unit.
[0019] In one embodiment, in the second optimization unit: replace the first C3K_2 and the second C3K_2 with bottleneck units.
[0020] In one embodiment, the lightweight model includes a backbone network, a neck network, and a detection network connected in sequence;
[0021] The backbone network includes multiple stacked MobileNetV4 network units and an SPPF unit, gradually extracts feature images at different levels from the input image, and inputs the feature images at multiple different levels into the neck network;
[0022] The neck network performs cross-level feature fusion on the feature images at multiple different levels based on the feature pyramid network structure, obtains enhanced key features at different levels, and inputs the multiple enhanced key features into the detection network;
[0023] The detection network uses a deep convolutional network to perform object detection respectively according to the enhanced key features at different levels.
[0024] In one embodiment, the neck network includes an upsampling unit, a channel splicing unit, a Shuffle attention mechanism unit, and the first optimization unit.
[0025] In one embodiment, when adjusting the parameters in the lightweight model using the training dataset under the knowledge distillation framework, the loss function used is expressed as:
[0026] ;
[0027] In the above formula, represents the diagonal distance of the smallest region that can simultaneously contain the predicted box and the ground truth box, and represent the center points of the predicted box and the ground truth box respectively, represents the Euclidean distance between the two center points, represents the weight function, represents the similarity used to measure the aspect ratio.
[0028] This application also provides a lightweight object detection device for a drone, and the device includes:
[0029] A complex model construction module, used to construct a complex model based on YOLOV11. In the backbone network of the complex model, a first optimization unit is constructed using a feature pyramid and dilated convolutions with different dilation rates to replace the CBS unit in the original backbone network, and the C3K_2 unit in the original backbone network is optimized and updated to a second optimization unit using a residual structure and a self-attention mechanism;
[0030] A lightweight model construction module, used to construct a lightweight model based on the MobileNetV4 network;
[0031] A complex model optimization module for sparsely training the complex model using a training data set, discriminating the importance of each channel in the complex model, pruning secondary channels using channel pruning, and obtaining a pruned and optimized complex model through fine-tuning training;
[0032] A lightweight object detection network training module for using the pruned and optimized complex model as a teacher network and the lightweight model as a student network to construct a knowledge distillation framework, adjusting the parameters in the lightweight model using the training data set under the knowledge distillation framework, obtaining a trained lightweight model, and using the trained lightweight model as a lightweight object detection network;
[0033] A real-time object detection module for obtaining a real-time image to be subjected to object detection and performing object detection on the real-time image using the lightweight object detection network.
[0034] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0035] Construct a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to construct a first optimization unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network to a second optimization unit;
[0036] Construct a lightweight model based on the MobileNetV4 network;
[0037] Sparsely train the complex model using the training data set, discriminate the importance of each channel in the complex model, prune secondary channels using channel pruning, and obtain a pruned and optimized complex model through fine-tuning training;
[0038] Use the pruned and optimized complex model as a teacher network and the lightweight model as a student network to construct a knowledge distillation framework, adjust the parameters in the lightweight model using the training data set under the knowledge distillation framework, obtain a trained lightweight model, and use the trained lightweight model as a lightweight object detection network;
[0039] Obtain a real-time image to be subjected to object detection and perform object detection on the real-time image using the lightweight object detection network.
[0040] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0041] Construct a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to construct a first optimization unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network to a second optimization unit;
[0042] Construct a lightweight model based on the MobileNetV4 network;
[0043] Use the training dataset to perform sparse training on the complex model, determine the importance of each channel in the complex model, and use channel pruning to cut off the secondary channels. After fine-tuning training, obtain the pruned and optimized complex model;
[0044] Use the pruned and optimized complex model as the teacher network and the lightweight model as the student network to construct a knowledge distillation framework. Use the training dataset to adjust the parameters in the lightweight model under the knowledge distillation framework to obtain the trained lightweight model, and use the trained lightweight model as the lightweight object detection network;
[0045] Obtain a real-time image to be detected for objects, and use the lightweight object detection network to perform object detection based on the real-time image.
[0046] The above lightweight object detection method, device, equipment and medium for drones, by constructing a complex model based on YOLOV11, using a feature pyramid and dilated convolutions with different dilation rates in the backbone network of the model to construct a first optimization unit to replace the CBS unit in the original backbone network, and using a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network to a second optimization unit, and constructing a lightweight model based on the MobileNetV4 network. After sequentially performing sparse training and channel pruning on the complex model, as the teacher network under the knowledge distillation framework, train the lightweight model as the student network to obtain the lightweight object detection network, and use the lightweight object detection network to achieve real-time detection of objects. This method can be carried on a drone to perform accurate real-time detection of objects. Description of the Drawings
[0047] Figure 1 It is a schematic flow chart of a lightweight object detection method for drones in an embodiment;
[0048] Figure 2 It is a schematic structural diagram of a DCBS unit in an embodiment;
[0049] Figure 3 It is a schematic structural diagram of a first optimization unit in an embodiment;
[0050] Figure 4 Schematic diagram of the structure of the second optimization unit in an embodiment;
[0051] Figure 5 Schematic diagram of the structure of the C3K_2 unit in an embodiment;
[0052] Figure 6 Schematic diagram of the structure of the bottleneck unit in an embodiment;
[0053] Figure 7 Schematic diagram of the structure of the complex model in an embodiment;
[0054] Figure 8 Schematic diagram of the structure of the lightweight model in an embodiment;
[0055] Figure 9 Schematic diagram of the process of channel pruning in an embodiment;
[0056] Figure 10 Schematic diagram of the process of knowledge label distillation training in an embodiment;
[0057] Figure 11 Schematic block diagram of the structure of the lightweight object detection device for drones in an embodiment;
[0058] Figure 12 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0059] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application.
[0060] In view of the existing field of drone object detection, traditional deep learning models usually have high accuracy and performance, but they have a large number of model parameters and are difficult to be directly deployed on resource-constrained drone vision sensors, which leads to the problem of restricting the real-time object detection ability of drones in complex environments. As Figure 1 shown, the present application provides a lightweight object detection method for drones, which specifically includes the following steps:
[0061] Step S100, constructing a complex model based on YOLOV11. In the backbone network of the complex model, a first optimization unit is constructed by using a feature pyramid and dilated convolutions with different dilation rates to replace the CBS unit in the original backbone network, and the C3K_2 unit in the original backbone network is optimized and updated to a second optimization unit by using a residual structure and a self-attention mechanism.
[0062] Step S110: Build a lightweight model based on the MobileNetV4 network.
[0063] Step S120: Use the training dataset to perform sparse training on the complex model, determine the importance of each channel in the complex model, and use channel pruning to cut off the secondary channels. After fine-tuning training, obtain the pruned and optimized complex model.
[0064] Step S130: Use the pruned and optimized complex model as the teacher network and the lightweight model as the student network to build a knowledge distillation framework. Use the training dataset to adjust the parameters in the lightweight model under the knowledge distillation framework to obtain the trained lightweight model, and use the trained lightweight model as the lightweight object detection network.
[0065] Step S140: Obtain the real-time image to be detected, and use the lightweight object detection network to perform object detection based on the real-time image.
[0066] In this embodiment, by optimizing the network structures of the complex model and the lightweight model, under the training of the knowledge distillation framework, a lightweight object detection network with high detection accuracy, low model complexity and low computational cost is obtained. This lightweight object detection network can be directly deployed in the AI chip of the drone terminal to achieve real-time detection of drone targets with high accuracy.
[0067] In step S100, when building the complex model, considering that the teacher model needs to have sufficient learning ability to extract the knowledge information of the target from a large amount of complex data, YOLOV11 (You Only Look Once) is selected as the main framework of the complex model. Optimize the original convolutional layer CBS unit and C3K_2 unit in the backbone network part of the YOLOV11 network.
[0068] In this embodiment, use the first optimization unit to replace the CBS unit. In the first optimization unit: Use three parallel dilated convolutional sub-branches with different dilation rates to perform convolution, batch normalization, and activation operations on the input feature map respectively to obtain feature maps of different scales. After performing a fusion operation on the feature maps of different scales, obtain the fused feature map, and use the fused feature map as the output data of the first optimization unit.
[0069] Specifically, in the first optimization unit, first use the dilated convolution unit (Dilated Conv) to replace the traditional convolution operation in the CBS unit, that is, the DCBS unit, as Figure 2As shown in the figure. Compared with the traditional convolution operation, dilated convolution can expand the receptive field without increasing the number of parameters by introducing the dilation rate parameter. Therefore, it can maintain the resolution and reduce the loss of important information, which is very important for small target detection of drones. Then, based on the DCBS unit, the DCBS_FPN, that is, the first optimization unit, is constructed by combining the feature pyramid structure. As Figure 3 shown, DCBS_FPN combines the feature pyramid and dilated convolutions with different dilation rates, which can effectively realize the fusion of different-scale and different-feature information, capture both local details and global context information, and be more adaptable to complex scenarios.
[0070] Specifically, in DCBS_FPN, p represents padding, d represents the dilation rate, and k represents the convolution kernel size. Among them, the stride in the three parallel dilated convolution sub-branches with different dilation rates is 2, and the dilation rate d is set to 2, 4, and 8 respectively.
[0071] In this embodiment, an attention mechanism is introduced into the original C3K_2 unit to construct the C3K2_2 unit, that is, the second optimization unit. In the second optimization unit, multiple C3K_2 units are included, and combined with the residual structure and bottleneck unit, feature splicing is performed on the input data, which can effectively prevent information loss and retain more detailed and semantic information. The introduction of the self-attention mechanism Transformer has many positive effects on drone target detection, especially showing significant advantages in dealing with key issues such as complex scenarios, small target detection, multi-scale adaptability, and real-time optimization.
[0072] As Figure 4 shown, in the second optimization unit, the first optimization unit is used to perform preliminary feature extraction on the input feature map to obtain a preliminary feature image. After the preliminary feature image is processed by the first C3K_2 and the second C3K_2 in sequence, it is concatenated with the scale-fused feature image and the data processed by the first C3K_2 along the channel dimension to obtain a concatenated image. After the concatenated image is processed by the first optimization unit again, it is then processed by the Transformer unit to obtain the output data of the second optimization unit.
[0073] Specifically, the structure of the C3K_2 unit is as Figure 5 shown, including a bottleneck unit (Bottel neck), a first optimization unit (DCBS_FPN), and a channel dimension concatenation unit (Cat). The input data first passes through the first optimization unit, and then is processed by two cascaded bottleneck units and the first optimization unit respectively. The channel dimension concatenation unit is used to concatenate the two feature images processed separately, and then is processed by the first optimization unit to obtain the output data of the C3K_2 unit.
[0074] Specifically, the structure of the bottleneck unit is as follows Figure 6 shown, including two first optimization units (DCBS_FPN) and an addition and concatenation unit (Add). The input data first enters the first first optimization unit for multi-scale feature fusion processing, and the processing result then flows into the second first optimization unit for further operations. At the same time, when "Shortcut=True", the original input data serves as a shortcut and is directly connected to the addition (Add) operation after the second first optimization unit, adding to the output feature map of the second first optimization unit to finally obtain the output data of the bottleneck unit.
[0075] In this embodiment, in the second optimization unit: two bottleneck units can also be used to replace the first C3K_2 and the second C3K_2.
[0076] As follows Figure 7 shown, is a schematic structural diagram of a complex model. In this embodiment, the proposed complex model includes a backbone network and a neck network. Among them, in the backbone network, the input data first sequentially passes through multiple first optimization units for multi-scale fusion processing of features, then further extracts features through multiple second optimization units, passes through the SPPF unit to enhance the feature expression ability, and finally passes through the C2PSA_2 unit to complete the feature extraction work of the backbone network. In the neck network, the features output by the backbone network enter the neck network. Some features first perform an upsampling (Upsample) operation to increase the resolution, and then are fused with other features through a channel concatenation (Cat) operation, and then are processed through the second optimization unit and the first optimization unit. The features processed by different paths continue to perform multiple channel concatenation and second optimization unit operations, etc. Finally, the three branches respectively output the processed feature maps for the detection task.
[0077] Specifically, in the backbone network, after the input data is first preliminarily processed by the first first optimization unit, the first optimization unit and the second optimization unit are used as a set of processing groups. The preliminarily processed data sequentially passes through two sets of processing groups to obtain a shallow-level feature map. The shallow-level feature map is further processed by a set of processing groups to obtain a middle-level feature map. The middle-level feature map passes through a set of processing groups and then sequentially passes through the SPPF unit and the C2PSA_2 unit to obtain a deep-level feature map. The output shallow-level, middle-level, and deep-level feature maps are input into the neck network. Among them, the shallow-level feature image retains more original detailed information; the middle-level feature image has certain semantic information and structural information; the deep-level feature image is rich in semantic information and is used for high-level tasks such as identifying object categories.
[0078] Specifically, in the neck network, after the deep - level feature image is upsampled and concatenated with the middle - level feature image, the first fused feature is obtained through the second optimization unit. After the first fused feature is upsampled and concatenated with the shallow - level feature image, the shallow - level fused feature image is obtained through the second optimization unit. The shallow - level fused feature is processed by the first optimization unit and concatenated with the first fused feature, and then after being processed by the second optimization unit, the middle - level fused feature image is obtained. The middle - level fused feature image is processed by the first optimization unit and concatenated with the deep - level feature image, and then after being processed by the second optimization unit, the deep - level fused feature image is obtained. Finally, object detection tasks are performed based on the shallow - level fused feature image, the middle - level fused feature image, and the deep - level feature image.
[0079] For the construction of the student network model, considering that the student model needs to be extremely lightweight for easy deployment on subsequent terminal devices, and at the same time, due to the limited computing power of terminal devices, the student model cannot have a large number of parameters and computational amounts.
[0080] In step S110, the MobileNetV4 network is used to construct the feature extraction backbone network of the student model, i.e., the lightweight model. At the same time, the lightweight convolutional attention mechanism Shuffle Attention is incorporated into the feature fusion module of the detection network. Finally, in order to further reduce the number of parameters of the network, depth - wise convolution is used to replace the traditional convolution.
[0081] In this embodiment, the lightweight model includes a backbone network, a neck network, and a detection network connected in sequence. The backbone network includes multiple stacked MobileNetV4 network units and an SPPF unit, which gradually extracts feature images of different levels according to the input image and inputs multiple feature images of different levels into the neck network. The neck network performs cross - level feature fusion on multiple feature images of different levels based on the feature pyramid network structure to obtain enhanced key features of different levels, and inputs multiple enhanced key features into the detection network. The detection network uses a depth - wise convolutional network to perform object detection according to the enhanced key features of different levels respectively.
[0082] As Figure 8 shown, in the backbone network of the lightweight model, the input data first passes through two MobileNetV4 network units in sequence to obtain shallow - level features. The shallow - level features pass through one MobileNetV4 network unit to obtain middle - level features. The middle - level features pass through one MobileNetV4 network unit, an SPPF unit, and one MobileNetV4 network unit in sequence to obtain deep - level features.
[0083] Furthermore, the neck network of the lightweight model includes an upsampling unit, a channel splicing unit, a Shuffle attention mechanism unit, and a first optimization unit. After the deep features are upsampled and channel-spliced with the middle features, the first enhanced features are obtained through the Shuffle attention mechanism unit. The first enhanced features and the shallow features are upsampled and then channel-spliced, and then the shallow enhanced key feature images are obtained through the Shuffle attention mechanism unit. The shallow enhanced key feature images are processed by the first optimization unit and then spliced with the first enhanced features, and then the middle enhanced key feature images are obtained through the processing of the Shuffle attention mechanism unit. The middle enhanced key feature images are processed by the first optimization unit and then spliced with the deep features, and then the deep enhanced key feature images are obtained through the processing of the Shuffle attention mechanism unit.
[0084] Furthermore, in the detection network of the lightweight model, the depthwise convolutional network (DW detection) is used to perform object detection based on the shallow enhanced key feature images, the middle enhanced key feature images, and the deep enhanced key feature images respectively.
[0085] In step S120, the complex model is sparsely trained and channel pruned.
[0086] In this embodiment, the purpose of sparse training is to discriminate the importance of each channel in the feature map of the convolutional layer, similar to the attention mechanism module in the convolutional neural network. Let the output of the previous layer be , m is the batch size of the training samples, and are the mean and variance of each batch of data, is the regularization parameter used to prevent the denominator from being zero during calculation, then the normalized output is:
[0087] ;
[0088] Furthermore, to improve the non-linear learning ability of the complex model, the distribution of the data can be reconstructed through the learnable scaling factor and the translation factor . is the reconstruction result, expressed as:
[0089] ;
[0090] Then, in the form of L1 regularization, the set of values corresponding to each channel in the convolutional layer is sparsely expressed. After sparsification, the feature channels with values tending to 0 are the secondary channels. The regularization formula introduces the loss function \(L\) of the network, and the loss function \(L\) is expressed as:
[0091] ;
[0092] In the above formula, is the loss function of the original convolutional neural network, are the input and target of training, \(W\) is the trainable weight, is the balance factor, is the sparse penalty term for the scaling factor, that is, .
[0093] Furthermore, channel pruning is performed on the complex model after sparse training. Based on the scaling factor after sparse training, it is connected to the corresponding channels in the feature map of the convolutional layer, and the absolute values are sorted. By setting the pruning ratio, the pruning threshold \(S\) is determined. If the corresponding to the feature channel is less than the threshold \(S\), it is determined as a secondary channel and pruned. However, to ensure the integrity of the network structure, the threshold should not be greater than the largest value in any BN layer to ensure matching with the dimensions of the backbone network. At the same time, the structure with residual connections in the network is not pruned to ensure that the dimensions of the diameter connection and the feature map of the residual layer are consistent. The specific process of channel pruning is as Figure 9 shown.
[0094] Finally, the weights are initialized with the pruned complex model for fine-tuning training to restore the detection accuracy of the pruned model, and the optimized complex model is obtained.
[0095] In step S130, the optimized complex model is used as the teacher model, the lightweight model is used as the student model, and a knowledge distillation framework is constructed. The teacher model and the student model are trained in the way of label knowledge distillation training.
[0096] In this embodiment, when training in the way of label knowledge distillation, label knowledge refers to transmitting knowledge through given labels or categories. The training process is as Figure 10 shown, where \(t\) represents the teacher network (i.e., the teacher model), \(s\) represents the student network (i.e., the student model), the hard label is the probability vector obtained without introducing the temperature \(T\) by the softmax function, and the soft label is the probability vector obtained after introducing the temperature \(T\) by the softmax function. Through distillation training, a student model with excellent performance can be obtained.
[0097] In this embodiment, when adjusting the parameters in the lightweight model under the knowledge distillation framework using the training dataset, the loss function used is expressed as:
[0098] ;
[0099] In the above formula, represents the diagonal distance of the smallest region that can simultaneously contain the predicted box and the ground truth box, and represent the center points of the predicted box and the ground truth box respectively, represents the Euclidean distance between the two center points, represents the weight function, represents the similarity used to measure the aspect ratio.
[0100] Furthermore, the weight function is expressed as:
[0101] ;
[0102] Furthermore, the similarity used to measure the aspect ratio is expressed as:
[0103] ;
[0104] In the above formula, and represent the width and height of the predicted box, and represent the width and height of the ground truth box.
[0105] By adjusting the hyperparameter Alpha, the detector can have greater flexibility in achieving different levels of bounding box regression accuracy and is more robust to small datasets and noise. Moreover, Alpha-CIoU can effectively solve the problem that some loss values are zero due to the inability of the predicted box and the ground truth box to overlap, and it has a certain effect on improving the detection performance of the network, shortening the training time, and optimizing the training process of the network.
[0106] In step S140, the student network trained through the above steps is directly deployed on the drone as a lightweight object detection network to perform real-time object detection on the images or videos collected in real time by the drone.
[0107] Specifically, under the correct environment and variable configuration, the target image or video of the drone to be detected is input into the lightweight object detection network, and using the student network weights and the forward propagation algorithm, the target type and location information in the image to be detected are regressed, and finally the high-precision and high-efficiency drone target detection task is achieved.
[0108] It should be noted here that the images or video images used by the drone are all optical images.
[0109] In the above lightweight object detection method for drones, it includes three steps: the channel pruning step, the knowledge distillation step, and the drone object detection step. In the channel pruning step, the complex model is sparsely trained on the drone image dataset, the importance of each channel in the model is judged, and the secondary channels are cropped, and the pruned and optimized complex model is obtained through fine-tuning training. In the knowledge distillation step, the pruned and optimized complex model is used as the teacher network to guide the training of the lightweight student network model, and a distillation loss is constructed on the outputs of the teacher network and the student network, so that the network parameters are continuously trained and updated, and a lightweight drone object detection weight model is obtained. Finally, the trained lightweight student model weight is used to detect the drone object under the drone object detection module. At the same time, this method also creatively proposes the structures of the teacher network and the student network. While lightweighting their network structures, it also ensures the object detection accuracy, enabling the drone to perform object detection on real-time collected images.
[0110] Through model pruning and sparse training, this method can obtain a pruned and optimized complex teacher network. Through knowledge distillation training, the complex teacher network can be used to assist the training of the simple student network, enabling the student network to not only learn the knowledge information in the dataset but also learn the target feature information of the object from the teacher network. And based on the above two networks, a lightweight drone object detection network with high detection accuracy, low model complexity, and low computational complexity can be obtained, meeting the requirements for deployment on the terminal AI chip.
[0111] It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in
[0112] may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps. Figure 11 In one embodiment, as
[0113] The complex model construction module 200 is used to construct a complex model based on YOLOV11. In the backbone network of the complex model, a feature pyramid and dilated convolutions with different dilation rates are used to construct a first optimization unit to replace the CBS unit in the original backbone network, and a residual structure and a self-attention mechanism are used to optimize and update the C3K_2 unit in the original backbone network into a second optimization unit;
[0114] The lightweight model construction module 210 is used to construct a lightweight model based on the MobileNetV4 network;
[0115] The complex model optimization module 220 is used to perform sparsification training on the complex model using a training data set, determine the importance of each channel in the complex model, and use channel pruning to cut off secondary channels, and obtain a pruned and optimized complex model through fine-tuning training;
[0116] The lightweight object detection network training module 230 is used to use the pruned and optimized complex model as a teacher network and the lightweight model as a student network to construct a knowledge distillation framework, and use the training data set to adjust the parameters in the lightweight model under the knowledge distillation framework to obtain a trained lightweight model, and use the trained lightweight model as a lightweight object detection network;
[0117] The real-time object detection module 240 is used to obtain a real-time image to be subjected to object detection, and use the lightweight object detection network to perform object detection based on the real-time image.
[0118] For the specific limitations of the lightweight object detection device for drones, reference can be made to the limitations of the lightweight object detection method for drones in the above text, which will not be elaborated here. Each module in the above lightweight object detection device for drones can be implemented in whole or in part by software, hardware, and their combinations. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0119] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 12As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a lightweight object detection method for unmanned aerial vehicles. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0120] Those skilled in the art can understand that Figure 12 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0121] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0122] Build a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to build a first optimization unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network to a second optimization unit;
[0123] Build a lightweight model based on the MobileNetV4 network;
[0124] Use the training dataset to perform sparse training on the complex model, discriminate the importance of each channel in the complex model, and use channel pruning to cut off the secondary channels. After fine-tuning training, obtain a pruned and optimized complex model;
[0125] Use the pruned and optimized complex model as the teacher network, use the lightweight model as the student network, build a knowledge distillation framework, and use the training dataset to adjust the parameters in the lightweight model under the knowledge distillation framework to obtain a trained lightweight model, and use the trained lightweight model as the lightweight object detection network;
[0126] Obtain a real-time image to be subjected to target detection, and use the lightweight target detection network to perform target detection based on the real-time image.
[0127] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0128] Construct a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to construct a first optimization unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network to a second optimization unit;
[0129] Construct a lightweight model based on the MobileNetV4 network;
[0130] Use a training data set to perform sparse training on the complex model, discriminate the importance of each channel in the complex model, and use channel pruning to cut off secondary channels. After fine-tuning training, obtain a pruned and optimized complex model;
[0131] Use the pruned and optimized complex model as the teacher network, use the lightweight model as the student network, construct a knowledge distillation framework, and use the training data set to adjust the parameters in the lightweight model under the knowledge distillation framework to obtain a trained lightweight model, and use the trained lightweight model as the lightweight target detection network;
[0132] Obtain a real-time image to be subjected to target detection, and use the lightweight target detection network to perform target detection based on the real-time image.
[0133] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0134] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0135] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A lightweight object detection method for drones, characterized in that, The method includes: Construct a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to construct a first optimization unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network to a second optimization unit. Among them, in the second optimization unit: use the first optimization unit to perform preliminary feature extraction on the input feature map to obtain a preliminary feature image. After the preliminary feature image is processed by the first C3K_2 and the second C3K_2 in sequence, it is concatenated with the scale-fused feature image and the data processed by the first C3K_2 along the channel dimension to obtain a concatenated image. After the concatenated image is processed by the first optimization unit again, it is then processed by a Transformer unit to obtain the output data of the second optimization unit; Construct a lightweight model based on the MobileNetV4 network; Use a training dataset to perform sparse training on the complex model, discriminate the importance of each channel in the complex model, and use channel pruning to cut off secondary channels. After fine-tuning training, obtain a pruned and optimized complex model; Use the pruned and optimized complex model as the teacher network and the lightweight model as the student network to construct a knowledge distillation framework. Use the training dataset to adjust the parameters in the lightweight model under the knowledge distillation framework to obtain a trained lightweight model, and use the trained lightweight model as a lightweight object detection network; Obtain a real-time image to be detected, and use the lightweight object detection network to perform object detection based on the real-time image.
2. The lightweight object detection method for an unmanned aerial vehicle according to claim 1, wherein, In the first optimization unit: Use three parallel dilated convolution sub-branches with different dilation rates to perform convolution, batch normalization, and activation operations on the input feature map respectively to obtain feature maps of different scales; After performing a fusion operation on the feature maps of different scales, obtain a fused feature map, and use the fused feature map as the output data of the first optimization unit.
3. The lightweight object detection method for an unmanned aerial vehicle according to claim 2, wherein In the second optimization unit: replace the first C3K_2 and the second C3K_2 with bottleneck units.
4. The lightweight object detection method for an unmanned aerial vehicle according to any one of claims 1-3, characterized in that The lightweight model includes a backbone network, a neck network, and a detection network connected in sequence; The backbone network includes multiple stacked MobileNetV4 network units and an SPPF unit, gradually extracts feature images of different levels from the input image, and inputs the feature images of multiple different levels into the neck network; The neck network performs cross-level feature fusion on the feature images of multiple different levels based on the feature pyramid network structure to obtain enhanced key features of different levels, and inputs the multiple enhanced key features into the detection network; The detection network uses a deep convolutional network to implement object detection based on the enhanced key features of different levels respectively.
5. The lightweight object detection method for an unmanned aerial vehicle according to claim 4, wherein The neck network includes an upsampling unit, a channel concatenation unit, a Shuffle attention mechanism unit, and the first optimization unit.
6. The lightweight object detection method for an unmanned aerial vehicle according to claim 5, wherein When adjusting the parameters in the lightweight model under the knowledge distillation framework using the training dataset, the loss function adopted is expressed as: In the above formula, represents the diagonal distance of the smallest region that can simultaneously contain the predicted box and the ground truth box, and represent the center points of the predicted box and the ground truth box respectively, represents the Euclidean distance between the two center points, represents the weight function, represents the similarity used to measure the aspect ratio.
7. A lightweight target detection device for a drone, characterized in that, The device includes: A complex model construction module, which is used to construct a complex model based on YOLOV11. In the backbone network of the complex model, a first optimization unit is constructed using a feature pyramid and dilated convolutions with different dilation rates to replace the CBS unit in the original backbone network, and the C3K_2 unit in the original backbone network is optimized and updated to a second optimization unit using a residual structure and a self-attention mechanism. Among them, in the second optimization unit: the input feature map is preliminarily feature-extracted by the first optimization unit to obtain a preliminary feature image. After the preliminary feature image is processed by the first C3K_2 and the second C3K_2 in sequence, it is concatenated with the feature image after scale fusion and the data after being processed by the first C3K_2 along the channel dimension to obtain a concatenated image. After the concatenated image is processed by the first optimization unit again, it is processed by a Transformer unit to obtain the output data of the second optimization unit; A lightweight model construction module, which is used to construct a lightweight model based on the MobileNetV4 network; A complex model optimization module, which is used to sparsely train the complex model using the training dataset, distinguish the importance of each channel in the complex model, and prune the secondary channels using channel pruning, and obtain a pruned and optimized complex model after fine-tuning training; A lightweight object detection network training module, which is used to use the pruned and optimized complex model as the teacher network and the lightweight model as the student network to construct a knowledge distillation framework, adjust the parameters in the lightweight model under the knowledge distillation framework using the training dataset to obtain a trained lightweight model, and use the trained lightweight model as the lightweight object detection network; A real-time object detection module, which is used to obtain a real-time image to be detected, and perform object detection on the real-time image using the lightweight object detection network.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Forest fire identification method
CN116469007A
SAR (Synthetic Aperture Radar) image multi-type target detection lightweight method and SAR image multi-type target detection lightweight device based on combination of channel pruning and knowledge distillation
CN116992940A