Lightweight target detection method and device for unmanned aerial vehicle, equipment and medium
By building complex and lightweight models based on YOLOV11 and MobileNetV4 on drones, and through sparse training and knowledge distillation framework optimization, the problem of high complexity of target detection models in the existing technology is solved, and efficient and stable drone target detection is achieved.
Patent Information
- Application Number
- CN202510518881.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing deep learning object detection model has large parameters and high computational complexity, making it difficult to achieve efficient and stable object detection on drones with limited computing resources, resulting in increased power consumption, shortened flight time and detection speed that are difficult to meet real-time requirements.
By constructing complex models based on YOLOV11 and introducing feature pyramids and expansion convolutions with different expansion rates into the backbone network, CBS and C3K_2 units are optimized. At the same time, lightweight models are built based on the MobileNetV4 network, and model complexity is reduced through sparse training and channel pruning. Finally, the lightweight model parameters are adjusted using the knowledge distillation framework to obtain a lightweight target detection network.
Accurate real-time object detection on drones is realized, which reduces model complexity and calculation amount, improves detection efficiency and accuracy, and avoids the problems of increased power consumption and shortened flight time.
Smart Images

Figure CN120047862A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer vision object detection, and particularly to a lightweight object detection method, device, equipment, and medium for unmanned aerial vehicles (UAVs). Background Art
[0002] With the rapid development of UAV technology, it has been widely used in many fields such as security monitoring, logistics distribution, agricultural plant protection, and environmental monitoring. In these application scenarios, UAVs need to detect and identify various target objects in real time and accurately, such as people and vehicles in security monitoring, pest areas and crop growth conditions in agricultural plant protection, etc.
[0003] Currently, common object detection technologies are mainly based on deep learning algorithms, such as object detection models based on convolutional neural networks (CNNs), like Faster R-CNN, YOLO series, etc. These algorithms can achieve relatively ideal detection accuracy in an environment with sufficient computing resources.
[0004] However, due to its own hardware limitations, such as limited computing power and insufficient battery life, UAVs have put forward lightweight requirements for the object detection methods carried. Existing deep learning object detection models have a large number of parameters and high computational complexity. When directly applied to UAVs, it will cause a significant increase in the power consumption of UAVs, seriously shortening the flight time. At the same time, the detection speed is also difficult to meet the real-time requirements, and even situations such as freezing and detection interruption may occur due to exhaustion of computing resources, making it impossible to achieve efficient and stable object detection, which greatly limits the further application and development of UAVs in related fields. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a lightweight object detection method, device, equipment, and medium for UAVs that can achieve accurate object detection.
[0006] A lightweight object detection method for UAVs, the method includes: Construct a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to construct a first optimization unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network into a second optimization unit; Construct a lightweight model based on the MobileNetV4 network; Use a training data set to perform sparsification training on the complex model, discriminate the importance of each channel in the complex model, and use channel pruning to cut off secondary channels, and obtain a pruned and optimized complex model through fine-tuning training; Use the complex model after pruning optimization as the teacher network, and the lightweight model as the student network to construct a knowledge distillation framework. Adjust the parameters in the lightweight model using the training dataset under the knowledge distillation framework to obtain the trained lightweight model, and use the trained lightweight model as the lightweight object detection network; Obtain a real-time image to be subjected to object detection, and use the lightweight object detection network to perform object detection based on the real-time image.
[0007] In one embodiment, in the first optimization unit: Use three parallel dilated convolutional sub-branches with different dilation rates to perform convolution, batch normalization, and activation operations on the input feature map respectively to obtain feature maps of different scales; After performing a fusion operation on the feature maps of different scales, obtain a fused feature map, and use the fused feature map as the output data of the first optimization unit.
[0008] In one embodiment, in the second optimization unit: Use the first optimization unit to perform preliminary feature extraction on the input feature map to obtain a preliminary feature image; After sequentially processing the preliminary feature image through the first C3K_2 and the second C3K_2, splice it with the scale-fused feature image and the data processed by the first C3K_2 along the channel dimension to obtain a spliced image; After processing the spliced image through the first optimization unit again, and then processing it using the Transformer unit, obtain the output data of the second optimization unit.
[0009] In one embodiment, in the second optimization unit: replace the first C3K_2 and the second C3K_2 with bottleneck units.
[0010] In one embodiment, the lightweight model includes a backbone network, a neck network, and a detection network connected in sequence; The backbone network includes multiple stacked MobileNetV4 network units and an SPPF unit, gradually extracts feature images of different levels according to the input image, and inputs the feature images of multiple different levels into the neck network; The neck network performs cross-level feature fusion on the feature images of multiple different levels based on the feature pyramid network structure to obtain enhanced key features of different levels, and inputs the multiple enhanced key features into the detection network; The detection network uses a depth convolutional network to perform object detection respectively according to the enhanced key features of different levels.
[0011] In one embodiment, the neck network includes an upsampling unit, a channel splicing unit, a Shuffle attention mechanism unit, and the first optimization unit.
[0012] In one embodiment, when adjusting the parameters in the lightweight model under the knowledge distillation framework using the training dataset, the loss function used is expressed as: ; In the above formula, represents the diagonal distance of the smallest region that can simultaneously contain the predicted box and the ground truth box, and represent the center points of the predicted box and the ground truth box respectively, represents the Euclidean distance between the two center points, represents the weight function, represents the similarity used to measure the aspect ratio.
[0013] This application also provides a lightweight object detection device for drones. The device includes: A complex model construction module, used to construct a complex model based on YOLOV11. In the backbone network of the complex model, a first optimization unit is constructed using a feature pyramid and dilated convolutions with different dilation rates to replace the CBS unit in the original backbone network, and the C3K_2 unit in the original backbone network is optimized and updated to a second optimization unit using a residual structure and a self-attention mechanism; A lightweight model construction module, used to construct a lightweight model based on the MobileNetV4 network; A complex model optimization module, used to sparsely train the complex model using the training dataset, discriminate the importance of each channel in the complex model, and prune the secondary channels using channel pruning, and obtain a pruned and optimized complex model through fine-tuning training; A lightweight object detection network training module, used to use the pruned and optimized complex model as the teacher network and the lightweight model as the student network to construct a knowledge distillation framework, and adjust the parameters in the lightweight model under the knowledge distillation framework using the training dataset to obtain a trained lightweight model, and use the trained lightweight model as the lightweight object detection network; A real-time object detection module, used to obtain a real-time image to be subjected to object detection, and perform object detection on the real-time image using the lightweight object detection network.
[0014] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: Build a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to construct a first optimized unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network into a second optimized unit; Build a lightweight model based on the MobileNetV4 network; Use the training dataset to perform sparsification training on the complex model, discriminate the importance of each channel in the complex model, and use channel pruning to cut off the secondary channels. After fine-tuning training, obtain a pruned and optimized complex model; Use the pruned and optimized complex model as the teacher network and the lightweight model as the student network to construct a knowledge distillation framework. Use the training dataset to adjust the parameters in the lightweight model under the knowledge distillation framework to obtain a trained lightweight model, and use the trained lightweight model as a lightweight object detection network; Obtain a real-time image to be subjected to object detection, and use the lightweight object detection network to perform object detection based on the real-time image.
[0015] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: Build a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to construct a first optimized unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network into a second optimized unit; Build a lightweight model based on the MobileNetV4 network; Use the training dataset to perform sparsification training on the complex model, discriminate the importance of each channel in the complex model, and use channel pruning to cut off the secondary channels. After fine-tuning training, obtain a pruned and optimized complex model; Use the pruned and optimized complex model as the teacher network and the lightweight model as the student network to construct a knowledge distillation framework. Use the training dataset to adjust the parameters in the lightweight model under the knowledge distillation framework to obtain a trained lightweight model, and use the trained lightweight model as a lightweight object detection network; Obtain a real-time image to be subjected to object detection, and use the lightweight object detection network to perform object detection based on the real-time image.
[0016] The above lightweight object detection method, device, equipment and medium for unmanned aerial vehicles (UAVs) construct a complex model based on YOLOV11. In the backbone network of this model, a feature pyramid and dilated convolutions with different dilation rates are used to construct a first optimized unit to replace the CBS unit in the original backbone network. The C3K_2 unit in the original backbone network is optimized and updated to a second optimized unit using a residual structure and a self-attention mechanism. A lightweight model is constructed based on the MobileNetV4 network. After sparsifying training and channel pruning the complex model in sequence, it is used as a teacher network in a knowledge distillation framework to train the lightweight model as a student network, obtaining a lightweight object detection network, and the lightweight object detection network is used to achieve real-time detection of objects. This method can be carried on a UAV to accurately detect objects in real time. Description of the Drawings
[0017] Figure 1 It is a schematic flow diagram of a lightweight object detection method for UAVs in an embodiment; Figure 2 It is a schematic structural diagram of a DCBS unit in an embodiment; Figure 3 It is a schematic structural diagram of a first optimized unit in an embodiment; Figure 4 It is a schematic structural diagram of a second optimized unit in an embodiment; Figure 5 It is a schematic structural diagram of a C3K_2 unit in an embodiment; Figure 6 It is a schematic structural diagram of a bottleneck unit in an embodiment; Figure 7 It is a schematic structural diagram of a complex model in an embodiment; Figure 8 It is a schematic structural diagram of a lightweight model in an embodiment; Figure 9 It is a schematic process diagram of channel pruning in an embodiment; Figure 10 It is a schematic process diagram of knowledge label distillation training in an embodiment; Figure 11 It is a structural block diagram of a lightweight object detection device for UAVs in an embodiment; Figure 12 It is an internal structural diagram of a computer device in an embodiment. Detailed Embodiments
[0018] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0019] In the existing field of UAV target detection, traditional deep learning models usually have high accuracy and performance, but they have a large number of model parameters and are difficult to be directly deployed on resource-constrained UAV vision sensors, which leads to the problem of restricting the real-time target detection ability of UAVs in complex environments, such as Figure 1 As shown, this application provides a lightweight target detection method for UAVs, which specifically includes the following steps: Step S100, construct a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to construct a first optimization unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network into a second optimization unit.
[0020] Step S110, construct a lightweight model based on the MobileNetV4 network.
[0021] Step S120, perform sparse training on the complex model using the training dataset, discriminate the importance of each channel in the complex model, and use channel pruning to cut off the secondary channels. After fine-tuning training, obtain a pruned and optimized complex model.
[0022] Step S130, use the pruned and optimized complex model as the teacher network and the lightweight model as the student network to construct a knowledge distillation framework. Use the training dataset to adjust the parameters in the lightweight model under the knowledge distillation framework to obtain a trained lightweight model, and use the trained lightweight model as the lightweight target detection network.
[0023] Step S140, obtain a real-time image to be target-detected, and use the lightweight target detection network to perform target detection based on the real-time image.
[0024] In this embodiment, by optimizing the network structures of the complex model and the lightweight model, and through training under the knowledge distillation framework, a lightweight target detection network with high detection accuracy, low model complexity and low computational cost is obtained. This lightweight target detection network can be directly deployed in the UAV terminal AI chip to achieve real-time detection of UAV targets with high accuracy.
[0025] In step S100, when constructing the complex model, considering that the teacher model needs to have sufficient learning ability to extract the knowledge information of the target from a large amount of complex data, YOLOV11 (You Only Look Once) is selected as the main framework of the complex model. Optimize the original convolutional layer CBS unit and C3K_2 unit in the backbone network part of the YOLOV11 network.
[0026] In this embodiment, the CBS unit is replaced by a first optimization unit. In the first optimization unit: three parallel dilated convolutional sub-branches with different dilation rates are used to perform convolution, batch normalization, and activation operations on the input feature map respectively to obtain feature maps of different scales. After performing a fusion operation on the feature maps of different scales, a fused feature map is obtained, and the fused feature map is used as the output data of the first optimization unit.
[0027] Specifically, in the first optimization unit, the dilated convolutional unit (Dilated Conv) is first used to replace the traditional convolution operation in the CBS unit, that is, the DCBS unit, as Figure 2 shown. Compared with the traditional convolution operation, dilated convolution can expand the receptive field without increasing the number of parameters by introducing the dilation rate parameter. Therefore, it can maintain the resolution and reduce the loss of important information, which is very important for small target detection of drones. Then, based on the DCBS unit, the DCBS_FPN, that is, the first optimization unit, is constructed by combining the feature pyramid structure, as Figure 3 shown. DCBS_FPN combines the feature pyramid and dilated convolutions with different dilation rates, can effectively realize the fusion of different-scale and different-feature information, can capture local details and global context information simultaneously, and is more adaptable to complex scenarios.
[0028] Specifically, in DCBS_FPN, p represents padding, d represents the dilation rate size, and k represents the convolution kernel size. Among them, the stride in the three parallel dilated convolutional sub-branches with different dilation rates is 2, and the dilation rate d sizes are set to 2, 4, and 8 respectively.
[0029] In this embodiment, an attention mechanism is introduced into the original C3K_2 unit to construct the C3K2_2 unit, that is, the second optimization unit. In the second optimization unit, multiple C3K_2 units are included, and the residual structure and bottleneck unit are combined to perform feature splicing on the input data, which can effectively prevent information loss and retain more detailed and semantic information. The introduction of the self-attention mechanism Transformer has many positive effects on drone target detection, especially showing significant advantages in dealing with key issues such as complex scenarios, small target detection, multi-scale adaptability, and real-time optimization.
[0030] Such as Figure 4As shown in the figure, in the second optimization unit, the input feature map is initially feature-extracted by the first optimization unit to obtain a preliminary feature image. After the preliminary feature image is processed successively by the first C3K_2 and the second C3K_2, it is concatenated with the feature image after scale fusion and the data after being processed by the first C3K_2 in the channel dimension to obtain a concatenated image. After the concatenated image is processed by the first optimization unit again, it is then processed by the Transformer unit to obtain the output data of the second optimization unit.
[0031] Specifically, the structure of the C3K_2 unit is as Figure 5 shown, and it includes a bottleneck unit (Bottel neck), a first optimization unit (DCBS_FPN), and a channel dimension concatenation unit (Cat). The input data first passes through the first optimization unit, and then is processed by two cascaded bottleneck units and the first optimization unit respectively. The channel dimension concatenation unit is used to concatenate the two feature images that have been processed separately, and then the result is processed by the first optimization unit to obtain the output data of the C3K_2 unit.
[0032] Specifically, the structure of the bottleneck unit (Bottel neck) is as Figure 6 shown, and it includes two first optimization units (DCBS_FPN) and an addition and concatenation unit (Add). The input data first enters the first first optimization unit for multi-scale feature fusion processing, and the processing result then flows into the second first optimization unit for further calculation. At the same time, when "Shortcut=True", the original input data will serve as a shortcut and be directly connected to the addition (Add) operation after the second first optimization unit to be added to the output feature map of the second first optimization unit, and finally the output data of the bottleneck unit is obtained.
[0033] In this embodiment, in the second optimization unit: two bottleneck units can also be used to replace the first C3K_2 and the second C3K_2.
[0034] Such as Figure 7As shown, it is a schematic structural diagram of a complex model. In this embodiment, the proposed complex model includes a backbone network and a neck network. Among them, in the backbone network, the input data first sequentially passes through multiple first optimization units to perform multi-scale fusion processing on the features, then further extracts features through multiple second optimization units, then passes through the SPPF unit to enhance the feature expression ability, and finally passes through the C2PSA_2 unit to complete the feature extraction work of the backbone network. In the neck network, the features output by the backbone network enter the neck network. Some features first perform an upsampling operation to improve the resolution, and then are fused with other features through a channel concatenation (Cat) operation, and then are processed by the second optimization unit and the first optimization unit. The features processed by different paths continue to perform multiple channel concatenation and second optimization unit operations. Finally, the three branches respectively output the processed feature maps for the detection task.
[0035] Specifically, in the backbone network, after the input data is first preliminarily processed by the first first optimization unit, the first optimization unit and the second optimization unit are used as a set of processing groups. The preliminarily processed data sequentially passes through two sets of processing groups to obtain a shallow-level feature map. The shallow-level feature map is further processed by a set of processing groups to obtain a middle-level feature map. The middle-level feature map passes through a set of processing groups and then sequentially passes through the SPPF unit and the C2PSA_2 unit to obtain a deep-level feature map. The output shallow-level, middle-level, and deep-level feature maps are input into the neck network. Among them, the shallow-level feature image retains more original detailed information; the middle-level feature image has certain semantic information and structural information; the deep-level feature image has rich semantic information and is used for high-level tasks such as identifying object categories.
[0036] Specifically, in the neck network, after the deep-level feature image is upsampled and concatenated with the middle-level feature image, it passes through the second optimization unit to obtain a first fusion feature. After the first fusion feature is upsampled and concatenated with the shallow-level feature image, it passes through the second optimization unit to obtain a shallow-level fusion feature image. The shallow-level fusion feature passes through the processing of the first optimization unit and is concatenated with the first fusion feature, and then passes through the processing of the second optimization unit to obtain a middle-level fusion feature image. The middle-level fusion feature image passes through the processing of the first optimization unit and is concatenated with the deep-level feature image, and then passes through the processing of the second optimization unit to obtain a deep-level fusion feature image. Finally, target detection tasks are performed based on the shallow-level fusion feature image, the middle-level fusion feature image, and the deep-level fusion feature image.
[0037] For the construction of the student network model, considering that the student model needs to be extremely lightweight for easy deployment on subsequent terminal devices, and at the same time, due to the limited computing power of terminal devices, the student model cannot have a large number of parameters and computational complexity.
[0038] In step S110, the MobileNetV4 network is used to construct the feature extraction backbone network of the student model, i.e., the lightweight model. At the same time, the lightweight convolutional attention mechanism Shuffle Attention is incorporated into the feature fusion module of the detection network. Finally, in order to further reduce the number of network parameters, depthwise convolution is used to replace traditional convolution.
[0039] In this embodiment, the lightweight model includes a backbone network, a neck network, and a detection network connected in sequence. The backbone network includes multiple stacked MobileNetV4 network units and an SPPF unit, which gradually extracts feature images at different levels from the input image and inputs multiple feature images at different levels into the neck network. The neck network performs cross-level feature fusion on multiple feature images at different levels based on the feature pyramid network structure to obtain enhanced key features at different levels, and inputs multiple enhanced key features into the detection network. The detection network uses a depthwise convolutional network to perform object detection according to enhanced key features at different levels respectively.
[0040] As Figure 8 shown, in the backbone network of the lightweight model, the input data first passes through two MobileNetV4 network units in sequence to obtain shallow features. The shallow features pass through one MobileNetV4 network unit to obtain middle features. The middle features pass through one MobileNetV4 network unit, an SPPF unit, and one MobileNetV4 network unit in sequence to obtain deep features.
[0041] Furthermore, the neck network of the lightweight model includes an upsampling unit, a channel concatenation unit, a Shuffle attention mechanism unit, and a first optimization unit. After the deep features are upsampled and concatenated with the middle features in channels, they pass through the Shuffle attention mechanism unit to obtain the first enhanced feature. The first enhanced feature is upsampled and concatenated with the shallow features in channels, and then passes through the Shuffle attention mechanism unit to obtain the shallow enhanced key feature image. The shallow enhanced key feature image is processed by the first optimization unit and concatenated with the first enhanced feature, and then passes through the Shuffle attention mechanism unit to obtain the middle enhanced key feature image. The middle enhanced key feature image is processed by the first optimization unit and concatenated with the deep features, and then passes through the Shuffle attention mechanism unit to obtain the deep enhanced key feature image.
[0042] Further, in the detection network of the lightweight model, a depth convolution network (DW detection) is used to perform object detection based on the shallow enhanced key feature image, the middle layer enhanced key feature image, and the deep layer enhanced key feature image respectively.
[0043] In step S120, sparse training and channel pruning are performed on the complex model.
[0044] In this embodiment, the purpose of sparse training is to distinguish the importance of each channel in the feature map of the convolutional layer, similar to the attention mechanism module in the convolutional neural network. Let the output of the previous layer be , m is the batch size of the training samples, and are the mean and variance of each batch of data, is the regularization parameter used to prevent the denominator from being zero during calculation, then the normalized output is: ; Further, to improve the non-linear learning ability of the complex model, the distribution of the data can be reconstructed through the learnable scaling factor and the translation factor . is the reconstruction result, expressed as: ; Then, in the form of L1 regularization, the set of values corresponding to each channel in the convolutional layer is sparsely expressed. After sparsification, the feature channels with values tending to 0 are the secondary channels. Introduce the regularization formula of into the loss function L of the network, and the loss function L is expressed as: ; In the above formula, is the loss function of the original convolutional neural network, is the input and target of the training, W is the trainable weight, is the balance factor, is the sparse penalty term for the scaling factor, that is .
[0045] Further, channel pruning is performed on the complex model after sparse training. Based on the scaling factor after sparse training, connect it to the corresponding channels in the feature map of the convolutional layer and sort by taking the absolute value. Determine the pruning threshold S by setting the pruning ratio. If the If it is less than the threshold S, it is determined as a secondary channel and cropping is applied. However, to ensure the integrity of the network structure, the threshold should not be greater than the maximum value in any BN layer to ensure matching with the backbone network dimensions. At the same time, pruning is not performed on the structures with residual connections in the network to ensure that the dimensions of the skip connections and the feature maps of the residual layers are consistent. The specific process of channel pruning is as follows. As shown. Figure 9 As shown.
[0046] Finally, the weights are initialized with the pruned complex model for fine-tuning training to restore the detection accuracy of the pruned model and obtain the optimized complex model.
[0047] In step S130, the optimized complex model is used as the teacher model, the lightweight model is used as the student model, and a knowledge distillation framework is constructed. The teacher model and the student model are trained by means of label knowledge distillation training.
[0048] In this embodiment, when training by means of label knowledge distillation, label knowledge refers to the transmission of knowledge through given labels or categories. The training process is as shown, where t represents the teacher network (i.e., the teacher model), s represents the student network (i.e., the student model), the hard label is the probability vector obtained without introducing the temperature T in the softmax function, and the soft label is the probability vector obtained after introducing the temperature T in the softmax function. A student model with excellent performance can be obtained through distillation training. Figure 10 As shown, where t represents the teacher network (i.e., the teacher model), s represents the student network (i.e., the student model), the hard label is the probability vector obtained without introducing the temperature T in the softmax function, and the soft label is the probability vector obtained after introducing the temperature T in the softmax function. A student model with excellent performance can be obtained through distillation training.
[0049] In this embodiment, when adjusting the parameters in the lightweight model under the knowledge distillation framework using the training dataset, the loss function used is expressed as: ; In the above formula, represents the diagonal distance of the smallest region that can simultaneously contain the predicted box and the ground truth box, and represent the center points of the predicted box and the ground truth box respectively, represents the Euclidean distance between the two center points, represents the weight function, represents the similarity used to measure the aspect ratio.
[0050] Furthermore, the weight function is expressed as: ; Furthermore, the similarity used to measure the aspect ratio is expressed as: ; In the above formula, and represent the width and height of the predicted bounding box, and represent the width and height of the ground-truth bounding box.
[0051] By adjusting the hyperparameter Alpha, the detector can have greater flexibility in achieving different levels of bounding box regression accuracy and is more robust to small datasets and noise. Moreover, Alpha-CIoU can effectively solve the problem that some loss values are zero due to the inability of the predicted bounding box and the ground-truth bounding box to overlap, and it has a certain effect on improving the detection performance of the network, shortening the training time, and optimizing the training process of the network.
[0052] In step S140, the student network trained through the above steps is directly deployed on the drone as a lightweight object detection network to perform real-time object detection on the images or videos collected in real time by the drone.
[0053] Specifically, under the correct configuration of the environment and variables, the target image or video of the drone to be detected is input into the lightweight object detection network. Using the student network weights and the forward propagation algorithm, the target type and location information in the image to be detected are regressed, and finally, a high-precision and high-efficiency drone target detection task is achieved.
[0054] It should be noted here that the images or video images used by the drone are all optical images.
[0055] The above lightweight object detection method for drones includes three steps: the channel pruning step, the knowledge distillation step, and the drone target detection step. In the channel pruning step, the complex model is sparsely trained on the drone image dataset to judge the importance of each channel in the model and prune the secondary channels, and a pruned and optimized complex model is obtained through fine-tuning training. In the knowledge distillation step, the pruned and optimized complex model is used as the teacher network to guide the training of the lightweight student network model, and a distillation loss is constructed on the outputs of the teacher network and the student network, so that the network parameters are continuously trained and updated to obtain a lightweight drone target detection weight model. Finally, the trained lightweight student model weights are used to detect the drone targets in the drone target detection module. At the same time, this method also creatively proposes the structures of the teacher network and the student network. While lightweighting their network structures, it also ensures the object detection accuracy, enabling the drone to perform object detection on the images collected in real time.
[0056] Through model pruning and sparse training, this method can obtain a complex teacher network optimized by pruning. Through knowledge distillation training, the complex teacher network can be used to assist the training of a simple student network, enabling the student network to not only learn the knowledge information in the dataset but also learn the target feature information from the teacher network. Based on the above two networks, a lightweight UAV target detection network with high detection accuracy, low model complexity, and low computational cost can be obtained, meeting the requirements for deployment on terminal AI chips.
[0057] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,
[0058] In one embodiment, as Figure 11 shown, a lightweight target detection device for UAVs is provided, including: a complex model construction module 200, a lightweight model construction module 210, a complex model optimization module 220, a lightweight target detection network training module 230, and a real-time target detection module 240, where: The complex model construction module 200 is used to construct a complex model based on YOLOV11. In the backbone network of the complex model, a first optimization unit is constructed using a feature pyramid and dilated convolutions with different dilation rates to replace the CBS unit in the original backbone network, and the C3K_2 unit in the original backbone network is optimized and updated to a second optimization unit using a residual structure and a self-attention mechanism; The lightweight model construction module 210 is used to construct a lightweight model based on the MobileNetV4 network; The complex model optimization module 220 is used to perform sparse training on the complex model using a training dataset, determine the importance of each channel in the complex model, and prune the secondary channels using channel pruning, and obtain a pruned and optimized complex model through fine-tuning training; The lightweight object detection network training module 230 is used to construct a knowledge distillation framework by using the pruned and optimized complex model as the teacher network and the lightweight model as the student network, and adjust the parameters in the lightweight model under the knowledge distillation framework by using the training dataset to obtain the trained lightweight model, and use the trained lightweight model as the lightweight object detection network; The real-time object detection module 240 is used to obtain a real-time image to be subjected to object detection, and perform object detection on the real-time image by using the lightweight object detection network.
[0059] For the specific limitations of the lightweight object detection device for drones, reference can be made to the limitations of the lightweight object detection method for drones in the above text, which will not be elaborated here. Each module in the above lightweight object detection device for drones can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0060] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 12 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a lightweight object detection method for drones. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0061] Those skilled in the art can understand that Figure 12 the structure shown in
[0062] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented: Build a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to build a first optimization unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network to a second optimization unit; Build a lightweight model based on the MobileNetV4 network; Use the training dataset to perform sparsification training on the complex model, discriminate the importance of each channel in the complex model, and use channel pruning to cut off the secondary channels. After fine-tuning training, obtain a pruned and optimized complex model; Use the pruned and optimized complex model as the teacher network, use the lightweight model as the student network, build a knowledge distillation framework, use the training dataset to adjust the parameters in the lightweight model under the knowledge distillation framework, obtain a trained lightweight model, and use the trained lightweight model as a lightweight object detection network; Obtain a real-time image to be subjected to object detection, and use the lightweight object detection network to perform object detection based on the real-time image.
[0063] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: Build a complex model based on YOLOV11. In the backbone network of the complex model, use a feature pyramid and dilated convolutions with different dilation rates to build a first optimization unit to replace the CBS unit in the original backbone network, and use a residual structure and a self-attention mechanism to optimize and update the C3K_2 unit in the original backbone network to a second optimization unit; Build a lightweight model based on the MobileNetV4 network; Use the training dataset to perform sparsification training on the complex model, discriminate the importance of each channel in the complex model, and use channel pruning to cut off the secondary channels. After fine-tuning training, obtain a pruned and optimized complex model; Use the pruned and optimized complex model as the teacher network, use the lightweight model as the student network, build a knowledge distillation framework, use the training dataset to adjust the parameters in the lightweight model under the knowledge distillation framework, obtain a trained lightweight model, and use the trained lightweight model as a lightweight object detection network; Obtain a real-time image to be subjected to object detection, and use the lightweight object detection network to perform object detection based on the real-time image.
[0064] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0065] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0066] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A lightweight target detection method for unmanned aerial vehicles, characterized in that: The method comprises: A complex model is built based on YOLOV11. In the backbone network of the complex model, the feature pyramid and dilated convolutions with different dilation rates are used to construct the first optimization unit to replace the CBS unit in the original backbone network, and the residual structure and self-attention mechanism are used to optimize and update the C3K_2 unit in the original backbone network to the second optimization unit. Build a lightweight model based on the MobileNetV4 network; The complex model is sparsely trained using the training data set to determine the importance of each channel in the complex model, and channel pruning is used to prune the secondary channels, and the pruned and optimized complex model is obtained through fine-tuning training; The complex model after pruning optimization is used as the teacher network, and the lightweight model is used as the student network, a knowledge distillation framework is constructed, and the parameters of the lightweight model are adjusted under the knowledge distillation framework using the training data set to obtain the trained lightweight model, and the trained lightweight model is used as the lightweight object detection network; A real-time image to be subjected to target detection is acquired, and the lightweight target detection network is used to perform target detection according to the real-time image.
2. The lightweight target detection method for unmanned aerial vehicles according to claim 1, characterized in that: In the first optimization unit: Using three parallel dilated convolution sub-branches with different dilation rates, convolution, batch normalization and activation operations are performed on the input feature map to obtain feature maps of different scales. After the feature maps of different scales are fused, a fused feature map is obtained, and the fused feature map is used as output data of the first optimization unit.
3. The lightweight target detection method for unmanned aerial vehicles according to claim 2, characterized in that: In the second optimization unit: Using the first optimization unit to perform preliminary feature extraction on the input feature map to obtain a preliminary feature image; After the preliminary feature image is processed by the first C3K_2 and the second C3K_2 in sequence, it is spliced with the feature image after scale fusion and the data processed by the first C3K_2 according to the channel dimension to obtain a spliced image; The stitched image is processed again by the first optimization unit and then processed by the Transformer unit to obtain output data of the second optimization unit.
4. The lightweight target detection method for unmanned aerial vehicles according to claim 3, characterized in that: In the second optimization unit: the first C3K_2 and the second C3K_2 are replaced by a bottleneck unit.
5. The lightweight target detection method for a drone according to any one of claims 1 to 4, characterized in that: The lightweight model includes a backbone network, a neck network, and a detection network connected in sequence; The backbone network includes multiple stacked MobileNetV4 network units and SPPF units, which gradually extract feature images of different levels according to the input image, and input the multiple feature images of different levels into the neck network; The neck network performs cross-level feature fusion on feature images of multiple different levels based on a feature pyramid network structure to obtain enhanced key features of different levels, and inputs the multiple enhanced key features into the detection network; The detection network uses a deep convolutional network to achieve target detection according to enhanced key features at different levels.
6. The lightweight target detection method for unmanned aerial vehicles according to claim 5, characterized in that: The neck network includes an upsampling unit, a channel splicing unit, a Shuffle attention mechanism unit and the first optimization unit.
7. The lightweight target detection method for unmanned aerial vehicles according to claim 6, characterized in that: When adjusting the parameters in the lightweight model using the training data set under the knowledge distillation framework, the loss function used is expressed as: In the above formula, It represents the diagonal distance of the minimum area that can contain both the predicted box and the real box. and Represent the center points of the predicted box and the real box respectively, represents the Euclidean distance between two center points, represents the weight function, Represents the similarity used to measure aspect ratio.
8. A lightweight target detection device for an unmanned aerial vehicle, characterized in that: The device comprises: A complex model construction module is used to build a complex model based on YOLOV11. In the backbone network of the complex model, the feature pyramid and dilated convolutions with different dilation rates are used to construct the first optimization unit to replace the CBS unit in the original backbone network, and the residual structure and self-attention mechanism are used to optimize and update the C3K_2 unit in the original backbone network to the second optimization unit; Lightweight model building module, used to build a lightweight model based on the MobileNetV4 network; A complex model optimization module is used to perform sparse training on the complex model using a training data set, determine the importance of each channel in the complex model, and use channel pruning to prune secondary channels, and obtain a pruned and optimized complex model through fine-tuning training; A lightweight target detection network training module is used to use the pruned and optimized complex model as a teacher network and the lightweight model as a student network to construct a knowledge distillation framework, use the training data set to adjust the parameters in the lightweight model under the knowledge distillation framework to obtain a trained lightweight model, and use the trained lightweight model as a lightweight target detection network; The real-time target detection module is used to obtain a real-time image to be subjected to target detection, and to perform target detection based on the real-time image using the lightweight target detection network.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Forest fire identification method
CN116469007A
SAR (Synthetic Aperture Radar) image multi-type target detection lightweight method and SAR image multi-type target detection lightweight device based on combination of channel pruning and knowledge distillation
CN116992940A
Lightweight pest detection method based on improved YOLOv7 and RKNPU2
CN117058552A
Substation foreign matter intrusion detection method and system based on improved YOLOv11 model
CN119722662A
Efficient fire smoke detection method based on lightweight convolution and cross channel feature fusion
CN119741646A
Cited By
Physical object traversal recognition detection method and device based on unmanned aerial vehicle equipment
CN120688985A
Coal mining machine roller tracking system and method
CN121010931A
Construction scene anomaly detection method and system based on unmanned aerial vehicle and lightweight model
CN121147578A