Rapid target detection and tracking method for unmanned aerial vehicle in complex environment

By collecting multimodal images on the drone and building a lightweight target detection and tracking model, the problem of drone target detection and tracking in complex environments is solved, and efficient and stable target detection and tracking is achieved, which is suitable for resource limitations of drones.

CN119992059AActive Publication Date: 2025-05-13JIANGSU WANWEI AISI NETWORK INTELLIGENT IND INNOVATION CENT
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510100434.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

In complex environments, the drone images have characteristics such as changing perspectives, variable scales, small targets or even extremely small targets, blurred motion, and complex lighting, resulting in object detection and tracking problems. At the same time, the drone's limited computing power and power carrying capacity limit the deployment of deep learning models.

Method used

A drone equipped with an optoelectronic pod collects visible light and infrared images, generates multimodal fusion images through an enhanced dual-scale image fusion algorithm, and builds a multimodal fusion target detection model based on Yolov5, and lightens the model through channel pruning, layer pruning and quantization operations. Combined with a SimaFc-based airborne end target tracking model, target detection and tracking are achieved.

Benefits of technology

In complex environments, the target detection accuracy and tracking stability of the drone are improved, the calculation load and power consumption are reduced, the model inference speed is improved, and the limited computing power and power resources of the drone are suitable for the finite computing power and power resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992059A_ABST
    Figure CN119992059A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle rapid target detection and tracking method in a complex environment. Obtaining a multi-modal fusion image M through an enhanced dual-scale image fusion algorithm; a multi-modal fusion target detection model based on Yolov5 is constructed, and pruning is carried out; an airborne end target tracking model based on SimaFc is constructed, and pruning is carried out; inputting the image M into a Yov5-based multi-modal fusion target detection model for processing to obtain an initial target detection result Y1; cutting from the image M according to the Y1 to generate a template image set Y2; according to the Y1, a large area intercepted near the target position determined in the image M is used as a search area Y3 of the tracking target; and inputting the Y2 and the Y3 into two convolutional neural network branches of an airborne end target tracking model based on SimaFc for feature extraction, and judging the part most similar to the template image in the search area by using similarity and distance measurement, thereby realizing target tracking. The target detection and tracking method provided by the scheme of the invention has relatively high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and in particular relates to a method for rapid target detection and tracking of unmanned aerial vehicles in complex environments. Background Art

[0002] With the advancement of science and technology, the application of drones is becoming more and more extensive. By equipping drones with high-definition cameras, combined with flight control systems and related algorithms, drones can recognize, detect and track photographed objects. At present, the domestic drone market has been developing for nearly 30 years. Domestic and foreign companies and institutions are actively researching and promoting this technology, and the application of drone target detection and tracking systems is also constantly expanding and developing. At present, domestic drone manufacturers have increased their investment in target detection and tracking system technology and have achieved certain research results. At the same time, some professional drone technology companies have also made great achievements in this field, and in civil fields such as security, disaster relief, agriculture and animal husbandry, electricity, aerial photography and mapping, various types of drones are flourishing.

[0003] Deeply combining drone technology with artificial intelligence technology, and using artificial intelligence technology and machine vision technology to give drone technology an intelligent core, can greatly improve the combat performance of drone technology in civil and military fields, while further liberating manpower. By equipping the drone platform with visual perception equipment, combined with the current intelligent perception technology and the lightweight and flexible characteristics of drones, the perception of the scene environment and targets in the environment can be effectively and comprehensively expanded.

[0004] In the field of UAV visual perception, the detection, recognition, tracking and positioning of targets in the environment is the basis of many tasks. For example, the perception of targets in a battlefield environment is the basis for subsequent target strikes and situation analysis tasks. Therefore, the study of UAV target detection has very important research significance.

[0005] In the field of drone image target detection, the current methods for drone vision target detection and recognition are mostly based on deep learning, mainly divided into Anchor-based and Anchor-free. The Anchor-based methods are divided into One-stage and Two-stage. In recent years, the Two-stage methods include Faster-RCNN, FPN, Light-RCNN and Cascade-RCNN, and the One-stage methods include YOLOv3, SSD, RetinaNet, RefineDet, etc. For the Two-stage method, the target positioning module and the target classification module are implemented independently and combined in sequence to jointly realize the target detection and recognition task. Therefore, the operation process of this type of algorithm goes through two stages: target positioning stage and target classification stage. Compared with the One-stage, the accuracy is better, but the speed is slower. The Anchor-free method avoids the problem that the manually designed target box based on the Anchor-based method is sensitive to the target size in different data sets, and also reduces the large amount of computing consumption generated by the regression calculation of the target box. However, when this method is applied to target detection in UAV vision, there are still some problems with UAV images, such as variable perspectives, variable scales, small or even extremely small targets, motion blur, and complex lighting conditions. It cannot directly solve the inherent problems of UAV images.

[0006] In the field of drone visual target tracking, the target tracking tasks for drone vision can be roughly divided into image target detection and tracking, video object detection and tracking, single target tracking and multi-target tracking. Image target tracking means that given a set of predefined object classes, such as aircraft, vehicles, and personnel, the algorithm will detect objects of these classes from a single image taken by the drone. From this point of view, image target detection and tracking is also image target detection. Video object detection and tracking is similar to image target detection and tracking. The algorithm will detect objects of predefined object classes from the video taken by the drone. Single-label tracking is to continuously track a single target in a video stream. There are currently three main methods for target tracking: one is based on correlation filters, such as ECO, C-COT, BACF, etc.; the second is based on semantic networks, such as SiamRPN++, Siam R-CNN, SiamMask, etc.; the third is based on CNN, such as MDNet, CFNet, ATOM, etc. For drone target tracking tasks, they will be affected by sudden motion, low resolution (small targets), complex lighting and occlusion, which may easily cause the drone to fail to accurately complete target tracking. This method cannot directly solve the inherent problems of drone images.

[0007] In view of the above problems, it is of great significance and value to study the realization of rapid target detection and tracking by drones based on images collected by drones equipped with optoelectronic pods. When drones perform detection and tracking tasks, the power supply of drones needs to support the power consumption of drones and the computing power in the field of machine vision technology. Therefore, it is difficult for drone computing power to bear the huge deep learning computing needs. Therefore, for the deep learning computing model on the drone airborne end, the neural network model needs to be lightweight to reduce the computing load and improve the computing speed. Summary of the invention

[0008] The purpose of the invention is to provide a method for rapid target detection and tracking of drones in complex environments. Based on the images collected by drones equipped with optoelectronic pods, a multimodal fusion algorithm based on infrared + visible light is used to achieve small target detection, and on this basis, continuous tracking of target objects is achieved, and finally the lightweight deployment of the model is achieved in combination with the power consumption of the drone and the computing power carrying capacity. In particular, the present invention mainly solves the problem of target detection and tracking of drone images in situations with variable perspectives, variable scales, small or even extremely small targets, motion blur, complex lighting, target occlusion, etc., as well as the problem of model deployment under the conditions of limited computing resources and power carrying capacity of drones.

[0009] Technical solution, a method for rapid target detection and tracking of unmanned aerial vehicles in complex environments, including the following steps:

[0010] Step S1, using a drone equipped with an optoelectronic pod to collect images, including visible light images and infrared images, and then performing data augmentation processing to expand the original available data volume and generate occluded target detection data type;

[0011] Step S2, converting the visible light image and the infrared image into a visible light video frame image M1 with a width of W and a height of H, and an infrared video frame M2 with a width of W and a height of H, and obtaining a multimodal fused image M through an enhanced dual-scale image fusion algorithm;

[0012] Step S3, construct a multimodal fusion target detection model based on Yolov5, and perform channel pruning and layer pruning on the multimodal fusion target detection model based on Yolov5 to ensure that the pruned target detection model can maintain a high detection accuracy while reducing the amount of calculation, and quantize the model in combination with the actual situation of hardware resources to further compress the model and improve the model reasoning speed;

[0013] Step S4, constructing an airborne target tracking model based on SimaFc, and performing path pruning and layer pruning on the above-mentioned airborne target tracking model based on SimaFc, to ensure that the pruned target tracking model will not have a significant impact on the target tracking effect while reducing the amount of calculation, and combining the actual situation of hardware resources, quantizing the model, further compressing the model, and improving the model reasoning speed;

[0014] Step S5, inputting the multimodal fusion image M obtained in step S2 into the multimodal fusion target detection model based on Yolov5 constructed in step S3 for processing to obtain an initial target detection result Y1;

[0015] Step S6, according to the target detection result Y1 obtained in step S5, crop from the multimodal fusion image M with a width of W and a height of H to generate a template image set Y2 containing only the target, where the set Y2 contains a series of initial target template images [m1, m2, ..., mn] contained in the image M';

[0016] Step S7, according to the target detection result Y1 obtained in step S5, a larger area near the target position determined in the multimodal fusion image M with a width of W and a height of H is intercepted as a search area for tracking the target, and the search area is recorded as a set Y3, and the set Y3 is a series of search area images [n1, n2, ..., nn] contained in the image M, corresponding to the template image set Y2;

[0017] Step S8: Input the template image set Y2 and the search area image set Y3 obtained in step S7 into the two convolutional neural network branches of the SimaFc-based airborne target tracking model constructed in step S4 for feature extraction, and use similarity and distance metrics to determine the part in the search area that is most similar to the template image, thereby achieving target tracking.

[0018] According to a further improvement of the present invention, the enhanced dual-scale image fusion algorithm in step S2 comprises the following steps:

[0019] Step S21: Input two images, one is a visible light image and the other is a corresponding infrared image. and Indicates that bilateral filtering is used to decompose the two images into a base layer and a detail layer respectively;

[0020] Step S22, the decomposed base layer and detail layer are fused respectively, wherein the base layer is fused by a weighted average method, and the detail layer is fused by a weighted average method using a fusion coefficient matrix constructed by using a significant feature map;

[0021] Step S23: After obtaining the fused base layer and detail layer, the base layer and detail layer are fused by simple addition, that is, the channel elements corresponding to the visible light image and the infrared image are concat-operated to obtain the final multimodal fusion data M.

[0022] According to a further improvement of the present invention, the multimodal fusion target detection model based on Yolov5 constructed in step S3 is an improved target detection model, including:

[0023] Step S31, construct the backbone network Backbone: It is composed of Darknet and CSP structure, and extracts features at different levels through layer-by-layer convolution, activation, batch normalization and other operations to provide rich low-level and high-level features for subsequent detection modules;

[0024] Step S32, constructing a feature fusion layer Neck: adopting a path aggregation network: structure, combining a feature pyramid network and a path aggregation network, and through multi-layer feature fusion, making the model more suitable for detecting targets of different sizes, and enhancing the recognition ability of small targets and long-distance targets;

[0025] Step S33, constructing the detection head: generating different Anchor boxes for each feature layer, and calculating the probability of the target category and the bounding box position for each Anchor. Three scales are usually used to predict large, medium and small targets respectively. The specific bounding box information of the generated feature map includes position, size, confidence and category. Through the non-maximum suppression process, the overlapping boxes are removed and the predicted box with the highest confidence is retained to obtain the final detection result. The C3-RS module is used to replace the C3 module in the original structure, R refers to the Bottleneck residual block in ResNet, and combined with the SE module, the channel attention is more focused and the recognition ability of small targets is improved.

[0026] According to a further improvement of the present invention, the SimaFc-based airborne target tracking model constructed in step S4 is an improved target tracking model, including:

[0027] Step S41, constructing a backbone network Backbone: based on MobileNet, it consists of two convolutional neural network branches with shared weights, and extracts features from the template image and the search area respectively;

[0028] Step S42, constructing a similarity measurement module: matching the template features with the search area features through a mutual correlation layer, and measuring the similarity between the two;

[0029] Step S43, constructing a response map generation module: generating a response map, i.e., a similarity distribution map, from the result of the cross-correlation, and locating the most likely position of the target by comparing the response values ​​at different positions in the response map. The peak position in the response map, also called the maximum response value, is the target position prediction;

[0030] Step S44, constructing an occlusion detection module: by calculating the change pattern in the response graph, when it is detected that the similarity suddenly drops sharply or becomes blurred, it is determined whether the target is occluded; when it is determined that occlusion has occurred, the occlusion detection module will trigger a relocation strategy to relocate the target through the response graph of the previous frame and the historical trajectory of the target;

[0031] Step S45: construct a target positioning and updating module, adjust the position of the target's bounding box according to the peak position of the response graph, and update the model based on the confidence level;

[0032] Step S46, constructing a post-processing module: performing smoothing and screening operations to improve the accuracy and stability of tracking in complex environments.

[0033] According to a further improvement of the present invention, in step S3, the multimodal fusion target detection model based on Yolov5 is pruned and quantized by a model lightweight method based on joint channel pruning and layer pruning, which includes the following steps:

[0034] S3a, channel pruning: customize the pruning script and implement basic channel pruning in models / common.py in the Yolov5 code framework;

[0035] S3b, layer pruning: The contribution of a layer is evaluated by measuring the importance of the layer, and the layers with low contribution are pruned first. Layer pruning is implemented based on the sparsity of the feature map. By counting the proportion of non-zero values ​​or the average absolute value in the activation output, the layer with high sparsity is found. That is, the feature map output of the layer is very sparse or even close to zero, indicating that the layer has a small effect and is the layer that needs to be pruned.

[0036] S3c, model fine-tuning: After channel pruning and layer pruning operations, the model is fine-tuned and trained to restore the accuracy of the model;

[0037] S3d, model quantization: quantize the weights in floating point FP32 format to integer INT8 format, and use a small number of real scene samples as calibration data, input into the model, collect the maximum and minimum value information of the activation layer, and then quantize the activation and weight of each layer in the model to INT8 representation according to the sampling range.

[0038] According to a further improvement of the present invention, in step S4, the SimaFc-based airborne target tracking model is pruned and quantized by a model lightweight method based on joint channel pruning and layer pruning, comprising the following steps:

[0039] Step S4a, channel pruning: by calculating the batch normalization layer scaling factor γ parameter of each channel, the scaling factor represents the activation amplitude of the channel, and judging the importance of the channel by analyzing these values; wherein, the larger the γ value is, the more important the feature of the channel is, and the smaller the γ value is, the smaller the influence of the channel on the feature in the model is. Channels with lower weights are selected as pruning candidates, and a certain proportion of channels are retained based on the set pruning rate, and 80% of the total number of channels are retained. After pruning, the number of channels of the model is adjusted to ensure that the input dimension of the downstream convolutional layer is consistent with the output dimension;

[0040] Step S4b, layer pruning: Evaluate the impact of deleting layers on model accuracy layer by layer, perform sensitivity analysis on each convolutional layer of MobileNet, and record the impact of each convolutional layer on model accuracy; use the validation set to perform a quick inference test to obtain model accuracy. If a layer has little impact on accuracy, it will be marked as a pruning candidate; the impact on accuracy is determined by setting an accuracy threshold. If the accuracy drops by less than 1% after removing a layer, it is considered that the layer can be removed; according to the results of the sensitivity analysis, select redundant layers from the deep convolutional layers with the least impact for pruning;

[0041] Step S4c, model fine-tuning: After performing channel pruning and layer pruning operations, fine-tune the model to restore the accuracy of the model;

[0042] Step S4d, model quantization: quantize the weights in floating point FP32 format to integer INT8 format, and use a small number of real scene samples as calibration data, input them into the model, collect the maximum and minimum value information of the activation layer, and then quantize the activation and weight of each layer in the model to INT8 representation according to the sampling range.

[0043] Beneficial effects: The method for rapid target detection and tracking of unmanned aerial vehicles in complex environments proposed by the present invention is a complete set of feasible solutions including target detection and tracking and model lightweight deployment; the multimodal fusion target detection method based on Yolov5 proposed by the present invention uses multimodal fusion images as input, utilizes the feature complementarity in visible light and infrared images, and combines the C3-Res-S module convolution design to enhance the detection capability of targets, especially small targets; the airborne target tracking method based on SimaFc proposed by the present invention uses a twin network structure with two branches, and adopts a lightweight convolutional network as a feature extractor to improve computational efficiency, ensure the real-time performance of target tracking, and enhances robustness by combining the occlusion detection module and multimodal information; the model lightweight method based on joint channel pruning and layer pruning proposed by the present invention is specially optimized for the deployment of resource-constrained airborne equipment, taking into account the computing power, storage space and real-time requirements of the equipment. Reduce model parameters and calculation amount, while avoiding excessive precision loss. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is the flow chart of the dual-scale image fusion algorithm of this invention.

[0045] Figure 2 It is a flow chart of the multi-channel data fusion method of the present invention.

[0046] Figure 3 This is a structural diagram of the multimodal fusion target detection model based on Yolov5 in the present invention.

[0047] Figure 4 It is a structural diagram of the airborne target tracking model based on SimaFc of the present invention.

[0048] Figure 5 It is a comparison chart of the results before and after compression and optimization of the model of the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention is clearly and completely described. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] Example 1

[0051] The method for rapid target detection and tracking of unmanned aerial vehicles in complex environments includes the following steps:

[0052] Step S1, using a drone equipped with an optoelectronic pod to collect images, including visible light images and infrared images, and then performing data augmentation processing to expand the original available data volume and generate occluded target detection data type;

[0053] Step S2, converting the visible light image and the infrared image into a visible light video frame image M1 with a width of W and a height of H, and an infrared video frame M2 with a width of W and a height of H, and obtaining a multimodal fused image M through an enhanced dual-scale image fusion algorithm;

[0054] Step S3, construct a multimodal fusion target detection model based on Yolov5, and perform channel pruning and layer pruning on the multimodal fusion target detection model based on Yolov5 to ensure that the pruned target detection model can maintain a high detection accuracy while reducing the amount of calculation, and quantize the model in combination with the actual situation of hardware resources to further compress the model and improve the model reasoning speed;

[0055] Step S4, constructing an airborne target tracking model based on SimaFc, and performing path pruning and layer pruning on the above-mentioned airborne target tracking model based on SimaFc, to ensure that the pruned target tracking model will not have a significant impact on the target tracking effect while reducing the amount of calculation, and combining the actual situation of hardware resources, quantizing the model, further compressing the model, and improving the model reasoning speed;

[0056] Step S5, inputting the multimodal fusion image M obtained in step S2 into the multimodal fusion target detection model based on Yolov5 constructed in step S3 for processing to obtain an initial target detection result Y1;

[0057] Step S6, according to the target detection result Y1 obtained in step S5, crop from the multimodal fusion image M with a width of W and a height of H to generate a template image set Y2 containing only the target, where the set Y2 contains a series of initial target template images [m1, m2, ..., mn] contained in the image M';

[0058] Step S7, according to the target detection result Y1 obtained in step S5, a larger area near the target position determined in the multimodal fusion image M with a width of W and a height of H is intercepted as a search area for tracking the target, and the search area is recorded as a set Y3, and the set Y3 is a series of search area images [n1, n2, ..., nn] contained in the image M, corresponding to the template image set Y2;

[0059] Step S8: Input the template image set Y2 and the search area image set Y3 obtained in step S7 into the two convolutional neural network branches of the SimaFc-based airborne target tracking model constructed in step S4 for feature extraction, and use similarity and distance metrics to determine the part in the search area that is most similar to the template image, thereby achieving target tracking.

[0060] According to a further improvement of the present invention, the enhanced dual-scale image fusion algorithm in step S2 comprises the following steps:

[0061] Step S21: Input two images, one is a visible light image and the other is a corresponding infrared image. and Indicates that bilateral filtering is used to decompose the two images into a base layer and a detail layer respectively;

[0062] Step S22, the decomposed base layer and detail layer are fused respectively, wherein the base layer is fused by a weighted average method, and the detail layer is fused by a weighted average method using a fusion coefficient matrix constructed by using a significant feature map;

[0063] Step S23: After obtaining the fused base layer and detail layer, the base layer and detail layer are fused by simple addition, that is, the channel elements corresponding to the visible light image and the infrared image are concat-operated to obtain the final multimodal fusion data M.

[0064] According to a further improvement of the present invention, the multimodal fusion target detection model based on Yolov5 constructed in step S3 is an improved target detection model, including:

[0065] Step S31, construct the backbone network Backbone: It is composed of Darknet and CSP structure, and extracts features at different levels through layer-by-layer convolution, activation, batch normalization and other operations to provide rich low-level and high-level features for subsequent detection modules;

[0066] Step S32, constructing a feature fusion layer Neck: adopting a path aggregation network: structure, combining a feature pyramid network and a path aggregation network, and through multi-layer feature fusion, making the model more suitable for detecting targets of different sizes, and enhancing the recognition ability of small targets and long-distance targets;

[0067] Step S33, constructing the detection head: generating different Anchor boxes for each feature layer, and calculating the probability of the target category and the bounding box position for each Anchor. Three scales are usually used to predict large, medium and small targets respectively. The specific bounding box information of the generated feature map includes position, size, confidence and category. Through the non-maximum suppression process, the overlapping boxes are removed and the predicted box with the highest confidence is retained to obtain the final detection result. The C3-RS module is used to replace the C3 module in the original structure, R refers to the Bottleneck residual block in ResNet, and combined with the SE module, the channel attention is more focused and the recognition ability of small targets is improved.

[0068] According to a further improvement of the present invention, the SimaFc-based airborne target tracking model constructed in step S4 is an improved target tracking model, including:

[0069] Step S41, constructing a backbone network Backbone: based on MobileNet, it consists of two convolutional neural network branches with shared weights, and extracts features from the template image and the search area respectively;

[0070] Step S42, constructing a similarity measurement module: matching the template features with the search area features through a mutual correlation layer, and measuring the similarity between the two;

[0071] Step S43, constructing a response map generation module: generating a response map, i.e., a similarity distribution map, from the result of the cross-correlation, and locating the most likely position of the target by comparing the response values ​​at different positions in the response map. The peak position in the response map, also called the maximum response value, is the target position prediction;

[0072] Step S44, constructing an occlusion detection module: by calculating the change pattern in the response graph, when it is detected that the similarity suddenly drops sharply or becomes blurred, it is determined whether the target is occluded; when it is determined that occlusion has occurred, the occlusion detection module will trigger a relocation strategy to relocate the target through the response graph of the previous frame and the historical trajectory of the target;

[0073] Step S45: construct a target positioning and updating module, adjust the position of the target's bounding box according to the peak position of the response graph, and update the model based on the confidence level;

[0074] Step S46, constructing a post-processing module: performing smoothing and screening operations to improve the accuracy and stability of tracking in complex environments.

[0075] According to a further improvement of the present invention, in step S3, the multimodal fusion target detection model based on Yolov5 is pruned and quantized by a model lightweight method based on joint channel pruning and layer pruning, which includes the following steps:

[0076] S3a, channel pruning: customize the pruning script and implement basic channel pruning in models / common.py in the Yolov5 code framework;

[0077] S3b, layer pruning: The contribution of a layer is evaluated by measuring the importance of the layer, and the layers with low contribution are pruned first. Layer pruning is implemented based on the sparsity of the feature map. By counting the proportion of non-zero values ​​or the average absolute value in the activation output, the layer with high sparsity is found. That is, the feature map output of the layer is very sparse or even close to zero, indicating that the layer has a small effect and is the layer that needs to be pruned.

[0078] S3c, model fine-tuning: After channel pruning and layer pruning operations, the model is fine-tuned and trained to restore the accuracy of the model;

[0079] S3d, model quantization: quantize the weights in floating point FP32 format to integer INT8 format, and use a small number of real scene samples as calibration data, input into the model, collect the maximum and minimum value information of the activation layer, and then quantize the activation and weight of each layer in the model to INT8 representation according to the sampling range.

[0080] According to a further improvement of the present invention, in step S4, the SimaFc-based airborne target tracking model is pruned and quantized by a model lightweight method based on joint channel pruning and layer pruning, comprising the following steps:

[0081] Step S4a, channel pruning: by calculating the batch normalization layer scaling factor γ parameter of each channel, the scaling factor represents the activation amplitude of the channel, and judging the importance of the channel by analyzing these values; wherein, the larger the γ value is, the more important the feature of the channel is, and the smaller the γ value is, the smaller the influence of the channel on the feature in the model is. Channels with lower weights are selected as pruning candidates, and a certain proportion of channels are retained based on the set pruning rate, and 80% of the total number of channels are retained. After pruning, the number of channels of the model is adjusted to ensure that the input dimension of the downstream convolutional layer is consistent with the output dimension;

[0082] Step S4b, layer pruning: Evaluate the impact of deleting layers on model accuracy layer by layer, perform sensitivity analysis on each convolutional layer of MobileNet, and record the impact of each convolutional layer on model accuracy; use the validation set to perform a quick inference test to obtain model accuracy. If a layer has little impact on accuracy, it will be marked as a pruning candidate; the impact on accuracy is determined by setting an accuracy threshold. If the accuracy drops by less than 1% after removing a layer, it is considered that the layer can be removed; according to the results of the sensitivity analysis, select redundant layers from the deep convolutional layers with the least impact for pruning;

[0083] Step S4c, model fine-tuning: After performing channel pruning and layer pruning operations, fine-tune the model to restore the accuracy of the model;

[0084] Step S4d, model quantization: quantize the weights in floating point FP32 format to integer INT8 format, and use a small number of real scene samples as calibration data, input them into the model, collect the maximum and minimum value information of the activation layer, and then quantize the activation and weight of each layer in the model to INT8 representation according to the sampling range.

[0085] Example 2

[0086] The following will describe the specific implementation method of the method for rapid target detection and tracking of unmanned aerial vehicles in complex environments proposed by the present invention. It is characterized by including three parts: multimodal fusion target detection based on Yolov5, airborne target tracking based on SimaFc, and model lightweighting based on joint channel pruning and layer pruning.

[0087] Among them, the multimodal fusion target detection based on Yolov5 includes the following steps:

[0088] Step 1_1: Take a visible light video frame image M1 with a width of W and a height of H, and an infrared video frame M2 with a width of W and a height of H, and obtain a multimodal fused image M through an enhanced dual-scale image fusion algorithm;

[0089] Step 1_2: Input the multimodal fusion image M with width W and height H obtained in the above step into the multimodal fusion target detection model based on Yolov5 to obtain the initial target detection result Y1. The format of Y1 is [class_id, x_center, y_center, width, height, confidence], where classes_id represents the category of the target, (x, y, width, height) represents the coordinates of the detection box, (x, y) is the center coordinate of the box, (width, height) is the width and height of the box, and confidence represents the confidence of the predicted bounding box, which is a value between 0 and 1.

[0090] Among them, the airborne target tracking based on SimaFc includes the following steps:

[0091] Step 2_1: Based on the target detection result Y1 obtained in the above step 1_2, crop from the multimodal fusion image M with a width of W and a height of H to generate a template image set Y2 containing only the target. The set Y2 contains a series of initial target template images [m1, m2, ..., mn] contained in the image M'.

[0092] Step 2_2: According to the target detection result Y1 obtained in the above step 1_2, a larger area near the target position determined in the multimodal fusion image M with a width of W and a height of H is intercepted as the search area for tracking the target. The search area is recorded as a set Y3. The Y3 set contains a series of search area images [n1, n2, ..., nn] contained in the image M, which corresponds to the template image set Y2.

[0093] Step 2_3: The template image set Y2 obtained in the above step 2_1 and the search area image set Y3 obtained in the above step 2_2 are respectively input into the two convolutional neural network branches of the airborne target tracking model based on SimaFC (Siamese Fully-Convolutional Network) for feature extraction. On this basis, the model will use similarity or distance measurement to determine the part in the search area that is most similar to the template image, thereby achieving target tracking.

[0094] Among them, the model lightweighting based on joint channel pruning and layer pruning includes the following steps:

[0095] Step 3_1: Perform channel pruning and layer pruning on the multimodal fusion target detection model based on Yolov5 to ensure that the pruned target detection model can maintain high detection accuracy while reducing the amount of calculation. In addition, based on the actual hardware resources, perform quantization operations on the model to further compress the model and improve the model reasoning speed.

[0096] Step 3_2: Prune the SimaFc-based airborne target tracking model to ensure that the target tracking model does not have a significant impact on the target tracking effect while reducing the amount of calculation. And based on the actual hardware resources, quantize the model to further compress the model and improve the model reasoning speed.

[0097] The following is a detailed description of the above steps:

[0098] In the above step 1_1, an enhanced dual-scale image fusion algorithm is used to fuse the visible light image and the infrared image. The specific implementation includes the following steps:

[0099] Image decomposition:

[0100] Input two pictures, one is a visible light image and the other is the corresponding infrared image, respectively. and Indicates that bilateral filtering (BF) is used to decompose the two images into a base layer and a detail layer, respectively. The calculation method is as follows:

[0101]

[0102] Where i = 1, 2, is the input image; Ω is a neighborhood range, a set of neighborhood pixels centered at (x, y); is the intensity weight function, controlling the effect of intensity differences; f s (||(x,y)-(u,v)||) is the spatial weight function, which controls the influence of spatial distance; W x,y is the normalization coefficient, which is used to ensure that the output brightness range is normal; Represents the base layer image.

[0103] After obtaining the shallow features, the detail layer is obtained by performing a difference operation between the original input image and the base layer image. The method for obtaining the detail layer is as follows:

[0104]

[0105] Where i = 1, 2, represents the original input image, represents the gradient magnitude image obtained in the previous step, that is, the base layer image, is the output detail layer image.

[0106] For the weight calculation of the detail layer, the input image and Two salient feature maps are generated by the spectral residual method (SR), and then each image is inverse Fourier transformed to obtain the salient feature maps ψ1(x, y) and ψ2(x, y) in the spatial domain. i The calculation process of (x,y) is as follows:

[0107] ψ i (x,y)=F -1 {e R(u,v) ·e jP(u.v)}

[0108] Where i = 1, 2, ψ i (x, y) is the saliency map in the spatial domain, R(u, v) is the spectral residual calculated by the spectral residual method, j is the imaginary unit, P(u, v) is the phase spectrum of the input image, and F-1 represents the inverse Fourier transform; eR(u, v) means adjusting the value of the amplitude spectrum by the spectral residual R(u, v), and converting it into nonlinear scaling through the properties of the exponential function to highlight the salient features.

[0109] Therefore, the fusion coefficient matrix of the detail layer is as follows:

[0110]

[0111] Image Fusion:

[0112] For the base layer, a weighted average strategy is used for fusion:

[0113]

[0114] Among them, B fused (x,y) represents the fused base layer, and They represent the base layers of visible light and infrared images respectively. The weight α ranges from 0 to 1 and can be adjusted according to the brightness characteristics of the image. Specifically, the value of α is dynamically adjusted according to different modes of day and night. For example, during the night, the quality of infrared images is better than that of visible light, and the value of α is 0.3-0.4. During the day, the quality of visible light images is higher, and the value of α is 0.6-0.7.

[0115] For the detail layer, a fusion strategy of weighted averaging is used using the fusion coefficient matrix constructed using the salient feature map:

[0116]

[0117] Among them, D fused (x, y) represents the fused detail layer, ξ1(x, y) and ξ2(x, y) are the weighting coefficients obtained by the spectral residual method mentioned above, and Represents the detail layers of visible light and infrared images respectively.

[0118] Image reconstruction:

[0119] After obtaining the final base layer and detail layer, the base layer and detail layer are fused by simple addition, that is, the channel elements corresponding to the visible light image and the infrared image are concat-operated to obtain the final multi-modal fusion data γ(x,y). The specific multi-channel data fusion process is detailed in the attached Figure 2 .

[0120] In the above steps 1_2, the multimodal fusion target detection model based on Yolov5 is obtained through training, and its training process includes the following steps:

[0121] Download the UAVDT-DatasetNinja dataset;

[0122] Divide the training data and test data, select 8000 pairs of visible light and infrared images from the UAVDT-DatasetNinja dataset as training data, and 2000 pairs of visible light and infrared images as test data. Ensure that the visible light image and the infrared image are strictly corresponding, and use the enhanced dual-scale image fusion algorithm to fuse the visible light image and the infrared image to obtain the fused training data and test data.

[0123] The multimodal fusion target detection model based on yolov5 is pre-trained using the prepared image data, and the model weight w1 obtained from the pre-training is saved. To ensure the detection accuracy of the model in real application scenarios, 1000 pairs of real scene data are used to fine-tune the model. Specifically, the 1000 pairs of real scene data are first divided into training data and test data in a ratio of 8:2, and the model is initialized using the pre-weight w1 obtained from the above training. The hyperparameters related to the model training are adjusted according to the training loss value, and the model weight w2 obtained from the second training fine-tuning is used as the final weight of the target detection model.

[0124] In the above step 2_1, the template image is cropped from the fixed size of the target area in the initial frame, and the size of the cropped template image is 127x127 pixels;

[0125] In the above step 2_2: the search area is a cropped image of a larger area in the next frame relative to the initial frame, and is used to detect the target. The image size of the search area is 255x255 pixels;

[0126] In the above steps 2_3, the airborne target tracking model based on SimaFc is obtained through training, and the training process includes the following steps:

[0127] According to the training set and validation set divided by the above UAVDT-DatasetNinja dataset, the template image and the corresponding search area image are cropped out according to the existing data labels as available data for model training.

[0128] The prepared image data is used to pre-train the SimaFc-based airborne target tracking model, and the model weight w3 obtained by pre-training is saved. To ensure the tracking accuracy of the model in real application scenarios, 1000 pairs of real scene data are used to fine-tune the model. Specifically, the 1000 pairs of real scene data are first divided into training data and test data in a ratio of 8:2, and the model is initialized using the pre-weight w3 obtained by the above training. The hyperparameters related to the model training are adjusted according to the training loss value, and the model weight w4 obtained by the second training fine-tuning is used as the final weight of the target tracking model.

[0129] In the above step 3_1, channel pruning, layer pruning and model quantization are performed on the multimodal fusion target detection model based on Yolov5. The specific steps are as follows:

[0130] Channel pruning: Customize the pruning script and implement basic channel pruning in models / common.py in the Yolov5 code framework. The channel pruning operation is mainly based on the L1 norm. The part with the smallest absolute value of the channel is considered to have the least impact on the model performance, that is, the feature channel that needs to be pruned.

[0131] Layer pruning: The contribution of a layer is evaluated by measuring the importance of the layer, and the layers with low contribution are pruned first. Layer pruning is mainly based on the sparsity of feature maps. By counting the proportion of non-zero values ​​or the average absolute value in the activation output, the layers with high sparsity are found. That is, the feature map output of the layer is very sparse or even close to zero, indicating that this layer has a small effect and is the layer that needs to be pruned.

[0132] Model fine-tuning: After channel pruning and layer pruning operations, the model needs to be fine-tuned and trained separately to restore the model's accuracy.

[0133] Model quantization: The quantization of the model aims to further reduce the size of the model, reduce memory consumption, and speed up inference. Specifically, the weights in the floating point FP32 format are quantized to the integer INT8 format, and a small number of real scene samples are used as calibration data and input into the model to collect the maximum and minimum value information of the activation layer. Then, according to the sampling range, the activation and weight of each layer in the model are quantized to INT8 representation.

[0134] In the above step 3_2, the SimaFc-based airborne target tracking model is pruned and pruned. The specific steps are as follows:

[0135] Channel pruning: By calculating the batch normalization (BN) scaling factor γ parameter of each channel, the scaling factor represents the activation amplitude of the channel, and analyzing these values ​​to determine the importance of the channel. The larger the γ value, the more important the characteristics of the channel, and the smaller the γ value, the smaller the channel's impact on the characteristics in the model. Channels with lower weights are selected as pruning candidates, and a certain proportion of channels are retained based on the set pruning rate, retaining 80% of the total number of channels. After pruning, adjust the number of channels of the model to ensure that the input dimension of the downstream convolutional layer is consistent with the output dimension.

[0136] Layer pruning: Evaluate the impact of removing layers on model accuracy layer by layer, especially on the accuracy of the target response map generation and occlusion detection modules. Perform sensitivity analysis on each convolutional layer of MobileNet and record the impact of each convolutional layer on model accuracy. Use the validation set to perform a quick inference test to get the model accuracy. If a layer has little impact on accuracy, it is marked as a pruning candidate. The impact on accuracy is determined by setting an accuracy threshold. If the accuracy drops by less than 1% after removing a layer, the layer is considered to be removable. Based on the results of the sensitivity analysis, select redundant layers from the deep convolutional layers with the least impact for pruning.

[0137] Model fine-tuning: After channel pruning and layer pruning operations, the model needs to be fine-tuned and trained separately to restore the model's accuracy.

[0138] Model quantization: Similar to the model quantization operation steps for target detection above, the weights in floating-point FP32 format are quantized to integer INT8 format, and a small number of real-world scene samples are used as calibration data and input into the model to collect the maximum and minimum value information of the activation layer. Then, according to the sampling range, the activations and weights of each layer in the model are quantized to INT8 representation.

[0139] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the present invention itself. Various changes in form and details may be made without departing from the spirit and scope of the present invention as defined in the appended claims.

Claims

1. A method for rapid target detection and tracking of unmanned aerial vehicles in complex environments, characterized in that: The following steps are involved: Step S1, using a drone equipped with an optoelectronic pod to collect images, including visible light images and infrared images, and then performing data augmentation processing to expand the original available data volume and generate occluded target detection data type; Step S2, converting the visible light image and the infrared image into a visible light video frame image M1 with a width of W and a height of H, and an infrared video frame M2 with a width of W and a height of H, and obtaining a multimodal fused image M through an enhanced dual-scale image fusion algorithm; Step S3, construct a multimodal fusion target detection model based on Yolov5, and perform channel pruning and layer pruning on the multimodal fusion target detection model based on Yolov5 to ensure that the pruned target detection model can maintain a high detection accuracy while reducing the amount of calculation, and quantize the model in combination with the actual situation of hardware resources to further compress the model and improve the model reasoning speed; Step S4, constructing an airborne target tracking model based on SimaFc, and performing path pruning and layer pruning on the above-mentioned airborne target tracking model based on SimaFc, to ensure that the pruned target tracking model will not have a significant impact on the target tracking effect while reducing the amount of calculation, and combining the actual situation of hardware resources, quantizing the model, further compressing the model, and improving the model reasoning speed; Step S5, inputting the multimodal fusion image M obtained in step S2 into the multimodal fusion target detection model based on Yolov5 constructed in step S3 for processing to obtain an initial target detection result Y1; Step S6, according to the target detection result Y1 obtained in step S5, crop from the multimodal fusion image M with a width of W and a height of H to generate a template image set Y2 containing only the target, where the set Y2 contains a series of initial target template images [m1, m2, ..., mn] contained in the image M'; Step S7, according to the target detection result Y1 obtained in step S5, a larger area near the target position determined in the multimodal fusion image M with a width of W and a height of H is intercepted as a search area for tracking the target, and the search area is recorded as a set Y3, and the set Y3 is a series of search area images [n1, n2, ..., nn] contained in the image M, corresponding to the template image set Y2; Step S8: Input the template image set Y2 and the search area image set Y3 obtained in step S7 into the two convolutional neural network branches of the SimaFc-based airborne target tracking model constructed in step S4 for feature extraction, and use similarity and distance metrics to determine the part in the search area that is most similar to the template image, thereby achieving target tracking.

2. The method for rapid target detection and tracking of unmanned aerial vehicles in complex environments according to claim 1 is characterized in that: The enhanced dual-scale image fusion algorithm in step S2 comprises the following steps: Step S21: Input two images, one is a visible light image and the other is a corresponding infrared image. and Indicates that bilateral filtering is used to decompose the two images into a base layer and a detail layer respectively; Step S22, the decomposed base layer and detail layer are fused respectively, wherein the base layer is fused by a weighted average method, and the detail layer is fused by a weighted average method using a fusion coefficient matrix constructed by using a significant feature map; Step S23: After obtaining the fused base layer and detail layer, the base layer and detail layer are fused by simple addition, that is, the channel elements corresponding to the visible light image and the infrared image are concat-operated to obtain the final multimodal fusion data M.

3. The method for rapid target detection and tracking of unmanned aerial vehicles in complex environments according to claim 1 is characterized in that: The multimodal fusion target detection model based on Yolov5 constructed in step S3 is an improved target detection model, including: Step S31, construct the backbone network Backbone: It is composed of Darknet and CSP structure, and extracts features at different levels through layer-by-layer convolution, activation, batch normalization and other operations to provide rich low-level and high-level features for subsequent detection modules; Step S32, constructing a feature fusion layer Neck: adopting a path aggregation network: structure, combining a feature pyramid network and a path aggregation network, and through multi-layer feature fusion, making the model more suitable for detecting targets of different sizes, and enhancing the recognition ability of small targets and long-distance targets; Step S33, constructing the detection head: generating different Anchor boxes for each feature layer, and calculating the probability of the target category and the bounding box position for each Anchor. Three scales are usually used to predict large, medium and small targets respectively. The specific bounding box information of the generated feature map includes position, size, confidence and category. Through the non-maximum suppression process, the overlapping boxes are removed and the predicted box with the highest confidence is retained to obtain the final detection result. The C3-RS module is used to replace the C3 module in the original structure, R refers to the Bottleneck residual block in ResNet, and combined with the SE module, the channel attention is more focused and the recognition ability of small targets is improved.

4. The method for rapid target detection and tracking of unmanned aerial vehicles in complex environments according to claim 1 is characterized in that: The SimaFc-based airborne target tracking model constructed in step S4 is an improved target tracking model, including: Step S41, constructing a backbone network Backbone: based on MobileNet, it consists of two convolutional neural network branches with shared weights, and extracts features from the template image and the search area respectively; Step S42, constructing a similarity measurement module: matching the template features with the search area features through a mutual correlation layer, and measuring the similarity between the two; Step S43, constructing a response map generation module: generating a response map, i.e., a similarity distribution map, from the result of the cross-correlation, and locating the most likely position of the target by comparing the response values ​​at different positions in the response map. The peak position in the response map, also called the maximum response value, is the target position prediction; Step S44, constructing an occlusion detection module: by calculating the change pattern in the response graph, when it is detected that the similarity suddenly drops sharply or becomes blurred, it is determined whether the target is occluded; when it is determined that occlusion has occurred, the occlusion detection module will trigger a relocation strategy to relocate the target through the response graph of the previous frame and the historical trajectory of the target; Step S45: construct a target positioning and updating module, adjust the position of the target's bounding box according to the peak position of the response graph, and update the model based on the confidence level; Step S46, constructing a post-processing module: performing smoothing and screening operations to improve the accuracy and stability of tracking in complex environments.

5. The method for rapid target detection and tracking of unmanned aerial vehicles in complex environments according to claim 1 is characterized in that: In step S3, the multimodal fusion target detection model based on Yolov5 is pruned and quantized by a model lightweight method based on joint channel pruning and layer pruning, which includes the following steps: S3a, channel pruning: customize the pruning script and implement basic channel pruning in models / common.py in the Yolov5 code framework; S3b, layer pruning: The contribution of a layer is evaluated by measuring the importance of the layer, and the layers with low contribution are pruned first. Layer pruning is implemented based on the sparsity of the feature map. By counting the proportion of non-zero values ​​or the average absolute value in the activation output, the layer with high sparsity is found. That is, the feature map output of the layer is very sparse or even close to zero, indicating that the layer has a small effect and is the layer that needs to be pruned. S3c, model fine-tuning: After channel pruning and layer pruning operations, the model is fine-tuned and trained to restore the accuracy of the model; S3d, model quantization: quantize the weights in floating point FP32 format to integer INT8 format, and use a small number of real scene samples as calibration data, input into the model, collect the maximum and minimum value information of the activation layer, and then quantize the activation and weight of each layer in the model to INT8 representation according to the sampling range.

6. The method for rapid target detection and tracking of unmanned aerial vehicles in complex environments according to claim 1 is characterized in that: In step S4, the SimaFc-based airborne target tracking model is pruned and quantized by a model lightweight method based on joint channel pruning and layer pruning, which includes the following steps: Step S4a, channel pruning: by calculating the batch normalization layer scaling factor γ parameter of each channel, the scaling factor represents the activation amplitude of the channel, and judging the importance of the channel by analyzing these values; wherein, the larger the γ value is, the more important the feature of the channel is, and the smaller the γ value is, the smaller the influence of the channel on the feature in the model is. Channels with lower weights are selected as pruning candidates, and a certain proportion of channels are retained based on the set pruning rate, and 80% of the total number of channels are retained. After pruning, the number of channels of the model is adjusted to ensure that the input dimension of the downstream convolutional layer is consistent with the output dimension; Step S4b, layer pruning: Evaluate the impact of deleting layers on model accuracy layer by layer, perform sensitivity analysis on each convolutional layer of MobileNet, and record the impact of each convolutional layer on model accuracy; use the validation set to perform a quick inference test to obtain model accuracy. If a layer has little impact on accuracy, it will be marked as a pruning candidate; the impact on accuracy is determined by setting an accuracy threshold. If the accuracy drops by less than 1% after removing a layer, it is considered that the layer can be removed; according to the results of the sensitivity analysis, select redundant layers from the deep convolutional layers with the least impact for pruning; Step S4c, model fine-tuning: After performing channel pruning and layer pruning operations, fine-tune the model to restore the accuracy of the model; Step S4d, model quantization: quantize the weights in floating point FP32 format to integer INT8 format, and use a small number of real scene samples as calibration data, input them into the model, collect the maximum and minimum value information of the activation layer, and then quantize the activation and weight of each layer in the model to INT8 representation according to the sampling range.

Citation Information

Patent Citations

  • An infrared small unmanned aerial vehicle target detection and tracking method under a complex background

    CN109816695A

  • Unmanned aerial vehicle visual identification method and device

    CN119152335A

  • Mobile gas and chemical imaging camera

    US20180077363A1

  • Infrared image processing method, apparatus, and device, and storage medium

    WO2024051067A1