Unmanned aerial vehicle target detection method based on long-tail distribution optimization

By early fusion and neural network optimization of the image and point cloud data collected by the drone, the long tail distribution problem in construction site scenarios is solved, the detection effect of low-frequency and high-risk targets is improved, and the robustness of the model in complex environments is enhanced.

CN120279247APending Publication Date: 2025-07-08XIAMEN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510333489.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing drone target detection technology has long tail distribution problems in construction site scenarios, resulting in poor detection of a few categories of samples. It is difficult for traditional methods to effectively deal with lighting changes and target occlusion problems in high-resolution images and complex environments.

Method used

By early fusion of image data collected by the drone and point cloud data, multi-scale feature extraction network is used to perform multi-scale feature extraction and adaptive reweighting, dynamically adjust the anchor box, optimize the sampling strategy of the category bias sampler, and combine the depth decoder for object detection to build a multi-modal joint confidence optimization loss function.

Benefits of technology

It significantly improves the detection sensitivity of low-frequency but high-risk targets in the construction site, improves the recognition ability of small targets and occlusion targets, and enhances the robustness of the model in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279247A_ABST
    Figure CN120279247A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle target detection method based on long-tail distribution optimization. The method comprises the steps of performing data fusion on acquired multi-modal data; inputting the fused data into a feature extraction network for feature extraction, and performing adaptive reweighting of channels and spaces during feature extraction; dynamically adjusting an anchor box according to the extracted feature map by adopting a regional proposal network to generate a candidate box; calculating a joint confidence coefficient according to the candidate box and the feature map so as to optimize a sampling strategy of a category bias sampler; respectively processing the sampled head feature frame and the tail feature frame to obtain a corresponding category prediction probability and a bounding box offset, and decoding a prediction depth value by adopting a depth decoder according to the candidate frame and the feature map; according to the category prediction probability, the bounding box offset and the prediction depth value, constructing a total loss function to perform model training; outputting a detection result by adopting the trained target detection model; therefore, the detection sensitivity of low-frequency and high-risk targets in the construction site is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent monitoring technology, and in particular, to an unmanned aerial vehicle (UAV) target detection method optimized based on long-tail distribution, a computer-readable storage medium, and a computer device. Background Art

[0002] In the related art, for the special working environment of construction sites where safety hazards are intensive and accidents occur frequently, the real-time monitoring of safety hazard projects is particularly important. Its importance is not only reflected in the aspect of personnel safety guarantee, but also deeply affects the efficiency of personnel and vehicle engineering management, economic benefits, and even other multi-dimensional considerations. Traditional construction site monitoring systems mainly rely on fixed cameras and manual inspections. However, considering the characteristics of construction sites, such as vast areas, a wide variety of safety hazards to be monitored, rapidly changing accident patterns, and urgent response times, such traditional monitoring methods inevitably expose limitations such as extensive coverage blind spots, lagging responses, and difficulty in comprehensively monitoring safety hazards.

[0003] In view of this, it is particularly important and urgent to use UAV target detection technology to implement real-time monitoring of safety hazards in construction site scenarios. The application of this technology can not only effectively make up for the above defects of traditional monitoring solutions, but also lay a solid foundation for the comprehensive deployment and efficient implementation of future smart construction sites, showing far-reaching strategic significance. However, although the UAV monitoring solution provides a new perspective and means to solve the construction site monitoring problem, when dealing with the complex and changeable monitoring requirements of construction sites, the existing UAV monitoring solutions still expose several obvious deficiencies, including insufficient long-tail detection ability. The existing object detection algorithms (such as Faster R-CNN, YOLO) detectors perform poorly when dealing with the data collected by UAVs, mainly facing three major challenges: (1) In high-resolution UAV images, the targets look very small; (2) The spatial distribution of targets in the images is uneven; (3) The imbalance problem of different-class target samples, which is called the long-tail distribution problem.

[0004] Facing the above problems, the current main optimization direction focuses on improving the quality of the input image. For example, developing an effective image cropping strategy (the DMnet algorithm is one of them). This method alleviates the problem of uneven spatial distribution, but ignores the imbalance problem of different-class target samples, which is called the long-tail distribution problem. During the UAV target detection process, due to the excessive number of detected target categories, in practical applications, a better recognition rate can be achieved for some categories with a larger number, while the detection of some tail-class targets with a smaller number cannot achieve an ideal effect. In classic UAV datasets such as VisDrone and UAVDT, there will be an obvious situation where the proportion of samples of a small number of categories (head classes) is large, and the samples of most categories (tail classes) are small, resulting in a large performance difference between the head classes and the tail classes during detection.

[0005] Regarding the long-tail distribution problem, currently two main methods are adopted to minimize the impact of the long-tail distribution in a two-stage detection manner: (1) Resampling: Balancing the data distribution by oversampling the tail classes or undersampling the head classes, but it is prone to overfitting or a decrease in generalization ability; (In the preprocessing stage of training data, it is completed through a sampler module. Here, it is necessary to increase the size of the training batch. However, the drone data occupies too much space, and only one or two images can be placed in each batch, making it difficult to increase the batch size.) (2) Reweighting: Assigning higher weights to the tail samples in the loss function, but it is difficult to handle large-scale data. However, both of these methods still have problems when facing the drone dataset. High-resolution images limit the batch size: Drone images have high resolutions (e.g., most VisDrone images are 2000×1500 pixels), and the batch size during training is usually only 1-2, making it difficult to meet the requirements of balanced sampling; The class imbalance within a single image is serious: A single image often contains hundreds of head-class targets, resulting in the failure of traditional image-level resampling. This causes the existing drone target detection frameworks to be unable to handle the long-tail distribution problem well when facing the huge class recognition requirements in the construction site. At the same time, the construction site is in an open-air state, affected by light changes, rain and snow weather. At the same time, due to the stacking habits of construction site consumables, problems such as target occlusion and small target density are common, and relying solely on visual data is prone to missing key risks. Summary of the Invention

[0006] This application aims to at least partly solve one of the technical problems in the above technologies. To this end, an object of this application is to propose a drone target detection method optimized based on the long-tail distribution. By early fusing the multi-modal data collected by drones and specifically optimizing each key functional module in the neural network, the detection sensitivity of low-frequency but high-risk targets in the construction site is significantly improved.

[0007] The second object of this application is to propose a computer-readable storage medium.

[0008] The third object of this application is to propose a computer device.

[0009] To achieve the above object, an embodiment of the first aspect of the present application proposes a drone target detection method optimized based on the long-tailed distribution, including the following steps: obtaining image data and point cloud data collected by a drone in a scene, and preprocessing and fusing the image data and the point cloud data to obtain fused data; inputting the fused data into a feature extraction network for feature extraction to obtain a multi-scale feature map, and performing adaptive reweighting of channels and space on the multi-scale feature map to obtain a multi-modal feature map; using a region proposal network to dynamically adjust anchor boxes according to the multi-modal feature map to generate candidate boxes; calculating a multi-modal joint confidence according to the candidate boxes and the multi-modal feature map to optimize the sampling strategy of a class bias sampler through the multi-modal joint confidence to obtain corresponding head feature boxes and tail feature boxes; inputting the head feature boxes into a head class classification head, and inputting the tail feature boxes into a tail class classification head for processing to obtain corresponding class prediction probabilities and bounding box offsets, and using a depth decoder to decode and predict depth values according to the candidate boxes and the multi-modal feature map; constructing a total loss function according to the class prediction probabilities, the bounding box offsets, and the predicted depth values for model training; using the trained target detection model for target detection to output detection results.

[0010] According to the drone target detection method optimized based on the long-tailed distribution in the embodiment of the present application, first, image data and point cloud data collected by a drone in a scene are obtained, and the image data and the point cloud data are preprocessed and fused to obtain fused data; then, the fused data is input into a feature extraction network for feature extraction to obtain a multi-scale feature map, and adaptive reweighting of channels and space is performed on the multi-scale feature map to obtain a multi-modal feature map; next, a region proposal network is used to dynamically adjust anchor boxes according to the multi-modal feature map to generate candidate boxes; then, a multi-modal joint confidence is calculated according to the candidate boxes and the multi-modal feature map to optimize the sampling strategy of a class bias sampler through the multi-modal joint confidence to obtain corresponding head feature boxes and tail feature boxes; then, the head feature boxes are input into a head class classification head, and the tail feature boxes are input into a tail class classification head for processing to obtain corresponding class prediction probabilities and bounding box offsets, and a depth decoder is used to decode and predict depth values according to the candidate boxes and the multi-modal feature map; finally, a total loss function is constructed according to the class prediction probabilities, the bounding box offsets, and the predicted depth values for model training; the trained target detection model is used for target detection to output detection results; thus, by performing early fusion on the collected multi-modal data and performing targeted optimization on each key functional module in the neural network, the detection sensitivity of low-frequency but high-risk targets in the construction site is significantly improved.

[0011] In addition, the drone target detection method optimized based on the long-tail distribution proposed in the above embodiments of the present application may further have the following additional technical features:

[0012] Optionally, preprocess and fuse the image data and the point cloud data to obtain fused data, including: synchronously aligning and normalizing the image data and the point cloud data to obtain an RGB image and a depth image with the same size; splicing the RGB image and the depth image along the channel dimension to obtain fused data.

[0013] Optionally, input the fused data into a feature extraction network for feature extraction to obtain a multi-scale feature map, and perform adaptive reweighting of channels and space on the multi-scale feature map to obtain a multi-modal feature map, including: inputting the fused data into a residual neural network in the feature extraction network for processing to obtain a multi-layer feature map; inputting the multi-layer feature map into a feature pyramid network in the feature extraction network for upsampling and lateral connection to obtain a multi-scale feature map; inserting a convolutional block attention module after the output of each layer of the feature pyramid network to perform adaptive reweighting of channels and space on the multi-scale feature map to obtain a multi-modal feature map.

[0014] Optionally, use a region proposal network to dynamically adjust the anchor boxes according to the multi-modal feature map to generate candidate boxes, including: extracting proxy depth values from the multi-modal feature map using a fixed rule mapping scheme; normalizing the depth proxy values to obtain normalized proxy depth values; obtaining anchor box parameters within the corresponding division intervals according to the normalized proxy depth values, and dynamically adjusting the anchor box size and aspect ratio according to the anchor box parameters to generate candidate boxes.

[0015] Optionally, calculate the multi-modal joint confidence according to the candidate boxes and the multi-modal feature map, including: using a region of interest alignment module to extract region of interest features according to the candidate boxes and the multi-modal feature map; splitting the region of interest features into RGB components and depth components, and predicting classification scores through independent fully connected layers; generating modal weights according to the sum of feature activations of the RGB components and the sum of feature activations of the depth components; calculating the multi-modal joint confidence according to the classification scores and the modal weights.

[0016] Optionally, when optimizing the sampling strategy of the class bias sampler through the multi-modal joint confidence, the class bias sampler preferentially samples the tail feature boxes, where the sampling probability is proportional to the multi-modal joint confidence.

[0017] Optionally, a depth decoder is used to decode and predict depth values based on the candidate bounding boxes and the multimodal feature maps, including: using a region of interest (ROI) alignment module to extract ROI features based on the candidate bounding boxes and the multimodal feature maps; generating depth features from the ROI features through a fully connected layer; and using an activation function to normalize the depth features to obtain the predicted depth values.

[0018] To achieve the above object, an embodiment of the second aspect of the present application provides a computer-readable storage medium, on which a target detection program optimized based on the long-tail distribution is stored. When the target detection program optimized based on the long-tail distribution is executed by a processor, the above-mentioned drone target detection method optimized based on the long-tail distribution is implemented.

[0019] To achieve the above object, an embodiment of the third aspect of the present application provides a computer device including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned drone target detection method optimized based on the long-tail distribution is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 FIG. is a flowchart of a drone target detection method optimized based on the long-tail distribution according to an embodiment of the present application;

[0021] Figure 2 FIG. is a schematic diagram of the overall network structure of a target detection model according to an embodiment of the present application;

[0022] Figure 3 FIG. is a schematic diagram of the path planning of a drone in a smart construction site safety monitoring scenario according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application, but should not be construed as limiting the present application.

[0024] To better understand the above technical solutions, the exemplary embodiments of the present application will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0025] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0026] Figure 1 FIG. is a schematic flowchart of an unmanned aerial vehicle (UAV) target detection method optimized based on a long-tailed distribution according to an embodiment of the present application. As Figure 1 shown, the UAV target detection method optimized based on the long-tailed distribution according to the embodiment of the present application includes the following steps:

[0027] S101, acquiring image data and point cloud data collected by a UAV in a scene, and performing preprocessing and fusion on the image data and the point cloud data to obtain fused data.

[0028] It should be noted that, as Figure 3 shown, in the UAV data acquisition stage, the UAV is equipped with a camera and a lidar device, and data is collected in the intelligent construction site scene according to preset inspection points to obtain RGB images and LiDAR point cloud data in the construction site scene; in the multi-modal data fusion stage, it is necessary to ensure that the spliced tensors have the same size in other dimensions (here, height H and width W) except the splicing dimension through data preprocessing, and then perform splicing fusion along the specified dimension.

[0029] As an example, performing preprocessing and fusion on the image data and the point cloud data to obtain fused data includes: performing synchronous alignment and normalization processing on the image data and the point cloud data to obtain an RGB image and a depth image with the same size; splicing the RGB image and the depth image together along the channel dimension to obtain fused data.

[0030] That is to say, the preprocessing process is divided into synchronous alignment and normalization processing. The data is aligned in time and space through hardware synchronization. The preprocessed and normalized RGB image (3 channels) with the same size and the depth image (1 channel) are spliced into a 4-channel input. In this way, the data of each modality is synthesized into a unified representation, completing multi-modal fusion at the data level. This method ensures that before inputting to the subsequent deep learning network (ResNet+FPN), the complementary information (color, texture, and depth / geometric information) obtained by different sensors has been effectively fused, providing richer and more accurate input data for target detection.

[0031] Specifically, first, the RGB image is processed in multiple stages: median filtering is used to remove noise, contrast is enhanced through histogram equalization, radial distortion correction is performed using camera intrinsic parameters, resolution is unified to the target size through bilinear interpolation, and pixel values ​​are normalized to the range of [0,1]. Next, the LiDAR point cloud data is transformed into the camera coordinate system in turn, a sparse depth map is generated through intrinsic matrix projection, bilinear interpolation is used to fill the missing areas, and the depth values ​​are normalized to [0,1]. Therefore, the spatial alignment of multimodal data at the pixel level is ensured through the spatiotemporal alignment algorithm, providing a spatiotemporal consistency basis for subsequent fusion. This preprocessing process effectively improves data quality and eliminates noise interference caused by sensor differences. Finally, the preprocessed RGB image (3 channels) and depth map (1 channel) are spliced ​​into a 4-channel tensor along the channel dimension to form a unified input representation; this fusion method can fully utilize the complementarity of multimodal information in the early stage of feature extraction and avoid the possible information loss problem in the later fusion. In the specific implementation, the RGB image with dimensions H×W×3 and the depth map with dimensions H×W×1 are merged into a fused tensor of H×W×4 through tensor concatenation operations (such as torch.cat of PyTorch), ensuring spatially aligned pixel-level information interaction.

[0032] It should be noted that in the data preprocessing stage, the synchronization and alignment between multimodal data is very important. In terms of time, the sensor's own timestamp (such as GPS, IMU) is used for synchronization, and the data is aligned to the same time base through interpolation or resampling. If the synchronization and calibration errors between sensors are large, the fusion effect may be poor. In terms of space, the point clouds of different coordinate systems are converted to the same coordinate system (camera coordinate system) through the external parameter matrix (rotation matrix R and translation vector t), and then the aligned depth map is generated by internal parameter matrix projection.

[0033] S102, inputting the fused data into a feature extraction network for feature extraction to obtain a multi-scale feature map, and performing channel and space adaptive reweighting on the multi-scale feature map to obtain a multimodal feature map.

[0034] That is to say, in the feature extraction stage, the feature extraction network is constructed through the ResNet+FPN+CBAM structure. FPN constructs a multi-scale feature pyramid through upsampling and lateral connection, so that the network can capture high-resolution detail information and low-resolution strong semantic information at the same time. After each layer of FPN outputs a multi-scale feature map, the CBAM module is introduced to adaptively re-weight the features in terms of channels and space, strengthen the multi-scale feature expression and attention focus, and improve the recognition ability of tail-type targets and small targets.

[0035] As an example, the fused data is input into a feature extraction network for feature extraction to obtain a multi-scale feature map, and the multi-scale feature map is adaptively re-weighted in terms of channels and space to obtain a multi-modal feature map, including: inputting the fused data into a residual neural network in the feature extraction network for processing to obtain a multi-layer feature map; inputting the multi-layer feature map into a feature pyramid network in the feature extraction network for upsampling and lateral connection to obtain a multi-scale feature map; inserting a convolutional block attention module after the outputs of each layer of the feature pyramid network to adaptively re-weight the multi-scale feature map in terms of channels and space to obtain a multi-modal feature map.

[0036] That is to say, the 4-channel data obtained by early fusion is input into a feature pyramid network with a ResNet residual neural network as the backbone network for feature extraction; among them, by modifying the input layer of the residual network to 4 channels, it can correctly input multi-modal data, and through convolution, batch normalization, activation function and pooling operations, a multi-layer feature map with decreasing resolution but increasing semantic information content is obtained; then, the multi-layer feature map output by the residual network is used to construct a multi-scale feature map through the upsampling and lateral connection of the feature pyramid network, enabling the network to simultaneously capture high-resolution detail information and low-resolution strong semantic information; inserting a CBAM (convolutional block attention module) after the outputs of each layer of the FPN to adaptively re-weight the features in terms of both channels and space can capture the dependencies and interactions between different channels, and by adjusting the channel weights, the key regions and important features can be more effectively highlighted, improving the recognition ability for small targets and long-tail targets.

[0037] Specifically, in the feature extraction stage, an improved ResNet-50 network is adopted. Its input layer is customized for 4-channel data: the number of input channels of the initial 7×7 convolutional layer is adjusted to 4, and the number of output channels is 64. The spatial resolution is reduced through a convolutional operation with a stride of 2; the subsequent max pooling layer further compresses the feature size to 1 / 4, laying a foundation for the feature extraction of subsequent residual blocks; the four residual stages (C2-C5) of ResNet-50 gradually increase the feature dimension (from 256 to 2048) through a bottleneck structure (1×1-3×3-1×1 convolution), and solve the gradient vanishing problem through skip connections. The FPN module is constructed using a top-down upsampling and lateral connection mechanism; starting from the C5 feature, a top-level feature P5 is generated through a 1×1 convolution, and then P5 is upsampled layer by layer and channel-aligned and element-added to the same-level ResNet features (C4-C2), and finally smoothed through a 3×3 convolution to obtain P4-P2; this design effectively integrates deep semantic features and shallow detail features, enhancing the multi-scale object detection ability. The CBAM module embedded after each layer of FPN optimizes the feature representation through the cascading of channel attention and spatial attention; the channel attention module generates channel weights through global average pooling and max pooling, and captures cross-channel dependencies through a shared MLP (with a hidden layer compression ratio r = 16); the spatial attention module generates spatial weights through channel-dimensional pooling, and a 7×7 convolution captures long-range spatial context; this attention mechanism dynamically adjusts the importance of each channel and spatial position, especially strengthening the complementary features of geometric information and RGB texture in the depth map, significantly improving the detection accuracy of small and occluded objects. The finally generated multi-modal feature maps P'2-P'5 retain multi-scale semantic information, providing a high-quality feature basis for subsequent long-tail distribution optimization based on the feature pyramid (such as FocalLoss improvement). It should be noted that the multi-modal feature pyramid generated using the FPN+CBAM structure is the key input point for providing an optimization direction for the CBS module and calculating the joint confidence, affecting the next optimization and adjustment direction.

[0038] S103, use the Region Proposal Network to dynamically adjust the anchor boxes according to the multi-modal feature maps to generate candidate boxes.

[0039] It should be noted that, compared with the classic Region Proposal Network (RPN), in view of the characteristics of multi-modal data, an adaptive anchor box design is added to the traditional framework. By extracting depth information and dynamically adjusting the size and aspect ratio of the anchor box according to the proxy depth value, while hardly increasing the computational cost, it can effectively improve the detection performance of small targets and occluded targets. Even when the target scale changes drastically, the dynamic anchor box can better match the actual target size. At the same time, compared with the single-modal feature input of the traditional RPN, the multi-modal RPN network accepts multi-modal features, has stronger semantic and geometric information, and the multi-modal data can complement each other, reducing false detections and missed detections, effectively increasing the robustness of the model, and can well adapt to small-target dense scenarios such as drone patrol and smart construction sites.

[0040] As an example, the Region Proposal Network is used to dynamically adjust the anchor box according to the multi-modal feature map to generate candidate boxes, including: extracting the proxy depth value from the multi-modal feature map using a fixed rule mapping scheme; normalizing the depth proxy value to obtain the normalized proxy depth value; obtaining the anchor box parameters within the corresponding division interval according to the normalized proxy depth value, and dynamically adjusting the size and aspect ratio of the anchor box according to the anchor box parameters to generate candidate boxes.

[0041] It should be noted that when using the Region Proposal Network to dynamically generate anchor boxes, instead of using the commonly used pre-defined fixed-size anchor boxes, the anchor box is dynamically adjusted through multi-modal feature fusion and depth proxy value. For distant targets (large depth value), a smaller anchor box is corresponding to reduce missed detections, and for nearby targets (small depth value), a larger anchor box is corresponding to reduce false detections; significantly improving the detection ability of small targets under the long-tail distribution. The multi-modal RPN takes the multi-modal feature map optimized by FPN+CBAM as input, which fuses RGB texture and LiDAR depth information, and strengthens the target region features through channel attention and spatial attention mechanisms. Compared with the traditional RPN, the multi-modal RPN introduces the proxy depth value in the anchor box generation stage to achieve dynamic adaptation of the anchor box parameters. It should be noted that the dynamic anchor box generated by the multi-modal RPN is the key information source for the subsequent depth consistency branch optimization link and post-processing screening link of BBH to make full use of multi-modal information.

[0042] Specifically, in the candidate box generation stage: design an RPN dynamic box adjustment mechanism, and extract depth-related information from the multi-modal feature map output by FPN as the basis for anchor box adjustment. Among them, the proxy depth value is extracted directly from the feature map through a fixed rule mapping scheme, and the accuracy is effectively improved without increasing additional computational cost; then, according to the division interval of the proxy depth value, the size and aspect ratio of the anchor box are dynamically adjusted to adaptively match the diverse aspect ratios of the targets in the construction site scenario (such as stacked building materials, irregular mechanical parts), reducing missed detections and false detections.

[0043] S104. Calculate the multi-modal joint confidence based on the candidate box and the multi-modal feature map, so as to optimize the sampling strategy of the class bias sampler through the multi-modal joint confidence, and obtain the corresponding head feature box and tail feature box.

[0044] That is to say, after generating the candidate regions (RPN stage), the sampling strategy is dynamically adjusted by the class bias sampler CBS. The tail classes are preferentially sampled (CBS(T)) to ensure that the tail class samples will not be overwhelmed by the head classes; the head classes are preferentially sampled (CBS(H)) to avoid overfitting of the head classes caused by excessive sampling of the tail classes, thereby forcibly balancing the sample numbers of the tail classes and the head classes. For multi-modal fusion data, the calculation of the multi-modal joint confidence is added during sampling. The method of sampling based on class frequency is combined with the multi-modal joint confidence for dynamic weighting, which can more accurately identify the difficult-to-detect tail class targets. At the same time, in complex scenarios (such as occlusion and illumination changes), the depth information compensates for the RGB visual defects, and the modal contributions can be effectively improved by adjusting the weights, thereby effectively improving the sampling reliability.

[0045] As an example, calculating the multi-modal joint confidence based on the candidate box and the multi-modal feature map includes: using the region of interest alignment module to extract the region of interest features according to the candidate box and the multi-modal feature map; splitting the region of interest features into RGB components and depth components, and predicting the classification scores through independent fully connected layers; generating modal weights according to the sum of the feature activations of the RGB components and the sum of the feature activations of the depth components; calculating the multi-modal joint confidence according to the classification scores and the modal weights.

[0046] That is to say, calculate the multi-modal joint confidence using the multi-modal data information, optimize the sampling strategy of the CBS class bias sampler, and dynamically adjust the sampling weights of the tail class targets (such as special warning signs and low-frequency devices) using the multi-modal information to alleviate the overfitting of the model to the head classes (common building materials), including: for the candidate boxes generated by RPN, extract the optimized multi-modal features through RoIAlign, split them into RGB and depth components (based on the original structure of the input four channels), predict the classification scores for the RGB and depth components respectively, generate modal weights according to the scene complexity (such as the depth sparsity of the occlusion area), and finally calculate the joint confidence. Optimize the tail class sampling strategy of the class bias sampler according to the joint confidence to dynamically balance the training sample distribution.

[0047] As an example, when optimizing the sampling strategy of the class bias sampler through the multi-modal joint confidence, the class bias sampler preferentially samples the tail feature boxes, where the sampling probability is proportional to the multi-modal joint confidence.

[0048] That is to say, the CBS module dynamically adjusts the sampling strategy through multi-modal confidence, calculates the joint confidence using the multi-modal features generated by FPN+CBAM. Among them, the weight λ is adaptively allocated according to the scene complexity. CBS(T) preferentially samples the candidate boxes of the tail class, and the probability is proportional to the confidence. When there is a shortage, it supplements the head class samples; the loss weight of the tail class is inversely related to the confidence to strengthen the learning of difficult samples. The dynamic anchor box adjusts the scale according to the proxy depth value to adapt to the change of the target size under the UAV perspective.

[0049] S105, input the head feature box into the head class classification head and the tail feature box into the tail class classification head for processing to obtain the corresponding class prediction probability and bounding box offset, and use the depth decoder to decode the predicted depth value according to the candidate box and the multi-modal feature map.

[0050] That is to say, after entering the classification and regression stage, while retaining the dual classification head, using the multi-modal fusion features from FPN+CBAM, adding a depth consistency constraint branch, so that the depth geometric information can be effectively utilized in the subsequent evaluation of the loss function, thereby providing more comprehensive features for the classification and regression of candidate boxes, enhancing the model's understanding of the target geometry, reducing the scale misjudgment caused by perspective changes (such as small targets in UAV aerial photography), and at the same time, in the occlusion scene, using the depth information to supplement the geometric features missing in RGB, improving the detection rate of tail class targets. Optimize the design of the loss function to alleviate the dominant position of the head classes (such as common building materials) in the loss calculation and ensure the effective backpropagation of the gradients of the tail classes (such as special equipment).

[0051] As an example, using the depth decoder to decode the predicted depth value according to the candidate box and the multi-modal feature map includes: using the region of interest alignment module to extract the region of interest features according to the candidate box and the multi-modal feature map; the region of interest features generate depth features through the fully connected layer; using the activation function to normalize the depth features to obtain the predicted depth value.

[0052] It should be noted that the BBH module improves the detection accuracy through the dual classification head and depth constraints: after the candidate box extracts features through RoIAlign, the tail class classification head (BBH(T)) and the head class classification head (BBH(H)) independently predict the class probability and regression offset; the depth decoder predicts the normalized depth value from the features, calculates the L1 loss with the LiDAR true depth to enhance geometric perception; in the total loss function, the weighting factor of the tail class classification loss is inversely proportional to the confidence to optimize the model bias under the long-tail distribution.

[0053] S106, construct the total loss function according to the class prediction probability, bounding box offset and predicted depth value for model training.

[0054] S107. Use the trained object detection model for object detection to output the detection results.

[0055] That is to say, in the BBH module, the total loss function is adjusted by calculating the depth loss, and the robustness in complex environments is enhanced by multi-modal fusion: the head and tail feature boxes are obtained from the class bias sampler respectively, and are input into the tail class classification head BBH(T) and the head class classification head BBH(H) for processing. Then, the bounding box offset is independently predicted for each classification head; the depth value is decoded from the RoI feature and used for calculating the depth consistency loss, so as to optimize the calculation of the loss function. The total loss function includes classification loss, regression loss and depth consistency loss. Finally, separate predictions and result aggregation are performed to screen out the final detection box.

[0056] It should be noted that in the post-processing of the test stage, the detection results are optimized through multi-modal confidence fusion and dynamic strategy. First, the prediction results of BBH(T) and BBH(H) are merged, and NMS is performed after sorting by multi-modal confidence. The IoU threshold is set to 0.5, and the Top-100 detection boxes are retained. Secondary NMS is performed for the tail class targets, and the IoU threshold is reduced to 0.3 to avoid missed detections. The confidence threshold is dynamically adjusted: for small targets at a long distance (proxy depth value D norm > 0.7), the threshold is reduced to 0.4, and 0.6 is maintained in other scenarios. The bounding box calibration combines the depth prediction value to optimize the scale and reduce the size misjudgment under the downward shooting angle of the drone. Finally, the prediction results of the tail class and the head class are fused through NMS to output the detection box.

[0057] As a specific embodiment, as Figure 2 shown, the drone object detection method based on long-tail distribution optimization proposed in this application mainly includes three parts: the first part is the dynamic anchor box optimization of the region proposal network, the second part is the optimization of sampling the tail class distribution according to the multi-modal joint confidence in the biased classification sampler, and the third part is the optimization of the loss function by the double bounding box head through the depth consistency constraint branch.

[0058] Specifically, the dynamic anchor box optimization of the region proposal network includes the following steps:

[0059] Step 1: Proxy depth generation

[0060] Generate the proxy depth using a fixed rule mapping scheme, perform depth interval division, and dynamically adjust the anchor box parameters. Extract the depth-related information from the multi-modal feature map output by the FPN as the basis for adjusting the anchor box. Feature map channel mean calculation: Calculate the channel mean of each spatial position (u, v) of each layer of the FPN as the proxy depth value:

[0061]

[0062] in, Represents the proxy depth map, reflecting the feature activation intensity; areas with high activation intensity (large mean) may correspond to nearby targets, while areas with low activation intensity may correspond to distant targets.

[0063] Normalization: Normalize the proxy depth value to the [0,1] range:

[0064]

[0065] in, Represents the proxy depth value of the (i)th layer FPN feature map. In the (i)th layer FPN feature map, the channel mean at position (u, v) is used to approximate the depth information of the position; represents the minimum proxy depth value of the FPN feature map of the (i)th layer, taking the minimum value of the mean of all position channels of the feature map of this layer; represents the maximum proxy depth value of the FPN feature map of the (i)th layer, taking the maximum value of the mean of all position channels of the feature map of this layer; Represents the normalized proxy depth value, mapping the original proxy depth value to the [0,1] interval for subsequent dynamic anchor frame adjustment.

[0066] Step 2: Depth interval division and anchor box parameter mapping

[0067] Divide the interval according to the proxy depth value and dynamically adjust the anchor box size and aspect ratio. Divided into 3 intervals, corresponding to different anchor box parameters:

[0068]

[0069] Step 3: Dynamic anchor box generation

[0070] Generate an anchor frame that fits the scene based on the proxy depth value. Use the three basic scales of traditional RPN (8 2 ,16 2 ,32 2 ) and 3 aspect ratios (1:1, 1:2, 2:1), for each position (u,v) of the feature map, according to Select the corresponding scaling factor and aspect ratio, and adjust the size of the anchor box:

[0071]

[0072] Among them, Base_Scale∈{8,16,32} represents the preset basic scale of each layer of FPN; Aspect Ratio represents the aspect ratio (such as 1:2 corresponds to 0.5); Represents the normalized proxy depth map.

[0073] Specifically, the optimization of the paranoid classifier stage includes the following steps:

[0074] Step 1: Calculate the multi-modal joint confidence

[0075] (1) RoIAlign feature extraction: For the candidate boxes Anchors generated by RPN i , extract the optimized multi-modal features through RoIAlign:

[0076] F RoI = RoIAling(P i′ , Anchors i ) ∈ R 7×7×256

[0077] Wherein, represents the multi-modal feature map optimized by FPN+CBAM; RoIAlign means aligning the candidate boxes to a fixed size (7×7).

[0078] (2) Modal feature separation: Split the RoI features into RGB and depth components (based on the original four-channel input structure):

[0079] F RGB = F RoI [:,:,0:192], F LiDAR = F RoI [:,:,192:256]

[0080] Channel splitting: Split the RoI feature F RoI ∈ R 7×7×256 into RGB components (the first 192 dimensions) and LiDAR components (the last 64 dimensions) according to a preset ratio, without an additional network.

[0081] (3) Classification score calculation: The RGB and LiDAR components respectively predict scores through independent fully connected layers (FC RGB , FC LiDAR ).

[0082] c RGB = Softmax(FC RGB (F RGB )), c LiDAR = Softmax(FC LiDAR (F LiDAR ))

[0083] Wherein, N class represents the total number of categories; represents an independent fully connected layer, outputting the probabilities of each category, and the number of parameters of the fully connected layer is 192×N class + 64×N class, similar to the original single-modal classification head; zero calculation for the separation operation: channel splitting only involves indexing operations and does not increase the computational burden.

[0084] (4) Dynamic weight assignment: Generate modal weights according to the scene complexity (such as the depth sparsity of the occlusion area):

[0085]

[0086] Among them, ∑F RGB represents the sum of feature activations of RGB components, reflecting the reliability of RGB information; ∑F LiDAR represents the sum of feature activations of the depth component, reflecting the reliability of depth information.

[0087] (5) Joint confidence calculation:

[0088] c multi = λ RGB ·c RGB + λ LiDAR ·c LiDAR

[0089] Among them, c RGB represents the classification score based on RGB features; c LiDAR represents the classification score based on LiDAR depth features; λ RGB , λ LiDAR represents the dynamic weight (adjusted according to the scene, such as increasing the LiDAR weight when there is occlusion).

[0090] Step 2: Optimize the tail class sampling strategy according to the confidence

[0091] (1) Tail class definition: Divide the tail class according to the number of samples of each category in the dataset (such as the number of samples N tail <100).

[0092] (2) Sampling probability adjustment: The sampling probability of the tail class candidate box is proportional to its joint confidence:

[0093]

[0094] Among them, c multi,j represents the joint confidence of the j-th tail class candidate box.

[0095] (3) Dynamic sampling:

[0096] CBS(T): Sample from the tail class candidate boxes according to p tail , and supplement the head class samples when the quantity is insufficient;

[0097] CBS(H): Uniformly sample from the head class candidate boxes to avoid overfitting.

[0098] Specifically, the BBH bidirectional box head classification stage includes the following steps:

[0099] Step 1: Dual-classification head feature processing

[0100] Input the head and tail feature boxes obtained by CBS sampling into the tail-class classification head BBH(T) and the head-class classification head BBH(H) respectively to improve classification and regression accuracy and enhance multi-modal feature expression.

[0101] BBH(T) (tail-class classification head):

[0102] F tail = FC tail1 (Flatten(F RoI )) ∈ R 1024

[0103]

[0104] BBH(H) (head-class classification head):

[0105] F head = FC head1 (Flatten(F RoI )) ∈ R 1024

[0106]

[0107] Among them, F RoI ∈ R 7×7×256 represents the candidate box features extracted from the FPN+CBAM feature map, with a size of 7×7×256; Flatten is a functional module (data processing operation), and Flatten(F RoI ) flattens the 7×7×256 feature map into a one-dimensional vector with a dimension of 7×7×256 = 12544, converting multi-dimensional features into a vector form that can be processed by the fully connected layer. The fully connected layer (FC) is a functional module (neural network layer), and FC tail1 is the first fully connected layer, with an input of 12544 dimensions and an output of 1024-dimensional feature F tail , extracting high-level semantic features and compressing the information dimension. FC tail2 is the second fully connected layer, with an input of 1024 dimensions and an output of N tail dimensions (the number of tail classes to be detected. Softmax is an activation function that normalizes the output into a probability distribution p tail , representing the probability that the candidate box belongs to each tail class.

[0108] Step 2: Establish a deep consistency loss branch

[0109] (1) Extract features of a fixed size from the feature map optimized by FPN+CBAM through RoIAlign:

[0110] For each candidate bounding box Anchors i , perform:

[0111] F RoI = RoIAling(P i′ , Anchors i ) ∈ R 7×7×256

[0112] First, for the diversity of candidate bounding box sizes and aspect ratios, RoIAlign uniformly maps them to a fixed size (such as 7×7), eliminating the dimensional difference problem during the processing of the fully connected layer. At the same time, combined with the multi-scale feature maps generated by FPN+CBAM, RoIAlign ensures the accurate matching of candidate bounding box features with the actual scale of the target through spatial alignment operations - the shallow features retain details such as edges and textures, and the deep features extract abstract semantics, thus taking into account the integrated expression of geometric details and semantic information. In addition, this operation retains geometric details through pixel-level alignment, effectively reducing the bounding box regression error caused by feature misalignment, especially having a significant optimization effect on small and occluded targets commonly found in drone images, ultimately achieving a comprehensive improvement in detection accuracy.

[0113] (2) Calculate the depth loss and perform depth consistency constraints:

[0114] Depth prediction, predict the normalized depth value through a neural network:

[0115] d pred = Sigmoid(FC depth (F RoI )) ∈ R 1

[0116] The input feature F Rol (the feature map obtained through the Rol operation) passes through the fully connected layer FC depth , generates an unnormalized depth feature, and uses the Sigmoid function to compress the output value to the range of 0 to 1 to obtain the normalized depth prediction value d pred .

[0117] Among them, F Rol is the input feature, the intermediate feature map obtained through the Rol operation, which may contain spatial or semantic information of the scene. FC depth is the fully connected layer dedicated to depth prediction, used to map high-dimensional features to scalar values. Sigmoid is the activation function to ensure that the output value is within the normalized range (0 to 1).

[0118] Depth loss calculation:

[0119]

[0120] By calculating the depth loss, the error between the predicted depth value and the true depth value is measured, which is used to optimize the model parameters. The L1 loss (mean absolute error) is adopted to directly calculate the sum of the absolute differences between the predicted value and the true value of each sample. This method is less sensitive to outliers and is suitable for depth estimation tasks (which may be affected by sensor noise).

[0121] Among them, d pred,i is the predicted depth value of the i-th sample, ranging from 0 to 1. d true,i is the true depth value of the i-th sample, coming from the normalized data preprocessed by LiDAR (also ranging from 0 to 1). L depth is the total loss value, which guides the model optimization through backpropagation.

[0122] Step 3: Bounding box regression prediction

[0123] Each classification head independently predicts the bounding box offset:[[]]

[0124] t eail = FC reg_tail (F tail ) ∈ R 4

[0125] t head = FC reg_head (F headl ) ∈ R 4

[0126] Among them, t tail , t head represent the predicted bounding box offsets (center coordinates, width and height scaling); FC reg represents the fully connected layer dedicated to regression.

[0127] Step 4: Loss function calculation

[0128] The total loss function includes classification loss, regression loss, and depth consistency loss:[[]]

[0129] L total = L cls + L reg + γ · L depth

[0130] (1) Classification loss (weighted cross-entropy):

[0131] Calculate the weighted cross-entropy for the tail class and the head class respectively:[[]]

[0132]

[0133] Among them, represents the weight of the tail class sample (c multi,i is the multi-modal joint confidence); ∈ = 1e-5 represents preventing division by zero error; represents the class probability distribution output by the tail class classification head; represents the class probability distribution output by the head class classification head; y i represents the true label, indicating whether sample i belongs to the current class (1 for belonging, 0 for not belonging).

[0134] (2) Regression loss (Smooth L1)

[0135]

[0136] Among them, represents the true bounding box offset;

[0137] SmoothL1 function:

[0138]

[0139] (3) Depth loss weight

[0140] L depth = ∑|d pred - d true |

[0141] Among them, γ = 0.1 represents the coefficient for balancing the depth loss.

[0142] Step 5: Prediction fusion in the test phase

[0143] Separate predictions:

[0144] BBH(T) only predicts the tail class probability p tail and the regression offset t tail ;

[0145] BBH(H) only predicts the head class probability p head and the regression offset t head .

[0146] Result aggregation:

[0147] Merge the two types of prediction results and filter the final detection boxes through NMS:

[0148] Detections = NMS(Concat(p tail , p head ), Concat(t tail , t head ))

[0149] In summary, according to the drone target detection method optimized based on the long-tail distribution proposed in this application, by combining multi-modal data with the drone long-tail target detection algorithm and applying it to the intelligent construction site safety monitoring problem, the high-resolution RGB image data and LiDAR point cloud data collected by the drone simultaneously are used for early fusion to construct a multi-modal input. By fusing visual texture and LiDAR depth information, the network can directly learn richer underlying features (such as the contour details of small targets and the geometric structure of occluded areas), and targeted optimization is carried out on key functional modules such as RPN, CBS, and BBH in the neural network, significantly improving the detection sensitivity to low-frequency but high-risk targets (such as loose scaffolding screws and damaged safety net edges) on the construction site. Even under complex environmental conditions (such as low light, haze, etc.), the depth information provided by multi-modal fusion makes up for the lack of visual texture and improves the robustness of the system.

[0150] In addition, an embodiment of the present application also proposes a computer-readable storage medium, on which a target detection program optimized based on the long-tail distribution is stored. When the target detection program optimized based on the long-tail distribution is executed by a processor, it implements the drone target detection method optimized based on the long-tail distribution as described above.

[0151] In addition, an embodiment of the present application also proposes a computer device including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the drone target detection method optimized based on the long-tail distribution as described above.

[0152] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0153] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 each process or multiple processes and / or blocks Figure 1means for the functions specified in one or more boxes.

[0154] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one Figure 1 one or more processes and / or boxes Figure 1 means for the functions specified in one or more boxes.

[0155] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one Figure 1 one or more processes and / or boxes Figure 1 means for the functions specified in one or more boxes.

[0156] It should be noted that, in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several distinct elements and by means of a suitably programmed computer. In a unit claim listing several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.

[0157] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0158] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

[0159] In the description of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.

[0160] In the present application, unless otherwise clearly defined and limited, the terms such as "mounted", "connected", "connected to", "fixed" and the like shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application may be understood according to specific circumstances.

[0161] In the present application, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely means that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "below" and "beneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely means that the horizontal height of the first feature is lower than that of the second feature.

[0162] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms should not be understood as necessarily referring to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0163] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. An unmanned aerial vehicle target detection method optimized based on long-tail distribution, characterized in that It includes the following steps: Obtain the image data and point cloud data collected by the drone in the scene, and preprocess and fuse the image data and the point cloud data to obtain fused data; Input the fused data into a feature extraction network for feature extraction to obtain a multi-scale feature map, and perform adaptive reweighting of channels and space on the multi-scale feature map to obtain a multi-modal feature map; Use a region proposal network to dynamically adjust the anchor boxes according to the multi-modal feature map to generate candidate boxes; Calculate the multi-modal joint confidence according to the candidate boxes and the multi-modal feature map, and optimize the sampling strategy of the class bias sampler through the multi-modal joint confidence to obtain corresponding head feature boxes and tail feature boxes; Input the head feature box into the head class classification head, and input the tail feature box into the tail class classification head for processing to obtain corresponding class prediction probabilities and bounding box offsets, and use a depth decoder to decode the predicted depth value according to the candidate boxes and the multi-modal feature map; Construct a total loss function according to the class prediction probability, the bounding box offset and the predicted depth value for model training; Use the trained object detection model for object detection to output the detection result.

2. The method for optimizing drone target detection based on long-tail distribution according to claim 1, wherein, Preprocess and fuse the image data and the point cloud data to obtain fused data, including: Perform synchronous alignment and normalization processing on the image data and the point cloud data to obtain an RGB image and a depth image with the same size; Stitch the RGB image and the depth image together along the channel dimension to obtain fused data.

3. The method for optimizing drone target detection based on long-tail distribution according to claim 1, wherein, Input the fused data into a feature extraction network for feature extraction to obtain a multi-scale feature map, and perform adaptive reweighting of channels and space on the multi-scale feature map to obtain a multi-modal feature map, including: Input the fused data into the residual neural network in the feature extraction network for processing to obtain a multi-layer feature map; Input the multi-layer feature map into the feature pyramid network in the feature extraction network for upsampling and lateral connection to obtain a multi-scale feature map; Insert a convolutional block attention module after the outputs of each layer of the feature pyramid network to perform adaptive reweighting of channels and space on the multi-scale feature map to obtain a multi-modal feature map.

4. The method for optimizing drone target detection based on long-tail distribution according to claim 1, wherein, Use a region proposal network to dynamically adjust the anchor boxes according to the multi-modal feature map to generate candidate boxes, including: Use a fixed rule mapping scheme to extract proxy depth values from the multi-modal feature map; Normalize the depth proxy values to obtain normalized proxy depth values; Obtain anchor box parameters within the corresponding division intervals according to the normalized proxy depth values, and dynamically adjust the anchor box size and aspect ratio according to the anchor box parameters to generate candidate boxes.

5. The method for optimizing drone target detection based on the long-tail distribution according to claim 1, wherein Calculate the multi-modal joint confidence according to the candidate boxes and the multi-modal feature map, including: Use a region of interest alignment module to extract region of interest features according to the candidate boxes and the multi-modal feature map; Split the region of interest features into RGB components and depth components, and predict the classification scores through independent fully connected layers; Generate modal weights based on the sum of feature activations of RGB components and the sum of feature activations of depth components; Calculate the multi-modal joint confidence according to the classification score and the modal weights.

6. The method for optimizing UAV target detection based on long-tail distribution according to claim 1, wherein, When optimizing the sampling strategy of the class-biased sampler through the multi-modal joint confidence, the class-biased sampler preferentially samples the tail feature boxes, where the sampling probability is proportional to the multi-modal joint confidence.

7. The method for optimizing UAV target detection based on long-tail distribution according to claim 1, wherein, Use a depth decoder to decode the predicted depth value according to the candidate box and the multi-modal feature map, including: Use the region of interest alignment module to extract region of interest features according to the candidate box and the multi-modal feature map; The region of interest features generate depth features through a fully connected layer; Use an activation function to normalize the depth features to obtain the predicted depth value.

8. A computer-readable storage medium, characterized in that, Stored thereon is an object detection program optimized based on the long-tail distribution. When the object detection program optimized based on the long-tail distribution is executed by a processor, it implements the method for optimizing the long-tail distribution-based unmanned aerial vehicle object detection according to any one of claims 1-7.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for optimizing the long-tail distribution-based unmanned aerial vehicle object detection according to any one of claims 1-7.

Citation Information

Cited By

  • Game interaction control method and device, computer equipment and storage medium

    CN120550406A

  • Methods, apparatus, equipment and media for histogram calculation of depth maps

    CN122573715A

  • Histogram calculation method and device for depth map, equipment and medium

    CN122573715B