Unmanned aerial vehicle multi-scale fusion detection method based on bimodal inconsistency suppression
By employing a cross-modal feature fusion method based on the BiFPN-CALNet model, the semantic conflict problem in multispectral images is resolved, improving the accuracy and robustness of UAV target detection, especially its detection capability in complex environments.
Patent Information
- Application Number
- CN202510971787.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-28
Smart Images

Figure CN120852933A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image detection technology, and in particular to a multi-scale fusion detection method for unmanned aerial vehicles based on dual-modal inconsistency suppression. Background Technology
[0002] With the rapid development of computer science and information technology, single sensors have certain limitations in scene information acquisition, especially in complex urban environments, severe weather conditions, and occlusion scenarios. Traditional image detection methods struggle to meet these challenges and are no longer sufficient for current application needs. In contrast, multi-sensor systems can capture more comprehensive information about the same scene. Therefore, target detection methods based on multi-source images have gradually attracted widespread attention.
[0003] However, multispectral images acquired by different sensors inevitably exhibit heterogeneity. This heterogeneity can lead to conflicts in complementary information within the multimodal data, affecting the accuracy of target classification and localization. Semantic conflicts arise when different modalities perceive the same object but assign it inconsistent semantic information. In such cases, fine-grained information crucial for classification may be masked or lost during fusion, resulting in confusion between the target category and the background. Semantic conflicts between different modalities can not only cause classification errors but also lead to target localization deviations, ultimately affecting the overall accuracy of detection and classification. Therefore, effectively resolving cross-modal semantic conflicts and eliminating semantic inconsistencies between different modalities is key to improving the accuracy of multimodal image detection. Summary of the Invention
[0004] The technical problem to be solved by this invention is to address the shortcomings of the prior art by providing a multi-scale fusion detection method for UAVs based on dual-modal inconsistency suppression, comprising the following steps:
[0005] Collect visible light and infrared images of complex environments, filter, register and label the collected visible light and infrared images of complex environments, and construct a cross-modal dataset;
[0006] Construct a BiFPN-CALNet model and train it using a cross-modal dataset to obtain a trained BiFPN-CALNet model.
[0007] The BiFPN-CALNet model includes a dual-stream backbone network, a CALNet network, a weighted bidirectional feature pyramid network BiFPN, and a detection head. The dual-stream backbone network includes two branches, each of which includes a bidirectional decoupling focusing layer and a feature extraction sub-network. The CALNet network includes a CCR cross-modal conflict correction module and an SCF selective cross-modal fusion module.
[0008] Based on the trained BiFPN-CALNet model, the target detection results are obtained by inputting the visible light image and infrared image of the UAV to be detected.
[0009] Optionally, visible light and infrared images of complex environments are acquired, and the acquired visible light and infrared images of complex environments are filtered, registered, and labeled to construct a cross-modal dataset, including:
[0010] Use drones equipped with dual-spectrum imaging devices to acquire visible light and infrared images of other drones flying in complex environments;
[0011] The acquired visible light and infrared images are filtered, and clear, unblurred visible light and infrared images covering the same area are selected to form one-to-one infrared-visible light data pairs. The visible light and infrared images in the infrared-visible light data pairs are then registered.
[0012] The registered infrared and visible light data pairs are labeled to obtain the labels of the registered infrared and visible light data pairs. A cross-modal dataset is constructed based on the registered infrared and visible light data pairs and their corresponding labels. The cross-modal dataset is divided into training set, test set and validation set according to the proportion. The labels include the bounding box coordinates of the target and the target category information.
[0013] Optionally, the visible light image and infrared image in the infrared-visible light data pair are registered, specifically by:
[0014] Global information features that can represent the entire infrared image and visible light image are extracted respectively;
[0015] To construct feature descriptors for the extracted global information features, including scale information, orientation information, and neighborhood information;
[0016] The similarity of global information features between infrared and visible light images is determined based on feature descriptors, and a correspondence between infrared and visible light images is established.
[0017] Based on the correspondence between infrared and visible light images, geometric transformations and linear interpolation are performed on the infrared and visible light images to obtain registered infrared and visible light images.
[0018] Optionally, a BiFPN-CALNet model is constructed and trained using a cross-modal dataset to obtain a trained BiFPN-CALNet model, including:
[0019] A dual-stream backbone network was constructed, and the network was used to extract visible light image features and infrared image features at multiple levels.
[0020] A CALNet network is constructed, and the visible light image features and infrared image features of the corresponding levels extracted by the dual-stream backbone network are input into the CALNet network for feature fusion to obtain cross-modal features of multiple levels.
[0021] The weighted bidirectional feature pyramid network BiFPN is used to perform multi-scale fusion of cross-modal features at multiple levels to obtain a fused feature map;
[0022] The fused feature map output by the weighted bidirectional feature pyramid network BiFPN is transmitted to the detection head to obtain the target detection results predicted by the BiFPN-CALNet model.
[0023] Optionally, a dual-stream backbone network is constructed, and multi-level visible light image features and multi-level infrared image features are extracted using the dual-stream backbone network. The specific method is as follows:
[0024] A dual-stream backbone network is constructed, consisting of two branches. Each branch includes a bidirectional decoupling focusing layer and a feature extraction subnetwork. The bidirectional decoupling focusing layer is constructed using a bidirectional decoupling focusing mechanism, and the feature extraction subnetwork includes multiple convolutional layers, outputting feature maps at multiple levels. The first branch of the dual-stream backbone network is used to extract visible light image features at multiple levels, and the second branch of the dual-stream backbone network is used to extract infrared image features at multiple levels.
[0025] The bidirectional decoupling and focusing mechanism includes:
[0026] The image of the input bidirectional decoupled focusing layer is sliced, and the pixels of the odd and even columns and rows are decoupled separately. The image of the input bidirectional decoupled focusing layer is divided into two different feature groups, and the features of each feature group are extracted.
[0027] In the channel dimension, the features of each feature group are concatenated with the image of the input bidirectional decoupled focusing layer, and the features of the image are extracted using depthwise separable convolution.
[0028] Optionally, a CALNet network is constructed, and the visible light image features and infrared image features of the corresponding levels extracted by the dual-stream backbone network are input into the CALNet network for feature fusion to obtain cross-modal features at multiple levels. The specific method is as follows:
[0029] Construct the CALNet network, including the CCR cross-modal conflict correction module and the SCF selective cross-modal fusion module;
[0030] The visible light and infrared image features of the corresponding levels extracted from the dual-stream backbone network are input into the CALNet network. The CCR cross-modal conflict correction module is used to correct the visible light and infrared image features of the corresponding levels. The corrected visible light and infrared modal features of the corresponding levels are then passed to the SCF selective cross-modal fusion module. The SCF selective cross-modal fusion module is used to fuse the corrected visible light and infrared modal features of the corresponding levels to obtain the cross-modal features of the corresponding levels.
[0031] Optionally, the CCR cross-modal conflict correction module is used to correct the corresponding levels of visible light image features and infrared image features. The specific method is as follows:
[0032] The visible light image features r and infrared image features t extracted from the dual-stream backbone network at the corresponding level are concatenated along the channel dimension to obtain the multispectral features f.
[0033] The multispectral feature f is segmented into multiple feature patches, and the multispectral feature f is mapped to a set of weight matrices, including: query matrix W. Q Key matrix W K Sum matrix W V ;
[0034] Use depthwise convolution to reduce the key matrix W K Sum matrix W V The dimension of the key matrix W is obtained by reducing its dimension. K Sum matrix W V ′;
[0035] Based on query matrix W Q and the reduced-dimensional bond matrix W K Calculate the similarity attention weight matrix A;
[0036] From the similarity attention weight matrix A, select k nearest feature neighbor patches to generate an adjacency matrix X. Then, use attention weighting to adjust the dimensionality-reduced value matrix W. V The intermediate features are obtained by fusing the features with the adjacency matrix X.
[0037] For intermediate features Perform backpropagation to obtain the fused feature f l+1 ;
[0038] Use the Split slicing operation to fuse features f l+1 The corresponding visible light mode features r′ and infrared mode features t′ are obtained after correction.
[0039] Optionally, the SCF selective cross-modal fusion module is used to perform feature fusion on the corrected visible light modal features and infrared modal features at the corresponding level. The specific method is as follows:
[0040] Normalize the visible light modal features r′ and infrared modal features t′ at the corresponding levels after correction, and calculate the channel weight coefficients of the visible light modal features r′ and infrared modal features t′ at the corresponding levels after correction. Adaptive weighting is performed on the normalized visible light modal features and infrared modal features at the corresponding levels to obtain the adaptively weighted visible light modal features r″ and infrared modal features t″ at the corresponding levels;
[0041] The visible light modal features r″ and infrared modal features t″ of the corresponding level are concatenated after adaptive weighting to obtain the intermodal features f′ of the corresponding level.
[0042] A cross-modal attention mechanism is used to perform fine-grained optimization of the inter-modal features f′ at the corresponding level. Average pooling (AvgPool) and max pooling (MaxPool) are applied to the inter-modal features f′ at the corresponding level, respectively. After processing by a multilayer perceptron, the features are concatenated and then passed through a sigmoid function to obtain the refined inter-modal features f′ at the corresponding level. c As shown in the formula below:
[0043] H c =Sigmoid(MLP(AvgPool(f′))+MLP(MaxPool(f′)))
[0044] f c =H c ·f′
[0045] Where MLP(·) is a multilayer perceptron, AvgPool(·) is an average pooling operation, MaxPool(·) is a max pooling operation, and H... c These can be considered as channel weight coefficients in a channel attention map;
[0046] The Sigmoid function and convolutional filters are used to further compress the corresponding level of refined inter-modal features f. c The dimensions are used to obtain the corresponding level of inter-modal fusion features. As shown in the formula below:
[0047] H s (f c =Sigmoid(C([AvgPool(f c MaxPool(f) c )]))
[0048] f = H s
[0049] Where C is the convolution filter, H s These can be considered as spatial weight coefficients of a spatial attention map;
[0050] By utilizing the inverse matrix operation of depthwise separable convolution, corresponding levels of intermodal features are fused. Transformed into visible light features with a consistent representation at the corresponding level. and infrared features As shown in the formula below:
[0051]
[0052] Where R(·) is the inverse matrix operation of depthwise separable convolution;
[0053] Visible light features with consistent representation at corresponding levels and infrared features Concatenating the features yields cross-modal features at the corresponding level.
[0054] Optionally, a weighted bidirectional feature pyramid network (BiFPN) is used to perform multi-scale fusion of cross-modal features at multiple levels to obtain a fused feature map. The specific method is as follows:
[0055] For cross-modal features at multiple levels, a bidirectional fusion strategy is adopted, starting with the higher-level cross-modal features. To lower-level cross-modal features The fusion is shown in the following formula:
[0056]
[0057] Where Conv represents the convolution operation;
[0058] Then from low-level cross-modal features To higher-level cross-modal features The feature map to be optimized is obtained by fusion. As shown in the formula below:
[0059]
[0060] UpSample is the upsampling operation;
[0061] Learnable weight coefficients are used to assign different weights to cross-modal features at different levels, and the fused feature map is obtained through multiple iterations of weighted fusion optimization.
[0062] Optionally, the BiFPN-CALNet model can be trained using a cross-modal dataset to obtain a trained BiFPN-CALNet model. The specific method is as follows:
[0063] Step 1: Randomly select a batch of registered infrared and visible light data pairs from the training set and input them into the BiFPN-CALNet model;
[0064] Step 2: Perform forward propagation, calculate the gradient of the loss function and update the model parameters accordingly to evaluate model performance;
[0065] Step 3: Perform backpropagation, calculate the gradient of the loss function, and update the model parameters;
[0066] Step 4: Repeat forward and backward propagation until the loss function converges;
[0067] Step 5: Repeat steps 1 to 4 until the entire training set has been traversed.
[0068] The beneficial effects of adopting the above technical solution are as follows: The UAV multi-scale fusion detection method based on bimodal inconsistency suppression provided by this invention effectively alleviates the heterogeneity problem of cross-modal images and improves the detection capability of UAV targets of different sizes through multi-scale feature fusion. With the help of cross-modal conflict correction and multi-scale feature transfer, this model can achieve more accurate target detection in complex environments, providing an efficient and reliable solution for UAV detection tasks. Attached Figure Description
[0069] Figure 1 This is a flowchart of a UAV multi-scale fusion detection method based on dual-modal inconsistency suppression provided in an embodiment of the present invention;
[0070] Figure 2 This is a diagram of the BiFPN-CALNet model structure provided in an embodiment of the present invention;
[0071] Figure 3 This is a schematic diagram of the bidirectional decoupling focusing mechanism provided in an embodiment of the present invention;
[0072] Figure 4 This is a diagram of the CALNet network structure provided in an embodiment of the present invention;
[0073] Figure 5 This is a structural diagram of the CCR cross-modal conflict correction module provided in an embodiment of the present invention;
[0074] Figure 6 This is a diagram of the weighted bidirectional feature pyramid network (BiFPN) structure provided in an embodiment of the present invention. Detailed Implementation
[0075] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0076] This embodiment presents a UAV multi-scale fusion detection method based on dual-modal inconsistency suppression, such as... Figure 1 As shown, it includes the following steps:
[0077] S1 acquires visible light and infrared images of complex environments, filters, registers, and labels the acquired visible light and infrared images of complex environments, and constructs a cross-modal dataset.
[0078] The S101 uses a drone equipped with a dual-spectrum imaging device to capture visible light and infrared images of other drones flying in complex environments;
[0079] In this embodiment, a DJI Mavic 3T drone equipped with visible light and infrared camera systems is used to collect visible light and infrared images of other drones flying in complex environments, so as to ensure that visible light and infrared images of complex environments under different spectral modes are acquired simultaneously.
[0080] S102 filters the acquired visible light and infrared images, selects clear and unblurred visible light and infrared images that cover the same area to form a one-to-one infrared visible light data pair, and registers the visible light and infrared images in the infrared visible light data pair.
[0081] The visible light image and infrared image in the infrared-visible light data pair are registered using the following method:
[0082] Global information features that can represent the entire infrared image and visible light image are extracted respectively;
[0083] To construct feature descriptors for the extracted global information features, including scale information, orientation information, and neighborhood information;
[0084] The similarity of global information features between infrared and visible light images is determined based on feature descriptors, and a correspondence between infrared and visible light images is established.
[0085] Based on the correspondence between infrared and visible light images, geometric transformations and linear interpolation are performed on the infrared and visible light images to obtain registered infrared and visible light images.
[0086] In this embodiment, the SIFT algorithm is used to register the visible light image and infrared image in an infrared-visible light data pair. SIFT (Scale Invariant Feature Transform) is a classic feature-based image registration method widely used in image registration. The SIFT algorithm first processes the image using a multi-kernel Gaussian filter function and constructs a Gaussian scale space. Then, it generates a difference scale space through the difference operation of adjacent convolutional images to extract extreme points. Subsequently, the extracted extreme points are filtered, feature point matching is performed, erroneous matching points are removed, and descriptors for the feature points are generated. Due to the high dimension of the descriptors, they can more completely express the image features. The SIFT algorithm has strong robustness in rotation transformations and scale changes, and also has a certain degree of noise resistance, thus exhibiting good stability in various complex environments. The SIFT algorithm is used to register the visible light image and infrared image in an infrared-visible light data pair, including the following steps:
[0087] 1) Construct the scale space L(x,y,σ) of the image to be registered using a Gaussian kernel function, as shown in the following formula:
[0088] L(x,y,σ)=G(x,y,σ)*I(x,y)
[0089]
[0090] Where G(x,y,σ) is the Gaussian kernel, I(x,y) is the image to be registered, x and y are the coordinates of any pixel in the image to be registered, σ is the scale space factor, and * is the convolution operation;
[0091] 2) To initially obtain the extreme points, a Difference of Gaussian (DoG) operation is performed on adjacent convolutional images to generate a difference scale space, thereby extracting the extreme points of the image. These extreme points are used as key points for initial localization. The DoG calculation formula is shown below:
[0092] D(x,y,σ)=L(x,y,kσ)-L(x,y,σ)
[0093] Where k is a constant multiplier;
[0094] 3) For each keypoint (x, y), calculate the orientation information θ(x, y) to obtain rotation invariance, as shown in the following formula:
[0095]
[0096] Where M(x,y) is the modulus of the keypoint (x,y), and L xL represents the scale space value in the horizontal direction of the keypoint (x,y). y The scale space value in the vertical direction of the key point (x,y);
[0097] 4) Generate SIFT feature T based on the orientation information of the key point (x,y), as shown in the following formula:
[0098] T = (t1, t2, ..., t W )
[0099] Where W is the dimension of SIFT feature T, t1, t2, ..., t W For the 1st, 2nd, ..., Wth descriptors;
[0100] In this embodiment, the generated SIFT feature T has a dimension of 128.
[0101] 5) Evaluate the similarity between image A and image B by performing Euclidean distance matching on the SIFT features of the two images to be registered, and calculate the i-th descriptor of image A. and the i-th descriptor of image B The Euclidean distance between them is shown in the following formula:
[0102]
[0103] 6) Compare the Euclidean distance between the descriptors of image A and image B, and use KNN nearest neighbor matching to select the two nearest neighbors as reliable feature points d1 and d2. Set a ratio threshold Z, and filter matching point pairs whose reliable feature point ratio is less than the threshold, as shown in the following formula:
[0104]
[0105] In this embodiment, the ratio threshold Z is set to 0.7;
[0106] 7) Remove erroneous matches from the matching points using the Random Sample Consensus (RANSAC) algorithm.
[0107] The homography transformation matrix F is estimated using the following formula:
[0108]
[0109] The random sampling consensus algorithm (RANSAC) is set to have N = 2000 iterations and an error threshold of ε = 3 pixels. All erroneous matching points are eliminated to achieve image registration.
[0110] S103 labels the registered infrared and visible light data pairs to obtain the labels of the registered infrared and visible light data pairs. Based on the registered infrared and visible light data pairs and their corresponding labels, a cross-modal dataset is constructed. The cross-modal dataset is divided into a training set, a test set, and a validation set according to the proportions. The labels include the bounding box coordinates of the target and the target category information.
[0111] In this embodiment, the Labelimg annotation tool is used to annotate the registered infrared and visible light data pairs. The registered infrared and visible light data pairs can share a single label, which includes the bounding box coordinates of the target and the target category information. A cross-modal dataset is constructed based on the registered infrared and visible light data pairs and their corresponding labels, and the cross-modal dataset is divided into a training set, a validation set, and a test set in a ratio of 8:1:1.
[0112] S2 constructs a BiFPN-CALNet model and trains it using a cross-modal dataset to obtain a trained BiFPN-CALNet model.
[0113] The BiFPN-CALNet model includes a dual-stream backbone network, a CALNet network, a weighted bidirectional feature pyramid network (BiFPN), and a detection head, such as... Figure 2 As shown, the dual-stream backbone network includes two branches, each branch including a bidirectional decoupling focusing layer and a feature extraction sub-network. The CALNet network includes a CCR cross-modal conflict correction module and an SCF selective cross-modal fusion module.
[0114] In this embodiment, to effectively address the semantic conflict problem in multispectral data, multiple levels of visible light image features and multiple levels of infrared image features output from the dual-stream backbone network are input into the CCR cross-modal conflict correction module of the CALNet network. By selecting the cross-modal pixels with the highest similarity for interaction, the semantic inconsistency between different modalities is reduced, and cross-modal conflicts are alleviated, resulting in corrected multiple levels of visible light modal features and infrared modal features. The multiple levels of visible light modal features and infrared modal features processed by the CCR module are then input into the SCF module to filter key features, further optimizing the fusion of cross-modal information and enhancing the discriminative ability of features, resulting in multiple levels of cross-modal features. The cross-modal features are then input into the weighted bidirectional feature pyramid network BiFPN for multi-scale feature fusion, improving the detection capability of the BiFPN-CALNet model for targets at different scales, resulting in a fused feature map. The fused feature map is then input into the detection head to predict the final target detection result.
[0115] S201 constructs a dual-stream backbone network, which includes two branches. Each branch includes a bidirectional decoupling focusing layer and a feature extraction sub-network. The bidirectional decoupling focusing layer is constructed using a bidirectional decoupling focusing mechanism. The feature extraction sub-network includes multiple convolutional layers and outputs feature maps at multiple levels. The first branch of the dual-stream backbone network is used to extract visible light image features at multiple levels, and the second branch of the dual-stream backbone network is used to extract infrared image features at multiple levels.
[0116] Two-way decoupling focusing mechanism such as Figure 3 As shown, it includes:
[0117] The image of the input bidirectional decoupled focusing layer is sliced, and the pixels of the odd and even columns and rows are decoupled separately. The image of the input bidirectional decoupled focusing layer is divided into two different feature groups, and the features of each feature group are extracted.
[0118] In the channel dimension, the features of each feature group are concatenated with the image of the input bidirectional decoupled focusing layer, and the features of the image are extracted using depthwise separable convolution.
[0119] The bidirectional decoupling focusing mechanism divides the image input to the bidirectional decoupling focusing layer into two feature groups. By decoupling sampling in the horizontal and vertical directions, it expands the receptive field and effectively avoids information loss in a single direction, thereby improving the detection capability of elongated or tilted targets. Combining depthwise convolution with multi-branch design, depthwise separable convolution reduces computational complexity while further extracting features, enhances the spatial adaptability of feature extraction, achieves coverage of a larger area, improves context awareness, and reduces the number of parameters, thus optimizing computational overhead while ensuring detection accuracy.
[0120] S202 constructs the CALNet network, including the CCR cross-modal conflict correction module and the SCF selective cross-modal fusion module;
[0121] The visible light image features and infrared image features extracted from the dual-stream backbone network at the corresponding levels are input into the CALNet network, such as... Figure 4 As shown, the CCR cross-modal conflict correction module is used to correct the visible light image features and infrared image features of the corresponding level. The corrected visible light modal features and infrared modal features of the corresponding level are then passed to the SCF selective cross-modal fusion module. The SCF selective cross-modal fusion module is used to fuse the corrected visible light modal features and infrared modal features of the corresponding level to obtain the cross-modal features of the corresponding level.
[0122] The CCR cross-modal conflict correction module projects multiple levels of visible light image features and multiple levels of infrared image features extracted from the dual-stream backbone network into a shared joint space, establishes nearest neighbor patch relationships, and extracts contextual information associations between visible light image features and multi-level infrared image features.
[0123] The CCR cross-modal conflict correction module is used to correct the corresponding levels of visible light image features and infrared image features, such as... Figure 5 As shown, the specific method is as follows:
[0124] The visible light image features r and infrared image features t extracted from the dual-stream backbone network at the corresponding level are concatenated along the channel dimension to obtain the multispectral features f, as shown in the following formula;
[0125] f = Concat(r,t)
[0126] Concat(·) is the concatenation operation;
[0127] The multispectral feature f is segmented into multiple feature patches, and the multispectral feature f is mapped to a set of weight matrices, including: query matrix W. Q Key matrix W K Sum matrix W V ;
[0128] The bond matrix W is reduced using a depthwise convolution with a kernel size of s×s and a stride of s. K Sum matrix W V The dimension of the key matrix W is obtained by reducing its dimension. K Sum matrix W V To reduce computational overhead, the following formula is used:
[0129]
[0130] Among them, W K ′ is the reduced-dimensional bond matrix, W V ' is the value matrix after dimensionality reduction, and DWConv(·) is the depthwise convolution. Let n = m × h × w be the image scale information input to the CCR cross-modal conflict correction module, m be the number of input images, h be the height of the input images, w be the width of the input images, and d be the width of the input images. k For coefficients;
[0131] The CCR cross-modal conflict correction module introduces a similarity attention mechanism to compute the query matrix of the i-th feature patch. The key matrix of the h-th feature patch The Euclidean distance between them is h≠i, and the k nearest neighbor feature patches with the smallest distance are selected. The information of the i-th feature patch and the most similar feature patch with a different modality from the i-th feature patch are fused. The similarity attention mechanism is calculated as follows:
[0132]
[0133] Among them, A i Let be the similarity attention weight matrix for the i-th feature patch. Let i be the query matrix for the i-th feature patch. For query matrix The key matrix of the k nearest neighbors, where d is the scaling factor. This is the set of nearest neighbor feature patches to the i-th feature patch;
[0134] Since calculating the Euclidean distance between the query matrix and all key matrices one by one is computationally complex, this embodiment uses matrix multiplication to optimize the similarity attention calculation process, thereby improving the computational efficiency of the similarity attention mechanism. This is based on the query matrix W. Q and the reduced-dimensional bond matrix W K The similarity attention weight matrix A is calculated as shown in the following formula:
[0135] A = f l W Q ·f l (W K ′) T
[0136] Among them, f l Input features for the similarity attention mechanism;
[0137] From the similarity attention weight matrix A, select k nearest neighbor notification patches to generate an adjacency matrix X. Then, use attention weighting to adjust the dimensionality-reduced value matrix W. V The intermediate features are obtained by fusing the features with the adjacency matrix X. This optimizes the performance of similarity attention, enabling more accurate capture of cross-modal contextual information, as shown in the following formula:
[0138]
[0139] Among them, [X] ij Let A be the element in the i-th row and j-th column of the adjacency matrix based on top-k sampling. ij Let be the element in the i-th row and j-th column of the similarity attention weight matrix, and let top-k(rowj) be the element in the similarity attention weight matrix A. ij The first k largest similarity values in the j-th row, Att knn (·) represents the attention-weighted operation;
[0140] For intermediate features Perform backpropagation to obtain the fused feature f l+1 As shown in the formula below:
[0141]
[0142] Where LN(·) is the layer normalization operation and Reverse(·) is the back propagation operation;
[0143] Use the Split slicing operation to fuse features f l+1 The corresponding visible light mode features r′ and infrared mode features t′ are obtained after correction.
[0144] In this embodiment, the CCR cross-modal conflict correction module aggregates a single-modal patch with some of the most similar patches from outside the modality, and uses external context information to reduce the heterogeneity between patches at the corresponding positions, thereby correcting cross-modal semantic conflicts.
[0145] The visible light modal features r′ and infrared modal features t′, after being corrected by the CCR cross-modal conflict correction module, are passed to the SCF selective cross-modal fusion module. The SCF selective cross-modal fusion module assigns weights to intra-modal features through an adaptive mechanism, filters out low-semantic features, selects semantically rich features, and combines inter-modal consistency information to enrich the semantic representation within the modality. It also mines inter-modal complementary information to fuse multi-modal features, thereby further reducing heterogeneity between cross-modal features.
[0146] The SCF selective cross-modal fusion module is used to fuse the visible light modal features and infrared modal features at the corresponding levels after correction. The specific method is as follows:
[0147] Normalize the visible light modal features r′ and infrared modal features t′ at the corresponding levels after correction, and calculate the channel weight coefficients of the visible light modal features r′ and infrared modal features t′ at the corresponding levels after correction. Adaptive weighting is applied to the normalized visible light modal features and infrared modal features at the corresponding levels to obtain the adaptively weighted visible light modal features r″ and infrared modal features t″ at the corresponding levels, as shown in the following formula:
[0148]
[0149] in, These are the weighting coefficients for the characteristic channels of the visible light modes. These are the weighting coefficients for the infrared modal characteristic channels. Let be the weighting coefficient of the m-th channel in the visible light modal feature r′. This is the sum of the weighting coefficients of all channels in the visible light modal feature r′. Let be the weighting coefficient of the m-th channel in the infrared modal feature t′. t′ is the sum of the weight coefficients of all channels in the infrared modal feature t′, BN(·) is the normalization process, and S(·) is the Sigmoid function;
[0150] The SCF selective cross-modal fusion module filters out redundant information and retains refined intramodal features by assigning selective weights to intramodal features under conditions such as low light.
[0151] The visible light modal features r″ and infrared modal features t″ of the corresponding level are concatenated after adaptive weighting to obtain the intermodal features f′ of the corresponding level, so as to further reduce the heterogeneity between modes;
[0152] A cross-modal attention mechanism is used to perform fine-grained optimization of the inter-modal features f′ at the corresponding level. Average pooling (AvgPool) and max pooling (MaxPool) are applied to the inter-modal features f′ at the corresponding level, respectively. After processing by a multilayer perceptron, the features are concatenated and then passed through a sigmoid function to obtain the refined inter-modal features f′ at the corresponding level. c As shown in the formula below:
[0153] H c =Sigmoid(MLP(AvgPool(f′))+MLP(MaxPool(f′)))
[0154] f c =H c ·f′
[0155] Where MLP(·) is a multilayer perceptron, AvgPool(·) is an average pooling operation, MaxPool(·) is a max pooling operation, and H... c These can be considered as channel weight coefficients in a channel attention map;
[0156] The Sigmoid function and convolutional filters are used to further compress the corresponding level of refined inter-modal features f. c The dimensions are used to obtain the corresponding level of inter-modal fusion features. As shown in the formula below:
[0157] H s (f c =Sigmoid(C([AvgPool(f c MaxPool(f) c )]))
[0158] f = H s
[0159] Where C is the convolution filter, H s These can be considered as spatial weight coefficients of a spatial attention map;
[0160] By utilizing the inverse matrix operation of depthwise separable convolution, corresponding levels of intermodal features are fused. Transform the corresponding level of visible light features with a consistent representation. and infrared features As shown in the formula below:
[0161]
[0162] Where R(·) is the inverse matrix operation of depthwise separable convolution;
[0163] Visible light features with consistent representation at corresponding levels and infrared features Concatenating the features yields cross-modal features at the corresponding level. As shown in the formula below:
[0164]
[0165] The SCF selective cross-modal fusion module calculates the importance of intra-modal feature weights and selects important features; it aggregates multimodal transport features using spatial and channel attention mechanisms; and it transforms inter-modal features into intra-modal features, ensuring that single-modal features contain consistent representations of multiple modalities. The SCF selective cross-modal fusion module improves cross-modal fusion by selecting rich semantic features. During feature interaction, it extracts rich intra-modal semantic information through feature selection mechanisms, while further enhancing inter-modal semantic consistency during feature fusion. This improves the accuracy and robustness of cross-modal target detection and enhances the performance of multispectral target detection.
[0166] S203 uses a weighted bidirectional feature pyramid network (BiFPN) to perform multi-scale fusion of cross-modal features at multiple levels to obtain a fused feature map;
[0167] The core idea of BiFPN is to perform bidirectional fusion of feature maps at different scales to obtain richer feature representation capabilities, such as... Figure 6 As shown, BiFPN employs a bidirectional feature fusion strategy that combines "top-down" and "bottom-up" approaches to fully utilize feature information at different scales. Specifically:
[0168] For cross-modal features at multiple levels, a bidirectional fusion strategy is adopted, fusing features from higher-level cross-modal features. To lower-level cross-modal features The fusion is performed as shown in the following formula:
[0169]
[0170] Where Conv represents the convolution operation;
[0171] Then from low-level cross-modal features To higher-level cross-modal features The feature map to be optimized is obtained by fusion. As shown in the formula below:
[0172]
[0173] UpSample is the upsampling operation;
[0174] The Bidirectional Feature Pyramid Network (BiFPN) effectively combines the advantages of features at different scales through its bidirectional feature fusion strategy. It preserves the global semantics of low-resolution features while utilizing the fine structural information of high-resolution features, thereby improving the accuracy of object detection. Furthermore, BiFPN employs a weighted feature fusion mechanism, assigning learnable weights to features at different resolutions during the feature fusion process to adaptively adjust the contribution of each feature map. This weighting method dynamically optimizes the fusion effect of features at different scales, reducing the decrease in detection accuracy caused by resolution differences and making the model more robust when handling multi-scale targets.
[0175] Learnable weight coefficients are used to assign different weights to cross-modal features at different levels, and the fused feature map is obtained through multiple iterations of weighted fusion optimization.
[0176] S204 transmits the fused feature map output by the weighted bidirectional feature pyramid network BiFPN to the detection head to obtain the target detection results predicted by the BiFPN-CALNet model;
[0177] S205 uses a cross-modal dataset to train the BiFPN-CALNet model, resulting in a trained BiFPN-CALNet model.
[0178] In this embodiment, a server is used to train the BiFPN-CALNet model. The training environment uses Python 3.9, the deep learning framework is PyTorch 1.10.1, CDUA is 11.3, the operating system is Ubuntu 20.04, the processor is a 16-core Intel i9-12900KF, the memory is 32GB, and the graphics card is an NVIDIA RTX3090Ti with 24GB of video memory. The training process includes:
[0179] Step 1: Randomly select a batch of registered infrared and visible light data pairs from the training set and input them into the BiFPN-CALNet model;
[0180] Step 2: Perform forward propagation, calculate the gradient of the loss function and update the model parameters accordingly to evaluate model performance;
[0181] Forward propagation refers to the process of passing the input data from the input layer through each hidden layer of the network in sequence, and finally outputting the prediction result. The specific steps are as follows:
[0182] Pass the input data to the network's input layer;
[0183] In each layer, a linear combination of the outputs of the previous layer is received, which is a weighted summation.
[0184] The linear combination z is used to generate a nonlinear output through the activation function f(z);
[0185] The output of each layer is used as the input of the next layer, and the activation of the linear transformation is performed sequentially until the output layer is reached;
[0186] The model's output is compared with the true labels, and the model's performance is measured by calculating the loss function, IOU_loss, as shown in the following formula:
[0187]
[0188] Among them, B gt B is the ground truth bounding box, and A is the predicted bounding box;
[0189] Step 3: Perform backpropagation, calculate the gradient of the loss function, and update the model parameters;
[0190] Backpropagation proceeds sequentially from the output layer to the input layer, calculating and storing the parameter gradients of the intermediate variables in the neural network. The specific steps are as follows:
[0191] Starting from the output layer, first calculate the gradient of the loss function IOU_loss with respect to the activation values of the output layer;
[0192] The gradient of the loss is propagated backward from the output layer to each layer, and the error δ of each layer is calculated.
[0193] The gradient of the loss function with respect to the parameters of the current layer is calculated based on the error δ. For each layer l, the weight W is calculated. l and bias b l The gradient;
[0194] The model parameters are updated based on the calculated gradient using gradient descent.
[0195] Step 4: Repeat forward and backward propagation until the loss function converges;
[0196] Step 5: Repeat steps 1 to 4 until the entire training set has been traversed;
[0197] S3 is based on a trained BiFPN-CALNet model. It takes a visible light image and an infrared image of the UAV to be detected as input and obtains the target detection result.
[0198] S301 extracts features from visible light and infrared images respectively using a dual-stream network;
[0199] To address the heterogeneity of images across different modalities, S302 uses the CALNet network for feature fusion, including:
[0200] The CCR cross-modal conflict correction module further processes features through spatial reduction attention and similarity attention mechanisms, helping the network to identify semantic conflicts between different modalities and correct the conflicts through contextual information, thereby obtaining semantically consistent cross-modal feature representations;
[0201] The features processed by the CCR cross-modal conflict correction module are input into the SCF selective cross-modal fusion module to filter key features, further optimize the fusion of cross-modal information, and enhance the distinguishability of features;
[0202] S303 inputs the fused multi-level features into the weighted bidirectional feature pyramid network BiFPN for multi-scale feature fusion, thereby improving the model's ability to detect targets at different scales.
[0203] S304 inputs the fused features into the detection head to obtain the target detection result.
[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the present invention.
Claims
1. A multi-scale fusion detection method for unmanned aerial vehicles (UAVs) based on dual-modal inconsistency suppression, characterized in that: Includes the following steps: Collect visible light and infrared images of complex environments, filter, register and label the collected visible light and infrared images of complex environments, and construct a cross-modal dataset; Construct a BiFPN-CALNet model and train it using a cross-modal dataset to obtain a trained BiFPN-CALNet model. The BiFPN-CALNet model includes a dual-stream backbone network, a CALNet network, a weighted bidirectional feature pyramid network BiFPN, and a detection head. The dual-stream backbone network includes two branches, each of which includes a bidirectional decoupling focusing layer and a feature extraction sub-network. The CALNet network includes a CCR cross-modal conflict correction module and an SCF selective cross-modal fusion module. Based on the trained BiFPN-CALNet model, the target detection results are obtained by inputting the visible light image and infrared image of the UAV to be detected.
2. The UAV multi-scale fusion detection method based on dual-modal inconsistency suppression according to claim 1, characterized in that: Acquire visible light and infrared images of complex environments, filter, register, and label the acquired images, and construct a cross-modal dataset, including: Use drones equipped with dual-spectrum imaging devices to acquire visible light and infrared images of other drones flying in complex environments; The acquired visible light and infrared images are filtered, and clear, unblurred visible light and infrared images covering the same area are selected to form one-to-one infrared-visible light data pairs. The visible light and infrared images in the infrared-visible light data pairs are then registered. The registered infrared and visible light data pairs are labeled to obtain the labels of the registered infrared and visible light data pairs. A cross-modal dataset is constructed based on the registered infrared and visible light data pairs and their corresponding labels. The cross-modal dataset is divided into training set, test set and validation set according to the proportion. The labels include the bounding box coordinates of the target and the target category information.
3. The UAV multi-scale fusion detection method based on dual-modal inconsistency suppression according to claim 2, characterized in that: The visible light image and infrared image in the infrared-visible light data pair are registered using the following method: Global information features that can represent the entire infrared image and visible light image are extracted respectively; To construct feature descriptors for the extracted global information features, including scale information, orientation information, and neighborhood information; The similarity of global information features between infrared and visible light images is determined based on feature descriptors, and a correspondence between infrared and visible light images is established. Based on the correspondence between infrared and visible light images, geometric transformations and linear interpolation are performed on the infrared and visible light images to obtain registered infrared and visible light images.
4. The UAV multi-scale fusion detection method based on dual-modal inconsistency suppression according to claim 1, characterized in that: Construct a BiFPN-CALNet model and train it using a cross-modal dataset to obtain a trained BiFPN-CALNet model, including: A dual-stream backbone network was constructed, and the network was used to extract visible light image features and infrared image features at multiple levels. A CALNet network is constructed, and the visible light image features and infrared image features of the corresponding levels extracted by the dual-stream backbone network are input into the CALNet network for feature fusion to obtain cross-modal features of multiple levels. The weighted bidirectional feature pyramid network BiFPN is used to perform multi-scale fusion of cross-modal features at multiple levels to obtain a fused feature map; The fused feature map output by the weighted bidirectional feature pyramid network BiFPN is transmitted to the detection head to obtain the target detection results predicted by the BiFPN-CALNet model.
5. The UAV multi-scale fusion detection method based on dual-modal inconsistency suppression according to claim 4, characterized in that: A dual-stream backbone network is constructed, and multi-level visible light image features and multi-level infrared image features are extracted using the dual-stream backbone network. The specific method is as follows: A dual-stream backbone network is constructed, consisting of two branches. Each branch includes a bidirectional decoupling focusing layer and a feature extraction subnetwork. The bidirectional decoupling focusing layer is constructed using a bidirectional decoupling focusing mechanism, and the feature extraction subnetwork includes multiple convolutional layers, outputting feature maps at multiple levels. The first branch of the dual-stream backbone network is used to extract visible light image features at multiple levels, and the second branch of the dual-stream backbone network is used to extract infrared image features at multiple levels. The bidirectional decoupling and focusing mechanism includes: The image of the input bidirectional decoupled focusing layer is sliced, and the pixels of the odd and even columns and rows are decoupled separately. The image of the input bidirectional decoupled focusing layer is divided into two different feature groups, and the features of each feature group are extracted. In the channel dimension, the features of each feature group are concatenated with the image of the input bidirectional decoupled focusing layer, and the features of the image are extracted using depthwise separable convolution.
6. The UAV multi-scale fusion detection method based on dual-modal inconsistency suppression according to claim 4, characterized in that: A CALNet network is constructed, and the visible light image features and infrared image features of corresponding levels extracted by the dual-stream backbone network are input into the CALNet network for feature fusion to obtain cross-modal features at multiple levels. The specific method is as follows: Construct the CALNet network, including the CCR cross-modal conflict correction module and the SCF selective cross-modal fusion module; The visible light and infrared image features of the corresponding levels extracted from the dual-stream backbone network are input into the CALNet network. The CCR cross-modal conflict correction module is used to correct the visible light and infrared image features of the corresponding levels. The corrected visible light and infrared modal features of the corresponding levels are then passed to the SCF selective cross-modal fusion module. The SCF selective cross-modal fusion module is used to fuse the corrected visible light and infrared modal features of the corresponding levels to obtain the cross-modal features of the corresponding levels.
7. The UAV multi-scale fusion detection method based on dual-modal inconsistency suppression according to claim 6, characterized in that: The CCR cross-modal conflict correction module is used to correct the corresponding levels of visible light image features and infrared image features. The specific method is as follows: The visible light image features r and infrared image features t extracted from the dual-stream backbone network at the corresponding level are concatenated along the channel dimension to obtain the multispectral features f. The multispectral feature f is segmented into multiple feature patches, and the multispectral feature f is mapped to a set of weight matrices, including: query matrix W. Q Key matrix W K Sum matrix W V ; Use depthwise convolution to reduce the key matrix W K Sum matrix W V The dimension of the key matrix W is obtained by reducing its dimension. K Sum matrix W V ′; Based on query matrix W Q and the reduced-dimensional bond matrix W K Calculate the similarity attention weight matrix A; From the similarity attention weight matrix A, select k nearest neighbor feature patches to generate the adjacency matrix X, and use attention weighting to adjust the dimensionality-reduced value matrix W. V The intermediate features are obtained by fusing the features with the adjacency matrix X. For intermediate features Perform backpropagation to obtain the fused feature f l+1 ; Use the Split slicing operation to fuse features f l+1 The corresponding visible light mode features r′ and infrared mode features t′ are obtained after correction.
8. The UAV multi-scale fusion detection method based on dual-modal inconsistency suppression according to claim 6, characterized in that: The SCF selective cross-modal fusion module is used to fuse the visible light modal features and infrared modal features at the corresponding levels after correction. The specific method is as follows: Normalize the visible light modal features r′ and infrared modal features t′ at the corresponding levels after correction, and calculate the channel weight coefficients of the visible light modal features r′ and infrared modal features t′ at the corresponding levels after correction. Adaptive weighting is performed on the normalized visible light modal features and infrared modal features at the corresponding levels to obtain the adaptively weighted visible light modal features r″ and infrared modal features t″ at the corresponding levels; The visible light modal features r″ and infrared modal features t″ of the corresponding level are concatenated after adaptive weighting to obtain the intermodal features f′ of the corresponding level. A cross-modal attention mechanism is used to perform fine-grained optimization of the inter-modal features f′ at the corresponding level. Average pooling (AvgPool) and max pooling (MaxPool) are applied to the inter-modal features f′ at the corresponding level, respectively. After processing by a multilayer perceptron, the features are concatenated and then passed through a sigmoid function to obtain the refined inter-modal features f′ at the corresponding level. c As shown in the formula below: H c =Sigmoid(MLP(AvgPool(f′))+MLP(MaxPool(f′))) f c =H c ·f′ Where MLP(·) is a multilayer perceptron, AvgPool(·) is an average pooling operation, MaxPool(·) is a max pooling operation, and H... c These can be considered as channel weight coefficients in a channel attention map; The Sigmoid function and convolutional filters are used to further compress the corresponding level of refined inter-modal features f. c The dimensions are used to obtain the corresponding level of inter-modal fusion features. As shown in the formula below: H s (f c )=Sigmoid(C([AvgPool(f c );MaxPool(f c )])) f=H s Where C is the convolution filter, H s These can be considered as spatial weight coefficients of a spatial attention map; By utilizing the inverse matrix operation of depthwise separable convolution, corresponding levels of intermodal features are fused. Transformed into visible light features with a consistent representation at the corresponding level. and infrared features As shown in the formula below: Where R(·) is the inverse matrix operation of depthwise separable convolution; Visible light features with consistent representation at corresponding levels and infrared features Concatenating the features yields cross-modal features at the corresponding level.
9. The UAV multi-scale fusion detection method based on dual-modal inconsistency suppression according to claim 4, characterized in that: The weighted bidirectional feature pyramid network BiFPN is used to perform multi-scale fusion of cross-modal features at multiple levels to obtain a fused feature map. The specific method is as follows: For cross-modal features at multiple levels, a bidirectional fusion strategy is adopted, starting with the higher-level cross-modal features. To lower-level cross-modal features The fusion is shown in the following formula: Where Conv represents the convolution operation; Then from low-level cross-modal features To higher-level cross-modal features The feature map to be optimized is obtained by fusion. As shown in the formula below: UpSample is the upsampling operation; Learnable weight coefficients are used to assign different weights to cross-modal features at different levels, and the fused feature map is obtained through multiple iterations of weighted fusion optimization.
10. The UAV multi-scale fusion detection method based on dual-modal inconsistency suppression according to claim 1, characterized in that: The BiFPN-CALNet model is trained using a cross-modal dataset to obtain a pre-trained BiFPN-CALNet model. The specific method is as follows: Step 1: Randomly select a batch of registered infrared and visible light data pairs from the training set and input them into the BiFPN-CALNet model; Step 2: Perform forward propagation, calculate the gradient of the loss function and update the model parameters accordingly to evaluate model performance; Step 3: Perform backpropagation, calculate the gradient of the loss function, and update the model parameters; Step 4: Repeat forward and backward propagation until the loss function converges; Step 5: Repeat steps 1 to 4 until the entire training set has been traversed.
Citation Information
Cited By
Wild animal detection method fusing unmanned aerial vehicle thermal infrared image and visible light image
CN121305621A
Multi-modal adaptive fine-grained feature fusion method based on view angle of unmanned aerial vehicle
CN121330556A
RGB (Red, Green, Blue) and thermal infrared fusion detection system and method for complex visual environment
CN122115891A