Space-time consistent multi-modal feature fusion air target detection method
Through the aerial target detection method of space-time consistent multimodal feature fusion, the problems of span modal feature coupling interference and insufficient feature response in the existing technology are solved, and higher detection accuracy and robustness are achieved, and aerial target detection in complex scenarios are suitable.
Patent Information
- Application Number
- CN202510420437.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-04
- Publication Date
- 2025-06-24
AI Technical Summary
Existing aerial object detection technology is difficult to effectively establish the context timing correlation between visible light and infrared modes, resulting in mutual interference when coupling across modal features. In addition, traditional feature fusion strategies lack cross-scale feature interaction mechanisms and insufficient response to edge features of weak targets.
The aerial target detection method using space-time consistent multimodal feature fusion is adopted, and the target motion trajectory characteristics are extracted through the time-domain enhancement module, and the air-domain fusion module combines the complementary characteristics of visible light and infrared images, and the detection network is optimized based on the attention mechanism.
It significantly improves the detection accuracy and robustness of aerial targets, reduces the difficulty of detection of complex weather and weak targets, and is suitable for drone supervision, low-altitude security and smart city air traffic management.
Smart Images

Figure CN120198650A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and target detection, and specifically to an airborne target detection method with spatio-temporal consistent multi-modal feature fusion, which is applicable to drone supervision, low-altitude security, and air traffic management in smart cities. Background Technique
[0002] With the rapid development of drone technology and the wide application of the low-altitude economy, airborne target detection faces severe challenges and significant opportunities. There are various means of airborne target detection, including radar detection, radio detection, acoustic detection, and optoelectronic detection, etc. Among them, the optoelectronic technology combined with artificial intelligence has become the mainstream technical solution due to its advantages of low cost, high precision, and flexible deployment. However, existing single-modal enhancement methods are difficult to effectively establish the context time-series correlation between the visible light and infrared modalities, resulting in easy mutual interference during cross-modal feature coupling; traditional feature fusion strategies lack a cross-scale feature interaction mechanism, have insufficient response to the edge features of small targets, and have weak dynamic adaptive ability in the channel dimension; the existing detection networks have limited ability to represent spatio-temporal consistent features, fail to effectively aggregate the information of multi-resolution feature pyramids, and lack an attention focusing mechanism for small target areas, resulting in difficulty in balancing the detection accuracy and computational efficiency of the model; in addition, the low-altitude security scenario has extremely high requirements for real-time performance, and the existing fusion algorithms have high computational complexity and are difficult to meet the deployment requirements of edge devices.
[0003] To address these problems, the present invention proposes an airborne target detection method with spatio-temporal consistent multi-modal feature fusion. By extracting the target motion trajectory features through a time-domain enhancement module, combining the complementary characteristics of visible light and infrared images using a spatial-domain fusion module, and optimizing the detection network based on the attention mechanism, the detection accuracy and robustness of airborne targets are significantly improved. At the same time, through lightweight design, the real-time performance requirements are met, providing efficient and reliable technical support for drone supervision, low-altitude security, and air traffic management in smart cities. Summary of the Invention
[0004] The present invention is to solve the above-mentioned deficiencies existing in the prior art, and proposes an airborne target detection method with spatio-temporal consistent multi-modal feature fusion, in order to efficiently fuse different modal features through the joint optimization of multi-source data complementarity and attention mechanism, so as to reduce the detection difficulty for complex weather, small targets, etc. in the air, thereby effectively solving the problem of weak airborne target features, improving the detection effect, and effectively enhancing the detection accuracy and robustness of airborne targets in complex scenarios, and is applicable to fields such as drone supervision and low-altitude security.
[0005] To achieve the above invention purpose, the present invention adopts the following technical solutions:
[0006] The characteristics of an air target detection method with spatio-temporal consistent multi-modal feature fusion according to the present invention are as follows, including the following steps:
[0007] Step 1: Use a binocular camera to synchronously acquire visible light images and infrared images of air targets in multiple scenarios and perform preprocessing, so as to obtain an air target visible light image dataset and an infrared image dataset , where represents the nth visible light image, represents the nth infrared image, n = 1, 2, …, N, and N is the total number of pairs of air visible light and infrared images; let or The true label of the target category in is denoted as , and ∈{1, 2, …, C}, where C represents the total number of target categories;
[0008] Step 2: Construct a spatio-temporal consistent air target detection network, including: a time-domain enhanced backbone network, a spatial domain fusion module, a bidirectional feature pyramid, and a Head module, and perform processing on and to obtain the probability distribution of the target category
[0009] Step 3: Construct a total loss function ;
[0010] Step 4: Use the backpropagation algorithm to optimize and train the target detection network model composed of the time-domain enhanced backbone network, the spatial domain fusion module, the bidirectional pyramid, and the Head module branches, and calculate the total loss function to update the model parameters until the total loss function converges, so as to obtain the optimal target detection model after training, which is used to predict the air target detection in infrared images and visible light images, and obtain an image with predicted target boxes and predicted categories.
[0011] The characteristics of an air target detection method with spatio-temporal consistent multi-modal feature fusion according to the present invention are also that the step 2 includes the following steps:
[0012] Step 2.1: Construct a time-domain enhanced backbone network, including: a visible light branch and three infrared branches;
[0013] Step 2.1.1: The visible light branch includes: three depthwise separable convolution modules, a temporal shift module, and a local attention module, and perform processing on to obtain the first feature map of the nth visible light, the enhanced feature , the second visible light feature map and the third visible light feature map ;
[0014] Step 2.1.2: The first infrared branch includes: a convolutional module, a C2f module, and a temporal adaptive module, and processes to obtain the nth infrared first feature map and the nth infrared enhanced image feature ;
[0015] Step 2.2: Construct a spatial domain fusion module, including: an edge feature extraction module and a feature fusion module, and processes and to obtain the nth first-stage fusion feature ; used to construct a fusion loss function ;
[0016] Step 2.3: The second infrared branch includes: a convolutional module, a C2f module, and splices with and then processes to obtain the nth infrared second feature map ;
[0017] Step 2.4: Input and into the spatial domain fusion module for processing, and output the nth second-stage fusion feature ;
[0018] Step 2.5: The third infrared branch includes: a convolutional module, a C2f module, and splices with and then processes to obtain the nth infrared third feature map ;
[0019] Step 2.6: Input and into the spatial domain fusion module for processing, and output the nth third-stage fusion feature ;
[0020] Step 2.7: Construct a bidirectional feature pyramid Bi-FPN, including: three upsampling modules, two splicing modules, and three channel spatial attention modules, and performs a top-down process on , and to output the nth first multi-scale feature map , the second multi-scale feature map and the third multi-scale feature map ;
[0021] Step 2.8: Construct a Head module, including: a classification branch and a detection branch, and process , and respectively, and output the probability distribution of the target class of the fused image of the nth visible light image and the nth infrared image as well as the coordinates and sizes of the target bounding boxes.
[0022] Furthermore, the said step 2.1.1 includes the following process:
[0023] Input it into the first depthwise separable convolution module in the visible light branch for processing, and output the first feature map of the nth visible light image ;
[0024] The temporal shift module performs a shift operation on to obtain the nth visible light feature map using Equation (1) :
[0025] (1)
[0026] In Equation (1), is the shift coefficient, represents the width of the image; represents taking the left width of which is the region of, represents taking the middle width of which is W of the region, represents taking the right width of which is the region of, represents concatenation;
[0027] Input it into the local attention module for processing to generate the nth visible light attention map , and after multiplying it element-wise with , obtain the nth visible light enhanced feature ;
[0028] Input it into the second depthwise separable convolution module for processing, and output the second feature map of the nth visible light image ; Input it into the third depthwise separable convolution module for processing to obtain the third feature map of the nth visible light image .
[0029] Furthermore, the said step 2.1.2 includes the following process:
[0030] Processed by the first convolutional module in the input infrared branch, and the nth infrared feature map is output ;
[0031] Split in the first C2f module to obtain the first split feature map and the second split feature map , and processed by multiple bottleneck residual blocks for to obtain the second residual feature map , and and are concatenated and then subjected to convolutional processing to output the nth infrared first feature map ;
[0032] The temporal adaptive module includes: a local branch and a global branch;
[0033] The local branch uses two layers of 1D temporal convolutional layers and ReLU non-linear operations to process to obtain the nth infrared local feature ;
[0034] The global branch uses two layers of fully connected layers and a Softmax normalization function to process to obtain the nth infrared global feature , and and are concatenated to obtain the nth infrared feature , and then processed by Gaussian filtering and residual learning for to output the nth infrared enhanced image feature .
[0035] Furthermore, step 2.2 includes the following process:
[0036] The edge feature extraction module uses the visible light convolution to be learned and the infrared convolution to extract the edge features of and and the edge features of respectively ;
[0037] The feature fusion module performs element-wise subtraction operations on and as well as and respectively to generate the nth visible light difference feature map and the nth infrared difference feature map ;
[0038] The feature fusion module respectively uses global average pooling for and to process, and correspondingly generates the nth visible light one-dimensional feature vector and the nth infrared one-dimensional feature vector . Then, the fully connected layer is used to respectively perform non-linear mapping on and to correspondingly obtain the nth visible light channel attention weight and the nth infrared channel attention weight . Multiply with as well as with respectively for channel multiplication operation to correspondingly generate the nth visible light fusion feature and the nth infrared fusion feature . Thus, and are stacked into the nth first-stage fusion feature .
[0039] Furthermore, in step 2.2, equation (2) is used to construct :
[0040] (2)
[0041] In equation (2), represents the intensity consistency loss, represents the gradient consistency loss, represents the structural similarity loss, and λ1, λ2, and λ3 represent three weight parameters, and there are:
[0042] (3)
[0043] (4)
[0044] (5)
[0045] In equations (3)-(5), represents the height of the image, represents the width of the image, max takes the maximum value, represents the image gradient, and SSIM represents the image similarity function.
[0046] Furthermore, step 2.7 includes the following steps:
[0047] After being processed by the first upsampling module, the nth first upsampling feature map ; After adjusting the number of channels using a 1×1 convolution operation on , it is input together with into the first splicing module for processing to obtain the nth first multi-scale fused feature map , and then input into the first channel-spatial attention module for processing to output the nth first multi-scale feature map ;
[0048] After being processed by the second upsampling module, the nth second upsampled feature map is output. After adjusting the number of channels using a 1×1 convolution operation on , it is input together with into the second splicing module for processing to obtain the nth second multi-scale fused feature map , and then input into the second channel-spatial attention module for processing to output the nth second multi-scale feature map ;
[0049] After being processed by the third upsampling module, the nth third upsampled feature map is output, and then input into the third channel-spatial attention module for processing to output the nth third multi-scale feature map .
[0050] Further, the first channel-spatial attention module obtains the nth first multi-scale feature map according to the following process :
[0051] Perform global average pooling processing without dimensionality reduction on to output the nth channel statistical vector , and then compress along the spatial dimension to finally generate the nth channel statistical vector ;
[0052] Use one-dimensional convolution to process to generate the nth channel weight matrix , and after multiplying it channel by channel with , output the nth channel enhanced feature ;
[0053] Perform axis average pooling processing on along the width and height axes respectively to correspondingly generate the nth statistical matrix on the width and the nth statistical matrix on the height , and then and After splicing, convolution processing is performed to generate the nth spatial weight matrix , finally, is multiplied by element by element, and the nth first multi-scale feature map is output .
[0054] Further, in step 3, formula (7) is used to construct :
[0055] (7)
[0056] In formula (7), and are two weight parameters; represents the classification loss, represents the regression loss, and there is:
[0057] (8)
[0058] (9)
[0059] In formula (8) and formula (9), represents the predicted probability of the cth category in; represents the true label of the cth category in, is the Euclidean distance, and are respectively the center point of the target box and the center point of the true box of the image after fusion of the nth visible light image and the nth infrared image, c is the diagonal length of the minimum bounding box, α is the weight coefficient, v is the penalty term of the aspect ratio, and IOU represents the intersection over union of the predicted box and the true box.
[0060] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the air target detection method, and the processor is configured to execute the program stored in the memory.
[0061] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0062] 1. The present invention adopts a dual-branch independent enhancement scheme to perform context temporal correlation on the visible light and infrared light channels respectively, enhancing the visible infrared image features, solving the problem of mutual interference that is prone to occur during cross-modal feature coupling, and thus better preparing for multi-modal fusion;
[0063] 2. The multi-modal fusion module adopted by the present invention fuses the features of infrared and visible light images through global average pooling, fully connected layers, and channel multiplication, solves the problem that traditional feature fusion strategies lack a cross-scale feature interaction mechanism, and the generated spatio-temporally consistent multi-modal feature maps can improve the detection effect.
[0064] 3. The attention mechanism detection network designed by the present invention uses a bidirectional pyramid to aggregate fusion features of different resolutions, and combines channel spatial attention to highlight small target regions, solves the problems of failing to effectively aggregate multi-resolution feature pyramid information and lacking an attention focusing mechanism for small target regions, so that an optimal model can be trained to detect the category and position of aerial targets with high precision. Brief Description of the Drawings
[0065] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0066] Figure 2 It is a structural diagram of the spatio-temporally consistent aerial target detection network of the present invention;
[0067] Figure 3 It is a structural diagram of the temporal adaptive module of the present invention;
[0068] Figure 4 It is a structural diagram of the spatial domain fusion module CAMF of the present invention;
[0069] Figure 5 It is a structural block diagram of the aerial target detection device of the present invention. Detailed Embodiment
[0070] In this embodiment, to solve the problems in the prior art such as low efficiency of multi-modal data fusion in complex aerial scenes, jumping of detection results due to spatio-temporal inconsistency, serious problems of missed detection and false detection of small targets, and insufficient real-time performance, an efficient detection scheme based on dynamic adaptive feature fusion and spatio-temporal joint optimization is proposed to achieve high-precision and low-latency aerial target perception. Specifically, as Figure 1 shown, an aerial target detection method based on spatio-temporally consistent multi-modal fusion includes the following steps:
[0071] Step 1: Use a binocular camera to synchronously acquire visible light images and infrared images of aerial targets in multiple scenarios and perform preprocessing to obtain an aerial target visible light image dataset and an infrared image dataset , where represents the nth visible light image, represents the nth infrared image, n = 1, 2,..., N, and N is the total number of pairs of aerial visible light and infrared images; let or The true label of the target category in is denoted as and ∈{1,2,…,C}, where C represents the total number of target categories;
[0072] The dataset of the present invention is obtained by synchronously collecting visible light and infrared modality image data through the mainstream device FLIR Duo Pro R, and constructing an aerial target dataset covering different weather conditions, light intensities, and complex backgrounds; in the data preprocessing stage, an adaptive histogram equalization and enhancement algorithm is used to perform low-light compensation on visible light images, and a non-uniformity correction technique is used to eliminate fixed pattern noise in infrared images.
[0073] Step 2: Construct a spatio-temporally consistent aerial target detection network, as shown in Figure 2 , which includes: a temporal enhancement backbone network, a spatial domain fusion module, a bidirectional feature pyramid, and a Head module, and processes and to obtain the probability distribution of the target category and the coordinates and sizes of the target bounding box;
[0074] Step 2.1: Construct a temporal enhancement backbone network, including: a visible light branch and three infrared branches;
[0075] Step 2.1.1: The visible light branch includes: three depthwise separable convolution modules, a temporal shift module, and a local attention module, and processes to obtain the nth visible light first feature map , the nth visible light enhanced feature , the visible light second feature map and the visible light third feature map ;
[0076] It is processed in the first depthwise separable convolution module in the visible light branch, and the nth visible light first feature map is output;
[0077] The temporal shift module performs a shift operation on , performs regional segmentation and shift operations in the spatial dimension in the width direction to enhance local context information, and thus obtains the nth visible light feature map using Equation (1):
[0078] (1)
[0079] In Equation (1), is the shift coefficient. In this embodiment, = 0.2, represents the width of the image; represents taking The left width is the area of indicating to take The middle width is the area of W, indicating to take The right width is the area of indicating splicing;
[0080] Input into the local attention module for processing to generate the nth visible light attention map , and after element-wise multiplication with , obtain the nth visible light enhanced feature ; further enhance the detail features of the low-contrast region, obtained and can be expressed as: , where σ is the Sigmoid function; Conv is convolution;
[0081] Input into the second depthwise separable convolution module for processing and output the nth visible light second feature map ; Input into the third depthwise separable convolution module for processing to obtain the nth visible light third feature map .
[0082] Step 2.1.2: The first infrared branch includes: a convolution module, a C2f module, and a temporal adaptive module, and processes to obtain the nth infrared first feature map and the nth infrared enhanced image feature ;
[0083] All convolution modules in the time domain enhanced backbone network of the present invention use convolution kernels, and the bottleneck residual blocks in the C2f module all sequentially use dimensionality reduction, convolution, and residual connection for processing;
[0084] Input into the first convolution module in the infrared branch for processing and output the nth infrared feature map ;
[0085] Input into the first C2f module for splitting to obtain the first split feature map and the second split feature map , use multiple bottleneck residual blocks to process to obtain the second residual feature map , and and After splicing and then performing convolution processing, the nth infrared first feature map is output ;
[0086] The temporal adaptive module aligns the features of adjacent frames through deformable convolution and fuses the temporal attention with weights. The specific implementation is as Figure 3 shown, including: a local branch and a global branch;
[0087] The local branch uses two layers of 1D temporal convolutional layers and ReLU non-linear operations to process, and obtains the nth infrared local feature ;
[0088] The global branch uses two layers of fully connected layers and the Softmax normalization function to process, and obtains the nth infrared global feature , and and are spliced to obtain the nth infrared feature , and then Gaussian filtering and residual learning are used to process, and the nth infrared enhanced image feature is output ;
[0089] Specifically, Gaussian smoothing can smooth the input signal and reduce the noise fluctuations caused by random thermal motion. Residual learning further optimizes the noise suppression effect by learning the difference between the input signal and the smoothed signal, while avoiding over-smoothing of the signal itself. This combined method performs well in dealing with thermal noise in infrared imaging and can retain the key features of the image while removing the noise;
[0090] This method performs well in terms of light weight. It uses depth-deformable convolution to replace standard convolution to reduce FLOPs; in addition, the temporal shift module and the temporal adaptive module are only inserted at key layers and are not globally used to reduce the computational amount. While ensuring real-time performance, it significantly improves the detection robustness of the target in dynamic scenes through temporal enhancement, especially suitable for complex environments with low light and thermal noise.
[0091] Step 2.2: Construct a spatial domain fusion module, as Figure 4 shown, including: an edge feature extraction module and a feature fusion module, and process and to obtain the nth first-stage fusion feature ; used to construct a fusion loss function ;
[0092] The edge feature extraction module EFM uses the visible light convolution to be learned And infrared convolution Extract separately Edge features And Edge features ;
[0093] The feature fusion module performs element-wise subtraction operations on And And And respectively to generate the nth visible light difference feature map And the nth infrared difference feature map ;
[0094] The feature fusion module uses global average pooling to process And respectively, and correspondingly generates the nth visible light one-dimensional feature vector And the nth infrared one-dimensional feature vector , and then passes through the fully connected layer to perform non-linear mapping on And respectively, and correspondingly obtains the nth visible light channel attention weight And the nth infrared channel attention weight . Multiply And And And respectively for channel multiplication operations, and correspondingly generates the nth visible light fusion feature And the nth infrared fusion feature , thereby And Are superimposed to form the nth first-stage fusion feature .
[0095] Construct using equation (2) :
[0096] (2)
[0097] In equation (2), Represents the intensity consistency loss, Represents the gradient consistency loss, Represents the structural similarity loss, , , Represents 3 weight parameters. In this embodiment, =2, =10, =1, and there is:
[0098] (3)
[0099] (4)
[0100] (5)
[0101] Equations (3) - (5), represents the height of the image, represents the width of the image, and max takes the maximum value, represents the image gradient, and SSIM represents the image similarity function.
[0102] Step 2.3: The second infrared branch includes: a convolutional module, a C2f module, and combines with After splicing and processing, the nth infrared second feature map is obtained .
[0103] Step 2.4: Input and into the spatial domain fusion module for processing, and output the nth second-stage fusion feature .
[0104] Step 2.5: The third infrared branch includes: a convolutional module, a C2f module, and combines with After splicing and processing, the nth infrared third feature map is obtained .
[0105] Step 2.6: Input and into the spatial domain fusion module for processing, and output the nth third-stage fusion feature .
[0106] Step 2.7: Construct a bidirectional feature pyramid Bi-FPN, including: three upsampling modules, two splicing modules, and three channel spatial attention modules, and perform top-down processing on , and to output the nth first multi-scale feature map , the second multi-scale feature map and the third multi-scale feature map ;
[0107] In this embodiment, the upsampling modules all achieve resolution improvement and channel adjustment through nearest neighbor upsampling and convolution;
[0108] After being processed by the first upsampling module, the nth first upsampling feature map is output ; After adjusting the number of channels using a 1×1 convolution operation on , it is input together with into the first splicing module for processing to obtain the nth first multi-scale fused feature map , and then input into the first channel-spatial attention module for processing to output the nth first multi-scale feature map ;
[0109] After being processed by the second upsampling module, the nth second upsampled feature map is output. After adjusting the number of channels using a 1×1 convolution operation on , it is input together with into the second splicing module for processing to obtain the nth second multi-scale fused feature map , and then input into the second channel-spatial attention module for processing to output the nth second multi-scale feature map ;
[0110] After being processed by the third upsampling module, the nth third upsampled feature map is output, and then input into the third channel-spatial attention module for processing to output the nth third multi-scale feature map ;
[0111] The first channel-spatial attention module obtains the nth first multi-scale feature map according to the following process :
[0112] Perform global average pooling without dimensionality reduction on to output the nth channel statistical vector , and then compress along the spatial dimension, and finally generate the nth channel statistical vector ;
[0113] Use one-dimensional convolution to process to generate the nth channel weight matrix , and after multiplying it channel by channel with , output the nth channel enhanced feature ;
[0114] Perform axis average pooling on along the width and height axes respectively to generate the nth statistical matrix on the width and the nth statistical matrix on the height , and then splice and and then perform convolution processing to generate the nth spatial weight matrix Finally, after multiplying and element by element, the n-th first multi-scale feature map is output with and output the n-th first multi-scale feature map .
[0115] Step 2.8: Construct the Head module, including: a classification branch and a detection branch, and process , and respectively, and output the probability distribution of the target class of the fused image of the n-th visible light image and the n-th infrared image , and respectively, and output the probability distribution of the target class of the fused image of the n-th visible light image and the n-th infrared image as well as the coordinates and sizes of the target boxes.
[0116] Step 3: Construct the total loss function ;
[0117] Construct using Equation (7): :
[0118] (7)
[0119] In Equation (7), and are two weight parameters. To maintain the dominance of the detection loss, here and are set to 1.0 and 0.5 respectively; represents the classification loss, represents the regression loss, and there is:
[0120] (8)
[0121] (9)
[0122] In Equations (8) and (9), represents the predicted probability of the c-th class in ; represents the true label of the c-th class in and are the center points of the target box and the true box of the fused image of the n-th visible light image and the n-th infrared image respectively, c is the diagonal length of the minimum bounding box, α is the weight coefficient, in this embodiment, α is defaulted to 0.05, v is the penalty term of the aspect ratio, and IOU represents the intersection over union of the predicted box and the true box.
[0123] Step 4: Use the backpropagation algorithm to optimize and train the object detection network model composed of the time-domain enhanced backbone network, the spatial domain fusion module, the bidirectional pyramid, and the Head module branches, and calculate the total loss function to update the model parameters until the total loss function converges, so as to obtain the optimal target detection model after training, which is used to predict the aerial target detection in infrared images and visible light images, and obtain an image with predicted target boxes and predicted categories.
[0124] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0125] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.
[0126] On the other hand, the present invention proposes an aerial target detection device based on spatio-temporal consistent multi-modal fusion, specifically as Figure 5 shown, including: a data acquisition unit 901, a model construction unit 902, and a detection output unit 903, where:
[0127] Data acquisition unit 901: This unit adopts an advanced integrated design, combines a high-resolution visible light camera with a high-sensitivity infrared thermal imager, and is equipped with the mainstream device FLIR Duo Pro R. Through hardware-level time synchronization and spatial calibration technology, it realizes the precise alignment and fusion of multi-source data, significantly improves the accuracy and efficiency of target detection, temperature monitoring, and scene analysis, and is applicable to various scenarios such as drone patrol, security monitoring, and industrial detection;
[0128] Model construction unit 902: Used to construct a spatio-temporal consistent multi-modal fusion detection model. The multi-modal fusion detection model includes: using a multi-branch convolutional neural network to extract multi-scale features of infrared images and visible light images, including shallow texture features and deep semantic features; aligning the feature dimensions of different modalities through a 1x1 convolutional layer to reduce the differences between modalities; using the ReLU activation function to enhance the non-linear expression ability of features. Design a cross-modal feature fusion module, generate channel attention weights through global average pooling and fully connected layers, dynamically adjust the fusion ratio of infrared and visible light features, and use element-wise addition and channel multiplication operations to fuse the complementary information of infrared and visible light features to generate spatio-temporal consistent multi-modal fusion features; retain the original feature information through residual connections to avoid information loss. Embed a spatio-temporal attention mechanism in the multi-scale feature maps in the bidirectional pyramid to enhance the spatio-temporal consistency representation of the target area, embed a channel spatial attention mechanism to enhance the response weight of small target areas, use a decoupled detection head to separate the tasks of target classification, localization, and trajectory prediction, and optimize the multi-task loss function through an adaptive weighting strategy;
[0129] Detection output unit 903: It is used to perform target detection and result output. The specific process is as follows: Input the preprocessed visible light and infrared image sequences into a pre-trained multi-modal fusion detection model. The model generates target detection results through end-to-end inference. It supports real-time processing of video streams of 1080p and above, and the inference time for a single frame is ≤20 ms. Transmit the detection results to the display terminal or control center in real time, with a delay ≤50 ms, and provide a visualization interface to superimpose detection frames and trajectory prediction lines on the original image, supporting the playback and analysis of historical data.
Claims
1. A method for aerial target detection by fusion of spatiotemporal consistent multimodal features, characterized in that: The following steps are involved: Step 1: Use a binocular camera to synchronously acquire visible light images and infrared images of aerial targets in multiple scenes and perform preprocessing to obtain a visible light image dataset of aerial targets. and infrared image datasets ,in, represents the nth visible light image, represents the nth infrared image, n=1,2,…,N, N is the total number of visible light and infrared images in the air; let or The true label of the target category in is recorded as ,and ∈{1,2,…,C}, C represents the total number of target categories; Step 2: Build a time-space consistent aerial target detection network, including: time domain enhancement backbone network, spatial domain fusion module, bidirectional feature pyramid, and Head module, and and Processing is performed to obtain the probability distribution of the target category And the coordinates and size of the target box; Step 3: Construct the total loss function ; Step 4: Use the back propagation algorithm to enhance the time domain backbone network, spatial domain fusion module, bidirectional pyramid and Head The target detection network model composed of module branches is optimized and trained, and the total loss function is calculated To update the model parameters until the total loss function The optimal mid-range target detection model is obtained after training until convergence, which is used to predict the detection of aerial targets in infrared images and visible light images, and obtain images with predicted target boxes and predicted category labels.
2. The method for aerial target detection by spatiotemporal consistent multimodal feature fusion according to claim 1, characterized in that: The step 2 comprises the following steps: Step 2.1: Construct a time domain enhancement backbone network, including a visible light branch and three infrared branches; Step 2.1.1: The visible light branch includes: three depth-separable convolution modules, a temporal shift module, a local attention module, and Processing is performed to obtain the nth visible light first feature map , the nth visible light enhanced feature , visible light second characteristic map and the third characteristic diagram of visible light ; Step 2.1.2: The first infrared branch includes: a convolution module, a C2f module, a timing adaptation module, and Processing is performed to obtain the nth infrared first feature map and the nth infrared enhanced image features ; Step 2.2: Construct a spatial fusion module, including edge feature extraction module and feature fusion module, and and Processing is performed to obtain the nth first-stage fusion feature ; Used to construct the fusion loss function ; Step 2.3: The second infrared branch includes: a convolution module, a C2f module, and and After splicing, the second infrared feature map is obtained. ; Step 2.4: and Input the spatial fusion module for processing and output the nth second-stage fusion feature ; Step 2.5: The third infrared branch includes: a convolution module, a C2f module, and and After splicing, the third infrared feature map is obtained. ; Step 2.6: and Input the spatial fusion module for processing, and output the nth third-stage fusion feature ; Step 2.7: Construct a bidirectional feature pyramid Bi-FPN, including three upsampling modules, two splicing modules, three channel space attention modules, and , and Perform top-down processing and output the nth first multi-scale feature map , the second multi-scale feature map And the third multi-scale feature map ; Step 2.8: Build the Head module, including the classification branch and the detection branch, and , and Processing is performed and the probability distribution of the target category of the image fused from the nth visible light image and the nth infrared image is output accordingly And the coordinates and size of the target box.
3. The method for aerial target detection by spatiotemporal consistent multimodal feature fusion according to claim 2, characterized in that: Step 2.1.1 The process includes: The input is processed in the first depth-separable convolution module in the visible light branch, and the nth visible light first feature map is output. ; The timing offset module Perform the offset operation, and then use formula (1) to obtain the nth visible light feature map : (1) In formula (1), is the offset coefficient, Indicates the width of the image; Indicates taking The left width is area, Indicates taking The middle width is The area of W, Indicates taking The right width is area, Indicates splicing; Input into the local attention module for processing to generate the nth visible light attention map , and After element-by-element multiplication, we get the nth visible light enhanced feature ; Input into the second depth-wise separable convolution module for processing, and output the nth visible light second feature map ; Input into the third depth-separable convolution module for processing to obtain the nth visible light third feature map .
4. The method for aerial target detection by spatiotemporal consistent multimodal feature fusion according to claim 3, characterized in that: The step 2.1.2 includes the following process: Input the first convolution module in the infrared branch for processing and output the nth infrared feature map ; Input into the first C2f module for splitting to obtain the first split feature map and the second split feature map , using multiple bottleneck residual blocks Processing is performed to obtain the second residual feature map ,Will and After splicing, convolution processing is performed to output the nth infrared first feature map ; The timing adaptive module includes: a local branch and a global branch; The local branch uses two layers of 1D time-domain convolutional layers and ReLU nonlinear operations. Processing is performed to obtain the nth infrared local feature ; The global branch uses two fully connected layers and a Softmax normalization function to Processing is performed to obtain the nth infrared global feature ,Will and After splicing, the nth infrared feature is obtained , and then use Gaussian filtering and residual learning to After processing, the nth infrared enhanced image feature is output .
5. The method for aerial target detection by spatiotemporal consistent multimodal feature fusion according to claim 4, characterized in that: The step 2.2 includes the following process: The edge feature extraction module uses the visible light convolution to be learned and infrared convolution Extract separately Edge features and Edge features ; The feature fusion module and as well as and Perform element-by-element subtraction operations to generate the nth visible light difference feature map and the nth infrared difference feature map ; The feature fusion module uses global average pooling to and Processing is performed to generate the nth visible light one-dimensional feature vector and the nth infrared one-dimensional feature vector , and then through the fully connected layer to and Perform nonlinear mapping and obtain the attention weight of the nth visible light channel accordingly And the attention weight of the nth infrared channel ,Will and as well as and Perform channel multiplication operations respectively to generate the nth visible light fusion feature And the nth infrared fusion feature ,thereby and Superimposed as the nth first stage fusion feature .
6. The method for aerial target detection by spatiotemporal consistent multimodal feature fusion according to claim 5, characterized in that: In step 2.2, formula (2) is used to construct : (2) In formula (2), represents the loss of strength consistency, represents the gradient consistency loss, represents the structural similarity loss, , , Represents 3 weight parameters, and has: (3) (4) (5) Formula (3) - Formula (5), Indicates the height of the image. Indicates the width of the image, max takes the maximum value, represents image gradient, and SSIM represents image similarity function.
7. The method for aerial target detection by spatiotemporal consistent multimodal feature fusion according to claim 6, characterized in that: The step 2.7 comprises the following steps: After being processed by the first upsampling module, the nth first upsampling feature map is output ; Use the convolution operation with a convolution kernel of 1×1 to After adjusting the number of channels, Input them together into the first splicing module for processing to obtain the nth first multi-scale fusion feature map , and then input into the first channel space attention module for processing, and output the nth first multi-scale feature map ; After being processed by the second upsampling module, the nth second upsampling feature map is output , using the convolution operation with a convolution kernel of 1×1 After adjusting the number of channels, Input them into the second splicing module for processing to obtain the nth second multi-scale fusion feature map , and then input into the second channel space attention module for processing, and output the nth second multi-scale feature map ; After being processed by the third upsampling module, the nth third upsampling feature map is output , and then input into the third channel space attention module for processing, and output the nth third multi-scale feature map .
8. The method for aerial target detection by spatiotemporal consistent multimodal feature fusion according to claim 7, characterized in that: The first channel space attention module obtains the nth first multi-scale feature map according to the following process : right Perform global average pooling without dimensionality reduction and output the nth channel statistics vector , and then along the spatial dimension Compress and finally generate the nth channel statistics vector ; Use one-dimensional convolution to Process and generate the nth channel weight matrix , and After channel-by-channel multiplication, the nth channel enhancement feature is output ; Will Axis average pooling is performed along the width and height axes respectively, and the nth statistical matrix on the width is generated accordingly. And the nth statistical matrix on height , then and After splicing, convolution is performed to generate the nth spatial weight matrix Finally, and After multiplying position by position, the nth first multi-scale feature map is output .
9. The method for aerial target detection by spatiotemporal consistent multimodal feature fusion according to claim 8, characterized in that: In step 3, formula (7) is used to construct : (7) In formula (7), and are 2 weight parameters; represents the classification loss, represents the regression loss, and has: (8) (9) In formula (8) and formula (9), express The predicted probability of the cth category in ; express The true label of the cth category in , is the Euclidean distance, and are the center point of the target box and the center point of the real box of the image fused with the nth visible light image and the nth infrared image, respectively. c is the diagonal length of the minimum bounding box. α is the weight coefficient. v is the penalty term for the aspect ratio. IOU represents the intersection-over-union ratio of the predicted box and the real box.
10. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the aerial target detection method described in any one of claims 1 to 9, and the processor is configured to execute the program stored in the memory.
Citation Information
Cited By
Target detection method and device based on multi-modal feature fusion and medium
CN120612476A
Multi-mode full-autonomous inspection method and system for electric unmanned aerial vehicle
CN120909339A
Multi-modal full autonomous inspection method and system for power unmanned aerial vehicle
CN120909339B
Multi-modal image fusion method and system, medium and program product
CN121259303A
Power equipment defect identification and alarm method and system based on deep learning
CN121280841A