Multi-modal Detection Method for Air-to-Ground Small Targets Based on Adaptive Modal Selection Strategy

By adopting a multimodal detection method with an adaptive modal selection strategy in the air-to-ground small object detection technology, the problems of small target size, modal spatial imbalance and information redundancy are solved, and the precise detection and robustness of small targets are achieved.

CN119229329BActive Publication Date: 2025-06-17ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411749968.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-06-17
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

The existing small-space target detection technology faces the problems of small target sizes, modal spatial imbalance and modal information redundant, especially in insufficient lighting or severe weather conditions.

Method used

Using the air-to-ground multimodal detection method based on the adaptive modal selection strategy, the initial framework model of the deep learning convolutional network, including a dual-channel convolutional neural network, a policy module, a fusion module and a detection head, feature extraction, screening and fusion of visible light and infrared images is carried out to remove invalid information and enhance the target detection performance.

Benefits of technology

It realizes accurate detection of long-distance micro-targets, has strong robustness, and can stably detect micro-targets under poor lighting conditions, overcomes the problems of modal spatial imbalance and information redundancy, and improves detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229329B_ABST
    Figure CN119229329B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal detection method for air-to-ground small targets based on an adaptive modal selection strategy, comprising the following steps: S1. Construct an initial framework model of a deep learning convolutional network, the initial framework model of the convolutional network comprising a dual-channel convolutional neural network, a strategy module, a fusion module, and a detection head arranged in sequence; S2. Use the training set in a drone crowd detection dataset based on visible light-infrared image pairs to train the initial framework model of the convolutional network to obtain a target detection model; S3. Obtain an image to be detected and input it into the target detection model, perform feature extraction, screening, and fusion on the image to be detected through the target detection model, and finally output the target position result. The present invention solves the problems of small scale of the target to be detected, unbalanced modal space, and redundant modal information existing in the existing air-to-ground small target detection technology, and can resist the influence of interference factors such as poor lighting conditions, partial occlusion, and overlap.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection, and relates to a method and device for multi-modal detection of small air-to-ground targets, in particular to a method for multi-modal detection of small air-to-ground targets based on an adaptive modal selection strategy. Background Art

[0002] Target detection is a basic technology in various computer vision applications such as monitoring, search and rescue, and fault monitoring. In recent years, with the development of neural network and artificial intelligence technologies, target detection technologies based on convolutional neural networks (CNNs) have made great progress. Currently, the mainstream target detectors can be divided into two categories: two-stage detectors and one-stage detectors. Generally speaking, two-stage detectors have higher detection accuracy but slower execution speed; while one-stage detectors have faster execution speed but slightly lower detection accuracy. However, these detectors are generally based on single-modal visible light images, so their detection performance will drop significantly under conditions such as insufficient light or bad weather (such as rain, fog, haze).

[0003] Therefore, target detectors that combine infrared and visible light multi-modal information have gradually received the favor of researchers as a feasible and effective improvement scheme. Experimental results show that fusing the information of visible light and infrared multi-modal can usually improve the target detection performance and can better resist the influence of adverse factors such as insufficient light and bad weather. However, the visible light and infrared multi-modal target detection technology for air-to-ground small targets is still in the exploratory stage and currently mainly faces the following three major challenges:

[0004] 1. The scale of the target to be detected is small. Since the images used in air-to-ground small target detection are all collected by drones, and when drones collect images, the height is high and the field of view is wide, resulting in the photographed target objects usually looking smaller, containing less details, and being more affected by the cluttered background. Therefore, it is quite challenging to overcome various interference factors in the air-to-ground small target detection task and accurately detect tiny targets with limited information.

[0005] 2. Modal space imbalance. The information contained in visible light images and infrared images has significantly different characteristics and can provide complementary clues for the target detection task. For example, visible light sensors can capture images with rich details, color information, and textures under sufficient light conditions. In contrast, infrared detectors can capture the temperature difference between the target and its surrounding environment and show strong performance under insufficient light conditions. Therefore, in different scenarios, the importance of the visible light modality and the infrared modality for the target detection task is not the same. How to effectively balance or fuse multi-modal features is still a difficult problem to be solved.

[0006] 3. Modal information redundancy. In a typical air-to-ground small target detection scenario, visible and infrared images captured by an unmanned aerial vehicle (UAV) contain a large amount of invalid information from the cluttered background. To improve the accuracy of detecting small targets, it is reasonable to remove this invalid input. In addition, in the visible and infrared images captured using a UAV platform, small targets usually look very similar. Therefore, it is feasible to reduce redundant information in multi-modal data without degrading the overall detection performance. However, how to design an effective scheme to remove invalid input and reduce redundant information remains to be explored. Summary of the Invention

[0007] In view of the above technical problems, the purpose of the present invention is to provide an air-to-ground small target multi-modal detection method based on an adaptive modal selection strategy, which selects effective modal information for fusion through an adaptive modal selection mechanism to improve the accuracy and robustness of small target detection.

[0008] The technical solution adopted by the present invention is as follows:

[0009] An air-to-ground small target multi-modal detection method based on an adaptive modal selection strategy, comprising the following steps:

[0010] S1. Construct an initial framework model of a deep learning convolutional network, where the initial framework model of the convolutional network includes a dual-path convolutional neural network, a strategy module, a fusion module, and a detection head arranged in sequence;

[0011] S2. Use the training set in the UAV crowd detection dataset (RGBTDronePerson) based on visible-infrared image pairs to train the initial framework model of the convolutional network to obtain a target detection model;

[0012] S3. Obtain the image to be detected and input it into the target detection model. The target detection model performs feature extraction, screening, and fusion on the image to be detected, and finally outputs the target position result.

[0013] Further, the dual-path convolutional neural network processes visible light images and infrared images respectively, extracts feature maps of the two modalities, and obtains dual-path feature maps.

[0014] Further, the strategy module is used to perform modal screening on the dual-path feature maps extracted by the dual-path convolutional neural network, and remove invalid or highly interfering modal information.

[0015] Further, the fusion module is used to integrate the complementary information of the dual-path feature maps screened by the strategy module.

[0016] Furthermore, the parameters of the dual-path convolutional neural network convolutional layer are all initialized with the weights and biases of the pre-trained Residual Network model (ResNet50) on the large-scale image recognition dataset (ImageNet), and the parameters of other convolutional layers are randomly initialized using the Gaussian normal distribution.

[0017] Furthermore, in step S2, when training the initial framework model of the convolutional network, a strategy of generating small batches of data in sequence is adopted for data input; the gradient descent method is used for parameter update, and the model parameters are adjusted by means of gradient clipping.

[0018] Furthermore, the fusion module identifies the key regions in the feature map through the spatial attention mechanism and reduces the influence of irrelevant regions.

[0019] An air-to-ground small target multi-modal detection system based on an adaptive modality selection strategy, comprising:

[0020] A model construction module: used to construct an initial framework model of a deep learning convolutional network, and the initial framework model of the convolutional network includes a dual-path convolutional neural network, a strategy module, a fusion module, and a detection head arranged in sequence;

[0021] A model training module: used to train the initial framework model of the convolutional network with the training set in the unmanned aerial vehicle crowd detection dataset based on visible light-infrared image pairs to obtain a target detection model;

[0022] An image detection module: used to obtain the image to be detected and input it into the target detection model, perform feature extraction, screening, and fusion on the image to be detected through the target detection model, and finally output the target position result.

[0023] A computer device, the computer device comprising:

[0024] One or more processors;

[0025] A memory for storing one or more programs;

[0026] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned air-to-ground small target multi-modal detection method based on the adaptive modality selection strategy.

[0027] A computer-readable storage medium storing computer instructions, when the computer instructions are executed by one or more processors, causing the one or more processors to execute the steps in the above method.

[0028] The beneficial effects of the present invention are:

[0029] 1. The present invention can make full use of the complementary information of the visible light and infrared modalities to achieve precise detection of distant small targets, and can still stably detect small targets under poor lighting conditions such as cloudy days and nights, with strong robustness.

[0030] 2. By introducing a strategy module, the present invention uses an adaptive modality screening mechanism to remove invalid or interference-rich modality information, greatly overcoming the problems of modality space imbalance and modality information redundancy, and improving the detection performance.

[0031] 3. By designing a specific fusion module, the present invention fuses the feature maps of the visible light and infrared modalities, highlights the target area and suppresses interference factors such as the background, and at the same time uses the spatial attention mechanism to enhance the detection performance for small targets.

[0032] 4. The present invention not only has excellent detection performance, but also has a relatively simple model, which is suitable for deployment on devices with limited computing power such as drones, and can meet the application scenarios of air-to-ground small target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a flowchart of the algorithm steps of an embodiment of the present invention.

[0034] Figure 2 is a schematic diagram of the overall structure of the convolutional neural network model adopted in an embodiment of the present invention.

[0035] Figure 3 is a schematic diagram of the structure of the strategy module in the convolutional neural network model of an embodiment of the present invention.

[0036] Figure 4 is a schematic diagram of the structure of the fusion module in the convolutional neural network model of an embodiment of the present invention.

[0037] Figure 5 is the target detection result (infrared channel display) in the residential area scene of the UAV crowd detection dataset based on visible light-infrared image pairs in an embodiment of the present invention.

[0038] Figure 6 is the target detection result (infrared channel display) in the urban road scene of the UAV crowd detection dataset based on visible light-infrared image pairs in an embodiment of the present invention.

[0039] Figure 7 is the target detection result (infrared channel display) in the park scene of the UAV crowd detection dataset based on visible light-infrared image pairs in an embodiment of the present invention.

[0040] Figure 8This is the target detection result (infrared channel display) of an embodiment of the present invention in a school playground scene in a drone crowd detection dataset based on visible light-infrared image pairs. DETAILED DESCRIPTION

[0041] The technical solution of the present invention will be further explained in detail below in conjunction with the accompanying drawings and specific embodiments.

[0042] The present invention provides an air-to-ground small target multi-modal detection method based on an adaptive modal selection strategy, comprising the following steps:

[0043] S1. Construct an initial framework model of a deep learning convolutional network, which includes a dual-channel convolutional neural network, a strategy module, a fusion module and a detection head arranged in sequence. The dual-channel convolutional neural network processes visible light images and infrared images respectively, performs dual-channel feature extraction, outputs feature maps corresponding to the two modalities, and obtains a dual-channel feature map. The dual-channel feature map is input into the strategy module for modality screening to remove invalid or interfering modal information. The screened dual-channel feature map is input into the fusion module for complementary information integration to obtain the fused feature map, and finally the fused feature map is input into the subsequent detection head to output the target position result finally predicted by the target detection model.

[0044] S2. The training set in the drone crowd detection dataset based on visible light-infrared image pairs is used to train the initial framework model of the convolutional network to obtain the target detection model; specifically, the training set in the drone crowd detection dataset based on visible light-infrared image pairs is used as training data to input into the initial framework model of the convolutional network for training, and the detection labels in the dataset are used as supervision information for supervision, so as to obtain the target detection model. During training, the convolutional layer parameters of the two-way convolutional neural network in the visible light image and infrared image feature extraction channels are initialized with the weights and biases of the residual network model (ResNet50) pre-trained on the large-scale image recognition dataset (ImageNet), while all other convolutional layers are randomly initialized using Gaussian normal distribution. When the training data of the visible light image and infrared image in the dataset are input into the convolutional neural network for training, the strategy of sequentially generating small batches of data is adopted, and the batch size is 2. The initial framework model of the convolutional network uses the gradient descent method to update parameters during the training process, and needs to be trained for several rounds. During training, the learning rate will gradually decay with the increase of training rounds, and the model parameters are adjusted by gradient clipping.

[0045] S3. Obtain the image to be detected and input it into the target detection model. The target detection model performs feature extraction, screening, and fusion on the image to be detected, and finally outputs the target position result. Among them, the input image to be detected needs to contain both visible light images and infrared images at the same time.

[0046] Embodiment

[0047] See Figure 1 , the specific implementation of the air-to-ground small target multi-modal detection method based on the adaptive modality selection strategy in the embodiment of the present invention includes the following steps:

[0048] Step 1, construct an initial framework model of a deep learning convolutional network.

[0049] Among them, for the structure of the initial framework model of the convolutional network, see Figure 2 , the initial architecture of the convolutional network includes a dual-path convolutional neural network (C1-C5), a policy module, a fusion module, and a detection head arranged in sequence. The dual-path convolutional neural network processes visible light images and infrared images respectively, extracts dual-channel features, and then inputs them into the policy module, the fusion module, and the detection head in sequence to obtain the output result.

[0050] The policy module is used to perform modality screening operations on the feature maps of the two modalities. For the structure of the policy module, see Figure 3 , and its working process is as follows:

[0051] First, input the dual-path feature maps into the adaptive average pooling layer, cascade the outputs, and then flatten them to obtain a one-dimensional vector :

[0052]

[0053] Among them and represent the feature maps corresponding to the input visible light and infrared modalities respectively; and represent the adaptive average pooling layers passed by the dual-path feature maps respectively; represents the cascading operation; represents flattening the result into a one-dimensional vector.

[0054] After that, input the vector into a series of fully connected layers to perform classification operations, and obtain discrete classification results:

[0055]

[0056]

[0057] Among them Corresponding to Figure 3 a series of fully connected layers in ; And respectively correspond to the classification results of the two modalities by the fully connected layer; Indicates that And The result after concatenation.

[0058] Since the classification results output by the fully connected layer are discrete, the model will be non-differentiable in subsequent training steps and the model parameters cannot be updated through backpropagation. Therefore, the Gumbel-Softmax distribution strategy is introduced in this module to approximate the discrete classification results with a continuous distribution, and the process can be expressed as:

[0059]

[0060] Where Indicates using the Gumbel normalization distribution to approximate the discrete result ; And respectively correspond to the strategy outputs of the visible light and infrared modalities; Indicates that the results are combined and output together.

[0061] Finally, multiply the feature maps ( And ) input by the two modalities with their corresponding strategy outputs ( And ) to obtain the final output feature maps ( And ):

[0062]

[0063] The fusion module is used to integrate the complementary information of the filtered dual-channel feature maps. The structure of the fusion module is shown in Figure 4 , which consists of an attention module and a multi-convolution module (C3 Module). The working process of the fusion module is as follows:

[0064] First, cascade the dual-channel feature maps of the visible light and infrared modalities output by the strategy module ( And ) and input them into the attention module, and the process can be expressed as:

[0065]

[0066]

[0067] in Indicates that the dual-path feature map and The result after cascading; and Represent the maximum pooling operation and the average pooling operation respectively; Represents a two-dimensional convolution operation; represents the activation function; Represents a multiplication operation; Represents the feature map output after the attention module.

[0068] Then After further feature extraction is performed in the input multi-convolution module, the feature map output by the fusion module is obtained. , the process can be expressed as:

[0069]

[0070] in Represents addition operation; express Figure 4 The operations corresponding to each normalized convolution (CBS) module in, ; Represents a joint operation consisting of multiple normalized convolution modules.

[0071] Step 2: Use the training set in the drone crowd detection dataset based on visible light-infrared image pairs as training data and input it into the convolutional network initial framework model for training, and use the detection labels in the dataset as supervision information for supervision, so as to obtain the target detection model.

[0072] The dataset used in the embodiment of the present invention is a drone crowd detection dataset based on visible light-infrared image pairs. Its training set consists of 4900 pairs of visible light and infrared image pairs, including 43006 pedestrian objects, 4869 rider objects and 7316 crowd objects.

[0073] When the training data of visible light images and infrared images in the dataset are input into the initial framework model of the convolutional network for training, the following specific operations are included but not limited to:

[0074] 1. During training, the input of visible light images and infrared images uses a method of sequentially generating small batches of data, with a batch size of 2.

[0075] 2. During training, the parameters of the convolutional layers in the dual-channel feature extraction channels for visible light and infrared images are initialized with the weights and biases of a residual network model pre-trained on a large-scale image recognition dataset, while all other convolutional layers are randomly initialized using a Gaussian normal distribution.

[0076] 3. The initial framework model of the convolutional network updates its parameters using the gradient descent method during training and requires several training rounds. The learning rate also gradually decays as the number of training rounds increases, and the model parameters are adjusted using gradient clipping.

[0077] The overall loss function used when training the model can be expressed as:

[0078]

[0079] where represents the classification loss for determining the category to which the target belongs; represents the regression loss for determining the location and size of the target; represents the policy loss for supervising the policy module; represents the proportionality factor that balances the weights of the policy loss and the other two losses.

[0080] In the embodiments of the present invention, the selected classification loss is the cross entropy loss, which is expressed as:

[0081]

[0082] where represents the true value of the class distribution probability, represents the predicted value of the class distribution probability.

[0083] In the embodiments of the present invention, the selected regression loss is the smooth L1 loss, which is expressed as:

[0084]

[0085] where represents the difference between the predicted box and the true box.

[0086] In the embodiments of the present invention, a policy loss is specifically designed to embed the policy module into the overall model for end-to-end joint training. The policy loss can be expressed as:

[0087]

[0088] where represents the total number of modalities, which is set to 2 in the formula; Corresponding to visible light and infrared modes; Indicates the output of the strategy module corresponding to the two modes; Represents the weight factors of the visible light mode and the infrared mode respectively. Since the visible light mode and the infrared mode are considered to be equivalent, the weight factors corresponding to the two modes are set to 1; Indicates the detection accuracy of the model; Represents the proportional factor between the first and second terms of the equilibrium strategy loss.

[0089] Step 3: Obtain the image to be detected and input it into the target detection model, extract, filter and fuse the features of the image to be detected through the target detection model, and finally output the target position result. The input image to be detected needs to include both visible light image and infrared image.

[0090] Step 3.1: Get the image to be detected and input it into the trained target detection model.

[0091] Since the deep learning convolutional network initial framework model constructed in the embodiment of the present invention receives both visible light images and infrared images as input, the image to be detected needs to contain images of both visible light and infrared modalities.

[0092] In step 3.2, the target detection model performs dual-channel feature extraction and outputs feature maps corresponding to the two modalities.

[0093] See also Figure 2 , the dual-channel convolutional neural network (C2 to C5) processes the input visible light image and infrared image respectively, and outputs the feature maps corresponding to the two modalities, the purpose of which is to integrate the features in the visible light image and the infrared image. In the process of extracting features from the two modal images, if no effective features are extracted from the image of one modality (for example, due to insufficient light in the night environment, the visibility of the visible light image taken is low and the target information is unclear), but the image of the other modality happens to provide the corresponding complementary features (because the temperature of the target to be detected at night is higher than the ambient temperature, the target features in the infrared image are more obvious), by integrating the features extracted from the two modalities, more comprehensive and complete target feature information can be obtained.

[0094] Step 3.3: Input the extracted feature maps of the visible light and infrared modes into the strategy module for mode screening.

[0095] For the feature maps extracted by the dual-path convolutional neural network, the strategy module proposed in the embodiments of the present invention adaptively selects the effective modal features among them for retention and discards the ineffective modal features. For example, if the feature map corresponding to the input visible light modality does not contain clear and effective target information, while the feature map corresponding to the infrared modality contains clear target information, then after the input dual-path feature map passes through the strategy module, the feature information corresponding to the visible light modality will be discarded, and the feature information corresponding to the infrared modality will be retained. It should be noted that the strategy module does not always select the feature map of only one modality each time, and its output may also be all retained or all discarded.

[0096] Step 3.4: Input the filtered feature maps into the fusion module to obtain the fused feature maps.

[0097] The fusion module adopted in the embodiments of the present invention identifies the key regions in the feature maps through the spatial attention mechanism and reduces the influence of irrelevant regions, thereby effectively highlighting the features of small targets and suppressing background interference, and improving the detection performance of small targets. In addition, the fusion module extracts target features of different scales by deploying multiple groups of convolutional layers, achieving better robustness and generalization ability.

[0098] Step 3.5: Input the fused feature maps into the subsequent detection head to output the final predicted target position results of the target detection model.

[0099] Figures 5 to 8 Part of the detection results of the architecture of the present invention on the test set of the UAV crowd detection dataset based on visible light-infrared image pairs are respectively shown, which include different lighting conditions (such as day and night) and different detection scenarios. The green boxes in the figure are true positive detection results, the red boxes are false positive detection results, and the blue boxes are false negative detection results. It is not difficult to find from the results that the algorithm proposed by the architecture of the present invention has high detection accuracy, fewer false detections and missed detections, and has excellent detection performance for multi-scale targets (especially small targets). For example, in Figure 8 the architecture of the present invention detected the vast majority of small and dense targets, indicating its strong ability to identify small targets at a distance and partially occluded targets, and high robustness to background interference.

[0100] The embodiments of the present invention solve the three problems of small scale of targets to be detected, modal space imbalance, and modal information redundancy existing in the existing air-to-ground small target detection technology by constructing an air-to-ground small target multi-modal detection method based on an adaptive modal selection strategy, and can resist the influence of interference factors such as poor lighting conditions, partial occlusion, and overlap. The architecture of the present invention can also be used for other multi-sensor-based visual analysis tasks, and promote high-quality path planning and target tracking for work such as UAV-based surveillance, search, and rescue.

[0101] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0102] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0103] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0104] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0105] After considering the specification and practicing the content disclosed herein, those skilled in the art will readily think of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the claims.

[0106] The above specific embodiments are used to explain the present invention rather than limit the present invention. Any modification and change made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.

Claims

1. A multi-modal detection method for small air-to-ground targets based on an adaptive modal selection strategy, characterized in that: The following steps are involved: S1. constructing an initial framework model of a deep learning convolutional network, wherein the initial framework model of the convolutional network includes a dual-path convolutional neural network, a strategy module, a fusion module and a detection head arranged in sequence; S2. Use the training set in the drone crowd detection dataset based on visible light-infrared image pairs to train the initial framework model of the convolutional network to obtain the target detection model; S3. Obtain the image to be detected and input it into the target detection model. The target detection model extracts, filters and fuses the features of the image to be detected, and finally outputs the target position result. Specifically: The dual-path convolutional neural network processes the visible light image and the infrared image respectively, extracts feature maps of the two modalities, and obtains a dual-path feature map; The strategy module is used to perform modal screening on the dual-path feature graph extracted by the dual-path convolutional neural network, adaptively select effective modal features to be retained, and remove invalid or more interfering modal information; the workflow of the strategy module is as follows: first, the dual-path feature graph is input into the adaptive average pooling layer and the output is cascaded and flattened to obtain a one-dimensional vector : , in and Represents the feature maps corresponding to the input visible light and infrared modes respectively; and Indicates the adaptive average pooling layer that each of the two feature maps passes through; Indicates a cascade operation; Indicates flattening the result into a one-dimensional vector; Then the vector Input a series of fully connected layers to perform classification operations and obtain discrete classification results: , , in Corresponding to a series of fully connected layers, ; and They correspond to the classification results of the fully connected layer for the two modalities; Indicates that and The result after cascading; Then, the Gumbel normalized distribution strategy is introduced in the strategy module to approximate the discrete classification results with continuous distribution. The process can be expressed as: , in Represents the use of Gumbel normalized distribution for discrete results Make an approximation; and The strategy outputs corresponding to visible light and infrared modalities respectively; Indicates that the results are combined and output together; Finally, the feature maps of the two modal inputs are and The corresponding policy output and Multiply them together to get the feature map of the final output of the strategy module. and : , The fusion module is composed of an attention module and a multi-convolution module, and is used to integrate the complementary information of the dual-path feature map after being filtered by the strategy module; the fusion module identifies the key areas in the feature map through the spatial attention mechanism, reduces the influence of irrelevant areas, highlights the features of small targets and suppresses background interference, and also extracts target features of different scales by deploying multiple groups of convolution layers; the workflow of the fusion module is as follows: first, the dual-path feature map of visible light and infrared modalities output by the strategy module is and After cascading, it is input into the attention module, and the process can be expressed as: , , in Indicates that the dual-path feature map and The result after cascading; and Represent the maximum pooling operation and the average pooling operation respectively; Represents a two-dimensional convolution operation; represents the activation function; Represents a multiplication operation; Represents the feature map output after the attention module; Then After further feature extraction is performed in the input multi-convolution module, the feature map output by the fusion module is obtained. , the process can be expressed as: , in Represents addition operation; Indicates the operations corresponding to each normalized convolution module in Figure 4, ; Represents a joint operation consisting of multiple normalized convolution modules.

2. The multi-modal detection method for small air-to-ground targets based on an adaptive modality selection strategy according to claim 1 is characterized in that: The convolutional layer parameters of the dual-path convolutional neural network are initialized using the weights and biases of a residual network model pre-trained on a large-scale image recognition dataset, and other convolutional layer parameters are randomly initialized using a Gaussian normal distribution.

3. The multi-modal detection method for small air-to-ground targets based on an adaptive modality selection strategy according to claim 1 is characterized in that: When the initial framework model of the convolutional network is trained, a strategy of sequentially generating small batches of data is adopted for data input; a gradient descent method is used for parameter update, and gradient clipping is used to adjust the model parameters.

4. A computer device, characterized in that: The computer device comprises: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-modal detection method for small air-to-ground targets based on an adaptive modal selection strategy as described in any one of claims 1-3.

5. A computer-readable storage medium storing computer instructions, characterized in that: When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • A deep neural network architecture of bounding box segmentation supervision for accurate real-time pedestrian detection of visible light and infrared images

    CN111209810A

  • Multi-mode intelligent customer service dialogue method and system

    CN118586498A