Multi-modal-based low-slow small target multi-dimensional detection method and system
By combining radar scanning and multi-band photoelectric imaging technologies with deep learning models, the problem of limited radar recognition capabilities for low, slow, and small targets in complex environments has been solved, enabling multi-modal and multi-dimensional target detection and improving detection accuracy and recognition robustness.
Patent Information
- Application Number
- CN202511130992.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-25
AI Technical Summary
In existing technologies, radar's ability to identify low, slow, and small targets in complex environments is limited, and single sensing methods are insufficient to achieve stable, accurate detection and rapid response.
The initial position is determined by radar omnidirectional scanning. Combined with the multi-band imaging technology of the photoelectric turntable, the target image is fused at the feature level through a multi-modal fusion deep learning model, including preprocessing and feature aggregation of infrared and visible light images, to build a multi-dimensional detection system.
In complex backgrounds and electromagnetic interference conditions, it achieves multimodal and multidimensional detection of low, slow, and small targets, improving detection accuracy and recognition robustness, and ensuring rapid and stable target identification and threat assessment.
Smart Images

Figure CN121008264A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of unmanned aerial vehicle (UAV) detection, specifically to a multi-modal multi-dimensional detection method and system for low-speed, small targets. Background Technology
[0002] In recent years, the development of low-altitude small aircraft has been rapid, and low-altitude, slow-moving, small targets (such as small drones, model aircraft, and gliders) are increasingly widely used in various fields. However, these targets typically have characteristics such as small radar cross-section, irregular flight trajectories, and low flight altitudes, posing a serious challenge to existing airspace monitoring and identification systems.
[0003] Currently, radar, as a core means of airspace monitoring, possesses all-weather, all-day detection capabilities, effectively completing initial airspace searches and target localization. However, in environments with complex terrain, strong electromagnetic interference, and dense low-altitude clutter, low-speed, small targets may exhibit weak echo characteristics in radar detection, especially when target features are indistinct or similar to background clutter, making subsequent target identification and classification difficult. Therefore, while radar can effectively detect and initially locate targets, more sophisticated sensing methods are needed for accurate identification and threat confirmation. Furthermore, although traditional photoelectric detection systems possess imaging capabilities, existing image recognition algorithms are unstable under conditions of multi-target interference, motion blur, and complex backgrounds due to limitations such as field of view, weather conditions, and lighting. This makes it difficult to achieve reliable detection and rapid response to low-speed, small targets in real-world scenarios relying solely on a single sensor or sensing method. Summary of the Invention
[0004] The purpose of this invention is to solve the problem mentioned in the background art that "radar has limited ability to identify low, slow and small targets in complex environments, and single sensing methods are difficult to achieve stable, accurate detection and rapid response", and to propose a multi-modal multi-dimensional detection method and system for low, slow and small targets.
[0005] A first aspect of this invention provides a multi-modal, low-speed, small target multi-dimensional detection method, the method comprising:
[0006] The radar is used to perform an all-round scan of the target airspace, and suspicious targets are identified by analyzing the echo signals, and the initial position information of the suspicious targets is obtained.
[0007] Based on the initial position information, the direction of the photoelectric turntable is adjusted, and the suspicious target is captured by multi-band imaging technology to obtain a first target image set; the first target image set includes infrared images and visible light images;
[0008] The first target image set is preprocessed to obtain the second target image set;
[0009] The second set of target images is used as input to a pre-trained target detection model to obtain the target detection results.
[0010] Optionally, the photoelectric turntable is equipped with a visible light sensor and an infrared sensor, and the relative positions of the two sensors are fixed.
[0011] The preprocessing of the first target image set to obtain the second target image set includes:
[0012] The visible light image and the infrared image are registered using a preset transformation matrix;
[0013] The registered images are resized to unify the two types of images to a preset resolution;
[0014] The image pixel values are normalized to obtain the second target image set.
[0015] Optionally, the target detection model is an improvement based on YOLOv8; specific improvements include:
[0016] A dual-branch backbone network is established to process infrared and visible light images separately; and a multimodal fusion module (MFM) is constructed to fuse feature maps from different stages of the dual branches to obtain fused features, which are then input into the neck network.
[0017] The C2f module is lightweighted and improved based on ghost convolution to obtain the C2fGhost module, and the C2f module in the backbone network is replaced by the C2fGhost module.
[0018] A feature aggregation module (FAM) is constructed based on an attention mechanism to interactively aggregate the low-dimensional features of the backbone network and the high-dimensional features of the neck network.
[0019] The path aggregation network structure of the neck network is replaced with a feature pyramid network structure.
[0020] Optionally, the computational expression of the multimodal fusion module (MFM) includes:
[0021]
[0022] Among them, F r It is a feature map of a visible light image; F i This is the feature map of the infrared image; `concat` indicates the stitching operation; `f` is the function operator, the subscript `Conv` indicates the convolution operation, and the superscripts `1×1` and `3×3` indicate the convolution kernel size; `BN` indicates batch normalization; `δ` is the ReLU activation function; OutputMFM It is the output of the multimodal fusion module.
[0023] Optionally, the computational expression of the feature aggregation module FAM includes:
[0024]
[0025] Where CAFM is the predefined attention module; Y is the input to the attention module; f is the function operator, with the subscript Conv indicating convolution operation and the superscript 1×1 indicating the kernel size; BN indicates batch normalization; δ is the ReLU activation function; GAP indicates global average pooling; SC and GC are the feature maps generated during the operation; σ is the Sigmoid activation function; Output MCA The output of the attention module is FB; the low-dimensional feature map output by the backbone network is FB; the high-dimensional feature map output by the higher-level module is FN; Y1 and Y2 are feature maps generated during the computation process; Output FAM This is the output of the feature aggregation module.
[0026] A second aspect of this invention provides a multi-modal, low-speed, small target multi-dimensional detection system, the system comprising:
[0027] The radar detection module is used to perform an all-round scan of the target airspace using radar, identify suspicious targets through echo signal analysis, and obtain the initial position information of the suspicious targets.
[0028] The optoelectronic coordination module is used to adjust the orientation of the optoelectronic turntable according to the initial position information, and to capture the suspicious target through multi-band imaging technology to obtain a first target image set; the first target image set includes infrared images and visible light images;
[0029] The preprocessing module is used to preprocess the first target image set to obtain the second target image set;
[0030] The target recognition module is used to take the second target image set as input to a pre-trained target detection model to obtain the target detection result.
[0031] Optionally, the photoelectric turntable is equipped with a visible light sensor and an infrared sensor, and the relative positions of the two sensors are fixed.
[0032] The preprocessing module includes:
[0033] The registration module is used to register visible light images with infrared images using a preset transformation matrix;
[0034] The size adjustment module is used to adjust the size of the registered images, unifying the two types of images to a preset resolution;
[0035] The normalization module is used to normalize the pixel values of the image to obtain the second target image set.
[0036] Optionally, the target detection model is an improvement based on YOLOv8; specific improvements include:
[0037] A dual-branch backbone network is established to process infrared and visible light images separately; and a multimodal fusion module (MFM) is constructed to fuse feature maps from different stages of the dual branches to obtain fused features, which are then input into the neck network.
[0038] The C2f module is lightweighted and improved based on ghost convolution to obtain the C2fGhost module, and the C2f module in the backbone network is replaced by the C2fGhost module.
[0039] A feature aggregation module (FAM) is constructed based on an attention mechanism to interactively aggregate the low-dimensional features of the backbone network and the high-dimensional features of the neck network.
[0040] The path aggregation network structure of the neck network is replaced with a feature pyramid network structure.
[0041] Optionally, the computational expression of the multimodal fusion module (MFM) includes:
[0042]
[0043] Among them, F r It is a feature map of a visible light image; F i This is the feature map of the infrared image; `concat` indicates the stitching operation; `f` is the function operator, the subscript `Conv` indicates the convolution operation, and the superscripts `1×1` and `3×3` indicate the convolution kernel size; `BN` indicates batch normalization; `δ` is the ReLU activation function; Output MFM It is the output of the multimodal fusion module.
[0044] Optionally, the computational expression of the feature aggregation module FAM includes:
[0045]
[0046] Where CAFM is the predefined attention module; Y is the input to the attention module; f is the function operator, with the subscript Conv indicating convolution operation and the superscript 1×1 indicating the kernel size; BN indicates batch normalization; δ is the ReLU activation function; GAP indicates global average pooling; SC and GC are the feature maps generated during the operation; σ is the Sigmoid activation function; Output MCAThe output of the attention module is FB; the low-dimensional feature map output by the backbone network is FB; the high-dimensional feature map output by the higher-level module is FN; Y1 and Y2 are feature maps generated during the computation process; Output FAM This is the output of the feature aggregation module.
[0047] The beneficial effects of this invention are:
[0048] By fusing radar omnidirectional scanning and multi-band photoelectric imaging technologies, multimodal and multi-dimensional detection of low-speed, small targets is achieved, effectively compensating for the insufficient recognition capabilities of single sensors in complex environments. Deep learning models are used to fuse features from multiple image types, enabling accurate and intelligent target recognition in complex backgrounds for threat assessment. This method significantly improves detection accuracy, environmental adaptability, and recognition robustness, ensuring that the system can still quickly and stably complete target detection even under electromagnetic interference and complex backgrounds, meeting practical needs. Attached Figure Description
[0049] Figure 1 A flowchart illustrating a multi-modal, low-speed, small target multi-dimensional detection method provided in an embodiment of the present invention;
[0050] Figure 2 This is a diagram of the original network architecture of the YOLOv8 model.
[0051] Figure 3 This is a network architecture diagram of a target detection model provided in an embodiment of the present invention;
[0052] Figure 4 This is a schematic diagram of the structure of a multimodal fusion module provided in an embodiment of the present invention;
[0053] Figure 5 This is a schematic diagram of the structure of a feature aggregation module provided in an embodiment of the present invention;
[0054] Figure 6 This is an architecture diagram of a multi-modal, low-speed, small target multi-dimensional detection system provided in an embodiment of the present invention. Detailed Implementation
[0055] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0056] This invention provides a multi-modal, multi-dimensional detection method for small, slow-moving targets. See also... Figure 1 , Figure 1 This is a flowchart illustrating a multi-modal, low-speed, small target multi-dimensional detection method provided in an embodiment of the present invention. The method includes the following steps:
[0057] S101 uses radar to perform an all-around scan of the target airspace, identifies suspicious targets through echo signal analysis, and obtains the initial position information of the suspicious targets.
[0058] S102, based on the initial position information, adjust the direction of the photoelectric turntable, and capture the suspicious target through multi-band imaging technology to obtain the first target image set.
[0059] S103, preprocess the first target image set to obtain the second target image set.
[0060] S104, the second target image set is used as input to the pre-trained target detection model to obtain the target detection result.
[0061] The first target image set includes infrared images and visible light images. The target detection model performs feature-level fusion on the infrared and visible light images.
[0062] This invention provides a multi-modal, multi-dimensional detection method for low-speed, small targets. By fusing omnidirectional radar scanning and multi-band photoelectric imaging technology, it achieves multi-modal and multi-dimensional detection of low-speed, small targets, effectively compensating for the insufficient recognition capability of single sensors in complex environments. A deep learning model is used to fuse features from multiple image types, enabling accurate and intelligent target recognition in complex backgrounds for threat assessment. This method significantly improves detection accuracy, environmental adaptability, and recognition robustness, ensuring that the system can still quickly and stably complete target detection even under electromagnetic interference and complex backgrounds, meeting practical needs.
[0063] One implementation employs a parallel operation mode of radar detection and wireless reconnaissance. The radar detection module utilizes advanced millimeter-wave or phased-array radar technology to perform an all-around scan of the low-altitude airspace, obtaining basic information such as the target's distance, azimuth, and speed through echo signal analysis. The wireless reconnaissance module monitors airborne wireless signals in real time, capturing the communication and navigation wireless signal characteristics of low-altitude, slow-moving, and small unmanned aerial vehicles (UAVs), such as radio frequency signals in specific frequency bands and UAV image transmission protocols. The two detection methods complement each other; radar detection provides target spatial location information, while wireless reconnaissance enables preliminary target identification, constructing a multi-dimensional target detection system and improving the probability of detecting low-altitude, slow-moving, and small targets in complex environments.
[0064] In one implementation, when radar or wireless detection detects a suspected target, it triggers the coordinated operation of the photoelectric tracking module. The photoelectric turntable is equipped with a multi-band sensor, which quickly adjusts its pointing based on the initial target position information provided by the radar, and accurately captures the target through multi-band imaging technology.
[0065] In one implementation, a visible light sensor and an infrared sensor are mounted on the photoelectric turntable, and the relative positions of the two sensors are fixed.
[0066] Step S103: Preprocess the first target image set to obtain the second target image set, including:
[0067] Step 1: Register the visible light image with the infrared image using a preset transformation matrix.
[0068] Step two: Adjust the size of the registered images to unify the two types of images to the preset resolution.
[0069] Step 3: Normalize the image pixel values to obtain the second target image set.
[0070] The fixed relative positions of the visible light and infrared sensors ensure that both types of images have a consistent coordinate reference system when capturing the target, which is beneficial to the accuracy of subsequent image registration and reduces spatial deviation. The transformation matrix is obtained through camera calibration and image registration techniques, using an affine transformation matrix.
[0071] In one embodiment, the object detection model is an improvement upon YOLOv8. See also... Figure 2 , Figure 2 This is a diagram of the original YOLOv8 network architecture. The YOLOv8 model consists of a backbone network, a neck network, and a head. The backbone network is used for feature extraction, the neck network for feature fusion, and the head for object recognition and classification. In the diagram, CBS is a convolutional layer (Conv) + batch normalized layer (BN) + SiLU activation function; C2f is a residual fusion module; SPPF is a fast spatial pyramid pooling module; UpSample represents the upsampling module; Concat is the feature connection module; and Detect is the head. These modules are inherent to the original YOLOv8, and their specific structures will not be elaborated upon here.
[0072] See Figure 3 , Figure 3 This is a network architecture diagram of a target detection model provided in an embodiment of the present invention. Specific improvements include:
[0073] A two-branch backbone network is established, with the two branches processing infrared (IR) images and visible light (RGB) images respectively.
[0074] A multimodal fusion module (MFM) is constructed to fuse feature maps from different stages of the two branches to obtain fused features, and the fused features from different stages are input into the neck network.
[0075] The Feature Aggregation Module (FAM) is constructed based on the attention mechanism to interactively aggregate the low-dimensional features of the backbone network and the high-dimensional features of the neck network.
[0076] The C2f module is lightweighted and improved based on ghost convolution to obtain the C2fGhost module, which replaces the C2f module in the backbone network.
[0077] The path aggregation network structure PANet of the neck network is replaced by the feature pyramid network structure FPN.
[0078] In one implementation, see [link to implementation details]. Figure 4 , Figure 4 This is a schematic diagram of a multimodal fusion module provided in an embodiment of the present invention. The operation process of the multimodal fusion module (MFM) can be expressed as follows:
[0079]
[0080] Among them, F r It is a feature map of a visible light image; F i This is the feature map of the infrared image; `concat` indicates the stitching operation; `f` is the function operator, the subscript `Conv` indicates the convolution operation, and the superscripts `1×1` and `3×3` indicate the convolution kernel size; `BN` indicates batch normalization; `δ` is the ReLU activation function; Output MFM It is the output of the multimodal fusion module.
[0081] Specifically, by concatenating the feature maps of the two modalities, the feature maps of the original visible light and infrared images, both of which were C-layers, are transformed into stacked feature maps of 2C layers. Then, the feature maps are fused and their dimensionality reduced using 1×1, 3×3, and 1×1 convolutional layers to obtain the fused feature map. As shown in the figure, this embodiment has three multimodal fusion modules, and the number of channels after dimensionality reduction can be flexibly set by the technician. In this embodiment, the number of output channels in the last layer of each of the multiple multimodal fusion modules is set to 256.
[0082] In one implementation, see [link to implementation details]. Figure 5 , Figure 5 This is a schematic diagram of the structure of a feature aggregation module provided in an embodiment of the present invention. Figure 5 (a) shows the overall structure of the Feature Aggregation Module (FAM). Figure 5 Figure (b) shows the specific structure of the attention module CAFM. In the figure, Add represents element-wise addition; Mul represents element-wise multiplication. The operation process of the feature aggregation module FAM can be expressed as:
[0083]
[0084] Among them, FB It is the low-dimensional feature map output by the backbone network, that is, the feature map output by the multimodal fusion module (MFE); F N Y1 and Y2 are high-dimensional feature maps output by the upper-level module, i.e., feature maps output by the upsampling module UpSample; Y1 and Y2 are feature maps generated during the computation process; Output FAM is the output of the feature aggregation module; CAFM is the preset attention module; Y is the input of the attention module; f is the function operator, the subscript Conv indicates convolution operation, and the superscript 1×1 indicates the convolution kernel size; BN indicates batch normalization; δ is the ReLU activation function; GAP indicates global average pooling; SC and GC are the feature maps generated during the operation; σ is the Sigmoid activation function; Output MCA This is the output of the attention module. (Symbol) Represents element-wise multiplication; symbol This indicates that the right-hand feature GC is expanded via broadcast and then added element-wise with the left-hand feature SC.
[0085] Specifically, the attention module extracts contextual information through global and local branches and dynamically adjusts feature weights.
[0086] Local branches ( Figure 5 The left branch of (b) uses a 1×1 convolution to perform channel interaction in spatial location. First, a 1×1 convolutional layer compresses the channels to C / r (r is the compression ratio, which can be flexibly set by the technician; in this embodiment, it is set to 2). Then, a 1×1 convolutional layer restores the channels to C, resulting in the local context SC, which has a shape of C×H×W.
[0087] Global branch ( Figure 5 The right branch of (b) extracts global information through global pooling and 1×1 convolution. First, global average pooling is performed on the input features to compress the spatial dimension to 1×1. Then, channel compression and channel restoration are performed through two 1×1 convolutional layers to obtain the context vector with shape C×1×1. Then, the spatial dimension is expanded to H×W through a broadcast mechanism to obtain the global context GC with shape C×H×W.
[0088] Finally, the global and local context elements are added together, and then the attention weights are generated using the Sigmoid activation function.
[0089] The feature aggregation module combines low-dimensional feature maps F by summing elements. B (Low-level texture features) and high-dimensional feature map F N (High-level semantic features) are integrated into Y1. Y1 is enhanced through an attention module, and the interaction between high- and low-level features is enhanced through iterative attention fusion, effectively preserving the fine-grained features of small targets.
[0090] In one implementation, the Feature Pyramid Network (FPN) employs a top-down hierarchical fusion mechanism, fusing high-level semantic features with low-level texture features through upsampling to ensure that shallow features retain sufficient semantic information. Compared to the path aggregation network structure PANet's dual-path approach of top-down + bottom-up, FPN reduces the complexity of feature flow, avoids dilution of small target features, and lowers computational overhead. Specifically, the final bottom-up path in the neck network is removed, and the feature map fused from the top-down path is directly imported into the detection head.
[0091] In one implementation, the standard convolution Conv in the bottleneck structure Bottleneck is replaced with a ghost convolution GhostConv, resulting in a lightweight bottleneck structure Bottleneck_Ghost. Replacing Bottleneck in C2f with Bottleneck_Ghost yields a lightweight C2fGhost module. This type of replacement is a common lightweight improvement in existing technologies and will not be elaborated upon further here.
[0092] To verify that the multimodal detection algorithm of this invention outperforms the single-modal detection algorithm for small, slow-moving targets, a comparative experiment was designed on a multimodal UAV dataset. YOLOv8 was used as the single-modal detection model, with RGB visible light images as input. To verify the effectiveness of each component, the performance of different component combinations was compared. Precision, Recall, mAP@0.5, and mAP@0.5:0.95 were used as evaluation metrics.
[0093] The experiments were conducted on a server equipped with an NVIDIA RTX 4090 GPU (24GB RAM). The operating system was Ubuntu 24.04, and the code used was Python, with dependencies including CUDA 11.6, PyTorch 1.12, and other commonly used deep learning and image processing libraries. The development environment was PyCharm. All models underwent the same number of training steps. The experimental results are shown in Table 1.
[0094] Table 1:
[0095]
[0096] Experimental results fully demonstrate the effectiveness and advancement of the proposed multimodal fusion structure and improved components in low-speed, small-target detection tasks. In particular, after fusing infrared and visible light information, the model exhibits stronger robustness and accuracy in scenarios with complex backgrounds, small targets, and weak features.
[0097] This invention provides a multi-modal, multi-dimensional detection method for small, slow-moving targets. See also... Figure 6 , Figure 6 This is an architecture diagram of a multi-modal, low-speed, small target multi-dimensional detection system provided in an embodiment of the present invention. The system includes:
[0098] The radar detection module is used to perform an all-round scan of the target airspace using radar, identify suspicious targets through echo signal analysis, and obtain the initial position information of the suspicious targets.
[0099] The optoelectronic coordination module is used to adjust the orientation of the optoelectronic turntable based on the initial position information, and to capture the suspicious target through multi-band imaging technology to obtain the first target image set.
[0100] The preprocessing module is used to preprocess the first target image set to obtain the second target image set.
[0101] The target recognition module is used to take the second set of target images as input to the pre-trained target detection model and obtain the target detection results.
[0102] The first target image set includes infrared images and visible light images. The target detection model performs feature-level fusion on the infrared and visible light images.
[0103] This invention provides a multi-modal, multi-dimensional detection system for low-speed, small targets. By fusing omnidirectional radar scanning and multi-band photoelectric imaging technologies, it achieves multi-modal and multi-dimensional detection of low-speed, small targets, effectively compensating for the insufficient recognition capabilities of single sensors in complex environments. Furthermore, it utilizes a deep learning model to fuse features from multiple image types, enabling accurate and intelligent target recognition in complex backgrounds.
[0104] In one embodiment, a visible light sensor and an infrared sensor are mounted on the photoelectric turntable, and the relative positions of the two sensors are fixed. The preprocessing module includes:
[0105] The registration module is used to register visible light images with infrared images using a preset transformation matrix.
[0106] The size adjustment module is used to adjust the size of the registered images, unifying the two types of images to a preset resolution.
[0107] The normalization module is used to normalize the pixel values of the image to obtain the second target image set.
[0108] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A multi-modal, low-speed, small target multi-dimensional detection method, characterized in that, The method includes: The radar is used to perform an all-round scan of the target airspace, and suspicious targets are identified by analyzing the echo signals, and the initial position information of the suspicious targets is obtained. Based on the initial position information, the direction of the photoelectric turntable is adjusted, and the suspicious target is captured by multi-band imaging technology to obtain a first target image set; the first target image set includes infrared images and visible light images; The first target image set is preprocessed to obtain the second target image set; The second set of target images is used as input to a pre-trained target detection model to obtain the target detection results.
2. The method for multi-modal, low-speed, small target multi-dimensional detection according to claim 1, characterized in that, The photoelectric turntable is equipped with a visible light sensor and an infrared sensor, and the relative positions of the two sensors are fixed. The preprocessing of the first target image set to obtain the second target image set includes: The visible light image and the infrared image are registered using a preset transformation matrix; The registered images are resized to unify the two types of images to a preset resolution; The image pixel values are normalized to obtain the second target image set.
3. The method for multi-modal, low-speed, small target multi-dimensional detection according to claim 1, characterized in that, The target detection model is an improvement based on YOLOv8; the specific improvements include: A dual-branch backbone network is established to process infrared and visible light images separately; and a multimodal fusion module (MFM) is constructed to fuse feature maps from different stages of the dual branches to obtain fused features, which are then input into the neck network. The C2f module is lightweighted and improved based on ghost convolution to obtain the C2fGhost module, and the C2f module in the backbone network is replaced by the C2fGhost module. A feature aggregation module (FAM) is constructed based on an attention mechanism to interactively aggregate the low-dimensional features of the backbone network and the high-dimensional features of the neck network. The path aggregation network structure of the neck network is replaced with a feature pyramid network structure.
4. The method for multi-modal, low-speed, small target multi-dimensional detection according to claim 3, characterized in that, The computational expressions of the multimodal fusion module (MFM) include: Among them, F r It is a feature map of a visible light image; F i This is the feature map of the infrared image; `concat` indicates the stitching operation; `f` is the function operator, the subscript `Conv` indicates the convolution operation, and the superscripts `1×1` and `3×3` indicate the convolution kernel size; `BN` indicates batch normalization; `δ` is the ReLU activation function; `Output`... MFM It is the output of the multimodal fusion module.
5. The method for multi-modal, low-speed, small target multi-dimensional detection according to claim 3, characterized in that, The computational expressions of the feature aggregation module FAM include: Where CAFM is the predefined attention module; Y is the input to the attention module; f is the function operator, with the subscript Conv indicating convolution operation and the superscript 1×1 indicating the kernel size; BN indicates batch normalization; δ is the ReLU activation function; GAP indicates global average pooling; SC and GC are the feature maps generated during the operation; σ is the Sigmoid activation function; Output MCA The output of the attention module is FB; the low-dimensional feature map output by the backbone network is FB; the high-dimensional feature map output by the higher-level module is FN; Y1 and Y2 are feature maps generated during the computation process; Output FAM This is the output of the feature aggregation module.
6. A multi-modal, low-speed, small target multi-dimensional detection system, characterized in that, The system includes: The radar detection module is used to perform an all-round scan of the target airspace using radar, identify suspicious targets through echo signal analysis, and obtain the initial position information of the suspicious targets. The optoelectronic coordination module is used to adjust the orientation of the optoelectronic turntable according to the initial position information, and to capture the suspicious target through multi-band imaging technology to obtain a first target image set; the first target image set includes infrared images and visible light images; The preprocessing module is used to preprocess the first target image set to obtain the second target image set; The target recognition module is used to take the second target image set as input to a pre-trained target detection model to obtain the target detection result.
7. A multi-modal, low-speed, small target multi-dimensional detection system according to claim 6, characterized in that, The photoelectric turntable is equipped with a visible light sensor and an infrared sensor, and the relative positions of the two sensors are fixed. The preprocessing module includes: The registration module is used to register visible light images with infrared images using a preset transformation matrix; The size adjustment module is used to adjust the size of the registered images, unifying the two types of images to a preset resolution; The normalization module is used to normalize the pixel values of the image to obtain the second target image set.
8. A multi-modal, low-speed, small target multi-dimensional detection system according to claim 6, characterized in that, The target detection model is an improvement based on YOLOv8; the specific improvements include: A dual-branch backbone network is established to process infrared and visible light images separately; and a multimodal fusion module (MFM) is constructed to fuse feature maps from different stages of the dual branches to obtain fused features, which are then input into the neck network. The C2f module is lightweighted and improved based on ghost convolution to obtain the C2fGhost module, and the C2f module in the backbone network is replaced by the C2fGhost module. A feature aggregation module (FAM) is constructed based on an attention mechanism to interactively aggregate the low-dimensional features of the backbone network and the high-dimensional features of the neck network. The path aggregation network structure of the neck network is replaced with a feature pyramid network structure.
9. A multi-modal, low-speed, small target multi-dimensional detection system according to claim 8, characterized in that, The computational expressions of the multimodal fusion module (MFM) include: Among them, F r It is a feature map of a visible light image; F i This is the feature map of the infrared image; `concat` indicates the stitching operation; `f` is the function operator, the subscript `Conv` indicates the convolution operation, and the superscripts `1×1` and `3×3` indicate the convolution kernel size; `BN` indicates batch normalization; `δ` is the ReLU activation function; Output MFM It is the output of the multimodal fusion module.
10. A multi-modal, low-speed, small target multi-dimensional detection system according to claim 8, characterized in that, The computational expressions of the feature aggregation module FAM include: Where CAFM is the predefined attention module; Y is the input to the attention module; f is the function operator, with the subscript Conv indicating convolution operation and the superscript 1×1 indicating the kernel size; BN indicates batch normalization; δ is the ReLU activation function; GAP indicates global average pooling; SC and GC are the feature maps generated during the operation; σ is the Sigmoid activation function; Output MCA The output of the attention module is FB; the low-dimensional feature map output by the backbone network is FB; the high-dimensional feature map output by the higher-level module is FN; Y1 and Y2 are feature maps generated during the computation process; Output FAM This is the output of the feature aggregation module.
Citation Information
Cited By
Multi-band image end-to-end multi-task intelligent detection method
CN121259306A
Automobile part detection method and system based on machine vision
CN121298745A