Infrared small target detection method based on multi-modal feature fusion
Through the convolutional neural network with multimodal feature fusion and multiple attention mechanisms, the detection problem of infrared dim targets in complex backgrounds is solved, high-precision real-time detection is achieved, and detection accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510760817.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Small infrared targets are difficult to distinguish against complex backgrounds, and are affected by noise, resulting in a low signal-to-noise ratio and poor detection accuracy. Existing methods make it difficult to achieve high-precision real-time detection.
A multimodal feature fusion method is adopted, including infrared image preprocessing, multi-dimensional feature fusion network, cross-layer multi-attention mechanism convolutional neural network, combined with non-local mean filtering and adaptive contrast enhancement algorithm to suppress background interference, extract multi-scale features and perform target detection.
The detection accuracy and real-time performance of infrared dim targets in complex backgrounds are improved, the robustness is enhanced, and the false alarm rate is reduced.
Smart Images

Figure CN120635648A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and in particular relates to an infrared small target detection method based on multimodal feature fusion. Background Art
[0002] Due to their small size and indistinct features, small infrared targets are easily lost in complex and changing background clutter and are also subject to noise interference. Therefore, detecting small infrared targets in complex backgrounds is a hot topic and a challenge in the field of target detection. Common methods for small target detection primarily target visible light images, with relatively few focusing on infrared images. Small infrared targets lack color information, differ significantly in scale from conventional targets, and rely more heavily on contextual information, making them difficult to effectively identify. Despite some progress, many challenges remain. Therefore, improving the detection capability of small infrared targets in complex backgrounds, reducing false alarm rates, and achieving real-time, high-precision detection of small infrared targets is of great academic research significance and engineering application value. Therefore, constructing a lightweight infrared target detection framework that effectively integrates multi-dimensional features, extracts multi-scale information, and suppresses background, while balancing detection accuracy and real-time performance, has become a key research issue. Summary of the Invention
[0003] In order to solve the above technical problems, the present invention proposes an infrared small target detection method based on multimodal feature fusion. Aiming at the problems that infrared weak small targets are difficult to distinguish from the background under complex background interference, low signal-to-noise ratio caused by noise interference, and poor detection accuracy, the present invention conducts research from the aspects of feature extraction capability, contrast difference and spatiotemporal feature extraction to make up for the shortcomings of single modal feature description, improve the detection performance of infrared weak small targets under complex background, and enhance the robustness of small target detection algorithm in different scenarios.
[0004] To achieve the above objectives, the present invention provides an infrared small target detection method based on multimodal feature fusion, comprising:
[0005] Collect infrared images;
[0006] Preprocessing the infrared image to obtain a preprocessed infrared image;
[0007] fusing the preprocessed infrared images using a multi-dimensional feature fusion network to obtain a fused image;
[0008] Performing multi-scale feature extraction on the fused image to obtain a small target multi-dimensional feature map;
[0009] The small target multidimensional feature map is input into the target detection module to obtain the detection result, wherein the target detection model is constructed by a convolutional neural network that integrates a cross-layer multi-attention mechanism.
[0010] Optionally, the infrared image set covers different weather conditions, illumination changes, large-area high-brightness cloud interference, ground building interference, and infrared small targets of various sizes.
[0011] Optionally, preprocessing the infrared image to obtain the preprocessed infrared image includes:
[0012] The infrared image is denoised using a non-local mean filtering algorithm to obtain a denoised infrared image.
[0013] According to the pixel distribution characteristics of the denoised infrared image, an adaptive contrast enhancement algorithm is used to perform local stretching processing on the low-contrast area to obtain a preprocessed infrared image.
[0014] Optionally, a non-local mean filtering algorithm is used to denoise the infrared image. The denoised infrared image is obtained by:
[0015] The infrared image is denoised using a non-local mean filtering algorithm to obtain a first image with reduced noise interference;
[0016] If the noise level of the first image is higher than a preset threshold, a secondary denoising process is performed on the first image using a median filtering algorithm to obtain a denoised infrared image.
[0017] Optionally, according to the pixel distribution characteristics of the denoised infrared image, an adaptive contrast enhancement algorithm is used to perform local stretching processing on low-contrast areas, and obtaining the preprocessed infrared image includes:
[0018] By analyzing the pixel distribution characteristics of the denoised infrared image, the histogram statistical method is used to determine the low-contrast area and obtain the region segmentation result.
[0019] According to the region segmentation result, an adaptive contrast enhancement algorithm is used to calculate local stretching parameters for the low-contrast region to generate a second image;
[0020] If the global contrast of the second image is lower than a preset threshold, a global histogram equalization algorithm is used to adjust the second image as a whole to obtain a third image;
[0021] By calculating the pixel distribution characteristics of the third image, determining whether the target contrast reaches a preset threshold, and obtaining a contrast verification result;
[0022] If the contrast verification result does not reach a preset threshold, adjusting the parameters of the adaptive algorithm according to the pixel distribution characteristics of the third image and regenerating the second image;
[0023] Repeating the global histogram equalization process according to the regenerated second image to generate a new third image;
[0024] The pre-processed infrared image is obtained by performing pixel distribution characteristic analysis on the new third image.
[0025] Optionally, fusing the preprocessed infrared images using a multi-dimensional feature fusion network to obtain a fused image includes:
[0026] A convolutional neural network is used to extract multi-scale features from the preprocessed infrared image to generate a feature set containing information at different scales.
[0027] An attention mechanism is used to analyze the multi-scale features in the feature set, and the weight of each scale feature is calculated to obtain a weighted feature set;
[0028] If there are low-weight features in the weighted feature set, filtering is performed through a preset threshold, retaining high-weight features to generate a streamlined feature set;
[0029] The simplified feature set is integrated to obtain the fused image.
[0030] Optionally, performing multi-scale feature extraction on the fused image to obtain multi-dimensional features of small objects includes:
[0031] The morphological features, radiation features and motion features of the fused image are extracted through a preset multi-scale feature extraction network to obtain multi-dimensional features of the small target.
[0032] Optionally, inputting the small target multidimensional feature map into a target detection module to obtain a detection result includes:
[0033] Determine whether the features in the small target multidimensional feature map exceed a preset threshold;
[0034] If at least one feature value in the multidimensional feature map of the small target exceeds a preset threshold, the target existence is judged by a preset classifier to obtain a target existence flag;
[0035] According to the target existence flag, a bounding box regression algorithm is used to determine the target position from the small target multidimensional feature map to obtain the target bounding box coordinates;
[0036] Extracting the target category from the small target multidimensional feature map using the target bounding box coordinates and the preset classifier output to obtain a target category label;
[0037] The image rendering technology is used to mark the target bounding box coordinates and target category labels on the small target multi-dimensional feature map to obtain the detection result.
[0038] Compared with the prior art, the present invention has the following advantages and technical effects:
[0039] The present invention improves detection accuracy and real-time performance through multi-stage processing. First, the input image is subjected to non-local mean filtering for denoising and adaptive contrast enhancement. Then, deep learning methods are used to suppress complex background interference. For low-saliency targets, an attention mechanism is used to fuse multi-scale information, and a multi-scale feature extraction network is used to obtain morphological, radiation, and motion features. Finally, a lightweight convolutional neural network is used for target detection, and online parameter updates are performed based on the confidence level. Through multi-stage optimization processing, the present invention effectively improves the detection accuracy of small targets in infrared images while meeting real-time requirements. It is suitable for infrared small target detection scenarios in complex backgrounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0041] Figure 1 This is a flow chart of an infrared small target detection method based on multimodal feature fusion according to an embodiment of the present invention. DETAILED DESCRIPTION
[0042] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0043] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0044] This embodiment proposes an infrared small target detection method based on multimodal feature fusion, such as Figure 1 As shown, the specific steps include:
[0045] Collect infrared images;
[0046] Preprocessing the infrared image to obtain a preprocessed infrared image;
[0047] The pre-processed infrared images are fused using a multi-dimensional feature fusion network to obtain a fused image;
[0048] Perform multi-scale feature extraction on the fused image to obtain a multi-dimensional feature map of the small target;
[0049] The multi-dimensional feature map of the small target is input into the target detection module to obtain the detection results. The target detection model is constructed by integrating a convolutional neural network with a cross-layer multi-attention mechanism.
[0050] Specifically, ① Research on infrared small target detection method based on multimodal fusion:
[0051] In order to address the problems of poor texture features, blurred edge contours, and low detection accuracy in infrared images, small infrared targets are modeled based on morphological features, contrast, infrared radiation, and motion features. A multimodal joint feature fusion network is constructed to extract features with strong representational capabilities and discrimination, thereby improving the utilization rate of small target feature information and making up for the shortcomings of single-modal feature description. The feature fusion network is optimized based on the YOLO detection framework to further improve the detection effect of small targets in complex scene changes and with strong mobility of the targets to be detected.
[0052] ② Research on the fusion of visual saliency and multi-scale local contrast enhancement algorithm:
[0053] To address complex background interference, such as large, high-brightness clouds and ground buildings, a fully convolutional neural network-based enhancement algorithm was developed that integrates visual saliency and multi-scale local contrast information. This algorithm uses visual saliency detection within visual attention to roughly extract candidate target regions. It then integrates spatial domain filtering and a multi-scale local contrast learning module to extract and fuse local contrast information at different scales, enhancing target energy and suppressing high-brightness clutter in complex spatial environments. Finally, a spatiotemporal feature extraction network was employed to accurately detect small infrared targets in complex environments, effectively improving the model's anti-interference capabilities and robustness.
[0054] ③ Research on lightweight infrared target detection algorithm based on multiple attention perception:
[0055] To improve the suppression of highlight stripes and corners in complex environments, increase target detection accuracy, address high false alarm rates, and achieve model lightweighting, a lightweight infrared small target real-time detection algorithm integrating multiple attention mechanisms was established based on YOLO. The backbone network was designed using lightweight modules to reduce the number of model parameters and weights. A multiple attention perception module was constructed by integrating coordinate attention mechanisms, spatial attention mechanisms, and channel-detail attention mechanisms, enabling adaptive adjustment of image features across different dimensions to capture more detailed information. A concurrent downsampling approach of convolution and attention was used to focus on and attend to target features, capturing more detailed information about small infrared targets, avoiding feature aliasing, and improving target detection accuracy and signal-to-noise ratio.
[0056] Furthermore, the infrared image set covers different weather conditions, illumination changes, large-area bright cloud interference, ground building interference, and small infrared targets of various sizes.
[0057] Furthermore, the infrared image is preprocessed to obtain the preprocessed infrared image, including:
[0058] The infrared image is denoised using a non-local mean filtering algorithm to obtain a denoised infrared image.
[0059] According to the pixel distribution characteristics of the denoised infrared image, an adaptive contrast enhancement algorithm is used to perform local stretching processing on the low-contrast area to obtain the preprocessed infrared image.
[0060] Furthermore, the non-local mean filtering algorithm is used to denoise the infrared image, and the denoised infrared image is obtained including:
[0061] The infrared image is denoised using a non-local mean filtering algorithm to obtain a first image with reduced noise interference;
[0062] If the noise level of the first image is higher than a preset threshold, a secondary denoising process is performed on the first image using a median filtering algorithm to obtain a denoised infrared image.
[0063] Specifically, the image acquisition module extracts an input image from an infrared image to obtain an initial image. For example, in an infrared thermal imager monitoring scenario, the device captures an image of the surface of a high-temperature device with a resolution of 640×480, where the pixel values reflect the temperature distribution. The original image may contain noise, such as thermal noise or sensor jitter. A non-local means filtering algorithm is used to denoise the initial image, producing a first image. Specifically, non-local means filtering replaces noise points with the mean value of pixels in similar regions of the image, preserving detail. For example, at the edge of a high-temperature device, the algorithm searches for similar texture blocks with a window size of 7×7 and a search range of 21×21, smoothing out noise while preserving boundary texture. This improves image clarity and facilitates subsequent analysis. If the noise level of the first image exceeds a preset threshold, a secondary denoising operation is performed using a median filter to obtain a denoised infrared image. In one possible implementation, the noise level is assessed using mean squared error, with a threshold of 20. If the noise level exceeds the threshold, a 3×3 window median filter is used to replace the center pixel value. For example, isolated noise points in areas of uneven heat dissipation on the device surface are smoothed, generating a cleaner, denoised infrared image. This method effectively suppresses salt and pepper noise and improves image quality.
[0064] Furthermore, based on the pixel distribution characteristics of the denoised infrared image, an adaptive contrast enhancement algorithm is used to perform local stretching processing on the low-contrast area. The preprocessed infrared image is obtained, including:
[0065] By analyzing the pixel distribution characteristics of the denoised infrared image, the histogram statistical method is used to determine the low-contrast area and obtain the region segmentation result.
[0066] Based on the region segmentation results, an adaptive contrast enhancement algorithm is used to calculate local stretching parameters for low-contrast regions to generate a second image;
[0067] If the global contrast of the second image is lower than a preset threshold, a global histogram equalization algorithm is used to adjust the second image as a whole to obtain a third image;
[0068] By calculating the pixel distribution characteristics of the third image, determining whether the target contrast reaches a preset threshold, and obtaining a contrast verification result;
[0069] If the contrast verification result does not reach the preset threshold, the parameters of the adaptive algorithm are adjusted according to the pixel distribution characteristics of the third image to regenerate the second image;
[0070] Repeating the global histogram equalization process according to the regenerated second image to generate a new third image;
[0071] The pre-processed infrared image is obtained by performing pixel distribution characteristic analysis on the new third image.
[0072] Specifically, in infrared image processing, when analyzing the pixel distribution characteristics of a denoised infrared image, histogram statistics can be used to identify areas of low contrast. The core of histogram statistics is to calculate the distribution of grayscale values, thereby determining which areas have excessively concentrated grayscale values, indicating insufficient contrast. For example, in an infrared image, assuming the grayscale value range is 0 to 255, statistically determining that the grayscale values in a certain area are primarily concentrated between 50 and 100 indicates low contrast. This analysis helps accurately locate areas requiring enhancement. In one possible implementation, an adaptive contrast enhancement algorithm is used based on the region segmentation results to target low-contrast areas. This adaptive algorithm dynamically calculates stretching parameters based on the grayscale mean and distribution characteristics of each region. For example, for areas with concentrated grayscale values, the algorithm may stretch the grayscale values to a wider range, such as from 50 to 100 to 30 to 120, to enhance the visibility of details within the area. The advantage of this method is that it preserves local image features and avoids detail loss caused by global adjustments. Specifically, if the global contrast of the second image falls below a preset threshold, global histogram equalization can be used to perform overall adjustments. The principle of global histogram equalization is to remap the image's grayscale values to make them more evenly distributed. For example, in an infrared image, if the original grayscale value distribution exhibits a unimodal characteristic, after equalization, the grayscale value distribution becomes flatter, improving the overall contrast. This method is particularly suitable for addressing the overall dimness of infrared images caused by insufficient illumination. It should be noted that verifying whether the fourth image's contrast meets the required standards can be accomplished through statistical analysis of pixel distribution characteristics. For example, after calculating the image's grayscale histogram, the standard deviation of the grayscale values can be examined. If the standard deviation falls below a certain threshold, such as 20, the contrast is still insufficient. Based on the analysis results, the parameters of the adaptive algorithm can be adjusted, such as by increasing the stretch factor or expanding the local adjustment range, and then the second image can be regenerated. This iterative optimization approach can gradually approach the ideal contrast effect. Preferably, after regenerating the second image, the global histogram equalization process is repeated to generate a new fourth image. For example, in one iteration, the grayscale value distribution of the adjusted second image becomes more even, but there may still be slight contrast deficiencies in some local areas. Through further equalization, the overall grayscale distribution of the new third image becomes smoother, and details are better rendered. This method ensures continuous improvement in image quality. In one embodiment, the contrast enhancement result is ultimately determined by analyzing the pixel distribution characteristics of the new third image. For example, the entropy value of the grayscale histogram is examined. If the entropy value increases significantly, such as from 5.0 to 6.5, it indicates that the image information content has increased and the contrast effect is ideal. This analysis method verifies the effectiveness of the algorithm through quantitative indicators, providing a reliable basis for subsequent image processing.
[0073] Furthermore, the pre-processed infrared images are fused using a multi-dimensional feature fusion network to obtain a fused image, including:
[0074] A convolutional neural network is used to extract multi-scale features from the preprocessed infrared image to generate a feature set containing information at different scales.
[0075] The attention mechanism is used to analyze the multi-scale features in the feature set, calculate the weights of the features at each scale, and obtain the weighted feature set;
[0076] If there are low-weight features in the weighted feature set, the preset threshold is used for filtering, and the high-weight features are retained to generate a streamlined feature set;
[0077] The simplified feature set is integrated to obtain a fused image.
[0078] Specifically, convolutional neural networks use multiple layers of convolution to capture features from low-level edges to high-level semantics. For example, the first layer of the network might extract the object's outline, the second layer captures texture details, and the third layer focuses on the object's overall shape. Assuming the sixth image is 256x256 pixels, the network can generate a multi-scale feature set with resolutions of 32x32, 64x64, and 128x128. This approach ensures that even small objects appear blurry at high resolution, salient features are preserved at low resolution. When analyzing multi-scale features, the attention mechanism assigns weights to each scale to highlight key information. Specifically, the attention mechanism calculates the global information of the feature map and generates a weight matrix. Assuming the multi-scale feature set includes three scales, the attention mechanism might assign a weight of 0.2 to low-resolution features, 0.5 to medium-resolution features, and 0.3 to high-resolution features. The weights reflect the contribution of each scale to object detection. Preferably, if the background is complex, low-resolution features may have higher weights because they better highlight the object's overall shape. This weighting scheme allows subsequent processing to focus on important features. If low-weight features exist in the weighted feature set, high-weight features are retained through threshold filtering. For example, by setting a threshold of 0.4, low-weight features, such as low-resolution features with a weight of 0.2, are eliminated, resulting in a reduced feature set. In one possible implementation, medium- and high-resolution features are retained after filtering to ensure a more compact feature set and reduce the interference of redundant information on subsequent fusion. This reduction method is particularly suitable for small target detection scenarios, as it avoids the influence of background noise. When integrating the reduced feature set, the feature fusion module aligns and merges features of different scales. It is understood that the fusion module may use a weighted summation or concatenation approach. For example, 64x64 and 128x128 feature maps are aligned to the same size through upsampling, and then weighted fusion is performed pixel by pixel to generate a fused feature map. This fusion method preserves multi-scale information and enhances the representation of small targets. Upsampling is often used to increase the resolution of the fused feature map to enhance the representation of small target features. For example, a 64x64 fused feature map is upsampled to 128x128 through bilinear interpolation to enhance the details of small targets. In one embodiment, the object edges are smoother after upsampling, which facilitates subsequent segmentation or tracking tasks.
[0079] Furthermore, multi-scale feature extraction is performed on the fused image to obtain multi-dimensional features of small targets, including:
[0080] The morphological features, radiation features and motion features of the fused image are extracted through a preset multi-scale feature extraction network to obtain multi-dimensional features of small targets.
[0081] Furthermore, the small target multi-dimensional feature map is input into the target detection module to obtain the detection results including:
[0082] Determine whether the features in the small target multidimensional feature map exceed the preset threshold;
[0083] If at least one eigenvalue in the multidimensional feature map of the small target exceeds a preset threshold, the target existence is judged by a preset classifier to obtain a target existence flag;
[0084] According to the target existence mark, the bounding box regression algorithm is used to determine the target position from the small target multidimensional feature map to obtain the target bounding box coordinates;
[0085] The target category label is obtained by extracting the target category from the small target multidimensional feature map through the target bounding box coordinates and the preset classifier output;
[0086] Image rendering technology is used to mark the target bounding box coordinates and target category labels on the small target multi-dimensional feature map to obtain the detection results.
[0087] Specifically, lightweight networks such as MobileNet use depthwise separable convolutions to reduce computational complexity, making them suitable for real-time applications. Specifically, a network consisting of three convolutional layers can be designed, with each layer extracting features at different levels, such as edges and textures, to generate a set of feature maps. For example, assuming the ninth image contains a small drone, the feature map set might include the drone's outline and surface texture features. This approach significantly reduces computational resource requirements while ensuring feature richness. When at least one feature value in the feature map set exceeds a preset threshold, a preset classifier is used to determine the presence of the target. Preferably, a support vector machine (SVM) can be used as the classifier, taking the feature values as input and outputting a binary classification result: target presence or absence. For example, if the threshold is set to 0.8, when the activation value of a feature map reaches 0.9, the classifier confirms the presence of the target and generates a target presence flag. This mechanism uses the threshold to filter out irrelevant information, ensuring that subsequent processing focuses on high-confidence targets. Based on the target presence flag, a bounding box regression algorithm is used to determine the target's location and generate the target's bounding box coordinates. In one embodiment, a region proposal-based regression algorithm can be used to predict the coordinates of the target's rectangular box from the feature map. For example, when a drone's feature map is input into a regression model, the output coordinates are (100, 150, 50, 50), representing the target's center point, width, and height. This precise positioning provides the foundation for subsequent annotation. Extracting the target category and generating a label using the target's bounding box coordinates and the classifier output is a key step. It's understandable that the classification head of a convolutional neural network can be combined to predict the target category based on the feature map. For example, after analyzing the drone's features, the classifier outputs the label "drone" rather than "bird." This refined classification ensures accurate target recognition. Image rendering techniques are used to annotate the bounding box and category labels on the small target's multidimensional feature map to generate detection results. Specifically, an image processing library can be used to draw a red rectangular box on the small target's multidimensional feature map and add the text label "drone." For example, the coordinates (100, 150, 50, 50) can be defined and the category labeled to generate an intuitive detection result image. This visualization facilitates quick understanding of the detection results. Image post-processing techniques are used to standardize the detection results to generate the final detection image. It should be noted that the detection results can be converted to a unified format, such as JPEG format with a resolution of 640x480, by adjusting the image resolution or color space. This standardization process ensures that the detection images are suitable for different display devices or storage requirements, improving compatibility.
[0088] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for detecting small infrared targets based on multimodal feature fusion, characterized in that: include: Collect infrared images; Preprocessing the infrared image to obtain a preprocessed infrared image; fusing the preprocessed infrared images using a multi-dimensional feature fusion network to obtain a fused image; Performing multi-scale feature extraction on the fused image to obtain a small target multi-dimensional feature map; The small target multidimensional feature map is input into the target detection module to obtain the detection result, wherein the target detection model is constructed by a convolutional neural network that integrates a cross-layer multi-attention mechanism.
2. The infrared small target detection method based on multimodal feature fusion according to claim 1 is characterized in that: The infrared image set covers different weather conditions, lighting changes, large-area bright cloud interference, ground building interference, and infrared small targets of various sizes.
3. The infrared small target detection method based on multimodal feature fusion according to claim 1 is characterized in that: Preprocessing the infrared image to obtain the preprocessed infrared image includes: The infrared image is denoised using a non-local mean filtering algorithm to obtain a denoised infrared image. According to the pixel distribution characteristics of the denoised infrared image, an adaptive contrast enhancement algorithm is used to perform local stretching processing on the low-contrast area to obtain a preprocessed infrared image.
4. The infrared small target detection method based on multimodal feature fusion according to claim 3 is characterized in that: The infrared image is denoised using the non-local mean filtering algorithm. The denoised infrared image includes: The infrared image is denoised using a non-local mean filtering algorithm to obtain a first image with reduced noise interference; If the noise level of the first image is higher than a preset threshold, a secondary denoising process is performed on the first image using a median filtering algorithm to obtain a denoised infrared image.
5. The infrared small target detection method based on multimodal feature fusion according to claim 3 is characterized in that: According to the pixel distribution characteristics of the denoised infrared image, an adaptive contrast enhancement algorithm is used to perform local stretching processing on low-contrast areas. The preprocessed infrared image is obtained, which includes: By analyzing the pixel distribution characteristics of the denoised infrared image, the histogram statistical method is used to determine the low-contrast area and obtain the region segmentation result. According to the region segmentation result, an adaptive contrast enhancement algorithm is used to calculate local stretching parameters for the low-contrast region to generate a second image; If the global contrast of the second image is lower than a preset threshold, a global histogram equalization algorithm is used to adjust the second image as a whole to obtain a third image; By calculating the pixel distribution characteristics of the third image, determining whether the target contrast reaches a preset threshold, and obtaining a contrast verification result; If the contrast verification result does not reach a preset threshold, adjusting the parameters of the adaptive algorithm according to the pixel distribution characteristics of the third image and regenerating the second image; Repeating the global histogram equalization process according to the regenerated second image to generate a new third image; The pre-processed infrared image is obtained by performing pixel distribution characteristic analysis on the new third image.
6. The infrared small target detection method based on multimodal feature fusion according to claim 1, characterized in that: The pre-processed infrared images are fused using a multi-dimensional feature fusion network to obtain a fused image, including: A convolutional neural network is used to extract multi-scale features from the preprocessed infrared image to generate a feature set containing information at different scales. An attention mechanism is used to analyze the multi-scale features in the feature set, and the weight of each scale feature is calculated to obtain a weighted feature set; If there are low-weight features in the weighted feature set, filtering is performed through a preset threshold, retaining high-weight features to generate a streamlined feature set; The simplified feature set is integrated to obtain the fused image.
7. The infrared small target detection method based on multimodal feature fusion according to claim 1 is characterized in that: Performing multi-scale feature extraction on the fused image to obtain multi-dimensional features of small targets includes: The morphological features, radiation features and motion features of the fused image are extracted through a preset multi-scale feature extraction network to obtain multi-dimensional features of the small target.
8. The infrared small target detection method based on multimodal feature fusion according to claim 1, characterized in that: Inputting the small target multidimensional feature map into the target detection module, and obtaining the detection result includes: Determine whether the features in the small target multidimensional feature map exceed a preset threshold; If at least one feature value in the multidimensional feature map of the small target exceeds a preset threshold, the target existence is judged by a preset classifier to obtain a target existence flag; According to the target existence flag, a bounding box regression algorithm is used to determine the target position from the small target multidimensional feature map to obtain the target bounding box coordinates; Extracting the target category from the small target multidimensional feature map using the target bounding box coordinates and the preset classifier output to obtain a target category label; The image rendering technology is used to mark the target bounding box coordinates and target category labels on the small target multi-dimensional feature map to obtain the detection result.