Double-branch infrared small target detection method based on neural low-rank background modeling
Through the dual-branch detection method with low-rank neural background modeling and multi-scale adaptive enhancement, the performance degradation of infrared small object detection in complex backgrounds is solved, and efficient and robust object detection effect is achieved.
Patent Information
- Application Number
- CN202510445090.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The existing infrared small object detection methods have deteriorated detection performance under complex backgrounds, limited feature extraction capabilities, and there is a risk of overfitting, making it difficult to adapt to complex backgrounds and low-contrast targets.
A dual branch detection method based on neural low-rank background modeling is adopted, including neural low-rank background modeling module, dual perception enhancement and resolution adaptive amplification module, multi-scale contrast adaptive enhancement module and multi-scale dynamic detector, combining physical size sensitive loss, spatial coordinate sensitive loss and focus loss to optimize model parameters.
It significantly improves the robustness to complex backgrounds, enhances target characteristics, reduces background noise interference, and improves detection accuracy and generalization capabilities.
Smart Images

Figure CN120374943A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a dual-branch infrared small target detection method based on neural low-rank background modeling. Background Art
[0002] Target detection is an important research direction in the field of computer vision, aiming to detect the position and contour of specific objects in images or videos. With the development of machine vision, target detection is widely used in scenarios such as security monitoring, drone patrol, and industrial automation, and has great strategic significance for the intelligent transformation of various fields. Among them, in the fields of military reconnaissance, night vision combat, and intelligent security, the detection technology for infrared small targets such as small drones and floating objects plays a crucial role. However, limited by the characteristics of infrared small targets, such as extremely low pixel values, extremely low contrast, complex backgrounds, and extremely close small targets, there are many unique challenges in infrared image target detection. Traditional infrared image target detection methods have the advantages of all-weather operation and strong robustness against complex backgrounds, but still face efficiency and accuracy problems when dealing with infrared small targets. Therefore, using infrared imaging assistance to achieve efficient and robust infrared small target detection is still a very challenging task and the focus of current academic research.
[0003] Chinese Patent Publication No. "CN116468928B", titled "An Infrared Small Target Detection Method Based on Data Augmentation and Weak Feature Enhancement", first uses an improved feature extraction network to reduce the loss of small target features, designs an efficient feature enhancement module to strengthen the representation of the target area without significantly increasing the computational complexity. At the same time, the method introduces a balanced sample training strategy to effectively solve the problem of extremely unbalanced positive and negative samples. Finally, the method further integrates feature information of different scales through a multi-level feature fusion structure to generate high-quality detection results.
[0004] However, this method still faces significant challenges when dealing with infrared small targets. First, the method does not adequately consider the adaptability to complex background scenes, and the detection performance will significantly decline when the complexity of the target and background distribution increases; second, the existing feature extraction backbone network has limited feature expression ability for low-contrast small targets, resulting in the loss of key target features; in addition, although a sample balancing strategy is adopted, there is still a risk of overfitting under extremely unbalanced data, affecting the generalization ability of the model in complex environments. Therefore, designing an infrared small target detection method that can effectively cope with complex background interference, reduce the loss of target features, and solve the overfitting problem of the detection model is an urgent problem to be solved by the present invention. Summary of the Invention
[0005] The technical solution of the present invention to solve the above technical problems is to provide a dual-branch infrared small target detection method based on neural low-rank background modeling, including the following steps:
[0006] S1. Prepare the dataset: Five publicly available datasets are used, where Dataset 1, Dataset 2, and Dataset 3 are used for network training and fine-tuning; Dataset 4 and Dataset 5 are used for model testing.
[0007] S2. Construct a dual-branch detection model: It includes a neural low-rank background modeling module, a dual perception enhancement and resolution adaptive amplification module, a multi-scale contrast adaptive enhancement module, and a multi-scale dynamic detector.
[0008] S3. Train the network model: Input the training data of S1 into the neural low-rank background modeling module to extract background features, enhance the features through the dual perception enhancement module and then input them into the fusion module; at the same time, input the training data into the dual perception enhancement module for enhancement, perform adaptive weight learning through the multi-scale contrast adaptive enhancement module, and then input it into the fusion module for multi-scale information fusion training.
[0009] S4. Design the loss function and evaluation metrics: Use a joint loss function composed of physical size sensitive loss, spatial coordinate sensitive loss, and focal loss to optimize the model parameters, and evaluate the model performance through multi-dimensional metrics.
[0010] S5. Fine-tune the model: Optimize the model parameters based on Dataset 2.
[0011] S6. Save the model: Solidify the network parameters after fine-tuning to obtain the final detection model.
[0012] Further, in S1: Dataset 1 is the MDvsFA dataset; Dataset 2 is the SIRST-Aug dataset; Dataset 3 is the IRSTD-1k dataset; Dataset 4 is the NUDT-SIRST dataset; Dataset 5 is the SIRST dataset.
[0013] Further, in S2, the neural low-rank background modeling module includes:
[0014] A neural function implemented by a multi-layer perceptron, representing the background tensor through a parameterized factor function, and the factor function is respectively defined in the height, width, and time dimensions.
[0015] A neural regularization module, which imposes local smoothness and temporal regularization constraints by calculating the three-dimensional space gradient of the background tensor, and its mathematical expression is:
[0016]
[0017] Among them, and respectively represent the height, width, and time dimension gradients.
[0018] Further, in the S2, the dual perception enhancement and resolution adaptive amplification module includes:
[0019] Perception block: After expanding the number of channels through convolution, it is followed by an SE module to dynamically adjust the channel weights. When the input and output channels are the same, a residual connection is adopted;
[0020] Upsampling module: The resolution is improved through channel recombination, and the bicubic interpolation method is adopted with a scale factor of 3;
[0021] Skip connection: Fuses the traditional interpolated image with the output features of the perception block to retain spatial information.
[0022] Further, in the S2, the multi-scale contrast adaptive enhancement module includes:
[0023] Feature segmentation, adopting a multi-scale pyramid segmentation strategy, and realizing the scale separation of the feature subspace through three-level parallel paths and a spatial attention mechanism;
[0024] Convolutional feature extraction, through a cascaded depthwise separable convolutional group and a residual connection, extracting multi-scale features while ensuring computational efficiency;
[0025] Contrast enhancement, based on a local contrast meter, through sliding window statistics, non-linear gain adjustment, and channel attention weighting, realizing the feature enhancement of small target regions;
[0026] Feature recombination and fusion, adopting a cross-scale aggregation mechanism, and outputting enhanced features through bilinear interpolation alignment, channel-level concatenation fusion, and spatial attention selection.
[0027] Further, in the S2, the multi-scale dynamic detector includes:
[0028] Channel excitation branch, realizing channel-level enhancement through average pooling, 1×1 convolution, and an activation function;
[0029] Channel spatial perception branch, capturing spatial context through indexing operations, 3×3 convolution, and bias adjustment;
[0030] Channel recalibration branch, generating a channel weight map through average pooling, a fully connected layer, and normalization.
[0031] Further, in the S4, the joint loss function is:
[0032] L PDSC =L PD +L SC ;
[0033] where: L PDPhysical size sensitive loss for small targets, L SC Spatial coordinate sensitive loss for small targets; L PD Calculate the weight by normalizing the target size difference, and the mathematical definition is:
[0034]
[0035] Among them, P and G respectively represent the predicted box pixel set and the ground truth box pixel set;
[0036] Spatial coordinate sensitive loss Calculate the center point distance and angle difference through polar coordinate transformation;
[0037] The focal loss is:
[0038] L Focal =-α t (1 - p t ) γ log(p t )
[0039] Furthermore, the evaluation metrics in S4 include:
[0040] Pixel-level metrics: IoU, recall rate;
[0041] Object-level metric: detection probability;
[0042] Comprehensive metric: F1 score, and the calculation formula is:
[0043]
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] (1) The present invention designs a new end-to-end infrared small target detection network model. By designing a neural low-rank background modeling method, it explicitly models the low-rank characteristics and spatio-temporal continuity of the background, enhancing the background modeling ability for complex backgrounds; at the same time, it develops neural regularization, and by imposing a three-dimensional total variation constraint, it effectively suppresses background noise interference and significantly improves the robustness of background modeling.
[0046] (2) The present invention designs a new end-to-end infrared small target detection network model and introduces a dual perception enhancement and resolution adaptive amplification module. This module realizes the adaptive enhancement of infrared images by parallel processing of feature information in the channel and spatial dimensions.
[0047] (3) The present invention designs a new end-to-end infrared small target detection network model, which introduces a multi-scale contrast adaptive enhancement module. This module innovatively combines an adaptive feature segmentation mechanism and a differential convolution operation to strengthen the high-frequency details of the target area while suppressing the noise interference in the background area, thereby significantly enhancing the contrast between the target and the background while maintaining the background consistency.
[0048] (4) The present invention designs a new end-to-end infrared small target detection network model, which introduces a multi-scale dynamic detector. This module realizes the adaptive weight allocation for features of different scales based on the attention mechanism. The module first extracts multi-scale feature representations through a multi-branch structure, then calculates the importance weights of each scale feature using the channel attention and spatial attention mechanisms, and finally generates an enhanced feature representation through weighted fusion.
[0049] (5) The present invention designs a new end-to-end infrared small target detection network model, which introduces a joint loss effectively combined by a physical size-sensitive loss, a spatial coordinate-sensitive loss, and a focal loss, helping the detection network pay more attention to smaller target areas and providing a stable gradient for the training process to make the training process more reliable.
[0050] (6) The detection model proposed by the present invention shows good results in both the NUDT-SIRST dataset and the SIRST dataset, and significant improvements have been achieved in the quantitative evaluation indicators, indicating that the detection method proposed in this paper has very strong generalization ability in single / multi-target and different complex backgrounds and can adapt to most small target detection tasks and scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0052] Figure 1 It is a flowchart of the steps of a dual-branch infrared small target detection method based on neural low-rank background modeling according to the present invention;
[0053] Figure 2 It is a structural diagram of the dual-branch infrared detection model constructed according to the present invention;
[0054] Figure 3 It is a structural diagram of the neural low-rank background modeling module according to the present invention;
[0055] Figure 4Structural diagram of the dual - perception enhancement module and resolution - adaptive amplification module of the present invention;
[0056] Figure 5 Structural diagram of the multi - scale contrast - adaptive enhancement module of the present invention;
[0057] Figure 6 Structural diagram of the multi - scale dynamic detector of the present invention;
[0058] Figure 7 Qualitative comparison effect diagram of the infrared small - target detection method described in the present invention and the existing methods;
[0059] Figure 8 Schematic diagram of the comparison of evaluation indexes of the infrared small - target detection method described in the present invention and the existing methods. Detailed implementation manners
[0060] The present invention proposes a dual - branch infrared small - target detection method based on neural low - rank background modeling, aiming to solve problems such as complex background interference, loss of target features, and poor model generalization ability.
[0061] The following will illustrate a dual - branch infrared small - target detection method based on neural low - rank background modeling proposed by the present invention in specific embodiments:
[0062] Embodiment 1:
[0063] A dual - branch infrared small - target detection method based on neural low - rank background modeling, as Figure 1 shown, includes the following steps:
[0064] S1. Prepare the dataset: Five publicly available datasets are adopted, where Dataset 1, Dataset 2, and Dataset 3 are used for network training and fine - tuning; Dataset 4 and Dataset 5 are used for model testing;
[0065] S2. Construct a dual - branch detection model: It includes a neural low - rank background modeling module, a dual - perception enhancement and resolution - adaptive amplification module, a multi - scale contrast - adaptive enhancement module, and a multi - scale dynamic detector;
[0066] S3. Train the network model: Input the training data of S1 into the neural low - rank background modeling module to extract background features, enhance the features through the dual - perception enhancement module and then input them into the fusion module; at the same time, input the training data into the dual - perception enhancement module for enhancement, perform adaptive weight learning through the multi - scale contrast - adaptive enhancement module, and then input them into the fusion module for multi - scale information fusion training;
[0067] S4. Design the loss function and evaluation indexes: Use a joint loss function composed of physical - size - sensitive loss, spatial - coordinate - sensitive loss, and focal loss to optimize the model parameters, and evaluate the model performance through multi - dimensional indexes;
[0068] S5, Fine-tuning the model: Optimize the model parameters based on Dataset 2;
[0069] S6, Saving the model: Solidify the fine-tuned network parameters to obtain the final detection model.
[0070] Furthermore, in S1: Dataset 1 is the MDvsFA dataset; Dataset 2 is the SIRST-Aug dataset; Dataset 3 is the IRSTD-1k dataset; Dataset 4 is the NUDT-SIRST dataset; Dataset 5 is the SIRST dataset.
[0071] Furthermore, in S2, the neural low-rank background modeling module includes:
[0072] A neural function implemented by a multi-layer perceptron, representing the background tensor through a parameterized factor function, where the factor function is defined in the height, width, and time dimensions respectively;
[0073] A neural regularization module, by calculating the gradients of the background tensor in three-dimensional space, imposing local smoothness and temporal regularization constraints, and its mathematical expression is:
[0074]
[0075] where and represent the gradients in the height, width, and time dimensions respectively.
[0076] Furthermore, in S2, the dual perception enhancement and resolution adaptive amplification module includes:
[0077] Perception block: Dynamically adjust the channel weights by expanding the number of channels through convolution and then connecting to the SE module, and use residual connection when the input and output channels are the same;
[0078] Upsampling module: Improve the resolution by channel recombination, using bicubic interpolation method with a scale factor of 3;
[0079] Skip connection: Fuse the traditional interpolated image and the output features of the perception block to retain spatial information.
[0080] Furthermore, in S2, the multi-scale contrast adaptive enhancement module includes:
[0081] Feature segmentation, adopting a multi-scale pyramid segmentation strategy, and realizing scale separation of the feature subspace through three-level parallel paths and spatial attention mechanism;
[0082] Convolutional feature extraction, through cascaded depthwise separable convolutional groups and residual connections, extracting multi-scale features while ensuring computational efficiency;
[0083] Contrast enhancement, based on a local contrast meter, realizes feature enhancement of small target regions through sliding window statistics, non-linear gain adjustment, and channel attention weighting;
[0084] Feature recombination and fusion, adopting a cross-scale aggregation mechanism, outputs enhanced features through bilinear interpolation alignment, channel concatenation fusion, and spatial attention selection.
[0085] Further, in the S2, the multi-scale dynamic detector includes:
[0086] Channel excitation branch, realizing channel-level enhancement through average pooling, 1×1 convolution, and activation function;
[0087] Channel spatial perception branch, capturing spatial context through indexing operation, 3×3 convolution, and bias adjustment;
[0088] Channel recalibration branch, generating channel weight maps through average pooling, fully connected layer, and normalization.
[0089] Further, in the S4, the joint loss function is:
[0090] L PDSC =L PD +L SC ;
[0091] Where: L PD is the physical size sensitive loss of small targets, and L SC is the spatial coordinate sensitive loss of small targets; L PD calculates the weight by normalizing the target size difference, and the mathematical definition is:
[0092]
[0093] Among them, P and G respectively represent the predicted box pixel set and the ground truth box pixel set;
[0094] Spatial coordinate sensitive loss calculates the center point distance and angle difference through polar coordinate transformation;
[0095] The focal loss is:
[0096] L Focal =-α t (1-p t ) γ log(p t ).
[0097] Further, the evaluation metrics in the S4 include:
[0098] Pixel-level metrics: IoU, recall rate;
[0099] Target-level metric: Detection probability;
[0100] Comprehensive metric: F1 score, calculated as:
[0101]
[0102] Example 2:
[0103] A flowchart of an infrared small target detection method based on data augmentation and weak feature enhancement, which specifically includes the following steps:
[0104] S1, Prepare the dataset: MDvsFA dataset one is the first publicly available infrared small target dataset, containing 11,000 infrared images, some of which are synthetic images. SIRST-Aug dataset two contains 9,070 images and is based on dataset SIRST with data augmentation (cropping, rotation, displacement, etc.). IRSTD-1k dataset three contains 1,001 images, covering a variety of scenarios and target distributions, with highly sparse targets. NUDT-SIRST dataset four contains 1,327 infrared small target images under complex backgrounds, with low target size and background contrast. SIRST dataset five contains 427 infrared images and 480 targets, and is the first real infrared small target dataset with high-quality images and labels.
[0105] S2, Construct a dual-branch infrared detection model: The detection model is as Figure 2 shown. The model introduces a dual-branch design: Branch one uses a ResNeSt feature extraction module to extract deep features, and improves the recognition accuracy of small targets through a dual perception enhancement and resolution adaptive amplification module and a multi-scale contrast adaptive enhancement module; Branch two uses a neural low-rank background modeling method for preprocessing to suppress background interference, and further enhances target details through a dual perception enhancement and resolution adaptive amplification module. Subsequently, the two-path decoder consists of a fusion module, which performs multi-scale feature fusion using inverse wavelet transform, integrates features from different network levels, and combines low-level details and high-level semantic information.
[0106] Neural low-rank background modeling module, as Figure 3As shown: The neural function of the multi-layer perceptron to implement the low-rank prior obtains the tensor of the background from the image, that is, represents the background through the parameterized function of the multi-layer perceptron; subsequently, the neural three-dimensional total variation regularization uses the output derivative of the multi-layer perceptron to enforce local smoothness and temporal regularization in the background components to achieve the purpose of outputting the background; in the background modeling process, the low-rank property is the core of capturing background redundant information; however, traditional low-rank decomposition methods (such as matrix decomposition and tensor decomposition) often perform weakly in dealing with non-linear background changes. To overcome this shortcoming, this embodiment introduces a neural low-rank tensor representation to model the background through the non-linear expression ability of the neural network. Specifically, the estimation of the background tensor can be expressed as the following optimization objective:
[0107]
[0108] where B represents the background tensor, ||B|| * is the nuclear norm, is the gradient regularization term in the time domain, which is used to capture the smoothness of the background in the time dimension. a is the regularization weight, which is used to balance the nuclear norm and the time gradient regularization term. To further enhance the expression ability of background modeling, the neural function f θ is used to model the background tensor, and its form is:
[0109]
[0110] where: B(i,j,k) respectively represent the dimension indices of the background tensor, and are factor functions defined on the height, width, and time dimensions, and each factor function is implemented by a multi-layer perceptron.
[0111] Although the low-rank tensor representation has significant advantages in capturing the global characteristics of the background, its ability to model local changes is relatively limited. In the background, local changes often have high complexity and uncertainty, and it is difficult to accurately model these details only relying on the low-rank constraint. For this reason, a neural three-dimensional total variation regularization method is proposed to explicitly enhance the local smoothness of the background in the spatial and temporal dimensions while retaining the target information. The mathematical definition of the neural three-dimensional total variation regularization is as follows:
[0112]
[0113] where, and respectively represent the gradients of the background tensor in the height, width, and time dimensions. By constraining the local gradients of the background, the neural three-dimensional total variation regularization N3DTV can explicitly enhance the smoothness of the background while avoiding the loss of the target signal caused by over-smoothing.
[0114] Dual Perception Enhancement and Resolution Adaptive Amplification Module, such as Figure 4 (1) shown, this module consists of three main components: perception block, upsampling block, and skip connection. As Figure 4 (2) shown is the perception block, aiming to enhance feature perception and detail enhancement. It first expands the number of channels through convolution to enhance the model's expressiveness, and then uses the SE block to dynamically adjust the channel weights to highlight key features. If the number of input and output channels matches, residual connection is adopted to maintain the information flow. Upsampling technology improves the image resolution by reorganizing the channels of the input feature map. Its advantage lies in being able to preserve spatial relationships, reduce artifacts, and is important for enhancing the resolution of infrared images. The skip connection combines the traditionally adjusted image and the features enhanced by the perception block, helping the network better learn the target features, while preserving spatial information and improving the overall perception ability. The bicubic interpolation method is used in this embodiment, and the scale factor is 3.
[0115] Multi-scale Contrast Adaptive Enhancement Module, such as Figure 5 shown, the multi-scale contrast adaptive enhancement module is the core component of the dual-branch infrared detection model, specifically designed to address the challenges of low contrast and variable target shapes in infrared small target detection. This module effectively enhances and extracts small target information by segmenting and processing the input features, significantly improving the detection accuracy and robustness. Specifically, this module enhances the small target features through three steps: first, segment the input features and enhance the contrast; second, extract convolutional features; contrast enhancement and feature recombination fusion. This process significantly improves the detection accuracy and robustness, enabling the dual-branch infrared small target detection model to more accurately detect infrared small targets in complex backgrounds.
[0116] Multi-scale Dynamic Detector, such as Figure 6 shown, the dynamic detector can adaptively focus on the scale-space task information of the object, better learn the relative importance of each semantic level and spatial information of the target, and adapt to different task forms. Specifically, the dynamic detector can be expressed as:
[0117] W(F) = f T (f Sp (f Sc (F)·F)·F)·F;
[0118] where f Sc , f Sp , f T represent the attention functions on scale, space, and task respectively; the scale attention f Sc (F)·F realizes dynamic feature fusion based on the importance of each layer of features:
[0119]
[0120] Among them, f(·) is a 1×1 convolutional layer, and σ(·) is a sigmoid function; the spatial perception attention f Sp (F)·F utilizes deformable convolution to fuse features of different levels at the same spatial position:
[0121]
[0122] Among them, K is the number of sparse sampling positions, p k +Δp k and Δm k are learned from the input features, p k +Δp k is the self-learned spatial offset, and Δm k is the self-learned important scalar at the p k position. The task perception attention f T (F)·F dynamically switches channels to support different tasks:
[0123] f T (F)·F = max(α 1 (F)·F c +β 1 (F), α 2 (F)·F c +β 2 (F));
[0124] Among them, [α 1 , α 2 , β 1 , β 2 is a hyperfunction that controls the threshold. Through average pooling, two fully connected layers, and a normalization layer, and finally normalized through a sigmoid activation function.
[0125] S3. Train the detection network model: Input the dataset one prepared in step S1 into the detection network model constructed in step S2 for training. The above network architecture and training process are implemented using Python 3.10 and PyTorch 2.3, and efficient parallel computing is performed on an NVIDIA Titan Xp GPU. During the training process, the Adam optimizer is selected, and parameters λ1 = 0.5 and λ2 = 0.999 are set to accelerate the convergence process and reduce oscillations. The number of training iterations is set to 500, and the batch size is set to 16. In the selection of the optimization method, the MultiStepLR learning rate decay strategy is adopted. This strategy dynamically adjusts the learning rate, provides a larger learning rate at the beginning of training to accelerate convergence, and then gradually decreases to avoid overfitting. The learning rate is set to 5×10 -4, the attenuation factor is set to 0.1. During the training process, after each round of training, validation is performed on the IRSTD-1k dataset.
[0126] S4. Select an appropriate loss function and determine the optimal evaluation metric for this method: The loss function is designed for the physical size and spatial coordinates of infrared small targets, solving the limitations shown by the traditional L IoU and L Dice losses in detecting small targets. Its mathematical definition is:
[0127] L PDSC = L PD + L SC ;
[0128] where L PD is the physical size sensitive loss of small targets, and L SC is the spatial coordinate sensitive loss of small targets. Then, for L PD weight the L Dice loss:
[0129]
[0130] where P and G represent the predicted box pixel set and the ground truth box pixel set respectively. The mathematical definition of the weight ω is:
[0131]
[0132] where Var is a function of scalar variance. The greater the difference between P and G, the smaller ω, resulting in a greater physical size sensitive loss. The basic principle behind the design of ω is that if there is a significant difference between P and G, the detector should allocate more attention to the target with a greater loss. That is, when the physical size difference between P and G is large, the weight ω will increase, thereby imposing a higher penalty on the target with a larger physical size deviation.
[0133] For L SC by averaging the coordinates of all pixels, the corresponding center points of P and G are obtained. Subsequently, the coordinates of these two center points are transformed into the polar coordinate system. For the corresponding center points, the corresponding distance and angle in the polar coordinate system are:
[0134]
[0135] The spatial coordinate sensitive loss L SC is calculated as follows:
[0136]
[0137] Meanwhile, to address the issue of class imbalance in the target-sparse infrared small target dataset, this paper adopts focal loss. Focal loss makes the model pay more attention to difficult-to-detect small targets by reducing the loss contribution of easy-to-classify samples and increasing the loss contribution of difficult-to-classify samples.
[0138] L Focal = -α t (1 - p t ) γ log(p t );
[0139]
[0140] Among them, p t is the probability of predicting the positive class, α t is the weight parameter for balancing positive and negative samples, with different values for positive samples (targets) and negative samples (backgrounds). γ is the focusing parameter for adjusting the loss of difficult-to-classify samples. α t and γ are set to the default values of 0.25 and 2 respectively.
[0141] S4. When selecting an appropriate loss function and determining the optimal evaluation metrics for this method, the evaluation metrics adopt the ratio of the intersection to the union of the predicted region and the ground truth region, IoU, and the recall rate as pixel-level evaluation metrics, and use the detection probability to evaluate the target-level performance. These different metrics provide insights into the detector performance from different perspectives. The recall rate and detection probability focus on recall and false alarms, while IoU considers both aspects simultaneously. The mathematical definition of IoU is:
[0142]
[0143] The mathematical definition of the recall rate is:
[0144]
[0145] Among them, P ture is the count of true positive pixels, and P all is the total number of pixels in the image.
[0146] The detection probability is given by:
[0147]
[0148] Among them, N ture is the number of correctly predicted targets, and N all is the total number of targets. In addition, the harmonic mean F1-score of precision and recall is also considered during evaluation. Through the joint analysis of these metrics, the detection performance of the dual-branch infrared small target detection model and its performance in different scenarios can be comprehensively analyzed.
[0149] S5, Fine-tuning the model: Use the IRSTD-1k dataset to train and fine-tune the model again. Set the learning rate to 0.005 and iterate 500 rounds in total, keeping other parameters unchanged to further improve the performance of the detection network.
[0150] S6, Saving the model: After the training in step S3 is completed, solidify the network parameters after fine-tuning. After fine-tuning the model in S5, determine the final detection network model. When performing the infrared small target detection task, the infrared small target image can be directly input into the trained end-to-end network model to obtain the final prediction result.
[0151] Qualitative and quantitative comparisons between the existing technology and the method proposed in the present invention are as Figure 7 、 Figure 8 shown, where Figure 8 The white bold frame represents the accurately detected infrared small target, the white dashed frame represents the misdetection, and the white frame represents the missed detection. It can be seen from the figure that the method proposed in the present invention has a higher normalized intersection over union and detection rate than the existing methods, and a better balance between precision and recall is achieved. These indicators further illustrate that the method proposed in the present invention achieves better infrared small target detection performance and obtains the expected results. Method 1 (ASTTV-NTLA) adopts the adaptive spatio-temporal total variation and non-local attention mechanism and achieves good background suppression ability at an accuracy of 87.00%. Method 2 (U-Net) is based on the classical encoder-decoder architecture, and each index is balanced but relatively mediocre. Method 3 (ALCNet) leads with an accuracy of 96.93% through attention-guided local contrast enhancement, but the recall rate fluctuates. Method 4 (ISNet) uses the iterative saliency detection mechanism and shows stable segmentation performance at an F1-score of 88.82%. In contrast, the method of the present invention is significantly superior to the comparative methods in terms of the two key indicators of recall rate (89.20%) and IoU (84.55%). In particular, the IoU index is improved by nearly 7 percentage points compared with the optimal comparative method (ALCNet), reflecting a more accurate target localization ability.
[0152] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A dual-branch infrared small target detection method based on neural low-rank background modeling, characterized in that It includes the following steps: S1. Prepare the dataset: Five publicly available datasets are adopted, where Dataset 1, Dataset 2, and Dataset 3 are used for network training and fine-tuning; Dataset 4 and Dataset 5 are used for model testing; S2. Construct a dual-branch detection model: It includes a neural low-rank background modeling module, a dual-perception enhancement and resolution adaptive amplification module, a multi-scale contrast adaptive enhancement module, and a multi-scale dynamic detector; S3. Train the network model: Input the training data in S1 into the neural low-rank background modeling module to extract background features, enhance the features through the dual-perception enhancement module and then input them into the fusion module; at the same time, input the training data into the dual-perception enhancement module for enhancement, perform adaptive weight learning through the multi-scale contrast adaptive enhancement module, and then input it into the fusion module for multi-scale information fusion training; S4. Design the loss function and evaluation metrics: Use a joint loss function composed of physical size-sensitive loss, spatial coordinate-sensitive loss, and focal loss to optimize the model parameters, and evaluate the model performance through multi-dimensional metrics; S5. Fine-tune the model: Optimize the model parameters based on Dataset 2; S6. Save the model: Solidify the network parameters after fine-tuning to obtain the final detection model.
2. The dual-branch infrared small target detection method based on neural low-rank background modeling according to claim 1, wherein In the above S1: Dataset 1 is the MDvsFA dataset; Dataset 2 is the SIRST-Aug dataset; Dataset 3 is the IRSTD-1k dataset; Dataset 4 is the NUDT-SIRST dataset; Dataset 5 is the SIRST dataset.
3. The dual-branch infrared small target detection method based on neural low-rank background modeling according to claim 1, wherein In the above S2, the neural low-rank background modeling module includes: A neural function implemented by a multi-layer perceptron, representing the background tensor through a parameterized factor function, and the factor function is defined in the height, width, and time dimensions respectively; A neural regularization module, which imposes local smoothness and temporal regularization constraints by calculating the gradients of the background tensor in three-dimensional space, and its mathematical expression is: Among them, and respectively represent the height, width, and time dimension gradients.
4. The dual-branch infrared small target detection method based on neural low-rank background modeling according to claim 1, wherein In the above S2, the dual-perception enhancement and resolution adaptive amplification module includes: A perception block: After expanding the number of channels through convolution, it is followed by an SE module to dynamically adjust the channel weights, and a residual connection is used when the input and output channels are the same; An upsampling module: Improve the resolution through channel recombination, using the bicubic interpolation method with a scale factor of 3; A skip connection: Fuse the traditional interpolated image and the output features of the perception block to retain spatial information.
5. The dual-branch infrared small target detection method based on neural low-rank background modeling according to claim 1, wherein In the above S2, the multi-scale contrast adaptive enhancement module includes: Feature segmentation, parallelly process the input features through 3×3 convolution and a dual 1×1 convolution branch, and after splicing and fusion, reduce the dimension and output to achieve multi-scale feature extraction; Convolutional feature extraction, including a three-branch convolutional structure and introducing a residual connection to retain the original information while extracting deep features; Contrast enhancement, based on a local contrast meter, through sliding window statistics, non-linear gain adjustment, and channel attention weighting, to achieve feature enhancement of small target regions; Feature recombination and fusion, splice the features at all levels and then compress and integrate them through 1×1 convolution, and finally output features of the target dimension.
6. The dual-branch infrared small target detection method based on neural low-rank background modeling according to claim 1, characterized in that In the above S2, the multi-scale dynamic detector includes: A channel excitation branch, which realizes channel-level enhancement through average pooling, 1×1 convolution, and an activation function; The channel spatial perception branch captures spatial context through indexing operations, 3×3 convolutions, and bias adjustments; The channel recalibration branch generates channel weight maps through average pooling, fully connected layers, and normalization.
7. The dual-branch infrared small target detection method based on neural low-rank background modeling according to claim 1, wherein In S4, the joint loss function is: L PDSC = L PD + L SC ; Where: L PD is the physical size sensitive loss of small targets, and L SC is the spatial coordinate sensitive loss of small targets; L PD Calculates the weight by normalizing the target size difference, and the mathematical definition is: where P and G represent the predicted box pixel set and the ground truth box pixel set respectively; Spatial coordinate sensitive loss Calculate the center point distance and angle difference through polar coordinate transformation; The focal loss is: L Focal = -α t (1 - p t ) γ log(p t ).
8. The double-branch infrared small target detection method based on neural low-rank background modeling according to claim 1, wherein The evaluation metrics in S4 include: Pixel-level metrics: IoU, recall rate; Object-level metric: detection probability; Comprehensive metric: F1 score, and the calculation formula is:
Citation Information
Patent Citations
A method for detecting small thermal infrared targets based on visual perception correlators
CN116468928B
Attention-guided hybrid double-branch spatial decomposition neural network infrared weak and small target detection method
CN118570439A
Hyperspectral target detection method of binary-classification encoder network based on momentum update
US20240386699A1
Object-level infrared-and-visible-light image fusion method based on fully convolutional neural network
WO2024174488A1
Cited By
Geological disaster change detection method and system based on improved twin U-Net and central surrounding double-flow network
CN121259596A
Geological disaster change detection method and system based on improved twin u-net and center-surround dual-stream network
CN121259596B