Contrast-based Yolo11 infrared weak and small target detection method

By using the Contrast-Yolo model, combined with the DC-DC module and the Attn-MLCL module, and based on contrast feature extraction and multi-scale learning, the problems of low detection accuracy and difficulty in detecting small targets in infrared weak target detection are solved, achieving higher detection accuracy and robustness.

CN120932061APending Publication Date: 2025-11-11XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510949975.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing infrared methods for detecting weak targets have low detection accuracy in complex backgrounds and are difficult to detect small targets. Traditional methods have advantages in real-time performance and hardware requirements, but low detection accuracy. Deep learning methods also have difficulties in detecting small targets.

Method used

A contrast-based YOLO11 infrared weak target detection method is adopted. By using the Contrast-YOLO model, combined with the DC-CDC module and the Attn-MLCL module, the difference between the target and the background is enhanced. Contrast feature extraction and multi-scale learning are used to optimize the loss function to improve detection accuracy.

Benefits of technology

It improves the robustness and accuracy of infrared detection of small targets, reduces background noise interference, and significantly enhances the detection capability of small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932061A_ABST
    Figure CN120932061A_ABST
Patent Text Reader

Abstract

The invention discloses a contrast-based Yolo11 infrared weak and small target detection method, and the method comprises the following specific steps: 1, obtaining an infrared image data set, marking an infrared image in the data set, obtaining a Mask image, sorting the Mask image into a Yolo data set label format, and dividing the sorted data set into a training set and a test set; step 2, constructing a Contrast-Yolo model, and inputting the training set in the step 1 into the Contrast-Yolo model for training; and step 3, inputting the test set into the Contrast-Yolo model trained in the step 2 for detection. The method is high in detection precision and can be suitable for small target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of infrared weak target detection methods, specifically relating to a contrast-based Yolo11 infrared weak target detection method. Background Technology

[0002] In the field of modern target detection, infrared small target detection technology has become a research hotspot in computer vision and signal processing due to its core application value in key areas such as military reconnaissance, security monitoring, and remote sensing. The "smallness" of a target manifests as a low signal-to-noise ratio, a small pixel ratio, and susceptibility to interference from complex background noise. Its accurate detection is crucial for scenarios such as early construction of multi-layered defense systems on the battlefield and all-weather monitoring in adverse weather conditions.

[0003] Currently, detection methods in this field fall into two main categories: The first is traditional infrared small target detection, which primarily employs techniques such as background suppression, filtering, and model-based analysis. These methods are simple, easy to implement, and highly targeted, but they have significant limitations. They are sensitive to background changes, easily leading to false positives, false negatives, and loss of key target information, resulting in low detection accuracy. The second approach involves introducing deep learning into this field. With its powerful feature extraction capabilities, deep learning can automatically learn target features from large amounts of data, reducing the complexity of manual feature extraction through end-to-end learning and significantly improving detection accuracy. However, deep learning models rely on large amounts of labeled data for training, and for targets that are too small, neural network downsampling operations struggle to extract effective feature information.

[0004] Traditional methods have advantages in real-time performance and hardware requirements, but suffer from low detection accuracy. While deep learning methods perform well in feature extraction and detection accuracy, they still face challenges in detecting small targets. Therefore, new methods are needed to address the problems of low detection accuracy under complex background interference and the difficulty in detecting small targets in infrared weak target detection. Summary of the Invention

[0005] The purpose of this invention is to provide a contrast-based method for detecting weak targets using YOLOv11 infrared technology, which solves the problems of low detection accuracy and difficulty in detecting small targets in existing methods.

[0006] The technical solution adopted in this invention is a contrast-based method for detecting weak targets using YOLOv11 infrared technology, and the specific steps are as follows: Step 1: Obtain the infrared image dataset, label the infrared images in the dataset to obtain mask images, and organize the mask images into the YOLO dataset label format. Divide the organized dataset into training set and test set. Step 2: Contrast-Yolo model. Input the training set from Step 1 into the Contrast-Yolo model for training. Step 3: Input the test set into the Contrast-Yolo model trained in Step 2 for detection. The invention is further characterized by: In step 1, the specific process of organizing the Mask image into the YOLO dataset label format is as follows: Step 1.1: Use OpenCV's imread function to read the Mask image in grayscale mode, obtain a two-dimensional array, and get the width and height of the Mask image; Step 1.2: Find the boundary points of all white areas in the Mask image to form a closed contour; Step 1.3: Generate a minimum bounding rectangle bounding box for each closed contour in Step 1.2; Step 1.4: Normalize the coordinates of the bounding box generated in Step 1.3; Step 1.5: Assign the category of each bounding box, the top-left corner coordinates of the bounding box after transformation in Step 1.4, and the width and height of the bounding box as the YOLO dataset label format.

[0007] In step 2, the Contrast-Yolo model is an improvement on the Yolo11 model, consisting of the Backbone, Neck, and Head; Specifically, the Backbone is as follows: the initial convolution in the YOLO11 model Backbone is replaced with a DC-DC module, and an Attn-MLCL module is added after each C3k2 block connected to the YOLO11 model Neck in the YOLO11 model Backbone. The Neck and head are completely identical to those in the YOLO11 model.

[0008] The processing procedure of the DCDC module is as follows: Step S1: Input the training set into the first regular convolution (conv) to extract features and obtain preliminary features; The kernel of the first regular convolutional conv is: The step size is 1; Step S2: The preliminary features obtained in step S1 are simultaneously divided into four parts, which are then input into the second regular convolutional module (conv), the CA_conv module, the third regular convolutional module (conv), and the CB_conv module, respectively, to obtain... , , , ; This is the output of the second regular convolution (conv). This is the output of the CA_conv module. This is the output of the third regular convolution, conv. This refers to the output of the CB_conv module; The kernels of the second and third regular convolutions are both... The step size is 1; Step S3, Calculate The dynamic weights (dynamic_weight) of the output of the CB_conv module are used to update the mask (update_mask) to obtain feature A. The dynamic weights (dynamic_weight) of the output of the CB_conv module are used to update the mask (update_mask) to obtain feature B. Step S4: Weight the output of the second regular convolution (conv) with feature A to obtain the feature. The output of the third regular convolution (conv) is weighted with feature B to obtain the feature. , will feature With features After concatenation, features are fused using a fourth regular convolutional layer (conv) to obtain fused features. ; The kernel of the fourth regular convolution conv is: The step size is 1; (2) (3) (4) In equations (2) and (3), This indicates batch normalization.

[0009] The processing procedure of the CA_conv module is as follows: The initial features input to the CA_conv module are padded with 1s to obtain feature X. padded The 3x3 difference mask CA is expanded into a 4D convolutional kernel CA that matches the input channels. expanded, Then perform a custom convolution to obtain the difference convolution result. ; The 3x3 differential mask CA is as follows: the center pixel has a weight of 4, the weights of the top, bottom, left, and right neighbors are all 1, and the weights of the top left, bottom left, top right, and bottom right corners are all 0. The expression for performing a custom convolution is: (1); The processing procedure of the CB_conv module is the same as that of the CA_conv module. In the CB_conv module, the 3x3 difference mask CB has a center pixel weight of 4, the weights of the top, bottom, left, and right neighbors are all 0, and the weights of the top left, bottom left, top right, and bottom right corners are all 1.

[0010] The processing procedure for each Attn-MLCL module is as follows: The feature map output from the C3k2 block is simultaneously input into the three branches. The outputs of the three branches are concatenated and then fed into the attention module. The attention module calculates different weights based on the importance of the features. The weights output by the attention module are weighted with the corresponding branch outputs. The weighted features are concatenated again. The concatenated features are then subjected to a 1×1 regular convolution with a stride of 1 to obtain the output features of the Attn-MLCL module.

[0011] The first branch consists of a 1×1 standard convolution with a stride of 1 and a 3×3 dilated convolution with a dilation rate of 1; the second branch consists of a 3×3 standard convolution with a stride of 1 and a 3×3 dilated convolution with a dilation rate of 3; and the third branch consists of a 5×5 standard convolution with a stride of 1 and a 3×3 dilated convolution with a dilation rate of 5.

[0012] The loss function during training is: Step A: Generate a target region mask based on the predicted bounding box and foreground mask. The target region mask is used to identify the target region of interest in the image. Calculate the mean and standard deviation of the target region and normalize the image data of the target region. The normalized expression is: (5) In equation (5), Indicates the input image. This represents the target region mask. Indicates the mean of the target region. Indicates the standard deviation of the target region. This represents a very small constant that prevents division by zero. , B is the batch size, i.e. the number of samples input into the network at one time; C is the number of channels, infrared images are usually single-channel; H and W are the image height and image width, respectively; b represents the image index in the batch; c represents the color channel index; and x and y represent the coordinates of the image in two-dimensional space. Step B: Calculate the contrast loss; The expression is: (6) In equation (6), This represents the average intensity of all target pixels in the b-th image of the batch. This is the target area mask, where a value of 1 represents the target area and 0 represents the background. Step C: Calculate the gradient loss; The expression is: (7) (8) (9) In equations (7) to (9), S x and S y G represents the convolution kernel of the Sobel operator in the x and y directions. x and G y This is the gradient map of the image in the x and y directions; Indicates the gradient magnitude of the target region; This represents the global mean of the gradient magnitude of the target region in the b-th image of the batch. Step D: Add the contrast loss and gradient loss according to their weights to obtain the contrast-gradient combined loss; The expression is: (10) In equation (10), The gradient weights are fixed; w(t) represents the time-varying weights. Step E: Calculate the total loss function; The expression is: (11) In equation (11), The value is 0.1; It consists of bounding box loss, classification loss, and distribution focusing loss.

[0013] The beneficial effects of this invention are: (1) The present invention is a contrast-based method for detecting weak targets in YOLO11 infrared. It uses a DC-DC module to enhance the difference between the target and the background, forces the learning of local contrast patterns in the shallow network, enhances the contrast features of small targets, effectively improves the feature extraction capability of the deep network, and reduces noise interference. (2) The contrast-based YOLO11 infrared weak target detection method of the present invention uses the Attn-MLCL module to perform deep mining of shallow contrast features. Through multi-scale contrast learning, it can generate richer feature representations, enabling the model to more sensitively capture the existence of small targets, reduce the interference of background noise, and improve the robustness of detection. (3) The present invention is a contrast-based method for detecting weak targets in YOLO11 infrared. The loss function is a composite loss function that integrates target region contrast optimization and gradient feature enhancement, which directly optimizes the salience of the target and improves the detection accuracy. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the Contrast-Yolo model in the contrast-based Yolo11 infrared weak target detection method of the present invention; Figure 2 This is a schematic diagram of the DCDC module in the contrast-based YOLO11 infrared weak target detection method of the present invention; Figure 3 This is a schematic diagram of the 3x3 differential mask CA in the contrast-based YOLO11 infrared weak target detection method of the present invention; Figure 4 This is a schematic diagram of the 3x3 differential mask CB in the contrast-based YOLO11 infrared weak target detection method of the present invention; Figure 5 This is a schematic diagram of the Attn-MLCL module in the contrast-based YOLO11 infrared weak target detection method of the present invention; Figure 6 ROC curves of the method of this invention and other existing methods on the IRST640 dataset; Figure 7 ROC curves of the method of this invention and other existing methods on the SIRST-v2 dataset; Figure 8 ROC curves of the method of this invention and other existing methods on the SIRST dataset; Figure 9 ROC curves of the method of this invention and other existing methods on the IRSTD-1K dataset; Figure 10 ROC curves of the method of this invention and other existing methods on the NCHU-SIRST dataset; Figure 11 This is a comparison of the detection results of the method of the present invention with other existing methods. Detailed Implementation

[0015] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0016] Example 1 The present invention provides a contrast-based method for detecting weak targets using the YOLOv11 infrared sensor. The specific steps are as follows: Step 1: Obtain the infrared image dataset, label the infrared images in the dataset to obtain mask images, and organize the mask images into the YOLO dataset label format. Divide the organized dataset into training set and test set. The specific process of organizing the Mask image into the YOLO dataset label format is as follows: Step 1.1: Use OpenCV's imread function to read the Mask image in grayscale mode, obtain a two-dimensional array, and get the width and height of the Mask image; Step 1.2: Find the boundary points of all white areas in the Mask image to form a closed contour; Step 1.3: Generate a minimum bounding rectangle bounding box for each closed contour in Step 1.2; Step 1.4: Normalize the coordinates of the bounding box generated in Step 1.3; Specifically, the coordinates of the bounding box generated in step 1.3 are converted into proportional values ​​relative to the image size; Step 1.5: Assign the category of each bounding box, the top-left corner coordinates of the bounding box after transformation in Step 1.4, and the width and height of the bounding box as the YOLO dataset label format; Step 2: Contrast-Yolo model. Input the training set from Step 1 into the Contrast-Yolo model for training. like Figure 1 As shown, the Contrast-Yolo model is an improvement on the YOLO11 model, consisting of a Backbone, Neck, and head. Specifically, the Backbone replaces the initial convolution in the YOLO11 model's Backbone with a DC-DC module, and adds an Attn-MLCL module after each C3k2 block (the second and third C3k2 blocks) connected to the Neck in the YOLO11 model's Backbone. The rest of the YOLO11 model's structure remains unchanged. Each convolutional layer Conv in the Backbone has a 3×3 kernel, a stride of 2, and padding of 1, ensuring that the size is halved after convolution, thus offsetting the edge loss of the 3×3 convolution. like Figure 2 As shown, the processing procedure of the DC-DC module is as follows: Step S1: Input the training set into the first regular convolution (conv) to extract features and obtain preliminary features; The kernel of the first regular convolutional conv is: The step size is 1; Step S2: The preliminary features obtained in step S1 are simultaneously divided into four parts, which are then input into the second regular convolutional module (conv), the CA_conv module, the third regular convolutional module (conv), and the CB_conv module, respectively, to obtain... , , , ; This is the output of the second regular convolution (conv). This is the output of the CA_conv module. This is the output of the third regular convolution, conv. This refers to the output of the CB_conv module; The kernels of the second and third regular convolutions are both... The step size is 1; The processing procedure of the CA_conv module is as follows: The initial features input to the CA_conv module are padded with 1s to obtain feature X. padded The 3x3 difference mask CA is expanded into a 4D convolutional kernel CA that matches the input channels. expanded Then perform a custom convolution to obtain the difference convolution result. ; like Figure 3 As shown, the 3x3 differential mask CA is as follows: the weight of the center pixel is 4, the weights of the top, bottom, left, and right neighbors are all 1, and the weights of the top left, bottom left, top right, and bottom right corners are all 0. 3x3 differential mask (CA) emphasizes the contrast relationship between lateral neighbors; The expression for performing a custom convolution is: (1); The processing procedure of the CB_conv module is the same as that of the CA_conv module, with the only difference being: Figure 4 As shown, in the 3x3 differential mask CB, the center pixel has a weight of 4, the weights of the top, bottom, left, and right neighbors are all 0, and the weights of the top left, bottom left, top right, and bottom right corners are all 1. 3x3 differential mask (CB) emphasizes the contrast relationship between vertical neighborhoods; Step S3, Calculate The dynamic weights (dynamic_weight) of the output of the CB_conv module are used to update the mask (update_mask) to obtain feature A. The dynamic weights (dynamic_weight) of the output of the CB_conv module are used to update the mask (update_mask) to obtain feature B. The process of obtaining feature A is as follows: Will The influence strength of the 3x3 differential mask CA is indirectly controlled by the parameter self.theta (which can be learned). theta is a trainable parameter nn.Parameter, which will be automatically adjusted during backpropagation to learn the weights that enhance the differential mask CA. The mask is updated through parameter range constraints. After each forward propagation, the mask parameters are forced to remain within a fixed range, with the center pixel weight limited to [-6, -2] and the neighboring pixel weights limited to [0.5, 2.0], to prevent weights from getting out of control during training, thus obtaining feature A. The process of obtaining feature B is the same as the process of obtaining feature A; Step S4: Weight the output of the second regular convolution (conv) with feature A to obtain the feature. The output of the third regular convolution (conv) is weighted with feature B to obtain the feature. , will feature With features After concatenation, features are fused using a fourth regular convolutional layer (conv) to obtain fused features. ; The kernel of the fourth regular convolution conv is: The step size is 1; (2) (3) (4) In equations (2) and (3), Indicates batch normalization; like Figure 5 As shown, the processing procedure for each Attn-MLCL module is as follows: The feature map output from the C3k2 block is simultaneously input into three branches. The outputs of the three branches are concatenated and then fed into the attention module. The attention module calculates different weights based on the importance of the features. The weights output by the attention module are weighted with the corresponding branch outputs. The weighted features are concatenated again, enabling the network to highlight important features and suppress unimportant features according to the needs of small object detection. The concatenated features are then subjected to a 1×1 regular convolution with a stride of 1 to obtain the output features of the Attn-MLCL module. The first branch consists of a 1×1 standard convolution with a stride of 1 and a 3×3 dilated convolution with a dilation rate of 1; the second branch consists of a 3×3 standard convolution with a stride of 1 and a 3×3 dilated convolution with a dilation rate of 3; and the third branch consists of a 5×5 standard convolution with a stride of 1 and a 3×3 dilated convolution with a dilation rate of 5. In the Contrast-Yolo model, the head outputs information such as the target's category, location, and confidence level. It also involves calculating the loss function and propagating it back to the backbone via the backpropagation algorithm. Specifically: Step A: Generate a target region mask based on the predicted bounding box and foreground mask. The target region mask is used to identify the target region of interest in the image. Calculate the mean and standard deviation of the target region and normalize the image data of the target region to enhance the contrast within the target region. The normalized expression is: (5) In equation (5), Indicates the input image. This represents the target region mask. Indicates the mean of the target region. Indicates the standard deviation of the target region. This represents a very small constant that prevents division by zero. , B is the batch size, i.e. the number of samples input into the network at one time; C is the number of channels, infrared images are usually single-channel; H and W are the image height and image width, respectively; b represents the image index in the batch; c represents the color channel index; x and y represent the coordinates of the image in two-dimensional space, i.e. the position of the pixel. Step B: Calculate the contrast loss; The expression is: (6) In equation (6), This represents the average intensity of all target pixels in the b-th image of the batch. This is the target area mask, where a value of 1 represents the target area and 0 represents the background. Step C: Convolve the image using Sobel kernels in the x and y directions respectively, calculate the gradient magnitude, calculate the standard deviation of the gradient in the target region, and calculate the gradient loss. The expression is: (7) (8) (9) In equations (7) to (9), S x and S y G represents the convolution kernel of the Sobel operator in the x and y directions. x and G y This is the gradient map of the image in the x and y directions; This represents the global mean of the gradient magnitude of the target region in the b-th image of the batch. Step D: Add the contrast loss and gradient loss according to their weights to obtain the contrast-gradient combined loss; The expression is: (10) In equation (10), is a fixed gradient weight coefficient; w(t) represents a time-varying weight that follows an exponential decay rule; Step E: Calculate the total loss function; The expression is: (11) In equation (11), The value of 0.1 is obtained from experience. It consists of bounding box loss, classification loss, and distribution focusing loss (i.e.) The bounding box loss (which is the sum of the classification loss and the distribution focus loss) is calculated by the model based on the prediction results and the true labels using the bounding box loss, classification loss, and distribution focus loss. It is used for backpropagation to optimize the model parameters, making the model's predictions more and more accurate. Step 3: Input the test set into the Contrast-Yolo model trained in Step 2 for detection.

[0017] Example 2 In the Contrast-Yolo model, which retains the same structure as the Yolo11 model, the Conv convolutional layer directly performs convolution operations on the image to extract basic features. The convolutional kernel is 3×3 to balance the receptive field and computational cost. The stride is 2, which is used for downsampling to halve the feature map size, such as from 640 to 320. The padding is 1 to ensure that the size is exactly halved after convolution, thus offsetting the edge loss of the 3×3 convolution. When the hyperparameter c3k is True, the C3k2 module replaces the bottleneck block with C3k, allowing for customizable convolution block sizes and more flexible extraction of features at different scales. The C3k2 module repeats multiple stages. Taking a stage with 128 input channels and 256 output channels as an example, the number of groups is 2, and the channels are split into 2 groups for convolution, reducing computation and enhancing feature diversity. The number of residual blocks is 3 to 5, increasing as the stage progresses. Features are repeatedly refined, and channels are transferred. First, 1×1 convolution reduces the number of channels, such as from 128 to 64. Then, there are "main branches (group convolution + residuals)" and "shortcut branches (direct jump connections)". Finally, the channels are merged back to 256 channels. The SPPF module (Spatial Pyramid Pooling Fast Module) fuses multi-scale feature information through pooling operations at different scales. The core parameters (taking 256 input channels and 256 output channels as an example) are: a 5×5 pooling kernel (multi-scale max pooling, which can also be fine-tuned to 3 / 7 to expand the receptive field); channel flow: 1×1 convolution to reduce channels (256→128); splicing together four scale features: "original image + 1st pooling + 2nd pooling + 3rd pooling"; and then upscaling back to 256 channels. The C2PSA module introduces PSA (Position-SensitiveAttention), which utilizes multi-head attention mechanisms and feedforward neural networks to enhance feature extraction capabilities. It can also selectively add residual structures to optimize gradient propagation and training effects, and map input features to a high-dimensional space to capture complex nonlinear relationships and learn richer feature representations. The core parameters (taking 256 input channels and 256 output channels as an example) are: 2×2 partitions, feature maps are cut into small blocks, local magnification is used to calculate attention, adapting to small infrared targets; 2 groups, channel grouping, linked with spatial partitioning, accurately capturing features; channel flow, 1×1 convolution to adjust channels, partition attention weighting, focusing on key regions; and 3×3 grouped convolution to enhance features.

[0018] Example 3 The model improvement effect is comprehensively and objectively evaluated using precision (P), recall (R), target-level F1 measure (F1), and mean precision (mAP).

[0019] The specific expression is: (12) (13) (14) (15) In equations (12) to (15), TP (TruePositive) is the true positive instance, i.e., the number of samples correctly predicted as positive by the model; FP (FalsePositive) is the false positive instance, i.e., the number of samples incorrectly predicted as positive by the model; FN (FalseNegative) is the false negative instance, i.e., the number of samples incorrectly predicted as negative by the model; AP is the average precision calculated separately for each class; N is the number of classes, AP i MAP is the average accuracy of the i-th class; MAP is often used to measure the overall performance of an object detection model across multiple classes.

[0020] Example 4 The ROC curves obtained using the method of this invention and Yolov5, Yolov6, Yolov8, Yolov9, Yolov10, and Yolov11 models on the IRST640 dataset are shown below. Figure 6 As shown, the ROC curves measured on the SIRST-v2 dataset are as follows: Figure 7 As shown, the ROC curves measured on the SIRST dataset are as follows: Figure 8 As shown, the ROC curves measured on the IRSTD-1K dataset are as follows: Figure 9 As shown, the ROC curves measured on the NCHU-SIRST dataset are as follows: Figure 10 As shown, the Contrast-Yolo model of this invention demonstrates significant performance advantages on all five datasets. The ROC curves show that the Contrast-Yolo model is more effective in handling challenging data, demonstrating its ability to solve complex scene and infrared small target detection problems.

[0021] Example 5 The Contrast-Yolo model of this invention was used to detect weak infrared targets. The results are shown in Tables 1 and 2. Figure 11 As shown.

[0022] Table 1

[0023] Table 2

[0024] As shown in Table 1, on the IRST640 dataset, R significantly improved, while P slightly decreased. However, the overall Contrast-Yolo model improved its F1 score by 0.130 and its Map50 score by 0.056. This demonstrates that the Contrast-Yolo method outperforms other models because it fully considers the complex features of object detection tasks and designs targeted optimization strategies to address the challenges. Table 2 shows that this Contrast-Yolo model generally ranks highly or leads in F1 and Map50 metrics, exhibiting excellent overall detection performance and adaptability to different datasets.

[0025] Example 6 The detection of weak infrared targets using the method of this invention and YOLOv5, YOLOv6, YOLOv8, YOLOv9, YOLOv10, and YOLOv11 models has been observed. Figure 11As shown. The model of this invention was implemented using PyTorch on a high-performance server equipped with an Intel(R) Xeon(R) Gold5218 CPU @ 2.30GHz and an NVIDIA A40 GPU. The input image was normalized to a size of 640×640. The initial weights of the model were randomly initialized to ensure unbiased learning. AdamW was chosen for training, with a momentum parameter set to 0.9 and a weight decay of 0.0005. 32 batches were used, the training protocol was extended to 300 iterations, and early stopping was set to 40.

Claims

1. A contrast-based method for detecting weak targets in YOLOv11 infrared infrared arrays, characterized in that, The specific steps are as follows: Step 1: Obtain the infrared image dataset, label the infrared images in the dataset to obtain mask images, and organize the mask images into the YOLO dataset label format. Divide the organized dataset into training set and test set. Step 2: Contrast-Yolo model. Input the training set from Step 1 into the Contrast-Yolo model for training. Step 3: Input the test set into the Contrast-Yolo model trained in Step 2 for detection.

2. The contrast-based YOLOv11 infrared weak target detection method according to claim 1, characterized in that, In step 1, the specific process of organizing the Mask image into the YOLO dataset label format is as follows: Step 1.1: Use OpenCV's imread function to read the Mask image in grayscale mode, obtain a two-dimensional array, and get the width and height of the Mask image; Step 1.2: Find the boundary points of all white areas in the Mask image to form a closed contour; Step 1.3: Generate a minimum bounding rectangle bounding box for each closed contour in Step 1.2; Step 1.4: Normalize the coordinates of the bounding box generated in Step 1.3; Step 1.5: Assign the category of each bounding box, the top-left corner coordinates of the bounding box after transformation in Step 1.4, and the width and height of the bounding box as the YOLO dataset label format.

3. The contrast-based YOLOv11 infrared weak target detection method according to claim 1, characterized in that, In step 2, the Contrast-Yolo model is an improvement on the Yolo11 model, consisting of the Backbone, Neck, and Head; Specifically, the Backbone is as follows: the initial convolution in the YOLO11 model Backbone is replaced with a DC-DC module, and an Attn-MLCL module is added after each C3k2 block connected to the YOLO11 model Neck in the YOLO11 model Backbone. The Neck and head are completely identical to those in the YOLO11 model.

4. The contrast-based YOLOv11 infrared weak target detection method according to claim 3, characterized in that, The processing procedure of the DCDC module is as follows: Step S1: Input the training set into the first regular convolution (conv) to extract features and obtain preliminary features; The kernel of the first regular convolutional conv is: The step size is 1; Step S2: The preliminary features obtained in step S1 are simultaneously divided into four parts, which are then input into the second regular convolutional module (conv), the CA_conv module, the third regular convolutional module (conv), and the CB_conv module, respectively, to obtain... , , , ; This is the output of the second regular convolution (conv). This is the output of the CA_conv module. This is the output of the third regular convolution, conv. This refers to the output of the CB_conv module; The kernels of the second and third regular convolutions are both... The step size is 1; Step S3, Calculate The dynamic weights (dynamic_weight) of the output of the CB_conv module are used to update the mask (update_mask) to obtain feature A. The dynamic weights (dynamic_weight) of the output of the CB_conv module are used to update the mask (update_mask) to obtain feature B. Step S4: Weight the output of the second regular convolution (conv) with feature A to obtain the feature. The output of the third regular convolution (conv) is weighted with feature B to obtain the feature. , will feature With features After concatenation, features are fused using a fourth regular convolutional layer (conv) to obtain fused features. ; The kernel of the fourth regular convolution conv is: The step size is 1; (2) (3) (4) In equations (2) and (3), This indicates batch normalization.

5. The contrast-based YOLOv11 infrared weak target detection method according to claim 4, characterized in that, The processing procedure of the CA_conv module is as follows: The initial features input to the CA_conv module are padded with 1s to obtain feature X. padded The 3x3 difference mask CA is expanded into a 4D convolutional kernel CA that matches the input channels. expanded Then perform a custom convolution to obtain the difference convolution result. ; The 3x3 differential mask CA is as follows: the center pixel has a weight of 4, the weights of the top, bottom, left, and right neighbors are all 1, and the weights of the top left, bottom left, top right, and bottom right corners are all 0. The expression for performing a custom convolution is: (1); The processing procedure of the CB_conv module is the same as that of the CA_conv module. In the CB_conv module, the 3x3 difference mask CB has a center pixel weight of 4, the weights of the top, bottom, left, and right neighbors are all 0, and the weights of the top left, bottom left, top right, and bottom right corners are all 1.

6. The contrast-based YOLOv11 infrared weak target detection method according to claim 3, characterized in that, The processing procedure for each Attn-MLCL module is as follows: The feature map output from the C3k2 block is simultaneously input into the three branches. The outputs of the three branches are concatenated and then fed into the attention module. The attention module calculates different weights based on the importance of the features. The weights output by the attention module are weighted with the corresponding branch outputs. The weighted features are concatenated again. The concatenated features are then subjected to a 1×1 regular convolution with a stride of 1 to obtain the output features of the Attn-MLCL module.

7. The contrast-based YOLOv11 infrared weak target detection method according to claim 6, characterized in that, The first branch consists of a 1×1 standard convolution with a stride of 1 and a 3×3 dilated convolution with a dilation rate of 1; the second branch consists of a 3×3 standard convolution with a stride of 1 and a 3×3 dilated convolution with a dilation rate of 3; and the third branch consists of a 5×5 standard convolution with a stride of 1 and a 3×3 dilated convolution with a dilation rate of 5.

8. The contrast-based YOLOv11 infrared weak target detection method according to claim 1, characterized in that, The loss function during training is: Step A: Generate a target region mask based on the predicted bounding box and foreground mask. The target region mask is used to identify the target region of interest in the image. Calculate the mean and standard deviation of the target region and normalize the image data of the target region. The normalized expression is: (5) In equation (5), Indicates the input image. This represents the target region mask. Indicates the mean of the target region. Indicates the standard deviation of the target region. This represents a very small constant that prevents division by zero. , B is the batch size, i.e. the number of samples input into the network at one time; C is the number of channels, infrared images are usually single-channel; H and W are the image height and image width, respectively; b represents the image index in the batch; c represents the color channel index; and x and y represent the coordinates of the image in two-dimensional space. Step B: Calculate the contrast loss; The expression is: (6) In equation (6), This represents the average intensity of all target pixels in the b-th image of the batch. This is the target area mask, where a value of 1 represents the target area and 0 represents the background. Step C: Calculate the gradient loss; The expression is: (7) (8) (9) In equations (7) to (9), S x and S y G represents the convolution kernel of the Sobel operator in the x and y directions. x and G y This is the gradient map of the image in the x and y directions; Indicates the gradient magnitude of the target region; This represents the global mean of the gradient magnitude of the target region in the b-th image of the batch. Step D: Add the contrast loss and gradient loss according to their weights to obtain the contrast-gradient combined loss; The expression is: (10) In equation (10), The gradient weights are fixed; w(t) represents the time-varying weights. Step E: Calculate the total loss function; The expression is: (11) In equation (11), The value is 0.

1. It consists of bounding box loss, classification loss, and distribution focusing loss.