Infrared small sample target detection method based on domain self-adaption

By adopting a cross-domain visual optimization model and an improved YOLOv8 object detection network in infrared object detection, the problem of lack of visual details and semantic information in infrared images is solved, and the detection accuracy and model adaptability are improved.

CN120148015APending Publication Date: 2025-06-13XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510210851.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Infrared images often lack rich visual details and semantic information, making it more difficult to extract effective features in deep learning models, and existing cross-domain algorithms have lower accuracy in infrared object detection.

Method used

The infrared small sample object detection method based on domain adaptation is adopted to convert visible light images into infrared-like images through cross-domain visual optimization model, and the YOLOv8 object detection network is improved, and the DCNv4_C2f module and multi-pooling hybrid attention mechanism are introduced to improve the accuracy of feature extraction and object detection.

Benefits of technology

It effectively overcomes the problem of difficulty in infrared image feature extraction, improves the accuracy of infrared object detection, enhances the model's adaptability to complex scenes, and reduces the workload of data labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148015A_ABST
    Figure CN120148015A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared small sample target detection method based on domain self-adaption, and solves the problems that an infrared image generally lacks rich visual details and semantic information, so that effective features are more difficult to extract in a deep learning model, and an existing cross-domain algorithm is lower in accuracy in infrared target detection. And realizing style migration from the visible light image to the infrared image through a cross-domain visual optimization model. A YOLOv8 target detection network is improved, a DCNv4C2f module is used for replacing a specific C2f layer in a backbone network, and deformable convolution is introduced, so that the features of objects with complex shapes and appearances can be captured more effectively. A new multi-pooling mixed attention mechanism is designed and introduced into a target detection network, cross-channel information can be effectively captured, direction information and position information can also be obtained, the detection capability of the multi-pooling mixed attention mechanism on a fuzzy target and a dense target in an infrared image is improved, and the problem of missing detection is effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technical method for image target detection in computer vision, and particularly to an infrared small-sample target detection method based on domain adaptation. Background Art

[0002] Infrared target detection is an important task in the field of computer vision, aiming to identify and locate target objects in infrared images. Different from visible light images, infrared images mainly reflect the thermal radiation information of objects and usually have better adaptability in low-light or harsh environments, such as at night, in smoke, haze, etc., and can clearly show the outline or temperature distribution of objects. At present, this field is still in a stage of rapid development, and the research on this task is not sufficient enough, especially the infrared image target detection algorithm under small samples. Exploring a reliable cross-domain small-sample target detection algorithm is expected to solve the current development bottleneck of the infrared small-sample target detection algorithm. Using this algorithm, the workload of annotating the infrared dataset can be greatly reduced to maximize the utilization of limited data, improve the training efficiency of the model, and enhance the adaptability of the system to complex scenarios.

[0003] In recent years, small-sample target detection algorithms combined with transfer learning training strategies have made significant progress. By transferring the knowledge of the base classes, these algorithms can effectively transfer the existing pre-trained models to small-sample classes, thus achieving accurate prediction and classification of small-sample targets. This method enables the effective improvement of the performance of target detection under limited sample sizes and solves the limitations of traditional target detection methods under small-sample conditions. For example, a small-sample target detection method based on the attention mechanism and transfer learning (publication number: CN118298147A) also adopts a transfer learning method, but it only uses the transfer learning fine-tuning method and does not consider the domain shift problem between the target domain and the source domain.

[0004] Since infrared images usually lack rich visual details and semantic information, it becomes more difficult to extract effective features in deep learning models. Summary of the Invention

[0005] The purpose of the present invention is to solve the problems that infrared images usually lack rich visual details and semantic information, resulting in more difficult extraction of effective features in deep learning models, and the low accuracy of existing cross-domain algorithms in infrared target detection, and to provide an infrared small-sample target detection method based on domain adaptation.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] The present invention provides an infrared small-sample target detection method based on domain adaptation, which is characterized in that it includes the following steps:

[0008] Step 1: Obtain a large-scale labeled visible light image and a small sample of unlabeled infrared images;

[0009] Step 2: Use a cross-domain visual optimization model to transform the visible light image into an infrared-like image as the source domain dataset, and use the unlabeled infrared image as the target domain dataset;

[0010] Step 3: Build an improved YOLOv8 object detection network, including a backbone network, a neck network, a detection head network, and a loss function module, and establish a multi-scale domain adaptation module;

[0011] Step 4: Input the source domain dataset and the target domain dataset into the backbone network of the improved YOLOv8 object detection network for feature extraction to obtain multi-scale image features;

[0012] Step 5: Input the multi-scale image features into the multi-scale domain adaptation module, predict the domain category to which each feature map belongs, and calculate the domain classification loss according to the prediction results;

[0013] Step 6: Input the multi-scale image features of the source domain dataset into the neck network and the detection head network of the improved YOLOv8 object detection network for object detection, and calculate the detection loss according to the detection results;

[0014] Step 7: Calculate the total loss of the improved YOLOv8 object detection network according to the detection loss and the domain classification loss, and optimize the parameters of the improved YOLOv8 object detection network according to the total loss;

[0015] Step 8: Detect the infrared image of the target to be detected through the optimized improved YOLOv8 object detection network to obtain the detection results.

[0016] Furthermore, in Step 2, the cross-domain visual optimization model includes a channel enhancement module and a cyclic generative adversarial network;

[0017] Step 2 is specifically as follows:

[0018] Step 2.1: Input the visible light image into the channel enhancement module to convert it into an HSV image, extract the V channel and then copy the V channel to generate a three-channel image x, and the expression is:

[0019] V = max(R, G, B);

[0020] x = (V, V, V);

[0021] Step 2.2: Use the cyclic generative adversarial network to convert the three-channel image x into an infrared-like image in the infrared domain.

[0022] Further, the improved YOLOv8 object detection network described in step 3 is an improvement on the YOLOv8 object detection network, and the specific improvements are as follows:

[0023] Replace the 2nd, 3rd, and 4th C2f modules in the backbone network of the original YOLOv8 object detection network with DCNv4_C2f modules;

[0024] Introduce a new multi-pooling hybrid attention mechanism in the 10th layer of the backbone network of the YOLOv8 object detection network, including a channel attention mechanism and a position attention mechanism, and add a max-pooling branch and a standard deviation pooling branch to the original attention mechanism;

[0025] At the same time, introduce multiple fully connected layer branches of the same size, which are used to calculate the attention weights of global average pooling, max-pooling, and standard deviation pooling respectively.

[0026] Further, the DCNv4_C2f module removes Softmax on the basis of the existing DCNv3 module, and adopts an adaptive aggregation window and dynamic aggregation weights with an unbounded value range;

[0027] The calculation formula is:

[0028]

[0029] y = concat([y(P 1 ), y(P 2 ),, y(P D )], axis=-1);

[0030] In the formula, P d represents the center sampling point position, y(P d ) represents the output of the standard convolution at the center sampling point position, D represents the total number of aggregation groups, N represents the total number of sampling points, n represents the enumerated sampling point, P n represents the enumerated sampling point position, m dn represents the modulation amount of the nth sampling point in the dth group, X d represents the sliced feature map, ΔP dn is the offset corresponding to the nth sampling point of the grid sampling in the dth group. After obtaining the output y(P d ) of each group, the outputs of these groups are concatenated in the channel dimension to form the final output feature map y.

[0031] Further, the multi-scale domain adaptation module described in step 5 includes a gradient reversal layer, 2 convolutional layers, a domain classifier, and a domain classification loss function. In the training stage, the multi-scale domain adaptation module receives feature maps of three scales as inputs, and these feature maps are processed by convolutional layers to predict the domain category to which each feature map belongs.

[0032] Furthermore, in order to enable the network to learn shared features between the source domain and the target domain, the multi-scale domain adaptation module uses the binary cross-entropy loss function to calculate the domain classification loss L da , specifically, this loss function compares the true domain label t i (where t i = 1 represents the source domain and t i = 0 represents the target domain) and the domain class probability of the model prediction at (x, y) on the i-th image The network is optimized by minimizing the loss so that it can obtain an invariant feature representation between different domains. The domain classification loss function is:

[0033]

[0034] Furthermore, in step 7, the calculation of the detection loss according to the detection result is specifically: inputting the detection result into the loss function module of the improved YOLOv8 object detection network to obtain the detection loss.

[0035] Furthermore, in step 8, the calculation of the total loss of the improved YOLOv8 object detection network according to the detection loss and the domain classification loss, the calculation formula is:

[0036] L = L det + αL da ;

[0037] where L represents the total loss, L det represents the detection loss, L da represents the domain classification loss, and α represents a hyperparameter.

[0038] Furthermore, in step 8, the optimization of the improved YOLOv8 object detection network according to the total loss is specifically: determining whether the total loss is less than the loss threshold. If it is less, the optimized improved YOLOv8 object detection network is output. Otherwise, the parameters in the improved YOLOv8 object detection network are modified and step 4 is returned.

[0039] Advantages of the present invention:

[0040] (1) The infrared small-sample object detection method based on domain adaptation of the present invention realizes the style transfer from visible light images to infrared images through a cross-domain vision optimization model. The channel enhancement method is easy to implement and further assists the cyclic generative adversarial network.

[0041] (2) In a method for infrared small-sample target detection based on domain adaptation of the present invention, the YOLOv8 target detection network is improved, and the specific C2f layer in the backbone network is replaced with the DCNv4_C2f module. This module introduces deformable convolution and can more effectively capture the features of objects with complex shapes and appearances.

[0042] (3) In a method for infrared small-sample target detection based on domain adaptation of the present invention, a multi-pooling hybrid attention mechanism is added, which can not only effectively capture cross-channel information, but also obtain direction information and position information, thereby enhancing the target localization ability of the model, improving its detection ability for blurred and dense targets in infrared images, and effectively avoiding missed detection problems. Description of the Drawings

[0043] Figure 1 is a flowchart of an embodiment of a method for infrared small-sample target detection based on domain adaptation of the present invention;

[0044] Figure 2 is a schematic diagram of the principle of a cross-domain visual optimization model in an embodiment of a method for infrared small-sample target detection based on domain adaptation of the present invention;

[0045] Figure 3 is a schematic diagram of the improved YOLOv8 target detection network structure in an embodiment of a method for infrared small-sample target detection based on domain adaptation of the present invention;

[0046] Figure 4 is a schematic diagram of the DCNv4_C2f module structure in an embodiment of a method for infrared small-sample target detection based on domain adaptation of the present invention;

[0047] Figure 5 is a schematic diagram of the multi-pooling hybrid attention mechanism structure in an embodiment of a method for infrared small-sample target detection based on domain adaptation of the present invention;

[0048] Figure 6 is a schematic diagram of the multi-scale domain adaptation module structure in an embodiment of a method for infrared small-sample target detection based on domain adaptation of the present invention. Detailed Embodiments

[0049] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the drawings and embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] This embodiment provides a method for infrared small-sample target detection based on domain adaptation, as Figure 1As shown in the figure, it includes the following steps:

[0051] Step 1: Obtain labeled visible light images and unlabeled infrared images;

[0052] In this embodiment, the visible light images are images extracted from the high-quality (1080P+) driving scene dataset SODA10M, and 8000 labeled dataset images are selected; the target domain infrared images are taken from the CTIR infrared dataset, and 2000 unlabeled image data are selected from it.

[0053] Step 2: Use a cross-domain visual optimization model to transform the visible light images into infrared-like images to form a source domain dataset; among them, the cross-domain visual optimization model includes a channel enhancement module and a cyclic generative adversarial network. The specific process is as follows:

[0054] Step 2.1: Input the visible light images into the channel enhancement module to convert them into HSV images, extract the V channel and then copy the V channel to generate a three-channel image (intermediate modality) x. The expression is:

[0055] V = max(R, G, B) (1)

[0056] x = (V, V, V) (2)

[0057] Step 2.2: Use the cyclic generative adversarial network to transform the three-channel image x into an infrared-like image in the infrared domain

[0058] In this embodiment, the cyclic generative adversarial network is as Figure 2 shown. The cyclic adversarial network model includes two generators and two discriminators. The generators are used to transform and generate images, and the discriminators are used to judge and feedback on the images. The corresponding image datasets are the intermediate modality in the X domain and the infrared images in the Y domain respectively.

[0059] Train the generator G so that the intermediate modality x in the X domain generates new infrared images through the generator G That is Then through the discriminator D Y Distinguish between real images and generated images; train the generator F so that the infrared image y in the Y domain generates new infrared images through the generator F That is Then through the discriminator D X Distinguish.

[0060] During the training process of the cyclic generative adversarial network model, update the model parameters according to the loss function. The loss function includes formulas (3), (4), (5), and (6):

[0061]

[0062] L(G,F,D X ,D Y ) = L GANl (G,D Y ,X,Y) + L GAN2 (F,D X ,Y,X) + L cyc (G,F) (6)

[0063] Among them, in formula (3), L GANl represents the adversarial loss of discriminator D y In formula (4), L GAN2 represents the adversarial loss of discriminator D x . represents the loss of the generated sample y in the Y domain, represents the loss of the generated sample x in the X domain, P data(y) and P data(x) respectively represent the distribution functions of the real data in the Y domain and the X domain. In formula (5), L cyc represents the cycle consistency loss. In formula (6), L represents the total loss of CycleGAN.

[0064] Step 3: Build an improved YOLOv8 object detection network, as Figure 3 shown, including a backbone network, a neck network, a detection head network, and a loss function module, and establish a multi-scale domain adaptation module, which is set at the output end of the backbone network feature extraction module;

[0065] In this embodiment, the improved YOLOv8 object detection network is specifically improved as follows:

[0066] Replace the 2nd, 3rd, and 4th C2f modules in the backbone network of the original YOLOv8 object detection network with DCNv4_C2f modules; The DCNv4_C2f module is as Figure 4 shown, and replace the ordinary convolutional layer of the Bottleneck module in the C2f module with a deformable convolution.

[0067] The DCNv4_C2f module removes Softmax on the basis of the existing DCNv3 module, and adopts an adaptive aggregation window and dynamic aggregation weights with an unbounded value range;

[0068] The calculation formula is:

[0069]

[0070] y = concat([y(P 1 ), y(P 2 ),, y(P D )], axis = -1);

[0071] In the formula, P d represents the position of the central sampling point, y(P d ) represents the output of the standard convolution at the position of the central sampling point, D represents the total number of aggregation groups, N represents the total number of sampling points, n represents the enumerated sampling point, P n represents the position of the enumerated sampling point, m dn represents the modulation amount of the nth sampling point in the dth group, X d represents the slice feature map, ΔP dn is the offset corresponding to the nth sampling point of the grid sampling in the dth group. After obtaining the output y(P d ) of each group, the outputs of these groups are concatenated in the channel dimension to form the final output feature map y.

[0072] In this embodiment, the DCNv4_C2f module improves the deficiencies of the regular convolution in terms of long-distance dependence and adaptive spatial aggregation, making the training data less and the model more effective.

[0073] In the backbone network of the YOLOv8 object detection network, a new multi-pooling hybrid attention mechanism is introduced at the 10th layer, as Figure 5 shown; on the basis of the original SENet framework, a max-pooling branch and a standard deviation pooling branch are introduced. The aim is to make up for the deficiencies of the network in capturing local features. Through this max-pooling branch, the network can pay more attention to the local features of small targets in the infrared image, effectively alleviating the problem of small target information loss caused by global average pooling. The standard deviation pooling branch further improves the model's perception ability of the target position and structure by capturing the degree of change in different regions of the feature map.

[0074] In this structure, first, channel-level feature compression is performed on the input feature x f . Then, through the average pooling, max-pooling, and standard deviation pooling branches respectively, three different feature representations are obtained, and these features are the global average pooling result x Avg , the max-pooling result x Max , and the standard deviation pooling result x Std .

[0075] Based on the above attention mechanism, multiple fully connected layer branches of the same size are also introduced, enabling the model to perform multiple feature extractions in parallel at the same level, so as to better capture information of different scales and different types to perform the excitation operation. Specifically, x Avg , x Max , and x StdAfter the compression part, that is, the Sq operation in Formula 9 passes through the fully connected layers FC1, FC2, FC3, and FC4, and then the attention weights of global average pooling, max pooling, and standard deviation pooling are calculated respectively through Formula (8), namely x a , x m , x s . Then, the three groups of weights are added together and passed through the Sigmoid activation function to obtain the final channel attention weight w c . Finally, the obtained channel attention weight is multiplied element-wise with the input feature map, thereby restoring the dimension of the original input and generating the channel attention feature map Output1. In Formula (8), Concat is the concatenation operation:

[0076]

[0077] w c = Sigmoid(x a + x m + x s ) (10)

[0078]

[0079] In this embodiment, an attention module for extracting position information is added after the channel attention to enhance feature aggregation, which can effectively capture the position information of the target and improve the detection ability. First, average pooling of Output1 in the horizontal and vertical directions is performed to obtain x' Avg , y' Avg to extract the global information along the x and y spatial dimensions. Then, the results of these two groups of pooling are concatenated to form a composite feature representation. Next, a 1×1 convolution operation K 1×1 is used to further capture the local cross-channel interaction information. Next, the split operation (Split) is used to split the convolution output into two independent tensors along the spatial dimension and send them into the Sigmoid activation function respectively to calculate the position attention weights w x , w y on two dimensions (x and y). Finally, these position attention weights are applied to Output1 to achieve weighted fusion and obtain the final output Output of the hybrid attention mechanism. The specific process of the position attention mechanism is shown in Formulas (12) to (14):

[0080] x conv = K 1×1 (Concat(x' Avg , y' Avg )) (12);

[0081] w x , wy = Sigmoid(Split(x conv )) (13);

[0082]

[0083] In this embodiment, as Figure 6 shown, the multi-scale domain adaptation module is set after the feature extraction module of the improved YOLOv8 object detection network; the feature extraction module of the improved YOLOv8 object detection network obtains three features F1, F2, F3. For each input feature, it first needs to pass through a Gradient Reversal Layer (GRL), and the GRL can optimize the parameters of both the domain classifier and the basic network simultaneously.

[0084] The multi-scale domain adaptation module includes a gradient reversal layer, two convolutional layers, a domain classifier, and a domain classification loss function.

[0085] The domain classification loss function is:

[0086]

[0087] In the formula, L da represents the domain classification loss, t i represents the true label of the i-th training image, where t i = 1 represents the source domain, and t i = 0 represents the target domain. represents the feature map of the i-th image at the position (x, y), that is, the probability of the domain classifier on the input.

[0088] Step 4: Input the source domain dataset and the target domain dataset into the backbone network of the improved YOLOv8 object detection network for feature extraction.

[0089] In this embodiment, the feature extraction is specifically as follows: set the number of training times to 200 and the initial learning rate to 0.01;

[0090] Input the source domain dataset and the target domain dataset into the improved YOLOv8 object detection network for feature extraction to obtain three feature layers F1, F2, F3.

[0091] Step 5: Input the three feature layers F1, F2, F3 into the multi-scale domain adaptation module. After passing through its core Gradient Reversal Layer (GRL), perform two convolutional operations, and then classify each feature point of each feature map through the domain classifier to determine whether it comes from the source domain dataset or the target domain dataset, and calculate the domain classification loss in combination with the domain classification loss function.

[0092] In this embodiment, when passing through the gradient reversal layer, the gradient direction is automatically reversed during the backpropagation process, and the identity transformation is achieved during the forward propagation process. Cross-domain distribution alignment is achieved through adversarial learning, and the process is described as in Equation (16):

[0093]

[0094] Among them, f b is the feature extraction module of the object detection model, that is, the YOLOv8 backbone network. S and T are the source domain and the target domain respectively, and x f is the feature vector after feature extraction. The goal of f b is to learn a feature to make the data in the target domain T and the source domain S as similar as possible in the feature space, that is, to minimize their distribution difference d H , and the goal of the classifier h() is to correctly distinguish whether the input comes from the source domain or the target domain. err s , err t represents the classification error of the classifier.

[0095] Step 6: Input the multi-scale features of the labeled source domain images into the neck network and the detection head network of the improved YOLOv8 object detection network for object detection, and input the detection results into the loss function module of the improved YOLOv8 object detection network to obtain the detection loss;

[0096] In this embodiment, the detection loss includes the classification loss and the bounding box regression loss, and the bounding box regression loss includes the distribution focal loss and the intersection over union loss.

[0097] Step 7: Calculate the total loss of the improved YOLOv8 object detection network according to the detection loss and the domain classification loss. The calculation formula is:

[0098] L = L det + αL da = L cls + λ 1 L CIOU + λ 2 L DFL + αL da (17)

[0099] Among them, L represents the total loss, L det represents the detection loss, L cls represents the classification loss, L CIOU represents the intersection over union loss, L DFL represents the distribution focal loss, L da represents the domain classification loss, α, λ 1 and λ 2 are hyperparameters.

[0100] Optimizing the improved YOLOv8 object detection network according to the total loss is specifically as follows: Determine whether the total loss is less than the loss threshold. If it is less, output the optimized improved YOLOv8 object detection network. Otherwise, modify the parameters in the improved YOLOv8 object detection network and return to step 4.

[0101] Step 8: Detect the infrared image of the target to be detected through the optimized improved YOLOv8 object detection network to obtain the detection result.

[0102] A method for infrared small-sample object detection based on domain adaptation in the present invention realizes style transfer from visible light images to infrared images through a cross-domain visual optimization model. The channel enhancement method is easy to implement and further assists the cyclic generative adversarial network. Thus, the problem that it is more difficult to extract effective features in the deep learning model due to the lack of rich visual details and semantic information is overcome. In addition, in the present invention, the YOLOv8 object detection network is improved by replacing a specific C2f layer in the backbone network with a DCNv4_C2f module. This module introduces deformable convolutions and can more effectively capture the features of objects with complex shapes and appearances. And a multi-pooling hybrid attention mechanism is added, which can not only effectively capture cross-channel information, but also obtain direction information and position information, thereby enhancing the object localization ability of the model, improving its detection ability for blurred and dense objects in infrared images, effectively avoiding missed detection problems, and improving the detection accuracy.

Claims

1. A domain adaptive infrared small sample target detection method, characterized in that: The following steps are involved: Step 1: Acquire large-scale labeled visible light images and small sample unlabeled infrared images; Step 2: Use a cross-domain visual optimization model to transform visible light images into infrared-like images as the source domain dataset, and use unlabeled infrared images as the target domain dataset; Step 3: Build an improved YOLOv8 target detection network, including the backbone network, neck network, detection head network and loss function module connected in sequence, and establish a multi-scale domain adaptation module at the output end of the backbone network feature extraction module; Step 4: Input the source domain dataset and the target domain dataset into the backbone network of the improved YOLOv8 target detection network for feature extraction to obtain multi-scale image features; Step 5: Input the multi-scale features of the image into the multi-scale domain adaptation module, predict the domain category to which each feature map belongs, and calculate the domain classification loss based on the prediction results; Step 6: Input the multi-scale features of the image of the source domain dataset into the neck network and the detection head network of the improved YOLOv8 target detection network to perform target detection, and calculate the detection loss based on the detection results; Step 7: Calculate the total loss of the improved YOLOv8 target detection network according to the detection loss and the domain classification loss, and optimize the parameters of the improved YOLOv8 target detection network according to the total loss; Step 8: Use the optimized improved YOLOv8 target detection network to detect the infrared image of the target to be detected and obtain the detection result.

2. According to claim 1, a domain adaptive infrared small sample target detection method is characterized in that: In step 2, the cross-domain visual optimization model includes a channel enhancement module and a recurrent generative adversarial network; Step 2 is as follows: Step 2.1: Input the visible light image into the channel enhancement module to convert it into an HSV image, extract the V channel and then copy the V channel to generate a three-channel image x, the expression is: V = max(R, G, B); x = (V, V, V); Step 2.2: Convert the three-channel image x into an infrared-like image in the infrared domain through a cyclic generative adversarial network.

3. The infrared small sample target detection method based on domain adaptation according to claim 1, characterized in that: The improved YOLOv8 target detection network described in step 3 is an improvement of the YOLOv8 target detection network, and the specific improvements are as follows: Replace the 2nd, 3rd, and 4th C2f modules in the backbone network of the original YOLOv8 target detection network with DCNv4_C2f modules; In the 10th layer of the backbone network of the YOLOv8 target detection network, a new multi-pooling hybrid attention mechanism is introduced, including the channel attention mechanism and the position attention mechanism, and the maximum pooling branch and the standard deviation pooling branch are added to the original attention mechanism. At the same time, multiple fully connected layer branches of the same size are added to calculate the attention weights of the global average pooling, maximum pooling and standard deviation pooling respectively.

4. According to claim 3, a domain adaptive infrared small sample target detection method is characterized in that: The DCNv4_C2f module removes Softmax from the existing DCNv3 module and adopts an adaptive aggregation window and a dynamic aggregation weight with an unbounded value range; The calculation formula is: and=concat([y(P1),y(P2),…,y(P D )],axis=-1); Where P d Indicates the position of the central sampling point, y(P d ) represents the output of the standard convolution at the center sampling point, D represents the total number of aggregation groups, N represents the total number of sampling points, n represents the enumerated sampling point, P n Indicates the enumeration of sampling point locations, m dn represents the modulation amount of the nth sampling point in the dth group, X d Represents the slice feature map, ΔP dn is the offset corresponding to the nth sampling point of the grid sample in the dth group.

5. The infrared small sample target detection method based on domain adaptation according to claim 1, characterized in that: The multi-scale domain adaptation module in step 5 includes a gradient reversal layer, two convolutional layers, a domain classifier, and a domain classification loss function connected in sequence.

6. The infrared small sample target detection method based on domain adaptation according to claim 5, characterized in that: The domain classification loss function is: Where, L da represents the domain classification loss, t i represents the true label of the i-th training image, where t i =1 indicates the source domain, t i =0 indicates the target domain, Represents the feature map of the ith image at position (x, y).

7. The infrared small sample target detection method based on domain adaptation according to claim 1, characterized in that: In step 6, the detection loss is calculated according to the detection result as follows: the detection result is input into the loss function module of the improved YOLOv8 target detection network to obtain the detection loss.

8. The method for infrared small sample target detection based on domain adaptation according to claim 7, characterized in that: In step 7, the total loss of the improved YOLOv8 target detection network is calculated based on the detection loss and the domain classification loss, and the calculation formula is: L=L det +αL da ; Among them, L represents the total loss, L det represents the detection loss, L da represents the domain classification loss and α represents a hyperparameter.

9. The infrared small sample target detection method based on domain adaptation according to claim 1, characterized in that: In step 7, the optimization of the improved YOLOv8 target detection network according to the total loss is specifically as follows: determining whether the total loss is less than the loss threshold, if so, outputting the optimized improved YOLOv8 target detection network, otherwise, modifying the parameters in the improved YOLOv8 target detection network, and returning to step 4.

Citation Information

Patent Citations

  • Small sample target detection method based on attention mechanism and transfer learning

    CN118298147A

Cited By

  • Language guidance feature decoupling infrared target detection method

    CN121330249A

  • Generative adversarial network-based distribution network defect sample enhancement and small sample high-precision identification method and system

    CN122135104A

  • Distribution defect sample enhancement and small sample high-precision identification method and system based on generative adversarial network

    CN122135104B