Complex scene infrared weak target detection method and system fused with multi-scale comparison features

By adding multiple attention perception modules and convolution kernel clustering and pruning technology to the YOLO model, a lightweight model is generated, which solves the accuracy and robustness of infrared weak target detection under complex backgrounds, and achieves efficient and real-time object detection.

CN120526271AInactive Publication Date: 2025-08-22JILIN VOCATIONAL COLLEGE OF IND & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510633477.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Under complex backgrounds, weak infrared targets are difficult to be accurately detected. The prior art is low in detection accuracy, high false alarm rate, and poor real-time performance when dealing with complex backgrounds, low signal-to-noise ratios, etc., and is difficult to effectively detect in night or inclement weather conditions.

Method used

Added multiple attention perception modules to the YOLO model, including channel and spatial attention mechanism submodules, and combined with convolutional kernel clustering and pruning technology, a lightweight infrared weak object detection model is generated, and the detection accuracy and robustness are improved through multi-scale comparison feature extraction and fusion.

Benefits of technology

It realizes high-precision and low error detection rate detection of weak infrared targets in complex backgrounds, reduces missed detection rates and false alarm rates, can run in real time on resource-constrained devices, and broadens application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526271A_ABST
    Figure CN120526271A_ABST
Patent Text Reader

Abstract

The invention relates to a complex scene infrared weak target detection method and system fused with multi-scale contrast features. The method comprises the following steps: acquiring an infrared image in a target area under a complex background; a multiple attention perception module is added in the YOLO model and used for self-adaptive adjustment in different dimensions to obtain feature information; the multiple attention perception module comprises a channel attention mechanism sub-module and a space attention mechanism sub-module; training a YOLO model containing multiple attention perception modules by using the training set, extracting convolution kernels of a convolution layer in the YOLO model in the training process, and performing lightweight processing on the convolution layer according to a convolution kernel group to obtain an infrared weak target detection model; and inputting the infrared image into an infrared weak target detection model to obtain a target detection result. According to the method, the defect of single-modal feature description is overcome, the detection performance of the infrared weak and small target under the complex background is improved, and the robustness of a small target detection algorithm in different scenes is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of infrared weak target detection, and in particular to a method and system for detecting infrared weak targets in complex scenes by integrating multi-scale contrast features. Background Art

[0002] Due to their small size and unclear features, small infrared targets are easily submerged in complex and changing background clutter and are also subject to noise interference. Therefore, the detection of small infrared targets under complex backgrounds is a research hotspot and difficulty in the field of target detection. Common methods in the field of small target detection mainly target visible light images, with fewer targeting infrared images. Small infrared targets do not contain color information, have a large scale difference from conventional targets, and are more dependent on contextual information. Traditional detection methods have difficulty in effectively identifying these targets. Although some progress has been made, many problems still need to be solved. Therefore, how to improve the detection capability of small infrared targets under complex background conditions, reduce the false alarm rate, and achieve real-time and high-precision detection of small infrared targets has important academic research significance and engineering application value.

[0003] Currently, existing technical solutions propose combining multi-scale filtering and image enhancement techniques to significantly improve the robustness and accuracy of the algorithm under noisy conditions. Existing technical solutions also optimize the performance of the detection algorithm based on comprehensive consideration of multi-scale information. Existing technical solutions also utilize generative adversarial networks (GANs) for synthetic data enhancement to improve the performance of infrared small target detection. By increasing data diversity, the model's generalization ability and detection accuracy in different scenarios are improved. Existing technical solutions also employ lightweight convolutional neural networks and deep learning methods to achieve more efficient extraction and detection of target features. Existing technical solutions also combine global and local information based on an improved global search for density peaks and the local contrast mechanism of human vision, effectively improving the algorithm's detection performance and robustness in complex scenes. In recent years, infrared small target detection technology has attracted widespread attention in China. The existing technical solution also proposes a lightweight network structure that integrates multiple heterogeneous filters, which improves the speed and accuracy of detection; the existing technical solution also proposes an improved detection algorithm for complex backgrounds, which improves the detection accuracy in complex environments; the existing technical solution also proposes a method that combines low-rank representation and reweighted sparse representation, which effectively improves the small target detection effect; the existing technical solution also proposes a detection algorithm that integrates multi-scale attention mechanism and separate decoupling head, which improves the detection accuracy and efficiency; the existing technical solution also proposes an infrared dim target detection method based on enhanced local contrast, which improves the accuracy and robustness of target detection; the existing technical solution also proposes to enhance the detection performance of infrared dim targets by utilizing prior knowledge of target position, thereby improving the generalization ability of the network; the existing technical solution also proposes an infrared dim target detection method based on global attention network and multi-scale feature fusion, which improves the accuracy and robustness of detection through global attention mechanism and multi-scale feature fusion technology. These research results have made significant progress in improving detection accuracy, real-time performance, and robustness. However, detecting small, weak infrared targets in complex scenes still faces numerous challenges. These results primarily rely on manual feature extraction, such as target grayscale, texture, and shape. While these methods are effective against simple backgrounds, their performance is often limited in complex, interfering environments. In recent years, the rise of deep learning technology has brought new opportunities for infrared target detection. However, existing deep learning-based detection methods still suffer from low detection accuracy, high false alarm rates, and poor real-time performance when dealing with complex backgrounds and low signal-to-noise ratios. This is particularly true at night or in adverse weather conditions, where small, weak infrared targets can be easily lost in complex and changing background clutter, making accurate detection difficult. Summary of the Invention

[0004] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a method and system for detecting infrared weak targets in complex scenes that integrates multi-scale contrast features. In view of the problems that infrared weak targets are difficult to distinguish from the background under complex background interference, low signal-to-noise ratio caused by noise interference, and poor detection accuracy, research is carried out from the aspects of feature extraction capability, anti-interference capability, contrast difference and spatiotemporal feature extraction to make up for the shortcomings of single modal feature description, improve the detection performance of infrared weak targets in complex backgrounds, and enhance the robustness of small target detection algorithms in different scenarios.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A method for detecting infrared weak targets in complex scenes by integrating multi-scale contrast features, including:

[0007] Acquire infrared images of target areas against complex backgrounds;

[0008] Add a multi-attention perception module to the YOLO model for adaptive adjustment in different dimensions to obtain feature information; the multi-attention perception module includes: a channel attention mechanism submodule and a spatial attention mechanism submodule;

[0009] The YOLO model including multiple attention perception modules is trained using the training set. During the training process, the convolution kernels of the convolution layer in the YOLO model are extracted. The convolution layer is lightweight processed according to the convolution kernel group to obtain an infrared weak target detection model.

[0010] The infrared image is input into the infrared weak target detection model to obtain a target detection result.

[0011] Optionally, obtaining the characteristic information includes:

[0012] The infrared image is preprocessed by using the input end of the infrared weak target detection model, and the preprocessed image is input into the backbone network. The preprocessed image is sliced ​​horizontally and vertically and then spliced ​​through the Focus module of the backbone network; the spliced ​​output is sequentially passed through the first convolution module, the first C3 module, the second convolution module, the second C3 module, the third convolution module, the third C3 module, and the fourth convolution module to obtain a feature map; the feature map is input into the SPPF module for maximum pooling, and the feature map of any size is converted into a first feature vector of a fixed size; the first feature vector is passed through the fourth C3 module to output a second feature vector; the second feature vector is sequentially input into the channel attention mechanism submodule and the spatial attention mechanism submodule of the multiple attention perception module to mark key feature information, which includes key detection targets and the positioning information of key detection targets, and obtains low-level features with details and positioning information and high-level features with semantic information;

[0013] Among them, the C3 module is used to perform target convolution learning on the residual features output by the convolution module.

[0014] Optionally, obtaining the infrared weak target detection model includes:

[0015] Step 1: Obtain the convolution kernel group of the current convolution layer in the YOLO model, and construct a convolution kernel feature group based on the convolution kernel features in the convolution kernel group;

[0016] Step 2: construct a hypergraph model based on the convolution kernel feature group and calculate the association matrix of the hypergraph model;

[0017] Step 3: Cluster the convolution kernel features based on the hypergraph model, the hypergraph association matrix, and the preset pruning rate to generate a convolution kernel clustering result;

[0018] Step 4: Prune the current convolutional layer according to the convolution kernel clustering result to generate a lightweight YOLO model, and train the lightweight YOLO model until the target network performance is achieved;

[0019] Step 5: Use the next convolutional layer as the new current convolutional layer and repeat steps 1 to 4 until all convolutional layers are pruned to obtain the infrared weak target detection model.

[0020] Optionally, constructing the convolution kernel feature group includes:

[0021] Extract multiple convolution kernels from any layer of the YOLO model, flatten the parameter matrix of each convolution kernel into a one-dimensional vector, extract the features of the one-dimensional vector, and generate the convolution kernel feature group.

[0022] Optionally, generating the convolution kernel clustering result includes:

[0023] Determine the number of categories for convolution kernel clustering based on the preset pruning rate;

[0024] Based on the hypergraph association matrix and hypergraph model, a new convolution kernel feature group is generated through the hypergraph structure learning algorithm.

[0025] A convolution kernel feature clustering operation is performed using the new convolution kernel feature group and the number of clustered categories to obtain the convolution kernel clustering result.

[0026] Optionally, generating the lightweight YOLO model includes:

[0027] Based on the convolution kernel clustering results, determine the cluster center of each cluster;

[0028] According to the cluster center, the convolution kernel closest to the cluster center is selected from each cluster and added to the convolution kernel group to be retained;

[0029] Based on the group of convolution kernels to be retained, other convolution kernels in the YOLO model are removed to generate the lightweight YOLO model.

[0030] To achieve the above objectives, the present invention further provides a complex scene infrared weak target detection system integrating multi-scale contrast features, comprising:

[0031] An image acquisition module is used to acquire infrared images of the target area under complex background;

[0032] A model improvement module is used to add a multiple attention perception module to the YOLO model for adaptive adjustment in different dimensions and acquisition of feature information; the multiple attention perception module includes: a channel attention mechanism submodule and a spatial attention mechanism submodule;

[0033] A model training module is used to train a YOLO model including a multiple attention perception module using a training set. During the training process, the convolution kernels of the convolution layer in the YOLO model are extracted, and the convolution layer is lightweighted according to the convolution kernel group to obtain an infrared weak target detection model.

[0034] The target detection module is used to input the infrared image into the infrared weak target detection model to obtain the target detection result.

[0035] Optionally, the model improvement module includes:

[0036] A feature information acquisition module is used to preprocess the infrared image using the input end of the infrared weak target detection model, input the preprocessed image into the backbone network, and slice and splice the preprocessed image horizontally and vertically through the Focus module of the backbone network; the spliced ​​output is sequentially passed through the first convolution module, the first C3 module, the second convolution module, the second C3 module, the third convolution module, the third C3 module, and the fourth convolution module to obtain a feature map; the feature map is input into the SPPF module for maximum pooling to convert a feature map of any size into a first feature vector of a fixed size; the first feature vector is passed through the fourth C3 module to output a second feature vector; the second feature vector is sequentially input into the channel attention mechanism submodule and the spatial attention mechanism submodule of the multiple attention perception module to mark key feature information, where the key feature information includes key detection targets and the positioning information of key detection targets, and low-level features with details and positioning information and high-level features with semantic information are obtained;

[0037] Among them, the C3 module is used to perform target convolution learning on the residual features output by the convolution module.

[0038] Optionally, the model training module includes:

[0039] The model training module is used to train the YOLO model including the multiple attention perception modules using the training set, and during the training process, obtain the convolution kernel group of the current convolution layer in the YOLO model, and construct a convolution kernel feature group according to the convolution kernel features in the convolution kernel group; construct a hypergraph model based on the convolution kernel feature group, and calculate the association matrix of the hypergraph model; cluster the convolution kernel features in combination with the hypergraph model, the hypergraph association matrix and the preset pruning rate to generate a convolution kernel clustering result; prune the current convolution layer according to the convolution kernel clustering result to generate a lightweight YOLO model, and train the lightweight YOLO model until the target network performance is achieved; take the next convolution layer as the new current convolution layer, repeat the lightweight processing until all convolution layers are pruned, and obtain the infrared weak target detection model.

[0040] The beneficial effects of the present invention are:

[0041] By adding multiple attention perception modules to the YOLO model, this paper can adaptively adjust across different dimensions, acquiring both low-level features containing detail and positioning information and high-level features containing semantic information. This allows for more accurate detection of small infrared targets in complex backgrounds, reducing missed detection rates and false detection rates. This allows for real-time and efficient detection under low signal-to-noise ratio conditions, improving target detection accuracy and anti-interference capabilities.

[0042] The present invention adopts a multimodal fusion method, which can stably detect weak infrared targets under different complex background conditions, such as illumination changes, noise interference, and background clutter, reducing the impact of environmental factors on detection results and improving the reliability and practicality of the system.

[0043] During training, the present invention clusters and prunes the convolution kernels of the YOLO model's convolutional layers to generate a lightweight infrared weak target detection model. This not only reduces the model's complexity and improves its operational efficiency, but also enables the detection method to run in real time on resource-constrained devices, broadening its application scenarios.

[0044] The present invention solves the problem of inaccurate detection of small infrared targets in actual scenarios. It has high practicality and broad application prospects. It can be applied to multiple fields such as industrial monitoring, security monitoring, medical diagnosis, forest fire prevention, smart cities, etc., and can generate huge economic and social benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0046] Figure 1 This is a flow chart of a method for detecting infrared weak targets in complex scenes by integrating multi-scale contrast features according to an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0048] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0049] like Figure 1As shown, this embodiment discloses a method for detecting infrared weak targets in complex scenes by integrating multi-scale contrast features, including: obtaining an infrared image in a complex background in a target area; adding a multiple attention perception module to a YOLO model for adaptive adjustment in different dimensions to obtain feature information; the multiple attention perception module includes: a channel attention mechanism submodule and a spatial attention mechanism submodule; using a training set to train the YOLO model including the multiple attention perception module, and during the training process, extracting the convolution kernel of the convolution layer in the YOLO model, and performing lightweight processing on the convolution layer according to the convolution kernel group to obtain an infrared weak target detection model; inputting the infrared image into the infrared weak target detection model to obtain a target detection result.

[0050] Furthermore, obtaining feature information includes: using the input end in the infrared weak target detection model to preprocess the infrared image, inputting the preprocessed image into the backbone network, and slicing the preprocessed image horizontally and vertically and then splicing it through the Focus module of the backbone network; passing the spliced ​​output in sequence through the first convolution module, the first C3 module, the second convolution module, the second C3 module, the third convolution module, the third C3 module, and the fourth convolution module to obtain a feature map; inputting the feature map into the SPPF module for maximum pooling, and converting a feature map of any size into a first feature vector of a fixed size; passing the first feature vector through the fourth C3 module to output a second feature vector; inputting the second feature vector into the channel attention mechanism submodule and the spatial attention mechanism submodule of the multiple attention perception module in sequence, marking key feature information, the key feature information including the key detection target and the positioning information of the key detection target, and obtaining the underlying features with details and positioning information and the high-level features with semantic information; wherein, the C3 module is used to perform target convolution learning on the residual features output by the convolution module.

[0051] Furthermore, the workflow of the multiple attention perception module: the multiple attention perception module is composed of a channel attention sub-module and a spatial attention sub-module, and its operation sequence is: first, the feature vector output by the fourth C3 module is input into the multiple attention perception module, and then passes through the channel attention sub-module and the spatial attention sub-module in turn, and finally the underlying features and high-level features are obtained to complete the feature extraction.

[0052] The channel attention submodule works as follows: The input feature vector first undergoes average pooling and max pooling operations to extract global statistical information, respectively. Within a local cross-channel interaction range of size k, a one-dimensional convolution operation with a kernel size of k is performed to aggregate channel information. The output values ​​are mapped to the interval (0, 1) using a sigmoid activation function to generate channel attention weights. The expand dimension expansion function is used to expand the channel attention weights to the same dimension as the input feature vector. The expanded channel attention weights are element-wise multiplied with the input feature vector to highlight key target detection information, resulting in a feature vector labeled with key detection targets.

[0053] The spatial attention submodule is used to obtain the location information of key detection targets. Its workflow is as follows: the feature vector marked with the key detection target is input into the spatial attention submodule and subjected to average pooling and maximum pooling operations, respectively. The results of the pooling operations are tensor-connected to merge the feature information of different pooling operations. The spatial attention weights are generated after the sigmoid activation function. The spatial attention weights are element-wise multiplied with the feature vector marked with the key detection target through a matrix dot multiplication operation, highlighting the location information of the key detection target. Finally, a feature vector marked with the key detection target and its location information is obtained. This feature vector contains low-level features with detailed and location information and high-level features with semantic information, completing feature extraction.

[0054] The refined low-level features and high-level features are input into the feature fusion network of the ESB-YOLO model to obtain fused features containing details, semantics, and positioning information at three detection depths.

[0055] First detection depth: The low-level and high-level features pass through the fifth convolutional module, then the upsampling module, and enter the first BiFPN module. In the first BiFPN module, they are fused with the output of the third C3 module in the backbone network to produce the first fused feature. The first fused feature passes through the fifth C3 module, the sixth convolutional module, and the upsampling module, and enters the second BiFPN module. In the second BiFPN module, they are fused with the output of the second C3 module in the backbone network to produce the second fused feature. The second fused feature passes through the sixth C3 module to produce the fused feature of the first detection depth, which is output to the model output. Second detection depth: The low-level and high-level features pass through the fifth convolutional module, then the upsampling module, and enter the first BiFPN module. In the first BiFPN module, they are fused with the output of the third C3 module in the backbone network to produce the first fused feature. The first fused feature passes through the fifth C3 module, the sixth convolutional module, and the upsampling module, and enters the second BiFPN module. In the second BiFPN module, they are fused with the output of the second C3 module in the backbone network to produce the second fused feature. The second fused feature passes through the sixth C3 module and the seventh convolution module in sequence before entering the third BiFPN module. In the third BiFPN module, it is fused with the output of the sixth convolution module to obtain the third fused feature. The third fused feature passes through the seventh C3 module to obtain the fused feature of the second detection depth, which is output to the output of the model. For the third detection depth, the low-level and high-level features pass through the fifth convolution module, the upsampling module, and enter the first BiFPN module. In the first BiFPN module, it is fused with the output of the third C3 module in the backbone network to obtain the first fused feature. The first fused feature passes through the fifth C3 module, the sixth convolution module, and the upsampling module in sequence before entering the second BiFPN module. In the second BiFPN module, it is fused with the output of the second C3 module in the backbone network to obtain the second fused feature. The second fused feature passes through the sixth C3 module and the seventh convolution module in sequence before entering the third BiFPN module. In the third BiFPN module, it is fused with the output of the sixth convolution module to obtain the third fused feature.

[0056] The third fused feature passes through the seventh C3 module and the eighth convolution module, and then enters the fourth BiFPN module. In the fourth BiFPN module, it is fused with the output of the fifth convolution module to obtain the fourth fused feature. The fourth fused feature passes through the eighth C3 module to obtain the third fused feature of the detection depth, which is output to the output of the model.

[0057] Finally, the fusion features of the above three detection depths are input into the output end of the ESB-YOLO model for result prediction and screening to obtain the final target detection results.

[0058] Specifically, the multiple attention perception module combines the channel attention sub-module and the spatial attention sub-module. The operation sequence in the multiple attention perception module is as follows: after the feature vector output by the fourth C3 module is input into the ECB AM module, it first passes through the channel attention sub-module and then passes through the spatial attention sub-module to obtain the underlying features and high-level features to complete the feature extraction.

[0059] The channel attention submodule specifically includes: average pooling operation, maximum pooling operation, one-dimensional convolution, Sigmiod activation function and expand dimension expansion function. The feature vector input to the multiple attention perception module first passes through average pooling and maximum pooling operations, and then within the local cross-channel interaction range of size k, a one-dimensional convolution operation with a convolution kernel size of k is used to realize channel information aggregation. Then, it passes through the Sigmiod activation function and the expand dimension expansion function in turn, and finally obtains the key target detection information. The key target detection information is element-wise multiplied with the feature vector to obtain a feature vector marked with the key detection target.

[0060] The spatial attention submodule is used to obtain the positioning information of the key detection target, which specifically includes average pooling operation, maximum pooling operation, tensor connection operation, Sigmiod activation function and matrix dot multiplication operation. First, the feature vector marked with the key detection target is input into the spatial attention submodule through average pooling and maximum pooling operations, and then the pooling operation result is tensor connected. Then, after passing through the Sigmiod activation function, the matrix dot multiplication operation is performed to obtain the positioning information of the key detection target. The positioning information of the key detection target is element-wise multiplied with the feature vector marked with the key detection target to obtain the feature vector marked with the key detection target and its positioning information. The feature vector includes low-level features with details and positioning information and high-level features with semantic information to complete feature extraction.

[0061] The low-level features and high-level features are input into the feature fusion network of the ESB-YOLO model to obtain fused features containing details, semantics, and positioning information at three detection depths. The fused features are input into the output of the ESB-YOLO model for result prediction and screening to obtain the final prediction results and achieve target detection. The feature fusion network inputs the low-level features and high-level features to obtain fused features at three detection depths.

[0062] The first detection depth obtains fusion features including: first, the bottom layer features and the high-level features are passed through the fifth convolution module, the upsampling module, and the first BiFPN module, and the features are fused with the output of the third C3 module in the backbone network in the first BiFPN module to obtain the first fusion feature; secondly, the first fusion feature passes through the fifth C3 module, the sixth convolution module, and the upsampling module in sequence, and enters the second BiFPN module, and the features are fused with the output of the second C3 module in the backbone network in the second BiFPN module to obtain the second fusion feature; thirdly, the second fusion feature passes through the sixth C3 module to obtain the fusion feature of the first detection depth, and is output to the output end by the sixth C3 module; The second detection depth obtains fusion features including: first, the bottom-level features and high-level features are passed through the fifth convolution module, the upsampling module, and the first BiFPN module, and the output of the third C3 module in the backbone network is fused in the first BiFPN module to obtain the first fusion feature; secondly, the first fusion feature passes through the fifth C3 module, the sixth convolution module, and the upsampling module in turn, and enters the second BiFPN module, and the output of the second C3 module in the backbone network is fused in the second BiFPN module to obtain the second fusion feature; thirdly, the second fusion feature passes through the sixth C3 module and the seventh convolution module in turn, and enters the third BiFPN module, and is fused in the third BiF The output of the sixth convolution module is subjected to feature fusion in the PN module to obtain the third fusion feature; finally, the third fusion feature passes through the seventh C3 module to obtain the fusion feature of the second detection depth, which is output to the output end by the seventh C3 module; the fusion feature obtained by the third detection depth includes: first, the bottom layer feature and the high layer feature pass through the fifth convolution module, pass through the upsampling module, and enter the first BiFPN module, and in the first BiFPN module, feature fusion is performed with the output of the third C3 module in the backbone network to obtain the first fusion feature; secondly, the first fusion feature passes through the fifth C3 module, the sixth convolution module, and the upsampling module in turn, and enters the second BiFPN module, and in the second BiFPN module The feature fusion is performed with the output of the second C3 module in the backbone network to obtain the second fused feature; again, the second fused feature passes through the sixth C3 module and the seventh convolution module in sequence, and enters the third BiFPN module, and is fused with the output of the sixth convolution module in the third BiFPN module to obtain the third fused feature; again, the third fused feature passes through the seventh C3 module and the eighth convolution module in sequence, and enters the fourth BiFPN module, and is fused with the output of the fifth convolution module in the fourth BiFPN module to obtain the fourth fused feature; finally, the fourth fused feature passes through the eighth C3 module to obtain the fused feature of the third detection depth, which is output to the output end by the eighth C3 module.

[0063] Furthermore, obtaining an infrared weak target detection model includes: step 1, obtaining a convolution kernel group of a current convolution layer in a YOLO model, and constructing a convolution kernel feature group based on the convolution kernel features in the convolution kernel group; step 2, constructing a hypergraph model based on the convolution kernel feature group, and calculating an association matrix of the hypergraph model; step 3, clustering the convolution kernel features based on the hypergraph model, the hypergraph association matrix, and a preset pruning rate to generate a convolution kernel clustering result; step 4, performing a pruning operation on the current convolution layer according to the convolution kernel clustering result to generate a lightweight YOLO model, and training the lightweight YOLO model until the target network performance is achieved; step 5, taking the next convolution layer as the new current convolution layer, repeating steps 1 to 4 until all convolution layers are pruned to obtain an infrared weak target detection model.

[0064] Constructing a convolution kernel feature group includes: extracting multiple convolution kernels from any layer of the YOLO model, flattening the parameter matrix of each convolution kernel into a one-dimensional vector, extracting the features of the one-dimensional vector, and generating a convolution kernel feature group.

[0065] Generating convolution kernel clustering results includes: determining the number of categories of convolution kernel clustering according to a preset pruning rate; generating new convolution kernel feature groups based on the hypergraph association matrix and hypergraph model through a hypergraph structure learning algorithm; and performing convolution kernel feature clustering operations using the new convolution kernel feature groups and the number of clustered categories to obtain convolution kernel clustering results.

[0066] Generating a lightweight YOLO model includes: determining the cluster center of each cluster based on the convolution kernel clustering results; selecting the convolution kernel closest to the cluster center from each cluster and adding it to the group of convolution kernels to be retained; based on the group of convolution kernels to be retained, removing other convolution kernels in the YOLO model to generate a lightweight YOLO model.

[0067] Specifically, obtaining the infrared weak target detection model includes:

[0068] Step 1: Extract the convolution kernel feature group: First, obtain the convolution kernel group of the current convolution layer in the YOLO model. For each convolution kernel in this group, flatten its parameter matrix into a one-dimensional vector and extract the features of these one-dimensional vectors to construct the convolution kernel feature group.

[0069] Step 2: Construct a hypergraph model and calculate the association matrix: Based on the convolution kernel feature group constructed above, further construct a hypergraph model and calculate the association matrix of the hypergraph model.

[0070] Step 3: Convolution kernel feature clustering: The convolution kernel features are clustered using the hypergraph model, hypergraph association matrix, and pre-set pruning rate to generate the convolution kernel clustering results. Specifically, the number of convolution kernel cluster categories is determined based on the pre-set pruning rate. Then, based on the hypergraph association matrix and hypergraph model, a hypergraph structure learning algorithm is used to calculate and generate new convolution kernel feature groups. Finally, the convolution kernel feature clustering operation is performed using the new convolution kernel feature groups and the number of cluster categories to obtain the convolution kernel clustering results.

[0071] Step 4: Convolutional Layer Pruning and Model Lightweighting: Based on the kernel clustering results, determine the cluster center for each cluster. Next, select the kernel closest to the cluster center from each cluster and add it to the group of kernels to be retained. Based on this group of kernels to be retained, remove all other kernels from the YOLO model to generate a lightweight YOLO model. Train the lightweight YOLO model until the target network performance is achieved.

[0072] Step 5. Repeat the operation until all convolutional layers are pruned: Use the next convolutional layer as the new current convolutional layer and repeat steps 1 to 4 above until all convolutional layers are pruned, and finally obtain the infrared weak target detection model.

[0073] This embodiment also provides a complex scene infrared weak target detection system that integrates multi-scale contrast features, including: an image acquisition module, used to acquire infrared images in a complex background in the target area; a model improvement module, used to add a multiple attention perception module to the YOLO model, used for adaptive adjustment in different dimensions to obtain feature information; the multiple attention perception module includes: a channel attention mechanism submodule and a spatial attention mechanism submodule; a model training module, used to use a training set to train the YOLO model containing the multiple attention perception module, and during the training process, extract the convolution kernel of the convolution layer in the YOLO model, and perform lightweight processing on the convolution layer according to the convolution kernel group to obtain an infrared weak target detection model; a target detection module, used to input the infrared image into the infrared weak target detection model to obtain the target detection result.

[0074] Furthermore, the model improvement module includes: a feature information acquisition module, which is used to use the input end in the infrared weak target detection model to preprocess the infrared image, input the preprocessed image into the backbone network, and slice and splice the preprocessed image horizontally and vertically through the Focus module of the backbone network; the spliced ​​output is sequentially passed through the first convolution module, the first C3 module, the second convolution module, the second C3 module, the third convolution module, the third C3 module, and the fourth convolution module to obtain a feature map; the feature map is input into the SPPF module for maximum pooling, and the feature map of any size is converted into a first feature vector of a fixed size; the first feature vector is passed through the fourth C3 module to output a second feature vector; the second feature vector is sequentially input into the channel attention mechanism submodule and the spatial attention mechanism submodule of the multiple attention perception module to mark key feature information, which includes key detection targets and the positioning information of key detection targets, and obtains underlying features with details and positioning information and high-level features with semantic information; wherein, the C3 module is used to perform target convolution learning on the residual features output by the convolution module.

[0075] Furthermore, the model training module includes: a model training module for training a YOLO model including a multiple attention perception module using a training set, and during the training process, obtaining the convolution kernel group of the current convolution layer in the YOLO model, and constructing a convolution kernel feature group based on the convolution kernel features in the convolution kernel group; constructing a hypergraph model based on the convolution kernel feature group, and calculating the association matrix of the hypergraph model; clustering the convolution kernel features in combination with the hypergraph model, the hypergraph association matrix and the preset pruning rate to generate a convolution kernel clustering result; performing a pruning operation on the current convolution layer according to the convolution kernel clustering result to generate a lightweight YOLO model, and training the lightweight YOLO model until the target network performance is achieved; taking the next convolution layer as the new current convolution layer, repeating the lightweight processing until all convolution layers are pruned to obtain an infrared weak target detection model.

[0076] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A method for detecting infrared weak targets in complex scenes by integrating multi-scale contrast features, characterized in that: include: Acquire infrared images of target areas against complex backgrounds; Add multiple attention perception modules to the YOLO model to adaptively adjust in different dimensions and obtain feature information; The multiple attention perception module includes: a channel attention mechanism submodule and a spatial attention mechanism submodule; The YOLO model including multiple attention perception modules is trained using the training set. During the training process, the convolution kernels of the convolution layer in the YOLO model are extracted. The convolution layer is lightweight processed according to the convolution kernel group to obtain an infrared weak target detection model. The infrared image is input into the infrared weak target detection model to obtain a target detection result.

2. The method for detecting weak infrared targets in complex scenes by integrating multi-scale contrast features according to claim 1, characterized in that: Acquiring the characteristic information includes: The infrared image is preprocessed by using the input end of the infrared weak target detection model, and the preprocessed image is input into the backbone network. The preprocessed image is sliced ​​horizontally and vertically and then spliced ​​through the Focus module of the backbone network; the spliced ​​output is sequentially passed through the first convolution module, the first C3 module, the second convolution module, the second C3 module, the third convolution module, the third C3 module, and the fourth convolution module to obtain a feature map; the feature map is input into the SPPF module for maximum pooling, and the feature map of any size is converted into a first feature vector of a fixed size; the first feature vector is passed through the fourth C3 module to output a second feature vector; the second feature vector is sequentially input into the channel attention mechanism submodule and the spatial attention mechanism submodule of the multiple attention perception module to mark key feature information, which includes key detection targets and the positioning information of key detection targets, and obtains low-level features with details and positioning information and high-level features with semantic information; Among them, the C3 module is used to perform target convolution learning on the residual features output by the convolution module.

3. The method for detecting weak infrared targets in complex scenes by integrating multi-scale contrast features according to claim 1, characterized in that: Obtaining the infrared weak target detection model includes: Step 1: Obtain the convolution kernel group of the current convolution layer in the YOLO model, and construct a convolution kernel feature group based on the convolution kernel features in the convolution kernel group; Step 2: construct a hypergraph model based on the convolution kernel feature group and calculate the association matrix of the hypergraph model; Step 3: Cluster the convolution kernel features based on the hypergraph model, the hypergraph association matrix, and the preset pruning rate to generate a convolution kernel clustering result; Step 4: Prune the current convolutional layer according to the convolution kernel clustering result to generate a lightweight YOLO model, and train the lightweight YOLO model until the target network performance is achieved; Step 5: Use the next convolutional layer as the new current convolutional layer and repeat steps 1 to 4 until all convolutional layers are pruned to obtain the infrared weak target detection model.

4. The method for detecting weak infrared targets in complex scenes by integrating multi-scale contrast features according to claim 3 is characterized in that: Constructing the convolution kernel feature group includes: Extract multiple convolution kernels from any layer of the YOLO model, flatten the parameter matrix of each convolution kernel into a one-dimensional vector, extract the features of the one-dimensional vector, and generate the convolution kernel feature group.

5. The method for detecting weak infrared targets in complex scenes by integrating multi-scale contrast features according to claim 3, characterized in that: Generating the convolution kernel clustering result includes: Determine the number of categories for convolution kernel clustering based on the preset pruning rate; Based on the hypergraph association matrix and hypergraph model, a new convolution kernel feature group is generated through the hypergraph structure learning algorithm. A convolution kernel feature clustering operation is performed using the new convolution kernel feature group and the number of clustered categories to obtain the convolution kernel clustering result.

6. The method for detecting weak infrared targets in complex scenes by integrating multi-scale contrast features according to claim 3, characterized in that: Generating the lightweight YOLO model includes: Based on the convolution kernel clustering results, determine the cluster center of each cluster; According to the cluster center, the convolution kernel closest to the cluster center is selected from each cluster and added to the convolution kernel group to be retained; Based on the group of convolution kernels to be retained, other convolution kernels in the YOLO model are removed to generate the lightweight YOLO model.

7. A complex scene infrared weak target detection system integrating multi-scale contrast features, characterized by: include: An image acquisition module is used to acquire infrared images of the target area under complex background; The model improvement module is used to add multiple attention perception modules to the YOLO model for adaptive adjustment in different dimensions and to obtain feature information; The multiple attention perception module includes: a channel attention mechanism submodule and a spatial attention mechanism submodule; A model training module is used to train a YOLO model including a multiple attention perception module using a training set. During the training process, the convolution kernels of the convolution layer in the YOLO model are extracted, and the convolution layer is lightweighted according to the convolution kernel group to obtain an infrared weak target detection model. The target detection module is used to input the infrared image into the infrared weak target detection model to obtain the target detection result.

8. The method for detecting weak infrared targets in complex scenes by integrating multi-scale contrast features according to claim 7, characterized in that: The model improvement module includes: A feature information acquisition module is used to preprocess the infrared image using the input end of the infrared weak target detection model, input the preprocessed image into the backbone network, and slice and splice the preprocessed image horizontally and vertically through the Focus module of the backbone network; the spliced ​​output is sequentially passed through the first convolution module, the first C3 module, the second convolution module, the second C3 module, the third convolution module, the third C3 module, and the fourth convolution module to obtain a feature map; the feature map is input into the SPPF module for maximum pooling to convert a feature map of any size into a first feature vector of a fixed size; the first feature vector is passed through the fourth C3 module to output a second feature vector; the second feature vector is sequentially input into the channel attention mechanism submodule and the spatial attention mechanism submodule of the multiple attention perception module to mark key feature information, where the key feature information includes key detection targets and the positioning information of key detection targets, and low-level features with details and positioning information and high-level features with semantic information are obtained; Among them, the C3 module is used to perform target convolution learning on the residual features output by the convolution module.

9. The method for detecting weak infrared targets in complex scenes by integrating multi-scale contrast features according to claim 7, characterized in that: The model training module includes: The model training module is used to train the YOLO model including the multiple attention perception modules using the training set, and during the training process, obtain the convolution kernel group of the current convolution layer in the YOLO model, and construct a convolution kernel feature group according to the convolution kernel features in the convolution kernel group; construct a hypergraph model based on the convolution kernel feature group, and calculate the association matrix of the hypergraph model; cluster the convolution kernel features in combination with the hypergraph model, the hypergraph association matrix and the preset pruning rate to generate a convolution kernel clustering result; prune the current convolution layer according to the convolution kernel clustering result to generate a lightweight YOLO model, and train the lightweight YOLO model until the target network performance is achieved; take the next convolution layer as the new current convolution layer, repeat the lightweight processing until all convolution layers are pruned, and obtain the infrared weak target detection model.