Cross-modal infrared small target detection method based on degeneration memory compensation
The degradation memory compensation network addresses cross-modal detection challenges by estimating and compensating for image degradation, improving detection accuracy and robustness in real-world environments.
Patent Information
- Application Number
- CN202510387996.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-31
AI Technical Summary
In the cross-modal infrared small object detection, the existing technology has problems such as unclear fusion of cross-modal image information, inadequate network learning process, decreased cross-modal detection accuracy and lack of degradation compensation mechanism. It is difficult to exert modal complementarity advantages in the case of significant distribution differences between different modalities.
Using a degradation memory compensation method, a multi-level Swin Transformer encoder and decoder is designed by building a degradation estimation module and a memory compensation network, combining metric learning and hybrid mask strategies, and optimizing the network with masked region loss and Soft-IoU loss is achieved to achieve cross-modal feature modulation and reconstruction.
It improves the accuracy and robustness of cross-modal infrared small object detection, improves detection performance in real environments, reduces false alarm rates and false alarm rates, and enhances information complementarity and network reconstruction capabilities.
Smart Images

Figure CN120318492A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular, to a cross-modal infrared small target detection method based on degradation memory compensation. Background Art
[0002] Infrared small target detection is an important research direction in the field of image processing, aiming to obtain the position information of small targets in a wide background by using data-driven and deep representation principles, and is widely applied to tasks such as military surveillance, UAV navigation, and target recognition in complex scenarios. However, the existing technologies are usually limited to a single modality (such as the infrared modality), and the few existing cross-modal detection technologies have the following obvious defects: (1) The purpose of cross-modal image information fusion is often not clear enough, and the learning process and cross-modal interpretability of the network are poor, and it is impossible to effectively control the influence of the degradation intensity between modalities on the detection accuracy. (2) Due to the "insufficient supervision signal for end-to-end training", the existing technologies often have the problem of a sharp drop in cross-modal detection accuracy in the real environment. (3) There is a lack of a degradation compensation mechanism for cross-modal detection. In the face of the significant difference in the distribution of targets between different modalities, the existing models are difficult to fully utilize the advantages of modality complementarity. Summary of the Invention
[0003] The technical solution of the present invention to solve the above technical problems is to provide a cross-modal infrared small target detection method based on degradation memory compensation, including the following steps:
[0004] S1: Prepare a data set, including data set one for training and data set two for testing. The data set one contains registered infrared and visible light image pairs, and the data set two contains a self-made registered test set;
[0005] S2: Construct a degradation memory compensation network, including a degradation estimation module (DEM) and a memory compensation network; the degradation estimation module estimates the blur degradation score (Pblur), noise degradation score (Pnoise), and visibility degradation score (Gvisible) of infrared and visible light images respectively through six metric spaces, and outputs them in the form of heat maps; the memory compensation network generates a weight heat map and a weight vector through a fully connected layer (FC), a Softmax layer, and a hist layer;
[0006] S3: Train the degradation memory compensation network, generate dual-modal control group (gc) and consistency reinforcement group (gr) degradation samples through a degradation sample generation model, and optimize the network by using an adaptive ranking loss function;
[0007] S4: Perform cross-modal image region masking processing, divide the registered infrared and visible light images into image blocks, randomly mask some image blocks and replace them with mask signals;
[0008] S5: Cross-modal image hybrid mask processing, replacing the masked infrared image patches with corresponding visible light image patches and the masked visible light image patches with corresponding infrared image patches to generate a hybrid mask image pair;
[0009] S6: Construct a small target detection network, including multi-level Swin Transformer encoders, a fusion strategy, and a decoder; the multi-level Swin Transformer encoders modulate and compensate cross-modal features using the weight heatmap and weight vector of the memory compensation network;
[0010] S7: Train the small target detection network, using an end-to-end framework and optimizing the network by combining masked region loss and Soft-IoU loss;
[0011] S8: Test and evaluate the model, and output the detection results and performance metrics.
[0012] Further, in step S1, the first dataset is the Anti-UAV dataset, and the second dataset contains infrared and visible light image pairs of multiple scenarios, and all image pairs have completed fine-grained registration.
[0013] Further, in step S2, the degradation estimation module (DEM) includes: a blur degradation sub-network, a noise degradation sub-network, and a visibility degradation sub-network. Each sub-network is composed of a spatial and channel attention module (CBAM) and outputs a degradation score heatmap using a Sigmoid function;
[0014] The noise degradation sub-network and the blur degradation sub-network are generated based on a probabilistic degradation model, and the generation is expressed by the formula:
[0015]
[0016] where n and k respectively represent the sizes of the noise kernel and the blur kernel, z n and z k respectively represent the noise and blur kernels, both of which follow a normal distribution; M n and M k respectively represent the noise and blur generation functions;
[0017] The visibility degradation sub-network is composed of a brightness and contrast degradation function;
[0018] The memory compensation network consists of a fully connected layer (FC), a Softmax, and a hist layer connected to the end of each DEM.
[0019] Further, in step S3, the adaptation ranking loss function includes:
[0020] For the Margin Ranking Loss of the bimodal control group, the difference in scores between the blurred degradation and the noise degradation is constrained, expressed as:
[0021]
[0022] For the Margin Ranking Loss of the consistency enhancement group, the difference in scores between the non-degraded samples and the severely degraded samples is constrained; expressed as:
[0023]
[0024] Furthermore, in step S4, the infrared and visible light images are respectively divided into 3×4 image patches, and 6 of them are randomly masked and replaced with mask signals.
[0025] Furthermore, in the above S5, the infrared masked image after the hybrid masking process is expressed as:
[0026]
[0027] The visible light masked image is expressed as:
[0028]
[0029] where {MASK} represents the mask signal, represents the image patch of the original visible light image at (i,j), represents the image patch of the original infrared image at (i,j).
[0030] Furthermore, in step S6, the small target detection network includes:
[0031] Encoder: 5 levels of Patch Merging and 5 levels of Swin Transformer blocks, and the features output at each level are fused after being modulated by the memory compensation weights;
[0032] Decoder: 5 levels of Patch Expanding and dual ResBlock blocks, used to reconstruct the detection image with target boxes.
[0033] Furthermore, in step S7, the masked area loss function is:
[0034]
[0035] where, represents the total number of pixels in the masked area to ensure uniform distribution of the loss; I f,i represents the pixel value of the reconstructed image at the i-th masked position; I ir / vi,i is the true pixel value of the original image at the i-th masked position.
[0036] Furthermore, in step S7,
[0037] The Soft-IoU loss function is expressed as:
[0038]
[0039] where P i,j represents the predicted probability at the position (i, j); G i,j represents the ground truth label at the pixel position (i, j); ε represents a smoothing term to prevent division-by-zero errors, usually a small constant (1e-6), which is used to ensure numerical stability.
[0040] Furthermore, in step S8, the performance metrics include the detection rate, false alarm rate, recall rate, IoU, and F1 score.
[0041] Compared with the prior art, the present invention provides a cross-modal infrared small target detection method based on degradation memory compensation, having the following beneficial effects:
[0042] 1. The present invention innovatively combines metric learning with a hybrid mask strategy, skillfully utilizes the image degradation modeling mechanism and the metric space mapping method, and introduces a mask strategy and a multi-level autoencoder for the target detection task, developing a new cross-modal small target detection framework.
[0043] 2. The present invention designs a controllable true degradation estimation module (DEM) based on metric learning, and for the first time introduces metric learning into the cross-modal infrared small target detection task; the present invention also proposes a complete metric degradation sample generation model, and trains the DEM to estimate the true degradation score by learning the similarity and difference between samples.
[0044] 3. The present invention proposes a multi-level Vit autoencoder cross-modal target detection network based on a hybrid mask; first, the traditional direct mask shielding strategy is improved, and a new infrared and visible light hybrid mask image pair is constructed for training a strongly robust detection model; then, a multi-level Vit autoencoder detection network is designed, and the estimated weights output by the memory compensation network are used to guide the network to reconstruct the detection image.
[0045] 4. For the above-mentioned hybrid mask shielding strategy, the present invention proposes a masked region loss (L mask ), by constraining the similarity between the reconstructed image and the masked regions in the original infrared and visible light images, and adapting to the Soft-IoU loss, promoting the model to learn better information representation and accurate cross-modal infrared small target detection ability. Description of the Drawings
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0047] Figure 1 It is a flowchart of the steps of a cross-modal infrared small target detection method based on degraded memory compensation according to the present invention;
[0048] Figure 2 It is a working principle diagram of a cross-modal infrared small target detection method based on degraded memory compensation according to the present invention;
[0049] Figure 3 It is a schematic diagram of the principle of the Degradation Estimation Module (DEM) of the present invention;
[0050] Figure 4 It is a structural diagram of the degradation compensation network of the present invention;
[0051] Figure 5 It is an effect diagram of the qualitative comparison between the cross-modal small target detection method of the present invention and the existing methods;
[0052] Figure 6 It is a schematic diagram of the quantitative comparison between the cross-modal small target detection method of the present invention and the existing methods. Specific Embodiments
[0053] The present invention proposes a cross-modal infrared small target detection method based on degraded memory compensation, aiming to design a cross-modal infrared small target detection method based on degraded memory compensation.
[0054] The following will illustrate the cross-modal infrared small target detection method based on degraded memory compensation proposed by the present invention in specific embodiments:
[0055] Embodiment 1:
[0056] A cross-modal infrared small target detection method based on degraded memory compensation, as Figure 1 shown, includes the following steps:
[0057] S1: Prepare the data sets, including data set one for training and data set two for testing. The data set one contains registered infrared and visible light image pairs, and the data set two contains a self-made registered test set;
[0058] Specifically, the prepared dataset one is the Anti-UAV dataset, which is used to train the degradation compensation network and the small target detection network. In this embodiment, 15,700 pairs of infrared and visible light image pairs of the Anti-UAV dataset are selected. The selection principle is to ensure the richness of the image background and remove overly similar background image pairs. The prepared dataset two is a self-made infrared and visible light image dataset, which is used to test the final small target detection model. The self-made dataset two in this embodiment contains 463 pairs of infrared and visible light image pairs in various scenarios such as ground-to-air, air-to-air, buildings, sky, and trees. All image pairs have achieved fine-grained registration.
[0059] S2: Construct a degradation memory compensation network, including a degradation estimation module (DEM) and a memory compensation network. The degradation estimation module estimates the blur degradation score (Pblur), noise degradation score (Pnoise), and visibility degradation score (Gvisible) of infrared and visible light images through six metric spaces respectively, and outputs them in the form of heatmaps. The memory compensation network generates a weight heatmap and a weight vector through a fully connected layer (FC), a Softmax layer, and a hist layer.
[0060] Specifically, the degradation memory compensation network includes a degradation estimation module (DEM) and a memory compensation network. The prepared dual-modal dataset in step S1 is input into the DEM to obtain two sets of degradation scores, and then the two sets of degradation scores are input into the compensation network to generate pseudo-supervised signals (including a set of weight heatmaps and weight vectors).
[0061] As Figure 3 shown, the input of the degradation estimation module (DEM) is the dual-modal control group degradation sample (gc) obtained by the degradation sample generation model and two sets of degradation samples (gr) of the consistency enhancement group. Six metric spaces are constructed for these two sets of degradation samples, corresponding to the blur degradation sub-network metric space (Pblur), noise degradation sub-network metric space (Pnoise), and visibility degradation sub-network metric space (Gvisible) of infrared and visible light images respectively. Each metric space sub-network is composed of one or more spatial and channel attention modules (CBAM), and uses a Sigmoid activation function as the output to convert the non-quantifiable degradation intensity into a degradation score. Specifically, each set of infrared and visible light images has three prediction branches to predict the Pnoise score, Pblur score, and Gvisible score respectively, and these three scores are output in the form of heatmaps. The six constructed metric spaces jointly estimate the degradation degree with noise degradation, blur degradation, and visibility degradation.
[0062] The noise degradation sub-network and the blur degradation sub-network are generated based on the probability degradation model, and the generation is represented by the formula:
[0063]
[0064] where n and k represent the sizes of the noise kernel and the blur kernel respectively, and z n and z k represent noise and the blur kernel respectively, both following a normal distribution; M n and M k represent the noise and the blur generation functions respectively;
[0065] The visibility degradation sub-network consists of the luminance and contrast degradation functions;
[0066] The memory compensation network, as shown in Figure 2 b, consists of fully connected layers (FC), Softmax, and hist layers connected to the end of each DEM; each FC layer of the memory compensation network takes a set of degraded score heatmaps output by the DEM as input, generates new weight heatmaps through Softmax, and then generates weight vectors through the hist layer; these two weights jointly guide the compensation as shown in Figure 2 (d) The training process of the fusion network.
[0067] S3: Train the degradation memory compensation network, generate dual-modal control group (gc) and consistency reinforcement group (gr) degradation samples through the degradation sample generation model, and optimize the network using the adaptive ranking loss function;
[0068] The degradation memory compensation network, as shown in Figure 4 , includes a degradation sample generation model, an adaptive ranking loss function, and network training. Use the dual-modal dataset, DEM module, and metric learning method prepared in S1 to train the memory compensation network model. The trained memory compensation model can output pseudo-supervision signals (including a set of weight heatmaps and weight vectors);
[0069] The training of the degradation memory compensation network includes:
[0070] First, to train a highly robust metric space, this embodiment generates two sets of degradation samples, namely the dual-modal control group (gc) and the consistency reinforcement group (gr).
[0071] Then, to simulate the noise degradation of infrared images due to insufficient light sources or extreme temperature conditions, and the blur degradation of visible light images due to haze or low contrast conditions, and degradation factors such as the blurred low-temperature boundary of the infrared image target and overexposure or underexposure of the visible light image caused by luminance / contrast degradation; this embodiment introduces two combined degradation methods of O:σ1 and U:σ2 in g c ; and introduces one degradation method of H:σ3 in g r , where σ i represents the combined degradation factor, and generates and There are three groups of degraded samples, each group of degraded samples contains several degraded images after infrared and visible light degradation; the acquisition of degraded samples can be described by the formula:
[0072]
[0073] σ3=d no ||d max ;
[0074] in, Compare With more severe O:σ1 degradation, Compare With more severe O:σ1 degradation; This rule also applies; represents a non-degenerate sample, Indicates the sample that suffered the most severe degradation of Pnoise, Pblur and Gvisible superposition;
[0075] Subsequently, two sets of degradation samples are input into the DEM to generate degradation scores. This process can be expressed as:
[0076]
[0077] Among them, f RDEM (.) represents the functional expression of the DEM module.
[0078] Finally, under the action of the loss function, each group of degradation scores continuously and accurately estimates the most accurate degradation score (degradation degree) of each degraded sample.
[0079] The adapted ranking loss function can be expressed as follows:
[0080] In order to distinguish samples with different degrees of degradation and train three metric spaces, this embodiment designs a multi-directional MarginRanking Loss as a loss function; this loss function distances positive samples from negative samples, constrains the distance between severely degraded negative samples after H:σ3 and non-degraded positive samples to be larger, and the distance between hierarchical degraded samples after O:σ1 and U:σ2 to be larger. Margin Ranking Loss can be defined as:
[0081]
[0082] in, κ is the boundary parameter that constrains the distance between two samples; N represents the number of training samples; and Represents the prediction score; when the present invention uses Margin Ranking Loss, and The corresponding true value s i Ranked behind s j So let ξ = 1. Therefore, in a certain metric space L MR Can be simplified to:
[0083]
[0084] Among them, →i represents the i-th training sample. In the task of this embodiment, there are three Margin RankingLoss losses. The first is the loss for the O:σ1 metric space in the bimodal control group g c Expressed as:
[0085]
[0086] The second is the loss for the U:σ2 metric space in the bimodal control group g c Expressed as:
[0087]
[0088] The third is the loss for the H:σ3 metric space in the consistency enhancement group g r For non-degraded samples, the loss function should limit the network from making any transformations or modifications to these samples. For samples suffering from severe degradation, the network requires stronger recovery operations. Defined as:
[0089]
[0090] Among them, Represents the l2 norm.
[0091] S4: Cross-modal image region masking processing, dividing the registered infrared and visible light images into image blocks, randomly masking some image blocks and replacing them with mask signals;
[0092] Specifically, cross-modal image region masking processing: includes cutting bimodal image blocks and randomly masking bimodal image blocks, dividing the infrared and visible light images into 3×4 image blocks respectively, randomly masking 6 of them and replacing them with mask signals.
[0093] S5: Cross-modal image hybrid masking processing, replacing the masked infrared image blocks with the corresponding visible light image blocks, and replacing the masked visible light image blocks with the corresponding infrared image blocks to generate a hybrid masked image pair;
[0094] Specifically, for cross-modal image hybrid mask processing: replace the masked infrared image patches in step S4 with the corresponding patches from the original visible light image; replace the masked visible light image patches in step S4 with the corresponding patches from the original infrared image; to obtain a new pair of dual-modal images after hybrid mask processing.
[0095] Existing multi-modal feature fusion methods often directly perform concatenation operations at the last or initial layer of the network, which undoubtedly weakens the powerful feature expression ability of the neural network; while the cross-modal mask method designed in this embodiment forces the network to learn how to reconstruct complementary information between infrared and visible light images under incomplete information by masking and mixing partial information of multiple modalities, solves the problem of imbalance between different modality image information, and improves the detection accuracy of small targets; as Figure 2 (c) shows the schematic diagram of cross-modal image region masking processing and hybrid mask principle designed in this embodiment.
[0096] For the cross-modal image region masking processing, at the input end, a set of infrared and visible light images are preprocessed and registered to the same size of 640×480, denoted as I ir and I vi ; Subsequently, the images are divided into image patches with a size of 3×4 and where H = 640, W = 480; Different from the traditional mask mechanism, a random image patch masking method is designed; randomly select 6 subsets of image patches IR from the infrared image I for masking, that is, replace them with the mask signal [MASK]; then randomly select 6 subsets of image patches VI from the infrared image I for masking, and also replace them with the mask signal [MASK].
[0097] For the cross-modal image hybrid mask processing, select the corresponding patches from the corresponding positions of the visible light image to replace the 6 subsets of image patches in the infrared image Then select the corresponding patches from the corresponding positions of the infrared image to replace the 6 subsets of image patches in the visible light image Finally, the infrared masked image I IR⊙mask can be expressed as:
[0098]
[0099] The visible light masked image I VI⊙mask can be expressed as:
[0100]
[0101] In this embodiment, through the above-mentioned mask mixing processing method, the masked infrared and masked visible light images not only contain the features of the infrared image I IR but also embed the information of the visible light image I VI . This not only strengthens the complementarity of information but also improves the network's reconstruction ability when some information is missing. In addition, the random mask mechanism ensures that the mask position changes in each iteration, further weakening the boundary effect caused by fixed image block division and avoiding the potential artifact risk brought by the fixed input mode. Especially for the dataset captured in a harsh environment and facing serious image degradation problems, the mask shielding method of this embodiment plays a crucial role.
[0102] S6: Construct a small target detection network, including multi-level Swin Transformer encoders, a fusion strategy, and a decoder; the multi-level Swin Transformer encoders use the weight heatmap and weight vector of the memory compensation network to modulate and compensate the cross-modal features;
[0103] Specifically, the small target detection network includes a designed autoencoder, a fusion strategy, and an auto decoder; the dual-modal image pair obtained in S5 is input into the small target detection network for feature extraction, feature fusion, and image reconstruction; as Figure 2 (d) shows the structure diagram of the small target detection network. The encoder consists of 5 levels of Patch Merging and 5 levels of Swin Transformer blocks (VitBlock); the decoder consists of 5 levels of Patch Expanding and 5 levels of dual ResBlock blocks;
[0104] The working principle of the small target detection network is as follows:
[0105] Before inputting into the encoder, first, the masked infrared image I IR⊙mask and the masked visible light image I VI⊙mask undergo feature mapping through a linear projection layer, and position information is embedded in the image to generate the initial feature representation of the masked infrared image and the initial feature representation of the masked visible light image
[0106] Then, the encoder is used for deep feature extraction. Each level of the encoder network processes the input features hierarchically, where the Patch Merging module is used to downsample the feature dimension of each layer; specifically, and After passing through the 1st level of Vit Block, the output is and The original image generates a degradation score after passing through the degradation estimation module (DEM), and then generates a weight heatmap w2 and a weight vector w1 through the memory compensation network. The modulation compensation operation with respect to w2w1 is completed as follows:
[0107]
[0108] Similarly, the outputs of the 2nd, 3rd, and 4th level Vit Blocks will all perform the above operation with the weights generated by the memory compensation network, and then be sent to the next-level network for further processing; the two outputs after the 5th level Vit Block use the Cat concatenation operation to output the features extracted by the encoder.
[0109] Finally, a decoder is used for feature reconstruction and target detection. Specifically, the Patch Expanding module is used to upsample the feature map layer by layer to gradually restore the low-dimensional features to the original resolution; the ResBlock block is used to retain high-level semantic and detailed features and perform non-linear transformation; the last layer of the decoder consists of a ReLU activation function and a convolutional layer, which is used to reconstruct the final fused image I with target boxes. f 。
[0110] S7: Train the small target detection network, adopt an end-to-end framework, and optimize the network by combining the masked area loss and the Soft-IoU loss;
[0111] Specifically, training the small target detection network includes the small target detection network training framework, designing the masked area loss function, and selecting the optimization method; the training framework is an end-to-end manner from the input bimodal image, degradation compensation pseudo-supervision to small target detection; in addition to using the masked area loss, the joint IoU loss (Soft-IoU loss) is also adapted to improve the accuracy of small target detection; use the bimodal image pairs obtained in S5 as the training data set to train the small target detection network constructed in S6.
[0112] The small target detection network training framework is the end-to-end Yolov10 framework; the masked area loss is used to constrain the reconstruction error of the cross-modal masked area; the design principle of the masked area loss is as follows:
[0113] First, since the small target detection network is trained by randomly masking some image patches and mixing image patches of another modality; for the masked areas, a masked area loss (L mask :Masked Area Loss) is designed to constrain the reconstruction error of the masked areas.
[0114] Suppose the index set of the masked areas is which contains the index positions of all masked image patches. For each masked area the mean squared error (MSE) is used to calculate the loss of the reconstructed image I f and the original image I ir / vi in this area, Lmask It can be expressed by the formula as follows:
[0115]
[0116] Wherein, represents the total number of pixels in the masked area to ensure uniform distribution of the loss; I f,i represents the pixel value of the reconstructed image at the i-th masked position; I ir / vi,i is the true pixel value of the original image at the i-th masked position.
[0117] To improve the cross-modal detection accuracy and ensure the global structural consistency of the reconstructed image, the joint IoU loss (Soft-IoU loss) is also adapted to further constrain the network training. The Soft-IoU loss is expressed as:
[0118]
[0119] Wherein, P i,j represents the predicted probability at the position (i,j); G i,j represents the true label at the pixel position (i,j); ε represents a smoothing term to prevent division-by-zero errors, usually a small constant (1e-6), which is used to ensure numerical stability.
[0120] S8: Test and evaluate the model, and output the detection results and performance metrics.
[0121] Specifically, after the training and validation in step S7 are completed, the network parameters are solidified, and the final small target detection model is saved; the model is tested using the dataset two in step S1 to obtain the result image marked with detection boxes; and the detection rate (P d ), false alarm rate (F a ), recall rate (R e ), area under the segmentation curve (IoU), and F1 score (F) are used to evaluate the performance and accuracy of the detection model.
[0122] The present invention constructs a cross-modal infrared small target detection method based on degradation memory compensation, innovatively combines metric learning with a hybrid mask strategy, solves the problem of cross-modal small target information loss through a controllable degradation estimation space, and uses the global mask self-supervision feature to solve the core problem of "insufficient end-to-end supervision" in traditional cross-modal small target detection tasks; as far as we know, the method designed by the present invention is one of the few joint image fusion, metric space, image degradation, mask self-encoding and other methods in the field of cross-modal small target detection. It rethinks a way of thinking about cross-modal detection from a new perspective and achieves excellent results in various indicators; the cross-modal small target detection method proposed in the present invention can not only realize controllable degradation estimation, but also achieve the best effect in small target detection tasks under real environments; the existing technology (comparison method 1, DNANet solves the problem of infrared small target loss in deep networks through densely nested structures and attention mechanisms. Its core components include densely nested interaction modules (D NIM) and cascaded attention module (CSAM), DNIM retains the target information flow through the progressive interaction of multi-layer features, and CSAM uses channel-space dual-path attention to dynamically enhance features. This method effectively suppresses background noise and strengthens the characteristics of small targets through repeated cross-layer feature fusion and mathematical modeling of attention weight allocation, and achieves stable target detection in low signal-to-noise ratio scenarios; Compared with method 2, ISNet detects small targets through Taylor finite difference edge blocks and bidirectional attention fusion mechanism. The TFD edge block enhances multi-level edge features through mathematical differential operators, and the TOAA module performs attention weighting in the row and column directions to fuse low-level details with high-level semantics. This method innovatively combines the mathematical edge detection principle with deep learning. Through the direction-sensitive feature aggregation mechanism, it effectively suppresses complex background interference while maintaining the integrity of the target geometric structure, and significantly improves the positioning accuracy of the target contour. Qualitative comparison of target detection on the self-made data set of the present invention) and the method proposed in the present invention Figure 5 shown.
[0123] Under the same conditions, the feasibility and superiority of this method are further verified by calculating the relevant indicators of the cross-modal small target detection results obtained by the existing method; the schematic diagram of the comparison of evaluation indicators between the existing technology and the method proposed in the present invention is shown in Figure 6 As shown in the figure, it can be seen that the method proposed in the present invention has higher detection rate and recall rate, lower false alarm rate and false alarm rate, etc.; in the testing stage, the time complexity and space complexity of the detection method proposed in the present invention decreased by about 7.5% and 16.8% percentage points respectively.
[0124] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A cross-modal infrared small target detection method based on degraded memory compensation, characterized in that It includes the following steps: S1: Prepare the data sets, including data set one for training and data set two for testing. Data set one contains registered infrared and visible light image pairs, and data set two contains a self-made registered test set; S2: Construct a degradation memory compensation network, including a degradation estimation module and a memory compensation network; the degradation estimation module estimates the blur degradation score, noise degradation score, and visibility degradation score of infrared and visible light images through six metric spaces respectively, and outputs them in the form of heat maps; the memory compensation network generates a weight heat map and a weight vector through a fully connected layer, a Softmax layer, and a hist layer; S3: Train the degradation memory compensation network, generate dual-modal control group and consistency reinforcement group degradation samples through a degradation sample generation model, and optimize the network using an adaptive ranking loss function; S4: Cross-modal image region masking processing, divide the registered infrared and visible light images into image patches, randomly mask some image patches and replace them with mask signals; S5: Cross-modal image mixed masking processing, replace the masked infrared image patches with corresponding visible light image patches, and replace the masked visible light image patches with corresponding infrared image patches to generate a mixed mask image pair; S6: Construct a small target detection network, including a multi-level Swin Transformer encoder, a fusion strategy, and a decoder; the multi-level Swin Transformer encoder modulates and compensates cross-modal features using the weight heat map and weight vector of the memory compensation network; S7: Train the small target detection network, adopt an end-to-end framework, and optimize the network by combining the masked region loss and the Soft-IoU loss; S8: Test and evaluate the model, and output the detection results and performance indicators.
2. The method according to claim 1, wherein In step S1, data set one is the Anti-UAV data set, and data set two contains infrared and visible light image pairs of multiple scenarios, and all image pairs have completed fine-grained registration.
3. The method according to claim 1, wherein In step S2, the degradation estimation module includes: a blur degradation sub-network, a noise degradation sub-network, and a visibility degradation sub-network. Each sub-network is composed of a spatial and channel attention module, and outputs a degradation score heat map using the Sigmoid function; The noise degradation sub-network and the blur degradation sub-network are generated based on a probability degradation model, and the generation is expressed using the formula: where n and k represent the sizes of the noise kernel and the blur kernel respectively, and z n and z k represent noise and the blur kernel respectively, both following a normal distribution; M n and M k represent the noise and blur generation functions respectively; The visibility degradation sub-network is composed of a brightness and contrast degradation function; The memory compensation network consists of a fully connected layer, a Softmax, and a hist layer connected to the end of each DEM.
4. The method according to claim 1, wherein In step S3, the adaptive ranking loss function includes: Margin Ranking Loss for the dual-modal control group, which constrains the difference between the blur degradation and noise degradation scores, and is expressed as: Margin Ranking Loss for the consistency reinforcement group, which constrains the score difference between non-degraded samples and severely degraded samples; and is expressed as:
5. The method according to claim 1, wherein In step S4, the infrared and visible light images are respectively divided into 3×4 image patches, and 6 of them are randomly masked and replaced with mask signals.
6. The method according to claim 1, characterized in that, In S5, the infrared masked image after mixed masking processing is expressed as: The visible light mask image is represented as: Among them, {MASK} represents a mask signal, represents the image block of the original visible light image at (i, j), represents the image block of the original infrared image at (i, j).
7. The method according to claim 1, wherein In step S6, the small target detection network includes: Encoder: 5 levels of Patch Merging and 5 levels of Swin Transformer blocks, and the features output at each level are fused after being modulated by the memory compensation weights; Decoder: 5 levels of Patch Expanding and double ResBlock blocks, which are used to reconstruct the detection image with target boxes.
8. The method according to claim 1, wherein In step S7, the masking area loss function is: Among them, Denote the total number of pixels in the masked region to ensure uniform distribution of the loss; I f,i Denote the pixel value of the reconstructed image at the i-th masked position; I ir / vi,i is the true pixel value of the original image at the i-th masked position.
9. The method according to claim 1, characterized in that, In step S7, The Soft-IoU loss function is represented as: Among them, P i,j represents the predicted probability at the position (i, j); G i,j represents the true label at the pixel position (i, j); ε represents a smoothing term to prevent division-by-zero errors.
10. The method according to claim 1, characterized in that, In step S8, the performance metrics include the detection rate, false alarm rate, recall rate, IoU, and F1 score.
Citation Information
Patent Citations
Multi-modal image fusion method based on high-order degradation model
CN117197627A
Low-illumination image enhancement method based on adaptive detail compensation
CN117408919A
Person Re-Identification Method Combining Random Batch Mask and Multi-Scale Representation Learning
JP6830707B1