Infrared small target detection method based on semantic alignment fusion and content-driven dynamic enhancement
Through adaptive selection of image-level enhancement operation and semantic alignment and fusion methods, the accuracy of infrared small object detection is improved, and the problems of few pixels and blurred edges in infrared small object detection are solved, achieving better detection results.
Patent Information
- Application Number
- CN202510569948.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-15
AI Technical Summary
The existing infrared small object detection methods fail to effectively utilize image enhancement technology, resulting in poor detection results, especially in scenarios with complex backgrounds and diverse target sizes.
A method based on semantic alignment and fusion and content-driven dynamic enhancement is adopted to adaptively select image-level enhancement operations through an end-to-end method, combining original and enhanced image features, and using semantic alignment and fusion modules to perform feature fusion to improve detection performance.
The accuracy and detection performance of infrared small object detection is improved, the problems of few pixels and blurred edges are solved, the discriminant characteristics of the image are enhanced, and the image quality decline caused by inappropriate enhancement is avoided.
Smart Images

Figure CN120495819A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of infrared small target detection, and in particular relates to an infrared small target detection method based on semantic alignment fusion and content-driven dynamic enhancement. Background Art
[0002] Infrared small target detection has been widely used in military applications, such as submarine detection and infrared guidance, as well as civilian applications, including medical anomaly diagnosis and industrial nondestructive testing. However, infrared imaging primarily relies on the temperature distribution of an object's surface, rather than its color or edge details. Therefore, infrared images typically contain few surface texture features. Furthermore, compared to general targets, small infrared targets occupy a much smaller area within the entire infrared image, resulting in fewer pixels and blurred edges, making their detection difficult.
[0003] In order to detect small infrared targets, researchers have proposed many traditional algorithms to address the detection challenges, including methods based on target features, background suppression, and image structural elements. Specifically, methods based on target features utilize the feature differences between the target and its background in a single frame of infrared image. For example, the human visual system method makes full use of the difference between small targets and their neighbors to increase the contrast between the foreground and the background. Background suppression methods focus on minimizing the image background through filtering such as spatial domain filtering and transform domain filtering. Methods based on image data structure are implemented by utilizing different structural features in infrared images, such as the sparsity of the target and the low-rank nature of the background. However, these methods are usually heavily dependent on prior knowledge, sensitive to hyperparameters, and generally perform poorly in scenes involving complex backgrounds and images with diverse target sizes.
[0004] Recently, with the development of deep learning methods, numerous deep learning-based infrared small target detection (IRSTD) methods have been proposed and achieved state-of-the-art performance. ACM introduced an asymmetric context modulation mechanism to trade high-level semantics for low-level details. ISNet incorporates target shape reconstruction into small infrared target detection. RepISDNet proposes a simple reparameterized network to balance detection accuracy and inference efficiency. MSAFFNet proposes a spatial pyramid pooling module (ASPPM) to focus on the global context of the target and a dual attention module (DAM) to focus on regions of interest in both shallow and deep feature layers to improve small target detection. SRNet proposed a method for infrared small target detection that learns shape-biased representations and uses a small number of 9×9 convolutions in the encoder to extract reliable shape knowledge from infrared images. In general, these methods focus solely on enhancing the network's feature processing, such as employing multi-scale feature fusion strategies or enhancing specific features such as shape. However, direct detection without enhancement for infrared small targets with small pixel occupancy and blurred edges does not achieve optimal detection results. Furthermore, inappropriate image-level enhancement can lead to even worse image quality. Therefore, it is necessary to select appropriate image enhancement operations for each image before detection.
[0005] In summary, great progress has been made in theoretical research on non-infrared small target detection. However, many researchers have not considered the impact of image enhancement on the detection performance of infrared small target detection, resulting in direct detection failing to achieve the best detection effect.
[0006] After searching, no prior art documents identical or similar to the present invention were found. Summary of the Invention
[0007] To fully utilize the information in images, this paper proposes a "dynamic content-specific target visualization" method. This method utilizes a joint optimization scheme to achieve optimal image-level enhancement for each image in an end-to-end manner. Furthermore, a semantic alignment fusion method is proposed to fuse features from the original and enhanced images to produce more discriminative features, improving the performance of infrared small target detection. Specifically, this method first uses a fully differentiable module to perceive image scale, edge information, and other information. Based on the given input, the most appropriate hyperparameters are selected. The image is then adaptively amplified and edge-sharpened based on the selected hyperparameters, enhancing the original image. This method enriches and sharpens the image, addressing the issues of infrared small targets with fewer pixels and blurred edges, thereby achieving more accurate detection results. Furthermore, the original image contains certain features that are not present in the enhanced image. The fusion of these two image features produces more discriminative features, improving the performance of infrared small target detection. However, due to the adaptive upscaling of the original image, the enhanced image is different in size. Directly downsampling the enhanced image features to match the original image feature size would lose some critical information. Therefore, we propose a Semantic Alignment Fusion (SAF) module to align and fuse features via semantic awareness instead of direct downsampling, enhancing the integration of two different feature types.
[0008] The present invention solves the practical problem by adopting the following technical solutions:
[0009] 1. A method for infrared small target detection based on semantic alignment fusion and content-driven dynamic enhancement, comprising the following steps:
[0010] Step 1: Extract the size and edge information of the image;
[0011] Step 2: Based on the size and edge information extracted in step 1, the most suitable enhancement operation is adaptively selected to enhance the image.
[0012] Step 3: Based on the enhanced image obtained in step 2, different features are effectively fused through the semantic alignment fusion module to enhance the image information;
[0013] Step 4: Use the semantically aligned features from step 3 to detect small infrared targets through the decoder.
[0014] 2. The infrared small target detection method based on semantic alignment fusion and content-driven dynamic enhancement according to claim 1 is characterized in that the step 1 extracts the size and edge information in the image, including: extracting contextual information of different scales in the image through multiple dilated convolution operations of different scales in dilated spatial pyramid pooling, thereby effectively capturing multi-scale features; considering that low-level features contain rich edge details, while high-level features carry rich semantic and position information, we combine low-level features and high-level features to perceive edge information to model edge information related to the target.
[0015] In the dilated spatial pyramid pooling module, the formula for dilated convolution is:
[0016]
[0017] Where x(m,n) is the pixel value of the input image, w(i,j) is the convolution kernel, d is the dilation rate, and y(i,j) is the output. By adjusting the dilation rate d, the receptive field of the convolution kernel can be flexibly adjusted.
[0018] Atrous spatial pyramid pooling uses multiple convolutional layers with different atrous rates to extract contextual information of different scales. Feature maps of different scales are extracted layer by layer through convolution and fused at the end:
[0019] f scale =concat(f1,f2,f3,f4,f5)
[0020] Among them, f1, f2, f3, f4, and f5 represent the convolution results under different void rates, including a global average pooling layer (used to capture global context information), and these feature maps are fused together through the concatenation operation (concat).
[0021] In order to extract edge information, low-level features and high-level features are combined to extract it, which can be specifically expressed as;
[0022]
[0023] Among them, DP represents the downsampling operation, Cat represents the splicing operation, and f edge Represents edge features.
[0024] 3. The infrared small target detection method based on semantic alignment fusion and content-driven dynamic enhancement according to claim 1 is characterized in that step 2 adaptively selects the most suitable enhancement operation based on the size and edge information extracted in step 1 to enhance the image.
[0025] During parameter learning, the goal is to learn a distribution and sample a value from it, rather than directly learning a specific value. This enhances the model's generalization and prevents it from falling into local optima. During testing, when the network has fully converged, the learned parameters are fixed to their mean. The value sampled from the learned distribution becomes the amplification factor, which is then passed back to the amplification operation.
[0026] Use three 3×3 convolutional layers to reduce the feature f scale The dimension of , and change the number of channels. By reducing f scale The generated feature map is then flattened and further reduced in dimension through a linear layer to learn the mean μ1 and variance σ1. In order to solve the problems of memory resource consumption and sample values deviating too far from the mean μ, the mean μ is limited to [1, 2] and the variance σ is limited to [0, 0.1]:
[0027]
[0028] Where: f scale represents the scale feature extracted by the ASPP module, φ(★) represents the stacking of multiple convolutional layers, MLP represents the fully connected layer, clamp(★; [1, 2]) limits the result between 1 and 2, and clamp(★; [0, 0.1]) limits the result between 0 and 0.1.
[0029] To keep the model differentiable, we use reparameterization techniques to perform random sampling within the distribution. Finally, to keep the probability distribution between [1, 2], we use a truncation operation to limit the final magnification factor m to this range, and then we can get an adaptively magnified image based on the magnification factor. The above can be expressed as:
[0030] ε1~N(0,1)
[0031] m=clamp(μ1+ε1×σ1;[0,1])
[0032] I m =U(I;m)
[0033] Where: ε represents the value sampled from the standard normal distribution, U(★;m) is a bilinear interpolation operation and is differentiable, I m is the adaptive upscaling of the image.
[0034] Referring to the adaptively enlarged image above, we also learn a distribution from the edge features and randomly sample the sharpening coefficient w from it, which usually ranges from 0 to 1. The formula used is as follows:
[0035]
[0036] w=clamp(μ2+ε2×σ2; [0,1])
[0037] I e =USM(I m )
[0038] Where: f edge Represents the extracted edge features, I e Indicates that only the sharpened image is applied. USM is a common image sharpening operation, which is widely used in image enhancement due to its wide applicability and ability to preserve details. It is implemented by adding the high-frequency components of the image to the original image.
[0039] 4. The infrared small target detection method based on semantic alignment fusion and content-driven dynamic enhancement according to claim 1 is characterized in that, in step 3, considering that traditional methods such as maximum pooling and average pooling rely on inherent rules for feature fusion, which is not suitable for adaptive processing of magnified images, a semantic alignment fusion module is proposed to effectively fuse different features to enhance image information. The specific steps include:
[0040] 4.1 Enhanced Image Features of Input Processing by Adaptive Pooling (ADP) Get a feature map with the same size as the original The same feature map The formula is:
[0041]
[0042] ADP stands for Adaptive Pooling.
[0043] 4.2 Features Add to original image features In this process, the size of the feature map remains unchanged, but the number of channels is modified to 4, and the feature is softmaxed to obtain the weight W. 1 ∈R 4×H×W The formula is:
[0044]
[0045] Among them, Conv represents the convolution operation and softmax represents the activation function.
[0046] 4.3 Use pixel_shuffle operation to obtain adaptive weight W∈R 1×2H×2W , and then upsample and interpolate to fixed-size features Its size is the original image feature Twice, the formula is:
[0047] W=Pixel_shuffle(W1)
[0048]
[0049] Pixel_shuffle refers to the operation of increasing spatial resolution by reducing the number of channels, and UP stands for bilinear interpolation.
[0050] 4.4 After obtaining the adaptive weight W, we will feature Multiply these weights W and multiply by a factor of 4. Subsequently, we apply a 2×2 average pooling (AMP) operation to obtain the final dynamic aligned features The formula used is as follows:
[0051]
[0052] Among them, AMP stands for Average Mean Pooling
[0053] 4.5 First, dynamically align the features The channel dimension is concatenated with the original image features fi to form a comprehensive representation. These concatenated features are input into a specially designed attention generation network, which consists of a series of convolutional layers, batch normalization layers (BN), and activation functions (ReLU). Through these operations, a two-channel feature map is generated, and the feature map is normalized through the Softmax activation layer to ensure that the generated attention map A i The probability distribution property is satisfied at each position, and the formula is as follows:
[0054]
[0055] Among them, φ represents the stacked "Conv-BN-ReLU" operation, and Cat(★) represents the splicing operation.
[0056] 4.6 Based on the previous step, for each scale m∈O,K, obtain the attention map aligned with the original image features and attention maps aligned with dynamic alignment features Finally, the input features are weighted and summed together with the attention map to generate the final fusion feature f' i .
[0057]
[0058] in, and Represent the adaptive weights of the two features respectively.
[0059] 5. The infrared small target detection method based on semantic alignment fusion and content-driven dynamic enhancement according to claim 1 is characterized in that step 4 uses the features after semantic alignment in step 3 to realize the detection of infrared small targets through a decoder.
[0060] Advantages and beneficial effects of the present invention:
[0061] 1. This invention addresses the problems of previous methods, which either fail to utilize image-level enhancement strategies or uniformly use predefined image-level enhancement strategies, resulting in inappropriate image enhancement. Both these approaches make it difficult to detect small infrared targets. To address this, the present invention proposes a new solution for infrared small target detection. This solution utilizes a joint optimization approach to achieve both enhancement and detection of small infrared targets. It adaptively learns hyperparameters based on image scale, edge information, and other information to select the appropriate image-level enhancement strategy for each image.
[0062] 2. This paper addresses the problem of poor visualization in infrared small target detection by introducing an image-level enhancement strategy, which improves the detection performance of the model. Furthermore, this excellent detection performance overcomes the shortcomings of existing work and enhances its practical application value.
[0063] 3. Unlike previous multi-scale feature fusion methods that directly downsample large-scale features or small-scale features to the same dimension, the present invention proposes a new fusion method SAF module, which first performs pixel-level semantic alignment on the two features of different scales and then fuses them. This will avoid semantic mismatch and better preserve the detailed information and semantic consistency of the features. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is an overall schematic diagram of the dynamic content specific target display method of the present invention;
[0065] Figure 2 is a schematic diagram of a dynamic content specific target enhancement module of the present invention;
[0066] Figure 3 is a schematic diagram of the semantic space alignment module of the present invention;
[0067] Figure 4 Schematic diagram of the cross-scale feature fusion module of the present invention. DETAILED DESCRIPTION
[0068] The embodiments of the present invention are further described below in conjunction with the accompanying drawings:
[0069] 1. A method for infrared small target detection based on semantic alignment fusion and content-driven dynamic enhancement, comprising the following steps:
[0070] Step 1: Extract the size and edge information of the image;
[0071] Step 2: Based on the size and edge information extracted in step 1, the most suitable enhancement operation is adaptively selected to enhance the image.
[0072] Step 3: Based on the enhanced image obtained in step 2, different features are effectively fused through the semantic alignment fusion module to enhance the image information;
[0073] Step 4: Use the semantically aligned features from step 3 to detect small infrared targets through the decoder.
[0074] 2. Extracting size and edge information from the image in step 1 includes: extracting contextual information of different scales in the image through multiple dilated convolution operations of different scales in dilated spatial pyramid pooling, thereby effectively capturing multi-scale features; considering that low-level features contain rich edge details, while high-level features carry rich semantic and position information, we combine low-level features and high-level features to perceive edge information and model edge information related to the target.
[0075] In the dilated spatial pyramid pooling module, the formula for dilated convolution is:
[0076]
[0077] Where x(m,n) is the pixel value of the input image, w(i,j) is the convolution kernel, d is the dilation rate, and y(i,j) is the output. By adjusting the dilation rate d, the receptive field of the convolution kernel can be flexibly adjusted.
[0078] Atrous spatial pyramid pooling uses multiple convolutional layers with different atrous rates to extract contextual information of different scales. Feature maps of different scales are extracted layer by layer through convolution and fused at the end:
[0079] f scale =concat(f1,f2,f3,f4,f5)
[0080] Among them, f1, f2, f3, f4, and f5 represent the convolution results under different void rates, including a global average pooling layer (used to capture global context information), and these feature maps are fused together through the concatenation operation (concat).
[0081] In order to extract edge information, low-level features and high-level features are combined to extract it, which can be specifically expressed as;
[0082]
[0083] Among them, DP represents the downsampling operation, Cat represents the splicing operation, and f edge Represents edge features.
[0084] 3. Step 2 adaptively selects the most suitable enhancement operation based on the size and edge information extracted in step 1 to enhance the image.
[0085] During parameter learning, the goal is to learn a distribution and sample a value from it, rather than directly learning a specific value. This enhances the model's generalization and prevents it from falling into local optima. During testing, when the network has fully converged, the learned parameters are fixed to their mean. The value sampled from the learned distribution becomes the amplification factor, which is then passed back to the amplification operation.
[0086] Use three 3×3 convolutional layers to reduce the feature f scale The dimension of , and change the number of channels. By reducing f scale The generated feature map is then flattened and further reduced in dimension through a linear layer to learn the mean μ1 and variance σ1. In order to solve the problems of memory resource consumption and sample values deviating too far from the mean μ, the mean μ is limited to [1, 2] and the variance σ is limited to [0, 0.1]:
[0087]
[0088] Where: f scale represents the scale feature extracted by the ASPP module, φ(★) represents the stacking of multiple convolutional layers, MLP represents the fully connected layer, clamp(★; [1, 2]) limits the result between 1 and 2, and clamp(★; [0, 0.1]) limits the result between 0 and 0.1.
[0089] To keep the model differentiable, we use reparameterization techniques to perform random sampling within the distribution. Finally, to keep the probability distribution between [1, 2], we use a truncation operation to limit the final magnification factor m to this range, and then we can get an adaptively magnified image based on the magnification factor. The above can be expressed as:
[0090] ε1~N(0,1)
[0091] m=clamp(μ1+ε1×σ1;[0,1])
[0092] I m =U(I;m)
[0093] Where: ε represents the value sampled from the standard normal distribution, U(★;m) is a bilinear interpolation operation and is differentiable, I m is the adaptive upscaling of the image.
[0094] Referring to the adaptively enlarged image above, we also learn a distribution from the edge features and randomly sample the sharpening coefficient w from it, which usually ranges from 0 to 1. The formula used is as follows:
[0095]
[0096]
[0097] w=clamp(μ2+ε2×σ2; [0,1])
[0098] I e =USM(I m )
[0099] Where: f edge Represents the extracted edge features, I e Indicates that only the sharpened image is applied. USM is a common image sharpening operation, which is widely used in image enhancement due to its wide applicability and ability to preserve details. It is implemented by adding the high-frequency components of the image to the original image.
[0100] 4. In step 3, considering that traditional methods such as maximum pooling and average pooling rely on inherent rules for feature fusion, which is not suitable for adaptively processing the enlarged image, a semantic alignment fusion module is proposed to effectively fuse different features to enhance image information. The specific steps include:
[0101] 4.1 Enhanced Image Features of Input Processing by Adaptive Pooling (ADP) Get a feature map with the same size as the original The same feature map The formula is:
[0102]
[0103] ADP stands for Adaptive Pooling.
[0104] 4.2 Features Add to original image features In this process, the size of the feature map remains unchanged, but the number of channels is modified to 4, and the feature is softmaxed to obtain the weight W. 1 ∈R 4×H×W The formula is:
[0105]
[0106] Among them, Conv represents the convolution operation and softmax represents the activation function.
[0107] 4.3 Use pixel_shuffle operation to obtain adaptive weight W∈R 1×2H×2W , and then upsample and interpolate to fixed-size features Its size is the original image feature Twice, the formula is:
[0108] W=Pixel_shuffle(W1)
[0109]
[0110] Pixel_shuffle refers to the operation of increasing spatial resolution by reducing the number of channels, and UP stands for bilinear interpolation.
[0111] 4.4 After obtaining the adaptive weight W, we will feature Multiply these weights W and multiply by a factor of 4. Subsequently, we apply a 2×2 average pooling (AMP) operation to obtain the final dynamic aligned features The formula used is as follows:
[0112]
[0113] Among them, AMP stands for Average Mean Pooling
[0114] 4.5 First, dynamically align the features The channel dimension is concatenated with the original image features fi to form a comprehensive representation. These concatenated features are input into a specially designed attention generation network, which consists of a series of convolutional layers, batch normalization layers (BN), and activation functions (ReLU). Through these operations, a two-channel feature map is generated, and the feature map is normalized through the Softmax activation layer to ensure that the generated attention map A i The probability distribution property is satisfied at each position, and the formula is as follows:
[0115]
[0116] Among them, φ represents the stacked "Conv-BN-ReLU" operation, and Cat(★) represents the splicing operation.
[0117] 4.6 Based on the previous step, for each scale m∈O,K, obtain the attention map aligned with the original image features and attention maps aligned with dynamic alignment features Finally, the input features are weighted and summed together with the attention map to generate the final fusion feature f' i .
[0118]
[0119] in, and Represent the adaptive weights of the two features respectively.
[0120] 4. Step 4 uses the semantically aligned features from step 3 to detect small infrared targets through a decoder.
[0121] In this embodiment, the overall schematic diagram of the dynamic content specific target display method is as follows: Figure 1 As shown, t includes the following groups Components: two U-net encoders, one U-net decoder, and Dynamic Content- SpecificObject Enhancement, DCE) module and semantic alignment fusion module.
[0122] The dynamic content specific target enhancement module constructed by the present invention is as follows Figure 2 As shown, it consists of two submodules to solve Solve the specific problem of infrared small targets: 1) Adaptive amplification enhancement = module to solve the problem of insufficient pixels of infrared small targets problem; 2) Dynamic edge-aware enhancement module to solve the problem of sharpening and enhance edge details in the image.
[0123] In addition, the semantic space alignment module such as Figure 3 As shown in the figure, on the one hand, it upsamples the enhanced features to the original feature map On the other hand, it uses these two features to generate a weight map. By combining the weight map with the upsampled features Multiply them together and use max pooling to adaptively select pixel values from 2x2 pixel blocks to obtain the aligned fusion features.
[0124] Figure 4 The figure shows a schematic diagram of cross-scale feature fusion. The core idea is to adaptively adjust the The combination of original image features and dynamically aligned features can improve the accuracy of feature fusion.
[0125] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0126] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A method for infrared small target detection based on semantic alignment fusion and content-driven dynamic enhancement, characterized by: The following steps are involved: Step 1: Extract the size and edge information of the image; Step 2: Based on the size and edge information extracted in step 1, the most suitable enhancement operation is adaptively selected to enhance the image. Step 3: Based on the enhanced image obtained in step 2, different features are effectively fused through the semantic alignment fusion module to enhance the image information; Step 4: Use the semantically aligned features from step 3 to detect small infrared targets through the decoder.
2. The infrared small target detection method based on semantic alignment fusion and content-driven dynamic enhancement according to claim 1 is characterized in that: The step 1 extracts the size and edge information from the image, including: extracting contextual information of different scales in the image through multiple dilated convolution operations of different scales in dilated spatial pyramid pooling, thereby effectively capturing multi-scale features; considering that low-level features contain rich edge details, while high-level features carry rich semantic and position information, we combine low-level features and high-level features to perceive edge information and model edge information related to the target. In the dilated spatial pyramid pooling module, the formula for dilated convolution is: Where x(m,n) is the pixel value of the input image, w(i,j) is the convolution kernel, d is the dilation rate, and y(i,j) is the output. By adjusting the dilation rate d, the receptive field of the convolution kernel can be flexibly adjusted. Atrous spatial pyramid pooling uses multiple convolutional layers with different atrous rates to extract contextual information of different scales. Feature maps of different scales are extracted layer by layer through convolution and fused at the end: f scale =concat(f1,f2,f3,f4,f5) Among them, f1, f2, f3, f4, and f5 represent the convolution results under different void rates, including a global average pooling layer (used to capture global context information), and these feature maps are fused together through the concatenation operation (concat). In order to extract edge information, low-level features and high-level features are combined to extract it, which can be specifically expressed as; Among them, DP represents the downsampling operation, Cat represents the splicing operation, and f edge Represents edge features.
3. The infrared small target detection method based on semantic alignment fusion and content-driven dynamic enhancement according to claim 1 is characterized in that: The step 2 adaptively selects the most suitable enhancement operation based on the size and edge information extracted in the step 1 to enhance the image. During parameter learning, the goal is to learn a distribution and sample a value from it, rather than directly learning a specific value. This enhances the model's generalization and prevents it from falling into local optima. During testing, when the network has fully converged, the learned parameters are fixed to their mean. The value sampled from the learned distribution becomes the amplification factor, which is then passed back to the amplification operation. Use three 3×3 convolutional layers to reduce the feature f scale The dimension of , and change the number of channels. By reducing f scale The generated feature map is then flattened and further reduced in dimension through a linear layer to learn the mean μ1 and variance σ1. In order to solve the problems of memory resource consumption and sample values deviating too far from the mean μ, the mean μ is limited to [1, 2] and the variance σ is limited to [0, 0.1]: Where: f scale represents the scale feature extracted by the ASPP module, φ(★) represents the stacking of multiple convolutional layers, MLP represents the fully connected layer, clamp(★; [1, 2]) limits the result between 1 and 2, and clamp(★; [0, 0.1]) limits the result between 0 and 0.
1. To keep the model differentiable, we use reparameterization techniques to perform random sampling within the distribution. Finally, to keep the probability distribution between [1, 2], we use a truncation operation to limit the final magnification factor m to this range, and then we can get an adaptively magnified image based on the magnification factor. The above can be expressed as: ε1~N(0,1) m=clamp(μ1+ε1×σ1;[0,1]) I m =U(I;m) Where: ε represents the value sampled from the standard normal distribution, U(★;m) is a bilinear interpolation operation and is differentiable, I m is the adaptive upscaling of the image. Referring to the adaptively enlarged image above, we also learn a distribution from the edge features and randomly sample the sharpening coefficient w from it, which usually ranges from 0 to 1. The formula used is as follows: w=clamp(μ2+ε2×σ2; [0,1]) I e =USM(I m ) Where: f edge Represents the extracted edge features, I e Indicates that only the sharpened image is applied. USM is a common image sharpening operation, which is widely used in image enhancement due to its wide applicability and ability to preserve details. It is implemented by adding the high-frequency components of the image to the original image.
4. The infrared small target detection method based on semantic alignment fusion and content-driven dynamic enhancement according to claim 1 is characterized in that: In step 3, considering that traditional methods such as maximum pooling and average pooling rely on inherent rules for feature fusion, which is not suitable for adaptively processing the enlarged image, a semantic alignment fusion module is proposed to effectively fuse different features. To enhance image information, the specific steps include: 4.1 Enhanced Image Features of Input Processing by Adaptive Pooling (ADP) Get a feature map with the same size as the original The same feature map The formula is: ADP stands for Adaptive Pooling. 4.2 Features Add to original image features In this process, the size of the feature map remains unchanged, but the number of channels is modified to 4, and the feature is softmaxed to obtain the weight W. 1 ∈R 4×H×W The formula is: Among them, Conv represents the convolution operation and softmax represents the activation function. 4.3 Use pixel_shuffle operation to obtain adaptive weight W∈R 1×2H×2W , and then upsample and interpolate to fixed-size features Its size is the original image feature Twice, the formula is: W=Pixel_shuffle(W1) Pixel_shuffle refers to the operation of increasing spatial resolution by reducing the number of channels, and UP stands for bilinear interpolation. 4.4 After obtaining the adaptive weight W, we will feature Multiply these weights W and multiply by a factor of 4. Subsequently, we apply a 2×2 average pooling (AMP) operation to obtain the final dynamic aligned features The formula used is as follows: in, AMP stands for Average Mean Pooling 4.5 First, dynamically align the features The channel dimension is concatenated with the original image features fi to form a comprehensive representation. These concatenated features are input into a specially designed attention generation network, which consists of a series of convolutional layers, batch normalization layers (BN), and activation functions (ReLU). Through these operations, a two-channel feature map is generated, and the feature map is normalized through the Softmax activation layer to ensure that the generated attention map A i The probability distribution property is satisfied at each position, and the formula is as follows: in, φ represents the stacked "Conv-BN-ReLU" operation, and Cat(★) represents the concatenation operation. 4.6 Based on the previous step, for each scale m∈O,K, obtain the attention map aligned with the original image features and attention maps aligned with dynamic alignment features Finally, the input features are weighted and summed together with the attention map to generate the final fusion feature f' i . in, and Represent the adaptive weights of the two features respectively.
5. The infrared small target detection method based on semantic alignment fusion and content-driven dynamic enhancement according to claim 1 is characterized in that: The step 4 uses the semantically aligned features in step 3 to detect small infrared targets through a decoder.
Citation Information
Cited By
Target tracking method and system for resisting unmanned aerial vehicle
CN121527654A