Image super-resolution adversarial defense method and system with position attention enhancement

By employing a location-attention-enhanced image super-resolution adversarial defense method, which utilizes region partitioning of frequency domain images and an adaptive mask module to adjust the self-attention mechanism of the SwinIR network, the robustness problem of remote sensing images under aggressive noise is solved, and high-quality remote sensing image reconstruction is achieved.

CN120746828BActive Publication Date: 2025-11-11XIAN AERONAUTICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511208879.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-11
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing super-resolution reconstruction techniques for remote sensing images lack robustness when facing aggressive noise, leading to artifacts and structural distortions in the reconstructed images, which affects the accuracy of subsequent analysis. In particular, the Transformer network model lacks effective adversarial robustness.

Method used

An image super-resolution adversarial defense method enhanced by location attention, including an adaptive adversarial noise random mask module and a self-attention mechanism model with an adjusted window self-attention layer, utilizes the low-frequency absolute protection, mid-frequency transition and high-frequency suppression regions of the frequency domain image to construct a baseline SwinIR network for super-resolution reconstruction of remote sensing images.

Benefits of technology

It improves the accuracy and robustness of remote sensing image reconstruction, reduces the impact of noise interference, ensures the integrity of complex texture restoration and edge sharpness, and enhances the model's noise resistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120746828B_ABST
    Figure CN120746828B_ABST
Patent Text Reader

Abstract

This invention relates to the field of remote sensing image processing technology, specifically to a location-attention-enhanced image super-resolution adversarial defense method and system. The method includes: acquiring the weight factor of the mask value corresponding to each location point in the remote sensing frequency domain image; acquiring a low-frequency threshold and a Gaussian attenuation width threshold; obtaining a low-frequency absolute protection zone, a mid-frequency transition zone, and a high-frequency suppression zone based on the low-frequency threshold and the Gaussian attenuation width threshold, and obtaining a low-frequency mask matrix, a mid-frequency mask matrix, a high-frequency mask matrix, and a noise sensitivity map to obtain a frequency-domain cleaned remote sensing image; constructing an adaptive adversarial noise random mask module from the process of acquiring the frequency-domain cleaned remote sensing image; constructing a baseline SwinIR network through the adaptive adversarial noise random mask module, and adjusting the self-attention mechanism model in the window self-attention layer of the baseline SwinIR network to obtain the final reconstructed remote sensing image. This invention improves the adversarial robustness of the model during remote sensing image reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing technology, and in particular to a position attention-enhanced image super-resolution adversarial defense method and system. Background Technology

[0002] With the development of remote sensing technology, the use of remote sensing images has become a common technique. However, when performing identification and analysis using remote sensing images, the clarity of the acquired images is insufficient due to noise interference, transmission conditions, and limitations of imaging equipment. This severely restricts the application potential of remote sensing images in civilian and military fields, especially in satellite and aerial imaging. Super-resolution reconstruction technology can address the problem of insufficient image clarity; however, conventional super-resolution reconstruction techniques are not very effective in restoring image details and clarity due to high hardware costs and complex natural environmental interference. In recent years, deep learning-based remote sensing image super-resolution reconstruction technology has gradually become the mainstream solution. Through hardware and software decoupling, it significantly improves the reconstruction quality of remote sensing images.

[0003] However, when using deep learning-based remote sensing image super-resolution reconstruction techniques to improve resolution, they are susceptible to carefully designed adversarial noise by attackers. This poses a serious threat to neural network models in deep learning. Attackers can add tiny adversarial perturbations to low-resolution images, causing the super-resolution model to output incorrect or distorted high-resolution images during reconstruction. Adversarial noise not only directly reduces model performance, leading to artifacts and structural distortions in the reconstructed images, but also destroys the integrity of feature extraction. Noise interference forces the model to rely on less contaminated local features for reconstruction, failing to effectively integrate global contextual information of the image. This ultimately results in incomplete recovery of complex textures, decreased edge sharpness, and other defects, seriously affecting the accuracy of subsequent analysis.

[0004] Currently, Transformer network models have achieved better performance than convolutional neural networks (CNNs) in most vision tasks. However, conventional techniques focus on the adversarial robustness of CNNs on popular threat models, and existing Transformer network models mainly focus on standard accuracy and computational cost, lacking research on the fundamental impact on model reliability and generalization. As a novel model architecture, directly applying adversarial defense methods from CNNs to Transformer networks does not provide effective robustness. Therefore, it is necessary to study the adversarial robustness of Transformer models from the perspective of their unique architectural components. Summary of the Invention

[0005] This invention provides a position-attention-enhanced image super-resolution adversarial defense method and system to solve existing problems.

[0006] The objective of this invention can be achieved through the following technical solutions:

[0007] The first aspect of this invention is to provide a position-attention-enhanced image super-resolution adversarial defense method, comprising:

[0008] Acquire remote sensing images after introducing adversarial noise;

[0009] The remote sensing image is transformed from the spatial domain to the frequency domain to obtain a remote sensing frequency domain image. The weighting factor for the mask value corresponding to each location point in the remote sensing frequency domain image is obtained. The low-frequency threshold and Gaussian attenuation width threshold are obtained. Based on the low-frequency threshold and Gaussian attenuation width threshold, the low-frequency absolute protection zone, the mid-frequency transition zone, and the high-frequency suppression zone in the remote sensing frequency domain image are determined. The low-frequency mask matrix is ​​obtained based on the mask values ​​of all locations in the low-frequency absolute protection zone. The mid-frequency mask is obtained based on the weighting factor, low-frequency threshold, and Gaussian attenuation width threshold for the mask value corresponding to each location point in the mid-frequency transition zone. The process involves: determining a noise sensitivity map based on the frequency domain amplitude in the remote sensing frequency domain image; obtaining a high-frequency mask matrix based on the mask values ​​of all points in the high-frequency suppression region and the noise sensitivity map; determining the final mask matrix using the low-frequency mask matrix, mid-frequency mask matrix, high-frequency mask matrix, and noise sensitivity map; obtaining an adjusted remote sensing image based on the final mask matrix and the remote sensing frequency domain image; fusing the adjusted remote sensing image and the original remote sensing image to obtain a frequency-domain purified remote sensing image; and constructing an adaptive adversarial noise random mask module for the process of acquiring the frequency-domain purified remote sensing image.

[0010] A baseline SwinIR network is constructed using an adaptive adversarial noise random mask module. The self-attention mechanism model in the window self-attention layer of the baseline SwinIR network is adjusted to obtain a new self-attention mechanism model in the window self-attention layer. The frequency domain cleaned remote sensing image is uniformly divided to obtain several non-overlapping windows. Each window in the frequency domain cleaned remote sensing image is used as input. The final reconstructed remote sensing image is obtained through the baseline SwinIR network and the new self-attention mechanism model in the window self-attention layer.

[0011] Furthermore, the step of converting the remote sensing image from the spatial domain to the frequency domain to obtain a remote sensing frequency domain image includes:

[0012] The spatial domain of the remote sensing image is converted to the frequency domain by discrete cosine transform, thus obtaining the frequency domain image of the remote sensing image.

[0013] Further, the weighting factor for obtaining the mask value corresponding to each location point in the remote sensing frequency domain image includes:

[0014] The top left corner of the remote sensing frequency domain image is recorded as the origin, and a reference coordinate system is constructed with the horizontal axis pointing to the right and the vertical axis pointing downwards.

[0015]

[0016]

[0017] In the formula, Represents the horizontal length of a remote sensing frequency domain image. Represents the vertical length of the remote sensing frequency domain image. This represents the maximum distance between all points in the remote sensing frequency domain image and the origin. Represents the location points in the remote sensing frequency domain image The distance to the origin Represents the location points in the remote sensing frequency domain image Weighting factors corresponding to mask values; location points in remote sensing frequency domain images The horizontal axis is represented as The vertical axis is The location point.

[0018] Further, the acquisition of the low-frequency threshold and the Gaussian attenuation width threshold; and the determination of the low-frequency absolute protection zone, the mid-frequency transition zone, and the high-frequency suppression zone in the remote sensing frequency domain image based on the low-frequency threshold and the Gaussian attenuation width threshold, include:

[0019]

[0020]

[0021] In the formula, and This represents the learnable and trainable hyperparameters. Indicates the low-frequency threshold. This represents the threshold width of the Gaussian decay. This represents the activation function, used for data normalization.

[0022] Will The area consisting of all corresponding location points is designated as the low-frequency absolute protection zone; The region consisting of all corresponding location points is denoted as the intermediate frequency transition region; The region consisting of all the corresponding location points is denoted as the high-frequency suppression region.

[0023] Further, the step of obtaining a low-frequency mask matrix based on the mask values ​​of all locations within the low-frequency absolute protection zone, and obtaining a mid-frequency mask matrix based on the weighting factor, low-frequency threshold, and Gaussian attenuation width threshold of the mask value corresponding to each location in the mid-frequency transition zone, includes:

[0024] The mask values ​​of all locations within the low-frequency absolute protection zone are assigned as 1, and the mask values ​​of all other locations outside the low-frequency absolute protection zone are assigned as 0, forming a matrix, denoted as the low-frequency mask matrix.

[0025]

[0026] In the formula, Indicates the location point in the intermediate frequency transition region. The mask value, Represents the location points in the remote sensing frequency domain image The weighting factor corresponding to the mask value, Indicates the low-frequency threshold. This represents the threshold width of the Gaussian decay. Represents an exponential function with the natural constant as its base;

[0027] The mask values ​​of all other locations outside the intermediate frequency transition region are assigned to 0. Then, the mask values ​​of all locations in the intermediate frequency transition region are combined to form a matrix, which is called the intermediate frequency mask matrix.

[0028] Further, the step of determining the noise sensitivity map based on the frequency domain amplitude in the remote sensing frequency domain image; and obtaining the high-frequency mask matrix based on the mask values ​​of all locations in the high-frequency suppression region and the noise sensitivity map, includes:

[0029]

[0030] In the formula, This represents the frequency domain amplitude in a remotely sensed frequency domain image. Represents the absolute value symbol. Represents the inverse discrete cosine transform. This indicates that remote sensing images are estimated using an estimator combined with a lightweight CNN network. Represents a noise sensitivity map;

[0031]

[0032] In the formula, Indicates the location point in the high-frequency suppression region The mask value, This indicates the location point in the corresponding high-frequency suppression region of the noise sensitivity diagram. The noise sensitivity value;

[0033] The mask values ​​of all other locations outside the high-frequency suppression region are assigned to 0. Then, the mask values ​​of all locations within the high-frequency suppression region are combined to form a matrix, which is denoted as the high-frequency mask matrix.

[0034] Further, the final mask matrix is ​​determined using a low-frequency mask matrix, a mid-frequency mask matrix, a high-frequency mask matrix, and a noise sensitivity map; an adjusted remote sensing image is obtained based on the final mask matrix and the remote sensing frequency domain image; the adjusted remote sensing image and the remote sensing image are fused to obtain a frequency-domain purified remote sensing image, including:

[0035]

[0036] In the formula, Represents the low-frequency mask matrix. Represents the intermediate frequency mask matrix. Represents the high-frequency mask matrix. The matrix represents the noise sensitivity values ​​of all points in the noise sensitivity map; ⊙ represents the Hadamard product. This represents the final mask matrix;

[0037] The adjusted remote sensing frequency domain image is obtained by using the Hadamard product between the final mask matrix and the corresponding values ​​of the remote sensing frequency domain image; the adjusted remote sensing image is then obtained by performing an inverse discrete cosine transform on the adjusted remote sensing frequency domain image.

[0038]

[0039] In the formula, This indicates adjusting the pixel values ​​in the remote sensing image. Represents the pixel values ​​in a remotely sensed image. This represents the pixel values ​​in the frequency domain purified remote sensing image. This indicates the preset scaling factor.

[0040] Further, the baseline SwinIR network is constructed using an adaptive adversarial noise random mask module. The self-attention mechanism model in the window self-attention layer of the baseline SwinIR network is adjusted to obtain a new self-attention mechanism model in the window self-attention layer. The frequency domain cleaned remote sensing image is uniformly divided to obtain several non-overlapping windows. Each window in the frequency domain cleaned remote sensing image is used as input. Through the baseline SwinIR network and the new self-attention mechanism model in the window self-attention layer, the final reconstructed remote sensing image is obtained, including:

[0041] Each window in the frequency-domain cleansed remote sensing image is used as input, and super-resolution reconstruction is performed through the baseline SwinIR network. During the reconstruction process of the baseline SwinIR network, convolutional layers are used to extract shallow features. Then, based on the shallow features, the residual Swin Transformer module RSTB is used to extract deep features. Finally, based on the deep features, PixelShuffle is used to perform upsampling to complete the super-resolution reconstruction and obtain the final reconstructed remote sensing image.

[0042] Each RSTB consists of a window self-attention layer, layer normalization, and a multilayer perceptron; and an adaptive adversarial noise random mask module is placed at the head and tail of each RSTB.

[0043] The self-attention mechanism model in the new window self-attention layer is specifically represented as follows:

[0044]

[0045] In the formula, Represents the query vector. Represents the key vector. Represents a value vector. The dimension of the key vector. Represents the sparse position importance matrix. The notation for calculating the sparse position importance matrix in a computer. This represents the adjusted vector. This represents the transpose of the key vector. This represents the activation function.

[0046] A second aspect of the present invention is to provide a position-attention-enhanced image super-resolution adversarial defense system, comprising:

[0047] Image acquisition module: used to acquire remote sensing images after introducing adversarial noise;

[0048] Image Analysis and Adjustment Module: This module performs spatial-domain to frequency-domain conversion on remote sensing images to obtain a frequency-domain image. It acquires the weighting factor for the mask value corresponding to each location point in the frequency-domain image; obtains the low-frequency threshold and Gaussian attenuation width threshold; determines the low-frequency absolute protection zone, mid-frequency transition zone, and high-frequency suppression zone in the frequency-domain image based on the low-frequency threshold and Gaussian attenuation width threshold; obtains the low-frequency mask matrix based on the mask values ​​of all locations in the low-frequency absolute protection zone; and acquires the low-frequency mask matrix based on the weighting factor, low-frequency threshold, and Gaussian attenuation width threshold for each location point in the mid-frequency transition zone. The process involves: obtaining the intermediate frequency mask matrix; determining the noise sensitivity map based on the frequency domain amplitude in the remote sensing frequency domain image; obtaining the high-frequency mask matrix based on the mask values ​​of all locations in the high-frequency suppression region and the noise sensitivity map; determining the final mask matrix using the low-frequency mask matrix, intermediate frequency mask matrix, high-frequency mask matrix, and noise sensitivity map; obtaining an adjusted remote sensing image based on the final mask matrix and the remote sensing frequency domain image; fusing the adjusted remote sensing image and the original remote sensing image to obtain a frequency-domain purified remote sensing image; and constructing an adaptive adversarial noise random mask module for the process of obtaining the frequency-domain purified remote sensing image.

[0049] Image reconstruction module: This module constructs a baseline SwinIR network using an adaptive adversarial noise random mask module, adjusts the self-attention mechanism model in the window self-attention layer of the baseline SwinIR network to obtain a new self-attention mechanism model in the window self-attention layer, uniformly divides the frequency domain cleaned remote sensing image to obtain several non-overlapping windows, uses each window in the frequency domain cleaned remote sensing image as input, and obtains the final reconstructed remote sensing image through the baseline SwinIR network and the new self-attention mechanism model in the window self-attention layer.

[0050] A third aspect of the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the aforementioned position attention-enhanced image super-resolution adversarial defense method.

[0051] Compared with existing technologies, the beneficial effects of this invention are: obtaining the weighting factor of the mask value corresponding to each location point in the remote sensing frequency domain image, improving the accuracy of noise-affected analysis of each location point in the remote sensing frequency domain image; obtaining the low-frequency threshold and Gaussian attenuation width threshold, determining the low-frequency absolute protection zone, mid-frequency transition zone, and high-frequency suppression zone in the remote sensing frequency domain image based on the low-frequency threshold and Gaussian attenuation width threshold, and obtaining the low-frequency mask matrix, mid-frequency mask matrix, high-frequency mask matrix, and noise sensitivity map to obtain a frequency-domain purified remote sensing image; reducing the degree of noise interference by processing different affected areas to different degrees through regional division; constructing an adaptive adversarial noise random mask module; and constructing an adaptive adversarial noise random mask module. A baseline SwinIR network was established, which improved the accuracy of SwinIR network in remote sensing image reconstruction. The self-attention mechanism model in the window self-attention layer of the baseline SwinIR network was adjusted to obtain a new self-attention mechanism model in the window self-attention layer, which improved the accuracy of position importance evaluation and reduced the impact of defects such as incomplete restoration of complex textures and decreased edge sharpness. The frequency domain cleaned remote sensing image was uniformly divided to obtain several non-overlapping windows. Each window in the frequency domain cleaned remote sensing image was used as input. Through the baseline SwinIR network and the new self-attention mechanism model in the window self-attention layer, the final reconstructed remote sensing image was obtained, which improved the adversarial robustness of the model in the remote sensing image reconstruction process. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 This invention provides a step-by-step flowchart of an image super-resolution adversarial defense method with enhanced positional attention.

[0054] Figure 2 This invention provides a schematic diagram of the module flow of an image super-resolution adversarial defense system with enhanced positional attention.

[0055] Figure 3 This is a schematic diagram illustrating the effects of different super-resolution models on remote sensing images subjected to a basic attack.

[0056] Figure 4 This is a schematic diagram illustrating the effects of different super-resolution models on remote sensing images subjected to general attacks.

[0057] Figure 5This is a schematic diagram of the ablation experiment effect on a remote sensing image subjected to a base attack of strength 8 / 255. Detailed Implementation

[0058] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0059] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0060] To address the problems existing in the background technology, a position attention-enhanced image super-resolution adversarial defense method and system were designed, which has important practical significance.

[0061] like Figure 1 As shown, the first aspect of the present invention is to provide a position-attention-enhanced image super-resolution adversarial defense method, comprising the following steps:

[0062] Step S001: Acquire the initial remote sensing image and obtain the remote sensing image after introducing adversarial noise.

[0063] It should be noted that in order to address the impact of adversarial interference from attackers, it is necessary to acquire remote sensing images and add adversarial noise of varying intensities to these images for experimental analysis.

[0064] Specifically, initial remote sensing images were acquired; these initial images consisted of 2100 aerial scene images from the United States Geological Survey (USGS) national map. Adversarial noise was generated in the initial remote sensing images using the fast gradient sign method, resulting in remote sensing images with adversarial noise introduced. The fast gradient sign method is a well-known technique and will not be described in detail here.

[0065] Thus, we obtain the remote sensing image after introducing adversarial noise.

[0066] Step S002: By converting the remote sensing image into the frequency domain space, a remote sensing frequency domain image is obtained. The remote sensing frequency domain image is divided into regions based on the feature distribution in the remote sensing frequency domain image. Different degrees of adjustment are made to the divided regions to obtain a frequency domain purified remote sensing image and construct an adaptive adversarial noise random mask module.

[0067] It should be noted that due to the influence of sensors, atmospheric interference, and the effects of transmission and storage during the acquisition of remote sensing images, the acquired remote sensing images are subject to significant noise interference. This noise interference is usually high-frequency noise. Therefore, the remote sensing images in the spatial domain can be converted into the frequency domain for analysis, and preliminary processing can be performed using the corresponding image information in the frequency domain.

[0068] Specifically, the remote sensing image is transformed from the spatial domain to the frequency domain by using Discrete Cosine Transform (DCT) to obtain the remote sensing frequency domain image; among which, Discrete Cosine Transform is a well-known technique and will not be described in detail here.

[0069] It should be noted that low-frequency information in the frequency domain reflects the smoothness of an image, while high-frequency information reflects details, sharp edges, and rapidly changing parts of the image, generally reflecting its texture. Therefore, when remote sensing images are subjected to noise interference, their high-frequency information is significantly affected, while mid- and low-frequency information is generally less affected. To analyze the overall situation of a remote sensing image, the characteristics of each frequency in the frequency domain image after frequency domain transformation can be used for analysis.

[0070] It should be further explained that, since low-frequency information in remote sensing images is less affected by noise, while high-frequency information is more affected, to enhance remote sensing images, frequencies less affected by noise should be retained, while frequencies more affected by noise should be removed to reduce the degree of noise impact. Therefore, by normalizing the radial distance and using dynamic threshold parameters, the two-dimensional frequency domain is divided into three regions with different processing strategies: a low-frequency absolute protection zone (preserving the true image content); a mid-frequency transition zone (partially attenuating potentially noisy components); and a high-frequency suppression zone (focusing on filtering and combating noise). Then, by setting mask matrices formed by these three different regions, the retention of image frequency information is analyzed. When the data in the mask matrix is ​​1, the corresponding frequency information should be retained, while when the data in the mask matrix is ​​0, the corresponding frequency information should be removed.

[0071] It should be further explained that, in the frequency domain, the upper left corner of the image represents low frequency, while the further away from the upper left corner, the higher the frequency information. That is, the coordinates of the origin represent low frequency information, and the further away from the origin, the higher the frequency information. Therefore, in order to distinguish high, medium and low frequency information, the positional relationship between all position points and the origin is used for differentiation.

[0072] Specifically, the top-left corner of the remote sensing frequency domain image is designated as the origin, and a reference coordinate system is constructed with the horizontal axis pointing to the right and the vertical axis pointing downwards. Based on the distances between all location points and the origin, a weighting factor for the mask value corresponding to each location point in the remote sensing frequency domain image is obtained. This weighting factor is specifically expressed by the formula:

[0073]

[0074]

[0075] In the formula, Represents the horizontal length of a remote sensing frequency domain image. Represents the vertical length of the remote sensing frequency domain image. This represents the maximum distance between all points in the remote sensing frequency domain image and the origin. Represents the location points in the remote sensing frequency domain image The distance to the origin Represents the location points in the remote sensing frequency domain image Weighting factors corresponding to mask values; location points in remote sensing frequency domain images The horizontal axis is represented as The vertical axis is Location point.

[0076] Specifically, the smaller the weight factor of the mask value corresponding to each location point in the remote sensing frequency domain image, the more it indicates low-frequency information, and the more low-frequency information should be preserved. Conversely, the larger the weight factor of the mask value corresponding to each location point in the remote sensing frequency domain image, the more it indicates high-frequency information, and the more high-frequency information should be removed appropriately.

[0077] It should be noted that since there is less noise in low-frequency information, low-frequency information below the threshold is directly retained by segmentation; mid-frequency information is appropriately removed; and high-frequency information is filtered and removed in a focused manner. Therefore, different processing is carried out by dividing the low-frequency absolute protection zone, mid-frequency transition zone, and high-frequency suppression zone.

[0078] Dynamic thresholds are generated by setting learnable hyperparameters, including a low-frequency threshold and a Gaussian decay width threshold; the low-frequency threshold is specifically expressed by the formula: The Gaussian attenuation width threshold is specifically expressed by the formula: ;

[0079] In the formula, and This represents the learnable and trainable hyperparameters. Indicates the low-frequency threshold. This represents the threshold width of the Gaussian decay. This represents the activation function, used for data normalization.

[0080] Based on the low-frequency threshold and the Gaussian attenuation width threshold, the low-frequency absolute protection zone, the mid-frequency transition zone, and the high-frequency suppression zone are determined; specifically: The area consisting of all corresponding location points is designated as the low-frequency absolute protection zone; The region consisting of all corresponding location points is denoted as the intermediate frequency transition region; The region consisting of all the corresponding location points is denoted as the high-frequency suppression region.

[0081] In this process, the mask values ​​of all locations within the low-frequency absolute protection zone are assigned as 1, and the mask values ​​of all other locations outside the low-frequency absolute protection zone are assigned as 0, forming a matrix denoted as the low-frequency mask matrix.

[0082] The mask value at each location point in the intermediate frequency transition region is specifically expressed by the following formula:

[0083]

[0084] In the formula, Indicates the location point in the intermediate frequency transition region. The mask value, Represents the location points in the remote sensing frequency domain image The weighting factor corresponding to the mask value, Indicates the low-frequency threshold. This represents the threshold width of the Gaussian decay. This represents an exponential function with the natural constant as its base.

[0085] The mask values ​​of all other locations outside the intermediate frequency transition region are assigned to 0. Then, the mask values ​​of all locations in the intermediate frequency transition region are combined to form a matrix, which is called the intermediate frequency mask matrix.

[0086] The mask value at each location point in the high-frequency suppression region is specifically expressed by the formula:

[0087]

[0088] In the formula, Indicates the location point in the high-frequency suppression region The mask value, This indicates the location point in the corresponding high-frequency suppression region of the noise sensitivity diagram. The noise sensitivity value.

[0089] The mask values ​​of all other locations outside the high-frequency suppression region are assigned to 0. Then, the mask values ​​of all locations within the high-frequency suppression region are combined to form a matrix, which is denoted as the high-frequency mask matrix.

[0090] The process of obtaining the noise sensitivity map is specifically described as follows:

[0091]

[0092] In the formula, This represents the frequency domain amplitude in a remotely sensed frequency domain image. Represents the absolute value symbol. Represents the inverse discrete cosine transform. This indicates that remote sensing images are estimated using an estimator combined with a lightweight CNN network. This represents a noise sensitivity map.

[0093] Thus, the low-frequency mask matrix, the mid-frequency mask matrix, and the high-frequency mask matrix are obtained.

[0094] The final mask matrix is ​​obtained based on the low-frequency mask matrix, mid-frequency mask matrix, high-frequency mask matrix, and the matrix corresponding to the noise sensitivity map; the final mask matrix is ​​specifically expressed by the formula:

[0095]

[0096] In the formula, Represents the low-frequency mask matrix. Represents the intermediate frequency mask matrix. Represents the high-frequency mask matrix. ⊙ represents the matrix composed of the noise sensitivity values ​​of all points in the noise sensitivity map; ⊙ represents the Hadamard product (i.e., the product of corresponding elements in the matrix). This represents the final mask matrix.

[0097] The adjusted remote sensing frequency domain image is obtained by using the Hadamard product between the corresponding values ​​in the final mask matrix and the remote sensing frequency domain image. The adjusted remote sensing image is then obtained by performing an inverse discrete cosine transform (IDCT) on the adjusted remote sensing frequency domain image. The inverse discrete cosine transform is a well-known technique and will not be described in detail here. Each data point in the final mask matrix corresponds to a location point in the adjusted remote sensing frequency domain image.

[0098] It should be noted that, in order to prevent excessive filtering from causing the loss of details in the image, the final image is obtained by adjusting and merging the remote sensing images proportionally.

[0099] Specifically, a frequency-domain purified remote sensing image is obtained by adjusting the remote sensing image and the remote sensing image; this is expressed by the following formula:

[0100]

[0101] In the formula, This indicates adjusting the pixel values ​​in the remote sensing image. Represents the pixel values ​​in a remotely sensed image. This represents the pixel values ​​in the frequency domain purified remote sensing image. This represents a preset scaling factor, where, in this embodiment, the preset scaling factor... In this embodiment, a preset proportional coefficient is used. No specific restrictions are imposed; implementers can decide based on the specific circumstances.

[0102] The frequency domain purified remote sensing image is composed of pixel values ​​from all locations.

[0103] The process of acquiring frequency-domain purified remote sensing images constitutes an adaptive adversarial noise random mask module.

[0104] Thus, we obtained the adaptive adversarial noise random mask module and the frequency domain purified remote sensing image.

[0105] Step S003: Construct a baseline SwinIR network, adjust the self-attention mechanism model in the window self-attention layer of the baseline SwinIR network to obtain a new self-attention mechanism model in the window self-attention layer; based on the frequency domain cleansed remote sensing image, obtain the final reconstructed remote sensing image through the baseline SwinIR network and the new self-attention mechanism model in the window self-attention layer.

[0106] It should be noted that frequency-domain cleansing of remote sensing images after adaptive adversarial noise random masking can be used for subsequent super-resolution reconstruction. A common technique for this is to use the SwinIR network as the baseline network for remote sensing image super-resolution reconstruction. In conventional SwinIR networks, attention computation is performed within a window, using relative position encoding to represent the sequential features of the input data. This encoding method typically embeds relative positions into the input representation, allowing the model to distinguish markers at different positions in the sequence. However, this method usually injects positional information into the attention matrix in a fully connected manner, leading to a quadratic increase in computational complexity with sequence length, which is not conducive to improving the model's robustness. The SwinIR network, on the other hand, is an image restoration model based on the Transformer architecture. It borrows some core ideas from the Swin Transformer (Shifted Window Transformer) and applies them to image restoration tasks such as image denoising, super-resolution, and deblurring.

[0107] It should be further noted that since the SwinIR network and the convolutional neural network are completely different models, in order to further improve the robustness of the SwinIR network model, we can start from the specific architectural information of the SwinIR network.

[0108] Specifically, the frequency domain purified remote sensing image is uniformly divided to obtain several non-overlapping windows.

[0109] Each window in the frequency-domain cleansed remote sensing image is used as input, and super-resolution reconstruction is performed through a baseline SwinIR network. During the reconstruction process of the baseline SwinIR network, convolutional layers are used to extract shallow features. Then, based on the shallow features, deep features are extracted through a residual Swin Transformer module (RSTB, Residual Swin Transformer Block). Finally, based on the deep features, upsampling is performed through PixelShuffle to complete the super-resolution reconstruction and obtain the final reconstructed remote sensing image.

[0110] In this embodiment, when using convolutional layers for shallow feature extraction, the following method is used: The convolution operation is performed using convolution kernels. In this embodiment, the size of the convolution kernel is not specifically limited, and the implementer can determine it according to the specific situation. Each RSTB consists of a window-based self-attention layer (WSA), a layer normalization layer (LN), and a multilayer perceptron (MLP). Furthermore, an adaptive adversarial noise random mask module is placed at the head and tail of each RSTB module to ensure that adversarial noise in the remote sensing image input to the network is randomly removed.

[0111] Specifically, the self-attention mechanism model in the new window self-attention layer is represented as follows:

[0112]

[0113] In the formula, Represents the query vector. Represents the key vector. Represents a value vector. The dimension representing the key vector (or query vector). Represents the sparse position importance matrix. The notation for calculating the sparse position importance matrix in a computer. This represents the adjusted vector. This represents the transpose of the key vector. This represents the activation function, used for data normalization. The query vector, key vector, and value vector can all be obtained using existing methods, which will not be elaborated upon here.

[0114] Here, an initial matrix is ​​formed by randomly selecting a number of data points within a random range, and then the initial matrix is ​​optimized through iterative learning to obtain the final matrix, which is denoted as the sparse position importance matrix. The random range for selecting random data in the initial matrix can be a uniformly distributed interval or an interval corresponding to a dataset that follows a standard normal distribution. In this embodiment, the random range is not specifically limited, and the implementer can determine it according to the specific circumstances.

[0115] It should be noted that by decoupling positional attention into a learnable sparse structure, an efficient and flexible solution is provided for the adversarial robustness of super-resolution reconstruction.

[0116] At this point, the final reconstructed remote sensing image is obtained.

[0117] It should be noted that, in order to improve the neural network model's resistance to noise interference, adversarial noise of varying intensities is added to the remote sensing images to verify the model's robustness. To verify the effectiveness of the present invention, the following two experiments are conducted to analyze its effectiveness: one on the robustness of image super-resolution reconstruction against basic attacks, and the other on the robustness of image super-resolution reconstruction against black-box attacks.

[0118] Experiment 1: Robustness of image super-resolution reconstruction against basic attacks.

[0119] Specifically, this embodiment evaluates the robustness of the proposed method on the UCM (University of California, Merced) dataset by employing a basic attack. The basic attack used in this embodiment is a white-box attack based on I-FGSM (Iterative Fast Gradient Sign Method). During the basic attack adversarial training process, this embodiment iterates twice and trains on the UCM images with an intensity of 8 / 255. During the evaluation, the number of iterations is increased to 10, and five intensities are used for testing, namely 0 / 255, 2 / 255, 4 / 255, 6 / 255, and 8 / 255.

[0120] Table 1 shows the performance of different super-resolution models under basic attacks on the UCM dataset. In the absence of attacks (0 / 255), all models maintain high performance, with our proposed solution (Ours) achieving the best performance with a PSNR of 36.14 dB and an SSIM of 0.9694. As the attack intensity increases, the performance of all models decreases to varying degrees, but our proposed solution exhibits stronger robustness: it maintains a PSNR of 35.63 dB and an SSIM of 0.9674 even under the strongest attack of 8 / 255, significantly outperforming other comparative models. Particularly noteworthy is the anomalous increase in SwinIR's PSNR (36.22 dB) under the 2 / 255 attack, possibly due to the adversarial example activating its attention mechanism. However, its performance rapidly degrades under higher-intensity attacks, while our proposed solution achieves stable performance maintenance through random frequency masking and positional attention enhancement strategies. Compared to traditional CNN (Convolutional Neural Network) models (EDSR, RCAN) and lightweight IMDN, Transformer-based LKFormer and SwinIR are more sensitive to attacks, verifying the vulnerability of the ViT (VisionTransformer) architecture in adversarial environments. The defense module proposed in this application effectively mitigates this deficiency. Specifically, EDSR (Enhanced Deep Super-Resolution network), RCAN (Residual Channel Attention Network), IMDN (Information Multi-distillation Network), LKFormer (Local and Global Transformer), SwinIR (Swin Transformer for Image Restoration), PSNR (Peak Signal-to-Noise Ratio), and SSIM (Structural Similarity Index Measure) are used; where Ours represents the network corresponding to the proposed solution.

[0121] Table 1. Comparison and evaluation of 2x SR results for the UCM dataset after the basic attack.

[0122]

[0123] according to Figure 3 The experimental results shown demonstrate that the method proposed in this invention exhibits significant advantages in robustness and reconstruction quality for remote sensing image SR tasks under basic adversarial attacks. Figure 3 The image shows multiple sub-images, each corresponding to the super-resolution (SR) reconstruction results after adversarial attacks processed by different algorithms or models. Compared to other methods, the image processed by the present invention (such as the magnified area within the red box) exhibits fewer artifacts and noise, with higher clarity in aircraft outlines and background details. This indicates that the present invention effectively suppresses the amplification effect of adversarial perturbations on high-frequency components by introducing a Stochastic Frequency Masking (SFM) module and a Position Attention Enhancement (PAE) module. Therefore, compared to the sub-images of other comparative algorithms, the remote sensing image reconstructed by the present invention has clearer texture and less adversarial noise interference. Overall, the model of the present invention provides more robust reconstruction results for remote sensing images subjected to basic adversarial attacks.

[0124] Experiment 2: Robustness of image super-resolution reconstruction against black-box attacks.

[0125] The SR results of the two datasets after being subjected to a general attack are shown below (×2).

[0126] Table 2. Comparison and evaluation of SR results (×2x) for the two datasets after being subjected to a general attack.

[0127]

[0128] Table 2 shows the performance of different super-resolution models on black-box adversarial attack datasets (AID-black and UCM-black). On the AID dataset, the proposed solution significantly outperforms other models with a PSNR of 27.35 dB and an SSIM of 0.8255. It also maintains a leading advantage on the UCM dataset with a PSNR of 28.48 dB and an SSIM of 0.8917. In contrast, the performance of Transformer-based models LKFormer and SwinIR fluctuates significantly, especially with a significant decrease in SSIM on the AID dataset (LKFormer only 0.7293). Traditional CNN models (EDSR, RCAN), while having a large number of parameters, lack robustness. The proposed solution, through the synergistic effect of random frequency masking and positional attention enhancement mechanisms, effectively improves the model's defense capability against unknown black-box attacks, providing a reliable solution for dealing with complex adversarial threats in practical applications. Among them, AID-black (AerialImage Dataset-black) and UCM-black (University of California, Merced-black).

[0129] like Figure 4 As shown, the present invention significantly improves the robustness and reconstruction quality of remote sensing image SR tasks under general adversarial attacks by combining frequency domain sanitization and position attention enhancement strategies. It can be seen that the aircraft frame in the IMDN result image is unclear and significantly affected by adversarial attacks. EDSR and RCAN, limited by their local receptive field, struggle to effectively recover high-frequency details, resulting in blurred edges. While Transformer-based methods (such as LKFormer and SwinIR) capture global dependencies, they still exhibit some noise and structural distortion under adversarial attacks, especially SwinIR, whose performance degrades significantly under high-intensity attacks. Compared to other methods, the results of the present invention show fewer artifacts and noise, sharper aircraft outlines, and more complete preservation of background details (such as ground textures and building outlines). This is attributed to the Stochastic Frequency Masking (SFM) module effectively suppressing the amplification of high-frequency components by adversarial perturbations, and the Position Attention Enhancement (PAE) module optimizing the attention mechanism and enhancing the model's ability to model complex textures and spatial relationships.

[0130] To better illustrate the improvement effect of the module proposed in this invention on the model, ablation experiments were conducted on the UCM dataset under a basic adversarial attack of strength 8 / 255. The results showed that SFM and PAS corresponded to different improvement effects of SwinIR at 2x SR results, respectively. This is combined with Table 3 and... Figure 5It can be seen that both modules of the present invention have good robustness to SR tasks.

[0131] Table 3. Comparison and evaluation of ×2 SR results for the UCM dataset after the basic attack.

[0132]

[0133] In this embodiment, the Adam (Adaptive Moment Estimation) optimizer is used to update the model during the network model optimization phase. The number of training iterations for optimization updates is set to 50,000. The trained model is evaluated using Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM). The various parameters and values ​​mentioned above are not specifically limited and can be determined by the implementer according to specific circumstances. SFM (Stochastic Frequency Masking) and PAE (Position Attention Enhancement) are also mentioned.

[0134] The technical methods used in the training and evaluation processes are all well-known technologies, and will not be described in detail here.

[0135] like Figure 2 As shown, a second aspect of the present invention is to provide a position-attention-enhanced image super-resolution adversarial defense system, comprising the following modules:

[0136] Image acquisition module 101: used to acquire remote sensing images after introducing adversarial noise;

[0137] Image analysis and adjustment module 102: This module performs spatial-domain to frequency-domain conversion on remote sensing images to obtain a remote sensing frequency-domain image; acquires the weighting factor of the mask value corresponding to each location point in the remote sensing frequency-domain image; acquires the low-frequency threshold and Gaussian attenuation width threshold; determines the low-frequency absolute protection zone, mid-frequency transition zone, and high-frequency suppression zone in the remote sensing frequency-domain image based on the low-frequency threshold and Gaussian attenuation width threshold; obtains the low-frequency mask matrix based on the mask values ​​of all locations in the low-frequency absolute protection zone; and determines the mask matrix based on the weighting factor, low-frequency threshold, and Gaussian attenuation width threshold of the mask value corresponding to each location point in the mid-frequency transition zone. The process involves obtaining an intermediate frequency mask matrix; determining a noise sensitivity map based on the frequency domain amplitude in the remote sensing frequency domain image; obtaining a high-frequency mask matrix based on the mask values ​​of all locations in the high-frequency suppression region and the noise sensitivity map; determining the final mask matrix using the low-frequency mask matrix, intermediate frequency mask matrix, high-frequency mask matrix, and noise sensitivity map; obtaining an adjusted remote sensing image based on the final mask matrix and the remote sensing frequency domain image; fusing the adjusted remote sensing image and the original remote sensing image to obtain a frequency-domain purified remote sensing image; and constructing an adaptive adversarial noise random mask module from the process of obtaining the frequency-domain purified remote sensing image.

[0138] Image reconstruction module 103: It is used to construct a baseline SwinIR network through an adaptive adversarial noise random mask module, adjust the self-attention mechanism model in the window self-attention layer of the baseline SwinIR network to obtain a new self-attention mechanism model in the window self-attention layer; uniformly divide the frequency domain cleaned remote sensing image to obtain several non-overlapping windows, take each window in the frequency domain cleaned remote sensing image as input, and obtain the final reconstructed remote sensing image through the baseline SwinIR network and the new self-attention mechanism model in the window self-attention layer.

[0139] A third aspect of the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements a position attention-enhanced image super-resolution adversarial defense method.

[0140] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) containing computer-usable program code.

[0141] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, systems, and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0142] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0143] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the present invention.

Claims

1. A position-attention-enhanced image super-resolution adversarial defense method, characterized in that, include: Acquire remote sensing images after introducing adversarial noise; The remote sensing image is converted from the spatial domain to the frequency domain to obtain the remote sensing frequency domain image; Obtain the weighting factor of the mask value corresponding to each location point in the remote sensing frequency domain image; obtain the low-frequency threshold and Gaussian attenuation width threshold; determine the low-frequency absolute protection zone, mid-frequency transition zone, and high-frequency suppression zone in the remote sensing frequency domain image based on the low-frequency threshold and Gaussian attenuation width threshold; obtain the low-frequency mask matrix based on the mask values ​​of all locations in the low-frequency absolute protection zone; obtain the mid-frequency mask matrix based on the weighting factor of the mask value corresponding to each location point in the mid-frequency transition zone, the low-frequency threshold, and the Gaussian attenuation width threshold. The noise sensitivity map is determined based on the frequency domain amplitude in the remote sensing frequency domain image; the high-frequency mask matrix is ​​obtained based on the mask values ​​and noise sensitivity map of all locations in the high-frequency suppression region; The final mask matrix is ​​determined using the low-frequency mask matrix, mid-frequency mask matrix, high-frequency mask matrix, and noise sensitivity map. Based on the final mask matrix and the remote sensing frequency domain image, the adjusted remote sensing image is obtained; The adjusted remote sensing image and the remote sensing image are fused to obtain a frequency domain cleaned remote sensing image; the process of obtaining the frequency domain cleaned remote sensing image constitutes an adaptive adversarial noise random mask module; A baseline SwinIR network is constructed using an adaptive adversarial noise random mask module. The self-attention mechanism model in the window self-attention layer of the baseline SwinIR network is adjusted to obtain a new self-attention mechanism model in the window self-attention layer. The frequency domain cleansing remote sensing image is uniformly divided into several non-overlapping windows. Each window in the frequency domain cleansing remote sensing image is used as input. Through the baseline SwinIR network and the new self-attention mechanism model in the window self-attention layer, the final reconstructed remote sensing image is obtained. The new self-attention mechanism model in the window self-attention layer is specifically expressed as follows: In the formula, Represents the query vector. Represents the key vector. Represents a value vector. The dimension of the key vector. Represents the sparse position importance matrix. The notation for calculating the sparse position importance matrix in a computer. This represents the adjusted vector. This represents the transpose of the key vector. This represents the activation function.

2. The image super-resolution adversarial defense method with position attention enhancement according to claim 1, characterized in that, The process of converting the remote sensing image from the spatial domain to the frequency domain to obtain a remote sensing frequency domain image includes: The spatial domain of the remote sensing image is converted to the frequency domain by discrete cosine transform, thus obtaining the frequency domain image of the remote sensing image.

3. The image super-resolution adversarial defense method with position attention enhancement according to claim 1, characterized in that, The weighting factor for obtaining the mask value corresponding to each location point in the remote sensing frequency domain image includes: The top left corner of the remote sensing frequency domain image is recorded as the origin, and a reference coordinate system is constructed with the horizontal axis pointing to the right and the vertical axis pointing downwards. In the formula, Represents the horizontal length of a remote sensing frequency domain image. Represents the vertical length of the remote sensing frequency domain image. This represents the maximum distance between all points in the remote sensing frequency domain image and the origin. Represents the location points in the remote sensing frequency domain image The distance to the origin Represents the location points in the remote sensing frequency domain image Weighting factors corresponding to mask values; location points in remote sensing frequency domain images The horizontal axis is represented as The vertical axis is Location point.

4. The image super-resolution adversarial defense method with position attention enhancement according to claim 3, characterized in that, The process involves acquiring a low-frequency threshold and a Gaussian attenuation width threshold; and determining, based on these thresholds, a low-frequency absolute protection zone, a mid-frequency transition zone, and a high-frequency suppression zone in the remote sensing frequency domain image, including: In the formula, and This represents the learnable and trainable hyperparameters. Indicates the low-frequency threshold. This represents the threshold width of the Gaussian decay. This represents the activation function, used for data normalization. Will The area consisting of all corresponding location points is designated as the low-frequency absolute protection zone; The region consisting of all corresponding location points is denoted as the intermediate frequency transition region; The region consisting of all the corresponding location points is denoted as the high-frequency suppression region.

5. The image super-resolution adversarial defense method with position attention enhancement according to claim 3, characterized in that, The low-frequency mask matrix is ​​obtained based on the mask values ​​of all locations within the low-frequency absolute protection zone. Based on the weighting factor, low-frequency threshold, and Gaussian attenuation width threshold of the mask value corresponding to each location point in the mid-frequency transition region, the mid-frequency mask matrix is ​​obtained, including: The mask values ​​of all locations within the low-frequency absolute protection zone are assigned as 1, and the mask values ​​of all other locations outside the low-frequency absolute protection zone are assigned as 0, forming a matrix, denoted as the low-frequency mask matrix. In the formula, Indicates the location point in the intermediate frequency transition region. The mask value, Represents the location points in the remote sensing frequency domain image The weighting factor corresponding to the mask value, Indicates the low-frequency threshold. This represents the threshold width of the Gaussian decay. Represents an exponential function with the natural constant as its base; The mask values ​​of all other locations outside the intermediate frequency transition region are assigned to 0. Then, the mask values ​​of all locations in the intermediate frequency transition region are combined to form a matrix, which is called the intermediate frequency mask matrix.

6. The image super-resolution adversarial defense method with position attention enhancement according to claim 1, characterized in that, The noise sensitivity map is determined based on the frequency domain amplitude in the remote sensing frequency domain image; Based on the mask values ​​and noise sensitivity maps of all locations in the high-frequency suppression region, a high-frequency mask matrix is ​​obtained, including: In the formula, This represents the frequency domain amplitude in a remotely sensed frequency domain image. Represents the absolute value symbol. Represents the inverse discrete cosine transform. This indicates that remote sensing images are estimated using an estimator combined with a lightweight CNN network. Represents a noise sensitivity map; In the formula, Indicates the location point in the high-frequency suppression region The mask value, This indicates the location point in the corresponding high-frequency suppression region of the noise sensitivity diagram. The noise sensitivity value; The mask values ​​of all other locations outside the high-frequency suppression region are assigned to 0. Then, the mask values ​​of all locations within the high-frequency suppression region are combined to form a matrix, which is denoted as the high-frequency mask matrix.

7. The image super-resolution adversarial defense method with position attention enhancement according to claim 1, characterized in that, The final mask matrix is ​​determined by using the low-frequency mask matrix, the mid-frequency mask matrix, the high-frequency mask matrix, and the noise sensitivity map. Based on the final mask matrix and the remote sensing frequency domain image, the adjusted remote sensing image is obtained; The adjusted remote sensing image and the fusion of the remote sensing images are used to obtain a frequency domain cleaned remote sensing image, including: In the formula, Represents the low-frequency mask matrix. Represents the intermediate frequency mask matrix. Represents the high-frequency mask matrix. The matrix represents the noise sensitivity values ​​of all points in the noise sensitivity map; ⊙ represents the Hadamard product. This represents the final mask matrix; The adjusted remote sensing frequency domain image is obtained by using the Hadamard product between the final mask matrix and the corresponding values ​​of the remote sensing frequency domain image; the adjusted remote sensing image is then obtained by performing an inverse discrete cosine transform on the adjusted remote sensing frequency domain image. In the formula, This indicates adjusting the pixel values ​​in the remote sensing image. Represents the pixel values ​​in a remotely sensed image. This represents the pixel values ​​in the frequency domain purified remote sensing image. This indicates the preset scaling factor.

8. The image super-resolution adversarial defense method with position attention enhancement according to claim 1, characterized in that, The process involves constructing a baseline SwinIR network using an adaptive adversarial noise random mask module, adjusting the self-attention mechanism model in the window self-attention layer of the baseline SwinIR network to obtain a new self-attention mechanism model in the window self-attention layer, uniformly dividing the frequency domain cleaned remote sensing image into several non-overlapping windows, using each window in the frequency domain cleaned remote sensing image as input, and passing it through the baseline SwinIR network and the new self-attention mechanism model in the window self-attention layer to obtain the final reconstructed remote sensing image, including: Each window in the frequency-domain cleansed remote sensing image is used as input, and super-resolution reconstruction is performed through the baseline SwinIR network. During the reconstruction process of the baseline SwinIR network, convolutional layers are used to extract shallow features. Then, based on the shallow features, the residual Swin Transformer module RSTB is used to extract deep features. Finally, based on the deep features, PixelShuffle is used to perform upsampling to complete the super-resolution reconstruction and obtain the final reconstructed remote sensing image. Each RSTB consists of a window self-attention layer, layer normalization, and a multilayer perceptron; and an adaptive adversarial noise random mask module is placed at the head and tail of each RSTB.

9. A position-attention-enhanced image super-resolution adversarial defense system, characterized in that, include: Image acquisition module: used to acquire remote sensing images after introducing adversarial noise; Image Analysis and Adjustment Module: This module converts remote sensing images from the spatial domain to the frequency domain to obtain a remote sensing frequency domain image; it acquires the weighting factor of the mask value corresponding to each location point in the remote sensing frequency domain image; it acquires the low-frequency threshold and the Gaussian attenuation width threshold; it determines the low-frequency absolute protection zone, the mid-frequency transition zone, and the high-frequency suppression zone in the remote sensing frequency domain image based on the low-frequency threshold and the Gaussian attenuation width threshold; it obtains the low-frequency mask matrix based on the mask values ​​of all locations in the low-frequency absolute protection zone; and it obtains the mid-frequency mask matrix based on the weighting factor of the mask value corresponding to each location point in the mid-frequency transition zone, the low-frequency threshold, and the Gaussian attenuation width threshold. The noise sensitivity map is determined based on the frequency domain amplitude in the remote sensing frequency domain image; the high-frequency mask matrix is ​​obtained based on the mask values ​​and noise sensitivity map of all locations in the high-frequency suppression region; The final mask matrix is ​​determined using the low-frequency mask matrix, mid-frequency mask matrix, high-frequency mask matrix, and noise sensitivity map. Based on the final mask matrix and the remote sensing frequency domain image, the adjusted remote sensing image is obtained; The adjusted remote sensing image and the remote sensing image are fused to obtain a frequency domain cleaned remote sensing image; the process of obtaining the frequency domain cleaned remote sensing image constitutes an adaptive adversarial noise random mask module; Image reconstruction module: This module constructs a baseline SwinIR network using an adaptive adversarial noise random mask module. It adjusts the self-attention mechanism model in the window self-attention layer of the baseline SwinIR network to obtain a new self-attention mechanism model in the window self-attention layer. The frequency-domain cleaned remote sensing image is uniformly divided into several non-overlapping windows. Each window in the frequency-domain cleaned remote sensing image is used as input. Through the baseline SwinIR network and the new self-attention mechanism model in the window self-attention layer, the final reconstructed remote sensing image is obtained. The new self-attention mechanism model in the window self-attention layer is specifically represented as follows: In the formula, Represents the query vector. Represents the key vector. Represents a value vector. The dimension of the key vector. Represents the sparse position importance matrix. The notation for calculating the sparse position importance matrix in a computer. This represents the adjusted vector. This represents the transpose of the key vector. This represents the activation function.

10. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the position attention-enhanced image super-resolution adversarial defense method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Super-resolution remote sensing image acquisition method and device

    CN119205511A

  • Hyperspectral image and laser radar image fusion method and system for field of remote sensing

    WO2024174314A1