A rotation-aware enhanced variable window attention super-resolution method and system

By introducing a learnable quadrilateral attention window mechanism and a rotation perception term, the window shape is dynamically adjusted, which solves the shortcomings of existing super-resolution methods in the processing of rotated and variable-scale images, and achieves efficient and high-precision remote sensing image reconstruction.

CN120563322BActive Publication Date: 2025-11-11XIAN AERONAUTICAL UNIV

Patent Information

Application Number
CN202510667098.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-11-11
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

Existing super-resolution methods suffer from blurred or geometrically distorted reconstruction results due to fixed window sizes when processing rotated and scaled images. They cannot adaptively process image structures of different scales and have high computational costs. Traditional attention mechanisms cannot capture feature changes of rotated targets.

Method used

A learnable quadrilateral attention window mechanism is adopted, which dynamically adjusts the window shape through the projective transformation matrix. The quadrilateral attention mechanism is introduced in combination with the rotation perception term, and the SwinIR network is embedded for feature extraction to establish a high-precision super-resolution model.

Benefits of technology

It significantly improves the reconstruction quality and efficiency of remote sensing images, can dynamically adapt to rotated and scaled images, reduces computational overhead, and improves the feature matching accuracy of rotated targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563322B_ABST
    Figure CN120563322B_ABST
Patent Text Reader

Abstract

This invention provides a rotation-aware enhanced variable window attention super-resolution method, belonging to the field of image processing technology. It includes acquiring and preprocessing the original remote sensing image of the target object, segmenting it into several basic windows, extracting feature vectors, using quadrilaterals to predict the basic parameters of the projective transformation matrix based on the extracted feature vectors to obtain the reconstructed projective transformation matrix, extracting the rotation angle difference between any two quadrilaterals as a rotation-aware term, establishing a variable window attention mechanism, embedding the variable window attention mechanism into a SwinIR network to build a high-precision super-resolution model architecture, training the high-precision super-resolution model, inputting real-time remote sensing images into the high-precision super-resolution model, and outputting a high-resolution image containing the target sampling area. By combining the variable window attention mechanism with the rotation-aware mechanism, the super-resolution reconstruction accuracy and geometric consistency of rotated objects are significantly improved while maintaining computational efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, specifically relating to a rotation-aware enhanced variable window attention super-resolution method and system. Background Technology

[0002] Single-Image Super-Resolution (SISR), a fundamental task in computer vision, aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs. In recent years, with the rapid development of deep learning technologies, especially the widespread application of Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), the performance of super-resolution techniques has been significantly improved. Early super-resolution methods (SR), such as Super-Resolution Convolutional Neural Network (SRCNN), Enhanced Deep Super-Resolution Network (EDSR), and Residual Channel Attention Network (RCAN), were mainly based on hierarchical CNN structures, learning texture detail features through local receptive fields. However, this type of method has two main limitations: on the one hand, the fixed-size convolutional kernel limits the network's ability to model long-range dependencies in images, and can only capture local features, resulting in inconsistent texture and structure of the image; on the other hand, since the convolution operation only has translational isomorphism and not rotational isomorphism, it has a weak ability to extract features of rotational patterns, which is particularly evident when processing images containing rotating objects.

[0003] With the successful application of Visual Transformers in computer vision, ViT-based SR methods, such as the Swin Transformer-based Image Restoration (SwinIR) and EDT, have significantly improved the model's ability to capture global contextual information by introducing self-attention mechanisms. These methods typically employ window-based attention mechanisms, dividing the image into non-overlapping, fixed-size rectangular windows to reduce computational complexity, and using window offset strategies to achieve cross-window information interaction. However, existing methods still have significant limitations when handling rotated and variable-scale structures.

[0004] First, a fixed-shape rectangular window is difficult to align with edges or texture features in any direction within an image, leading to problems such as blurriness or geometric distortion in the reconstruction results. Second, a fixed window size cannot adaptively handle image structures of different scales; for example, it performs poorly when processing both fine textures and large objects simultaneously.

[0005] While some improved methods, such as DAT and Deformable DETR, attempt to incorporate deformable convolution into the attention mechanism, these methods still have significant drawbacks. They primarily focus on modeling translational offsets, lacking explicit handling of rotational and shearing transformations. Furthermore, these methods typically require dense sampling point computation, leading to high memory consumption and computational overhead, making them unsuitable for high-resolution super-resolution tasks. Another key issue is that current window-based methods, such as SwinIR, require manual design of window size and offset strategies. This not only limits the model's generalization ability across different datasets but also introduces a significant amount of unnecessary computational overhead, especially when dealing with images containing irregular textures.

[0006] Furthermore, in traditional attention mechanisms, the size and position of the fixed window are predefined in the model and cannot be dynamically adjusted according to the characteristics of the input data; ordinary variable windows are usually adjusted based on local features or heuristic rules, lacking global context dependencies, which leads to inaccurate window adjustments, causing important information to be ignored or redundant information to be included, failing to capture global dependencies, and being unable to adapt to multi-size feature relationships; the attention weights in existing traditional attention mechanisms are usually calculated by calculating the dot product between the query (Q) and the key (K) This is obtained by means of the similarity between features. When an object rotates, its feature representation will also change, but traditional attention mechanisms cannot capture this change, resulting in inaccurate feature matching of rotating targets.

[0007] Based on the above background, this invention proposes a rotation-aware enhanced variable window attention super-resolution network, which innovatively introduces a learnable quadrilateral attention window mechanism. Summary of the Invention

[0008] The purpose of this invention is to overcome the above-mentioned shortcomings and provide a variable window attention super-resolution method and system with rotation-aware enhancement.

[0009] To achieve the above objectives, the present invention adopts the following technical solution:

[0010] This invention provides a rotation-aware enhanced variable window attention super-resolution method, comprising the following steps:

[0011] Acquire raw remote sensing images of the target object and perform enhancement preprocessing on the acquired raw remote sensing images;

[0012] The preprocessed original remote sensing image is divided into several basic windows, and feature vectors are extracted from each basic window.

[0013] Define the projective transformation matrix, use quadrilaterals to predict the surrogate parameters of scaling, rotation, shearing, translation and projection in the projective transformation matrix based on the extracted feature vectors, and decompose the projective transformation matrix into the corresponding transformation matrix based on the predicted surrogate parameters.

[0014] Reconstruct the projective transformation matrix based on the decomposed transformation matrix, and use the reconstructed projective transformation matrix to dynamically convert each base window into an arbitrary quadrilateral;

[0015] Based on the dynamic transformation of the base window, the rotation angle difference between pixels at the same position in any two base windows after transformation is extracted from the rotation transformation matrix. Based on the quadrilateral attention mechanism, the extracted rotation angle difference is introduced as a rotation perception term into the quadrilateral attention mechanism to establish a variable window attention mechanism.

[0016] A variable window attention mechanism is embedded into the SwinIR network and applied to the feature extraction layer to establish the architecture of a high-precision super-resolution model.

[0017] Collect a dataset of remote sensing images of the target object, divide the dataset into a training set and a test set, use the training set to input the architecture of the high-precision super-resolution model to obtain the trained high-precision super-resolution model, and use the test set to test and evaluate the trained high-precision super-resolution model.

[0018] The system acquires remote sensing images of the target object in real time, inputs these real-time remote sensing images into a high-precision super-resolution model, and outputs a high-resolution image containing the target sampling area.

[0019] Optionally, the enhancement preprocessing of the acquired raw remote sensing image includes:

[0020] Based on the acquired raw remote sensing images, grayscale processing is performed on the raw remote sensing images to obtain grayscale images;

[0021] The grayscale image is decomposed into wavelet coefficients of different frequencies based on wavelet transform, and the high-frequency coefficients are thresholded. The image is then reconstructed using the thresholded high-frequency coefficients to obtain the original denoised remote sensing image.

[0022] Optionally, the step of segmenting the preprocessed original remote sensing image into several basic windows and extracting feature vectors from each basic window includes:

[0023] Based on a predefined window size, the preprocessed original remote sensing image is segmented into multiple basic windows of the same size;

[0024] For each base window, features are extracted, and the extracted features are transformed into query tokens, key tokens, and value tokens through linear projection. The function expression for linear projection is:

[0025]

[0026] In the formula, Q represents the query token, K represents the key token, and V represents the value token. Represents a linear transformation function. Represents the features within the base window;

[0027] The query token, key token, and value token corresponding to each basic window are concatenated to obtain the feature vector corresponding to the basic window.

[0028] Optionally, the definition of the projective transformation matrix, based on the extracted feature vectors, uses quadrilaterals to predict surrogate parameters for scaling, rotation, shearing, translation, and projection in the projective transformation matrix, including:

[0029] Define the projective transformation matrix, and its functional expression is:

[0030]

[0031] In the formula, Define scaling, rotation, and shearing transformations. Define translation transformation. The projection vector defines how the observer perceives changes in the target as the depth dimension changes.

[0032] Existing remote sensing images with known projective transformation matrices were collected, and a neural network model consisting of average pooling layers, LeakyReLU activation layers, and convolutional layers was established. Features were extracted from the existing remote sensing images and input into the neural network model for training to obtain a quadrilateral prediction model.

[0033] The feature vectors extracted from the base window are input into the quadrilateral prediction model, which outputs surrogate parameters for scaling, rotation, shearing, translation, and projection. The functional expressions for the surrogate parameters are as follows:

[0034]

[0035] In the formula, For the nth type of proxy parameter, It is a convolutional layer. For activation layer, For average pooling layer, The eigenvectors are the linearly projected features.

[0036] Optionally, the step of reconstructing the projective transformation matrix based on the decomposed transformation matrix and using the reconstructed projective transformation matrix to dynamically convert each base window into an arbitrary quadrilateral includes:

[0037] The predicted surrogate parameters for scaling, rotation, shearing, translation, and projection are applied to the defined projective transformation matrix, respectively, to obtain the scaling transformation matrix, rotation transformation matrix, shearing transformation matrix, translation transformation matrix, and projection transformation matrix in sequence.

[0038] The functional expression for the scaling transformation matrix is:

[0039]

[0040] In the formula, For scaling transformation matrix, This is the scaling factor of the base window along the x-axis. The scaling factor of the base window in the y-axis direction;

[0041] The functional expression for the rotation transformation matrix is:

[0042]

[0043] In the formula, Let be the rotation transformation matrix. The rotation angle of the base window;

[0044] The functional expression for the shear transformation matrix is:

[0045]

[0046] In the formula, Let be the rotation transformation matrix. The shearing factor of the base window in the x-axis direction. The shearing factor of the base window in the y-axis direction;

[0047] The functional expression for the translation transformation matrix is:

[0048]

[0049] In the formula, The translation transformation matrix is... This is the amount of translation of the base window along the x-axis. This is the amount of translation of the base window in the y-axis direction. The translation factor;

[0050] The functional expression for the projection transformation matrix is:

[0051]

[0052] In the formula, Let be the projection transformation matrix. This is the projection of the base window along the x-axis. It is the projection of the base window along the y-axis.

[0053] Based on the decomposed scaling, rotation, shearing, translation, and projection transformation matrices, the reconstructed projective transformation matrix is ​​established. The function expression for the reconstructed projective transformation matrix is ​​as follows:

[0054]

[0055] In the formula, This is the reconstructed projective transformation matrix.

[0056] Optionally, the step of dynamically converting each base window into an arbitrary quadrilateral using the reconstructed projective transformation matrix includes:

[0057] Identify the four vertices on the base window that are to be converted into an arbitrary quadrilateral, and obtain the original coordinates of the four vertices;

[0058] The original coordinates of the four vertices are converted into three-dimensional homogeneous coordinates, and the reconstructed projective transformation matrix is ​​used to calculate the three-dimensional homogeneous coordinates of the base window vertices before the transformation, so as to obtain the three-dimensional homogeneous coordinates of the base window vertices after the transformation.

[0059] Based on the three-dimensional homogeneous coordinates of the transformed vertices of the base window, determine the positions of the four vertices of the base window after the transformation, and after the positions of the transformed vertices are determined, obtain the original coordinates of any pixel point within the base window.

[0060] The original coordinates of any pixel are converted into 3D homogeneous coordinates. Then, the reconstructed projective transformation matrix is ​​used to calculate the 3D homogeneous coordinates of any pixel within the base window before the transformation, resulting in the transformed 3D homogeneous coordinates of any pixel within the base window. The function expression for the transformation of the pixel coordinates is as follows:

[0061]

[0062]

[0063] In the formula, The transformed coordinates of any pixel within the base window. The transformed 3D homogeneous coordinates of the k-th pixel within the base window. The three-dimensional homogeneous coordinates of any pixel within the base window before transformation;

[0064] By combining the four vertices of the base window and the transformed three-dimensional homogeneous coordinates of all pixels within the base window, an arbitrary quadrilateral after the transformation of the base window can be obtained.

[0065] Optionally, the step of introducing the extracted rotation angle difference as a rotation perception term into the quadrilateral attention mechanism to establish a variable window attention mechanism includes:

[0066] Extract the rotation angle from the rotation transformation matrix corresponding to each basic window;

[0067] The rotation angle difference is calculated based on the rotation angles of any two base windows. The functional expression for the rotation angle difference is:

[0068]

[0069] In the formula, The difference in rotation angle, Let be the rotation angle after the transformation of the i-th base window. Let be the rotation angle of the j-th base window after transformation;

[0070] The calculated rotation angle difference is used as the rotation perception term, and combined with other features, a quadrilateral attention mechanism is introduced to obtain the variable window attention mechanism. The functional expression of the variable window attention mechanism is:

[0071]

[0072] In the formula, For variable window attention mechanism, Let d be the normalization function, and d be the vector dimension. The learnable rotational sensitivity coefficient, To query the dot product of the token and the key token.

[0073] The present invention also provides a rotation-aware enhanced variable window attention super-resolution system, comprising:

[0074] The raw remote sensing image processing module is used to acquire raw remote sensing images of the target object and perform enhancement preprocessing on the acquired raw remote sensing images;

[0075] The feature vector extraction module is used to segment the preprocessed original remote sensing image into several basic windows and extract feature vectors from each basic window.

[0076] The projective transformation matrix decomposition module is used to define the projective transformation matrix, predict the surrogate parameters of scaling, rotation, shearing, translation and projection in the projective transformation matrix using quadrilaterals based on the extracted feature vectors, and decompose the projective transformation matrix into the corresponding transformation matrix based on the predicted surrogate parameters.

[0077] The basic window transformation module is used to reconstruct the projective transformation matrix based on the decomposed transformation matrix, and then use the reconstructed projective transformation matrix to dynamically transform each basic window into an arbitrary quadrilateral.

[0078] The attention mechanism establishment module is used to extract the rotation angle difference between pixels at the same position in any two base windows after transformation from the rotation transformation matrix based on the dynamic transformation of the base window. Based on the quadrilateral attention mechanism, the extracted rotation angle difference is introduced into the quadrilateral attention mechanism as a rotation perception term to establish a variable window attention mechanism.

[0079] The network architecture building module is used to embed the variable window attention mechanism into the SwinIR network and apply it to the feature extraction layer to build the architecture of a high-precision super-resolution model.

[0080] The model building module is used to collect and acquire remote sensing image datasets of the target object. The remote sensing image dataset is divided into training set and test set. The architecture of the high-precision super-resolution model is input using the training set to obtain the trained high-precision super-resolution model. The trained high-precision super-resolution model is tested and evaluated using the test set.

[0081] The image resolution processing module is used to acquire remote sensing images of the target object in real time, input the acquired real-time remote sensing images into the high-precision super-resolution model, and output a high-resolution image containing the target sampling area.

[0082] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of a rotation-aware enhanced variable window attention super-resolution method.

[0083] The present invention also provides a storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of a rotation-aware enhanced variable window attention super-resolution method.

[0084] Compared with the prior art, the present invention has the following beneficial effects:

[0085] This paper achieves dynamic processing of remote sensing images by integrating image preprocessing, projective transformation, feature extraction, attention mechanisms, and super-resolution techniques. First, image quality is enhanced through preprocessing. Then, the image is segmented into multiple base windows, and feature vectors are extracted from them. Next, these feature vectors are used to predict surrogate parameters of the projective transformation matrix, decomposing it into transformation matrices for scaling, rotation, shearing, and translation. The reconstructed projective transformation matrix is ​​used to dynamically transform the base windows into arbitrary quadrilaterals. Based on this, a quadrilateral attention mechanism is used to extract rotation angle differences, which are then introduced into a variable window attention mechanism to enhance the model's dynamic adaptability. Finally, this mechanism is embedded into the SwinIR network to construct a high-precision super-resolution model. Evaluation on training and testing sets shows that the model can process remote sensing images in real time, outputting high-resolution images of the target area, significantly improving the efficiency and accuracy of remote sensing image analysis. Attached Figure Description

[0086] Figure 1 A flowchart of the variable window attention super-resolution method provided in an embodiment of the present invention;

[0087] Figure 2 A schematic diagram comparing the variable window attention super-resolution method provided in this embodiment of the invention with existing methods in a super-resolution reconstruction task;

[0088] Figure 3 This is a schematic diagram illustrating the average segmentation of the original remote sensing image using a normal window within a variable window, as provided in an embodiment of the present invention.

[0089] Figure 4 This is a schematic diagram illustrating the arbitrary segmentation of an original remote sensing image using an irregularly shaped window within a variable window, as provided in an embodiment of the present invention.

[0090] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0091] The markings in the attached diagram are as follows:

[0092] 100. Electronic equipment; 101. Memory; 102. Processing gas; 103. Computer program; 104. Communication bus. Detailed Implementation

[0093] To further understand the content of this invention, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are merely illustrative and not limiting of the invention.

[0094] This invention provides a rotation-aware enhanced variable window attention super-resolution method, such as... Figure 1 As shown, it includes the following steps:

[0095] S1. Acquire the original remote sensing image of the target object and perform enhancement preprocessing on the acquired original remote sensing image;

[0096] S2. Divide the preprocessed original remote sensing image into several basic windows and extract feature vectors from each basic window;

[0097] S3. Define the projective transformation matrix. Based on the extracted feature vectors, use quadrilaterals to predict the surrogate parameters for scaling, rotation, shearing, translation, and projection in the projective transformation matrix. Then, decompose the projective transformation matrix into the corresponding transformation matrix based on the predicted surrogate parameters.

[0098] S4. Reconstruct the projective transformation matrix based on the decomposed transformation matrix, and use the reconstructed projective transformation matrix to dynamically convert each base window into an arbitrary quadrilateral.

[0099] S5. Based on the dynamic transformation of the base window, extract the difference in rotation angle between pixels at the same position in any two base windows after the transformation from the rotation transformation matrix, and based on the quadrilateral attention mechanism, introduce the extracted rotation angle difference as a rotation perception term into the quadrilateral attention mechanism to establish a variable window attention mechanism.

[0100] S6. Embed the variable window attention mechanism into the SwinIR network and apply it to the feature extraction layer to establish the architecture of a high-precision super-resolution model;

[0101] S7. Collect and acquire the remote sensing image dataset of the target object, divide the remote sensing image dataset into a training set and a test set, use the training set to input the architecture of the high-precision super-resolution model to obtain the trained high-precision super-resolution model, and use the test set to test and evaluate the trained high-precision super-resolution model.

[0102] S8. Acquire remote sensing images of the target object in real time, input the acquired real-time remote sensing images into the high-precision super-resolution model, and output a high-resolution image containing the target sampling area.

[0103] SwinIR is an image super-resolution (SR) network based on the Swin Transformer. It achieves dynamic processing of remote sensing images by integrating image preprocessing, projective transformation, feature extraction, attention mechanisms, and super-resolution techniques. First, image quality is enhanced through preprocessing. Then, the image is segmented into multiple base windows, and feature vectors are extracted from them. Next, these feature vectors are used to predict surrogate parameters of the projective transformation matrix, decomposing it into transformation matrices for scaling, rotation, shearing, and translation. The reconstructed projective transformation matrix is ​​used to dynamically transform the base windows into arbitrary quadrilaterals. Based on this, a quadrilateral attention mechanism is used to extract rotation angle differences, which are then introduced into a variable window attention mechanism to enhance the model's dynamic adaptability. Finally, this mechanism is embedded into the SwinIR network to construct a high-precision super-resolution model. Evaluation on training and test sets shows that this model can process remote sensing images in real time, outputting high-resolution images of the target area, significantly improving the efficiency and accuracy of remote sensing image analysis.

[0104] Furthermore, the acquired raw remote sensing images undergo enhancement preprocessing, including:

[0105] Based on the acquired raw remote sensing images, grayscale processing is performed on the raw remote sensing images to obtain grayscale images;

[0106] The grayscale image is decomposed into wavelet coefficients of different frequencies based on wavelet transform, and the high-frequency coefficients are thresholded. The image is then reconstructed using the thresholded high-frequency coefficients to obtain the original denoised remote sensing image.

[0107] Furthermore, the preprocessed original remote sensing image is segmented into several basic windows, and a feature vector is extracted from each basic window, including:

[0108] Based on a predefined window size, the preprocessed original remote sensing image is segmented into multiple basic windows of the same size;

[0109] For each base window, features are extracted, and the extracted features are transformed into query tokens, key tokens, and value tokens through linear projection. The function expression for linear projection is:

[0110]

[0111] In the formula, Q represents the query token, K represents the key token, and V represents the value token. Represents a linear transformation function. Represents the features within the base window;

[0112] The query token, key token, and value token corresponding to each basic window are concatenated to obtain the feature vector corresponding to the basic window.

[0113] Furthermore, a projective transformation matrix is ​​defined, and surrogate parameters for scaling, rotation, shearing, translation, and projection in the projective transformation matrix are predicted using quadrilaterals based on the extracted feature vectors, including:

[0114] Define the projective transformation matrix, and its functional expression is:

[0115]

[0116] In the formula, Define scaling, rotation, and shearing transformations. Define translation transformation. The projection vector defines how the observer perceives changes in the target as the depth dimension changes.

[0117] Existing remote sensing images with known projective transformation matrices were collected, and a neural network model consisting of average pooling layers, LeakyReLU activation layers, and convolutional layers was established. Features were extracted from the existing remote sensing images and input into the neural network model for training to obtain a quadrilateral prediction model.

[0118] The feature vectors extracted from the base window are input into the quadrilateral prediction model, which outputs surrogate parameters for scaling, rotation, shearing, translation, and projection. The functional expressions for the surrogate parameters are as follows:

[0119]

[0120] In the formula, For the nth type of proxy parameter, It is a convolutional layer. For activation layer, For average pooling layer, The eigenvectors are the linearly projected features.

[0121] Furthermore, the projective transformation matrix is ​​reconstructed based on the decomposed transformation matrix, and the reconstructed projective transformation matrix is ​​used to dynamically transform each base window into an arbitrary quadrilateral, including:

[0122] The predicted surrogate parameters for scaling, rotation, shearing, translation, and projection are applied to the defined projective transformation matrix, respectively, to obtain the scaling transformation matrix, rotation transformation matrix, shearing transformation matrix, translation transformation matrix, and projection transformation matrix in sequence.

[0123] The functional expression for the scaling transformation matrix is:

[0124]

[0125] In the formula, For scaling transformation matrix, This is the scaling factor of the base window along the x-axis. The scaling factor of the base window in the y-axis direction;

[0126] The functional expression for the rotation transformation matrix is:

[0127]

[0128] In the formula, Let be the rotation transformation matrix. The rotation angle of the base window;

[0129] The functional expression for the shear transformation matrix is:

[0130]

[0131] In the formula, Let be the rotation transformation matrix. The shearing factor of the base window in the x-axis direction. The shearing factor of the base window in the y-axis direction;

[0132] The functional expression for the translation transformation matrix is:

[0133]

[0134] In the formula, The translation transformation matrix is... This is the amount of translation of the base window along the x-axis. This is the amount of translation of the base window in the y-axis direction. The translation factor;

[0135] The functional expression for the projection transformation matrix is:

[0136]

[0137] In the formula, Let be the projection transformation matrix. This is the projection of the base window along the x-axis. It is the projection of the base window along the y-axis.

[0138] Based on the decomposed scaling, rotation, shearing, translation, and projection transformation matrices, the reconstructed projective transformation matrix is ​​established. The function expression for the reconstructed projective transformation matrix is ​​as follows:

[0139]

[0140] In the formula, This is the reconstructed projective transformation matrix.

[0141] The process involves treating a base window as a reference and, based on its features, transforming it into a target quadrilateral using a projective transformation matrix. This covers regions of varying positions, sizes, orientations, and shapes, thus adapting to diverse targets. Because the projective transformation does not preserve parallelism, length, or orientation, the generated quadrilateral is highly flexible in terms of position, size, orientation, and shape, effectively covering targets of different sizes, orientations, and shapes. The projective transformation matrix is ​​typically represented by an eight-parameter transformation matrix.

[0142] Furthermore, the reconstructed projective transformation matrix is ​​used to dynamically transform each base window into an arbitrary quadrilateral, including:

[0143] Identify the four vertices on the base window that are to be converted into an arbitrary quadrilateral, and obtain the original coordinates of the four vertices;

[0144] The original coordinates of the four vertices are converted into three-dimensional homogeneous coordinates, and the reconstructed projective transformation matrix is ​​used to calculate the three-dimensional homogeneous coordinates of the base window vertices before the transformation, so as to obtain the three-dimensional homogeneous coordinates of the base window vertices after the transformation.

[0145] Based on the three-dimensional homogeneous coordinates of the transformed vertices of the base window, determine the positions of the four vertices of the base window after the transformation, and after the positions of the transformed vertices are determined, obtain the original coordinates of any pixel point within the base window.

[0146] The original coordinates of any pixel are converted into 3D homogeneous coordinates. Then, the reconstructed projective transformation matrix is ​​used to calculate the 3D homogeneous coordinates of any pixel within the base window before the transformation, resulting in the transformed 3D homogeneous coordinates of any pixel within the base window. The function expression for the transformation of the pixel coordinates is as follows:

[0147]

[0148]

[0149] In the formula, The transformed coordinates of any pixel within the base window. The transformed 3D homogeneous coordinates of the k-th pixel within the base window. The three-dimensional homogeneous coordinates of any pixel within the base window before transformation;

[0150] By combining the four vertices of the base window and the transformed three-dimensional homogeneous coordinates of all pixels within the base window, an arbitrary quadrilateral after the transformation of the base window can be obtained.

[0151] Furthermore, the extracted rotation angle difference is introduced as a rotation perception term into the quadrilateral attention mechanism to establish a variable window attention mechanism, including:

[0152] Extract the rotation angle from the rotation transformation matrix corresponding to each basic window;

[0153] The rotation angle difference is calculated based on the rotation angles of any two base windows. The functional expression for the rotation angle difference is:

[0154]

[0155] In the formula, The difference in rotation angle, Let be the rotation angle after the transformation of the i-th base window. Let be the rotation angle of the j-th base window after transformation;

[0156] The calculated rotation angle difference is used as the rotation perception term, and combined with other features, a quadrilateral attention mechanism is introduced to obtain the variable window attention mechanism. The functional expression of the variable window attention mechanism is:

[0157]

[0158] In the formula, For variable window attention mechanism, Let d be the normalization function, and d be the vector dimension. The learnable rotational sensitivity coefficient, To query the dot product of the token and the key token.

[0159] Among them, a Deformable Quadrangle Attention (DQA) mechanism is proposed, which enables the model to dynamically determine the position, size, orientation and shape of each window.

[0160] When the two windows rotate in similar directions, that is When attention scores are enhanced, feature fusion is promoted for targets with similar orientations; when the rotation directions differ significantly, Attention scores decay as irrelevant features are suppressed.

[0161] Finally, a rotation-aware enhanced variable window attention mechanism is embedded into the SwinIR network to replace the original fixed window attention calculation, achieving high-precision super-resolution reconstruction of rotating features. Specifically, arbitrary quadrilateral windows are dynamically generated through projective transformation, and the output features of all windows are stitched together to obtain the target sampling region.

[0162] The present invention also provides a rotation-aware enhanced variable window attention super-resolution system, comprising:

[0163] The raw remote sensing image processing module is used to acquire raw remote sensing images of the target object and perform enhancement preprocessing on the acquired raw remote sensing images;

[0164] The feature vector extraction module is used to segment the preprocessed original remote sensing image into several basic windows and extract feature vectors from each basic window.

[0165] The projective transformation matrix decomposition module is used to define the projective transformation matrix, predict the surrogate parameters of scaling, rotation, shearing, translation and projection in the projective transformation matrix using quadrilaterals based on the extracted feature vectors, and decompose the projective transformation matrix into the corresponding transformation matrix based on the predicted surrogate parameters.

[0166] The basic window transformation module is used to reconstruct the projective transformation matrix based on the decomposed transformation matrix, and then use the reconstructed projective transformation matrix to dynamically transform each basic window into an arbitrary quadrilateral.

[0167] The attention mechanism establishment module is used to extract the rotation angle difference between pixels at the same position in any two base windows after transformation from the rotation transformation matrix based on the dynamic transformation of the base window. Based on the quadrilateral attention mechanism, the extracted rotation angle difference is introduced into the quadrilateral attention mechanism as a rotation perception term to establish a variable window attention mechanism.

[0168] The network architecture building module is used to embed the variable window attention mechanism into the SwinIR network and apply it to the feature extraction layer to build the architecture of a high-precision super-resolution model.

[0169] The model building module is used to collect and acquire remote sensing image datasets of the target object. The remote sensing image dataset is divided into training set and test set. The architecture of the high-precision super-resolution model is input using the training set to obtain the trained high-precision super-resolution model. The trained high-precision super-resolution model is tested and evaluated using the test set.

[0170] The image resolution processing module is used to acquire remote sensing images of the target object in real time, input the acquired real-time remote sensing images into the high-precision super-resolution model, and output a high-resolution image containing the target sampling area.

[0171] As described above, and in conjunction with Tables 1 and 2, to verify the effectiveness of the proposed method, this invention comprehensively compares it with current mainstream state-of-the-art super-resolution methods based on Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs). In terms of experimental setup, two super-resolution scaling factors, 2x and 4x, were used for evaluation, and pre-trained models provided by the original authors of each method were directly used to ensure fairness. Among CNN-based methods, we selected EDSR (Enhanced Deep Super-Resolution), RCAN (Recursive Convolutional Network), DPSR (Deep Parametric Super-Resolution), and IMDN (Image Multi-Scale Deep Network). Among Transformer-based methods, we compared LKFormer (Long-Range former with Long Distance Attention Mechanism), SwinIR (Swin Image Restoration), and TTST (Spatiotemporal Transformer). For evaluation metrics, we used Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) for quantitative analysis. We objectively measured the performance of each method by calculating the PSNR and SSIM values ​​between the high-resolution ground image and the super-resolution reconstructed image. Experimental results show that our proposed method based on the variable window attention mechanism outperforms these comparative methods in terms of rotation perception capability and super-resolution reconstruction quality while maintaining computational efficiency, especially when processing remote sensing images with complex geometries.

[0172] Table 1. Comparison of reconstruction performance of the present invention and existing super-resolution methods at 2x downsampling rate.

[0173]

[0174] Quantitative Evaluation: The rotation-aware enhanced variable window attention super-resolution method proposed in this invention demonstrates significant advantages in quantitative evaluation. As shown in Table 1, in the 2x downsampling task of the UC Merced Land-Use dataset, this method achieves a PSNR of 36.35dB and an SSIM of 0.9871 with only 2.15M parameters. While maintaining a lightweight model, the PSNR is improved by 0.46dB compared to the traditional CNN method EDSR and by 0.21dB compared to the Transformer benchmark SwinIR, with the SSIM value reaching the optimal level. Particularly noteworthy is that on the NWPU dataset, this method, with an SSIM of 0.9619, significantly outperforms SwinIR's 0.9386, validating its ability to reconstruct complex terrain structures. Compared to the TTST dataset (18.94M) with a larger parameter count, this method maintains competitiveness while achieving an 8.8x improvement in model efficiency. These data fully demonstrate that by combining projective transformation matrix decomposition and rotation sensing mechanism, this invention effectively improves the super-resolution reconstruction accuracy and geometric consistency of rotating features while maintaining computational efficiency.

[0175] Table 2. Comparison of reconstruction performance of the present invention and existing super-resolution methods for different land features at 2x downsampling rate.

[0176]

[0177] As shown in Table 2, the rotation-aware enhanced variable window attention super-resolution method proposed in this invention demonstrates comprehensive advantages in the quantitative evaluation of 13 typical terrain scenes. It surpasses all comparison models with an average PSNR of 34.29 dB, and improves upon traditional CNN methods EDSR (34.11 dB) and RCAN (34.03 dB) by 0.18 dB and 0.26 dB, respectively. Furthermore, it achieves optimal performance in 9 specific scene types (such as buildings 35.20 dB, golf courses 37.13 dB, highways 33.01 dB, etc.), particularly in structured target reconstruction. The surface performance is outstanding (buildings +0.22dB, highways +0.29dB), which verifies the strong adaptability of the projective transformation matrix decomposition and rotation perception mechanism to geometrically deformed targets. Although the improvement is small (0.06-0.09dB) in some natural scenes (such as forest 32.31dB, park 29.75dB), the method still maintains a stable advantage. Moreover, with a lightweight parameter count of only 2.15M, it is significantly better than the Transformer model TTST (18.94M) with a larger parameter count, which fully demonstrates its dual advantages in model efficiency and reconstruction accuracy.

[0178] like Figures 2-4As shown, the reconstruction results of the method of this invention are compared with those of existing super-resolution methods at a 2x downsampling rate. Qualitative evaluation: The rotation-aware enhanced variable window attention method proposed in this paper significantly improves the reconstruction quality of rotating ground objects in super-resolution reconstruction tasks. Taking the aircraft target in the dataset as an example ( Figure 3 and Figure 4 This method not only preserves clearer texture details (such as wing skin seams and engine air intake structures) but also effectively avoids the contour blurring problem of the IMDN method and the blocky artifacts of LKFormer. Compared with Transformer-type methods (SwinIR and TTST), this method, due to its explicit modeling of projective transformation and rotation-aware mechanisms, exhibits stronger geometric fidelity for objects with obvious directional characteristics, such as aircraft. The reconstruction results are closer to real high-resolution images in terms of edge sharpness and structural integrity. This advantage is particularly evident in the transition area between the aircraft fuselage and the background, verifying the adaptability of the variable quadrilateral window to rotating targets.

[0179] This invention also discloses a rotation-aware enhanced variable window attention super-resolution system, comprising:

[0180] The raw remote sensing image processing module is used to acquire raw remote sensing images of the target object and perform enhancement preprocessing on the acquired raw remote sensing images;

[0181] The feature vector extraction module is used to segment the preprocessed original remote sensing image into several basic windows and extract feature vectors from each basic window.

[0182] The projective transformation matrix decomposition module is used to define the projective transformation matrix, predict the surrogate parameters of scaling, rotation, shearing, translation and projection in the projective transformation matrix using quadrilaterals based on the extracted feature vectors, and decompose the projective transformation matrix into the corresponding transformation matrix based on the predicted surrogate parameters.

[0183] The basic window transformation module is used to reconstruct the projective transformation matrix based on the decomposed transformation matrix, and then use the reconstructed projective transformation matrix to dynamically transform each basic window into an arbitrary quadrilateral.

[0184] The attention mechanism establishment module is used to extract the rotation angle difference between pixels at the same position in any two base windows after transformation from the rotation transformation matrix based on the dynamic transformation of the base window. Based on the quadrilateral attention mechanism, the extracted rotation angle difference is introduced into the quadrilateral attention mechanism as a rotation perception term to establish a variable window attention mechanism.

[0185] The network architecture building module is used to embed the variable window attention mechanism into the SwinIR network and apply it to the feature extraction layer to build the architecture of a high-precision super-resolution model.

[0186] The model building module is used to collect and acquire remote sensing image datasets of the target object. The remote sensing image dataset is divided into training set and test set. The architecture of the high-precision super-resolution model is input using the training set to obtain the trained high-precision super-resolution model. The trained high-precision super-resolution model is tested and evaluated using the test set.

[0187] The image resolution processing module is used to acquire remote sensing images of the target object in real time, input the acquired real-time remote sensing images into the high-precision super-resolution model, and output a high-resolution image containing the target sampling area.

[0188] like Figure 5 As shown, the present invention also provides an electronic device 100 for a rotation-aware enhanced variable window attention super-resolution method; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.

[0189] The memory 101 can be used to store the computer program 103. The processor 102 implements the steps of the rotation-aware enhanced variable window attention super-resolution method described in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device 100 (such as audio data), etc. In addition, the memory 101 may include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.

[0190] The at least one processor 102 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or any conventional processor. The processor 102 is the control center of the electronic device 100, connecting various parts of the electronic device 100 via various interfaces and lines.

[0191] The memory 101 in the electronic device 100 stores multiple instructions to implement a rotation-aware, variable window attention super-resolution method.

[0192] The present invention also provides a method for storing modules / units integrated in the electronic device 100, if implemented as software functional units and sold or used as independent products, in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).

[0193] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0194] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0195] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0196] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A rotation-aware enhanced variable window attention super-resolution method, characterized in that, Includes the following steps: Acquire raw remote sensing images of the target object and perform enhancement preprocessing on the acquired raw remote sensing images; The preprocessed original remote sensing image is divided into several basic windows, and feature vectors are extracted from each basic window. Define the projective transformation matrix, use quadrilaterals to predict the surrogate parameters of scaling, rotation, shearing, translation and projection in the projective transformation matrix based on the extracted feature vectors, and decompose the projective transformation matrix into the corresponding transformation matrix based on the predicted surrogate parameters. Reconstruct the projective transformation matrix based on the decomposed transformation matrix, and use the reconstructed projective transformation matrix to dynamically convert each base window into an arbitrary quadrilateral; Based on the dynamic transformation of the base window, the rotation angle difference between pixels at the same position in any two base windows after transformation is extracted from the rotation transformation matrix. Based on the quadrilateral attention mechanism, the extracted rotation angle difference is introduced as a rotation perception term into the quadrilateral attention mechanism to establish a variable window attention mechanism. A variable window attention mechanism is embedded into the SwinIR network and applied to the feature extraction layer to establish the architecture of a high-precision super-resolution model. Collect a dataset of remote sensing images of the target object, divide the dataset into a training set and a test set, use the training set to input the architecture of the high-precision super-resolution model to obtain the trained high-precision super-resolution model, and use the test set to test and evaluate the trained high-precision super-resolution model. The system acquires remote sensing images of the target object in real time, inputs these real-time remote sensing images into a high-precision super-resolution model, and outputs a high-resolution image containing the target sampling area.

2. The rotation-aware enhanced variable window attention super-resolution method according to claim 1, characterized in that, The enhancement preprocessing of the acquired raw remote sensing images includes: Based on the acquired raw remote sensing images, grayscale processing is performed on the raw remote sensing images to obtain grayscale images; The grayscale image is decomposed into wavelet coefficients of different frequencies based on wavelet transform, and the high-frequency coefficients are thresholded. The image is then reconstructed using the thresholded high-frequency coefficients to obtain the original denoised remote sensing image.

3. The rotation-aware enhanced variable window attention super-resolution method according to claim 1, characterized in that, The process of segmenting the preprocessed original remote sensing image into several basic windows and extracting feature vectors from each basic window includes: Based on a predefined window size, the preprocessed original remote sensing image is segmented into multiple basic windows of the same size; For each base window, features are extracted, and the extracted features are transformed into query tokens, key tokens, and value tokens through linear projection. The function expression for linear projection is: In the formula, Q represents the query token, K represents the key token, and V represents the value token. Represents a linear transformation function. Represents the features within the base window; The query token, key token, and value token corresponding to each basic window are concatenated to obtain the feature vector corresponding to the basic window.

4. The rotation-aware enhanced variable window attention super-resolution method according to claim 3, characterized in that, The defined projective transformation matrix uses quadrilateral prediction of surrogate parameters for scaling, rotation, shearing, translation, and projection in the projective transformation matrix based on the extracted feature vectors, including: Define the projective transformation matrix, and its functional expression is: In the formula, Define scaling, rotation, and shearing transformations. Define translation transformation. The projection vector defines how the observer perceives changes in the target as the depth dimension changes. Existing remote sensing images with known projective transformation matrices were collected, and a neural network model consisting of average pooling layers, LeakyReLU activation layers, and convolutional layers was established. Features were extracted from the existing remote sensing images and input into the neural network model for training to obtain a quadrilateral prediction model. The feature vectors extracted from the base window are input into the quadrilateral prediction model, which outputs surrogate parameters for scaling, rotation, shearing, translation, and projection. The functional expressions for the surrogate parameters are as follows: In the formula, For the nth type of proxy parameter, It is a convolutional layer. For activation layer, For average pooling layer, The eigenvectors are the linearly projected features.

5. The rotation-aware enhanced variable window attention super-resolution method according to claim 4, characterized in that, The step of reconstructing the projective transformation matrix based on the decomposed transformation matrix, and then using the reconstructed projective transformation matrix to dynamically convert each base window into an arbitrary quadrilateral, includes: The predicted surrogate parameters for scaling, rotation, shearing, translation, and projection are applied to the defined projective transformation matrix, respectively, to obtain the scaling transformation matrix, rotation transformation matrix, shearing transformation matrix, translation transformation matrix, and projection transformation matrix in sequence. The functional expression for the scaling transformation matrix is: In the formula, For scaling transformation matrix, This is the scaling factor of the base window along the x-axis. The scaling factor of the base window in the y-axis direction; The functional expression for the rotation transformation matrix is: In the formula, Let be the rotation transformation matrix. The rotation angle of the base window; The functional expression for the shear transformation matrix is: In the formula, Here is the shear transformation matrix. The shearing factor of the base window in the x-axis direction. The shearing factor of the base window in the y-axis direction; The functional expression for the translation transformation matrix is: In the formula, The translation transformation matrix is... This is the amount of translation of the base window along the x-axis. This is the amount of translation of the base window in the y-axis direction. The translation factor; The functional expression for the projection transformation matrix is: In the formula, Let be the projection transformation matrix. This is the projection of the base window along the x-axis. It is the projection of the base window along the y-axis. Based on the decomposed scaling, rotation, shearing, translation, and projection transformation matrices, the reconstructed projective transformation matrix is ​​established. The function expression for the reconstructed projective transformation matrix is ​​as follows: In the formula, This is the reconstructed projective transformation matrix.

6. The rotation-aware enhanced variable window attention super-resolution method according to claim 5, characterized in that, The process of dynamically converting each base window into an arbitrary quadrilateral using the reconstructed projective transformation matrix includes: Identify the four vertices on the base window that are to be converted into an arbitrary quadrilateral, and obtain the original coordinates of the four vertices; The original coordinates of the four vertices are converted into three-dimensional homogeneous coordinates, and the reconstructed projective transformation matrix is ​​used to calculate the three-dimensional homogeneous coordinates of the base window vertices before the transformation, so as to obtain the three-dimensional homogeneous coordinates of the base window vertices after the transformation. Based on the three-dimensional homogeneous coordinates of the transformed vertices of the base window, determine the positions of the four vertices of the base window after the transformation, and after the positions of the transformed vertices are determined, obtain the original coordinates of any pixel point within the base window. The original coordinates of any pixel are converted into 3D homogeneous coordinates. Then, the reconstructed projective transformation matrix is ​​used to calculate the 3D homogeneous coordinates of any pixel within the base window before the transformation, resulting in the transformed 3D homogeneous coordinates of any pixel within the base window. The function expression for the transformation of the pixel coordinates is as follows: In the formula, The transformed coordinates of any pixel within the base window. The transformed 3D homogeneous coordinates of the k-th pixel within the base window. The three-dimensional homogeneous coordinates of any pixel within the base window before transformation; By combining the four vertices of the base window and the transformed three-dimensional homogeneous coordinates of all pixels within the base window, an arbitrary quadrilateral after the transformation of the base window can be obtained.

7. The rotation-aware enhanced variable window attention super-resolution method according to claim 5, characterized in that, The step of introducing the extracted rotation angle difference as a rotation perception term into the quadrilateral attention mechanism to establish a variable window attention mechanism includes: Extract the rotation angle from the rotation transformation matrix corresponding to each basic window; The rotation angle difference is calculated based on the rotation angles of any two base windows. The functional expression for the rotation angle difference is: In the formula, The difference in rotation angle, Let be the rotation angle after the transformation of the i-th base window. Let be the rotation angle of the j-th base window after transformation; The calculated rotation angle difference is used as the rotation perception term, and combined with other features, a quadrilateral attention mechanism is introduced to obtain the variable window attention mechanism. The functional expression of the variable window attention mechanism is: In the formula, For variable window attention mechanism, Let d be the normalization function, and d be the vector dimension. The learnable rotational sensitivity coefficient, To query the dot product of the token and the key token.

8. A rotation-aware enhanced variable window attention super-resolution system, based on the rotation-aware enhanced variable window attention super-resolution method according to any one of claims 1 to 7, characterized in that, include: The raw remote sensing image processing module is used to acquire raw remote sensing images of the target object and perform enhancement preprocessing on the acquired raw remote sensing images; The feature vector extraction module is used to segment the preprocessed original remote sensing image into several basic windows and extract feature vectors from each basic window. The projective transformation matrix decomposition module is used to define the projective transformation matrix, predict the surrogate parameters of scaling, rotation, shearing, translation and projection in the projective transformation matrix using quadrilaterals based on the extracted feature vectors, and decompose the projective transformation matrix into the corresponding transformation matrix based on the predicted surrogate parameters. The basic window transformation module is used to reconstruct the projective transformation matrix based on the decomposed transformation matrix, and then use the reconstructed projective transformation matrix to dynamically transform each basic window into an arbitrary quadrilateral. The attention mechanism establishment module is used to extract the rotation angle difference between pixels at the same position in any two base windows after transformation from the rotation transformation matrix based on the dynamic transformation of the base window. Based on the quadrilateral attention mechanism, the extracted rotation angle difference is introduced into the quadrilateral attention mechanism as a rotation perception term to establish a variable window attention mechanism. The network architecture building module is used to embed the variable window attention mechanism into the SwinIR network and apply it to the feature extraction layer to build the architecture of a high-precision super-resolution model. The model building module is used to collect and acquire remote sensing image datasets of the target object. The remote sensing image dataset is divided into training set and test set. The architecture of the high-precision super-resolution model is input using the training set to obtain the trained high-precision super-resolution model. The trained high-precision super-resolution model is tested and evaluated using the test set. The image resolution processing module is used to acquire remote sensing images of the target object in real time, input the acquired real-time remote sensing images into the high-precision super-resolution model, and output a high-resolution image containing the target sampling area.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of a rotation-aware enhanced variable window attention super-resolution method according to any one of claims 1 to 7.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of a rotation-aware enhanced variable window attention super-resolution method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Hierarchical Transform high-resolution remote sensing image semantic segmentation method and system

    CN116258976A

  • Remote sensing image reconstruction method and device based on multi-scale window space channel attention

    CN120013761A

Cited By

  • A multi-frame fusion night vision image super-resolution enhancement method

    CN122415332A