Satellite image restoration method based on edge condition expected attention
By introducing edge condition expectation attention mechanism in satellite image repair, using the Canny-Sobel hybrid edge sampler to extract edge information, the Transformer model's shortcomings in high-frequency edge information extraction are solved, and a more stable and accurate satellite image repair effect is achieved.
Patent Information
- Application Number
- CN202510255797.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
The existing satellite image repair method based on Transformer lacks the ability to extract high-frequency edge information, resulting in poor repair results.
A satellite image repair method based on edge condition desired attention is proposed. Multi-scale edge information is extracted through the Canny-Sobel hybrid edge sampler and coupled it to the self-attention calculation process, enhancing the weight of edge information between similar image blocks and ensuring that the global context focuses on key texture information in satellite images.
Effectively utilize conditions to capture long-distance dependencies, reduce excessive deviations, ensure a balance between edge enhancement and non-edge area smoothing, improve the stability and accuracy of satellite image repair, and provide visually coherent repair effects.
Smart Images

Figure CN120182142A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of satellite image restoration, and particularly relates to a satellite image restoration method based on edge conditional expectation attention. Background Art
[0002] Satellite images are crucial in analyzing human activities and the natural environment. The advancement of computing technology has promoted the wide application of satellite images in various fields, including promoting the sustainable development of human society, agricultural development, and ecosystem management. However, due to factors such as the harsh electromagnetic environment in space and the mechanical jitter of the satellite itself, satellite images usually have various defects or partial information loss, which greatly reduces the usability of satellite images. For example, Landsat-7, one of the most commonly used satellites, had a permanent failure of its scanline corrector (SLC) of the Enhanced Thematic Mapper Plus (ETM+) sensor on May 31, 2003, resulting in overlapping stripes in its scanned images. The width of the stripes gradually increased from 1 pixel to 14 pixels, and a huge wedge-shaped gap formed from the middle to both sides of the image, with approximately 20% of the data lost. As Figure 1 shown in the image taken by the Landsat-7 satellite in the area near 119.58° east longitude and 33.21° north latitude (area: 23,552 square kilometers) under the condition of scanline corrector failure. The data loss has seriously affected the application of Landsat-7 satellite images in various research fields. In addition, damaged historical satellite images have also affected the research on aspects such as geological evolution, urban expansion, and forest cover change.
[0003] Given these challenges, to ensure the availability of satellite images in a wide range of applications, the improvement of image inpainting techniques becomes crucial. Traditional image inpainting methods are generally divided into diffusion-based and patch-based methods. Diffusion-based methods propagate adjacent information to the missing regions of the image (M. Bertalmio, G. Sapiro, V. Caselles, and C. Ballester, “Image inpainting,” in Proceedings of the 27th annual conference on Computer graphics and interactive techniques, pp. 417–424, 2000). The literature (S. Esedoglu and J. Shen, “Digital inpainting based on the mumford–shah–euler image model,” European Journal of Applied Mathematics, vol. 13, no. 4, pp. 353–370, 2002) introduced the Euler elastica operator for image smoothing and applied the Mumford-Shah segmentation model to image inpainting. The literature (D. Liu, X. Sun, F. Wu, S. Li, and Y.-Q. Zhang, “Image compression with edge-based inpainting,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 17, no. 10, pp. 1273–1287, 2007) proposed to extract the binary edge information of the image within the encoder and repair the damaged regions with distance-related structures by means of the decoder. However, diffusion-based methods are limited to the locally available information near the missing image regions. These methods are unable to generate meaningful structures within the missing regions and are not effective enough when dealing with large missing regions. To address this problem, patch-based methods recover the missing parts by transferring information from similar regions within the source image (A. Criminisi, P. Perez, and K. Toyama, “Region filling and object′ removal by exemplar-based image inpainting,” IEEE Transactions on image processing, vol. 13, no. 9, pp. 1200–1212, 2004).However, this process is computationally expensive because it requires calculating the similarity between the image patches to be repaired and all the image patches in the set. To address the problem of excessive computational requirements, the literature (C. Barnes, E. Shechtman, A. Finkelstein, and D. B. Goldman, “Patchmatch: A randomized correspondence algorithm for structural image editing,” ACM Trans. Graph., vol. 28, no. 3, p. 24, 2009.) proposed a nearest neighbor algorithm to quickly find the corresponding image patches and then fuse these patches into the background. However, if the missing content does not exist in the source image, the accuracy of the repair will be greatly reduced. Moreover, due to the limited ability of these technologies to interpret image semantics, they face difficulties in recovering images from large-area data loss, especially for satellite images with complex terrain features.
[0004] In remote sensing image restoration, spectral-based restoration methods can utilize the multi-spectral or hyperspectral characteristics of satellite images to capture data at different wavelengths (X. Li, H. Shen, L. Zhang, H. Zhang, and Q. Yuan, “Dead pixel completion of aqua modis band 6 using a robust m-estimator multiregression,” IEEE Geoscience and Remote Sensing Letters, vol. 11, no. 4, pp. 768–772, 2014.). These methods rely on the correlation between spectral bands to reconstruct missing or damaged pixels. However, spectral-based methods also have limitations. They depend on complete information being available for all spectral bands. But in Landsat-7 scan line corrector failure images, there are pixel rasterization losses in each band. Although the loss situations at the edges of different bands may vary slightly, it severely restricts the practical application of spectral-based restoration methods. On the other hand, time series-based restoration methods utilize the time series of images collected in the same area (C. Zeng, H. Shen, and L. Zhang, “Recovering missing pixels for landsat etm+ slc-off imagery using multi-temporal regression analysis and a regularization method,” Remote Sensing of Environment, vol. 131, pp. 182–194, 2013.). By comparing the changes in images over time, numerical predictions can be made based on previous observations, and the predicted values are then used to fill in the missing data. Although time series-based restoration methods have achieved certain results, when the time series data is limited or sudden environmental changes occur, time series-based restoration methods will be restricted, making it challenging to restore areas without an accurate time pattern.
[0005] Thanks to the progress of deep learning technology, a Context Encoder (CE) based on the encoder-decoder architecture (D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. A. Efros, “Context encoders: Feature learning by inpainting,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2536–2544, 2016.) has been proposed in the field of image inpainting. The encoder maps the missing region to a low-dimensional feature space, and the decoder uses this space to construct the output image. However, the inpainted regions in the output image often appear smooth and blurred, which is attributed to the information bottleneck problem in the fully connected layer. A new model based on the context encoder adopts a fully convolutional network, a global discriminator, and a local discriminator (S. Iizuka, E. Simo-Serra, and H. Ishikawa, “Globally and locally consistent image completion,” ACM Transactions on Graphics (ToG), vol. 36, no. 4, pp. 1–14, 2017.). The fully convolutional network can process inputs of any size, and the global and local discriminators make the generated image content more realistic at the global and local scales, respectively. In addition, to expand the receptive field of the model, the dilated convolution layer is introduced for the first time to replace the fully connected layer in the context encoder model. However, the dilated convolution layer generates many sparse matrices, significantly increasing the computation time. Partial convolution only operates on non-missing pixels and performs averaging, effectively preventing the convolution filter from accumulating too many zeros when processing the missing information region (G. Liu, F. A. Reda, K. J. Shih, T.-C. Wang, A. Tao, and B. Catanzaro, “Image inpainting for irregular holes using partial convolutions,” in Proceedings of the European conference on computer vision (ECCV), pp. 85–100, 2018.).Generators and discriminators with encoder and decoder structures have been designed based on convolutional neural networks (CNNs) and generative adversarial networks (GANs) (I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Advances in neural information processing systems, vol. 27, 2014). This process involves inputting the damaged image into the generator, while the discriminator compares the generated image with the real image. This system utilizes loss functions such as the L1 loss function, adversarial loss function, style loss function, and perceptual loss function to enable the generator to fully analyze and extract image features. Through this mechanism, CNNs and GANs demonstrate the ability to capture the semantic features of pictures and effectively repair damaged regions. However, the inherent limitations of the CNN mechanism may affect the repair quality. On the one hand, the local receptive field wastes a large amount of spatially related information. The long-range dependence information in satellite images is more diverse and irregular, such as rivers, roads, and vegetation, which often extend, disperse, or aggregate throughout the image in various ways. The local processing method of CNNs lacks the ability to capture large-scale spatial relationships and global context in satellite images. On the other hand, parameter sharing is a double-edged sword. The fixed convolutional filter weights result in a consistent response to damaged and undamaged regions of the image, and during the repair process, over-smoothing may occur in regions that should not be smoothed, or details may be lost at key details.Although some methods adopt dynamic filters to make the weights of convolutional filters change dynamically during training (X. Jia, B. De Brabandere, T. Tuytelaars, and L. V. Gool, “Dynamic filter networks,” Advances in neural information processing systems, vol. 29, 2016.), there are also some methods that use non-local blocks or CCNet to address the long-range dependence problem inherent in convolutional neural networks (X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7794–7803, 2018.).
[0006] Based on the multi-head self-attention mechanism, Transformer establishes global self-attention through the dynamically learnable correlations between queries and keys and the weighted sum with values, playing a key role in the field of natural language processing (A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017.)。The Vision Transformer (i.e., ViTs) model marks the first application of Transformer in the field of computer vision (A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020.)。ViTs dynamically learns queries and keys to generate global weights and performs a weighted sum of values to produce global context features. The self-attention layer can compute exact context mappings, and then the feed-forward layer assigns the results of these context mappings to the desired output values (C. Yun, S. Bhojanapalli, A. S. Rawat, S. Reddi, and S. Kumar, “Are transformers universal approximators of sequence-to-sequence functions?,” in International Conference on Learning Representations, 2020.)。Given the inherent property of ViTs to obtain global attention by computing an explicit dot product matrix, the Transformer method is more suitable for satellite image inpainting than CNNs. Compared with convolutional neural networks, ViTs can capture long-range dependencies in satellite images.On this basis, the Bidirectional Autoregressive Transformer optimizes the ICT architecture by incorporating bidirectional and autoregressive Transformers, thus enhancing the model's inpainting ability (Y.Yu, F.Zhan, R.Wu, J.Pan, K.Cui, S.Lu, F.Ma, X.Xie, and C.Miao, “Diverse image inpainting with bidirectional and autoregressive transformers,” in Proceedings of the 29th ACM International Conference on Multimedia, pp.69–78, 2021.). Tfill designs restricted CNNs to extract non-overlapping tokens within the receptive field and introduces an attention-aware layer to calculate the attention of the known region and the generated region respectively (C.Zheng, T.-J.Cham, J.Cai, and D.Phung, “Bridging global context interactions for high-fidelity image completion,” in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp.11512–11522, 2022.). Restormer adopts per-channel self-attention combined with depthwise separable convolution for extensive context feature extraction (S.W.Zamir, A.Arora, S.Khan, M.Hayat, F.S.Khan, and M.-H.Yang, “Restormer: Efficient transformer for high-resolution image restoration, in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pp.5728–5739, 2022.).The T-former innovates in the attention function, using a second-order Taylor series to approximate the softmax function, reducing the computational complexity of the dot product matrix (Y. Deng, S. Hui, S. Zhou, D. Meng, and J. Wang, “T-former: An efficient transformer for image inpainting,” in Proceedings of the 30th ACM International Conference on Multimedia, pp. 6559–6568, 2022.). However, different from CNNs, the Transformer backbone is global, and its self-attention will reduce the high-frequency signals in the image, acting as a low-pass filter (N. Park and S. Kim, “How do vision transformers work?,” in International Conference on Learning Representations, 2022.). Therefore, existing Transformer-based satellite image inpainting methods lack the ability to extract high-frequency edge information. To enhance the high-frequency edge information extraction ability of Transformer-based satellite image inpainting methods, it is an urgent problem to propose a new satellite image inpainting method. Summary of the Invention
[0007] The purpose of the present invention is to solve the problem that existing Transformer-based satellite image inpainting methods lack the ability to extract high-frequency edge information, and a satellite image inpainting method based on edge conditional expectation attention is proposed.
[0008] The technical solution adopted by the present invention to solve the above technical problems is: a satellite image inpainting method based on edge conditional expectation attention, and the method specifically includes the following steps:
[0009] Step 1: Obtain a dataset of real undamaged satellite images and a dataset of loss mask matrices The mask matrices in the mask matrix dataset only contain elements 0 and 1, where 1 represents that the pixel at the corresponding position in the satellite image is damaged, and 0 represents that the pixel at the corresponding position in the satellite image is undamaged;
[0010] According to and obtain a dataset of damaged satellite images dataset The way to obtain any damaged satellite image X in the is:
[0011]
[0012] Among them, X G1 represents any one image in the data set M G1 represents any one mask matrix in the data set , represents the multiplication of the elements at the corresponding positions of matrix X G1 and 1 - M G1 ;
[0013] For each satellite image in the data set , edge extraction is performed respectively to obtain the real undamaged edge image set ε G , for each satellite image in the data set , edge extraction is performed respectively to obtain the damaged edge image set ε I , and the data sets ε G and ε I are used to form the training data set;
[0014] Step 2: Construct a generative adversarial network for satellite image restoration. The generative adversarial network includes a generator and a discriminator. The working process of the generator is as follows:
[0015] For any damaged satellite image in the data set H0 and W0 represent the height and width of image X. The damaged edge image corresponding to X in the data set ε is denoted as I After passing through the first PE layer, the damaged edge image E is reshaped into Then is passed through the first downsampling unit, and the output of the first downsampling unit is denoted as
[0016] is used as the input of the second downsampling unit, and the output of the second downsampling unit is denoted as
[0017] is used as the input of the third downsampling unit, and the output of the third downsampling unit is denoted as
[0018]
[0019] is used as the input of the first residual unit, and the output of the first residual unit is used as the input of the first upsampling unit. The output of the first upsampling unit is denoted as
[0019] is used asAs the input of the second upsampling unit, denote the output of the second upsampling unit as
[0020] Denote As the input of the third upsampling unit, denote the output of the third upsampling unit as
[0021] Denote the damaged satellite image As the input of the second PE layer, and after passing through the second PE layer, it is reshaped into Let \(H\times W\) denote the number of image patches into which image \(X\) is divided. Denote As the input of the first encoding unit;
[0022] Take the output of the first encoding unit as the input of the fourth downsampling unit, and denote the output of the fourth downsampling unit And the output of the first downsampling unit As the input of the second encoding unit;
[0023] Take the output of the second encoding unit as the input of the fifth downsampling unit, and denote the output of the fifth downsampling unit And the output of the second downsampling unit As the input of the third encoding unit;
[0024] Take the output of the third encoding unit as the input of the sixth downsampling unit, and denote the output of the sixth downsampling unit And the output of the third downsampling unit as As the input of the fourth encoding unit;
[0025] Take the output of the fourth encoding unit as the input of the fourth upsampling unit, and denote the output of the fourth upsampling unit Concatenate it with the output of the third encoding unit to obtain the concatenation result \(a\). Pass the concatenation result \(a\) through the first convolutional layer, and denote the output of the first convolutional layer and the output of the first upsampling unit As the input of the first decoding unit;
[0026] Take the output of the first decoding unit as the input of the fifth upsampling unit, and denote the output of the fifth upsampling unit Concatenate it with the output of the second encoding unit to obtain the concatenation result \(b\). Pass the concatenation result \(b\) through the second convolutional layer;
[0027] Denote the output of the second convolutional layer and the output of the second upsampling unit As the input of the second decoding unit, and take the output of the second decoding unit as the input of the sixth upsampling unit;
[0028] Denote the output of the sixth upsampling unit Concatenate with the output of the first coding unit to obtain the concatenation result c, and pass the concatenation result c through the third convolutional layer;
[0029] Take the output of the third convolutional layer and the output of the third upsampling unit as the input of the third decoding unit. Pass the output of the third decoding unit through the fourth convolutional layer, the second residual unit, the seventh upsampling unit, and the eighth upsampling unit in sequence, and take the output of the eighth upsampling unit as the output of the generator;
[0030] Step 3: Use the obtained training dataset to train the constructed generative adversarial network until the loss function converges, and then stop training to obtain the trained generative adversarial network;
[0031] Step 4: Input the damaged satellite image to be restored into the generator in the trained generative adversarial network, and use the generator to output the restored image.
[0032] The beneficial effects of the present invention are as follows:
[0033] Each part of the present invention is based on the edge-conditioned expectation attention model. The Canny-Sobel hybrid edge sampler can capture more prominent edges and intensity changes by multi-scale sampling of the edge information of the satellite image, and couple the satellite image edge information into the self-attention calculation process. It is used as the conditional expectation in the self-attention calculation, amplifying the weight of the edge information between similar image patches, ensuring that image patches with similar content and edges can obtain higher attention, thereby focusing the global context on the key texture information in the satellite image. The edge-conditioned expectation attention model can effectively utilize the conditional expectation to capture long-distance dependencies in satellite image restoration, and applies the Chebyshev inequality in the edge-conditioned expectation attention mechanism to control the deviation to reduce excessive deviation, ensuring the balance between edge enhancement and non-edge region smoothing, and helping to reduce the risks of excessive edge enhancement and excessive smoothing in non-edge regions.
[0034] Experimental data and image results show that the method of the present invention performs excellently in restoring key features. Based on the extracted high-frequency edge information, it can ensure more stable and accurate satellite image restoration and provide a visually coherent restoration effect. Description of the Drawings
[0035] Figure 1 is an image taken by the Landsat-7 satellite in the case of a scan line corrector failure;
[0036] The original image taken is on the left, and the enlarged view within the red frame is on the right;
[0037] Figure 2 is the structural diagram of the generator of the present invention;
[0038] Residual Block represents the residual unit, Down represents the downsampling unit, Up represents the upsampling unit, Encoder represents the encoding unit, and Decoder represents the decoding unit;
[0039] Figure 3 It is a comparison chart of the effects between the traditional attention mechanism and the edge conditional expectation attention mechanism of the present invention;
[0040] The present invention calculates the score matrix obtained by performing q1(K + E) on Patch 1. T Compared with the score matrix calculated based on q1K, T the weight of dissimilar patches at the edge is reduced. In this case, except that Patch 1 has the highest similarity with itself, the similarity at the edge between Patch 1 and Patch 2 increases the weight of the softmax score. On the contrary, the similarity at the edge between Patch 1, Patch 3, and Patch 4 is low, reducing the weight of the final score;
[0041] Figure 4 It is obtained through and to obtain the comparison chart of attention features;
[0042] Among them, each row represents a group of comparison charts, indicating that the bright area in the attention feature map is the fusion of the former two;
[0043] Figure 5 It is a comparison chart of the results obtained by performing edge extraction using the Canny algorithm, Sobel operator, and CSHE of the present invention;
[0044] Figure 6 It is a schematic diagram of the expectation adjustment process based on Chebyshev;
[0045] Figure 7 It is a comparison chart of the Token distribution between different image patches before and after the expectation adjustment based on Chebyshev;
[0046] Figure 8 It is a repair result chart of the strip loss of Landsat-7 images;
[0047] (a) is the real satellite image; (b) is the damaged satellite image; (c) is the edge after downsampling and upsampling; (d) is the repaired output satellite image;
[0048] Figure 9 It is a repair result chart of the irregular missing part in the SpaceNet image;
[0049] (a) is a real satellite image; (b) is a damaged satellite image; (c) is the edge after downsampling and upsampling; (d) is the restored output satellite image;
[0050] Figure 10 It is a comparison chart of the restoration results of the model of the present invention and other methods on the Landsat-7 dataset;
[0051] Figure 11 It is a comparison chart of the restoration results of the model of the present invention and other methods on the SpaceNet dataset;
[0052] Figure 12(a) is a box plot of the PSNR values of each model at different mask ratios on the Landsat-7 dataset;
[0053] Figure 12(b) is a box plot of the SSIM values of each model at different mask ratios on the Landsat-7 dataset;
[0054] Figure 13 It is the influence of different Chebyshev expectation adjustment hyperparameters k on the image restoration results in the SpaceNet dataset;
[0055] (a) is the image restoration result; (b) is the difference E diff [v];
[0056] Figure 14 It is a graph of the image restoration results using different edge detectors;
[0057] From top to bottom, they are the restoration result graphs of the non-edge, Canny, Sobel, and Canny-Sobel hybrid edge (Canny-Sobel Hybrid Edge, CSHE) samplers. Detailed implementation manner
[0058] Detailed implementation manner one: Combine Figure 2 To illustrate this implementation manner. A satellite image restoration method based on edge-conditioned expectation attention described in this implementation manner specifically includes the following steps:
[0059] Step one: Obtain a dataset of real undamaged satellite images And a dataset of loss mask matrices The mask matrices in the mask matrix dataset Only contain elements 0 and 1, where 1 represents that the pixel at the corresponding position in the satellite image is damaged, and 0 represents that the pixel at the corresponding position in the satellite image is not damaged;
[0060] According to And Obtain a dataset of damaged satellite images Data set The acquisition method of any damaged satellite image X in the data set is as follows:
[0061]
[0062] Among them, X G1 represents any image in the data set M G1 represents any mask matrix in the data set , represents the matrix X G1 and 1 - M G1 multiply the corresponding elements (i.e., Hadamard element-wise multiplication), and 1 represents a matrix with all elements being 1;
[0063] For each satellite image in the data set perform edge extraction respectively to obtain the set of real undamaged edge images ε G , for each satellite image in the data set perform edge extraction respectively to obtain the set of damaged edge images ε I , use the data set ε G and ε I to form a training data set;
[0064] Step 2: Construct a generative adversarial network for satellite image repair. The generative adversarial network includes a generator and a discriminator. The working process of the generator is as follows:
[0065] For any damaged satellite image in the data set H0 and W0 represent the height and width of the image X. Denote the damaged edge image corresponding to X in the data set ε as I After passing through the first PE layer (Patch Embedding), the damaged edge image E is reshaped into Then After passing through the first downsampling unit, denote the output of the first downsampling unit as
[0066] Use as the input of the second downsampling unit, and denote the output of the second downsampling unit as
[0067] Use as the input of the third downsampling unit, and denote the output of the third downsampling unit as
[0068] Use As the input of the first residual unit, the output of the first residual unit is used as the input of the first upsampling unit, and the output of the first upsampling unit is denoted as
[0069] Use as the input of the second upsampling unit, and the output of the second upsampling unit is denoted as
[0070] Use as the input of the third upsampling unit, and the output of the third upsampling unit is denoted as (Output Restore to the original size before inputting to the first PE layer);
[0071] Use the damaged satellite image as the input of the second PE layer, and after passing through the second PE layer, it is reshaped into (The channel dimension of the output result of the second PE layer is the same as that of the output result of the first PE layer), H×W represents the number of image blocks into which image X is divided. Use as the input of the first encoding unit;
[0072] Use the output of the first encoding unit as the input of the fourth downsampling unit, and the output of the fourth downsampling unit and the output of the first downsampling unit as the input of the second encoding unit;
[0073] Use the output of the second encoding unit as the input of the fifth downsampling unit, and the output of the fifth downsampling unit and the output of the second downsampling unit as the input of the third encoding unit;
[0074] Use the output of the third encoding unit as the input of the sixth downsampling unit, and the output of the sixth downsampling unit and the output of the third downsampling unit are denoted as as the input of the fourth encoding unit;
[0075] Use the output of the fourth encoding unit as the input of the fourth upsampling unit, and the output of the fourth upsampling unit is concatenated with the output of the third encoding unit to obtain the concatenation result a. The concatenation result a passes through the first convolutional layer, and the output of the first convolutional layer and the output of the first upsampling unit are used as the input of the first decoding unit;
[0076] Use the output of the first decoding unit as the input of the fifth upsampling unit, and the output of the fifth upsampling unit Concatenate with the output of the second encoding unit to obtain the concatenation result b, and pass the concatenation result b through the second convolutional layer;
[0077] Take the output of the second convolutional layer and the output of the second upsampling unit as the input of the second decoding unit, and take the output of the second decoding unit as the input of the sixth upsampling unit;
[0078] Take the output of the sixth upsampling unit Concatenate with the output of the first encoding unit to obtain the concatenation result c, and pass the concatenation result c through the third convolutional layer;
[0079] Take the output of the third convolutional layer and the output of the third upsampling unit as the input of the third decoding unit, and pass the output of the third decoding unit through the fourth convolutional layer, the second residual unit, the seventh upsampling unit, and the eighth upsampling unit in sequence. Take the output of the eighth upsampling unit as the output of the generator;
[0080] Step 3: Use the obtained training dataset to train the constructed generative adversarial network until the loss function converges, and then stop training to obtain the trained generative adversarial network;
[0081] Step 4: Input the damaged satellite image to be restored into the generator in the trained generative adversarial network, and use the generator to output the restored image.
[0082] In downstream tasks, satellite images are usually very large. Therefore, it is necessary to segment these large images and stitch them after restoration. However, the stitching after restoration by traditional methods may cause the checkerboard effect. Therefore, the present invention proposes an overlapping segmentation method. First, determine the geometric center of the original satellite image, and then crop the largest square image outward from the geometric center. Denote the size of the cropped image as (i, i); then, segment the cropped image with a step size of s. The size of each image patch obtained by segmentation is (p, p), and the number of image patches that can be completely segmented along a single dimension is The number of remaining pixels after overlapping segmentation along one dimension is:
[0083] r = i - p + s(1 - n′), r ≥ 0
[0084] If r = 0, perfect segmentation will be achieved. In a more common case, r > 0. Use the mirror reflection filling method to fill the remaining pixels to a size of p. The total number of image patches n is given by the following formula:
[0085]
[0086] After each image patch is inpainted, the image patches are stitched together according to their original positions. In the width and height dimensions, the overlapping region of size p-s is averaged, and the mirror padding regions are removed in both directions. This method eliminates the checkerboard effect without changing the structure of the (i,i) sized image.
[0087] Specific Embodiment 2: The difference between this embodiment and Specific Embodiment 1 is that each satellite image in the dataset is separately subjected to edge extraction to obtain a set of real edge images ε G , and the specific process is as follows:
[0088] For any undamaged satellite image in the dataset , the Canny edge detection algorithm is used to perform edge extraction on this undamaged satellite image to obtain a first edge image, and then the Sobel operator is used to perform edge extraction on this undamaged satellite image to obtain a second edge image;
[0089] The average of the first edge image and the second edge image is taken, and the average result is used as the edge image corresponding to this undamaged satellite image.
[0090] Other steps and parameters are the same as those in Specific Embodiment 1.
[0091] The present invention introduces the Canny-Sobel Hybrid Edge (CSHE) method. The Canny algorithm effectively reduces noise through Gaussian smoothing, detects edges using a double-threshold process, and then performs edge connection to obtain continuous edges. When the Canny edge detection algorithm is used to perform edge extraction on an undamaged satellite image, the value of each pixel point in the obtained edge image is either 0 or 1; the Sobel operator, as a first-order derivative filter, can highlight significant gradient changes and is particularly effective in capturing prominent edges. When the Sobel operator is used to perform edge extraction on an undamaged satellite image, the value of each pixel point in the obtained edge image is between 0 and 1, and the first edge image and the second edge image are of the same size. The average of the values of the pixel points at the same position in the two edge images is taken, and the average result is used as the value of the pixel point in the final edge image. This embodiment obtains the final edge extraction result by averaging the edge extraction results of the Canny edge detection algorithm and the Sobel operator, which can create a complementary edge map to generate a more balanced edge representation and strengthen the self-attention mechanism of the Transformer.
[0092] Such as Figure 5As shown, the Canny algorithm can capture sharp features, but the generated edge intensities are relatively uniform; while Sobel performs well in detecting significant geographical features with varying edge intensities, but has difficulty in suppressing noise. The combination of (Canny + Sobel) / 2 utilizes the advantages of both methods, being able to retain large-scale edges with different intensities and reduce noise. This enables the Transformer to better capture spatial features at different levels of detail, thereby enhancing the effects of tasks such as image reconstruction and restoration. Similarly, the method of this embodiment is also used for edge extraction of damaged satellite images.
[0093] Specific Embodiment 3: The difference between this embodiment and Specific Embodiment 1 or 2 is that there is 1 Transformer block in the first encoding unit, 2 serially connected Transformer blocks in the second encoding unit, 3 serially connected Transformer blocks in the third encoding unit, 4 serially connected Transformer blocks in the fourth encoding unit, 1 Transformer block in the first decoding unit, 2 serially connected Transformer blocks in the second decoding unit, and 3 serially connected Transformer blocks in the third decoding unit; where:
[0094] The attention layers in each Transformer block of the second encoding unit, the third encoding unit, the fourth encoding unit, the first decoding unit, the second decoding unit, and the third decoding unit are all Edge-Conditional Expectation Attention (ECEA) layers, and a Chebyshev expectation adjustment layer is connected after the edge-conditional expectation attention layer.
[0095] Other steps and parameters are the same as those in Specific Embodiment 1 or 2.
[0096] The Transformer block of this embodiment is obtained by improving the encoder of the traditional Transformer model. The encoder of the traditional Transformer model is used as the basic Transformer block, and then the attention layer in the basic Transformer block is replaced with an edge-conditioned expectation attention layer, and a Chebyshev expectation adjustment layer is connected after each edge-conditioned expectation attention layer to obtain an improved Transformer block. In this embodiment, except for the Transformer block in the first coding unit (the attention layer in the first coding unit is a standard attention layer), each other Transformer block is an improved Transformer block. In the cascade structure, the output of the previous Transformer block is the input of the next Transformer block. For example, for the second coding unit, the output of the first downsampling unit will be used as the input of the edge-conditioned expectation attention layer in each Transformer block in the second coding unit. In the Transformer block, the output of the layer before the edge-conditioned expectation attention layer is transformed to obtain the query Q, key K, and value V, and the output of the first downsampling unit is used as the edge E. Then, the ECEA attention calculation in the Transformer block is performed using the transformed query Q, key K, value V, and the edge E output by the first downsampling unit. The Chebyshev expectation adjustment layer is used to adjust the ECEA attention calculation result in the Transformer block, and the adjustment result will be used as the input of the next layer.
[0097] The i-th coding unit is connected to the (4 - i)-th decoding unit through a 1×1 convolution to help the decoding unit generate image content. After the last decoding unit, a 9×9 convolution is first used, followed by eight residual blocks and a 2-level 2× upsampling unit to restore the image to the size of H0×W0×3, and the image restoration is completed with high-fidelity pixels.
[0098] Specific Embodiment 4: The difference between this embodiment and one of Specific Embodiments 1 to 3 is that the working process of the edge-conditioned expectation attention layer is as follows:
[0099] Taking the edge-conditioned expectation attention layer in the second coding unit as an example
[0100] According to and Get the query Q, key K, value V, and edge
[0101]
[0102] Among them, Represents the output of a layer immediately preceding the edge conditional expectation attention layer within the second coding unit, W q , W k , W v and W E represent learnable transformation matrices, R represents real numbers, and N×C represents the dimensions of the query Q, key K, value V, and edge E;
[0103] Then the output A(q i , K, V|e j ) is:
[0104]
[0105] where q i represents the i-th row of the query Q, k j represents the j-th row of the key K, k l represents the l-th row of the key K, v j represents the j-th row of the value V, e j represents the j-th row of the edge , and the superscript T represents transpose, and C represents the channel dimension of the damaged edge image E after passing through the first PE layer.
[0106] Other steps and parameters are the same as those in any one of the specific embodiments one to three.
[0107] Three-channel image After passing through the PE layer, it is divided into H×W image patches. Assuming N = H×W, the original edge image is reshaped into C represents the new channel dimension after passing through the PE layer. In the standard attention mechanism, is the feature map used for self-attention calculation. Through the transformation matrix, the learnable query key and value Let q i , k i , v i ∈R 1×C respectively represent each row of the query Q, key K, and value V. Then the standard attention mechanism can be expressed as:
[0108]
[0109] where, represents taking the i-th image patch as the query and performing a dot product operation with the key. This means that when an image patch is used as the query, if it is more similar to or has common features with other image patches in the global image, then the dot product between the vector of this image patch and the vectors of the corresponding similar image patches will be larger. However, each q iThe attention result can be expressed as the expectation of all v j , which exhibits spatial smoothness characteristics and reduces the ability to effectively capture high-frequency signals.
[0110] Edges, as high-frequency signals in satellite images, have long-range dependencies. To improve the ability to capture high-frequency edge signals, the edge-conditioned expectation attention proposed in the present invention utilizes the conditional probability distribution P(q i , k j |e j ) to adjust the attention weights according to the similarity of edge features between adjacent image patches. This conditional distribution is more inclined to image patches that are more similar to the edge structure of the current image patch, thereby maintaining the structural integrity of the image during the image inpainting process. The present invention adjusts the attention mechanism by introducing a conditional term , where is the transpose of the edge feature of a single image patch obtained after passing the edge information extracted by the CSHE sampler through the PE, that is This enhances the consistency of edge features between adjacent image patches. This helps to better retain the original structural edges when repairing damaged areas. Assuming that an image patch as a query is more similar to other image patches in terms of edge features, then, as Figure 3 shown, the dot product between the image patch vector and the corresponding similar image patch vector will be larger.
[0111] The attention mechanism of the present invention can filter out image patches with low correlation with edge information and pay more attention to the key structures of the image. As Figure 4 shown, it not only retains the pairwise interactions between input tokens in , but also maintains permutation equivariance (C. Yun, S. Bhojanapalli, A. S. Rawat, S. Reddi, and S. Kumar, “Are transformers universal approximators of sequence-to-sequence functions?,” in International Conference on Learning Representations, 2020.). This ensures that even after incorporating edge information, the ECEA-based model can still maintain the context understanding of all image patches in the image and maintain the sensitivity to global information. At the same time, the dynamically learnable query q i will gradually learn more parameters that are beneficial to edge repair during the training process. The improved conditional expectation of the present invention tends to emphasize edge information, which means restoring clear landform features.
[0112] Specific Embodiment 5: Different from one of Specific Embodiments 1 to 4, the convolution kernel sizes of the first convolution layer, the second convolution layer, and the third convolution layer are all 1×1, and the convolution kernel size of the fourth convolution layer is 9×9.
[0113] Other steps and parameters are the same as those in one of Specific Embodiments 1 to 4.
[0114] Specific Embodiment 6: Different from one of Specific Embodiments 1 to 5, both the first residual unit and the second residual unit include 8 residual blocks connected in series.
[0115] Other steps and parameters are the same as those in one of Specific Embodiments 1 to 5.
[0116] Specific Embodiment 7: Different from one of Specific Embodiments 1 to 6, the first downsampling unit is a convolution layer with a stride of 2 and a convolution kernel size of 3×3, and the second to sixth downsampling units are the same as the first downsampling unit;
[0117] The first upsampling unit includes a nearest neighbor interpolation layer and a convolution layer with a stride of 1 and a convolution kernel size of 3×3 (in the upsampling unit, the input first passes through the nearest neighbor interpolation layer and then through the convolution layer), and the second to sixth upsampling units are the same as the first upsampling unit.
[0118] Other steps and parameters are the same as those in one of Specific Embodiments 1 to 6.
[0119] After passing through the i-th downsampling unit, the shape of the edge image is Each upsampling unit uses a nearest neighbor interpolation layer and a convolution layer with a stride of 1 and a convolution kernel size of 3×3. After passing through the i-th upsampling unit, the shape of the edge image is It should be noted that before sampling, the edge map needs to be reshaped to adapt to the convolution operation, and after sampling, it needs to be reshaped again for the Transformer module calculation to integrate the results of each sampling into the ECEA.
[0120] Specific Embodiment 8: Combine Figure 6 to illustrate this embodiment. Different from one of Specific Embodiments 1 to 7, the working process of the Chebyshev expectation adjustment layer is as follows:
[0121]
[0122] Among them, represents the output of the edge conditional expectation attention layer, k is a hyperparameter, represents the output of the Chebyshev expectation adjustment layer.
[0123] The other steps and parameters are the same as those in any one of Embodiments 1 to 7.
[0124] The Transformer block based on the marginal conditional expectation attention layer can effectively enhance the edge information. However, as shown in the boxed area in Figure 4 , it may cause over-enhancement in the edge area or over-smoothing in the non-edge area. Therefore, the present invention introduces the Chebyshev inequality into the attention mechanism to control the possible deviation in the output value. The Chebyshev inequality provides a probability bound to ensure that the deviation of the output from the input v of the current marginal conditional expectation attention layer j is controlled within a certain range of the expected value. Specifically, for each output , the following inequality is applied:
[0125]
[0126] where This constraint ensures that the deviation of the value from the expected value does not exceed , thereby maintaining a balance between edge enhancement and non-edge area smoothing. In actual operation, the output of the previous Transformer block will be used as the input of the next Transformer block. Inside the next Transformer block, the output of the layer immediately before the marginal conditional expectation attention layer will undergo a new linear transformation to obtain the updated v j . The updated v j is used as the input of the marginal conditional expectation attention layer inside the next Transformer block, as shown in Figure 6 . After the attention calculation, is subsequently adjusted using the method of this embodiment.
[0127] By adjusting the control output, excessive deviation in the edge and non-edge areas can be prevented. As shown in Figure 7 is the attention expectation map obtained from a satellite image with a size of 3×256×256. After Patch Embedding, the shape becomes 256×768. Calculate the difference E diff [v]=E current [v]-E adjust [v] (k = 3 in the present invention). For the image patches with significant differences, kernel density estimation is performed along the Token dimension. Patch 10 (marked with a red circle) does not reach the adjustment limit in terms of attention expectation, and v j and E current [v jThe difference between them is extremely small. In contrast, Patch151 (marked with a yellow circle) and Patch 231 (marked with a green circle) have obvious differences. After adjustment, E adjust [v j is closer to v j . This adjustment alleviates the sharp deviation caused by excessive edge enhancement and non-edge smoothing, making the distribution more balanced.
[0128] Specific Embodiment Nine: The difference between this embodiment and any one of Embodiments One to Eight is that the loss function is specifically:
[0129] L = λ a L adv + λpL per + λ s L sty + λ r L re + λ e L edge
[0130] Among them, L is the loss function, L adv represents the adversarial loss, L per represents the perceptual loss, L sty represents the style loss, L re represents the reconstruction loss, L edge represents the edge loss, λ a , λ p , λ s , λ r and λ e are all weights;
[0131] The adversarial loss is:
[0132]
[0133] Among them, I gt represents the real undamaged satellite image, I com represents the restored satellite image, D represents the PatchGAN discriminator implemented using spectral normalization, and represent taking the expectation;
[0134] In the restoration network, if the shape of a feature map is H i × W i × C i , the perceptual loss is:
[0135]
[0136] Among them, represents I gtThe output of the activation function layer of the $i$-th layer after inputting the pre-trained VGG-19 network on the ImageNet dataset; Denote $I$ com The output of the activation function layer of the $i$-th layer after inputting the pre-trained VGG-19 network on the ImageNet dataset; $N$ i $=H$ i $\times W$ i $\times C$ i $, N$ i represents the total number of elements in the feature map of the $i$-th layer in the VGG-19 network; $\|\cdot\|_1$ represents the L1 norm;
[0137] The style loss is:
[0138]
[0139] where denotes the Gram matrix constructed by and denotes the Gram matrix constructed by ;
[0140] The reconstruction loss and the edge loss are respectively:
[0141] $L$ re $=\|\ I$ gt $-I$ com $\|_1$
[0142] $L$ edge $=\|\ I$ ge $-I$ ce $\|_1$
[0143] where $I$ ge represents the true undamaged edge image, and $I$ ce represents the repaired edge image.
[0144] The reconstruction loss can measure the L1 distance between the true undamaged satellite image $I$ gt and the repaired satellite image $I$ com ; the edge loss can measure the L1 distance between the true undamaged edge image $I$ ge and the repaired satellite edge image $I$ ce . During the upsampling process of the CSHE sampler, only the $L$ edge loss is used to generate finer edges without using generative adversarial networks (GANs). This method is used to minimize the error propagation from multi-stage GANs and improve the training efficiency, so that the edge information can be more focused on supporting the ECEA-based model.
[0145] Other steps and parameters are the same as those in any one of the specific embodiments one to eight.
[0146] Embodiment Ten: Different from any one of Embodiments One to Nine, the weight λ a = 1, λ p = 0.5, λ s = 1000, λ r = 100, λ e = 80.
[0147] Other steps and parameters are the same as any one of Embodiments One to Nine.
[0148] These weight values are selected through experiments with the aim of balancing the effects of different loss functions, ensuring that each component can be appropriately weighted during training to achieve a more balanced image restoration effect.
[0149] Experimental Section
[0150] A. Dataset
[0151] To evaluate the performance of the model proposed by the method of the present invention, we used two datasets: the LANDSAT / LE07 / C02 / T1_TOA dataset (N. Gorelick, M. Hancher, M. Dixon, S. Ilyushchenko, D. Thau, and R. Moore, “Google earth engine: Planetary-scale geospatial analysis for everyone,” Remote sensing of Environment, vol. 202, pp. 18–27, 2017.) and the SpaceNet / AOI_4_shanghai dataset (A. Van Etten, D. Lindenbaum, and T. M. Bacastow, “Spacenet: A remote sensing dataset and challenge series,” arXiv preprint arXiv:1807.01232, 2018.). The Landsat-7 dataset contains 100 satellite images, each of which is centrally cropped to a size of 5120×5120 pixels, covering the non-polar regions between 60° south latitude and 60° north latitude. These images were acquired before May 31, 2003, with a spatial resolution of 30 meters and contain eight spectral bands. We focused on bands B1, B2, and B3, mapped the pixel values to between 0 and 0.4, and applied a gamma correction of 1.4. To ensure a cleaner ground truth image, only images with a cloud cover of less than 1% were selected. Each image was segmented into 400 image patches of 256×256 pixels, covering an area of approximately 59 square kilometers. The SpaceNet dataset contains 4582 satellite images of the Shanghai area, with a spatial resolution of approximately 0.26 - 0.3 meters. For this dataset, we selected bands 5, 3, and 2 and resampled each image to a size of 256×256 pixels. Both datasets were subjected to 5-fold cross-validation, dividing the training set and the test set in an 80%-20% ratio. The Landsat-7 dataset used the QA_RADSAT missing pixel layer as a mask, while the SpaceNet dataset was evaluated using an irregular mask (G. Liu, F. A. Reda, K. J. Shih, T.-C. Wang, A. Tao, and B. Catanzaro, “Image inpainting for irregular holes using partial convolutions,” in The European Conference on Computer Vision (ECCV), 2018.). As Figure 8The figure shows the result of repairing the stripe loss of Landsat-7 images, as Figure 9 The figure shows the result of repairing the irregular missing parts in SpaceNet images.
[0152] B. Implementation Details
[0153] The image inpainting network of the present invention is developed based on PyTorch and runs on 2 NVIDIA RTX 3090 (24GB) graphics cards, with the Batch size set to 6. The first to fourth encoding units contain 1, 2, 3, and 4 encoders of the Transformer model respectively, and the first to fourth decoding units contain 3, 2, and 1 encoders of the Transformer model respectively. The dropout ratio in the linear mapping layers of the ECEA and Transformer modules is set to 0.2, and layer normalization (LayerNorm, LN) is applied to normalize the feature maps. During training, the generator uses the Adam optimizer with the learning rate set to 10 -4 , and the learning rate of the discriminator is set to 10 -5 . In the CSHE sampler, the thresholds of the Canny algorithm are set to 50 and 120, and the Sobel operator extracts edges from the horizontal, vertical, and 45° diagonal directions.
[0154] C. Comparison Methods and Evaluation Metrics
[0155] The model of the present invention will be compared with advanced models in the field of image inpainting in recent years as follows:
[0156] GC: An image inpainting model based on CNNs, which introduces a gating mechanism to guide the model to pay more attention to the repair of damaged areas of the image (J. Yu, Z. Lin, J. Yang, X. Shen, X. Lu, and T. S. Huang, “Freeform image inpainting with gated convolution,” in Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 4471–4480, 2019.).
[0157] EC: A two-stage image inpainting model based on CNNs. In the first stage, it preferentially restores the image edges through adversarial generation, and in the second stage, it completes the inpainting of the overall image based on the restored edges (K. Nazeri, E. Ng, T. Joseph, F. Z. Qureshi, and M. Ebrahimi, “Edgeconnect: Generative image inpainting with adversarial edge learning,” arXiv preprint, 2019.).
[0158] ICT: An image inpainting method that combines Transformer and CNNs. It uses Transformer to reconstruct the key texture structures of the image during downsampling and uses CNNs to supplement texture details during upsampling (Z. Wan, J. Zhang, D. Chen, and J. Liao, “High-fidelity pluralistic image completion with transformers,” in Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 4692–4701, 2021.).
[0159] MAT: A Transformer-based image inpainting method that introduces a binary mask layer in the attention calculation process to weight the valid tokens, thereby reducing the impact of invalid tokens on the reconstruction of the missing region in the early stage of training (W. Li, Z. Lin, K. Zhou, L. Qi, Y. Wang, and J. Jia, “Mat: Mask aware transformer for large hole image inpainting,” in Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pp. 10758–10768, 2022.).
[0160] CMT: Based on MAT, it incorporates a mask update mechanism to reduce the possibility of tokens being misclassified as the network depth increases (K. Ko and C.-S. Kim, “Continuously masked transformer for image inpainting,” in Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 13169–13178, 2023.).
[0161] GC and EC are CNNs-based methods, while ICT is a method for feature extraction based on the Transformer architecture. MAT and CMT respectively use their adjusted Transformer modules as the basic network units. We modified the input tensor size of MAT and CMT to 256×256.
[0162] In addition, our study adopted evaluation metrics widely used in the field of image inpainting to evaluate the comparative models, mainly including Peak Signal To-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Frechet Inception Distance (FID). PSNR and SSIM are standard metrics for evaluating image similarity. FID comprehensively evaluates image quality by comparing the features of the inpainted image set and the real image set in the feature space extracted by the pre-trained Inception network.
[0163] D. Qualitative Comparison
[0164] On the Landsat-7 and SpaceNet datasets, the model proposed in the present invention was compared with the GC, EC, ICT, MAT, and CMT models. Three evaluation metrics were used for the comparison: PSNR, SSIM, and FID.
[0165] Figure 8 and Figure 9 Shows the inpainting results of the model of the present invention under different mask coverages. Even in severely damaged cases, the model of the present invention can successfully learn a reasonable feature distribution from the dataset and provide coherent content for the missing regions, effectively restoring the main structure of the satellite image.
[0166] Figure 10 and Figure 11The repair results of the model of the present invention and GC, EC, ICT, MAT, and CMT are respectively shown under strip-shaped and irregular missing regions. The GC model can complete the repair of missing regions in satellite images by utilizing context semantics. However, when the missing regions are large and isolated, it is difficult for the GC model to capture long-distance dependencies and predict reasonable content. For example, in Figure 11 the first row, the GC method distorts the texture structure within the region highlighted by the red box. The EC method can restore the key structure of the image, but for irregular losses with large areas of missing such as the irregular masks used in the SpaceNet dataset, EC cannot effectively predict the main structure. In contrast, the strip-shaped missing regions of Landsat-7 are easier for EC to handle. EC successfully reconstructs Figure 10 the river landform in the first row, but in Figure 11 it fails to reconnect the obvious road features. The ICT model can capture global features and generate reasonable textures and structures. However, when the missing regions lack reliable surrounding guidance, the artifacts in the disappearing regions tend to be minimized but not eliminated. As shown in Figure 10 the second row, the noise artifacts caused by the mask are still visible.
[0167] MAT and CMT are the latest advanced image inpainting methods, which can learn high-frequency details in the early stage of training and show strong fine texture reconstruction capabilities. However, as highlighted by the red boxes in Figure 10 and Figure 11 the first row, the roads covered by the missing regions are still not fully repaired, while the model proposed by the present invention can smoothly complete and reconnect these features. The ECEA-based model of the present invention shows satisfactory qualitative results in satellite image inpainting and can smoothly and effectively connect key features.
[0168] E. Quantitative comparison
[0169] As shown in Table 1, for the Landsat-7 dataset, within the range of mask ratios of 0 - 10% and 10 - 20%, the GC model performs excellently in all three evaluation metrics, followed by the CMT model and the model proposed by the present invention. In contrast, for the SpaceNet dataset with more complex semantic features and mask patterns, the MAT and CMT models perform well when the degree of image loss is small. However, when the damage degree of the satellite image exceeds 20%, the ECEA-based model of the present invention gradually outperforms other methods on both datasets, highlighting its effectiveness in dealing with more severe image damage.
[0170] Table 1: Comparison of the model of the present invention with other models in terms of PSNR, SSIM, and FID
[0171]
[0172] The higher the metric values of PSNR and SSIM, the better, and the lower the metric value of FID, the better. The metric values of the method of the present invention are shown in bold, and the best values under the current metrics are underlined.
[0173] As shown in FIGS. 12(a) and 12(b), box plots of PSNR and SSIM were used to conduct a comparative analysis of the performance of various models under different mask ratios on the Landsat-7 dataset. When the mask ratio was 0-10% and 10-20%, the GC model was always superior to other models, with higher medians of PSNR and SSIM. However, when the mask ratio exceeded 20%, the model proposed by the present invention performed better, especially when the mask ratio was between 30-50%. The CMT model followed closely in terms of performance. But the model of the present invention achieved a higher median and more stable results, manifested as a narrower interquartile range, especially in the range of higher mask ratios. This indicates that the model proposed by the present invention has stronger robustness in dealing with more severe image degradation. In addition, under higher mask ratio conditions, the model of the present invention has fewer outliers than other methods, further demonstrating its stability in these scenarios.
[0174] F. Ablation Study
[0175] 1) Performance of the ECEA mechanism: Verify the effectiveness of the model proposed by the present invention on the SpaceNet dataset. In the ablation experiment, we mainly explored two parts:
[0176] (1) The role of the conditional expectation in attention calculation.
[0177] Its main part is denoted as ξ. Removing ξ is equivalent to the ordinary attention mechanism;
[0178] (2) The role of using the Chebyshev inequality for expectation adjustment in the attention mechanism, denoted as
[0179] Table 2 Comparison of the model of the present invention and ablation variables in terms of PSNR, SSIM, and FID on the SpaceNet dataset
[0180]
[0181] "w / o ξ" means without considering ξ, "w / o " means without considering "w / o " means without considering As can be clearly seen from Table 2, removing the edge feature ξ will severely weaken the model's ability to capture key structures, resulting in a decrease in the PSNR and SSIM values and an increase in the FID score, indicating a deterioration in the image inpainting performance. Removing the Chebyshev-based expectation adjustment will also lead to a performance decline. Removing both ξ and will result in the worst performance, with the lowest PSNR and SSIM values and the highest FID value. In contrast, after incorporating these two features, the entire model achieves the best results in all metrics, indicating that it is effective in dealing with more complex image degradation problems.
[0182] 2) Performance based on Chebyshev adjustment: The value of the hyperparameter k affects the upper and lower limits of the Chebyshev adjustment. Figure 13 The inpainting results for k ∈ [0.5, 1, 3, 5] are shown. When k = 0.5, it is easy to reach the expectation adjustment limit which leads to significant modification of the feature maps. As shown in Figure 13 the figure corresponding to k = 0.5 in row (b), the overly frequent adjustment makes the outputs of multiple Transformer modules change minimally during the training process, making it difficult to learn parameters beneficial for image inpainting. When k = 5, E diff [v] graph is almost empty, and rarely reaches the Chebyshev adjustment limit, basically equivalent to abandoning the Chebyshev expectation adjustment, which is comparable to the case without in Table 2. It can also be seen from the inpainting results of k = 5 that the filling effect of the edge line of the roof is not good. When k = 1, the problems of over-enhanced edges and over-smoothed non-edge regions are not fully alleviated. The Chebyshev expectation adjustment imposes feature constraints on the attention expectation and variance of each layer from a probability perspective to prevent the feature distribution shift caused by edge enhancement. In addition, image features do not necessarily follow a Gaussian distribution or other specific distributions. In the present invention, better inpainting results are obtained when k = 3, and in practical applications, the value of k should be flexibly adjusted according to the dataset.
[0183] 3) Performance of CSHE: Next, the effects of no edge detection, Canny, Sobel, and CSHE on the image inpainting performance of the SpaceNet and Landsat-7 datasets are studied. When no edge detection is used, the model cannot effectively capture the main structural features of the image, resulting in relatively poor image inpainting performance, as shown in Figure 14As shown in the first row. This result is quantitatively equivalent to the case of "w / oξ" in Table 2. The results produced by the other three methods are similar but slightly different. Among them, the CSHE method proposed in the present invention makes up for the deficiency of the uniform edges generated by Canny. It suppresses the noisy edges generated by Sobel and can provide better detail restoration results than the other two methods, as Figure 14 shown by the red box in
[0184] G. Computational efficiency
[0185] The present invention evaluates the computational efficiency of different models in terms of multiply-accumulate operations (MACs), the number of trainable parameters, and the inference speed. The number of parameters refers to the trainable weights of the generator (excluding the discriminator). For models with multiple generator stages, the weights of all stages are added together. The inference speed is measured in frames per second (FPS), and the test environment is an NVIDIA RTX 3090 GPU (24GB) with a batch size of 1 and an input size of 3×256×256. As shown in Table 3, since the model of the present invention adopts an independent processing path for edge upsampling and downsampling and the increased complexity based on Chebyshev expectation adjustment, it shows relatively higher MACs and more parameters compared with other models. Although this results in a moderately reduced inference speed (39.3 FPS), in tasks that require more accurate terrain feature reconstruction, this model performs better, making this trade-off worthwhile.
[0186] Table 3 Complexity metrics of different models
[0187]
[0188] The above examples of the present invention are only for illustrating in detail the computational model and computational process of the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is impossible to list all the implementation manners here. Any obvious changes or variations derived from the technical solutions of the present invention still fall within the protection scope of the present invention.
Claims
1. A satellite image restoration method based on edge condition expected attention, characterized in that: The method specifically comprises the following steps: Step 1: Obtain a real undamaged satellite image dataset And the loss mask matrix dataset Mask Matrix Dataset The mask matrix in contains only elements 0 and 1, where 1 represents that the pixel at the corresponding position in the satellite image is damaged, and 0 represents that the pixel at the corresponding position in the satellite image is not damaged; according to and Get the damaged satellite imagery dataset Dataset The way to obtain any damaged satellite image X in is: Among them, X G1 Representation dataset Any image in M G1 Representation dataset Any mask matrix in Represents the matrix X G1 and 1-M G1 Multiply the elements at corresponding positions; For the dataset Each satellite image in the image is subjected to edge extraction to obtain the true undamaged edge image set ε G , for the dataset Each satellite image in is subjected to edge extraction to obtain the damaged edge image set ε I , using the dataset ε G and ε I Composing a training data set; Step 2: Construct a generative adversarial network for satellite image restoration. The generative adversarial network includes a generator and a discriminator. The working process of the generator is as follows: For the dataset Any damaged satellite image in H0 and W0 represent the height and width of image X, and X is included in the dataset ε I The corresponding damaged edge image is recorded as The damaged edge image E is reshaped after the first PE layer as Then After the first downsampling unit, the output of the first downsampling unit is recorded as Will As the input of the second downsampling unit, the output of the second downsampling unit is recorded as Will As the input of the third downsampling unit, the output of the third downsampling unit is recorded as Will As the input of the first residual unit, the output of the first residual unit is used as the input of the first upsampling unit, and the output of the first upsampling unit is recorded as Will As the input of the second upsampling unit, the output of the second upsampling unit is recorded as Will As the input of the third upsampling unit, the output of the third upsampling unit is recorded as Damaged satellite imagery As the input of the second PE layer, it is reshaped into H×W represents the number of image blocks into which the image X is divided. as input to the first encoding unit; The output of the first encoding unit is used as the input of the fourth down-sampling unit, and the output of the fourth down-sampling unit is used as the input of the fourth down-sampling unit. and the output of the first downsampling unit as input to a second encoding unit; The output of the second encoding unit is used as the input of the fifth down-sampling unit, and the output of the fifth down-sampling unit is used as the input of the fifth down-sampling unit. and the output of the second downsampling unit as input to a third encoding unit; The output of the third encoding unit is used as the input of the sixth down-sampling unit, and the output of the sixth down-sampling unit is used as the input of the sixth down-sampling unit. And the output of the third downsampling unit is recorded as as input to a fourth encoding unit; The output of the fourth encoding unit is used as the input of the fourth up-sampling unit, and the output of the fourth up-sampling unit is used as the input of the fourth up-sampling unit. The output of the third encoding unit is concatenated to obtain the concatenated result a, and the concatenated result a is passed through the first convolution layer. The output of the first convolution layer and the output of the first upsampling unit are concatenated. as input to the first decoding unit; The output of the first decoding unit is used as the input of the fifth upsampling unit, and the output of the fifth upsampling unit is used as the input of the fifth upsampling unit. The output of the second encoding unit is concatenated to obtain a concatenated result b, and the concatenated result b is passed through the second convolutional layer; Combine the output of the second convolutional layer with the output of the second upsampling unit as an input of the second decoding unit, and using the output of the second decoding unit as an input of the sixth up-sampling unit; The output of the sixth upsampling unit The output of the first encoding unit is concatenated to obtain a concatenated result c, and the concatenated result c is passed through a third convolutional layer; Combine the output of the third convolutional layer with the output of the third upsampling unit As the input of the third decoding unit, the output of the third decoding unit passes through the fourth convolutional layer, the second residual unit, the seventh upsampling unit and the eighth upsampling unit in sequence, and the output of the eighth upsampling unit is used as the output of the generator; Step 3: Use the acquired training data set to train the constructed generative adversarial network until the loss function converges, and then stop the training to obtain a trained generative adversarial network. Step 4: Input the damaged satellite image to be restored into the generator in the trained generative adversarial network, and use the generator to output the restored image.
2. According to claim 1, a satellite image restoration method based on edge condition expected attention is characterized in that: The pair of data sets Each satellite image in the image is edge extracted to obtain the real edge image set ε G , the specific process is: For the dataset For any undamaged satellite image in the image, the Canny edge detection algorithm is used to extract the edge of the undamaged satellite image to obtain the first edge image, and then the Sobel operator is used to extract the edge of the undamaged satellite image to obtain the second edge image; The first edge image and the second edge image are averaged, and the average result is used as the edge image corresponding to the undamaged satellite image.
3. According to claim 2, a satellite image restoration method based on edge condition expected attention is characterized in that: The first encoding unit includes 1 Transformer block, the second encoding unit includes 2 Transformer blocks connected in series, the third encoding unit includes 3 Transformer blocks connected in series, the fourth encoding unit includes 4 Transformer blocks connected in series, the first decoding unit includes 1 Transformer block, the second decoding unit includes 2 Transformer blocks connected in series, and the third decoding unit includes 3 Transformer blocks connected in series; wherein: The attention layers in each Transformer block of the second encoding unit, the third encoding unit, the fourth encoding unit, the first decoding unit, the second decoding unit and the third decoding unit are edge conditional expectation attention layers, and a Chebyshev expectation adjustment layer is connected after the edge conditional expectation attention layer.
4. According to claim 3, a satellite image restoration method based on edge condition expected attention is characterized in that: The working process of the edge conditional expectation attention layer is: Take the edge conditional expectation attention layer in the second encoding unit as an example according to and Get query Q, key W, value V and edge in, represents the output of the layer immediately before the edge-conditional desired attention layer in the second encoding unit, W q , W k , W v and W E represents the transformation matrix; Then the edge condition expects the output of the attention layer A(q i ,K,V|e j )for: Among them, q i represents the i-th row of query Q, k j represents the jth row of key K, k l represents the lth row of key K, v j represents the jth row of value V, e j Representing the edge In the j-th row of , the superscript T represents transpose, and C represents the channel dimension of the damaged edge image E after passing through the first PE layer.
5. According to claim 4, a satellite image restoration method based on edge condition expected attention is characterized in that: The convolution kernel sizes of the first convolution layer, the second convolution layer, and the third convolution layer are all 1×1, and the convolution kernel size of the fourth convolution layer is 9×9.
6. The satellite image restoration method based on edge condition expected attention according to claim 5 is characterized in that: The first residual unit and the second residual unit each include 8 residual blocks connected in series.
7. The satellite image restoration method based on edge condition expected attention according to claim 6 is characterized in that: The first downsampling unit is a convolution layer with a step size of 2 and a convolution kernel size of 3 3, and the second downsampling unit to the sixth downsampling unit are the same as the first downsampling unit; The first upsampling unit includes a nearest neighbor interpolation layer and a convolution layer with a step size of 1 and a convolution kernel size of 3 3. The second upsampling unit to the sixth upsampling unit are the same as the first upsampling unit.
8. The satellite image restoration method based on edge condition expected attention according to claim 7 is characterized in that: The working process of the Chebyshev expectation adjustment layer is: in, represents the output of the edge conditional expected attention layer, k is a hyperparameter, Represents the output of the Chebyshev expectation adjustment layer.
9. The satellite image restoration method based on edge condition expected attention according to claim 8 is characterized in that: The loss function is specifically: L=λ a L adv +λ p L per +λ s L sty +λ r L re +λ e L edge Among them, L is the loss function, L adv represents the adversarial loss, L per represents the perceptual loss, L sty represents the style loss, L re represents the reconstruction loss, L edge represents the marginal loss, λ a , p , s , r and λ e All are weights; The adversarial loss is: Among them, I gt represents the real undamaged satellite image, I com represents the restored satellite image, D represents the PatchGAN discriminator, and Expressing hope; The perceptual loss is: in, Indicates that I gt The output of the activation function layer of the i-th layer after inputting the VGG-19 network pre-trained on the ImageNet dataset; Indicates that I com The output of the activation function layer of the i-th layer after inputting the VGG-19 network pre-trained on the ImageNet dataset; N i represents the total number of elements in the feature map of the i-th layer in the VGG-19 network; ||·||1 represents the L1 norm; The style loss is: in, Indicated by The constructed Gram matrix is Indicated by Constructed Gram matrix; The reconstruction loss and edge loss are: L re =||I gt -I com ||1 L edge =||I ge -I ce ||1 Among them, I ge represents the true undamaged edge image, I ce Represents the inpainted edge image.
10. The satellite image restoration method based on edge condition expected attention according to claim 9 is characterized in that: described a =1,λ p =0.5,λ s =1000, l r =100,λ e = 80.