Synthetic Aperture Optical Image Restoration Method Based on Local-Global Feature Enhancement
By constructing a synthetic aperture optical image restoration network with local-global features, using the LAG-Transformer layer and the GRM-Convolution layer, the imaging blur problem of synthetic aperture optical system is solved, and efficient image restoration effect is achieved, adapting to different data distributions, and reducing gradient explosion phenomenon.
Patent Information
- Application Number
- CN202411697443.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-11-26
AI Technical Summary
合成孔径光学系统成像模糊问题导致成像结果降质,传统方法难以有效复原高分辨率图像的真实结构信息,且易受振铃现象影响。
The synthetic aperture optical image restoration method based on local-global feature enhancement is adopted, and the global and local information attention is enhanced by using the LAG-Transformer layer and the GRM-Convolution layer. Combined with the adaptive scale feature enhancement module, an encoded and decoding network is built for image restoration.
It significantly improves the recovery effect of high-resolution optical image of synthetic aperture, improves the ability to extract high-level semantic information, reduces gradient explosion, enhances the generalization ability of the network, and adapts to different data distributions.
Smart Images

Figure CN119205581B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a synthetic aperture optical image restoration method based on local-global feature enhancement, and belongs to the technical field of synthetic aperture optical system image restoration. Background Art
[0002] With the progress of technology, the human requirements for the resolution of optical imaging systems such as atmospheric observation and environmental monitoring are increasing day by day. According to the Rayleigh criterion, when the working wavelength of the optical system is fixed, the system aperture is positively correlated with its resolution. However, due to limitations such as processing technology and cost, it is difficult to increase the aperture without limit. Therefore, the optical synthetic aperture imaging (OSAI) technology is proposed, which uses multiple sub-apertures to form an array to synthesize a large-aperture system to achieve high-resolution imaging. Since several sub-apertures jointly fill a single large aperture, the total light-transmitting area of the entire system decreases, and there are inevitably translational errors between the apertures, resulting in changes in the point spread function (PSF) of the system and a decrease in the middle and low frequencies of the modulation transfer function (MTF), leading to degraded and blurred imaging results. In severe cases, imaging may even be impossible.
[0003] In order to solve the problem of imaging blur in synthetic aperture optical systems, traditional methods use image restoration techniques such as Wiener filtering, blind deconvolution method and maximum likelihood method to compensate for OSAI. The development of deep learning has greatly promoted the improvement of image restoration technology. For traditional blurred images, Tao et al. (X. Tao, H. Gao, J. Wang, et al., "Scale-Recurrent Network for Deep Image Deblurring," 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition, 8174–8182 (2018)), Zamir et al. (SW Zamir, A. Arora, L. Khan, et al., "Multi-Stage Progressive Image Restoration," 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14816–14826 (2021)), Mou et al. (C. Mou, Q. Wang, and J. Zhang, "Deep Generalized Unfolding Networks for Image Restoration," 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17378–17389 (2022) proposed the use of a multi-scale network, which is no longer limited by the blur kernel, and cross-fused features at different stages, optimizing the transmission of information between different scales. Hui et al. (M. Hui, Y. Wu, Y. Li, et al., “Image restoration for synthetic aperture systems with a non-blind deconvolution algorithm via a deep convolutional neural network,” Opt. Express 28(7), 9929–9943 (2020)) further combined the variational physical model and prior-based methods with deep learning, achieving better results, but still inseparable from the design of traditional methods. Unlike motion blur and jitter blur in traditional images, the imaging blur problem of synthetic aperture optical systems presents a specific situation described by the point spread function, and the blur shape is not uniformly distributed.In addition, affected by the sub-aperture phase arrangement, the imaging of synthetic aperture optical systems is prone to ringing phenomena. When using traditional methods for restoration, it is susceptible to the influence of ringing phenomena and generates artifacts, resulting in poor restoration effects. Therefore, how to effectively focus on the real structural information of high-resolution images and restore them has become a key issue in the restoration of synthetic aperture optical high-resolution images. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a synthetic aperture optical image restoration method based on local-global feature enhancement, which better focuses on the real structural information of high-resolution images and significantly improves the restoration effect of synthetic aperture high-resolution optical images.
[0005] The present invention adopts the following technical solutions to solve the above technical problems:
[0006] A synthetic aperture optical image restoration method based on local-global feature enhancement includes the following steps:
[0007] Step 1, obtain a synthetic aperture optical image dataset and perform simulation degradation to obtain a simulated degraded image dataset, and the images in the synthetic aperture optical image dataset and the simulated degraded image dataset correspond one by one. The corresponding image dataset is randomly divided into a training set, a validation set, and a test set;
[0008] Step 2, construct a synthetic aperture optical image restoration network, including an encoding network and a decoding network. The encoding network includes a first module to a fifth module connected in sequence. An input mapping module is connected between the first module and the original input image. The decoding network includes a sixth to tenth module connected in sequence. The input end of the sixth module is connected to the output end of the fifth module, and the output end of the tenth module is connected to a second gated residual convolutional layer. The output of the gated residual convolutional layer is used as the final output of the synthetic aperture optical image restoration network;
[0009] The structures of the first module to the tenth module are the same, and each includes a linear gated Transformer layer, a first gated residual convolutional layer, and an adaptive scale feature enhancement module connected in sequence. The output of the input mapping module and the output of the second gated residual convolutional layer are skip-connected. The output of the adaptive scale feature enhancement module in the first module and the output of the adaptive scale feature enhancement module in the tenth module are skip-connected. The output of the adaptive scale feature enhancement module in the second module and the output of the adaptive scale feature enhancement module in the ninth module are skip-connected. The output of the adaptive scale feature enhancement module in the third module and the output of the adaptive scale feature enhancement module in the eighth module are skip-connected. The output of the adaptive scale feature enhancement module in the fourth module and the output of the adaptive scale feature enhancement module in the seventh module are skip-connected. The output of the adaptive scale feature enhancement module in the fifth module and the output of the adaptive scale feature enhancement module in the sixth module are skip-connected;
[0010] The input mapping module maps the original input image into an original input feature sequence. The first to fifth modules respectively output feature sequences with lengths of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original input feature sequence, and the sixth to tenth modules respectively output feature sequences with lengths of 1 / 32, 1 / 16, 1 / 8, 1 / 4, and 1 / 2 of the original input feature sequence;
[0011] Step 3: Input the training set and validation set obtained in Step 1 into the synthetic aperture optical image restoration network constructed in Step 2 for training, calculate the loss function and perform backpropagation to update the network parameters, and obtain the best parameter model after training;
[0012] Step 4: Input the test set obtained in Step 1 into the best parameter model trained in Step 3 to output the restored image of the synthetic aperture optical system.
[0013] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0014] 1. The synthetic aperture optical image restoration network based on local-global feature enhancement proposed by the present invention consists of an encoding network and a decoding network. Among them, the encoding and decoding network uses the LAG-Transformer layer to replace the downsampling layer in the traditional U-Net network, and constructs the GRM-Convolution layer and the ASFE layer to enhance the network's attention ability to global information and local information.
[0015] 2. The synthetic aperture optical image restoration network based on local-global feature enhancement proposed by the present invention uses the LAG-Transformer layer to pay attention to the global information of the extracted feature sequence, while reducing the sequence length to extract high-level semantic information. At the same time, the GRM-Convolution layer and the ASFE layer are used to pay attention to the local information of the features, making up for the deficiency of the Transformer in capturing local information. After obtaining the high-level semantic information, the final restored image is obtained through the decoding area.
[0016] 3. The gated mechanism-based residual convolution layer proposed by the present invention uses depthwise separable convolution to perform ordinary extraction operations on feature information, and uses The depthwise separable convolution controls the number of feature channels and uses a simple gating mechanism to divide the features into two parts in the channel dimension, and then multiplies them pixel by pixel, which can effectively replace the GELU activation function, help reduce the occurrence of gradient explosion phenomenon, and has a low computational complexity. At the same time, the deformable convolution is used to perform deformation extraction operations on the input features. The deformable convolution introduces learnable offsets in the receptive field. Compared with traditional convolution, it deforms the fixed-size convolution kernel, and the position of each convolution kernel will be adaptively adjusted according to the learnable offsets, realizing more flexible extraction of local features, enabling GRM-Convolution to better adapt to the local changes of input features, improving the generalization ability of GRM-Convolution, and making it able to adapt to different data distributions while capturing local information of different data features. After two deformable convolution extractions, the features are connected in residual with the features after the gating mechanism, and then a simple gating operation is performed on them to improve the network's ability to express features. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of the synthetic aperture optical image restoration method based on local-global feature enhancement of the present invention;
[0018] Figure 2 is a structural diagram of the synthetic aperture optical image restoration network model constructed by the present invention;
[0019] Figure 3 is a structural diagram of the linear gating Transformer layer (LAG-Transformer) proposed by the present invention;
[0020] Figure 4 is a structural diagram of the gating residual convolution layer (GRM-Convolution) proposed by the present invention;
[0021] Figure 5 is a structural diagram of the adaptive scale feature enhancement module proposed by the present invention;
[0022] Figure 6 is a schematic diagram of the comparison of the synthetic aperture optical system image restoration results between the method proposed by the present invention and other methods. Among them, (a) is the original image, (b) is the degraded image, (c) is WF, (d) is DeblurGan, (e) is Uformer-T, (f) is Uformer-S, (g) is Stripformer, and (h) is the present invention;
[0023] Figure 7Schematic diagram of comparison of degraded images under the influence of different ringing gain factors and restoration results of different methods. Among them, (a) is the degraded image under the influence of different ringing gain factors, and (b) is the comparison of restoration results of different methods and the comparison of local enlarged views under 40dB - 50dB noise;
[0024] Figure 8 Schematic diagram of comparison of degraded images under the influence of different levels of noise and restoration results of different methods. Among them, (a) is the degraded image under the influence of different levels of noise, and (b) is the comparison of restoration results of different methods;
[0025] Figure 9 Comparison diagram of detections of the YOLOv7 method on the unrecovered and recovered images. Among them, (a) is the detection of cars, and (b) is the detection of ships. Specific implementation manners
[0026] The following details the implementation manners of the present invention, and the examples of the implementation manners are shown in the drawings. The implementation manners described below with reference to the drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation of the present invention.
[0027] For the unique blur type and the unique ringing phenomenon of the high-resolution degraded images of synthetic aperture optical systems, traditional deep learning image restoration is not well applicable. The high-resolution image restoration method for synthetic aperture optical systems based on local-global feature enhancement Transformer proposed by the present invention is as Figure 1 shown, and the specific steps are as follows:
[0028] Step 1: Obtain the DOTA remote sensing dataset, select 7000 images with obvious features after segmentation. Use the system point spread function actually measured by the synthetic aperture optical system with an eight-aperture annular array structure to perform a simulation degradation experiment on the dataset to obtain degraded images based on the DOTA remote sensing dataset, and divide them into a training set, a validation set, and a test set according to a set ratio.
[0029] Among them, the DOTA remote sensing dataset contains 2806 ultra-high-resolution remote sensing images, with the highest pixel reaching , and a total of 15 common targets. Select 7000 images with obvious features after segmentation.
[0030] The specific process of simulating the degradation of the dataset is as follows: Measure the parameters of the actual synthetic aperture optical system, obtain the system point spread function, save it in txt format, and then use the pycharm software to read the point spread function and perform convolution with the DOTA remote sensing dataset to obtain the blurred image of simulated degradation.
[0031] The specific process of preprocessing the dataset is as follows: Crop the DOTA remote sensing dataset and the simulated degraded images into Images of the same size are then paired one by one as the input image and the target image.
[0032] Specifically, the dataset is divided as follows: the paired images are randomly divided into three parts, where the training set accounts for 80%, the validation set accounts for 10%, and the test set accounts for 10%.
[0033] Step 2: As Figure 2 shown, build a high-resolution image restoration network for synthetic aperture optical systems based on Transformer with local-global feature augmentation (Transformer Networks with Local-Global Feature Augmentation, LGFAformer), including a linear gating Transformer layer (Transformer Based on Linear Attention and Gating Mechanism, LAG-Transformer), a gating residual convolution layer (Convolutional Layer Based on Gating and Residual Mechanism, GRM-Convolution), and an adaptive scale feature enhancement module (Adaptive Scale Feature Enhancement Module, ASFE).
[0034] Step 2.1: Input the blurred and degraded image into the LGFAformer restoration network. The input image is first mapped to a feature sequence by a convolutional layer and then input into the LAG-Transformer layer to extract the image features from the bottom-level details to the high-level features and obtain feature images of different scales. Specifically: first, the input degraded image is mapped to a feature sequence and then input into the LAG-Transformer layer. After the feature sequence undergoes layer normalization, it is divided into windows and input into the focused linear attention module (FLA). After calculating the focused linear attention, it undergoes layer normalization again. Subsequently, it is input into the forward propagation module based on the gating mechanism FPBG, and the global features extracted are re-extracted using depthwise separable convolution. Then, to reduce the interference of windowing operations on the extraction of global information, the divided windows are shifted, re-stitched into a new window, and input into the FLA and FPBG modules again for secondary extraction of global information. Finally, the output is the feature sequence with a length 1 / 4 of the original after adding with the residual structure.
[0035] Step 2.2: Input the feature sequence with a length of 1 / 4 of the original length obtained in 2.1 into the GRM-Convolution layer. After converting the feature sequence into a feature map, first use deformable convolution to extract feature information with spatial position offsets and then perform layer normalization. Perform deformable convolution again to further enhance the extraction of spatial position information. After performing layer normalization on the original feature map, use the depthwise separable convolution of to control the number of channels, and then use the depthwise separable convolution of to extract its local features. Then use the depthwise separable convolution of to adjust the number of channels and add it to the feature information extracted by deformable convolution. Subsequently, use the depthwise separable convolution of and a simple gating mechanism to fuse the two while ensuring the number of channels and finally output.
[0036] Step 2.3: Input the feature map output in Step 2.2 into the ASFE layer. One way uses deformable convolution to obtain a feature weight map , and the other way uses depthwise separable convolution to obtain a feature weight map . After splicing the feature information extracted from the two paths and readjusting the number of channels, use average pooling operation to obtain a feature weight map . Subsequently, fuse the three different feature weight maps respectively to obtain the final feature weight map Q . Under the control of the weight factor M , use the feature weight map Q to perform secondary learning on the original input feature map to obtain a feature map with a resolution of 1 / 2 of the original input image.
[0037] Step 2.4: Re-convert the feature map with a resolution of 1 / 2 of the original input image obtained in Step 2.3 into a feature sequence, repeat Steps 2.1 - 2.3, and obtain feature sequences with lengths of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original input feature sequence respectively. Finally, input the extracted high-level semantic feature sequence into the decoder.
[0038] Step 2.5: Pass the high-level semantic feature sequence with a length of 1 / 32 of the original input feature sequence obtained in Step 2.1 through the decoding region to re-obtain feature images of different scales. The decoding region and the encoding region together form an encoding-decoding region, and the upsampled feature recovery is gradually completed through the LAG-Transformer layer to realize the restoration of the synthetic aperture high-resolution optical degraded image. Specifically: First, pass the high-level semantic feature sequence with a length of 1 / 32 of the original input feature sequence output by the encoder through the LAG-Transformer layer, and use layer normalization and the FPBG module to increase the length of its feature sequence to make it a high-level semantic feature sequence with a length of 1 / 16 of the original input feature sequence, and then perform a simple skip connection operation with the feature image of the corresponding size in the encoding layer. Subsequently, after inputting the GRM-Convolution layer and the ASFE layer, perform similar operations until finally restoring a generated image with 3 channels and a size of 512×512.
[0039] The LGFAformer network is mainly composed of an encoding network and a decoding network. The encoding-decoding network uses the LAG-Transformer layer to replace the downsampling layer in the traditional U-Net network, and constructs the GRM-Convolution layer and the ASFE module to enhance the network's attention ability to global and local information.
[0040] As Figure 3 shown, the LAG-Transformer layer proposed in the present invention is composed of a forward propagation module based on a gating mechanism (The Forward Propagation Module Based on The Gating Mechanism, FPBG), a focused linear attention module (Focused linear attention module, FLA), and a normalization layer (LayerNorm). Among them, the focused linear attention module FLA is composed of a Relu activation function, a focused linear attention function, and a fully connected layer; the forward propagation module FPBG based on the gating mechanism is composed of a depthwise separable convolution, an effective channel attention module (Effectivechannel attention module, ECA), and a Gelu activation function;
[0041] The aggregation mapping function is expressed as:
[0042] ,
[0043] Among them, represents the similarity function, represent query vectors and key vectors of different dimensions respectively, Relu represents the Relu activation function, represents Element power 。
[0044] The LAG-Transformer layer takes the degraded image as input. After performing layer normalization on the input features, it divides the feature information into multiple windows, calculates the attention for each window respectively, reduces the computational cost on high-resolution images, and establishes long-range dependencies within the windows. To highlight important feature information, multiple lightweight learnable tensors (with size , where W is the window size and C is the feature channel dimension) are set. Each tensor is added as a shared weight to the non-overlapping windows before linear attention processing, achieving a significant performance improvement with a small increase in parameters.
[0045] The Focused Linear Attention module FLA uses a new linear calculation method to replace the Softmax in the original Multi-head Attention (MHA), reducing the computational complexity from to . First, the input tensor is linearly transformed to obtain the query vector Q, key vector K, and value vector V respectively. Activation operations are performed on the query vector and key vector to ensure they are non-zero, and then they are scaled using a learnable factor. To prevent the lack of a sharp attention distribution similar to Softmax in traditional linear attention, an aggregation mapping function is adopted to adjust the directions of the features of each query and key vector, making similar query vector and key vector pairs closer, while pushing dissimilar vector pairs apart to make the difference between them more obvious, simulating the sharp attention distribution in Softmax. After linearly calculating the key-value pairs (Key-Value Pair) using the aggregation mapping function for the key vector and value vector, they are linearly calculated with the query vector, and finally fused with the value vector extracted after depthwise separable convolution. To prevent overfitting, after passing the final result through a linear layer, the Dropout operation is used to set some elements of the tensor to zero.
[0046] The Forward Propagation module FPBG first adjusts the number of channels of the features through a convolution, and then uses two depthwise separable convolutions. One is used to expand the channel features (expansion factor is 2), and the other is used to reduce the channels to the original input dimension. A depthwise separable convolution is used to extract the spatial adjacent pixel position information, which is beneficial for the LAG layer to learn the local image structure to restore the global image. One path is implemented using the GELU activation function, and the other path uses lightweight ECA attention to focus on the relationship between feature channels. After the two branches perform an element-wise product operation, they are then processed using The convolution adjusts the number of channels to help the LAE module obtain more context information. The implementation of the gating mechanism is represented by calculating the element-wise product between two branches.
[0047] As Figure 4 shown, the residual convolutional layer based on the gating mechanism proposed by the present invention mainly consists of deformable convolution, depthwise separable convolution, normalization layer, and a simple gating mechanism. The depthwise separable convolution of is used to perform ordinary extraction operations on feature information, and the depthwise separable convolution of is used to control the number of feature channels. Using a simple gating mechanism, the features are divided into two parts in the channel dimension and then multiplied pixel by pixel, which can effectively replace the GELU activation function, help reduce the occurrence of gradient explosion, and has a lower computational complexity. At the same time, the deformable convolution is used to perform deformation extraction operations on the input features. The deformable convolution introduces learnable offsets in the receptive field. Compared with traditional convolution, it deforms the fixed-size convolution kernel, and the position of each convolution kernel is adaptively adjusted according to the learnable offsets, realizing more flexible extraction of local features, enabling GRM-convolution to better adapt to local changes in input features, improving the generalization ability of GRM-convolution, and making it capable of capturing local information of different data features while also adapting to different data distributions. After two deformable convolution extractions, the features are connected residually with the features after the gating mechanism, and then a simple gating operation is performed on them to improve the network's ability to express features.
[0048] As Figure 5 shown, the adaptive scale feature enhancement module proposed by the present invention mainly consists of depthwise separable convolution, deformable convolution, and average pooling operations. The output features of 2.2 are used to obtain a feature weight map based on the current features by depthwise separable convolution and deformable convolution and then fused to sharpen the structural information of the extracted features and enhance the network's ability to restore the image structure. First, the input features are converted into the form of feature tensors, and two deformable convolutions are used to extract features. Utilizing its property with offsets to focus on information of different shapes on the features, and then the depthwise separable convolution of is used to adjust its number of channels to obtain the feature weight map . At the same time, two depthwise separable convolutions are used to obtain the feature weight map . After the features extracted by the depthwise separable convolution and the deformable convolution are concatenated, the average pooling operation is used to obtain the feature weight map . Finally, the weight maps are fused to obtain the final weight map Q . A learnable weight parameter M is set (initial M= 0), multiply the input features with the weight map under the control of the weight factor, and finally obtain features with sharper structural information. Q Perform dot product to finally obtain features with sharper structural information.
[0049] Step 3: Input the training set and validation set that have been preprocessed and simulated and degraded in Step 1 into the synthetic aperture optical image restoration network LGFAformer in Step 2 for training, calculate the loss function and perform backpropagation, update the network parameters, and obtain the optimal parameter model. Specifically:
[0050] Randomly initialize the parameters of the synthetic aperture optical system image restoration network based on local-global feature enhancement. Input the preprocessed training set and validation set data in Step 1 into the synthetic aperture optical system image restoration network based on local-global feature enhancement in Step 2, output the restored optical image, calculate the loss with the original image, use the minimization of the discriminative loss as the optimization goal, automatically update the network parameters, obtain the optimal parameter model and save it.
[0051] In Step 3, the Charbonnier loss is used as the loss function of the overall network, and the specific definition is:
[0052] ,
[0053] Among them, represents the loss function, represents the generated image, represents the original image, is a constant.
[0054] Step 4: Input the test set that has been simulated and degraded in Step 1 into the optimal parameter model trained in Step 3, and output the restored image of the synthetic aperture optical system.
[0055] To prove the effectiveness of the synthetic aperture optical image restoration method based on local-global feature enhancement proposed in the present invention, after using DOTA remote sensing data for simulation and degradation, the model is trained, validated and tested. From Figure 6 in (a)-(h) and Table 1, it can be seen that the results of this embodiment are basically higher than the compared methods in terms of peak signal-to-noise ratio, structural similarity and perceptual similarity. The restored result is closest to the original image. In the case of being weaker than a certain SOTA method, the parameter quantity of the model of the present invention is also much smaller than that of this method, proving that this model is more lightweight and is not much different from SOTA. And for the ringing phenomenon that appears in the synthetic aperture optical system imaging in this embodiment, it has excellent performance compared with the traditional deep learning deblurring algorithm, as shown in Figure 7 in (a) and (b) and Table 2. Figure 8(a) and (b) in [the reference] and Table 3 show the comparison results of degraded images under different degrees of noise influence and the restoration results of different methods. In addition, in this embodiment, the YOLOv7 method is used to perform detection and comparison on the non-restored and restored images. As shown in (a) and (b) in [the reference], the results prove that this restoration method helps to improve the detection rate of the YOLOv7 method by about 40%, demonstrating the practicability of the method of the present invention. Figure 9 As shown in (a) and (b) in [the reference], the results prove that this restoration method helps to improve the detection rate of the YOLOv7 method by about 40%, demonstrating the practicability of the method of the present invention.
[0056] Table 1 Quantitative experimental results of different methods on 512×512 resolution images
[0057]
[0058] Table 2 Quantitative experimental results of different methods on images affected by ringing factors
[0059]
[0060] Table 3 Quantitative experimental results of different methods on images affected by 40dB - 50dB noise
[0061]
[0062] Based on the same inventive concept, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the aforementioned synthetic aperture optical image restoration method based on local-global feature enhancement are implemented.
[0063] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the aforementioned synthetic aperture optical image restoration method based on local-global feature enhancement are implemented.
[0064] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0065] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0066] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0067] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0068] The above embodiments are only for illustrating the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the present invention.
Claims
1. A synthetic aperture optical image restoration method based on local-global feature enhancement, characterized in that It includes the following steps: Step 1: Obtain a synthetic aperture optical image dataset and perform simulation degradation to obtain a simulated degraded image dataset. The images in the synthetic aperture optical image dataset and the simulated degraded image dataset correspond one by one. Randomly divide the corresponding image dataset into a training set, a validation set, and a test set; Step 2: Construct a synthetic aperture optical image restoration network, including an encoding network and a decoding network. The encoding network includes a first module to a fifth module connected in sequence. An input mapping module is connected between the first module and the original input image. The decoding network includes a sixth to tenth module connected in sequence. The input end of the sixth module is connected to the output end of the fifth module. The output end of the tenth module is connected to the second gated residual convolutional layer, and the output of the gated residual convolutional layer is used as the final output of the synthetic aperture optical image restoration network; The structures of the first module to the tenth module are the same, and each includes a linear gated Transformer layer, a first gated residual convolutional layer, and an adaptive scale feature enhancement module connected in sequence. The output of the input mapping module is skip-connected to the output of the second gated residual convolutional layer. The output of the adaptive scale feature enhancement module in the first module is skip-connected to the output of the adaptive scale feature enhancement module in the tenth module. The output of the adaptive scale feature enhancement module in the second module is skip-connected to the output of the adaptive scale feature enhancement module in the ninth module. The output of the adaptive scale feature enhancement module in the third module is skip-connected to the output of the adaptive scale feature enhancement module in the eighth module. The output of the adaptive scale feature enhancement module in the fourth module is skip-connected to the output of the adaptive scale feature enhancement module in the seventh module. The output of the adaptive scale feature enhancement module in the fifth module is skip-connected to the output of the adaptive scale feature enhancement module in the sixth module; The input mapping module maps the original input image into an original input feature sequence. The first to fifth modules respectively output feature sequences with lengths of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original input feature sequence. The sixth to tenth modules respectively output feature sequences with lengths of 1 / 32, 1 / 16, 1 / 8, 1 / 4, and 1 / 2 of the original input feature sequence; The linear gated Transformer layer includes a first to fourth normalization layer, a first to second forward propagation module, a window multi-head focused linear attention module, and a shifted window multi-head focused linear attention module; The input \(x\) of the linear gated Transformer layer is successively passed through the first normalization layer and the window multi-head focused linear attention module to obtain the output of the window multi-head focused linear attention module. The output of the window multi-head focused linear attention module is added to \(x\) to obtain the feature sequence \(x1\); \(x1\) is successively passed through the second normalization layer and the first feed-forward module to obtain the output of the first feed-forward module. The output of the first feed-forward module is added to \(x1\) to obtain the feature sequence \(x2\); \(x2\) is successively passed through the third normalization layer and the shifted window multi-head focused linear attention module to obtain the output of the shifted window multi-head focused linear attention module. The output of the shifted window multi-head focused linear attention module is added to \(x2\) to obtain the feature sequence \(x3\); \(x3\) is successively passed through the fourth normalization layer and the second feed-forward module to obtain the output of the second feed-forward module. The output of the second feed-forward module is added to \(x3\) as the output of the linear gated Transformer layer; Both the window multi-head focused linear attention module and the shifted window multi-head focused linear attention module include a focused linear attention module; The window multi-head focused linear attention module divides the output of the first normalization layer into multiple windows, calculates the attention for each window using the focused linear attention module respectively, and then concatenates the attentions of all windows as the output of the window multi-head focused linear attention module; The shifted window multi-head focused linear attention module first shifts the output of the third normalization layer and then divides it into multiple windows, calculates the attention for each window using the focused linear attention module respectively, and then concatenates the attentions of all windows as the output of the shifted window multi-head focused linear attention module; The input of the focused linear attention module is linearly transformed into a query vector, a key vector, and a value vector. The query vector is scaled after passing through the Relu activation function and then solved by the first Euclidean norm; the key vector is scaled after passing through the Relu activation function and then solved by the second Euclidean norm; the results of the first Euclidean norm and the second Euclidean norm enter the focusing function to obtain the focusing result, and the focusing result is transposed to obtain the transposed result; the transposed result is added to the value vector by Einstein summation to obtain the K-V key-value pair. The K-V key-value pair is added to the focusing result by Einstein summation to obtain the first result; the value vector is added to the Einstein result after passing through a 3×3 depthwise separable convolution, and the second result is obtained. The second result is successively passed through a fully connected layer and a dropout layer to obtain the output of the focused linear attention module; The formula of the focusing function is as follows: Sim(Q i ,K j ) = φ p (Q i )φ p (K j ) T φ p (Q i ) = f p (Relu(Q i )) Among them, Sim(Q i ,K j ) represents the similarity function, Q i ,K j represent the query vector and key vector in different dimensions respectively, φ p (Q i ) represents the calculation of the focused linear attention on the activation function of Q i , φ p (K j ) represents the calculation of the focused linear attention on the activation function of K j , Relu represents the Relu activation function, f p (·) represents the focused linear attention calculation function, Relu(Q i ) **p represents the element power p of Relu(Q i ); Step 3: Input the training set and the validation set obtained in Step 1 into the synthetic aperture optical image restoration network constructed in Step 2 for training, calculate the loss function and perform backpropagation to update the network parameters, and obtain the best parameter model after training; Step 4: Input the test set obtained in Step 1 into the best parameter model trained in Step 3 to output the restored image of the synthetic aperture optical system.
2. The synthetic aperture optical image restoration method based on local-global feature enhancement according to claim 1, wherein The specific process of Step 1 is as follows: Obtain the synthetic aperture optical image dataset, measure the parameters of the actual synthetic aperture optical system, obtain the system point spread function, save it in txt format, and then use the PyCharm software to read the point spread function and perform convolution with the synthetic aperture optical image dataset to obtain the simulated degraded image dataset; Crop the images in the synthetic aperture optical image dataset and the simulated degraded image dataset into images of size 512×512, and then correspond the images in the synthetic aperture optical image dataset and the simulated degraded image dataset one by one as the input image and the target image; Randomly divide the corresponding image dataset into three parts: the training set accounts for 80%, the validation set accounts for 10%, and the test set accounts for 10%.
3. The synthetic aperture optical image restoration method based on local-global feature enhancement according to claim 1, characterized in that The structures of the first forward propagation module and the second forward propagation module are the same, and both include a 1×1 convolutional layer, the first to third 1×1 depthwise separable convolutional layers, the first to second 3×3 depthwise separable convolutional layers, the Gelu activation function, and the effective channel attention module; After the input of each forward propagation module passes through the 1×1 convolutional layer, the first feature sequence output by the 1×1 convolutional layer is converted into a feature map; the feature map passes through two branches respectively. In the first branch, the feature map passes through the first 1×1 depthwise separable convolutional layer, the first 3×3 depthwise separable convolutional layer, and the effective channel attention module in sequence to obtain the output of the first branch; In the second branch, the feature map passes through the second 1×1 depthwise separable convolutional layer, the second 3×3 depthwise separable convolutional layer, and the Gelu activation function in sequence to obtain the output of the second branch; The output of the second branch is multiplied pointwise with the output of the first branch and the feature map. After the result of the pointwise multiplication passes through the third 1×1 depthwise separable convolutional layer, the feature map output by the third 1×1 depthwise separable convolutional layer is converted into a second feature sequence, and the second feature sequence is added to the first feature sequence as the output of each forward propagation module.
4. The synthetic aperture optical image restoration method based on local-global feature enhancement according to claim 1, wherein The structures of the first gated residual convolutional layer and the second gated residual convolutional layer are the same, and both include the fifth to seventh normalization layers, the fourth to seventh 1×1 depthwise separable convolutional layers, the third 3×3 depthwise separable convolutional layer, the first to second deformable convolutional layers, and the simple gating mechanism; Convert the feature sequence input to each gated residual convolutional layer into a feature map. After the feature map passes through the fifth normalization layer, the fourth 1×1 depthwise separable convolutional layer, the third 3×3 depthwise separable convolutional layer, the fifth 1×1 depthwise separable convolutional layer, and the sixth normalization layer in sequence, the output F of the sixth normalization layer is obtained. w After the feature map passes through the first deformable convolutional layer, the seventh normalization layer, and the second deformable convolutional layer in sequence, the output F of the second deformable convolutional layer is obtained. f ; F f and F w are added together to obtain the feature map F'. F' passes through the sixth 1×1 depthwise separable convolutional layer, a simple gating mechanism, and the seventh 1×1 depthwise separable convolutional layer in sequence to obtain the output F" of the seventh 1×1 depthwise separable convolutional layer. F" and F' are added together to obtain the feature map F'''. Convert F''' into a feature sequence as the output of each gated residual convolutional layer.
5. The synthetic aperture optical image restoration method based on local-global feature enhancement according to claim 1, wherein, The adaptive scale feature enhancement module includes the eighth to eleventh 1×1 depthwise separable convolutional layers, the fourth to sixth 3×3 depthwise separable convolutional layers, the third to fourth deformable convolutional layers, and the average pooling layer; Convert the feature sequence input by the adaptive scale feature enhancement module into a feature map F1. F1 passes through two branches. In the first branch, F1 is successively passed through the third deformable convolutional layer, the fourth deformable convolutional layer, and the eighth 1×1 depthwise separable convolutional layer to obtain the feature map F2. In the second branch, F1 is successively passed through the fourth 3×3 depthwise separable convolutional layer, the fifth 3×3 depthwise separable convolutional layer, and the tenth 1×1 depthwise separable convolutional layer to obtain the feature map F3. F3 and F2 are concatenated to obtain the feature map F4. F2 passes through the ninth 1×1 depthwise separable convolutional layer to obtain the feature weight map q1, F3 passes through the eleventh 1×1 depthwise separable convolutional layer to obtain the feature weight map q3, and F4 is successively passed through the sixth 3×3 depthwise separable convolutional layer and the average pooling layer to obtain the feature weight map q2. q2 is added to q1 to obtain the feature weight map q 2’ , q 2’ is added to q3 to obtain the feature weight map Q. Q is multiplied by F1 pointwise to obtain the feature map N, and N is converted into a feature sequence as the output of the adaptive scale feature enhancement module.
6. The method for restoring a synthetic aperture optical image based on local-global feature enhancement according to claim 1, wherein In step 3, the loss function is the Charbonnier loss, and the formula is as follows: Among them, represents the loss function, I' represents the generated image, represents the original image, ε is a constant, ε = 0.
001.
7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the synthetic aperture optical image restoration method based on local-global feature enhancement according to any one of claims 1 to 6.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the synthetic aperture optical image restoration method based on local-global feature enhancement according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image semantic segmentation method based on Transform architecture
CN115482382A
Plant leaf lesion identification method based on improved Transform model
CN116311186A