A near-field SAR image enhancement method based on a three-dimensional convolutional neural network
Patent Information
- Application Number
- CN202311379812.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-11
- Filing Date
- 2023-10-23
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2043-10-23
AI Technical Summary
然而,由于现有的基于深度学习的SAR图像质量增强技术主要针对远场SAR和二维SAR图像设计,并未考虑三维SAR图像重构问题,使得三维SAR图像重构精度仍然存在进一步提升的空间,研究充分挖掘三维SAR图像特性的图像增强技术成为了一个关键性的问题
[0078] The innovation of this invention lies in constructing a three-dimensional convolutional neural network based on the original three-dimensional convolution, which realizes the full utilization of three-dimensional structural features in near-field SAR image reconstruction, and makes the image reconstruction model in this invention have superior reconstruction accuracy.
Smart Images

Figure CN117422830B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of Synthetic Aperture Radar (SAR) image enhancement technology, and relates to a near-field SAR image enhancement method based on a three-dimensional convolutional neural network. Background Technology
[0002] Synthetic Aperture Radar (SAR) is a radar system with all-weather, day-and-night imaging capabilities. It can image at any time, whether it's day or night, clear or rainy / snowy, overcoming the limitations of optical and infrared systems that cannot image at night or in complex weather conditions. However, traditional SAR has a long observation range, small relative viewing angle change with the target, and a small synthetic aperture angle, resulting in limited azimuth resolution of the image.
[0003] Unlike traditional SAR, near-field SAR operates in the near-field region of the target. In this region, the relative viewing angle change with the observed target is significantly increased, resulting in a larger synthetic aperture angle and a significantly improved azimuth resolution, reaching centimeter-level resolution. Furthermore, near-field SAR systems can typically use stepped-frequency signals as transmission signals, allowing the range resolution to reach the same order of magnitude as the azimuth resolution. Therefore, near-field SAR enables fine-grained target imaging and has wide-ranging applications in numerous fields, including human security imaging, concealed object detection, building deformation detection, autonomous driving, and scattering characteristic measurement.
[0004] Currently, near-field SAR imaging typically employs a backprojection algorithm based on matched filtering. However, due to finite aperture sampling, target sidelobes are unavoidable in near-field SAR imaging results. In addition, clutter from the ground and surrounding environment is also present. These sidelobes and clutter negatively impact subsequent image applications. For example, they are easily misidentified as real targets, leading to a loss of accuracy in target detection and recognition. Therefore, it is necessary to enhance the imaging results using image enhancement methods to obtain target images with lower target sidelobes and less clutter.
[0005] Current SAR image enhancement techniques have achieved good imaging results from different perspectives and have high reconstruction accuracy. However, since existing deep learning-based SAR image enhancement techniques are mainly designed for far-field SAR and 2D SAR images and do not consider the 3D SAR image reconstruction problem, there is still room for improvement in the reconstruction accuracy of 3D SAR images. Therefore, researching image enhancement techniques that fully exploit the characteristics of 3D SAR images has become a key issue.
[0006] Therefore, to address the aforementioned problems, this invention proposes a near-field SAR image enhancement method based on a three-dimensional convolutional neural network. This method proposes a novel three-dimensional convolutional network structure that enables semantic feature extraction and effective reconstruction of the extracted semantic features across three dimensions of the image. This feature encoding and decoding facilitates the extraction of richer SAR image information and the reconstruction of better near-field SAR images, ensuring excellent near-field three-dimensional SAR image reconstruction accuracy. Summary of the Invention
[0007] This invention belongs to the field of synthetic aperture radar (SAR) image enhancement technology, and discloses a near-field SAR image enhancement method based on a three-dimensional convolutional neural network to enhance the quality of near-field SAR images. The method mainly includes five parts: preparing the dataset, constructing a three-dimensional convolutional neural network, establishing an image enhancement model, testing the image enhancement model, and evaluating the image enhancement model. Based on the original three-dimensional convolution, this method constructs a three-dimensional convolutional neural network, thereby optimizing the network structure to improve reconstruction accuracy. Experimental results on a near-field three-dimensional SAR image dataset show that, compared with the original near-field SAR imaging algorithm, this invention achieves higher SAR image reconstruction accuracy.
[0008] To facilitate the description of the present invention, the following terms are defined first:
[0009] Definition 1: Synthetic Aperture Radar
[0010] Synthetic Aperture Radar (SAR) is a high-resolution microwave imaging radar with the advantages of all-weather and all-day operation. It has been widely applied in various fields, such as topographic mapping, guidance, environmental remote sensing, and resource exploration. A crucial prerequisite for SAR applications and the main goal of signal processing is to acquire high-resolution, high-precision microwave images through imaging algorithms. See "Pi Yiming, Yang Jianyu, Fu Yusheng, Yang Xiaobo. Principles of Synthetic Aperture Radar Imaging [M]. University of Electronic Science and Technology of China Press. 2007."
[0011] Definition 2: Backward Projection Algorithm
[0012] The back projection algorithm uses the radar platform's position information to calculate the historical distance between the platform and scene pixels. Then, by traversing the distance history, it finds the corresponding echoes in the data after echo pulse compression interpolation, performs phase compensation for the corresponding distances, coherently accumulates the data, and projects the accumulated result into the image space to complete the imaging process. The back projection algorithm mainly includes the following steps: range-direction matched filtering, range-direction zero-padding interpolation, range-direction echo indexing, azimuth phase compensation, and azimuth coherent accumulation. See "Shi Jun. Research on the Principle and Imaging Technology of Bistatic SAR and Linear Array SAR [D]. Doctoral Dissertation, University of Electronic Science and Technology of China. 2009". Definition 3: Near-field 3D SAR Image Dataset
[0013] The Near-Field 3D SAR Image Dataset refers to a near-field 3D SAR image reconstruction dataset, which can be used to train deep learning models and for researchers to evaluate the performance of their algorithms on this unified dataset. The Near-Field 3D SAR Image Dataset contains 1200 images and 10 types of aircraft, with an average of 120 images per aircraft. This publicly available dataset for near-field 3D SAR image enhancement utilizes electromagnetic simulation software to simulate the scattering characteristics of aircraft and uses a back-projection algorithm to generate near-field 3D SAR images. It can simultaneously obtain SAR images and corresponding ground truth values, making method testing and evaluation more convenient. The dataset can be obtained from the reference "Zhang W, Zhang X, et al. Near-Field SAR Image Restoration Framework Via Deep Learning[C] / / IGARSS2023-2023IEEE International Geoscience and Remote Sensing Symposium.IEEE,2023".
[0014] Definition 4: Classical Convolutional Neural Network Methods
[0015] Classical convolutional neural networks (CNNs) refer to a class of feedforward neural networks that include convolutional computations and have a deep structure. CNNs are constructed by mimicking the visual perception mechanisms of biological systems, enabling both supervised and unsupervised learning. The shared parameters of the convolutional kernels within their hidden layers and the sparsity of inter-layer connections allow CNNs to extract features with relatively low computational cost. In recent years, CNNs have made rapid progress in computer vision, natural language processing, and speech recognition, and their powerful feature learning capabilities have attracted widespread attention from experts and scholars both domestically and internationally. For details on classic CNN methods, please refer to the literature “Zhang Suofei, Feng Ye, Wu Xiaofu. Progress in target detection algorithms based on deep convolutional neural networks [J / OL]. Journal of Nanjing University of Posts and Telecommunications (Natural Science Edition), 2019(05):1-9. https: / / doi.org / 10.14132 / j.cnki.1673-5439.2019.05.010.”
[0016] Definition 5: Classic CNN Feature Extraction Method
[0017] Classical CNN feature extraction involves using a CNN to extract features from the original input image. In short, the original input image is transformed into a series of feature maps through convolutional operations on different features. In a CNN, the convolutional kernels in the convolutional layers continuously slide across the image for computation. Simultaneously, the max-pooling layer is responsible for taking the maximum value of each local block in the inner product result. Therefore, CNNs implement image feature extraction methods through convolutional layers and max-pooling layers. For a detailed explanation of classic CNN feature extraction, please refer to the website "https: / / blog.csdn.net / qq_30815237 / article / details / 86703620".
[0018] Definition 6: Convolutional Kernel
[0019] In image processing, a convolution kernel is a weighted average of pixels in a small region of an input image, which is then used to produce the corresponding pixels in the output image. The weights are defined by a function called the convolution kernel. The role of the convolution kernel is to extract features. A larger convolution kernel size means a larger receptive field, but of course, more parameters are required. As early as 1998, LeCun's LetNet-5 model showed that there are local correlations in the spatial domain of an image, and the convolution process is a way to extract these local correlations. For details on how to set the convolution kernel, please refer to the literature "LeCun Y, Bottou L, Bengio Y, et al. Gradient-based learning applied to document recognition[J]. Proceedings of the IEEE, 1998, 86(11):2278-2324.".
[0020] Definition 7: Classic method for setting convolution kernel size
[0021] The kernel size refers to the length, width, and depth of the convolution kernel, denoted as L×W×D, where L represents the length, W represents the width, and D represents the depth. Setting the kernel size means determining the specific values of L, W, and D. Generally, to achieve the same receptive field, the smaller the kernel, the fewer parameters and computational cost are required. Specifically, the length and width of the kernel must be greater than 1 to increase the receptive field. Even with symmetrical zero-padding, kernels with even-sized kernels cannot guarantee that the input and output feature spectrum sizes remain unchanged. Generally, 3 is used as the kernel size. For details on setting the kernel size, please refer to the literature "Lecun Y, Bottou L, Bengio Y, et al. Gradient-based learning applied to document recognition[J]. Proceedings of the IEEE, 1998, 86(11):2278-2324.".
[0022] Definition 8: Classic method for setting the stride of a convolution kernel
[0023] The kernel stride refers to the length of the convolution kernel in each movement, denoted as S. Setting the kernel stride means determining the specific value of S. Generally, the larger the stride, the fewer features are extracted; conversely, the more features are extracted. Convolutional layers typically use a stride of 1, while max pooling layers typically use a stride of 2. For classic kernel stride setting methods, please refer to the literature "Lecun Y, Bottou L, Bengio Y, et al. Gradient-based learning applied to document recognition[J]. Proceedings of the IEEE, 1998, 86(11):2278-2324."
[0024] Definition 9: Classical Convolutional Layer
[0025] A convolutional layer consists of several convolutional units, each with parameters optimized using the backpropagation algorithm. The purpose of convolution is to extract different features from the input. The first convolutional layer may only extract low-level features such as edges, lines, and corners, while more layers can iteratively extract more complex features from these low-level features. For a detailed explanation of classic convolutional layers, please refer to the website "https: / / www.zhihu.com / question / 49376084".
[0026] Definition 10: Classic Max Pooling Layer
[0027] Max pooling layers are used to extract the maximum value of all neurons within a region of the previous network layer. This is so that during backpropagation, the gradient value can be propagated to the location of the corresponding maximum value. Max pooling layers can reduce the shift in the estimated mean caused by convolutional layer parameter errors, preserving more texture information. For a detailed explanation of classic max pooling layers, see the reference "Lin M, Chen Q, Yan S. Network in network[J]. arXiv preprint arXiv:1312.4400,2013."
[0028] Definition 11: 3D Convolution
[0029] 3D convolution adds a depth dimension compared to 2D convolution, performing a sliding window operation in the image's height, width, and channels. Its convolution kernel is an L×W×D 3D matrix, where L represents length, W represents width, and D represents depth. The input data for 3D convolution must be 3D data, and the corresponding output is also 3D data. For details on 3D convolution, see the reference "Tran D, Bourdev L, Fergus R, et al. Learning spatiotemporal features with 3d convolutional networks[C] / / Proceedings of the IEEE international conference on computervision.2015:4489-4497."
[0030] Definition 12: Upsampling
[0031] Upsampling primarily uses image interpolation to enlarge images, thereby achieving higher resolution. There are various upsampling methods; this invention employs bilinear interpolation, specifically implemented by calling the `torch.nn.Upsample` function in Python. For details on upsampling, please refer to the literature "Zhao R, Li Q, Wu J, et al. Anested U-shape network with multi-scaleupsample attention for robust retinal vascular segmentation[J]. PatternRecognition,2021,120:107998."
[0032] Definition 13: The classic Adam algorithm
[0033] The classic Adam algorithm is a first-order optimization algorithm that can replace the traditional stochastic gradient descent process. It iteratively updates the weights of a neural network based on training data. The Adam algorithm differs from traditional stochastic gradient descent. Stochastic gradient descent updates all weights with a single learning rate that remains unchanged during training. Adam, however, designs independent adaptive learning rates for different parameters by calculating the first and second moment estimates of the gradient. See the literature "Kingma, D.; Ba, J. Adam: A Method for Stochastic Optimization. arXiv 2014, arXiv:1412.6980." for details.
[0034] Definition 14: Standard testing method for network detection
[0035] The standard method for testing detection networks refers to performing a final test on a test set to obtain the detection results of the detection model on the test set. See the literature "C. Lu, and W. Li, "Ship Classification in High-Resolution SAR Images via Transfer Learning with Small Training Dataset," Sensors, vol. 19, no. 1, pp. 63, 2018."
[0036] Definition 15: Standard Evaluation Index Calculation Method
[0037] Normalized mean square error (NMSE) refers to the error between a predicted and a true measurement. NMSE is defined as follows: Where n represents a number, Y i Indicates the predicted value. Represents the actual value;
[0038] Structural similarity (SSIM) measures the degree of similarity between a reconstructed image and the original image. SSIM is defined as follows: Where C1 and C2 represent constants, μ x μ y σ represents the mean pixel value between the reconstructed image and the real image. x σ y σ represents the pixel variance between the reconstructed image and the real image. xy This represents the pixel covariance between the reconstructed image and the original image;
[0039] For details on how to calculate the above parameter values, please refer to the reference "Li Hang. Statistical Learning Methods [M]. Beijing: Tsinghua University Press, 2012."
[0040] This invention provides a dual-polarization SAR ship detection method based on grouped hybrid attention, which includes the following steps:
[0041] Step 1: Prepare the dataset
[0042] For the near-field three-dimensional SAR image dataset provided in Definition 3, the order of SAR images in the dataset is adjusted using a random method to obtain a new dataset, denoted as SAR_new;
[0043] The SAR_new dataset is divided into two parts in a 7:3 ratio to obtain a training set and a test set. The training set is denoted as Train_SAR, and the test set is denoted as Test_SAR.
[0044] Step 2: Construct a feature extraction module for a neural network based on 3D convolution.
[0045] Step 2.1: First Layer Feature Extraction
[0046] The input layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the first layer of a three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f1. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C1 and M1 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C1 is set to 3×3×8 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C1 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M1 is set to 2 using the classic kernel stride setting method in Definition 8.
[0047] Using the classic CNN feature extraction method in Definition 5, a SAR image in the Training_SAR set obtained in step 1 is processed to obtain the first layer feature output, denoted as A1;
[0048] Step 2.2: Second-layer feature extraction
[0049] The intermediate layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the second layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f2. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C2 and M2 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C2 is set to 3×3×16 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C2 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M2 is set to 2 using the classic kernel stride setting method in Definition 8.
[0050] Using the classic CNN feature extraction method in Definition 5, the first layer feature output A1 obtained in step 2.1 is processed to obtain the second layer feature output, denoted as A2;
[0051] Step 2.3: Feature Extraction at Layer 3
[0052] The intermediate layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the third layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f3. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C3 and M3 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C3 is set to 3×3×32 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C3 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M3 is set to 2 using the classic kernel stride setting method in Definition 8.
[0053] Using the classic CNN feature extraction method in Definition 5, the second layer feature output A2 obtained in step 2.2 is processed to obtain the third layer feature output, denoted as A3;
[0054] Step 2.4: Feature Extraction at Layer 4
[0055] The intermediate layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the fourth layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f4. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C4 and M4 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C4 is set to 3×3×64 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C4 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M4 is set to 2 using the classic kernel stride setting method in Definition 8.
[0056] Using the classic CNN feature extraction method in Definition 5, the feature output A3 of the third layer obtained in step 2.3 is processed to obtain the feature output of the fourth layer, denoted as A4;
[0057] Finally, we obtain the constructed feature extraction module and the feature outputs of all layers, denoted as Encoder and A, respectively. s ,s=1,...,5. Step 3: Construct an image reconstruction module based on 3D convolution to build a neural network.
[0058] Step 3.1: Image Reconstruction at Layer 1
[0059] The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the fifth layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f5. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C5 and U5 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C5 is set to 3×3×64 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C5 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U5 is set to upsample the image to twice its original size in the height and width directions.
[0060] Using the classic CNN feature extraction method in Definition 5, the feature output A4 obtained in step 2.4 is processed to obtain the first layer image reconstruction result, denoted as A5;
[0061] Step 3.2: Second Layer Image Reconstruction
[0062] The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the 6th layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f6. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C6 and U6 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C6 is set to 3×3×32 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C6 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U6 is set to upsample the image to twice its original size in the height and width directions.
[0063] Using the classic CNN feature extraction method in Definition 5, the sum of the feature output A3 obtained in step 2.3 and the reconstruction result A5 obtained in step 3.1 is processed to obtain the second layer image reconstruction result, denoted as A6;
[0064] Step 3.3: Image Reconstruction at Layer 3
[0065] The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the 7th layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f7. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C7 and U7 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C7 is set to 3×3×16 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C7 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U7 is set to upsample the image to twice its original size in the height and width directions.
[0066] Using the classic CNN feature extraction method in Definition 5, the sum of the feature output A2 obtained in step 2.2 and the reconstruction result A6 obtained in step 3.1 is processed to obtain the image reconstruction result of the third layer, denoted as A7;
[0067] Step 3.4: Image Reconstruction at Layer 4
[0068] The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the 8th layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f8. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C8 and U8 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C8 is set to 3×3×8 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C8 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U8 is set to upsample the image to twice its original size in the height and width directions.
[0069] Using the classic CNN feature extraction method in Definition 5, the sum of the feature output A1 obtained in step 2.1 and the reconstruction result A7 obtained in step 3.3 is processed to obtain the final image reconstruction result, denoted as A8;
[0070] This completes the construction of the three-dimensional convolutional neural network.
[0071] Step 4: Establish an image reconstruction model
[0072] Using the training set Train_SAR obtained in step 1 as input, the classic Adam algorithm in definition 13 is used to train the three-dimensional convolutional neural network completed in step 3. After training, the image reconstruction model is obtained, denoted as 3D-CNN.
[0073] Step 5: Test the image reconstruction model
[0074] Using the test set Test_SAR obtained in step 1, the image reconstruction model 3D-CNN obtained in step 3 is tested using the standard detection network testing method in Definition 14, and the test results of the test set on the image reconstruction model are obtained, denoted as Result.
[0075] Step 6: Evaluate the image reconstruction model
[0076] Using the test result Result of the image reconstruction model obtained in step 5 as input, the mean squared error and structural similarity are calculated using the standard evaluation index calculation method in definition 15, and denoted as MSE and SSIM respectively.
[0077] This concludes the entire method.
[0078] The innovation of this invention lies in constructing a three-dimensional convolutional neural network based on the original three-dimensional convolution, which realizes the full utilization of three-dimensional structural features in near-field SAR image reconstruction, and makes the image reconstruction model in this invention have superior reconstruction accuracy.
[0079] The advantage of this invention is that it uses a three-dimensional convolutional neural network to fully utilize the three-dimensional structural features of near-field SAR images, and can provide a method for image enhancement of three-dimensional near-field SAR images to solve the problem of insufficient reconstruction accuracy of existing three-dimensional near-field SAR images. Attached Figure Description
[0080] Figure 1 This is a flowchart illustrating the three-dimensional near-field SAR image reconstruction method of the present invention.
[0081] Figure 2 This table presents a numerical comparison of the mean square error and structural similarity between the three-dimensional near-field SAR image reconstruction method and the back projection algorithm in this invention. Detailed Implementation
[0082] The following is in conjunction with the appendix Figure 1 The present invention will be described in further detail below.
[0083] Step 1: Prepare the dataset
[0084] For the near-field three-dimensional SAR image dataset provided in Definition 3, the order of SAR images in the dataset is adjusted using a random method to obtain a new dataset, denoted as SAR_new;
[0085] The SAR_new dataset is divided into two parts in a 7:3 ratio to obtain a training set and a test set. The training set is denoted as Train_SAR, and the test set is denoted as Test_SAR.
[0086] Step 2: Construct a feature extraction module for a neural network based on 3D convolution.
[0087] Step 2.1: First Layer Feature Extraction
[0088] The input layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the first layer of a three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f1. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C1 and M1 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C1 is set to 3×3×8 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C1 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M1 is set to 2 using the classic kernel stride setting method in Definition 8.
[0089] Using the classic CNN feature extraction method in Definition 5, a SAR image in the Training_SAR set obtained in step 1 is processed to obtain the first layer feature output, denoted as A1;
[0090] Step 2.2: Second-layer feature extraction
[0091] The intermediate layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the second layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f2. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C2 and M2 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C2 is set to 3×3×16 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C2 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M2 is set to 2 using the classic kernel stride setting method in Definition 8.
[0092] Using the classic CNN feature extraction method in Definition 5, the first layer feature output A1 obtained in step 2.1 is processed to obtain the second layer feature output, denoted as A2;
[0093] Step 2.3: Feature Extraction at Layer 3
[0094] The intermediate layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the third layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f3. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C3 and M3 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C3 is set to 3×3×32 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C3 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M3 is set to 2 using the classic kernel stride setting method in Definition 8.
[0095] Using the classic CNN feature extraction method in Definition 5, the second layer feature output A2 obtained in step 2.2 is processed to obtain the third layer feature output, denoted as A3;
[0096] Step 2.4: Feature Extraction at Layer 4
[0097] The intermediate layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the fourth layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f4. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C4 and M4 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C4 is set to 3×3×64 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C4 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M4 is set to 2 using the classic kernel stride setting method in Definition 8.
[0098] Using the classic CNN feature extraction method in Definition 5, the feature output A3 of the third layer obtained in step 2.3 is processed to obtain the feature output of the fourth layer, denoted as A4;
[0099] Finally, we obtain the constructed feature extraction module and the feature outputs of all layers, denoted as Encoder and A, respectively. s ,s=1,...,5. Step 3: Construct an image reconstruction module based on 3D convolution to build a neural network.
[0100] Step 3.1: Image Reconstruction at Layer 1
[0101] The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the fifth layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f5. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C5 and U5 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C5 is set to 3×3×64 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C5 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U5 is set to upsample the image to twice its original size in the height and width directions.
[0102] Using the classic CNN feature extraction method in Definition 5, the feature output A4 obtained in step 2.4 is processed to obtain the first layer image reconstruction result, denoted as A5;
[0103] Step 3.2: Second Layer Image Reconstruction
[0104] The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the 6th layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f6. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C6 and U6 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C6 is set to 3×3×32 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C6 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U6 is set to upsample the image to twice its original size in the height and width directions.
[0105] Using the classic CNN feature extraction method in Definition 5, the sum of the feature output A3 obtained in step 2.3 and the reconstruction result A5 obtained in step 3.1 is processed to obtain the second layer image reconstruction result, denoted as A6;
[0106] Step 3.3: Image Reconstruction at Layer 3
[0107] The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the 7th layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f7. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C7 and U7 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C7 is set to 3×3×16 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C7 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U7 is set to upsample the image to twice its original size in the height and width directions.
[0108] Using the classic CNN feature extraction method in Definition 5, the sum of the feature output A2 obtained in step 2.2 and the reconstruction result A6 obtained in step 3.1 is processed to obtain the image reconstruction result of the third layer, denoted as A7;
[0109] Step 3.4: Image Reconstruction at Layer 4
[0110] The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the 8th layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f8. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C8 and U8 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C8 is set to 3×3×8 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C8 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U8 is set to upsample the image to twice its original size in the height and width directions.
[0111] Using the classic CNN feature extraction method in Definition 5, the sum of the feature output A1 obtained in step 2.1 and the reconstruction result A7 obtained in step 3.3 is processed to obtain the final image reconstruction result, denoted as A8;
[0112] This completes the construction of the three-dimensional convolutional neural network.
[0113] Step 4: Establish an image reconstruction model
[0114] Using the training set Train_SAR obtained in step 1 as input, the classic Adam algorithm in definition 13 is used to train the three-dimensional convolutional neural network completed in step 3. After training, the image reconstruction model is obtained, denoted as 3D-CNN.
[0115] Step 5: Test the image reconstruction model
[0116] Using the test set Test_SAR obtained in step 1, the image reconstruction model 3D-CNN obtained in step 3 is tested using the standard detection network testing method in Definition 14, and the test results of the test set on the image reconstruction model are obtained, denoted as Result.
[0117] Step 6: Evaluate the image reconstruction model
[0118] Using the test result Result of the image reconstruction model obtained in step 5 as input, the mean squared error and structural similarity are calculated using the standard evaluation index calculation method in definition 15, and denoted as MSE and SSIM respectively.
[0119] This concludes the entire method.
[0120] like Figure 2 As shown, the structural similarity achieved by this invention on a near-field SAR image dataset is 0.86. Furthermore, this invention achieves higher imaging quality compared to existing backprojection techniques, demonstrating its ability to reconstruct high-quality near-field 3D SAR images.
Claims
1. A near-field SAR image enhancement method based on a three-dimensional convolutional neural network, characterized by its... Includes the following steps: Step 1: Prepare the dataset For the near-field three-dimensional SAR image dataset provided in Definition 3, the order of SAR images in the dataset is adjusted using a random method to obtain a new dataset, denoted as SAR_new; The SAR_new dataset is divided into two parts in a 7:3 ratio to obtain a training set and a test set. The training set is denoted as Train_SAR and the test set is denoted as Test_SAR. Step 2: Construct a feature extraction module for a neural network based on 3D convolution. Step 2.1: First Layer Feature Extraction The input layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the first layer of a three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f1. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C1 and M1 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C1 is set to 3×3×8 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C1 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M1 is set to 2 using the classic kernel stride setting method in Definition 8. Using the classic CNN feature extraction method in Definition 5, a SAR image in the Training_SAR set obtained in step 1 is processed to obtain the first layer feature output, denoted as A1; Step 2.2: Second-layer feature extraction The intermediate layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the second layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f2. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C2 and M2 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C2 is set to 3×3×16 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C2 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M2 is set to 2 using the classic kernel stride setting method in Definition 8. Using the classic CNN feature extraction method in Definition 5, the first layer feature output A1 obtained in step 2.1 is processed to obtain the second layer feature output, denoted as A2; Step 2.3: Feature Extraction at Layer 3 The intermediate layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the third layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f3. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C3 and M3 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C3 is set to 3×3×32 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C3 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M3 is set to 2 using the classic kernel stride setting method in Definition 8. Using the classic CNN feature extraction method in Definition 5, the second layer feature output A2 obtained in step 2.2 is processed to obtain the third layer feature output, denoted as A3; Step 2.4: Feature Extraction at Layer 4 The intermediate layer of the feature extraction module is established using the classic convolutional neural network method in Definition 4, resulting in the fourth layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f4. This layer consists of the classic convolutional layer in Definition 9 and the classic max pooling layer in Definition 10, denoted as C4 and M4 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C4 is set to 3×3×64 using the classic kernel size setting method in Definition 7, the kernel stride of the three-dimensional convolution of C4 is set to 1 using the classic kernel stride setting method in Definition 8, and the kernel stride of M4 is set to 2 using the classic kernel stride setting method in Definition 8. Using the classic CNN feature extraction method in Definition 5, the feature output A3 of the third layer obtained in step 2.3 is processed to obtain the feature output of the fourth layer, denoted as A4; Finally, we obtain the constructed feature extraction module and the feature outputs of all layers, denoted as Encoder and A, respectively. s ,s=1,...,5; Step 3: Construct an image reconstruction module based on 3D convolution to build a neural network. Step 3.1: Image Reconstruction at Layer 1 The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the fifth layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f5. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C5 and U5 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C5 is set to 3×3×64 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C5 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U5 is set to upsample the image to twice its original size in the height and width directions. Using the classic CNN feature extraction method in Definition 5, the feature output A4 obtained in step 2.4 is processed to obtain the first layer image reconstruction result, denoted as A5; Step 3.2: Second Layer Image Reconstruction The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the 6th layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f6. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C6 and U6 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C6 is set to 3×3×32 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C6 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U6 is set to upsample the image to twice its original size in the height and width directions. Using the classic CNN feature extraction method in Definition 5, the sum of the feature output A3 obtained in step 2.3 and the reconstruction result A5 obtained in step 3.1 is processed to obtain the second layer image reconstruction result, denoted as A6; Step 3.3: Image Reconstruction at Layer 3 The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the 7th layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f7. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C7 and U7 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C7 is set to 3×3×16 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C7 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U7 is set to upsample the image to twice its original size in the height and width directions. Using the classic CNN feature extraction method in Definition 5, the sum of the feature output A2 obtained in step 2.2 and the reconstruction result A6 obtained in step 3.1 is processed to obtain the image reconstruction result of the third layer, denoted as A7; Step 3.4: Image Reconstruction at Layer 4 The input layer of the image reconstruction module is established using the classic convolutional neural network method in Definition 4, resulting in the 8th layer of the three-dimensional convolutional neural network composed of classic convolutional neural networks, denoted as f8. This layer consists of the classic convolutional layer in Definition 9 and the classic upsampling layer in Definition 12, denoted as C8 and U8 respectively. According to the three-dimensional convolution principle in Definition 11, the kernel size of the three-dimensional convolution of C8 is set to 3×3×8 using the classic kernel size setting method in Definition 7, and the kernel stride of the three-dimensional convolution of C8 is set to 1 using the classic kernel stride setting method in Definition 8. The upsampling layer U8 is set to upsample the image to twice its original size in the height and width directions. Using the classic CNN feature extraction method in Definition 5, the sum of the feature output A1 obtained in step 2.1 and the reconstruction result A7 obtained in step 3.3 is processed to obtain the final image reconstruction result, denoted as A8; This completes the construction of the three-dimensional convolutional neural network; Step 4: Establish an image reconstruction model The training set Train_SAR obtained in step 1 is used as input, and the classic Adam algorithm in definition 13 is used to train the three-dimensional convolutional neural network completed in step 3. After training, the image reconstruction model is obtained, which is denoted as 3D-CNN. Step 5: Test the image reconstruction model Using the test set Test_SAR obtained in step 1, the image reconstruction model 3D-CNN obtained in step 3 is tested using the standard detection network testing method in Definition 14, and the test results of the test set on the image reconstruction model are obtained, denoted as Result; Step 6: Evaluate the image reconstruction model Using the test result Result of the image reconstruction model obtained in step 5 as input, the mean squared error and structural similarity are calculated using the standard evaluation index calculation method in definition 15, and denoted as MSE and SSIM respectively. This concludes the entire method.