Pseudo-SAR Ship Adaptive Target Classification Method, Storage Medium and Computer Program Product

The generation of pseudo-high-resolution SAR ship images through RIR-GAN and SD-Net networks solves the problem of insufficient attention to SAR image reconstruction and high-frequency detailed information, and improves the accuracy and model scalability of SAR ship classification.

CN119048820BActive Publication Date: 2025-07-11SOUTHWEST JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411133649.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2025-07-11
Estimated Expiration
2044-08-19

AI Technical Summary

Technical Problem

The existing technology cannot effectively utilize the problem that the low-resolution data volume of SAR ship images is large but the information is lacking. The existing super-segment reconstruction network cannot be directly applicable to the super-resolution reconstruction of SAR images, resulting in insufficient attention to image reconstruction and high-frequency detailed information, which is difficult to adapt to the application needs of SAR ATR systems.

Method used

The nested residual connection adversarial network RIR-GAN based on SRGAN is used to generate pseudo-high-resolution SAR ship images, and feature extraction is combined with the convolutional dense connection network SD-Net. The game format optimization generator is optimized through the nested residual connection generator and discriminator, and the high-frequency detailed information is used to improve the high-frequency detail information, and a convolutional dense connection network SD-Net is built to alleviate the resolution loss problem.

Benefits of technology

It improves the recovery ability of high-frequency details of SAR images, reduces domain gaps, achieves higher model accuracy and generalization capabilities, and significantly improves the accuracy and scalability of SAR ship classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119048820B_ABST
    Figure CN119048820B_ABST
Patent Text Reader

Abstract

The present invention discloses a pseudo - SAR ship adaptive target classification method, a storage medium, and a computer program product, belonging to the technical field of artificial intelligence and synthetic aperture radar target classification, and solving the problem that common classification networks perform poorly on small - sample SAR images during the SAR ship target classification process. The present invention includes pre - processing SAR ship images in a real high - resolution SAR data set and a corresponding generated low - resolution SAR data set; constructing a nested residual connection module connecting an SRGAN feature encoder and a feature decoder and forming an RIR - GAN network with a discriminator, training the RIR - GAN network with the pre - processed SAR ship images, and using the trained RIR - GAN network to generate pseudo - high - resolution SAR images of the to - be - converted low - resolution SAR ship images, inputting the pseudo - high - resolution SAR images into a constructed convolutional densely - connected network SD - Net for feature extraction, and performing pseudo - SAR ship adaptive target classification in the pseudo - high - resolution SAR images. The present invention is used for image super - resolution reconstruction and SAR ship target classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] A method for pseudo-SAR ship adaptive target classification, a storage medium and a computer program product, which are used for image super-resolution reconstruction and SAR ship target classification, and belong to the technical field of artificial intelligence and synthetic aperture radar target classification. Background Art

[0002] SAR and optical sensors are the most commonly used earth observation devices at present, and the image data they generate are important means for realizing intelligent interpretation of ground objects. Optics has the characteristics of being massive, easy to obtain, and easy to understand the image content.

[0003] In recent years, the excellent performance of SAR in detection has led to its increasing application. Given the importance of SAR, the automatic interpretation of military and civilian targets has always been a hot topic of concern to scholars at home and abroad. Automatic target recognition (ATR) under small sample conditions is an important research direction for SAR image interpretation and has also been a hot topic of concern in recent years.

[0004] Traditional target recognition methods, such as those based on templates, models, and features, all have a common drawback: they require a large amount of manpower and material resources. Researchers with professional knowledge and rich experience can extract the target features of SAR samples one by one and perform manual interpretation. This traditional method is extremely inefficient and costly.

[0005] In recent years, object recognition technologies based on convolutional neural networks (CNNs) have been able to better mine the internal information of samples and reveal the internal connections between samples.

[0006] Deep learning is a new type of machine learning algorithm. Its core idea is to continuously learn and fit a large number of samples and learn and fit their internal laws to obtain the optimal test results. The data-driven deep learning algorithm gives full play to the advantages of deep learning in multiple fields, such as optical big data, and has shown good performance in multiple fields. Its end-to-end technology greatly reduces the manpower, material resources, and time required in the training process and can effectively improve the overall efficiency of training. There are currently various typical SAR target recognition methods based on CNN.

[0007] Although the CNN method has obtained good results in SAR target recognition, its training method still requires a large number of effective samples. In the absence of training samples, the performance of deep learning algorithms will deteriorate rapidly. Deep learning technology requires massive data and highly labeled samples, and there is an urgent need to study SAR target recognition technology in a small number of scenarios.

[0008] Few-shot learning can be roughly divided into methods based on data augmentation, methods based on model improvement, and methods based on algorithm optimization. The basic idea of data augmentation methods is sample expansion, which can be achieved by adding weakly labeled data or similar data to a small dataset. Currently, in the field of SAR target recognition, there are three common sample expansion methods: image processing-based, subband decomposition-based, and generative adversarial network (GAN)-based. First, the common image processing-based SAR sample expansion method maps the original image from the pixel level to a new coordinate position through geometric or noise methods without destroying the original image information, which is the simplest and most commonly used sample expansion method. Generally, to ensure the characteristics of SAR image samples and keep the sample size unchanged, three methods, namely translation, rotation, and noise, can be used to expand the samples. Although the image processing-based SAR sample expansion method does not destroy the effective information of the original SAR samples, it only retains most of the effective information of the original SAR samples and does not provide additional knowledge for the SAR target recognition task. Essentially, it has little effect on improving the accuracy of SAR ATR. These image-based operations may perform well on specific datasets but are not generally applicable.

[0009] Second, for high-resolution SAR images with both amplitude and phase, subband decomposition is used for sampling and denoising. The SAR system refers to tracking a target through multiple echoes over a long period of time. This method uses multiple low-resolution echo sequences of a synthetic aperture radar to obtain a high-resolution SAR echo signal, thus obtaining the echo signal of the high-resolution SAR. The idea of subband decomposition is to decompose the large-aperture radar echo signal into multiple sub-apertures, and after relevant processing, obtain the sampling points of the sub-aperture SAR to complete sampling expansion. During the subband decomposition process, it is necessary to start from the SAR image containing both amplitude and phase information.

[0010] However, most of the existing open-source SAR data only contains amplitude information, and the bandwidth of some sub-aperture signals is narrower than that of the original band, resulting in generally low resolution.

[0011] In terms of data generation, some researchers currently use data from the same domain as the target domain training dataset to train GAN to generate data similar to real samples. This method uses GAN to achieve the purpose of expanding and balancing the target domain training dataset and improves the performance of the fault diagnosis model through data augmentation. The above method achieves the purpose of expanding the data volume but does not bring additional auxiliary information to the few-shot dataset. The common data conversion method for the same domain is called super-resolution reconstruction, and its function can convert low-resolution and shallow-information image data into high-resolution and rich-information image data in the same domain.

[0012] In the field of super-resolution reconstruction, the most well-known SRGAN is based on a perceptual loss function and has achieved excellent super-resolution reconstruction results in practical scenarios. However, the above-mentioned super-resolution reconstruction algorithms based on real-world scenarios are no longer universal in SAR images.

[0013] SRGAN (Super-Resolution Generative Adversarial Network) is an ideal isomorphic transformation method from low-resolution to high-resolution (LR2HR). Existing research has found that: 1) Existing GAN super-resolution reconstructions are mostly based on real-world scenarios and there are huge differences from real-world scenarios, resulting in the current design of super-resolution reconstruction GAN networks being inapplicable to the low-resolution to high-resolution SAR image super-resolution reconstruction task we proposed.

[0014] 2) SRGAN is the current existing super-resolution reconstruction GAN, which has high stability, a relatively simple structure, and good reconstruction ability. There is still a large room for improvement in the LR2HR task. However, due to its block structure with stacked residuals, it ignores detail information during reconstruction, resulting in blurred reconstruction results. This method is very different from the high-resolution SAR images that humans can perceive.

[0015] Since super-resolution reconstruction networks such as SRGAN are all developed from real-world images, and SAR images, due to their unique imaging principles, present information such as textures and structures different from real-world images.

[0016] Therefore, the existing technologies have the following technical problems:

[0017] 1. There are many low-resolution SAR ship images, but the information is lacking; high-resolution images are rich in information but scarce in quantity. Since the advantage of a large amount of low-resolution data cannot be fully utilized, and more high-resolution data is required during the network training process to achieve the training effect, this not only increases the cost of data collection and annotation, but also limits the scalability of the model in large-scale applications;

[0018] 2. Through super-resolution reconstruction technology, the missing information in the original low-resolution images can be restored or supplemented, making the features and details of the target clearer and more accurate. In this way, more abundant and informative data can be used when training the model, improving the accuracy and generalization ability of the model; however, the currently open-source super-resolution reconstruction networks cannot be directly applied to the SAR image super-resolution reconstruction task, resulting in problems such as blurred image reconstruction and insufficient attention to high-frequency detail information, thus greatly affecting the quality of SAR image reconstruction;

[0019] 3. The current domain adaptation methods still lack the ability to deeply mine rich information when extracting features of SAR images, and at the same time, the network parameters of such methods are redundant, making it difficult to meet the application requirements of the SAR ATR system;

[0020] 4. At present, in the field of SAR ship recognition, there is a lack of intelligent classification by combining a low-resolution dataset with a quantity advantage and a domain adaptation network with the characteristics of SAR ship images. Summary of the Invention

[0021] The purpose of the present invention is to provide a pseudo-SAR ship adaptive target classification method, a storage medium, and a computer program product, to solve the problem that in the actual training process of the network, there are many SAR low-resolution images but the information is fuzzy; the currently open-source super-resolution reconstruction network cannot be directly applied to the SAR image super-resolution reconstruction task, which will cause problems such as blurred image reconstruction and insufficient attention to high-frequency detail information, thus greatly affecting the quality of SAR image reconstruction; the current domain adaptation methods still lack the ability to deeply mine rich information when extracting SAR image features, and at the same time, the network parameters of such methods are redundant, making it difficult to meet the application requirements of the SAR ATR system; the problem that in the current SAR ship recognition field, there is a lack of intelligent classification by combining a low-resolution dataset with a quantity advantage and a domain adaptation network with the characteristics of SAR ship images.

[0022] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0023] A pseudo-SAR ship adaptive target classification method includes the following steps:

[0024] S1. Preprocess the SAR ship images in the target domain high-resolution SAR dataset and the SAR ship images in the source domain low-resolution SAR dataset generated from the corresponding target domain high-resolution SAR dataset.

[0025] S2. Based on the SRGAN feature encoder and feature decoder, construct a nested residual connection module connecting the feature encoder and the feature decoder to obtain a nested residual connection generator network model, and form an RIR-GAN network with the discriminator.

[0026] S3. Input the SAR ship images in the source domain low-resolution SAR dataset and the SAR ship images in the target domain high-resolution SAR dataset obtained by preprocessing into the RIR-GAN network for training, and use the trained RIR-GAN network to generate pseudo-high-resolution SAR images of the low-resolution SAR ship images to be converted.

[0027] S4. Construct a convolutional dense connection network SD-Net to extract features from the pseudo-high-resolution SAR images to obtain a one-dimensional feature vector, and perform pseudo-SAR ship adaptive target classification in the pseudo-high-resolution SAR images by calculating the distance between the one-dimensional vector and the prototype representations of different classification ship query sets.

[0028] Furthermore, the specific steps of step S1 are as follows:

[0029] S1.1. Obtain a source domain low-resolution SAR dataset of size 64*64 through linear interpolation using the acquired high-resolution SAR ship images. Among them, the low-resolution SAR ship images in the source domain low-resolution SAR dataset are paired with the high-resolution SAR ship images, and the high-resolution SAR ship images are used as the SAR ship images in the target domain to construct a target domain high-resolution SAR dataset;

[0030] S1.2. Center-crop each SAR ship image in the acquired target domain high-resolution SAR dataset to a size of 128*128, and then use the linear interpolation method to enlarge the cropped image to a size of 256*256.

[0031] Furthermore, the nested residual connection module includes a 9*9 Conv layer, five CRISR modules, a BB bottleneck block, two upsampling layers US, and a 9*9 Conv layer connected in sequence. Among them, the output of the first 9*9 Conv layer and the output of the BB bottleneck block are added and then used as the input of the first upsampling layer US;

[0032] Each CRISR module includes a first channel extraction module, a CA_IR block, a second channel extraction module, and a spatial attention layer SALayer connected in sequence. The input of the first channel extraction module and the output of the spatial attention layer SALayer are merged in the channel dimension, and finally the output of each CRISR module is obtained;

[0033] The BB bottleneck block includes a Conv layer with a convolution kernel size of 3*3 and an instance normalization layer connected in sequence;

[0034] The upsampling layer US includes a 3*3 Conv layer, a pixel shuffle layer PixelShuffle, and an activation function layer connected in sequence.

[0035] Furthermore, the CA_IR block includes a 3*3 Conv layer, an instance normalization layer, an activation function layer, a depthwise separable convolution module, and a channel attention layer CALayer connected in sequence. The input of the 3*3 Conv layer and the output of the channel attention layer CALayer are merged in the channel dimension, and finally the output of each CA_IR block is obtained;

[0036] The first channel extraction module and the second channel extraction module include a 1*1 Conv layer, an instance normalization layer, and an activation function layer connected in sequence;

[0037] The separable convolution module includes a depthwise convolution layer Conv_C, an instance normalization layer, an activation function layer, a pointwise convolution layer Conv_P, and an instance normalization layer connected in sequence;

[0038] The channel attention layer CALayer includes a parallel average pooling layer and max pooling layer, a multi-layer perceptron with multiple hidden layers connected to the average pooling layer and max pooling layer respectively, and uses element-wise summation to combine the output feature vectors of the multi-layer perceptron, that is, to obtain the channel attention map, and the final output layer outputs the result of multiplying the input feature map by the channel attention map;

[0039] The spatial attention layer SALayer includes a max pooling layer and an average pooling layer connected in sequence, a convolutional layer that concatenates the output of the average pooling layer, and an output layer that encodes the output of the convolutional layer to obtain the spatial attention map and outputs the result of multiplying the spatial attention map by the input feature map element-wise.

[0040] Furthermore, the game form expression of the nested residual connection generator network model and discriminator in the RIR-GAN network is:

[0041]

[0042] Among them, represents the loss function of the generator, and the goal is to optimize the generator by minimizing this loss so that the generated high-resolution SAR ship image is closer to the real high-resolution SAR ship image. The generator has parameters , and its input is the low-resolution SAR ship image , and the output is the generated high-resolution SAR ship image . The discriminator has parameters , and its input is the generated high-resolution SAR ship image , and the output is the probability that this image is a "real" image. The generator hopes to maximize the probability given by the discriminator so that the discriminator believes that the generated high-resolution image is "real". represents the number of input low-resolution SAR ship images .

[0043] Furthermore, the convolutional densely connected network SD-Net includes a downsampling module DS, a first densely connected module DM, a first residual separable convolution module SCR, a second densely connected module DM, a second residual separable convolution module SCR, a third densely connected module DM, a third residual separable convolution module SCR, a fourth densely connected module DM, and a pooling module connected in sequence;

[0044] The downsampling module DS includes a Conv2 convolutional layer, a batch normalization layer, an activation function layer, and a max pooling layer that are connected in sequence. Among them, the Conv2 layer represents a general convolution operation with a convolution kernel size of 7*7 and a stride of 2;

[0045] The first dense connection module DM, the second dense connection module DM, the third dense connection module DM, and the fourth dense connection module DM each include 6, 12, 24, and 16 connected dense connection modules DL. Each dense connection module DL includes a batch normalization layer, an activation function layer, a first Conv1 convolutional layer, a batch normalization layer, an activation function layer, and a second Conv1 convolutional layer that are connected in sequence. The concatenation operation result of the input of the first batch normalization layer and the output of the second Conv1 convolutional layer in each dense connection module DL is used as the final output of each dense connection module DL. Among them, the first Conv1 convolutional layer represents a general convolution operation with a convolution kernel of 1 and a stride of 1, and the second Conv1 convolutional layer represents a general convolution operation with a convolution kernel of 3 and a stride of 1;

[0046] The first residual separation convolution module SCR, the second residual separation convolution module SCR, and the third residual separation convolution module SCR include a separable convolution module S-Conv with a stride of 2, and a feature extraction pooling module connected to the separable convolution module S-Conv through a residual connection;

[0047] The separable convolution module S-Conv includes a batch normalization layer, an activation function layer, a depthwise convolution layer Conv_C, a batch normalization layer, an activation function layer, and a pointwise convolution layer Conv_P that are connected in sequence;

[0048] The feature extraction pooling module includes a batch normalization layer, an activation function layer, a Conv1 convolutional layer, and a max pooling layer that are connected in sequence. Among them, the Conv1 layer represents a general Conv layer with a convolution kernel of 1 and a stride of 1;

[0049] The pooling module includes a batch normalization layer, an activation function layer, an average pooling layer, and a Flatten operation layer that are connected in sequence.

[0050] A pseudo-SAR ship adaptive target classification system includes a computer program that implements a pseudo-SAR ship adaptive target classification method.

[0051] A computer-readable storage medium stores a computer program that implements a pseudo-SAR ship adaptive target classification method.

[0052] A computer program product includes a computer program that implements a pseudo-SAR ship adaptive target classification method. Compared with the prior art, the advantages of the present invention are as follows:

[0053] Based on SRGAN, the present invention proposes a nested residual connection adversarial network RIR-GAN to generate pseudo high-resolution SAR ship images. Among them, the CRISR module is the core of its improvement. The CA_IR module is designed in the SA_R module of the CRISR module. Based on the SAR super-resolution reconstruction method of SRGAN, the high-frequency detail information of the SAR image is improved, which is specifically reflected in:

[0054] First, in the case where the number of low-resolution SAR ship images is large but the information is lacking, while the high-resolution images are rich in information but scarce in quantity, the present invention can make full use of the advantage of a large amount of low-resolution data. During the network training process, it is not necessary to use more high-resolution data to achieve the training effect, which not only increases the cost of data collection and annotation but also improves the scalability of the model in large-scale applications.

[0055] Second, the features of the pseudo high-resolution SAR images generated by the RIR-GAN network in the present invention are highly confused with the features of real high-resolution SAR ship images (that is, the information missing in the original low-resolution images can be restored or supplemented, so that the features and details of the target are more clear and accurate. When training the model, more abundant and informative data can be used to improve the accuracy and generalization ability of the model), significantly reducing the domain gap and effectively realizing the super-resolution reconstruction task, that is, effectively avoiding the problems of blurred image reconstruction and insufficient attention to high-frequency detail information.

[0056] Third, the separable convolutional densely connected network SD-Net in the present invention has fewer parameters with a slight increase in computational complexity, and has a high feature reuse rate, alleviating the resolution loss problem during downsampling of feature maps in the pooling layer. In comparison with other common popular methods, this backbone network performs better than general feature extraction networks.

[0057] Fourth, the small-sample target classification method based on the RIR-GAN network and the separable convolutional densely connected network SD-Net in the present invention has excellent performance in the field of SAR ship classification. Compared with popular small-sample classification methods, the classification accuracy has been improved. Description of the Drawings

[0058] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0059] Figure 1 is the overall framework diagram of the present invention. Among them, in the RIR-GAN generator, Conv represents the convolution operation, CRISR represents the CRISR module in the nested residual connection generator network model, IN_BB represents the BB bottleneck block with an InstanceNorm2d (IN) layer, US represents the upsampling layer using the PixelShuffle method. Secondly, in the feature extraction process, the real classification dataset and the pseudo-SAR dataset are mixed to add auxiliary knowledge to form a new GAN dataset; in the SD-Net, DS represents the downsampling module, DM represents the dense connection module, SCR represents the residual separation convolution module with a downsampling function, AP represents the average pooling layer, F represents the Flatten operation to flatten the feature map into a one-dimensional feature vector. Finally, in the domain adaptation process, the prototype network is selected as the domain adaptation method. In the prototype network, represents the one-dimensional feature vector representation of each known class sample after feature extraction, represents the summation operation, Mean represents the averaging operation, and the result obtained after summation and averaging will be used as the prototype representation of the known class result. The sample to be classified obtains a one-dimensional feature vector after feature extraction, calculates the distance between this feature vector and the prototype representations representing each known class, and converts the distance into a probability form through Softmax to determine the class to which each ship sample in the image to be classified belongs;

[0060] Figure 2 is the structural diagram of the nested residual connection generator network model in the RIR-GAN network of the present invention. Among them, IN represents instance normalization, ReLU6 represents the activation function layer, Conv_C represents the per-channel convolution layer, and Conv_P represents the pointwise convolution layer;

[0061] Figure 3 is the structural diagram of the SAlayer in the CRISR of the RIR-GAN of the present invention;

[0062] Figure 4 is the structural diagram of the CAlayer in the CRISR of the RIR-GAN of the present invention;

[0063] Figure 5It is the structural diagram of the Convolutional Dense Connection Network SD-Net used for feature extraction in the classification network part of the present invention. Among BN / ReLU6 / Conv1 / BN / ReLU6 / Conv1, BN represents the batch normalization layer, ReLU6 represents the activation function layer, the first Conv1 represents the first Conv1 convolutional layer, which is a normal convolutional operation with a convolution kernel of 1 and a stride of 1. The second Conv1 represents the second Conv1 convolutional layer, representing a normal convolutional operation with a convolution kernel of 3 and a stride of 1. Conv_C represents the channel-wise convolutional layer, Conv_P represents the point-wise convolutional layer, MaxPooling represents the max pooling layer, and Avgpooling represents the average pooling layer;

[0064] Figure 6 It is the schematic diagram of the FID and KID values of the ablation experiment of the RIR-GAN generative adversarial network in the present invention. By calculating the distance of the multivariate normal distribution, FID compares the generated data with the training data at the feature level. KID measures the difference between two sets of samples by calculating the square of the maximum mean difference between Inception representations. The smaller the FID and KID values, the better the reconstruction effect. Among them, C_BB_D means that only the feature conversion module RB in SRGAN is replaced by the CRISR module of the present invention, C_IB_D means that the bottleneck layer replaces BN normalization with IN normalization to avoid the co-use of BN and IN, and C_BB_IND means that the discriminator is also replaced by the IN normalization method to improve the image recognition performance of the discriminator;

[0065] Figure 7 It is the comparison of the super-resolution reconstruction images of SRGAN and RIR-GAN obtained from the high-resolution original images (HR original images) and bicubic interpolation images (Bicubic linear interpolation) during the model training of the present invention at epoch1, epoch5, and epoch100 (where 𝑒 represents epoch, that is, the number of network training times). In addition, three comparison diagrams of the details of the reconstructed images and the high-resolution original images are given;

[0066] Figure 8 It is the parameter comparison between SD-Net and other popular networks in the present invention, intuitively showing the comparison of the parameters, calculations, and memory usage of each backbone network.

[0067] Figure 9 It is the comparison result of the classification accuracy presented by the present invention under the small-sample condition and the classification accuracies of the popular small-sample classification methods RelationNet and Meta DeepBDC. Among them, 3way-1shot means that the training set and the classification set contain three categories, and each category contains only 1 image; 3way-5shot means that the training set and the classification set contain three categories, and each category contains only 5 images. Detailed implementation mode

[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0069] For the synthetic aperture radar of active microwave remote sensing, this case intends to expand the design idea of traditional radar by using pulse compression and synthetic aperture theory, realize high-resolution and high-resolution imaging of targets, and extract information such as the amplitude and phase of the targets after their interaction with environmental elements. A remote sensing image sample was obtained. However, due to its insensitivity to factors such as meteorology, natural imaging mechanism, and ground object scattering mechanism, the content analysis and marking of SAR images mostly rely on expert experience and require huge human and material resources, making it difficult to meet the actual application requirements.

[0070] As Figure 1 and Figure 2 shown, the present invention provides a pseudo-SAR ship adaptive target classification method, including the following steps:

[0071] S1. Preprocess the SAR ship images in the high-resolution SAR dataset of the target domain and the SAR ship images in the source domain low-resolution SAR dataset generated corresponding to the high-resolution SAR dataset of the target domain. The specific steps are as follows:

[0072] S1.1. Use the obtained high-resolution SAR ship images to obtain a source domain low-resolution SAR dataset of 64*64 size through linear interpolation (here refers to the process required for training the network, that is, the pre-training process. The low-resolution images must be obtained using real high-resolution images because the network needs to learn the paired high-resolution and low-resolution images to learn the relationship between the two. After the network learns this relationship mapping, in the application, only the real low-resolution image needs to be input to generate a pseudo-high-resolution image and perform adaptive classification). Among them, the low-resolution SAR ship images in the source domain low-resolution SAR dataset are paired with the high-resolution SAR ship images, and the high-resolution SAR ship images are used as the SAR ship images in the target domain to construct the high-resolution SAR dataset of the target domain;

[0073] S1.2. Crop the center of each SAR ship image in the obtained high-resolution SAR dataset of the target domain to a size of 128*128, so as to reduce the difference in the proportion of the main body of the ship in the images between the low-resolution dataset of the source domain and the high-resolution SAR dataset of the target domain, and then use the linear interpolation method to enlarge the cropped image to a size of 256*256.

[0074] S2. Keep the original feature encoder and decoder of SRGAN unchanged, construct a nested residual connection module, and generate a nested residual connection generator network model; that is, based on the SRGAN feature encoder and feature decoder, construct a nested residual connection module connecting the feature encoder and feature decoder to obtain a nested residual connection generator network model, and form an RIR-GAN network with the discriminator; the generator structure network model of RIR-GAN is as Figure 2 shown.

[0075] The nested residual connection module includes a 9*9 Conv layer (a 9*9 ordinary convolutional layer) connected in sequence, five CRISR modules, a BB bottleneck block, two upsampling layers US, and a 9*9 Conv layer (a 9*9 ordinary convolutional layer). Among them, the output of the first 9*9 Conv layer and the output of the BB bottleneck block are added and then used as the input of the first upsampling layer US; each CRISR module includes a first channel extraction module, a CA_IR block, a second channel extraction module, and a spatial attention layer SALayer connected in sequence. The input of the first channel extraction module and the output of the spatial attention layer SALayer are merged in the channel dimension to finally obtain the output of each CRISR module; the BB bottleneck block includes a Conv layer with a convolution kernel size of 3*3 (a 3*3 ordinary convolutional layer) and an instance normalization layer connected in sequence; the upsampling layer US includes a 3*3 Conv layer (a 3*3 ordinary convolutional layer), a pixel shuffle layer PixelShuffle, and an activation function layer connected in sequence. In the structural description of the present invention, although some layers or modules have the same description, their connection positions are different, and the received inputs and outputs are also different.

[0076] The CA_IR block includes a 3*3 Conv layer, an instance normalization layer, an activation function layer, a separable convolution module, and a channel attention layer CALayer connected in sequence. The input of the 3*3 Conv layer and the output of the channel attention layer CALayer are merged in the channel dimension to finally obtain the output of each CA_IR block; the channel extraction module includes a 1*1 Conv layer, an instance normalization layer, and an activation function layer connected in sequence; the separable convolution module includes a depthwise convolution layer Conv_C, an instance normalization layer, an activation function layer, a pointwise convolution layer Conv_P, and an instance normalization layer connected in sequence; as Figure 4As shown, the channel attention layer CALayer includes a parallel average pooling layer and max pooling layer, a multi-layer perceptron with multiple hidden layers connected to the average pooling layer and max pooling layer respectively, and the output results of the multi-layer perceptron are merged using element-wise summation to obtain the channel attention map. Finally, the output layer outputs the result of the dot product of the input feature map and the channel attention map, as Figure 3 As shown, the spatial attention layer SALayer includes a max pooling layer and an average pooling layer connected in sequence, a convolutional layer that concatenates the outputs of the average pooling layer, and an output layer that encodes the output of the convolutional layer to obtain the spatial attention map and outputs the result of the element-wise multiplication of the spatial attention map and the input feature map.

[0077] The CRISR module is a nested residual module designed in the present invention with the IN method. This module designs a channel attention inverted residual block in the spatial attention residual path, which solves the problem of insufficient attention to high-frequency detail information in the SRGAN for SAR ship super-resolution reconstruction tasks. The BB layer consists of a normal convolutional layer with a convolutional kernel size of 33 and an IN layer. Subsequently, the feature map before the first CRISR module will be added to the feature map after passing through the BB bottleneck layer, that is, the output of its first 9*9 Conv layer is added to the output of the BB bottleneck block and used as the input of the first upsampling layer US to achieve maximum retention of low-level information in a skip connection manner. At this time, the number of feature maps remains 646464. Then, the feature map passes through two upsampling layers to achieve 4-fold super-resolution reconstruction of the LR image. Finally, the feature map reduces the number of channels in a normal convolutional layer with a convolutional kernel size of 9*9, and outputs a pseudo high-resolution image of 325*256.

[0078] The nested residual connection module designs CA_IR in the CRISR module. To cooperate with the super-resolution reconstruction task, it uses the normalization layer in the CRISR module to perform instance normalization on the data.

[0079] In the first channel extraction module in the CRISR module, a 1*1 Conv, a normalization layer, and an activation function layer are used to expand the input channel number by 2 times to achieve channel number expansion, so as to pay attention to the channel information in the subsequent CA_IR block. Then, a 1*1 Conv, a normalization layer, and an activation function layer in the second channel extraction module are used to reduce the channel number to 1 / 2, and the focus of position information is achieved in the subsequent spatial attention layer (SALayer). In the CA_IR block, a 3*3 Conv layer, a normalization layer, an activation function layer, a separable convolution module, and a channel attention layer CALayer are used to expand the number of feature channels to 𝑛 times the original. In this experiment, 𝑛 is set to 2.

[0080] During the inter-channel convolution process, the number of channels remains unchanged, and the number of parameters and the computational amount are reduced through grouped convolution. Pointwise convolution reduces the number of channels to 1 / 2 and then restores the number of channels. Finally, the channel attention layer CAlayer is connected to achieve attention to channel information. Subsequently, the channel extraction module is continued to reduce the number of channels to 1 / 2, and attention to position information is achieved in the subsequent spatial attention layer SALayer.

[0081] In the channel attention layer CALayer, both average pooling and max pooling operations are used to aggregate the channel information of the feature map to improve the representation ability of the network. The generated average pooling feature and max pooling feature are forwarded to a shared network composed of a multi-layer perceptron with multiple hidden layers to generate a channel attention map , and finally element-wise summation is used to merge the output feature vectors, that is, the channel attention map is obtained Element-wise multiplication with the input feature map results in the output feature .

[0082] In the spatial attention layer SALayer, the spatial relationship of the features is used to generate a spatial attention map. This layer first applies average pooling and max pooling along the channel axis direction to effectively highlight the information region of the input feature map , then concatenates the feature descriptions of the two, and applies a convolutional layer to generate a spatial attention map , and encodes the emphasized or suppressed regions to obtain the spatial attention map Element-wise multiplication with the feature results in the output feature .

[0083] The CRISR module adaptively refines the intermediate feature map extracted by the feature transformer by successively applying channel and position attention mechanisms. While realizing the function of locating the main features of the ship, it maintains the number of network parameters of the nested residual connection and does not bring a burden of a large increase in training time to the network model.

[0084] The expression of the game form of the nested residual connection generator network model and the discriminator in the RIR-GAN network is:

[0085]

[0086] Among them, represents the loss function of the generator. The goal is to optimize the generator by minimizing this loss, so that the generated high-resolution SAR ship image is closer to the real high-resolution SAR ship image. The generator has parameters , and its input is a low-resolution SAR ship image , the output is the generated high-resolution SAR ship image , discriminator with parameters , its input is the generated high-resolution SAR ship image , the output is the probability that the image is a "real" image. The generator hopes to maximize the probability given by the discriminator such that the discriminator believes that the generated high-resolution image is "real". denotes the low-resolution SAR ship image of the input quantity.

[0087] S3. Simultaneously input the SAR ship images in the preprocessed source-domain low-resolution SAR dataset and the SAR ship images in the target-domain high-resolution SAR dataset into the RIR-GAN network for training, and use the trained RIR-GAN network to generate the pseudo-high-resolution SAR images of the low-resolution SAR ship images to be converted.

[0088] The implementation logic of the training steps of the RIR-GAN network is as follows:

[0089] The SAR ship images in the preprocessed source-domain low-resolution dataset and the SAR ship images in the target-domain high-resolution SAR dataset are simultaneously input to train the nested residual connection generator network model. During the training process, the confrontation between the corresponding nested residual connection generator network model and the discriminator will indirectly improve the performance of the generator. Among them, the specific process after the image is input into the generator includes: downsampling through the feature encoder - feature transformation through the nested residual connection module - pixel reconstruction through the feature decoder.

[0090] The generation steps of the pseudo-high-resolution SAR images are as follows:

[0091] Using the trained generator, input the LR image (low-resolution SAR ship image) to be converted to generate the pseudo-high-resolution SAR image (pseudo-SAR ship image). The specific process is: the LR image is input into the feature encoder for downsampling - feature transformation through the nested residual connection converter - pixel reconstruction through the feature decoder.

[0092] Obtain the pseudo-high-resolution SAR images to expand the target-domain SAR dataset, and train the SD-Net based on the expanded target-domain SAR dataset for the SAR ship target classification in the SAR ship images to be classified.

[0093] S4. Construct a convolutional densely connected network SD-Net to extract features from the pseudo-high-resolution SAR images, obtaining a one-dimensional feature vector. Adaptive target classification of pseudo-SAR ships in the pseudo-high-resolution SAR images is performed by calculating the distances between this one-dimensional vector and the prototype representations of ship query sets of different classifications. In this backbone network, the convolutional densely connected network SD-Net has a residual separation convolutional module SCR with a bottleneck structure to match the feature reuse characteristics of the densely connected structure and alleviate the resolution loss problem when downsampling feature maps in the pooling layer. The 1*1 convolutional operation in the residual structure preserves more details and improves the feature reuse rate through dense connections, thereby enhancing the feature extraction performance of the network. The number of parameters of the convolutional densely connected network SD-Net is significantly reduced, and the computational amount and memory consumption increase less. The convolutional densely connected network SD-Net is as Figure 5 shown.

[0094] The convolutional dense connection network SD-Net includes a downsampling module DS, a first dense connection module DM, a first residual separable convolution module SCR, a second dense connection module DM, a second residual separable convolution module SCR, a third dense connection module DM, a third residual separable convolution module SCR, a fourth dense connection module DM, and a pooling module, which are connected in sequence. The downsampling module DS includes a Conv2 convolutional layer, a batch normalization layer, an activation function layer, and a max pooling layer, which are connected in sequence. Among them, the Conv2 layer represents a normal convolution operation with a convolution kernel size of 7*7 and a stride of 2. The first dense connection module DM, the second dense connection module DM, the third dense connection module DM, and the fourth dense connection module DM each include 6, 12, 24, and 16 connected dense connection modules DL. Each dense connection module DL includes a batch normalization layer, an activation function layer, a first Conv1 convolutional layer, a batch normalization layer, an activation function layer, and a second Conv1 convolutional layer, which are connected in sequence. The input of the first batch normalization layer in each dense connection module DL and the output of the second Conv1 convolutional layer are concatenated as the final output of each dense connection module DL. Among them, the first Conv1 convolutional layer represents a normal convolution operation with a convolution kernel of 1 and a stride of 1, and the second Conv1 convolutional layer represents a normal convolution operation with a convolution kernel of 3 and a stride of 1. DL does not have the ability to change the size of the feature map. The first residual separable convolution module SCR, the second residual separable convolution module SCR, and the third residual separable convolution module SCR include a separable convolution module S-Conv with a stride of 2 and a feature extraction pooling module connected to the separable convolution module S-Conv through a residual connection. The separable convolution module S-Conv includes a batch normalization layer, an activation function layer, a depthwise convolution layer Conv_C, a batch normalization layer, an activation function layer, and a pointwise convolution layer Conv_P. The feature extraction pooling module includes a batch normalization layer, an activation function layer, a Conv1 convolutional layer, and a max pooling layer, which are connected in sequence. Among them, the Conv1 layer represents a normal Conv layer with a convolution kernel of 1 and a stride of 1. The pooling module includes a normalization layer, an activation function layer, an average pooling layer, and a Flatten operation layer.

[0095] For an input image of size 3×256×256, after entering the convolutional densely connected network SD-Net feature extraction network: First, the image will be downsampled twice through the downsampling module DS. The size of the feature map changes from 256×256 to 64×64, and the number of channels changes from the original 3 channels to 64 channels. The Conv2 convolutional layer in the downsampling module represents a common convolutional operation with a kernel size of 7×7 and a stride of 2. Through this layer, the original input image is downsampled to a size of 32×128×128. The max pooling layer further downsamples the image to a size of 64×64×64, achieving fast extraction of image features and rapid reduction of image size. Then, the feature map is input into a cross-linked network structure composed of 4 densely connected modules DM and 3 residual separable convolutional modules SCR to efficiently extract features. Every time passing through a DL layer, the output feature map of this layer is concatenated with the input feature map, connecting the two feature maps in the channel dimension, enabling subsequent layers to access the feature information of previous layers simultaneously. Therefore, the DL layer does not have the ability to change the size of the feature map. Each time the input feature map passes through DL, the result obtained from this layer is combined with the input feature map to achieve feature reuse, improve the feature reuse rate, and reduce feature redundancy.

[0096] Using the separable convolution S-Conv model of depthwise separable convolution to obtain better feature extraction performance. Finally, a feature map with a size of 1024×8×8 is obtained.

[0097] Finally, after passing through the batch normalization layer and activation function in the pooling module, the feature map is transformed into a feature map with a size of 1024×1×1 through the average pooling layer with a kernel size of 8, and is flattened into a one-dimensional feature vector through the Flatten operation of the Flatten operation layer. Thus, the feature extraction and one-dimensional vector generation process of the convolutional densely connected network SD-Net is completed. Adaptive target classification is achieved by calculating the distance between this one-dimensional vector and the prototype representations of different classifications.

[0098] In this embodiment, in order to illustrate the superiority of the generative adversarial network RIR-GAN network of this case, ablation experiments were first conducted, and the experimental results are shown as Figure 6 shown. It can be seen that the RIR-GAN network has the lowest FID and KID in various categories and the best translation performance. Figure 7 It is a comparison of the super-resolution reconstruction images of SRGAN and RIR-GAN obtained from the HR original images and bicubic interpolation images at epoch1, epoch5, and epoch100 during model training. In addition, three comparison graphs of the details of the reconstructed images and the original HR images are given.

[0099] To further verify the effectiveness of the present invention in the classification task of generating high-resolution images HR from low-resolution images LR of ships during the LR2HR process, we compared the computational cost, number of parameters, memory occupancy, and recognition performance of the convolutional dense connection network SD-Net with the baseline network DenseNet121, and used the accuracy metric to evaluate the network's feature extraction and recognition performance. Figure 8 Intuitively shows the comparison of the parameters, computations, and memory usage of each backbone network. From the statistical graphs of the number of parameters and computational cost, the convolutional dense connection network SD-Net has relatively low values among popular backbone networks. Although the memory occupancy of this model is larger than that of other networks, overall, it is not very different from or is superior to currently popular feature extraction networks. In summary, the convolutional dense connection network SD-Net has a relatively excellent model design, but whether a model is excellent depends on its performance in actual tasks.

[0100] Finally, to verify the rationality and effectiveness of the pseudo-SAR domain generated based on the ship LR2HR transfer task in solving the problems of scarce effectively labeled samples and class imbalance in the SAR ship ATR deep learning network model, the ability of the pseudo-SAR images to improve the classification accuracy of popular ship classification networks was tested. Compared with three popular few-shot methods, an improvement in accuracy was achieved under the premise of the same parameter settings, demonstrating the rationality, effectiveness, and application value of the pseudo-SAR domain generated by the present invention in solving the problems of scarce effectively labeled samples and class imbalance in the SAR ship ATR deep learning network model.

[0101] The present invention discloses a new network RIR-GAN more suitable for the ship LR2HR transfer task based on the well-known image super-resolution reconstruction network RIR-GAN. The problem of insufficient attention to high-frequency detail information in super-resolution reconstruction is solved using CRISR design. Finally, according to the preset training parameters and loss function, the RIR-GAN network is trained, and the trained RIR-GAN network is used to generate pseudo-SAR images to construct a pseudo high-resolution SAR domain to drive the SAR ship target classification task. This network can demonstrate the most superior performance in the ship LR2HR transfer task with the lowest FID value and KID value. At the same time, we constructed a separate convolutional dense connection network SD-Net. The S-Conv structure has the function of reducing the number of network parameters, and the core-designed SCR module has the function of reducing the resolution loss during the feature map pooling subsampling process and can retain more details. Finally, the few-shot target classification method formed based on these two network models is highly reasonable and has application value for solving the problems of scarce effectively labeled samples and class imbalance in the SAR ship ATR deep learning network model.

Claims

1. A pseudo-SAR ship adaptive target classification method, characterized in that It includes the following steps: S1. Preprocess the SAR ship images in the target domain high-resolution SAR dataset and the SAR ship images in the source domain low-resolution SAR dataset generated from the corresponding target domain high-resolution SAR dataset; S2. Based on the SRGAN feature encoder and feature decoder, construct a nested residual connection module connecting the feature encoder and the feature decoder to obtain a nested residual connection generator network model, and form an RIR-GAN network with the discriminator; The nested residual connection module includes a 9*9 Conv layer, five CRISR modules, a BB bottleneck block, two upsampling layers US, and a 9*9 Conv layer connected in sequence. Among them, the output of the first 9*9 Conv layer and the output of the BB bottleneck block are added and used as the input of the first upsampling layer US; Each CRISR module includes a first channel extraction module, a CA_IR block, a second channel extraction module, and a spatial attention layer SALayer connected in sequence. The input of the first channel extraction module and the output of the spatial attention layer SALayer are merged in the channel dimension, and finally the output of each CRISR module is obtained; The BB bottleneck block includes a Conv layer with a convolution kernel size of 3*3 and an instance normalization layer connected in sequence; The upsampling layer US includes a 3*3 Conv layer, a pixel shuffle layer PixelShuffle, and an activation function layer connected in sequence; The CA_IR block includes a 3*3 Conv layer, an instance normalization layer, an activation function layer, a separable convolution module, and a channel attention layer CALayer connected in sequence. The input of the 3*3 Conv layer and the output of the channel attention layer CALayer are merged in the channel dimension, and finally the output of each CA_IR block is obtained; S3. Input the SAR ship images in the source domain low-resolution SAR dataset and the SAR ship images in the target domain high-resolution SAR dataset obtained by preprocessing into the RIR-GAN network for training, and use the trained RIR-GAN network to generate a pseudo-high-resolution SAR image of the low-resolution SAR ship image to be converted; S4. Construct a convolutional dense connection network SD-Net to extract features from the pseudo-high-resolution SAR image to obtain a one-dimensional feature vector, and perform adaptive target classification of pseudo-SAR ships in the pseudo-high-resolution SAR image by calculating the distance between the one-dimensional feature vector and the prototype representations of different classified ship query sets.

2. The adaptive target classification method for ships based on pseudo-SAR according to claim 1, wherein The specific steps of step S1 are as follows: S1.

1. Use the obtained high-resolution SAR ship images to obtain a source domain low-resolution SAR dataset of size 64*64 through linear interpolation. Among them, the low-resolution SAR ship images in the source domain low-resolution SAR dataset are paired with the high-resolution SAR ship images, and the high-resolution SAR ship images are used as the SAR ship images in the target domain to construct a target domain high-resolution SAR dataset; S1.

2. Crop the center of each SAR ship image in the obtained high-resolution SAR dataset of the target domain to a size of 128*128, and then use the linear interpolation method to enlarge the cropped image to a size of 256*256.

3. The method for pseudo-SAR ship adaptive target classification according to claim 1, wherein: The first channel extraction module and the second channel extraction module include a 1*1 Conv layer, an instance normalization layer, and an activation function layer connected in sequence. The separable convolution module includes a depthwise convolution layer Conv_C, an instance normalization layer, an activation function layer, a pointwise convolution layer Conv_P, and an instance normalization layer connected in sequence. The channel attention layer CALayer includes a parallel average pooling layer and max pooling layer, a multi-layer perceptron with multiple hidden layers connected to the average pooling layer and max pooling layer respectively, and the output results of the multi-layer perceptron are combined using element-wise summation to obtain the channel attention map, and the final output layer outputs the result of multiplying the input feature map by the channel attention map. The spatial attention layer SALayer includes a max pooling layer and an average pooling layer connected in sequence, a convolution layer that concatenates the outputs of the average pooling layer, and an output layer that encodes the output of the convolution layer to obtain the spatial attention map and outputs the result of multiplying the spatial attention map by the input feature map element-wise.

4. The pseudo-SAR ship adaptive target classification method according to claim 3, wherein The game form expression of the nested residual connection generator network model and discriminator in the RIR-GAN network is: Among them, represents the loss function of the generator. The goal is to optimize the generator by minimizing this loss so that the generated high-resolution SAR ship images are closer to the real high-resolution SAR ship images. The generator has parameters , and its input is the low-resolution SAR ship image , and the output is the generated high-resolution SAR ship image . The discriminator has parameters , and its input is the generated high-resolution SAR ship image , and the output is the probability that the image is a "real" image. The generator hopes to maximize the probability given by the discriminator so that the discriminator believes that the generated high-resolution image is "real". represents the number of input low-resolution SAR ship images .

5. A pseudo-SAR ship self-adaptive target classification method according to claim 1, characterized in that The convolutional dense connection network SD-Net includes a downsampling module DS, a first dense connection module DM, a first residual separable convolution module SCR, a second dense connection module DM, a second residual separable convolution module SCR, a third dense connection module DM, a third residual separable convolution module SCR, a fourth dense connection module DM, and a pooling module connected in sequence. The downsampling module DS includes a Conv2 convolution layer, a batch normalization layer, an activation function layer, and a max pooling layer connected in sequence, where the Conv2 layer represents a normal convolution operation with a convolution kernel size of 7*7 and a stride of 2. The first dense connection module DM, the second dense connection module DM, the third dense connection module DM, and the fourth dense connection module DM each include 6, 12, 24, and 16 connected dense connection modules DL. Each dense connection module DL includes a batch normalization layer, an activation function layer, a first Conv1 convolution layer, a batch normalization layer, an activation function layer, and a second Conv1 convolution layer connected in sequence. The concatenation operation result of the input of the first batch normalization layer and the output of the second Conv1 convolution layer in each dense connection module DL is used as the final output of each dense connection module DL. Among them, the first Conv1 convolution layer represents a normal convolution operation with a convolution kernel of 1 and a stride of 1, and the second Conv1 convolution layer represents a normal convolution operation with a convolution kernel of 3 and a stride of 1. The first residual separable convolution module SCR, the second residual separable convolution module SCR, and the third residual separable convolution module SCR include a separable convolution module S-Conv with a stride of 2 and a feature extraction pooling module connected to the separable convolution module S-Conv through a residual connection; The separable convolution module S-Conv includes a batch normalization layer, an activation function layer, a depthwise convolution layer Conv_C, a batch normalization layer, an activation function layer, and a pointwise convolution layer Conv_P connected in sequence; The feature extraction pooling module includes a batch normalization layer, an activation function layer, a Conv1 convolution layer, and a max pooling layer connected in sequence, where the Conv1 layer represents a normal Conv layer with a kernel size of 1 and a stride of 1; The pooling module includes a batch normalization layer, an activation function layer, an average pooling layer, and a Flatten operation layer connected in sequence.

6. A computer-readable storage medium, characterized in that: A computer program for implementing the method according to any one of claims 1-5 is stored.

7. A computer program product, characterized in that: It includes a computer program for implementing the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Image super-resolution method based on improved generative adversarial network

    CN114463181A