A method, device, equipment, medium and product for ship target recognition
Through the S3GAN network model, the SAR ship image is converted from the amplitude domain to the spectrum domain and feature fusion recognition is performed, which solves the problem of the lack of spectrum information in the SAR ship image dataset and achieves high-precision ship target recognition.
Patent Information
- Application Number
- CN202411237492.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-09-05
AI Technical Summary
In the prior art, SAR ship image data set lacks spectrum information, resulting in low ship target recognition accuracy, especially in small samples, it is difficult to achieve high-precision recognition.
The S3GAN network model is used to convert the amplitude domain to the spectrum domain. By generating the feature encoder of the adversarial network U-GAT-IT, the spectrum transformation expansion module and the multi-scale channel comprehensive processing module are added before the feature encoder of the adversarial network U-GAT-IT, the amplitude ship image is converted into a 128-channel image rich in deep features, and the feature fusion recognition is performed based on the spectrum information.
It improves the accuracy of ship target recognition, makes full use of the rich and complex physical scattering information of SAR ship images, and improves the recognition effect under small sample conditions.
Smart Images

Figure CN119206638B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and synthetic aperture radar target recognition, and particularly to a ship target recognition method, device, equipment, medium and product. Background Art
[0002] In ship target recognition, synthetic aperture radar (SAR) is more suitable for maritime target recognition due to its characteristics such as all-weather observation and strong penetration. In recent years, with the highly development of deep learning, deep neural networks have shown extremely excellent performance in target detection, target recognition, and image segmentation. Currently, it is very popular and effective to use SAR images and deep learning methods for ship target recognition in the field of ship automatic target recognition. However, in deep learning methods, if the dataset is small in scale, imbalanced in categories, and poor in representativeness, the effect of target recognition is unsatisfactory. Unfortunately, due to reasons such as the difficulty of obtaining SAR images, the high cost of manual annotation, and the military restrictions on ship targets, the sample size of SAR ship images is small and the categories are imbalanced. Therefore, how to achieve high-precision SAR ship target recognition under the premise of small samples is very important.
[0003] The main development directions of small sample target recognition are mainly divided into two categories:
[0004] 1) Improvement of the algorithm structure.
[0005] 2) Improvement of the dataset itself.
[0006] The improvement of the algorithm structure mainly uses methods such as transfer learning and meta-learning. Essentially, they all apply the prior knowledge learned from large-scale datasets to small sample datasets, and do not really process the small sample datasets themselves; the improvement of the dataset itself is mainly through ways such as expanding the data scale, improving the data quality, and enhancing the supervision information. Data augmentation methods such as rotating, scaling, and cropping the original data are used to expand the original data scale, which was also the most popular data augmentation method in previous years.
[0007] In recent years, with the popularity of generative models, expanding the sample scale by sampling and generating samples through generative models is also a common data processing method. Since the resolution of currently public SAR datasets is generally low, super-resolution reconstruction of them to improve their data quality is also a method for improving the dataset.
[0008] However, these methods are all improvement methods for SAR amplitude single-channel grayscale images and do not utilize the rich spectral information of SAR images. Due to the imaging mechanism of SAR images being different from that of optical images, even two completely different targets may be extremely similar in the amplitude image. Simply using the SAR amplitude image for ship target recognition is not completely accurate, incomplete, and even one-sided. This may make it difficult to distinguish objects with similar textures but with discriminative scattering patterns.
[0009] Unfortunately, in the field of ships, almost all the currently publicly available SAR ship datasets only have amplitude information and lack spectral information. This fails to fully utilize the unique characteristics of SAR images with rich and complex physical scattering information, resulting in low accuracy in subsequent ship target recognition. Summary of the Invention
[0010] The purpose of this application is to provide a ship target recognition method, device, equipment, medium, and product to solve the problem of low accuracy in ship target recognition.
[0011] To achieve the above purpose, this application provides the following solutions:
[0012] In the first aspect, this application provides a ship target recognition method, including:
[0013] Training an S3GAN network model based on the amplitude ship images in the source domain SAR ship image dataset and the SAR ship phase data in the target domain SAR ship image dataset; the S3GAN network model includes a first generator network model for performing the conversion task from the amplitude domain to the spectral domain, a second generator network model for performing the conversion task from the spectral domain to the amplitude domain, and a discriminator; among them, the first generator network model is constructed by adding a spectral transformation expansion module in front of the feature encoder of the generative adversarial network U-GAT-IT; the second network model is constructed by adding a multi-scale channel comprehensive processing module in front of the feature encoder of the generative adversarial network U-GAT-IT; the frequency domain transformation expansion module is used to convert the single-channel or three-channel amplitude ship image into a 128-channel amplitude ship image rich in deep features; the multi-scale channel comprehensive processing module is used to adjust the output image to the same scale as the input image to extract the depth information of the input image at the frequency domain level;
[0014] Inputting the amplitude ship image to be converted into the trained S3GAN network model to generate the pseudo-SAR spectral information of the amplitude ship image to be converted;
[0015] Input the pseudo-SAR spectrum information and the amplitude ship images in the source-domain SAR ship image dataset into the feature fusion recognition network for feature fusion to identify SAR ship targets in the fused image; the feature fusion recognition network is trained based on the amplitude ship images in the source-domain SAR ship image dataset and the SAR ship phase data in the target-domain SAR ship image dataset.
[0016] In a second aspect, the present application provides a ship target recognition device, including:
[0017] An S3GAN network model training module, configured to train an S3GAN network model according to the amplitude ship images in the source-domain SAR ship image dataset and the SAR ship phase data in the target-domain SAR ship image dataset; the S3GAN network model includes a first generator network model for performing the conversion task from the amplitude domain to the spectrum domain, a second generator network model for performing the conversion task from the spectrum domain to the amplitude domain, and a discriminator; wherein, the first generator network model is constructed by adding a spectrum transformation expansion module in front of the feature encoder of the generative adversarial network U-GAT-IT; the second network model is constructed by adding a multi-scale channel comprehensive processing module in front of the feature encoder of the generative adversarial network U-GAT-IT; the frequency domain transformation expansion module is used to convert a single-channel or three-channel amplitude ship image into a 128-channel amplitude ship image rich in deep features; the multi-scale channel comprehensive processing module is used to adjust the output image to the same scale as the input image to extract the depth information of the input image at the frequency domain level;
[0018] A pseudo-SAR spectrum information generation module, configured to input the amplitude ship image to be converted into the trained S3GAN network model to generate the pseudo-SAR spectrum information of the amplitude ship image to be converted;
[0019] An SAR ship target recognition module, configured to input the pseudo-SAR spectrum information and the amplitude ship images in the source-domain SAR ship image dataset into the feature fusion recognition network for feature fusion to identify SAR ship targets in the fused image; the feature fusion recognition network is trained based on the amplitude ship images in the source-domain SAR ship image dataset and the SAR ship phase data in the target-domain SAR ship image dataset.
[0020] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the ship target recognition method described in any one of the above.
[0021] Fourthly, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the ship target recognition method described in any one of the above is implemented.
[0022] Fifthly, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the ship target recognition method described in any one of the above is implemented.
[0023] According to the specific embodiments provided by the present application, the following technical effects are disclosed: The present application trains an S3GAN network model with a first generator network model for performing the amplitude domain to frequency domain conversion task, a second generator network model for performing the frequency domain to amplitude domain conversion task, and a discriminator based on the amplitude ship images in the source domain SAR ship image dataset and the SAR ship phase data in the target domain SAR ship image dataset. Among them, the first generator network model is constructed by adding a frequency domain transformation expansion module in front of the feature encoder of the generative adversarial network U-GAT-IT; the frequency domain transformation expansion module is used to convert the single-channel or three-channel amplitude ship image into a 128-channel amplitude ship image rich in deep features, so as to realize the expansion of the number of channels and alleviate the problem of difficult domain conversion caused by the large modal difference between the amplitude map and the frequency spectrum information; the second network model is constructed by adding a multi-scale channel comprehensive processing module in front of the feature encoder of the generative adversarial network U-GAT-IT; the multi-scale channel comprehensive processing module is used to adjust the output image to the same scale as the input image to extract the depth information of the input image at the frequency domain level, and solve the problem of difficult information processing caused by too many channels of frequency spectrum information. The present application combines amplitude information and frequency spectrum information for SAR ship target recognition, makes full use of the unique characteristics of SAR ship images with rich and complex physical scattering information, and improves the accuracy of ship target recognition. Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0025] Figure 1 It is a flowchart of a ship target recognition method in an embodiment of the present application;
[0026] Figure 2 It is a general framework diagram of the ship target recognition method in an embodiment of the present application;
[0027] Figure 3It is the structural diagram of the generator network model in the S3GAN network in an embodiment of the present application;
[0028] Figure 4 It is the detailed structural diagram of the multi-scale channel comprehensive processing module in an embodiment of the present application;
[0029] Figure 5 It is the schematic diagram of the data flow and loss function of S3GAN in an embodiment of the present application;
[0030] Figure 6 It is the detailed structural diagram of the feature fusion recognition network in an embodiment of the present application;
[0031] Figure 7 It is the detailed structural diagram of the spectral processing attention module in an embodiment of the present application;
[0032] Figure 8 It is the schematic diagram of S3GAN generating pseudo-spectrum information and the heat map during the generation process in an embodiment of the present application. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0034] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0035] If a deep learning network can learn the feature mapping between the SAR ship target amplitude information domain and the spectrum information domain, it can generate pseudo-spectrum information data for any amplitude map, and can fuse the features of both for target recognition, ultimately achieving a better effect than the target recognition result obtained by simply extracting the features of the amplitude map. Currently, the most common methods between domains include autoencoders, generative adversarial networks, etc. Due to the high robustness and quality of the generated results provided by generative adversarial networks, they have gradually become the most popular generation method between domains. For example, in the U-GAT-IT network, an encoder is incorporated into the generator for feature extraction, and by introducing an adaptive attention mechanism, the generator and discriminator can focus on different regions of the image, thereby achieving high-quality generation effects even with fewer samples.
[0036] However, these popular GANs are designed for single-channel or 3-channel image data. Even if the spatial information of the spectral information data rich in complex information is discarded, its three-dimensional form still has 128 channels. The morphological differences between single-channel amplitude information data and 128-channel spectral information data are significant, and it is not applicable to directly substitute both into common popular GAN networks and target recognition networks.
[0037] It can be seen that there are the following technical problems in implementing the information conversion from SAR amplitude to SAR spectral domain using the U-GAT-IT network (i.e., directly using the open-source generative adversarial network):
[0038] 1. The number of channels is not equal. During domain conversion, the conversion between a single channel and 128 channels cannot achieve a one-to-one pixel point relationship mapping between channels.
[0039] 2. The number of channels of spectral information is too large. Ordinary domain conversion networks for image data have limited vision and cannot comprehensively and effectively extract deeper feature information of spectral information data.
[0040] The embodiments of the present application provide a method for ship target recognition. This method is executed by a computer device, which can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiments of the present application, this method includes the following steps 101 to step 103. Among them:
[0041] Step 101: Train the S3GAN network model according to the amplitude ship images in the source domain SAR ship image dataset and the SAR ship phase data in the target domain SAR ship image dataset; As Figure 2 - Figure 3 shown, the S3GAN network model includes a first generator network model for performing the conversion task from the amplitude domain to the spectral domain, a second generator network model for performing the conversion task from the spectral domain to the amplitude domain, and a discriminator; Among them, the first generator network model is constructed by adding a spectral transformation expansion module in front of the feature encoder of the generative adversarial network U-GAT-IT; the second network model is constructed by adding a multi-scale channel comprehensive processing module in front of the feature encoder of the generative adversarial network U-GAT-IT; the frequency domain transformation expansion module is used to convert a single-channel or three-channel amplitude ship image into a 128-channel amplitude ship image rich in deep features; the multi-scale channel comprehensive processing module is used to adjust the output image to the same scale as the input image to extract the depth information of the input image at the frequency domain level. Figure 3 In a En s represents the encoder, De i represents the decoder, ω represents En aThe k-th activation map in, α i represents the feature map after weight activation.
[0042] Step 102: Input the amplitude ship image to be converted into the trained S3GAN network model to generate the pseudo-SAR spectrum information of the amplitude ship image to be converted.
[0043] Step 103: Input the pseudo-SAR spectrum information and the amplitude ship images in the source domain SAR ship image dataset into the feature fusion recognition network for feature fusion to identify the SAR ship targets in the fused image; the feature fusion recognition network is trained according to the amplitude ship images in the source domain SAR ship image dataset and the SAR ship phase data in the target domain SAR ship image dataset.
[0044] In an exemplary embodiment, before step 101, it further includes: converting the amplitude ship image to the same size using the linear interpolation method; performing joint time-domain analysis processing on the SAR ship phase data and discarding the spatial information of the data after time-domain analysis processing to determine the spectrum information with 128 channels of the target domain SAR ship image.
[0045] Furthermore, the amplitude ship image is uniformly converted to a size of 224*224.
[0046] In an exemplary embodiment, the frequency domain transformation expansion module specifically includes: a 7×7 convolutional layer, a normalization layer, an activation function layer, a frequency domain processing layer, and a fusion processing layer connected in sequence; wherein, the frequency domain processing layer includes a Fourier transform layer, a tensor shape expansion layer, an attention mechanism layer, and an inverse Fourier transform layer connected in sequence;
[0047] The single-channel or three-channel amplitude ship image enters the frequency domain processing layer after being segmented at different scales by the 7×7 convolutional layer, the normalization layer, and the activation function layer; the fusion processing layer is used to perform feature fusion on the amplitude ship images passing through the frequency domain processing layer at different scales in the channel dimension to obtain the output amplitude ship image of the frequency domain exchange expansion module, and this output amplitude ship image is the input of the feature encoder in the first generator network model.
[0048] Furthermore, first, the input data is input into the convolutional layer for preliminary feature extraction.
[0049] Then, after performing Fourier transform, the real and imaginary part frequency domain data are concatenated and merged in the frequency domain to obtain the frequency domain data, enabling further feature enhancement operations. At the same time, after global pooling, channel attention weighting is performed, and after weighting, convolutional operations are performed to obtain the final spectral space representation.
[0050] Finally, the frequency-domain data is converted into time-domain data by using the inverse Fourier transform, and spatial scale transformation is performed to adjust the data to the same scale as the input data, realizing the extraction of depth information from the input data at the frequency-domain level.
[0051] Next, a part of the input data is taken to extract the low-frequency components, and the result is obtained through the same processing as the above steps.
[0052] Finally, the data after frequency-domain conversion is added to the original input data and the result of low-frequency component processing, and the number of channels is adjusted to 128 through a convolution kernel with a kernel size of 1.
[0053] With such a design, the spectrum transformation expansion module can convert single-channel or three-channel input data into 128-channel data rich in deep features, thus realizing the expansion of the number of channels, providing a richer feature representation for subsequent domain conversion tasks, and alleviating the feature correspondence problem caused by a large difference in the number of channels.
[0054] In an exemplary embodiment, the multi-scale channel comprehensive processing module specifically includes: three processing branches; the input of each branch is the same spectrum information with 128 channels, and the output is the result obtained by splicing the feature information processed by the three branches in the channel dimension; among them, the first branch includes a normalization layer and an activation function layer connected in sequence, and the output obtained by the first branch is the global information of the spectrum information; the second branch includes a random selection layer, a 1×1 convolution layer, a normalization layer, and an activation function layer connected in sequence. The second branch extracts features by randomly selecting half of the channel spectrum information, improving the robustness and stability of the network in extracting spectrum information features; the third branch includes a channel attention layer, a normalization layer, and an activation function layer connected in sequence. Among them, the channel attention layer includes a global pooling layer and a weight calculation layer, and the third branch extracts the key information in the spectrum information through channel attention.
[0055] Furthermore, as Figure 4 shown, first, a convolution kernel with a kernel size of 3 and a padding of 1 is used. This convolution layer performs convolution operations on all input channels, and the output feature map preserves the information of all channels. Then, for multi-channel data, half of the number of channels is randomly selected as the input number of channels in a random selection manner, and the information of the selected channels is obtained using the convolution kernel of the same size. This process makes the model have a certain degree of randomness when processing input data, which helps the generalization ability of the model.
[0056] Subsequently, after performing global pooling on the real spectral data, the weight information of each channel is calculated to enable the network to learn the weighted feature information, so as to obtain the feature information that enhances the response of important channels and suppresses the response of unimportant channels.
[0057] Finally, the feature information processed by full-channel convolution, convolution with randomly selected half of the channels, and the channel attention mechanism is concatenated, and a convolution kernel with a kernel size of 1 is used for further extraction of the feature information and adjustment of the number of channels. The generator incorporating this module demonstrates flexibility in processing data with a large number of channels by extracting features from the original single data at different scales, improving the efficiency and performance of the model.
[0058] In an exemplary embodiment, as Figure 5 shown, the game form between the generator network model and the discriminator in the S3GAN network model is:
[0059]
[0060] where, is the min-max game between the generator network model and the discriminator; G is the generator network model, and the generator network model is the first generator network model or the second generator network model; D is the discriminator, that is, the discriminator; is the loss of the real data, and the goal of this term is to maximize D(v), making the output of the discriminator for the real data close to 1, so as to classify it as real data; E is the expectation; v ∼ P r (v) is the real data v and the data distribution P r (v); D(v) is the discriminant output of the discriminator for the real data, indicating the probability that the discriminator believes the input data is real; is the loss of the fake data; z ∼ P g (z) is the fake data z and the probability distribution P g (z); D(G(z)) represents the discriminant output of the discriminator for the fake data G(z) generated by the generator network model, indicating the probability that the discriminator believes the input data is real. The goal of this term is to minimize D(G(z)), making the output of the discriminator for the generated data close to 0, so as to classify it as fake data. The goal of the generator is to deceive the discriminator and generate realistic target domain images, while the goal of the discriminator is to distinguish as accurately as possible between the fake images generated by the generator and the real target domain images. This competition prompts the generator to learn a better image conversion mapping.
[0061] The total loss function expression of the S3GAN network for the conversion task from amplitude to spectral domain is:
[0062]
[0063] Among them, α, β, and γ respectively represent and weights; represents the adversarial loss, represents the auxiliary classification loss, represents the cycle-consistency loss, D represents the discriminator, and G represents the generator, represents the Class Activation Map Loss (CAMLoss) function used in the spectral information discriminator.
[0064] The adversarial loss is adopted to reduce the difference in feature distribution between the pseudo-target and the real target samples:
[0065]
[0066] where x ∼ X s represents the real spectral domain sample, x ∼ X a represents the real amplitude domain sample, G a→s (x) represents the generator network model in the formation process of the adversarial game for the amplitude domain to spectral domain conversion task, D s represents the discriminator for real SAR spectral information and pseudo SAR spectral information, X a and X s represent the ship amplitude images and ship spectral information from the ship amplitude image datasets in the source domain SAR ship image dataset and the target domain SAR ship image dataset respectively, represents the expected value of the data sampled from the real spectral data distribution, represents the expected value of the data sampled from the real amplitude data distribution, D s (x) represents the discriminator's discriminative output for the real spectral information x, D s (G a→s (x)) represents the discriminator's discriminative output for the pseudo spectral information generated from the amplitude information.
[0067] The function is defined as:
[0068]
[0069] where G s→a represents the generator network model in the formation process of the adversarial game for the spectral domain to amplitude domain task, G s→a (G a→s (x)) represents the pseudo spectral information generated from the real amplitude information passing through G s→aThe pseudo-amplitude information generated by the generator, where |·|1 represents the L1 norm, i.e., the absolute value error. This term measures the conversion loss from the amplitude domain to the frequency spectrum domain and back to the amplitude domain. By minimizing this loss, it is ensured that the data remains the same when going from the amplitude domain to the frequency spectrum domain and then back to the amplitude domain.
[0070] It is mainly used to generate heatmaps in image classification tasks to visualize the regions of the image that the convolutional neural network focuses on when making decisions.
[0071]
[0072] Among them, η s (x) represents the auxiliary classifier in the generator.
[0073] In an exemplary embodiment, as Figure 6 shown, the feature fusion recognition network specifically includes: two processing branches, a feature fusion module, and an object recognition module.
[0074] Among them, as Figure 7 shown, the first processing branch includes a spectral processing attention module and a first feature extraction module; the first processing branch is used to process the pseudo-SAR spectral information; the spectral attention module includes a global feature extraction module and a channel attention feature extraction module; the inputs of the global feature extraction module and the channel attention feature extraction module are the same pseudo-SAR spectral information, and the outputs of the global feature extraction module and the channel attention feature extraction module are concatenated in the channel dimension to obtain the output of the spectral processing attention module; the global feature extraction module includes a 7×7 convolutional layer, a batch normalization layer, an activation function layer, a pooling layer, and a 1×1 convolutional layer connected in sequence; the channel attention feature extraction module includes a global pooling layer and a weight acquisition layer, a 1×1 convolutional layer, a spectral normalization layer, a batch normalization layer, an activation function layer, a max pooling layer, and a 1×1 convolutional layer connected in sequence.
[0075] The second processing branch includes a second feature processing module; the second processing branch is used to process the amplitude ship image to be converted.
[0076] The feature fusion module is used to fuse the pseudo-SAR spectral information processed by the first processing branch and the amplitude ship image processed by the second processing branch to obtain a fused image.
[0077] The object recognition module is used to recognize the SAR ship target in the fused image.
[0078] Furthermore, first, the present application performs global average pooling on the input spectral data and obtains the weights of a total of one channel, designed to enable the network to continuously update the weights to learn the differences in the importance of different channels, and based on this, focuses on the channels beneficial to the classification effect among a large number of channels.
[0079] Then, the present application adds a spectral feature convolutional layer group, which is in turn a common two-dimensional convolution, spectral normalization layer, batch normalization layer, and ReLU provided by PyTorch. Spectral normalization helps reduce the problems of gradient explosion and gradient disappearance during model training by constraining the spectral norm of the weight matrix, and improves the stability and generalization ability of the model during the processing of complex spectral data.
[0080] Subsequently, the present application adjusts the spatial size and number of channels of the processed feature information through a max-pooling layer and a convolutional kernel with a kernel size of 1, and adjusts the spatial size and number of channels of the original input data after feature extraction through a convolutional layer with a kernel_size = 7, stride = 2, and padding = 3, and then also through batch normalization processing and activation by a ReLU activation function, and finally fuses and obtains the output.
[0081] Common object recognition networks are divided into feature extraction and classification output. The feature fusion recognition network of the present application performs feature extraction on the amplitude information and spectral information respectively, then concatenates them in the channel dimension, and then performs classification output through a fully connected layer.
[0082] The present application realizes the reasonable feature fusion of two types of data in different modalities after feature fusion network object recognition of the pseudo-SAR spectral information generated by the S3GAN network and the amplitude information, and uses the spectral processing attention module to extract the effective features of complex spectral information, enhancing the training stability and generalization ability of the model, improving the learning ability of complex spectral information, and finally achieving a better object recognition effect than using only amplitude maps for object recognition.
[0083] To illustrate the effectiveness of the proposed generative adversarial network S3GAN network of the present application, the real amplitude map, the heat map of focusing on key features during the S3GAN training process, and the effect diagram of the generated spectral information are shown, as Figure 8 shown.
[0084] To verify the rationality and effectiveness of the ship target recognition method in solving the problems of scarce effectively labeled samples and lack of spectral information in the SAR ship target recognition deep learning network model, the present application tests the ability of the pseudo-spectral information fused with amplitude images to improve the recognition accuracy of popular classification networks.
[0085] The object recognition accuracy of the feature fusion recognition network after feature fusion has an average improvement of 0.12 under the conditions of three popular classification networks, three types of ships, and 10 samples for each type, which proves the rationality, effectiveness, and application value of this application in solving the problems of scarce effectively labeled samples and lack of spectral information in the deep learning network model for SAR ship automatic target recognition.
[0086] An S3GAN proposed in this application can generate its pseudo-spectral information according to SAR ship amplitude data, and a multi-channel feature fusion automatic recognition framework that supports feature fusion of amplitude data and pseudo-spectral data is proposed based on this S3GAN. A multi-scale channel comprehensive processing module and a frequency domain transformation expansion module are designed to address problems such as the lack of focus in the training of the generative adversarial network due to the large number of data channels in the spectral information and the difficulty of feature mapping due to the large difference in the number of channels from amplitude information to spectral information; at the same time, in order to achieve the optimized processing of spectral information by the feature fusion recognition network, a spectral processing attention module is designed to perform feature extraction processing on important channels at the spectral level; finally, according to the preset training parameters and loss function, the S3GAN network is trained, and the trained S3GAN network is used to generate pseudo-spectral information and fuse the amplitude Figure 1 to perform the SAR ship target recognition task. The pseudo-spectral information generated by this application and the improved fusion feature network are highly reasonable and have application value in solving the problems of scarce effectively labeled samples and lack of spectral information in the deep learning network model for SAR ship automatic target recognition.
[0087] Based on the same inventive concept, the embodiments of this application also provide a ship target recognition device for implementing the ship target recognition method involved above. The implementation solutions provided by this device to solve problems are similar to those recorded in the above method, so the specific limitations in one or more embodiments of the ship target recognition device provided below can refer to the limitations on the ship target recognition method in the above text and will not be repeated here.
[0088] In an exemplary embodiment, as Figure 5 shown, a ship target recognition device is provided, including:
[0089] The S3GAN network model training module is used to train the S3GAN network model according to the amplitude ship images in the source domain SAR ship image dataset and the SAR ship phase data in the target domain SAR ship image dataset; the S3GAN network model includes a first generator network model for performing the conversion task from the amplitude domain to the frequency domain, a second generator network model for performing the conversion task from the frequency domain to the amplitude domain, and a discriminator; wherein, the first generator network model is constructed by adding a spectral transformation expansion module in front of the feature encoder of the generative adversarial network U-GAT-IT; the second network model is constructed by adding a multi-scale channel comprehensive processing module in front of the feature encoder of the generative adversarial network U-GAT-IT; the frequency domain transformation expansion module is used to convert the single-channel or three-channel amplitude ship images into 128-channel amplitude ship images rich in deep features; the multi-scale channel comprehensive processing module is used to adjust the output image to the same scale as the input image to extract the depth information of the input image at the frequency domain level.
[0090] The pseudo-SAR spectrum information generation module is used to input the amplitude ship image to be converted into the trained S3GAN network model to generate the pseudo-SAR spectrum information of the amplitude ship image to be converted.
[0091] The SAR ship target recognition module is used to input the pseudo-SAR spectrum information and the amplitude ship images in the source domain SAR ship image dataset into the feature fusion recognition network for feature fusion to identify the SAR ship targets in the fusion image; the feature fusion recognition network is trained according to the amplitude ship images in the source domain SAR ship image dataset and the SAR ship phase data in the target domain SAR ship image dataset.
[0092] In an exemplary embodiment, a computer device is provided, which can be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through the system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store ship target recognition data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a ship target recognition method.
[0093] In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0094] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0095] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0096] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0097] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0098] In each of the embodiments provided in this application, the database involved may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., and is not limited thereto. In each of the embodiments provided in this application, the processor involved may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., and is not limited thereto.
[0099] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0100] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for identifying ship targets, characterized in that, The ship target recognition method includes: Training an S3GAN network model based on the amplitude ship images in the source domain SAR ship image dataset and the SAR ship phase data in the target domain SAR ship image dataset; the S3GAN network model includes a first generator network model for performing the conversion task from the amplitude domain to the frequency spectrum domain, a second generator network model for performing the conversion task from the frequency spectrum domain to the amplitude domain, and a discriminator; wherein, the first generator network model is constructed by adding a frequency spectrum transformation expansion module in front of the feature encoder of the generative adversarial network U-GAT-IT; the second network model is constructed by adding a multi-scale channel comprehensive processing module in front of the feature encoder of the generative adversarial network U-GAT-IT; the frequency domain transformation expansion module is used to convert the single-channel or three-channel amplitude ship image into a 128-channel amplitude ship image rich in deep features; the multi-scale channel comprehensive processing module is used to adjust the output image to the same scale as the input image to extract the depth information of the input image at the frequency domain level; Inputting the amplitude ship image to be converted into the trained S3GAN network model to generate the pseudo-SAR frequency spectrum information of the amplitude ship image to be converted; Inputting the pseudo-SAR frequency spectrum information and the amplitude ship images in the source domain SAR ship image dataset into a feature fusion recognition network for feature fusion to identify the SAR ship targets in the fused image; the feature fusion recognition network is trained based on the amplitude ship images in the source domain SAR ship image dataset and the SAR ship phase data in the target domain SAR ship image dataset.
2. The ship target recognition method according to claim 1, wherein Before training the S3GAN network model based on the amplitude ship images in the source domain SAR ship image dataset and the SAR ship phase data in the target domain SAR ship image dataset, it further includes: Using the linear interpolation method to convert the amplitude ship images to the same size; Performing joint time domain analysis processing on the SAR ship phase data, and discarding the spatial information of the data after the time domain analysis processing to determine the frequency spectrum information with 128 channels of the target domain SAR ship image.
3. The ship target recognition method according to claim 1, characterized in that The frequency domain transformation expansion module specifically includes: a 7×7 convolutional layer, a normalization layer, an activation function layer, a frequency domain processing layer, and a fusion processing layer connected in sequence; wherein, the frequency domain processing layer includes a Fourier transform layer, a tensor shape expansion layer, an attention mechanism layer, and an inverse Fourier transform layer connected in sequence; The single-channel or three-channel amplitude ship image enters the frequency domain processing layer through different-scale segmentation by the 7×7 convolutional layer, the normalization layer, and the activation function layer; the fusion processing layer is used to perform feature fusion on the amplitude ship images of different scales passing through the frequency domain processing layer in the channel dimension to obtain the output amplitude ship image of the frequency domain exchange expansion module, and this output amplitude ship image is the input of the feature encoder in the first generator network model.
4. The ship target recognition method according to claim 1, characterized in that, The multi-scale channel comprehensive processing module specifically includes: three processing branches; the input of each branch is the same spectral information with 128 channels, and the output is the result obtained by concatenating the feature information processed by the three branches in the channel dimension; among them, the first branch includes a normalization layer and an activation function layer connected in sequence, and the output obtained by the first branch is the global information of the spectral information; the second branch includes a random selection layer, a 1×1 convolutional layer, a normalization layer and an activation function layer connected in sequence, and the second branch extracts features by randomly selecting half of the channels of the spectral information; the third branch includes a channel attention layer, a normalization layer and an activation function layer connected in sequence, where the channel attention layer includes a global pooling layer and a weight calculation layer, and the third branch extracts key information in the spectral information through channel attention.
5. The ship target recognition method according to claim 1, characterized in that, In the S3GAN network model, the game form between the generator network model and the discriminator is: Among them, is the min-max game between the generator network model and the discriminator; G is the generator network model, and the generator network model is the first generator network model or the second generator network model; D is the discriminator; is the loss of real data; E is the expectation; v ~ P r (v) is the real data v and the data distribution P r (v); D(v) is the discriminant output of the discriminator for the real data; is the loss of fake data; z ~ P g (z) is the fake data z and the probability distribution P g (z); D(G(z)) represents the discriminant output of the discriminator for the fake data G(z) generated by the generator network model.
6. The ship target recognition method according to claim 1, characterized in that The feature fusion and recognition network specifically includes: two processing branches, a feature fusion module, and an object recognition module; Among them, the first processing branch includes a spectral processing attention module and a first feature extraction module; the first processing branch is used to process the pseudo-SAR spectral information; the spectral attention module includes a global feature extraction module and a channel attention feature extraction module; the input of the global feature extraction module and the channel attention feature extraction module is the same pseudo-SAR spectral information, and the outputs of the global feature extraction module and the channel attention feature extraction module are concatenated in the channel dimension to obtain the output of the spectral processing attention module; the global feature extraction module includes a 7×7 convolutional layer, a batch normalization layer, an activation function layer, a pooling layer and a 1×1 convolutional layer connected in sequence; the channel attention feature extraction module includes a global pooling layer and a weight acquisition layer, a 1×1 convolutional layer, a spectral normalization layer, a batch normalization layer, an activation function layer, a max pooling layer and a 1×1 convolutional layer connected in sequence; The second processing branch includes a second feature processing module; the second processing branch is used to process the amplitude ship image to be converted; The feature fusion module is used to fuse the pseudo-SAR spectral information processed by the first processing branch and the amplitude ship image processed by the second processing branch to obtain a fused image; The object recognition module is used to recognize the SAR ship target in the fused image.
7. A ship target recognition device, characterized in that The ship target recognition device includes: The S3GAN network model training module is used to train the S3GAN network model according to the amplitude ship images in the source domain SAR ship image dataset and the SAR ship phase data in the target domain SAR ship image dataset; the S3GAN network model includes a first generator network model for performing the conversion task from the amplitude domain to the frequency spectrum domain, a second generator network model for performing the conversion task from the frequency spectrum domain to the amplitude domain, and a discriminator; wherein, the first generator network model is constructed by adding a frequency spectrum transformation expansion module in front of the feature encoder of the generative adversarial network U-GAT-IT; the second network model is constructed by adding a multi-scale channel comprehensive processing module in front of the feature encoder of the generative adversarial network U-GAT-IT; the frequency domain transformation expansion module is used to convert the single-channel or three-channel amplitude ship image into a 128-channel amplitude ship image rich in deep features; the multi-scale channel comprehensive processing module is used to adjust the output image to the same scale as the input image to extract the depth information of the input image at the frequency domain level; The pseudo-SAR frequency spectrum information generation module is used to input the amplitude ship image to be converted into the trained S3GAN network model to generate the pseudo-SAR frequency spectrum information of the amplitude ship image to be converted; The SAR ship target recognition module is used to input the pseudo-SAR frequency spectrum information and the amplitude ship images in the source domain SAR ship image dataset into the feature fusion recognition network for feature fusion to recognize the SAR ship targets in the fusion image; the feature fusion recognition network is trained according to the amplitude ship images in the source domain SAR ship image dataset and the SAR ship phase data in the target domain SAR ship image dataset.
8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the ship target recognition method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the ship target recognition method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the ship target recognition method according to any one of claims 1-6.
Citation Information
Patent Citations
Single ship target SAR image generation method based on generative adversarial network
CN112052899A
SAR image arbitrary direction ship target generation method based on generative adversarial network
CN116468880A