Optical image translation method based on ViT-Pix2Pix
By adopting the target translation network model combined with ViT-Pix2Pix in SAR image translation, the problem of low translation quality from SAR image to optical image is solved, high-quality optical image generation and authenticity discrimination are achieved, and the interpretation of SAR image is assisted.
Patent Information
- Application Number
- CN202210779801.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-07-04
AI Technical Summary
The prior art is difficult to translate SAR images into optical images in a comprehensive and high-quality manner, resulting in difficulty in interpreting data.
Using the optical image translation method based on ViT-Pix2Pix, by constructing a target translation network model combined with Vision Transformer and Pix2Pix, neural network training and optimization are used to generate high-quality pseudo-optical images and judge the authenticity of the optical image.
It improves the translation quality of SAR images to optical images, enhances the accuracy of authenticity discrimination of optical images, assists in the interpretation of SAR images, and improves the quality of generated images and the performance of discriminators.
Smart Images

Figure CN115272787B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image translation, and in particular to an optical image translation method based on ViT-Pix2Pix. Background Art
[0002] Synthetic Aperture Radar (SAR) is a high-resolution imaging radar that can work all day and all weather. It can obtain high-resolution radar images under extremely low visibility weather conditions, which is difficult to achieve with optical remote sensing. Therefore, SAR images have different alternative roles in many fields. However, the imaging principles of SAR and optical remote sensing are completely different, making SAR images less intuitive than optical images and data interpretation more difficult. With the increasing demand for SAR data analysis, there is a need for a method that can convert input SAR images into optical output images to assist in the judgment of raw data.
[0003] In the prior art, since the information content in SAR images and optical images is partially overlapping and partially incompatible, that is, the two sensors can only observe part of the information, and each sensor observes other information that the other sensor cannot observe, the translation task from SAR images to optical images cannot be achieved comprehensively and with high quality.
[0004] Therefore, there is an urgent need for a SAR image to optical image translation method that can generate high-quality images. Summary of the invention
[0005] Based on this, it is necessary to provide an optical image translation method based on ViT-Pix2Pix to address the above technical problems.
[0006] An optical image translation method based on ViT-Pix2Pix comprises the following steps: obtaining a SAR image to be tested; constructing an initial target translation network model, and optimizing the parameters of the initial target translation network model through paired SAR images and optical images to obtain a target translation network model, wherein the target translation network model is a model combining Vision Transformer and Pix2Pix, and comprises a generator and a discriminator, wherein the generator is used to translate the SAR image into a pseudo optical image, and the discriminator is used to judge whether an input optical image is a true optical image matched by the SAR image, and the generator and the discriminator complete neural network training optimization in an adversarial form; and the SAR image to be tested is input into the target translation network model to obtain a target optical image.
[0007] In one of the embodiments, the initial target translation network is constructed, and the parameters of the initial target translation network are optimized through paired SAR images and optical images to obtain the target translation network, specifically including: taking Pix2Pix as the basic model and combining it with Vision Transformer to form a ViT-Pix2Pix initial target translation network model; optimizing the parameters of the initial target translation network model through paired SAR images and optical images; inputting the SAR image into a generator to output a pseudo-optical image corresponding to the SAR image; performing data enhancement on the SAR image and the true optical image, and the SAR image and the pseudo-optical image; inputting the data-enhanced image pair into a discriminator, dividing the image pair into small blocks of fixed size and non-overlapping, flattening them into linear embedding for processing, and outputting the probability that the optical image is a real image and matches the SAR image; optimizing the parameters of the generator and the discriminator through a cross entropy loss function, an L1 loss function and a balanced consistency regularization method to obtain a target translation model.
[0008] In one of the embodiments, the step of inputting a SAR image into a generator and outputting a pseudo optical image corresponding to the SAR image specifically includes: obtaining a SAR image as a training sample and a corresponding true optical image as an image pair; inputting the image pair into a generator, performing feature extraction through the generator, and obtaining a pseudo optical image, wherein the generator is U-Net.
[0009] In one embodiment, the training optimization of the generator specifically includes: calculating the L1 loss of the true optical image and the pseudo optical image, and the classification loss applied by the generator, respectively, according to the SAR image, the true optical image and the pseudo optical image, and the formula is:
[0010] L L1 (G) = E x,y [||yG(x)||1]
[0011] L cGAN (G) = -E x [logD(x,G(x))]
[0012] Where x represents the SAR image, G represents the generator for generating optical images from SAR images, G(x) represents the pseudo optical image generated by the generator, y represents the true optical image, ||·||1 represents the sum of the absolute values of the differences between the corresponding pixels of the two images, and E x,y [·] represents the expectation after calculating the loss for all image pairs (x, y), and the final loss is obtained, E x [·] represents the expectation after calculating the loss of the SAR image; according to the L1 loss and classification loss, the total loss of the generator is calculated as:
[0013] L(G)=L cGAN (G)+λ L1 L L1 (G)
[0014] In the formula, λ L1 is a configurable hyperparameter; according to the total loss of the generator, the back-propagation algorithm is used to update the neural network training parameters of the generator to optimize the generator.
[0015] In one of the embodiments, the training optimization of the discriminator specifically includes: inputting the image pair into the Vision Transformer network model to determine the authenticity of the optical image and whether the two images in the image pair match; merging the two images in the image pair into a multi-channel input, and dividing them into small blocks of fixed size and non-overlapping, obtaining a linearly arranged embedding through a fully connected layer, and adding a classification symbol at the beginning of the sequence; after adding position information encoding to the linearly arranged embedding, completing the processing in the Transformer encoder of the improved self-attention layer; inputting the output features of the classification symbol into a multi-layer perceptron to complete the discrimination, and obtaining a true optical image and a false optical image.
[0016] In one embodiment, the self-attention layer of the Vision Transformer network model is improved, specifically including: using L2 distance to replace the dot product operation in the self-attention process, and using it to query and input the weight binding of the projection matrix of the self-attention, and the improved self-attention layer is calculated as:
[0017]
[0018] Where W q =W k , W q , W k and W v are the projection matrices for query, key, and value, respectively. d(·,·) computes the vectorized L2 distance between two sets of points. is the characteristic size of each head; the spectral normalization method is used to optimize the improved VisionTransformer network model.
[0019] In one embodiment, the classification loss generated during the training of the discriminator is: the improved Vision Transformer network model is applied to the discriminator, and the classification loss between the true optical image and the pseudo optical image is obtained according to the correctness or error of the judgment result of the discriminator:
[0020] L cGAN (D)=-E x,y [logD(x,y)]-Ex [1-logD(x,G(x))]
[0021] Where D(x, y) represents inputting x and y into the discriminator D to obtain the discrimination result; D(x, G(x)) represents inputting x and G(x) into the discriminator D to obtain the discrimination result; G(x) is the pseudo optical image generated by the x input generator; the correct judgment results of all SAR images are expected to obtain the classification loss between optical images.
[0022] In one embodiment, a balanced consistency regularization method is used to obtain the L2 loss of the discriminator, specifically including: performing data enhancement on the SAR image and pseudo-optical image pairs, and the SAR image and real optical image pairs, respectively, and inputting the image pairs after data enhancement into the improved Vision Transformer network model; using the L2 loss requires that the output result of the image pair after data enhancement is consistent with the unenhanced result, and the formula is:
[0023] L bcR_fake =||D(x,G(x))-D(T(x,G(x)))||2
[0024] L bCR_real =||D(x, y)-D(T(x, y))||2
[0025] Where x represents the SAR image, G represents the generator for generating optical images from SAR images, G(x) represents the pseudo optical image generated by the generator, y represents the real optical image, D represents the discriminator for inputting the SAR image and the optical image pair, D(x, G(x)) and D(x, y) represent the results of the SAR image and the pseudo optical image pair, and the SAR image and the real optical image pair input to the discriminator, respectively, T represents the transformation of data augmentation, and ||·||2 represents the minimum square error between the corresponding pixels of the two images.
[0026] In one embodiment, the total loss of the discriminator is calculated based on the classification loss and the balanced consistency regularization method, and the formula is:
[0027] L(D)=L cGAN (D)+λ bCR_fake L bCR_fake +λ bCR_real L bCR_real
[0028] In the formula, λ bCR_fake , bCR_real is a configurable hyperparameter; based on the total loss, the back propagation algorithm is used to update the neural network training parameters to achieve the optimization of the discriminator.
[0029] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: the present invention acquires a SAR image to be tested, constructs an initial target translation network model, and optimizes the parameters of the initial target translation network model through paired SAR images and related part images to acquire a target translation network model, wherein the target translation network model is a model combining Vision Transformer and Pix2Pix, and comprises a generator and a discriminator, wherein the generator is used to translate the SAR image into a pseudo-optical image, and the discriminator is used to judge whether the input optical image is a real optical image matched by the SAR image, and the generator and the discriminator complete the neural network training optimization in an adversarial form; the SAR image to be tested is input into the target translation network model to acquire the target optical image, and the overall structural information of the image can be taken into account, the authenticity of the optical image can be more accurately identified, a clearer optical image can be generated, the interpretation of the SAR image can be assisted, the quality of the generated image and the performance of the discriminator are improved, and the stability of the network training is also guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A schematic diagram of a flow chart of an optical image translation method based on ViT-Pix2Pix in one embodiment:
[0031] Figure 2 A schematic diagram of a process of converting a SAR image to an optical image in one embodiment;
[0032] Figure 3 A schematic diagram of the structure of a Vision Transformer network model in one embodiment;
[0033] Figure 4 Schematic diagram of an improved scheme for the self-attention layer of a Vision Transformer network model in one embodiment. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific implementation methods in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0035] In one embodiment, Figures 1 to 4 As shown, an optical image translation method based on ViT-Pix2Pix is provided, comprising the following steps:
[0036] Step S101, obtaining a SAR image to be measured.
[0037] Specifically, a SAR image to be tested is acquired through a synthetic aperture radar for subsequent translation of the SAR image to an optical image, and a number of sample SAR images and corresponding true optical images are acquired at the same time to train an initial target translation network model.
[0038] Step S102, constructing an initial target translation network model, and optimizing the parameters of the initial target translation network model through paired SAR images and optical images to obtain a target translation network model, wherein the target translation network model is a model combining VisionTransformer and Pix2Pix, including a generator and a discriminator, wherein the generator is used to translate the SAR image into a pseudo-optical image, and the discriminator is used to determine whether the input optical image is a true optical image matched by the SAR image, and the generator and the discriminator complete the neural network training optimization in an adversarial form.
[0039] Specifically, an initial target translation network model is constructed based on Vision Transformer and Pix2Pix. Pix2Pix is used as the basic model. Taking into account the overall structural information of the image, the generated image is clearer. The model can perform learning loss according to specific tasks and data, and is suitable for a variety of settings. The discriminator uses Vision Transformer. According to the convergence condition of the generative adversarial network, the spectral normalization method is applied and the self-attention layer of Vision Transformer is improved. At the same time, the balanced consistency regularization method is applied to improve the training stability. The Vision Transformer self-attention architecture has better effect than the convolutional architecture, which improves the performance of the discriminator, thereby improving the accuracy of optical image authenticity discrimination. In the process of confrontation with the generator, the generator is prompted to generate more realistic pseudo-optical images, thereby assisting the interpretation of SAR images.
[0040] Among them, step S102 specifically includes: taking Pix2Pix as the basic model and combining it with Vision Transformer to form a ViT-Pix2Pix initial target translation network model; optimizing the parameters of the initial target translation network model through paired SAR images and optical images; inputting the SAR image into the generator, and outputting a pseudo-optical image corresponding to the SAR image; performing data enhancement on the SAR image and the true optical image, and the SAR image and the pseudo-optical image; inputting the image pairs after data enhancement into the discriminator, dividing the image pairs into small blocks of fixed size and non-overlapping, and flattening them into linear embedding for processing, and outputting the probability that the optical image is a real image and matches the SAR image; optimizing the parameters of the generator and the discriminator through the cross entropy loss function, the L1 loss function and the balanced consistency regularization method to obtain the target translation model.
[0041] Specifically, in order to realize the translation from SAR images to optical images, the Pix2Pix model is used as the basic model and combined with VisionTransformer to form the ViT-Pix2Pix initial target translation network model, and the parameters of the initial target translation network are optimized through paired SAR images and optical images. First, the SAR image is input into the generator, and the pseudo-optical image corresponding to the SAR image is output; data enhancement is performed on the SAR image and true optical image pairs, and the SAR image and pseudo-optical image pairs, and the image pairs are input into the discriminator. The image pairs are divided into small blocks of fixed size and non-overlapping, and flattened into linear embedding for processing through the improved Vi, and the output optical image is the probability that the real image matches the SAR image; the cross entropy loss function, L1 loss function and balanced consistency regularization method are used to optimize the parameters of the generator and the discriminator to obtain the target translation network model, and the target optical image is obtained by inputting the SAR image to be tested into the target translation network.
[0042] It should be understood that the above optimization steps are iterative. The objective function of the generator requires generating a pseudo-optical image that is as close to the actual situation as possible, while the objective function of the discriminator requires judging as much as possible whether the input optical image is a pseudo-optical image generated by the generator or a true optical image. The zero-sum game between the two makes the model gradually tend to the optimal.
[0043] The SAR image is input into a generator, and a pseudo-optical image corresponding to the SAR image is output, which specifically includes: obtaining a SAR image as a training sample and a corresponding true optical image as an image pair; inputting the image pair into a generator, performing feature extraction through the generator, and obtaining a pseudo-optical image, and the generator is U-Net.
[0044] Specifically, a SAR image and a corresponding true optical image as a training sample are obtained as an image pair, and the paired SAR image and true optical image are input into a U-Net with an adapted number of layers for feature extraction to obtain a pseudo optical image corresponding to the SAR image. The selection of the generator model can be any existing method, as long as the input original image can output the corresponding pseudo optical image.
[0045] The discriminator in this embodiment learns the loss according to the specific task and data, and has high parameter efficiency and better discrimination effect compared with the discriminator based on the convolutional architecture.
[0046] Among them, when training and optimizing the discriminator, it specifically includes: inputting the image pair into the Vision Transformer network model to judge the authenticity of the optical image and whether the two images in the image pair match; merging the two images in the image pair into a multi-channel input, and dividing them into small blocks of fixed size and non-overlapping, and obtaining a linearly arranged embedding through a fully connected layer, and adding a classification symbol at the beginning of the sequence; after adding position information encoding to the linearly arranged embedding, the processing is completed in the Transformer encoder of the improved self-attention layer; the output features of the classification symbol are input into the multi-layer perception discriminator to complete the discrimination, and obtain true optical images and false optical images.
[0047] Specifically, the matching SAR image and the optical image are combined in pairs and input into the Vision Transformer network model to determine the authenticity of the optical image and whether the two input images match. Specifically, the two images are merged into a multi-channel input, divided into small blocks of fixed size and non-overlapping, and a linearly arranged embedding is obtained through a fully connected layer. A special classification symbol is added to the beginning of the sequence for subsequent discrimination, which is similar to the processing of symbols in the field of natural language processing. The discriminator adds position information encoding to the embedding of the linear arrangement, completes the processing in the Vision Transformer network model with an improved self-attention layer, and then takes the output features of the classification symbol and inputs the multi-layer perceptron to complete the discrimination, that is, to separate the true and false optical images into two categories, and obtain true optical images and false optical images.
[0048] Among them, the self-attention layer of the Vision Transformer network model is improved, specifically including: using L2 distance to replace the dot product operation in the self-attention process, and using it to query and input the weight binding of the projection matrix of self-attention. The improved self-attention layer is calculated as:
[0049]
[0050] Where W q =W k , W q , W k and W v are the projection matrices for query, key, and value, respectively. d(·,·) computes the vectorized L2 distance between two sets of points. is the characteristic size of each head; the spectral normalization method is used to optimize the improved VisionTransformer network model.
[0051] Specifically, the improvement scheme of the self-attention layer of the Vision Transformer network model is as follows: Figure 4As shown in the figure, the continuity of the Lipschitz continuity condition affects the existence of the optimal discriminant function of the discriminator and the existence of the shifted Nash equilibrium, while the Lipschitz continuity condition constant in the standard dot product self-attention layer may be unbounded, which destroys the Lipschitz continuity in the Vision Transformer network model. In order to enhance the Lipschitz continuity of the discriminator, the L2 distance is used to replace the dot product operation in the self-attention process of the Vision Transformer network model, and the weights of the projection matrix used for query and input self-attention are bounded to obtain the Vision Transformer network model with improved self-attention layer.
[0052] Similarly, the spectral normalization method is applied to the Vision Transformer network model to further enhance the Lipschitz continuity and improve the stability of the Vision Transformer network model in generative adversarial network training.
[0053] Among them, the classification loss in the discriminator training process is: the improved Vision Transformer network model is applied to the discriminator, and the classification loss between the true optical image and the pseudo optical image is obtained according to the correctness or error of the discriminator's judgment result:
[0054] L cGAN (D)=-E x,y [logD(x,y)]-E x [1-logD(x,G(x))]
[0055] Where D(x, y) represents inputting x and y into the discriminator D to obtain the discrimination result; D(x, G(x)) represents inputting x and G(x) into the discriminator D to obtain the discrimination result; G(x) is the pseudo optical image generated by the x input generator; the correct judgment results of all SAR images are expected to obtain the classification loss between optical images.
[0056] Specifically, the improved Vision Transformer network model is applied as the discriminator in the initial target translation network model, and the classification loss L of the true optical image or the pseudo optical image is obtained according to whether the judgment of the discriminator is correct or not. cGAN (D), and then take the expected discrimination results of all images to obtain the classification loss between optical images. The choice of loss function can also be adjusted according to specific needs. In addition to the cross entropy function used in this embodiment, the least square loss, Wasserstein (bulldozer distance) distance and other loss functions can also be used.
[0057] It should be noted that, according to the idea of zero-sum game, the discriminator needs to judge as correctly as possible whether the input image is a true optical image or a pseudo optical image. At the same time, the generator needs to generate an image as close to the real image as possible, so as to make the discriminator make wrong judgments. The generator and discriminator can be trained according to the classification results of the discriminator. Therefore, this part of the loss is regarded as the classification loss.
[0058] For the discriminator D, since y is a real image, D(x, y) should be as close to 1 as possible; and since G(x) is a pseudo-optical image, D(x, G(x)) should be as close to 0 as possible. When the classification is correct, L cGAN (D) will become smaller, otherwise it will become larger, thereby guiding the discriminator D to train.
[0059] Among them, the balanced consistency regularization method is used to obtain the L2 loss of the discriminator, which specifically includes: data enhancement is performed on the SAR image and pseudo-optical image pairs, and the SAR image and real optical image pairs, and the image pairs after data enhancement are input into the improved Vision Transformer network model; the L2 loss requires that the output results of the image pairs after data enhancement are consistent with the unenhanced results, and the formula is:
[0060] L bcR_fake =||D(x,G(x))-D(T(x,G(x)))||2
[0061] L bCR_real =||D(x, y)-D(T(x, y))||2
[0062] Where x represents the SAR image, G represents the generator for generating optical images from SAR images, G(x) represents the pseudo optical image generated by the generator, y represents the real optical image, D represents the discriminator for inputting the SAR image and the optical image pair, D(x, G(x)) and D(x, y) represent the results of the SAR image and the pseudo optical image pair, and the SAR image and the real optical image pair input to the discriminator, respectively, T represents the transformation of data augmentation, and ||·||2 represents the minimum square error between the corresponding pixels of the two images.
[0063] Specifically, considering that the paired SAR image and optical image data sets are limited, in order to prevent the discriminator from overfitting during the training process, the balanced consistency regularization method is applied to perform data enhancement on the SAR image and pseudo-optical image pairs, and the SAR image and true optical image pairs, respectively. The image pairs after data enhancement are input into the Vision Transformer network model, so that the L2 loss requires that the output results of the enhanced image pairs are consistent with the unenhanced results. The balanced consistency regularization method requires that the enhancements applied to the same input image pair produce the same output, that is, L bCR_fake , LbCR_real As small as possible.
[0064] Among them, according to the classification loss and balanced consistency regularization method, the total loss of the discriminator is calculated, and the formula is:
[0065] L(D)=L cGAN (D)+λ bCR_fake L bCR_fake +λ bCR_real L bCR_real
[0066] In the formula, λ bCR_fake , bCR_real It is a configurable hyperparameter; according to the total loss, the back propagation algorithm is used to update the neural network training parameters to optimize the discriminator.
[0067] Specifically, the above process is combined, and the total loss of the discriminator is calculated according to the classification loss and L2 loss of the discriminator. According to the total loss of the discriminator, the back propagation algorithm is used to update the neural network training parameters to optimize the discriminator. In addition, in order to improve the training efficiency, a pre-trained model can be obtained in a large data set.
[0068] The training optimization of the generator specifically includes: calculating the L1 loss of the true optical image and the pseudo optical image, and the classification loss of the generator application according to the SAR image, the true optical image and the pseudo optical image, respectively. The formula is:
[0069] L L1 (G) = E x,y [||yG(x)||1]
[0070] L cGAN (G) = -E x [logD(x,G(x))]
[0071] Where x represents the SAR image, G represents the generator for generating optical images from SAR images, G(x) represents the pseudo optical image generated by the generator, y represents the true optical image, ||·||1 represents the sum of the absolute values of the differences between the corresponding pixels of the two images, and E x,y [·] represents the expectation after calculating the loss for all image pairs (x, y), and the final loss is obtained, E x [·] represents the expectation after calculating the loss of the SAR image; according to the L1 loss and classification loss, the total loss of the generator is calculated as:
[0072] L(G)=L cGAN (G)+λ L1 L L1 (G)
[0073] In the formula, λL1 is a configurable hyperparameter; according to the total loss of the generator, the back-propagation algorithm is used to update the neural network training parameters of the generator to optimize the generator.
[0074] Specifically, according to the SAR image, true optical image and pseudo optical image, the L1 loss function is used to calculate the L1 loss of the true optical image and the pseudo optical image, and the classification loss function is used to obtain the classification loss of the true optical image or the pseudo optical image. The total loss of the generator is calculated according to the L1 loss and the classification loss, and according to the total loss of the generator, the back propagation algorithm is used to update the neural network training parameters of the generator to achieve optimization of the generator.
[0075] Among them, the L1 loss function is also called the minimum absolute deviation and absolute value loss function, which minimizes the sum of the absolute differences between the target value and the estimated value. Among them, the classification loss function can adopt the cross entropy loss function.
[0076] The above steps are performed iteratively. Through continuous iteration, the training optimization of the generator and discriminator is achieved to obtain the optimal generator and discriminator.
[0077] Step S103: input the SAR image to be measured into the target translation network model to obtain the target optical image.
[0078] Specifically, after the target translation network model is trained, the SAR image to be tested is input to generate a corresponding optical image, which can be used to assist in the interpretation of the SAR image or provide additional information.
[0079] In this embodiment, by acquiring a SAR image to be tested, an initial target translation network model is constructed, and the parameters of the initial target translation network model are optimized through paired SAR images and related part images to obtain a target translation network model, which is a model combining Vision Transformer and Pix2Pix, and includes a generator and a discriminator, wherein the generator is used to translate the SAR image into a pseudo-optical image, and the discriminator is used to judge whether the input optical image is a real optical image matched by the SAR image, and the generator and the discriminator complete the neural network training optimization in an adversarial form; the SAR image to be tested is input into the target translation network model to obtain a target optical image, which can take into account the overall structural information of the image, more accurately identify the authenticity of the optical image, generate a clearer optical image, assist in the interpretation of the SAR image, improve the quality of the generated image and the performance of the discriminator, and also ensure the stability of the network training.
[0080] Obviously, those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a computer storage medium (ROM / RAM, magnetic disk, optical disk) and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than that here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Therefore, the present invention is not limited to any specific combination of hardware and software.
[0081] The above contents are further detailed descriptions of the present invention in combination with specific implementation methods, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the scope of protection of the present invention.
Claims
1. An optical image translation method based on ViT-Pix2Pix, characterized in that: The following steps are involved: Acquire the SAR image to be tested; An initial target translation network model is constructed, and parameters of the initial target translation network model are optimized through paired SAR images and optical images to obtain a target translation network model, wherein the target translation network model is a model combining VisionTransformer and Pix2Pix, and includes a generator and a discriminator, wherein the generator is used to translate the SAR image into a pseudo-optical image, and the discriminator is used to determine whether the input optical image is a true optical image matched by the SAR image, and the generator and the discriminator complete the neural network training optimization in an adversarial form; The training optimization of the discriminator specifically includes: inputting the first image pair into the Vision Transformer network model to discriminate the authenticity of the optical image and whether the two images in the first image pair match; merging the two images in the first image pair into a multi-channel input, and dividing them into small blocks of fixed size and non-overlapping, obtaining a linearly arranged embedding through a fully connected layer, and adding a classification symbol at the beginning of the sequence; after adding position information encoding to the linearly arranged embedding, completing the processing in the Transformer encoder of the improved self-attention layer; inputting the output features of the classification symbol into a multi-layer perceptron to complete the discrimination, and obtaining a true optical image and a false optical image; The self-attention layer of the Vision Transformer network model is improved, specifically including: using L2 distance to replace the dot product operation in the self-attention process, and using it to query and input the weight binding of the projection matrix of the self-attention. The improved self-attention layer is calculated as: Where W q =W k , W q , W k and W v are the projection matrices for query, key, and value, respectively. d(·,·) computes the vectorized L2 distance between two sets of points. is the characteristic size of each head; the spectral normalization method is used to optimize the improved VisionTransformer network model; The SAR image to be measured is input into the target translation network model to obtain a target optical image.
2. The optical image translation method based on ViT-Pix2Pix according to claim 1, characterized in that: The step of constructing an initial target translation network and optimizing parameters of the initial target translation network through paired SAR images and optical images to obtain a target translation network specifically includes: Based on Pix2Pix as the basic model, combined with Vision Transformer, the ViT-Pix2Pix initial target translation network model is formed; Optimizing parameters of the initial target translation network model using paired SAR images and optical images; Inputting the SAR image into a generator and outputting a pseudo optical image corresponding to the SAR image; Perform data enhancement on SAR image and true optical image pairs, and SAR image and pseudo optical image pairs; The data-enhanced image pairs are input into the discriminator, which divides the image pairs into small blocks of fixed size and non-overlapping, flattens them into linear embeddings for processing, and outputs the probability that the optical image is a real image and matches the SAR image. The generator and discriminator parameters are optimized through the cross entropy loss function, L1 loss function and balanced consistency regularization method to obtain the target translation model.
3. The optical image translation method based on ViT-Pix2Pix according to claim 2, characterized in that: The step of inputting the SAR image into a generator and outputting a pseudo optical image corresponding to the SAR image specifically includes: Acquire a SAR image as a training sample and a corresponding true optical image as a first image pair; The first image pair is input into a generator, and features are extracted by the generator to obtain a pseudo optical image, where the generator is a U-Net.
4. The optical image translation method based on ViT-Pix2Pix according to claim 3, characterized in that: The training optimization of the generator specifically includes: According to the SAR image, the true optical image and the pseudo optical image, the L1 loss of the true optical image and the pseudo optical image, and the classification loss applied by the generator are calculated respectively, and the formula is: L L1 (G)=E x,y [‖y-G(x)‖1] L cGAN (G)=-E x [logD(x,G(x))] Where, L L1 (G) represents L1 loss, L cGAN (G) represents the classification loss, x represents the SAR image, G represents the generator that generates the optical image from the SAR image, G(x) represents the pseudo optical image generated by the generator, D(x, G(x)) represents inputting x and G(x) into the discriminator D to obtain the discrimination result, y represents the true optical image, ‖·‖1 represents the sum of the absolute values of the differences between the corresponding pixels of the two images, and E x,y [·] represents the expectation after calculating the loss for all image pairs (x, y), and the final loss is obtained, E x [·] represents the expectation after calculating the loss of the SAR image; According to the L1 loss and classification loss, the total loss of the generator is calculated as: L(G)=L cGAN (G)+λ L1 L L1 (G) In the formula, λ L1 is a configurable hyperparameter; According to the total loss of the generator, the back-propagation algorithm is used to update the neural network training parameters of the generator to optimize the generator.
5. The optical image translation method based on ViT-Pix2Pix according to claim 4, characterized in that: The classification loss generated during the discriminator training process is: The improved Vision Transformer network model is applied to the discriminator, and the classification loss between the true optical image and the pseudo optical image is obtained according to the correctness or error of the discriminator's judgment result: L cGAN (D)=-E x,y [logD(x,y)]-E x [1-logD(x,G(x))] Where, L cGAN (D) represents the classification loss of the discriminator, and D(x,y) represents the input of (x,y) into the discriminator D to obtain the discrimination result; The expected correct judgment results of all SAR images are taken to obtain the classification loss between optical images.
6. The optical image translation method based on ViT-Pix2Pix according to claim 5, characterized in that: The balanced consistency regularization method is used to obtain the L2 loss of the discriminator, which includes: Data enhancement is performed on the SAR image and pseudo-optical image pairs, and the SAR image and real optical image pairs, and the image pairs after data enhancement are input into the improved Vision Transformer network model; The L2 loss requires that the output results of the image after data enhancement are consistent with the unenhanced results. The formula is: L bCR_fake =||D(x,G(x))-D(T(x,G(x)))||2 L bCR_real =||D(x,y)-D(T(x,y))||2 Where T represents the data augmentation transformation and ‖·‖2 represents the minimum square error between corresponding pixels of two images.
7. The optical image translation method based on ViT-Pix2Pix according to claim 6, characterized in that: According to the classification loss and balanced consistency regularization method, the total loss of the discriminator is calculated as follows: L(D)=L cGAN (D)+λ bCR_fake L bCR_fake +λ bCR_real L bCR_real In the formula, λ bCR_fake , bCR_real is a configurable hyperparameter; According to the total loss, the back propagation algorithm is used to update the neural network training parameters to achieve the optimization of the discriminator.
Citation Information
Patent Citations
SAR and optical image bidirectional translation method based on cascade residual generative adversarial network
CN111784560A
Methods and systems for model based automatic target recognition in SAR data
US20170350974A1