An image conversion method, system, electronic device and computer storage medium
By introducing the regional differential structure consistency correction module and asymmetric relaxation contrast loss function in image conversion method, the challenges of semantic structure maintenance, positive and negative sample selection and contrast loss function optimization in image conversion are solved, and a higher quality image conversion effect is achieved.
Patent Information
- Application Number
- CN202311333250.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-16
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-10-16
AI Technical Summary
Existing image conversion methods have challenges in maintaining the semantic structure of the target domain image, selecting positive and negative samples, and optimizing the contrast loss function, resulting in limited feature expression capabilities of the image conversion model.
A method of image conversion based on asymmetric relaxation contrast learning is proposed. By constructing a regional differential structure consistency correction module and asymmetric relaxation contrast loss function, the parameters of the generator and discriminator are optimized to improve the image conversion quality.
This method can maintain structural consistency between cross-domain images at a deeper level, improve the efficiency of loss function optimization, and generate higher quality target domain images.
Smart Images

Figure CN117314738B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image conversion, and in particular, to an image conversion method, system, electronic device, and computer storage medium. Background Art
[0002] Unsupervised image conversion usually requires cross-domain learning to maximize the mutual information between source-domain and target-domain images, which requires the generator to stably maintain the source-domain image structure and prevent unnecessary modifications. In recent years, due to the powerful learning ability of self-supervised contrastive learning in cross-domain image feature conversion and modeling, it has been widely used in image conversion tasks, such as: image conversion from a horse to a zebra, seasonal transformation of landscape images, image denoising and dehazing, super-resolution reconstruction, old photo restoration, black-and-white image coloring, makeup on a natural face photo, real image style conversion (converting a real photo to an oil painting or a cartoon, etc.), day-to-night conversion, etc. However, most existing methods are often challenged by maintaining the semantic structure of the generated target-domain image, selecting positive and negative samples, and efficiently optimizing the contrast loss function when applied to image conversion tasks, resulting in limited feature expression ability of the image conversion model. Summary of the Invention
[0003] The purpose of the present invention is to provide an image conversion method, system, electronic device, and computer storage medium, which can improve the conversion quality of images.
[0004] To achieve the above purpose, the present invention provides the following solutions:
[0005] An image conversion method, comprising:
[0006] Obtaining a training data set; the training data set includes a source-domain data set and a target-domain data set;
[0007] Using the source-domain data set and the target-domain data set as inputs of an initial adversarial network, and using the converted target-domain image as the output of the initial adversarial network, training the initial adversarial network by using a generator loss function and a discriminator loss function to obtain an image conversion model; the initial adversarial network includes a generator and a discriminator connected to the generator; the generator includes an encoder, a decoder, and a regional differential structure consistency correction module; the decoder and the regional differential structure consistency correction module are respectively connected to the encoder;
[0008] Performing image conversion on a sample to be converted by using the image conversion model to obtain a converted image.
[0009] Optionally, using the source domain dataset and the target domain dataset as the input of the initial adversarial network, and the transformed target domain image as the output of the initial adversarial network, training the initial adversarial network using the generator loss function and the discriminator loss function to obtain an image conversion model, specifically including:
[0010] Using the encoder to perform image feature encoding on the sample data in the training dataset to obtain the intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images;
[0011] Using the regional differential structure consistency correction module to perform tensor extraction on the intermediate layer image feature maps of the source domain and target domain images to obtain the regional differential tensor and the image feature vector;
[0012] Determining the generator loss function according to the regional differential tensor and the image feature vector;
[0013] Performing decoding processing on the image encoding feature maps of the source domain and target domain images to obtain the target domain image corresponding to the content of the source domain image;
[0014] Using the discriminator to perform authenticity discrimination on the samples in the target domain dataset in the training dataset and the transformed target domain images corresponding to the content of the source domain images to obtain the discriminator loss function;
[0015] Optimizing the parameters of the generator and the discriminator according to the generator loss function and the discriminator loss function to obtain an image conversion model.
[0016] Optionally, optimizing the parameters of the generator and the discriminator according to the generator loss function and the discriminator loss function to obtain an image conversion model, specifically including:
[0017] Judging whether the generator loss function satisfies the first convergence condition to obtain the first judgment result;
[0018] If the first judgment result is yes, then judging whether the discriminator loss function satisfies the second convergence condition to obtain the second judgment result;
[0019] If the second judgment result is yes, then determining the adversarial network at the current iteration as the image conversion model;
[0020] If the second judgment result is no, then return to the step of "Using the encoder to perform image feature encoding on the sample data in the training dataset to obtain the intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images";
[0021] If the first judgment result is negative, return to the step of "using the encoder to perform image feature encoding on the sample data in the training dataset to obtain the intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images".
[0022] Optionally, optimize the parameters of the generator and discriminator according to the generator loss function and the discriminator loss function to obtain an image conversion model, specifically including:
[0023] Judge whether the current iteration number is greater than the set training number to obtain a third judgment result;
[0024] If the third judgment result is negative, return to "using the encoder to perform image feature encoding on the sample data in the training dataset to obtain the intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images";
[0025] If the third judgment result is positive, determine that the adversarial network at the current iteration number is the image conversion model.
[0026] Optionally, the expression of the generator loss function is:
[0027] L(G, D, X, Y) = L RelativeLSGAN (G, D, X, Y) + λ id L id (G) + λ X L contrastive (G, En, X) + λ Y L contrastive (G, En, y)
[0028] Among them, L(G, D, X, Y) represents the generator loss function, X and Y respectively represent the source domain image and the generated target domain image; L RelativeLSGAN (G, D, X, Y) represents the least squares relative adversarial loss function; λ id is the identity loss function weight parameter; L id (G) represents the identity loss; G and D respectively represent the generator and the discriminator; λ X 、λ Y both represent the weight parameters of the loss function; L contrastive (G, En, X) and L contrastive (G, En, Y) respectively represent the image patch multi-layer contrast loss functions of the source domain and the target domain, and En represents the region difference structure consistency correction module.
[0029] The present invention also provides an image conversion system, including:
[0030] An acquisition module, configured to acquire a training dataset; the training dataset includes a source domain dataset and a target domain dataset;
[0031] A training module, configured to use the source domain dataset and the target domain dataset as the input of an initial adversarial network, and the transformed target domain image as the output of the initial adversarial network, and train the initial adversarial network by using a generator loss function and a discriminator loss function to obtain an image conversion model; the initial adversarial network includes a generator and a discriminator connected to the generator; the generator includes an encoder, a decoder, and a regional differential structure consistency correction module; the decoder and the regional differential structure consistency correction module are respectively connected to the encoder;
[0032] An image conversion module, configured to perform image conversion on a sample to be converted by using the image conversion model to obtain a converted image.
[0033] The present invention further provides an electronic device, including:
[0034] One or more processors;
[0035] A storage device, on which one or more programs are stored;
[0036] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method as described above.
[0037] The present invention further provides a computer storage medium, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the method as described above.
[0038] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0039] The present invention obtains a training dataset; the training dataset includes a source domain dataset and a target domain dataset; uses the source domain dataset and the target domain dataset as the input of an initial adversarial network, and the transformed target domain image as the output of the initial adversarial network, and trains the initial adversarial network by using a generator loss function and a discriminator loss function to obtain an image conversion model; the initial adversarial network includes a generator and a discriminator connected to the generator; the generator includes an encoder, a decoder, and a regional differential structure consistency correction module; the decoder and the regional differential structure consistency correction module are respectively connected to the encoder; performs image conversion on a sample to be converted by using the image conversion model to obtain a converted image. The generator loss function is calculated by using the regional differential structure consistency correction module, thereby improving the conversion quality of the image. Description of the Drawings
[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0041] Figure 1 It is a structure diagram of the regional difference structure consistency correction module;
[0042] Figure 2 It is an architecture diagram of the generator;
[0043] Figure 3 It is a comparison chart of the experimental results of the method of the present invention and the baseline method on the horse2zebra dataset;
[0044] Figure 4 It is a comparison chart of the experimental results of the method of the present invention and the baseline method on the Cityscapes dataset;
[0045] Figure 5 It is a comparison chart of the experimental results of the method of the present invention and the baseline method on the monet2photo dataset;
[0046] Figure 6 It is a flowchart of the image conversion method provided by the present invention. Detailed implementation manners
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0048] The purpose of the present invention is to provide an image conversion method, system, electronic device and computer storage medium, which can improve the conversion quality of images.
[0049] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific implementation manners.
[0050] The present invention proposes an image conversion method (Asymmetric Slack Contrastive Learning, ASCo) and system based on asymmetric relaxation contrast learning. While constructing a regional differential structure consistency correction module to solve the local semantic consistency in the image conversion process, a relaxation adjustment factor is introduced, and an asymmetric relaxation contrast loss function is designed in an asymmetric manner. Based on the theoretical analysis of it, its ability to adaptively optimize and adjust the similarity scores of each sample in the latent space is illustrated. This method can maintain the structural consistency between cross-domain images at a deeper level, and has more advantages in establishing a real image domain mapping relationship, can effectively improve the optimization efficiency of the loss function, and thus generate higher-quality target domain images.
[0051] As Figure 6 shown, an image conversion method of the present invention includes:
[0052] Step 101: Obtain a training data set; the training data set includes a source domain data set and a target domain data set.
[0053] Step 102: Using the source domain data set and the target domain data set as the input of the initial adversarial network, and the converted target domain image as the output of the initial adversarial network, training the initial adversarial network using a generator loss function and a discriminator loss function to obtain an image conversion model; the initial adversarial network includes a generator and a discriminator connected to the generator; the generator includes an encoder, a decoder, and a regional differential structure consistency correction module; the decoder and the regional differential structure consistency correction module are respectively connected to the encoder.
[0054] Step 102 specifically includes: using the encoder to perform image feature encoding on the sample data in the training data set to obtain the intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images; using the regional differential structure consistency correction module to perform tensor extraction on the intermediate layer image feature maps of the source domain and target domain images to obtain a regional differential tensor and an image feature vector; determining the generator loss function according to the regional differential tensor and the image feature vector; performing decoding processing on the image encoding feature maps of the source domain and target domain images to obtain the target domain image corresponding to the content of the source domain image; using the discriminator to perform true / false discrimination on the samples of the target domain data set in the training data set and the generated images corresponding to the content of the source domain images to obtain the discriminator loss function; optimizing the parameters of the generator and the discriminator according to the generator loss function and the discriminator loss function to obtain an image conversion model.
[0055] Optimizing the parameters of the generator and the discriminator according to the generator loss function and the discriminator loss function to obtain an image conversion model specifically includes:
[0056] Determine whether the generator loss function satisfies the first convergence condition to obtain a first determination result.
[0057] If the first determination result is yes, then determine whether the discriminator loss function satisfies the second convergence condition to obtain a second determination result.
[0058] If the second determination result is yes, then determine the adversarial network at the current iteration number as the image conversion model.
[0059] If the second determination result is no, then return to the step of "using the encoder to perform image feature encoding on the sample data in the training dataset to obtain the intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images".
[0060] If the first determination result is no, then return to the step of "using the encoder to perform image feature encoding on the sample data in the training dataset to obtain the intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images".
[0061] Alternatively, optimize the parameters of the generator and discriminator according to the generator loss function and the discriminator loss function to obtain an image conversion model, specifically including:
[0062] Determine whether the current iteration number is greater than the set training number to obtain a third determination result. If the third determination result is no, then return to "using the encoder to perform image feature encoding on the sample data in the training dataset to obtain the intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images"; if the third determination result is yes, then determine the adversarial network at the current iteration number as the image conversion model.
[0063] Step 103: Perform image conversion on the sample to be converted using the image conversion model to obtain the converted image.
[0064] The expression of the generator loss function is:
[0065] L(G, D, X, Y) = L RelativeLSGAN (G, D, X, Y) + λ id L id (G) + λ X L contrastive (G, En, X) + λ Y L contrastive (G, En, y)
[0066] Wherein, L(G, D, X, Y) represents the generator loss function, X and Y respectively represent the source domain image and the generated target domain image; L RelativeLSGAN (G, D, X, Y) represents the least squares relative adversarial loss function; λid is the weight parameter of the identity loss function; L id (G) represents the identity loss; G and D represent the generator and discriminator respectively; λ X and λ Y both represent the weight parameters of the loss function; L contrastive (G, En, X) and L contrastive (G, En, Y) represent the multi-layer contrast loss functions of the image patches in the source domain and the target domain respectively, and En is the regional differential structure consistency correction module.
[0067] The present invention also provides a specific working process of an image conversion method, and the steps are as follows:
[0068] Step S1: Construct a training data set, which contains two domains: the source domain X = {x ∈ X} and the target domain Y = {y ∈ Y} (for example, in the conversion of horses and zebras, the source domain is the data set containing horses, and the target domain is the data set containing zebras).
[0069] Step S2: Divide the generator in the generative adversarial network into an encoder, a decoder, and a regional differential structure consistency correction module to obtain an initial adversarial network.
[0070] The generative adversarial network is divided into a generator and a discriminator. The input of the generator is the source domain image sample and the target domain image sample to be converted, and the output is the generated image corresponding to the content of the source domain image (for example, inputting the image of a horse <source domain image to be converted> and the image of a zebra <target domain image> into the horse-to-zebra image conversion model, the corresponding image of a zebra <converted target domain image> can be obtained); the input of the discriminator is the target domain image and the generated target domain image, and the output is the true / false judgment probability values of the two images respectively.
[0071] Divide the generator in the generative adversarial network into an encoder, a decoder, and a regional differential structure consistency correction module:
[0072] The input of the encoder is the source domain image sample and the target domain image sample to be converted, and the output is the intermediate layer image feature map and the image encoding feature map. The input of the decoder is the image encoding feature map output by the encoder, and the output is the converted target domain image corresponding to the content of the source domain image; as Figure 1 shown, the input of the regional differential structure consistency correction module is the intermediate layer image feature map output by the encoder, and the output is the regional differential tensor of the image (used to calculate the regional differential structure consistency correction loss function Δl ReDisc ) and the image feature vector (used to calculate the multi-layer contrast loss function l AS-NCE ). In Figure 1Among them, B, C, H, and W respectively represent the batch size, number of channels, height, and width of the input feature map; L and k_s respectively represent the number of extracted blocks and the size of a single block.
[0073] The specific calculation steps are as follows:
[0074] Step 1: For the regional difference tensor of the image, divide the input image feature map into regions. Given the size of the divided regions, divide the feature map into L parts, and go to Step 2; for the image feature vector, go to Step 5.
[0075] Step 2: According to the block size, continuously extract the neighborhood feature tensor f2_neighbor and the central feature vector f2_centre along the channel dimension from the obtained region-divided feature tensor f1.
[0076] Step 3: The difference calculation is to subtract the neighborhood feature vector from the central vector corresponding to this neighborhood. The obtained regional difference tensor f3 is used for the calculation of the similarity between positive and negative samples.
[0077] Step 4: Input the regional difference tensor into the shared linear layer 2 to obtain the output of the positive and negative sample feature vectors.
[0078] Step 5: Randomly select L feature vectors along the channel dimension from the obtained feature map and input them into the shared linear layer 1 to obtain the output of the positive and negative sample feature vectors.
[0079] Step S3: Use the training data set to train and optimize the initial adversarial network to obtain an image conversion model; for Step S3:
[0080] Step S31: Input the training data set into the initial adversarial network for training, calculate the generator loss value using the generator loss function formula, and calculate the discriminator loss function value using the discriminator loss function.
[0081] Step S32: Based on the generator loss value, determine whether the first convergence condition is satisfied (the first convergence condition is that the difference between the generator network loss function values in two adjacent times is less than the set threshold, or the generator network loss function value is within the set range): if the first convergence condition is satisfied, then execute "Step S33"; if the first convergence condition is not satisfied, then return to "Step S31".
[0082] Step S33: Based on the discriminator loss function value, determine whether the second convergence condition is satisfied (the second convergence condition is that the difference between the discriminator network loss function values in two adjacent times is less than the set threshold, or the discriminator network loss function value is within the set range): if the second convergence condition is satisfied, then use the trained initial adversarial network as the image conversion model; if the second convergence condition is not satisfied, then return to "Step S31".
[0083] Or step S31: the maximum number of training times set;
[0084] Step S32: Input the training data set into the initial adversarial network for training:
[0085] Step S33: Determine whether the number of iterations is less than or equal to the maximum number of training times: If the number of iterations is less than or equal to the maximum number of training times, calculate the generator loss function value using the generator loss function formula and calculate the discriminator loss function value using the discriminator loss function formula. Then, update the network parameters using the Adam optimization algorithm; if the number of iterations is greater than the maximum number of training times, use the trained initial adversarial network as the image conversion model. In the present invention, it is recommended to set the learning rate lr to 0.002, the first-order momentum β1 to 0.5, and the second-order momentum β2 to 0.999.
[0086] The generator loss function is described as:
[0087] L(G, D, X, Y) = L RelativeLSGAN (G, D, X, Y) + λ id L id (G) + λ x L contrastive (G, En, X) + λ Y L contrastive (G, En, Y)
[0088] Wherein, X and Y respectively represent the source image and the generated target domain image;
[0089] L RelativeLSGAN (G, D, X, Y) represents the least squares relative adversarial loss function; λ id is the identity loss function weight parameter, and it is recommended to take 0.5 during the calculation; L id (G) represents the identity loss; G and D respectively represent the generator and the discriminator; λ X 、λ Y represent the weight parameters of the loss function. During the training process, it is recommended to take λ X = λ Y = 1; L contrastive (G, En, X) and L contrastive (G, En, Y) respectively represent the multi-layer contrast loss functions of the image patches in domain X and domain Y.
[0090] 1) For L RelativeLSGAN(G, D, X, Y), the present invention selects PatchGAN as the discriminator model, which divides the input image into multiple image patches of size N×N, makes a true / false judgment on each region, and selects the relative discriminative adversarial loss as the objective of adversarial learning. At the same time, to further improve the quality and authenticity of the generated image, the present invention combines the least squares method, replaces the logarithmic operation with a residual operation, and strictly controls the direction of gradient descent. Therefore, the adversarial loss of the generator of the present invention is the least squares adversarial loss that fuses relative discrimination, which is described as:
[0091]
[0092] Among them, G(x) represents the transformed image obtained after inputting an image x in the source domain of the training set into the generation network. D(G(x)) represents the discrimination probability value obtained after inputting the image G(x) into the discrimination network. represents taking the expectation of the output result of the generator.
[0093] 2) For L contrastive (G, En, X), represents the multi-layer contrast loss function of the image patches in domain X, which is described as:
[0094]
[0095] In the formula, En represents a two-layer multi-layer perceptron network; E x~X represents taking the expectation of all feature calculation results; L represents the number of inputs to the contrast loss calculation layer in the selected generator network; l AS-NCE represents the asymmetric relaxation contrast loss function proposed by the present invention; Δl ReDisc represents the contrast loss function corrected by the regional difference structure consistency proposed by the present invention. The present invention selects L layers from the encoder in the generator network and sends them to two different En networks to calculate the difference vector and the asymmetric relaxation vector Among them represents the output of the l-th layer. represents the difference vector output by the difference structure consistency correction module; represents the asymmetric relaxation vector output by the difference structure consistency correction module; then, index the l∈{1, 2, …, L} layers and define the position encoding serial number s∈{1, 2, …, S l}, where S l represents the total number of spatial positions of each layer. The corresponding features are called positive samples, and the remaining features are called negative samples, and the cross-entropy loss function is used to calculate the two loss functions. The calculation formula for l AS-NCE is as follows.
[0096]
[0097] where s p = v·v + represents the positive sample similarity score; represents the negative sample similarity score; v, v + , v - represent the corresponding query sample, positive sample, and negative sample. ε ∈ (0, 1) represents the relaxation adjustment factor; the temperature coefficient η is recommended to be taken as 100 / 7 during the calculation process.
[0098] For Δl ReDisc the calculation formula is:
[0099]
[0100] where Δs p (positive sample pair) and Δs n (negative sample pair) are the similarity scores calculated for the positive and negative samples of the differential vectors in the original features; τ represents the reciprocal of the temperature coefficient.
[0101] 3) For L contrastive (G, En, Y), which represents the multi-layer contrast loss function of the image patch in domain Y. Similar to 2), only the output image y' of the generator is input into the encoder for re-encoding to obtain and
[0102] 4) For L id (G), considering that since the generator cannot well distinguish the complex image domain information in the image during the unsupervised image conversion process, in order to avoid some irrelevant domain information being tampered with during image conversion, it is necessary to improve the discriminative ability of the generator for the mutually converted image domains. Therefore, the present invention introduces self-reconstruction consistency constraints to further guide the direction of image conversion. Considering that the overall color of the image controlled by pixel-level pixel points is relatively easy to change as irrelevant domain information, in the selection of the loss function, the pixel-level L1 loss function is used as the identity loss of the generator to reduce the change of irrelevant domain features, and its description is:
[0103]
[0104] represents the expectation of the identity loss calculation result.
[0105] In addition, for the discriminator loss function, corresponding to the generator loss function, its description is:
[0106]
[0107] Meanwhile, the multi-layer contrast loss function l of the generator loss functionAS-NCE , and its derivation process is described as follows.
[0108] In the CUT original text, the proposed contrastive loss function is described as follows:
[0109]
[0110] In the formula, l(v, v + , v - ) represents the contrastive loss function in CUT; τ is the reciprocal of the temperature coefficient, which is used to scale the distance between the query and other samples, and is taken as 0.07 during training; sim represents the cosine similarity between the query and other samples, and the calculation formula is as follows:
[0111] is the query sample; u T is the positive / negative sample. Let s p = v · v + and where s p and respectively represent the similarity score between the query sample and the positive sample and the similarity score with the j-th negative sample. The above formula can be transformed into:
[0112]
[0113] In the formula, η = 1 / τ, which represents the temperature coefficient. Let Y 1 = exp(ηs p ), The gradients of the contrastive loss function with respect to s p and are respectively:
[0114]
[0115]
[0116] For N negative samples, the sum of the gradients of the loss function for all negative samples is:
[0117]
[0118] That is:
[0119]
[0120] The above formula shows that the original contrastive loss function PatchNCE optimizes the similarity between positive sample pairs and negative sample pairs in an anti-symmetric manner (the absolute values of the gradients are the same, and the ultimate goal of optimization is to make s p larger, Smaller), the "closeness" between the query sample and the positive sample is equal to the "distance" from the negative sample.
[0121] In the actual optimization process, the similarity score s between positive sample pairs p reflects the similarity between samples of the same type, so the larger the better, while the similarity score between negative sample pairs should be the smaller the better. Therefore, to effectively distinguish between positive and negative similarities, the present invention believes that s p and can "pull closer" and "push away" positive and negative samples adaptively according to their own magnitudes, that is, for positive samples, when s p has a very large value, it means that the query sample is very close to the positive sample. At this time, the optimization degree can be appropriately reduced during optimization; while when s n has a relatively large value, at this time the query sample is relatively close to the negative sample. To achieve the purpose of "pushing away" the negative sample, it is necessary to increase the adjustment of the loss function for negative similarity during the optimization process. Therefore, the present invention introduces an asymmetric adjustment weight β p and to adjust the parameter gradients of positive and negative similarities respectively during the optimization process, and constructs an asymmetric contrast loss AsymNCE function accordingly:
[0122]
[0123] To reduce the additional impact of the asymmetric weight hyperparameters β p and on network training, according to their properties, β p is defined as the negative correlation function of the similarity score s p of positive sample pairs, and is defined as the positive correlation function of the similarity score of negative sample pairs; at the same time, considering that the cosine similarity calculation method is generally used when calculating the similarity score, and the similarity scores between positive and negative sample pairs are in the range of 0 to 1, therefore, let the asymmetric weight hyperparameters β p and be:
[0124]
[0125] Since the positive and negative samples are selected randomly, to tolerate pseudo-negative examples in the feature space and other cases where the separability of other feature representations is weak, the present invention introduces a relaxation variable ε to reduce the sensitivity of the gradient weight parameter to the similarity score. Therefore, the asymmetric hyperparameters β p and with relaxation terms are defined as:
[0126]
[0127] where ε ∈ (0, 1), λ p is the asymmetric relaxation adjustment factor for the positive similarity score, is the asymmetric relaxation adjustment factor for the negative similarity score. Substituting the asymmetric weight adjustment factor with the relaxation term gives the asymmetric relaxation contrastive loss AS-NCE function proposed in the present invention:
[0128]
[0129] From this analysis, after introducing the asymmetric weight factor with the relaxation variable, the optimization of the AS-NCE contrastive loss function for positive and negative similarities is carried out in an asymmetric manner. When the positive similarity score s p is closer to 1 and the negative similarity score is closer to 0, the corresponding weight update factor can dynamically decrease, making the optimization process smoother. In addition, by comparing the calculation formula of the AS-NCE loss function with that of the original InfoNCE loss function, it can be seen that InfoNCE directly uses the inner product of the MLP output as the factor for calculating the loss function value, while AS-NCE uses a quadratic polynomial after the inner product calculation. Therefore, in terms of computational complexity, AS-NCE is equal to the original InfoNCE function, that is, using AS-NCE as the loss function of the network model will not increase additional time costs.
[0130] The region difference structure consistency correction loss function Δl of the generator loss function ReDisc , and its derivation process is described as follows.
[0131] Under the contrastive learning idea, the contrastive loss function in CUT maximizes the mutual information between the input source domain image patches and the corresponding generated target domain image patches. Although this can solve the problem of content preservation in the cycle consistency loss, enabling the current unsupervised image conversion method to establish a good mapping relationship between different image domains, there are still some problems. First, in order to improve the performance of the generation network, there are certain limitations on the size of the convolutional kernels in the convolutional operations in the network; second, in the process of calculating the PatchNCE contrastive loss function, positive and negative samples are selected at the pixel level. Therefore, during the training process, the generated images can only focus on local images and will ignore the global correlation, resulting in incomplete transformation of the converted images in the specified image domain, and the overall coordination and authenticity of the images are poor.
[0132] On this basis, the present invention believes that the global consistency of the generated image can be improved by local consistency, that is, by a larger range of local perception, to find a similar relationship between the change in the feature information represented between the neighboring positions of the generated image during the image conversion process and the information change at the same position in the source domain image. Through such an adjustment of local neighbor change similarity, local distortion is reduced to improve the visual quality of the image. However, considering that increasing the receptive field of the convolutional kernel in the convolutional operation will increase the number of model parameters, therefore, the present invention designs a regional difference structure consistency calculation module (ReDisc_Block) to increase the local perception ability of the model with the result of slightly increasing the number of model parameters. The calculation steps are as follows:
[0133] 1) Divide the obtained feature map into regions. Given the size of the divided regions, the feature map is divided into L parts, as shown in the following formula.
[0134]
[0135] In the formula, represents the regional feature map after division; F (B,C,H,W) is the feature map; Regional(·) represents the grouping operation of feature vectors; B, C, H, and W respectively represent the batch size, number of channels, height, and width of the input feature map.
[0136] 2) Take out the neighborhood feature tensor f2_neighbor and the central feature vector f2_centre continuously along the channel dimension from the obtained region-divided feature tensor f1 according to the block size, as shown in the following formula.
[0137]
[0138]
[0139] In the formula, is the neighborhood feature tensor; is the central feature vector; The Unfolg(·) function represents the block extraction function along the channel dimension; L_s and k_s respectively represent the number of extracted blocks and the size of a single block.
[0140] 3) The differential calculation is to subtract the neighborhood feature vector from the central vector corresponding to the neighborhood, and the obtained regional difference tensor f3 is used for the calculation of positive and negative sample similarity, as shown in the following formula.
[0141]
[0142] In the formula, the Differential(·) function represents the differential calculation; is the regional difference tensor.
[0143] Based on the view that the differential vectors at the same position should be the most relevant in the latent space, in the ReDisc module, the mutual information between "positive" differential vector pairs is maximized, only between a pair of source domain differential tensors and the differential vectors at the same position in the target domain, while the remaining differential vectors can be regarded as "negative" samples.
[0144] Let the region differential vectors calculated from the source domain images and the generated target domain images be f 3-source and f 3 -target respectively, and they are projected by a two-layer MLP network to calculate the corresponding similarity scores. When using the similarity of the region differential vector structure consistency for correction, the similarity scores calculated for the positive and negative samples in the original features are Δs p (positive sample pair), Δs n (negative sample pair); in the calculation formula of the contrast loss function, let s p = v·v + and s n = v·v - . Therefore, calculate the block image structure consistency matrix of each feature map and the corresponding contrast loss function value, as the correction term of the original PatchNCE contrast loss function, that is, the similarity correction contrast loss function of the feature map region differential structure consistency proposed by the present invention. The corrected calculation formula is:
[0145]
[0146] where λ ReDisc is the correction parameter, which is taken as 0.5 in this experiment; l ReDiscNCE (v, v + , v - ) is the similarity correction contrast loss function of the feature map region differential structure consistency after correction.
[0147] By regionally dividing the feature maps obtained during the convolution process, on the basis of calculating the random region differential structure similarity of cross-domain images, the PatchNCE contrast loss function is corrected. By adding the structure similarity loss function value of local features in addition to the pixel-level similarity score, the generated target domain images can pay attention to the correlation within the local range of the image on a larger scale and maintain the consistency of the image spatial structure within the local feature range with the source domain images, thereby improving the coordination and authenticity of the generated target domain images.
[0148] Step S4: Input the sample to be converted into the image conversion model for image conversion to obtain the converted image (for example, input the image of a horse into the horse-to-zebra image conversion model to obtain the corresponding image of a zebra).
[0149] The present invention also provides an image conversion system, including:
[0150] An acquisition module, configured to acquire a training data set; the training data set includes a source domain data set and a target domain data set.
[0151] A training module, configured to use the source domain data set and the target domain data set as inputs of an initial adversarial network, and use a converted target domain image as an output of the initial adversarial network, and train the initial adversarial network by using a generator loss function and a discriminator loss function to obtain an image conversion model; the initial adversarial network includes a generator and a discriminator connected to the generator; the generator includes an encoder, a decoder, and a regional differential structure consistency correction module; the decoder and the regional differential structure consistency correction module are respectively connected to the encoder.
[0152] An image conversion module, configured to perform image conversion on a sample to be converted by using the image conversion model to obtain a converted image.
[0153] The present invention also provides an electronic device, including: one or more processors; a storage device storing one or more programs thereon; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method as described above.
[0154] The present invention also provides a computer storage medium, having a computer program stored thereon, wherein when the computer program is executed by a processor, the method as described above is implemented.
[0155] Methods such as NEGCUT and F-LSeSim are improvements to the CUT method to make full use of contrastive learning. However, NEGCUT requires additional training for the negative sample generator from random vectors, which may not guarantee following the true negative sample distribution, thus bringing the possibility of unstable training; F-LSeSim only focuses on the structural features of images, ignoring the semantic relationships between patches and treating all negative sample patches as equal negative samples. The AsCo proposed in the present invention comprehensively considers these two types of problems, namely: first, the semantic structure preservation of the generated image, how to better utilize the effective information between the given source domain image and the target domain image to avoid local distortion of the generated image and obtain better generation quality; second, the selection of positive and negative samples and the optimization of the contrastive loss function, how to correctly handle the relationship between positive and negative samples to ensure more efficient optimization of the contrastive loss function and generate a more satisfactory target domain image. The Regional Differential Structural Consistency Block (ReDisc_Block) and the Asymmetric Slack Contrastive Learning (ASCL) are respectively designed to improve the above problems.
[0156] Figure 2 In (a) is the regional differential structural consistency calculation module proposed in the present invention, which performs differential calculation on the regional division of the feature map and introduces the differential vector into a shared MLP network to achieve greater local structural consistency of the image; Figure 2 Part (b) is also a two-layer shared MLP network (different from the previous MLP). After obtaining the feature vector, the AS-NCE loss function value can be calculated and combined with the ReDiscNCE loss function value as the total contrastive loss function of the model. Figure 2 Part (c) is the backbone network of the generator in the unsupervised image conversion task, consisting of an encoder and a decoder, and is used to generate and output the generated image.
[0157] The present invention also provides an application of the image conversion method, a method for training an image conversion model, including:
[0158] Obtaining sample data, where the sample data includes: source domain images and target domain images;
[0159] Inputting the source domain image and the target domain image into a preset unsupervised image conversion model of original asymmetric slack contrastive learning, where the unsupervised image conversion model of original asymmetric slack contrastive learning includes: a generator (encoder + decoder + regional differential structural consistency correction module) and a discriminator;
[0160] Each of the encoders performs image feature encoding on the source domain image samples and target domain image samples to be converted, obtaining intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images;
[0161] The regional differential structure consistency correction module performs tensor extraction on the intermediate layer image feature maps of the source domain and target domain, obtaining a regional differential tensor and an image feature vector, and calculating a regional differential structure consistency correction loss function and an image patch multi-layer contrast loss function. Among them, the regional differential tensor is used to calculate the regional differential structure consistency correction loss function Δl ReDisc and the image feature vector is used to calculate the image patch multi-layer contrast loss function l AS-NCE ;
[0162] The decoder performs decoding processing on the image encoding feature maps of the source domain and target domain images, obtaining a generated image corresponding to the source domain image content;
[0163] The discriminator performs true / false discrimination on the target domain sample images and the images generated by the decoder, obtaining the adversarial loss function value of the generator and the discriminator;
[0164] According to the loss function, the parameters of the generator and the discriminator are optimized and adjusted, obtaining an image conversion model after one training;
[0165] According to the source domain images and the target domain sample image dataset, the original image conversion model is iteratively trained and parameter adjusted. When the iteration end condition is satisfied (reaching the maximum number of training times), a target image conversion model is obtained;
[0166] The target image data to be converted is input into the target image conversion model for image conversion, obtaining a target image.
[0167] The regional differential structure consistency correction module includes:
[0168] A regional differential tensor extraction network and an image feature vector extraction network;
[0169] The regional differential tensor extraction network performs regional differential tensor extraction on the intermediate layer image feature maps output by the encoder, for calculating the regional differential structure consistency correction loss function value;
[0170] The image feature vector extraction network performs image feature vector extraction on the intermediate layer image feature maps output by the encoder, for calculating the image patch multi-layer contrast loss function value.
[0171] The regional differential tensor extraction network includes:
[0172] Two layers of weight-sharing linear layer neural networks;
[0173] Divide the image feature map of the intermediate layer output by the encoder into regions. Given the size of the divided regions, divide the feature map into L parts to obtain the region-divided feature tensor f1.
[0174] Continuously extract the neighborhood feature tensor f2_neighbor and the central feature vector f2_centre along the channel dimension from the region-divided feature tensor f1 according to the block size.
[0175] Subtract the neighborhood feature tensor from its corresponding central feature vector to obtain a region difference tensor for positive and negative sample similarity calculation;
[0176] The obtained region difference tensor is input into the shared linear layer for neural network calculation to obtain positive and negative sample feature vectors as output for calculating the region difference structure consistency correction loss function value;
[0177] The image feature vector extraction network includes:
[0178] Two layers of weight-sharing linear layer neural networks (different from those in the region difference tensor extraction network);
[0179] Randomly select L feature vectors from the image feature map of the intermediate layer output by the encoder along the channel dimension and input them into the shared linear layer to obtain positive and negative sample feature vectors for calculating the image patch multi-layer contrast loss function value.
[0180] Adjusting the parameters of the original image conversion model according to the source domain image and the target domain sample image to obtain the target image conversion model includes:
[0181] Calculating the loss between the target sample image and the reference image through a preset loss function to obtain the target loss function value;
[0182] Adjust the parameters of the original image conversion model according to the target function value to obtain the target image conversion model.
[0183] The loss function includes:
[0184] The least squares relative discriminative adversarial loss function, identity loss function of the generator and discriminator, and the asymmetric relaxation contrast loss function between domain X and domain Y, where the asymmetric relaxation contrast loss function includes the region difference structure consistency correction loss function Δl ReDisc and the image patch multi-layer contrast loss function l AS-NCE .
[0185] Calculating the loss function between the source domain image and the target domain image through a preset loss function to obtain the target loss function value, including:
[0186] Calculate the adversarial loss of the target sample image and the reference image through the least - squares relative discriminative adversarial loss function of the generator and the discriminator to obtain the adversarial loss function value of the generator and the discriminator;
[0187] Calculate the identity loss of the target sample image and the reference image through the identity loss function to obtain the identity loss function value of the source - domain image;
[0188] Calculate the contrast loss of the target sample image and the reference image through the asymmetric relaxation contrast loss function to obtain the contrast loss function value of the source - domain image and the target - domain image;
[0189] Sum up the adversarial loss function value, the identity loss function value, and the asymmetric relaxation contrast loss function value according to the preset weight to obtain the objective function value.
[0190] The least - squares relative discriminative adversarial loss function of the generator and the discriminator is calculated as follows:
[0191] The present invention selects PatchGAN as the discriminator model. It divides the input image into multiple image patches of size N×N, makes true - false judgments on each region, and selects the relative discriminative adversarial loss [6] As the purpose of adversarial learning, at the same time, to further improve the quality and authenticity of the generated image, the present invention combines the least - squares method, replaces the logarithmic operation with the residual operation, and strictly controls the direction of gradient descent. Therefore, the adversarial loss of the generator in the present invention is the least - squares adversarial loss that fuses relative discrimination, which is described as:
[0192]
[0193] where \(G(x)\) represents the transformed image obtained after inputting an image \(x\) in the source domain of the training set into the generation network; \(D(G(x))\) represents the discrimination probability value obtained after inputting the image \(G(x)\) into the discrimination network; \(E\) represents the expected value.
[0194] The adversarial loss of the discriminator corresponds to the generator loss function, which is described as:
[0195]
[0196] The identity loss function includes:
[0197] Considering that the generator cannot well distinguish complex image domain information in the network during unsupervised image conversion, when converting images, in order to avoid tampering with some irrelevant domain information, it is necessary to improve the discriminative ability of the generator for the mutually converted image domains. Therefore, the present invention introduces self-reconstruction consistency constraints to further guide the direction of image conversion. Considering that the overall color of the image controlled by pixel-level pixel points is relatively easy to change as irrelevant domain information, in the selection of the loss function, the pixel-level L1 loss function is used as the identity loss of the generator to reduce the change of irrelevant domain features, which is described as:
[0198]
[0199] The asymmetric relaxation contrast loss function of the domain X includes:
[0200] The multi-layer contrast loss function of the image patches in the domain X, which is described as:
[0201]
[0202] In the formula, En represents a two-layer multi-layer perceptron network; L represents the number of inputs from the selected generator network to the contrast loss calculation layer; l AS-NCE represents the asymmetric relaxation contrast loss function proposed by the present invention; Δl ReDisc represents the contrast loss function corrected by the regional difference structure consistency proposed by the present invention. The present invention selects L layers from the encoder in the generator network and sends them to two different MLP networks to calculate the difference vector and the asymmetric relaxation vector where represents the output of the l-th layer. Then, the layers l ∈ {1, 2,..., L} are indexed, and s ∈ {1, 2,..., Sl} is defined, where S l represents the total number of spatial positions in each layer. The corresponding features are called positive samples, and the remaining features are called negative samples. The cross-entropy loss function is used to calculate the two loss functions.
[0203] The asymmetric relaxation contrast loss function of the domain Y corresponds to the domain, except that the output image y′ of the generator is input into the encoder for re-encoding to obtain and
[0204] The present invention proposes an image conversion algorithm based on asymmetric slack contrastive learning (ASCo). While constructing a regional differential structure consistency correction module to solve the local semantic consistency during the image conversion process, a relaxation adjustment factor is introduced, and an asymmetric slack contrastive loss function is designed in an asymmetric manner. Based on the theoretical analysis of it, its ability to adaptively optimize and adjust the similarity scores of each sample in the latent space is illustrated.
[0205] Specifically, it is reflected in the training and optimization of the initial adversarial network in step S3, and the L contrastive (G, En, X) and L contrastive (G, En, Y) are designed to maintain the structural consistency between cross-domain images at a deeper level, and have more advantages in establishing a real image domain mapping relationship, which can effectively improve the optimization efficiency of the loss function, and then generate target domain images of higher quality. The loss function formula is as follows:
[0206]
[0207]
[0208] The advantages of the present invention are as follows:
[0209] (1) A regional differential structure consistency correction module is designed. To expand the local perception range of the model, by finding the similar relationship between the change of the feature information represented between adjacent positions on the generated image in the image conversion task and the information change at the same position in the source domain image, local distortion is reduced to improve the visual quality of the image;
[0210] (2) An asymmetric slack contrastive loss function is constructed. By introducing a relaxation adjustment factor and in an asymmetric way, it can adaptively determine the corresponding optimization amplitude according to the size of the similarity scores in the latent spaces of positive and negative samples to improve the optimization efficiency and generate target domain images of higher quality;
[0211] (3) An image conversion algorithm model based on asymmetric slack contrastive learning and structure correction is constructed. By using the asymmetric slack contrastive loss and regional differential structure consistency correction, higher-quality image generation is achieved in the image conversion task;
[0212] (4) Six existing algorithms are compared on the publicly available image conversion task dataset to evaluate the model proposed by the present invention. The experimental results and a large number of ablation experiment analyses verify the effectiveness and advancement of the proposed method.
[0213] The present invention conducts experiments on the following three datasets:
[0214] (1) horse2zebra: This dataset contains images of horses and zebras from the ImageNet image dataset, with 1,067 horse images and 1,334 zebra images respectively, and was first used in CycleGAN. The image resolution is 256×256;
[0215] (2) Cityscapes: This dataset contains street views from German cities, with a total of 2,975 training images and 500 validation images. The model was trained at a resolution of 256×256 in the experiment. Different from other datasets, this dataset has corresponding labels. Therefore, this dataset can be used to measure the ability of the algorithm of the present invention to capture the structural correspondence between images in the dataset;
[0216] (3) monet2photo: This dataset includes two types of images, Monet's oil painting style images and landscape style images taken by cameras. The Monet oil painting style training set in the training set consists of 1,337 images, and the landscape style training set consists of 3,671 images. The Monet oil painting style test set in the test set contains 271 images, and the landscape style test set has 751 images. The large resolution of all images is 256×256.
[0217] The following provides data for verification.
[0218] The experiment is divided into two stages: training and testing.
[0219] During training, the source domain and target domain images in the training dataset are first scaled to 286×286, and then randomly cropped to 256×256 to enhance the robustness of the model. In the experiment, the Adam optimizer is used as the optimization algorithm for gradient descent (Momentum parameters are 0.5 and 0.999), the learning rate is set to 0.002, and the batch size is set to 1. The entire training process is trained for up to 400 epochs. However, during the experiment, the method proposed in the present invention will reach the best result in advance around the 300th epoch.
[0220] In the testing stage, only the trained generator model needs to be used, and inputting the test dataset can obtain the generated images that conform to the target domain images.
[0221] For the quality evaluation of the generated images, the present invention uses the Frechet Inception Distance (FID) as the evaluation criterion and evaluates it separately on all datasets. For the Cityscapes dataset, the present invention performs a conversion from the labeled images, and thereby can also measure the correspondence between the segmented images of the output images and their true segmented images. Specifically as follows: First, use the DRN model to train the semantic segmentation network, and calculate the mean average precision (mAP), pixel accuracy (pixAcc), and mean class accuracy (classAcc). The batch size in the DRN network is 32, the learning rate is 0.01, and it is trained for 250 epochs on images with a resolution of 256×128. After that, use bicubic downsampling to resize the images generated from 500 test labels to 256×128, and input them into the trained DRN network, and compare them with the ground-truth real images downsampled to the same size using the nearest neighbor sampling method.
[0222] Table 1 shows the evaluation results of the present method and all baselines on two datasets, horse2zebra and Cityscapes, and their visual effects are shown as Figure 3 and Figure 4 shown. Figure 5 Shows the qualitative comparison results of the single-sample conversion of the present method and the state-of-the-art unpaired method on the monet2photo dataset.
[0223] Table 1 Comparison of the method of the present invention with all baselines
[0224]
[0225] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0226] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An image conversion method, characterized in that, it includes: Obtain a training data set; The training data set includes a source domain data set and a target domain data set; Using the source domain data set and the target domain data set as the input of the initial adversarial network, and the converted target domain image as the output of the initial adversarial network, training the initial adversarial network using a generator loss function and a discriminator loss function to obtain an image conversion model; the initial adversarial network includes a generator and a discriminator connected to the generator; the generator includes an encoder, a decoder, and a regional differential structure consistency correction module; the decoder and the regional differential structure consistency correction module are respectively connected to the encoder; Perform image conversion on the sample to be converted using the image conversion model to obtain the converted image; The encoder is used to perform feature encoding on the input image sample and output an intermediate layer feature map and an encoded feature map of the image; the decoder is used to convert the encoded feature map of the image obtained by the encoder into a target image; the regional differential structure consistency correction module is used to perform deep feature extraction on the intermediate layer feature map of the image obtained by the encoder and output a regional differential tensor and an image feature tensor of the image; The expression of the generator loss function is: L(G, D, X, Y) = L RelativeLSGAN (G, D, X, Y) + λ id L id (G) +λ X L contrastive (G, En, X) + λ Y L contrastive (G, En, Y) Among them, \(L(G, D, X, Y)\) represents the generator loss function, where \(X\) and \(Y\) represent the source domain image and the generated target domain image respectively; \(L\) RelativeLSGAN (G, D, X, Y) represents the least squares relative adversarial loss function; \(\lambda\) id is the weight parameter of the identity loss function; \(L\) id (G) represents the identity loss; \(G\) and \(D\) represent the generator and the discriminator respectively; \(\lambda\) X , \(\lambda\) Y both represent the weight parameters of the loss function; \(L\) contrastive (G, En, X) and \(L\) contrastive (G, En, Y) represent the multi-layer contrast loss functions of image patches in the source domain and the target domain respectively, including the asymmetric relaxation contrast loss function and the contrast loss function corrected by the regional differential structure consistency; \(En\) represents the multi-layer perceptron network in the differential structure consistency correction module; For L contrastive (G, En, X), which represents the multi-layer contrast loss function of the image block in the domain X, is described as follows: wherein, En represents a multi-layer perceptron network with two layers; E x~X represents the expectation of the calculation results of all features; L represents the number of inputs from the selected generator network to the contrast loss calculation layer; l AS-NCE represents an asymmetric relaxation contrast loss function; Δl ReDisc represents a contrast loss function for region difference structure consistency correction; For L contrastive (G, En, Y), which represents the multi-layer contrast loss function of the image block in domain Y, is implemented in the same way as L contrastive (G, En, X); similarly The adversarial loss of the generator is the least squares adversarial loss that fuses relative discrimination, which is described as: Among them, \(G(x)\) represents the transformed image obtained after inputting an image \(x\) in the training set source domain into the generation network; \(D(G(x))\) represents the discrimination probability value obtained after inputting the image \(G(x)\) into the discrimination network; denotes taking the expectation of the output result of the generator; For the discriminator loss function, corresponding to the generator loss function, it is described as:
2. The image conversion method according to claim 1, characterized in that, Using the source domain data set and the target domain data set as the input of the initial adversarial network, and the converted target domain image as the output of the initial adversarial network, training the initial adversarial network using a generator loss function and a discriminator loss function to obtain an image conversion model, specifically including: Using the encoder to perform image feature encoding on the sample data in the training data set to obtain intermediate layer image feature maps and encoded feature maps of the source domain and target domain images; Using the regional differential structure consistency correction module to perform tensor extraction on the intermediate layer image feature maps of the source domain and target domain images to obtain a regional differential tensor and an image feature vector; Determine the generator loss function according to the regional differential tensor and the image feature vector; Perform decoding processing on the encoded feature maps of the source domain and target domain images to obtain a target domain image corresponding to the content of the source domain image; Using the discriminator to perform true / false discrimination on the samples of the target domain data set in the training data set and the converted target domain images corresponding to the content of the source domain images to obtain the discriminator loss function; Optimize the parameters of the generator and the discriminator according to the generator loss function and the discriminator loss function to obtain an image conversion model.
3. The image conversion method according to claim 2, characterized in that, Optimize the parameters of the generator and the discriminator according to the generator loss function and the discriminator loss function to obtain an image conversion model, specifically including: Judge whether the generator loss function satisfies the first convergence condition to obtain a first judgment result; If the first judgment result is yes, then judge whether the discriminator loss function satisfies the second convergence condition to obtain a second judgment result; If the second judgment result is yes, then determine that the adversarial network at the current iteration number is an image conversion model; If the second judgment result is no, then return to the step "using the encoder to perform image feature encoding on the sample data in the training dataset to obtain the intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images"; If the first judgment result is no, then return to the step "using the encoder to perform image feature encoding on the sample data in the training dataset to obtain the intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images".
4. The image conversion method according to claim 2, wherein, optimizing the parameters of the generator and the discriminator according to the generator loss function and the discriminator loss function to obtain an image conversion model, specifically including: judging whether the current iteration number is greater than the set training number to obtain a third judgment result; If the third judgment result is no, then return to "using the encoder to perform image feature encoding on the sample data in the training dataset to obtain the intermediate layer image feature maps and image encoding feature maps of the source domain and target domain images"; If the third judgment result is yes, then determine that the adversarial network at the current iteration number is an image conversion model.
5. An image conversion system, wherein, comprising: an acquisition module, configured to acquire a training dataset; the training dataset includes a source domain dataset and a target domain dataset; a training module, configured to use the source domain dataset and the target domain dataset as the input of an initial adversarial network, and use the converted target domain image as the output of the initial adversarial network, and train the initial adversarial network by using a generator loss function and a discriminator loss function to obtain an image conversion model; the initial adversarial network includes a generator and a discriminator connected to the generator; the generator includes an encoder, a decoder, and a regional differential structure consistency correction module; the decoder and the regional differential structure consistency correction module are respectively connected to the encoder; an image conversion module, configured to perform image conversion on a sample to be converted by using the image conversion model to obtain a converted image; The expression of the generator loss function is: L(G, D, X, Y) = L RelativeLSGAN (G, D, X, Y) + λ id L id (G) +λ X L contrastive (G, En, X) + λ Y L contrastive (G, En, Y) Among them, \(L(G, D, X, Y)\) represents the generator loss function, where \(X\) and \(Y\) represent the source domain image and the generated target domain image respectively; \(L\) RelativeLSGAN (G, D, X, Y) represents the least squares relative adversarial loss function; \(\lambda\) id is the weight parameter of the identity loss function; \(L\) id (G) represents the identity loss; \(G\) and \(D\) represent the generator and discriminator respectively; \(\lambda\) X , \(\lambda\) Y both represent the weight parameters of the loss function; \(L\) contrastive (G, En, X) and \(L\) contrastive (G, En, Y) represent the multi-layer contrast loss functions of image patches in the source domain and target domain respectively, including the asymmetric relaxation contrast loss function and the contrast loss function corrected by regional differential structure consistency; \(En\) represents the multi-layer perceptron network in the differential structure consistency correction module; For L contrastive (G, En, X), representing the multi-layer contrast loss function of the image block in the domain X, is described as follows: Wherein, En represents a two-layer multi-layer perceptron network; E x~X represents the expectation of the calculation results of all features; L represents the number of inputs from the selected generator network to the contrast loss calculation layer; l AS-NCE represents the asymmetric relaxation contrast loss function; Δl ReDisc represents the contrast loss function for regional differential structure consistency correction; For L contrastive (G, En, Y), which represents the multi-layer contrast loss function of the image block in domain Y, is implemented in the same way as L contrastive (G, En, X); the same applies to The adversarial loss of the generator is the least squares adversarial loss that fuses relative discrimination, which is described as: Among them, G(x) represents the transformed image obtained after inputting an image x in the training set source domain into the generation network; D(G(x)) represents the discrimination probability value obtained after inputting the image G(x) into the discrimination network; represents taking the expectation of the output result of the generator; For the discriminator loss function, corresponding to the generator loss function, it is described as:
6. An electronic device, wherein, comprising: one or more processors; a storage device, on which one or more programs are stored; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any one of claims 1 to 4.
7. A computer storage medium, wherein, a computer program is stored thereon, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Image conversion method and system
CN114331821A