A high-quality image generation method and device

By improving the discriminator structure and neighbor comparison loss function of the CycleGAN model, the diversity and detail fidelity of the generated images are improved, and the balance problem between diversity and quality of the existing models is solved, and high-quality image generation effect is achieved.

CN120182424BActive Publication Date: 2025-08-05NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510660976.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-08-05
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The existing CycleGAN model has shortcomings in generating image diversity and detailed expression, especially in tasks such as seasonal changes and facial expression migration. The generated images lack rich diversity and detailed texture fidelity, resulting in the generated samples becoming single and the pattern crash problem is obvious.

Method used

Improve the discriminator structure of the CycleGAN model, introduce a multi-scale feature fusion module and a channel attention module, and combine the neighbor comparison loss function to calculate the total model loss by counter-loss, cyclic consistency loss and neighbor comparison loss, improving the diversity of generated images and local detail fidelity.

Benefits of technology

The generated images can present diverse style characteristics, maintain good authenticity in details, ensure the stability of image quality, and are suitable for image conversion application scenarios such as seasonal conversion and facial expression migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182424B_ABST
    Figure CN120182424B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-quality image generation method and device in the field of image generation technology. The generation method includes: obtaining an image to be processed; obtaining a generated image with a target image style by using a pre-trained image generation model according to the image to be processed; the image generation model adopts an improved CycleGAN model; the discriminator of the improved CycleGAN model includes a multi-scale feature fusion module and a channel attention module; the improved CycleGAN model calculates the total model loss through adversarial loss, cycle consistency loss and neighbor contrast loss. The present invention can improve the diversity and detail fidelity of the generated image while maintaining the quality of the generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a high-quality image generation method and apparatus, belonging to the technical field of image generation. Background Art

[0002] Currently, with the development of Generative Adversarial Networks (GAN) technology, significant progress has been made in the image-to-image translation task. As a classic unsupervised learning framework, Cyclic Generative Adversarial Networks (CycleGAN) is suitable for inter-domain image conversion without paired data. Although the CycleGAN model preserves the content features of the input image when generating target-domain images, there are still deficiencies in diversity and detail expression. Especially when applied to image conversion tasks such as seasonal changes and facial expression transfer, the generated images often lack rich diversity, and the fidelity of detail textures is poor, resulting in a tendency for generated samples to be single and obvious mode collapse problems.

[0003] To address the limitations of CycleGAN in generating image diversity, current methods mainly improve the loss function, such as using neighbor constraints in the latent space to achieve diverse generation of image content while maintaining consistent image styles. However, existing improvement methods often affect the fidelity of images to a certain extent, making the quality of generated images unstable.

[0004] How to balance the diversity of image generation and the stability of image quality is the next research focus in the technical field of image generation. Summary of the Invention

[0005] The object of the present invention is to provide a high-quality image generation method and apparatus. On the one hand, the discriminator of the CycleGAN model is improved by adding multi-scale feature fusion and channel attention mechanisms. On the other hand, the neighbor contrast loss function is introduced into the traditional CycleGAN framework, and the parameters of the neighbor contrast loss function are improved to perform fine-grained similarity constraints on the generated image and the target-domain image in the feature space, effectively enhancing the diversity of the generated image and the fidelity of local details, and achieving a better balance between the diversity and quality of the generated image.

[0006] To achieve the above object, the present invention is implemented by the following technical solutions:

[0007] On the one hand, the present invention provides a high-quality image generation method, including:

[0008] Obtain the image to be processed;

[0009] According to the image to be processed, a generated image with the style of the target image is obtained by using a pre-trained image generation model;

[0010] The image generation model adopts an improved CycleGAN model; the discriminator of the improved CycleGAN model includes a multi-scale feature fusion module and a channel attention module; the total loss of the model is calculated by adversarial loss, cycle consistency loss and neighbor contrast loss.

[0011] Combined with the first aspect, further, the image preprocessing includes size unification, pixel value normalization and data augmentation.

[0012] Combined with the first aspect, further, the multi-scale feature fusion module is used for:

[0013] Extract the shallow features, middle-level features and deep features of the generated image through multiple convolutional layers respectively;

[0014] Downsample the shallow features and middle-level features to the same spatial size as the deep features through adaptive average pooling to complete the spatial alignment of the shallow features and the deep features;

[0015] Concatenate the spatially aligned shallow features, middle-level features and deep features along the channel dimension through a feature fusion layer to obtain the fused features;

[0016] Reduce the dimension of the fused features to the dimension of the generated image to obtain the dimension-reduced fused features.

[0017] Combined with the first aspect, further, the channel attention module is used for:

[0018] Perform global average pooling on the feature map of each channel output by the multi-scale feature fusion module to obtain the description value of the feature map of each channel, and form a channel description vector; the description value of the

[0019] ;

[0020] where, is the spatial height, is the spatial width, represents the input feature map in the th channel at the spatial position is the activation value at the index in the height direction,

[0021] Input the channel description vector into the first fully-connected layer, and compress and activate the channel description vector through the ReLU activation function to obtain an intermediate variable;

[0022] Restore the intermediate variable to the number of channels C of the output feature map of the multi-scale feature fusion module through the second fully-connected layer, and then constrain the output to the interval [0,1] through the Sigmoid function to obtain the channel weight vector , and the expression is as follows:

[0023] ;

[0024] where, , are learnable parameters, is the ReLU activation function, is the Sigmoid function;

[0025] Multiply the channel weight vector and the output feature map of the multi-scale feature fusion module channel by channel to obtain the weighted feature map;

[0026] Integrate the weighted feature map into the discriminator to perform true and false image discrimination.

[0027] Combined with the first aspect, further, perform global average pooling on the fused feature after dimensionality reduction, and then perform normalization processing through a fully-connected layer to obtain the multi-scale fused feature;

[0028] According to the multi-scale fused feature, obtain the improved neighbor contrast loss function, and the expression is as follows:

[0029] ;

[0030] where, is the neighbor contrast loss function, is the multi-scale fused feature, represents the th neighbor positive sample feature, is the index of the neighbor positive sample feature, represents the th negative sample feature, is the index of the negative sample feature, is the temperature parameter, is the number of neighbors, is the number of negative samples.

[0031] Combined with the first aspect, further, input the multi-scale fused feature into a fully-connected network including an input layer, a hidden layer and an output layer, and this fully-connected network outputs the temperature scaling factor ;

[0032] Using a temperature scaling factor Calculating the temperature parameter , and the calculation formula is as follows:

[0033] .

[0034] In combination with the first aspect, further, the improved CycleGAN model calculates the total model loss through adversarial loss, cycle consistency loss, and neighbor contrast loss, and the expression is as follows:

[0035] ;

[0036] where represents the total model loss, represents the adversarial loss function, represents the cycle consistency loss function, represents the neighbor contrast loss function, represents the generator from the source domain to the target domain , represents the discriminator of the target domain , represents the generator from the target domain to the source domain , represents the weight coefficient of the cycle consistency loss, represents the weight coefficient of the neighbor contrast loss.

[0037] In combination with the first aspect, further, according to the requirements of the image generation task, a set of source images and target images are obtained and image preprocessing is performed;

[0038] The image generation model is trained using the preprocessed source images and target images to obtain a trained image generation model.

[0039] In combination with the first aspect, further, training the image generation model using the preprocessed source images and target images to obtain a trained image generation model includes:

[0040] When the number of iteration steps is less than the first number of steps, the weight of the neighbor contrast loss is 0. As the number of iteration steps increases, the learning rate linearly increases from 0 to the initial value. According to the preprocessed source images and target images, the image generation model is trained using the adversarial loss and the cycle consistency loss;

[0041] When the number of iteration steps is greater than the first number of steps and less than the second number of steps, as the number of iteration steps increases, the weight of the neighbor contrast loss linearly increases from 0 to the preset weight threshold. According to the preprocessed source images and target images, the image generation model is trained using the adversarial loss, the cycle consistency loss, and the neighbor contrast loss;

[0042] When the number of iteration steps is greater than the second number of steps, the weight of the neighbor contrast loss is fixed to a preset weight threshold, the learning rate decays to n times of the initial value, the discriminator parameters are frozen, and the generator is trained using the adversarial loss, cycle consistency loss, and neighbor contrast loss based on the preprocessed source image and target image; where 0 < n < 1;

[0043] When the total model loss is less than the preset loss threshold, the iteration ends and the trained image generation model is output;

[0044] Where the first number of steps and the second number of steps are preset values.

[0045] In combination with the first aspect, further, when the number of iteration steps is greater than the first number of steps and less than the second number of steps, negative samples are selected in batches through the dynamic target domain feature library, and hard negative sample screening is initiated;

[0046] The data structure of the dynamic target domain feature library adopts a circular queue. After each iteration ends, the features of the current target domain real image are updated to the target domain feature library with a momentum coefficient m; when inserting new features, if the queue of the target domain feature library is not full, the new features are inserted sequentially; if the queue is full, the oldest feature in the queue is replaced;

[0047] Hard negative samples refer to the top 10% of negative samples with the highest similarity to the generated images.

[0048] In a second aspect, the present invention provides a high-quality image generation device, including:

[0049] An image preprocessing module for obtaining an image to be processed;

[0050] An image generation module for obtaining a generated image with the style of the target image according to the image to be processed by using a pre-trained image generation model;

[0051] The image generation model adopts an improved CycleGAN model; the discriminator of the improved CycleGAN model includes a multi-scale feature fusion module and a channel attention module; the improved CycleGAN model calculates the total model loss through the adversarial loss, cycle consistency loss, and neighbor contrast loss.

[0052] Compared with the prior art, the beneficial effects achieved by the present invention:

[0053] The present invention proposes a method and device for generating high-quality images. An improved CycleGAN model is used as the image generation model to learn the features of the target image and generate a generated image with the style of the target image. In the improved CycleGAN model, the present invention improves the quality of the generated image by introducing a neighbor contrast loss function to constrain the similarity of the generated image, enhancing the detail fidelity and texture consistency of the model. At the same time, the present invention also systematically improves the discriminator structure of CycleGAN, designs a multi-scale feature fusion module and a channel attention module, significantly improving the discriminator's ability to capture semantic features at different levels through multi-scale feature fusion, ensuring global style consistency of the generated image, and weighting the fused features through the channel attention mechanism to make the local detail texture expression more realistic. At the same time, the present invention also proposes a dynamic temperature parameter prediction strategy and a difficult negative sample screening mechanism in the neighbor contrast loss function, strengthening the effectiveness and pertinence of the contrast loss training.

[0054] The images generated by the present invention can present diverse style features, maintain good authenticity in details, ensure the stability of image quality, and achieve excellent results in image conversion application scenarios such as season conversion and face expression transfer. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 The figure shows a schematic diagram of the steps of a method for generating high-quality images provided by an embodiment of the present invention;

[0056] Figure 2 The figure shows a schematic diagram of the structure of the generator of an image generation model in an embodiment of the present invention;

[0057] Figure 3 The figure shows a schematic diagram of the structure of the discriminator of an image generation model in an embodiment of the present invention;

[0058] Figure 4 The figure shows a schematic diagram of the neighbor contrast loss function in an embodiment of the present invention;

[0059] Figure 5 The figure shows a schematic diagram of the structure of a device for generating high-quality images provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] The technical solution of the present invention will be described in detail below through the accompanying drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present invention are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.

[0061] Embodiment 1

[0062] This embodiment introduces a high-quality image generation method that generates images based on CycleGAN combined with neighbor contrast loss. Multilevel feature extraction is introduced, and the discriminator is fused from a single scale to multiple scales, as Figure 1 shown. The method specifically includes the following steps:

[0063] Step A: Obtain a set of source images and target images, and perform image preprocessing. Among them, the source image is the original image that needs to be style-transformed or image-restored, and the target image is the corresponding target style or the complete image after restoration.

[0064] Step A01: According to the image generation task to be executed, obtain a set of source images and target images related to the task, and unify the sizes of all input images to the same size for subsequent neural network processing.

[0065] Step A02: Use bilinear interpolation to scale the source images and target images to 256×256 pixels, maintaining the aspect ratio (scaling the short side and centering and cropping the long side).

[0066] Map the pixel values of the source images and target images from [0, 255] linearly to the range of [0, 1] or [-1, 1] through a normalization algorithm to ensure stability during the neural network training process.

[0067] Step A03: Adopt data augmentation techniques such as rotation, cropping, and flipping to perform data augmentation on the source images and target images to improve the generalization ability of the model.

[0068] By unifying the formats and qualities of the input images, reducing differences in aspects such as image size and brightness, the subsequent training of the model becomes more stable and efficient.

[0069] Step B: Obtain an improved CycleGAN model as the image generation model.

[0070] The improved CycleGAN model includes a generator and a discriminator . The generator is responsible for converting the style of the source image into that of the target image to obtain a generated image. Specifically, the task of the generator is to learn the mapping function from the source image to the target image. The task of the discriminator is to distinguish between the image generated by the generator and the real target image and output a probability value representing the authenticity of the image.

[0071] The generator generally includes an encoder, a bottleneck layer, and a decoder. The encoder extracts low-level and high-level features of the image through multiple convolutional layers, batch normalization, and activation functions (such as ReLU). The bottleneck layer is between the encoder and the decoder. The generator compresses features and transfers information through the bottleneck layer, which contains the representation of the latent space for generating the target style image. The decoder restores the latent representation of the bottleneck layer to the target style image through transposed convolutional (deconvolution) layers. During the decoding process, the goal of the generator is to minimize the difference from the target image while preserving the main structure and content of the source image.

[0072] In the embodiment of the present invention, the architecture of the generator is a U-Net variant based on ResNet-9, and its network structure is as Figure 2 shown. It mainly includes 3 downsampling layers, 9 residual blocks, and 3 upsampling layers. The 3 downsampling layers correspond to convolutional layer 1, convolutional layer 2, and convolutional layer 3 in the figure. The convolutional kernels of the 3 downsampling layers are 4×4, the stride is 2, and the padding is 1. The output channels of the 3 downsampling layers are 64, 128, and 256 in sequence. The 3 upsampling layers correspond to upsampling layer 1, upsampling layer 2, and upsampling layer 3 in the figure. The convolutional kernels of the 3 upsampling layers are 4×4, the stride is 2, and the padding is 1. The output channels of the 3 upsampling layers are 128, 64, and 3 (RGB) in sequence. The input of the generator is a 256×256 source image, and the output is a 256×256 generated image. Its initial weights are initialized with He normal distribution, and the biases are initialized to 0.

[0073] The discriminator is composed of multiple convolutional layers and fully connected layers. The present invention optimizes the structure of the discriminator, mainly in two parts. The first part is that the present invention introduces a multi-scale feature fusion module into the discriminator. The second part is that the present invention introduces a channel attention module (Squeeze-and-Excitation Block, SE Block) into the discriminator. The improved discriminator discriminates real and fake images through the multi-scale feature fusion module and the channel attention module.

[0074] The network structure of the improved discriminator is as Figure 3 shown. Convolutional layer 1, convolutional layer 2, convolutional layer 3, convolutional layer 4, adaptive average pooling, feature concatenation, and 1×1 convolution in the figure all belong to the structure of the multi-scale feature fusion module. Convolutional layer 5 is used to fuse the output features of the multi-scale feature fusion module and the channel attention module. In Figure 3Among them, the convolutional kernels of convolutional layer 1, convolutional layer 2, convolutional layer 3, and convolutional layer 4 are 4×4, the stride is 2, the padding is 1, and the output channels are 64, 128, 256, and 512 in sequence. The convolutional kernel of convolutional layer 5 is 4×4, the stride is 2, the padding is 1, and the output channels are 64, 128, 256, 512, 1 in sequence. The output channel of the 1×1 convolution is 128. The input of the discriminator is a 256×256 RGB image output by the generator, and the output of the discriminator is a 30×30 probability map.

[0075] 1. Traditional PatchGAN discriminators usually only use the output of the last convolutional layer as features. Such single-scale features are difficult to balance local details (such as leaf textures, snow grains) and global structures (such as mountain contours, lake shapes). For example, deep features may lose fine-grained information, while shallow features lack semantic abstraction ability. To solve the above problems, the present invention designs a multi-scale feature fusion module in the discriminator to capture multi-granularity information from different levels of convolutions and enhance the representation ability of features through fusion.

[0076] The structure of the multi-scale feature fusion module includes multiple convolutional layers and a feature fusion layer. Among them, the first convolutional layer (Conv1) outputs a size of 128×128, capturing low-level features and containing rich low-level information (such as edges, color gradients). The second convolutional layer (Conv2) outputs a size of 64×64, capturing intermediate-level features (such as object shapes, local textures). The third convolutional layer (Conv3) outputs a size of 32×32, still capturing intermediate-level features but deeper intermediate-level information. The fourth convolutional layer (Conv4) outputs a size of 16×16, containing high-level semantics (such as scene categories, overall layouts).

[0077] The operations of the multi-scale feature fusion module are as follows:

[0078] (1) Extract feature maps of different scales through multiple convolutional layers. For example, the first layer outputs 64-channel features with a resolution of 128×128 (capturing edges, colors), the second layer outputs 128-channel features of 64×64 (capturing shapes), and the fourth layer outputs 512-channel features of 16×16 (capturing semantics). In theory, both the third layer and the second layer features can be selected as intermediate-level features, but considering that the third layer features may lose fine-grained information, the present invention preferentially selects the second layer for feature fusion to capture multi-granularity information at multiple levels as much as possible.

[0079] (2) Downsample the shallow features (Conv1) and intermediate-level features (Conv2) to the same spatial size (16×16) as the deep features (Conv4) through adaptive average pooling to complete spatial alignment. For example, pool the 128×128 Conv1 features into 16×16, retaining key region information.

[0080] (3) After spatially aligning, concatenate the features of Conv1, Conv2, and Conv4 along the channel dimension to form a fused feature tensor. The number of channels of the fused feature tensor is: 64 + 128 + 512 = 704.

[0081] (4) Use a 1×1 convolution to compress the fused feature tensor from 704 dimensions to 128 dimensions, reducing redundancy and retaining core information, thus completing the dimensionality reduction projection. The size of the feature map at this time is 16×16×128, which not only retains multi-scale information but also reduces the computational complexity.

[0082] 2. In the multi-scale fused features, the contribution degrees of different channels to the season conversion task are different. For example, the channels describing the snow depth may be more important than the channels representing the vegetation color. However, traditional models treat all channels equally, resulting in the key information being submerged. To solve the above problems, the present invention introduces a channel attention module in the discriminator, which adaptively learns to assign weights to each channel, highlighting important features and suppressing redundant information.

[0083] The operations of the channel attention module are as follows:

[0084] (1) Perform global average pooling on the feature maps of each channel of the multi-scale feature fusion module to compress the spatial information into a scalar, obtaining a channel description vector.

[0085] Assume the input feature map has a dimension of , where is the batch size, is the number of channels, is the spatial size. Perform average pooling on all spatial positions of each channel to compress the two-dimensional feature map of each channel into a scalar value. The expression is as follows:

[0086] ;

[0087] Where represents the feature map description value of the th channel, is the spatial height, is the spatial width, represents the activation value of the th channel in the input feature map at the spatial position , is the index in the height direction, is the index in the width direction.

[0088] The feature map description values of all channels together form a channel description vector , the channel description vector reflects the contribution of different channels in the overall task. For example, in the task of generating winter scenery, the channel describing the snow depth may have a relatively high value, while the channel describing the vegetation color may have a relatively low value.

[0089] (2) Learn the relationship between channels through a two-layer fully connected network. The first layer compresses the number of channels (Reduction = 16), and the second layer restores the original number of channels and applies the Sigmoid function to generate weights between 0 and 1.

[0090] Input the channel description vector into the first fully connected layer (the number of neurons is C / 16), compress and activate the channel description vector through the ReLU activation function to obtain an intermediate variable, reducing the dimension and capturing the non-linear relationship between channels.

[0091] Restore the intermediate variable to the original number of channels C through the second fully connected layer, and then constrain the output to the interval [0,1] through the Sigmoid function to obtain the channel weight vector , and the expression of the channel weight vector is as follows:

[0092] ;

[0093] where , are learnable parameters, is the ReLU activation function, and is the Sigmoid function.

[0094] The channel weight vector represents the importance score of each channel. For example, when generating winter scenery, the channel describing the snow texture may be assigned a weight close to 1, while the channel weight of the channel describing summer vegetation approaches 0.

[0095] (3) Multiply the learned channel weight vector by the original features channel by channel to achieve dynamic weighting and obtain the weighted feature map.

[0096] Expand the channel weight vector to the dimension of to align it with the input feature map .

[0097] Multiply the expanded channel weight vector by the input feature map channel by channel to obtain the weighted feature map :

[0098] ;

[0099] Among them, represents per-channel multiplication.

[0100] By adding weights to each channel of the feature map, important channels can be enhanced while redundant channels can be suppressed.

[0101] Integrate the weighted feature map into the discriminator, and insert an SE Block after the deep features (Conv4) to enhance the focus on key channels.

[0102] In the CycleGAN framework, the design of the loss function directly determines the quality and diversity of the generated images. Traditional methods mainly rely on the Adversarial Loss and the Cycle-Consistency Loss, but it is difficult for these two to effectively constrain the semantic alignment of the generated images in the feature space. Especially in the season conversion task, problems such as detail distortion or single pattern are likely to occur. Therefore, this invention introduces the Neighbor Contrastive Loss. By measuring the similarity between the generated image and the target image through the neighbor contrast loss, the common function of the neighbor contrast loss is in the form of InfoNCE (Information Noise-Contrastive Estimation), and its calculation formula is as follows:

[0103] ;

[0104] Among them, is the neighbor contrast loss function, is the feature of the generated image, is the feature of the positive sample in the target domain, is the feature of the negative sample in the target domain, is the temperature parameter, represents the sample feature within the target domain.

[0105] In the image generation task based on the neighbor contrast loss, the reasonable definition of positive and negative samples is the core to ensure the effectiveness of contrastive learning. Its core idea is to force the features of the generated image to be as close as possible to the real samples in the target domain and far from irrelevant samples through the similarity constraint in the feature space.

[0106] The positive sample is the k-Nearest Neighbors (k-NN) real sample of the generated image in the target domain feature library. Specifically, for the generated image , extract its 128-dimensional feature vector , and then search for the real sample features with the highest cosine similarity in the target domain feature library as the positive sample.

[0107] The value selection can be adjusted according to the task complexity, usually taking . A smaller value can enhance local alignment but may introduce noise; a larger value can improve robustness but may dilute key features.

[0108] The negative samples are randomly sampled non-neighbor real samples from the target domain feature library, used to form a contrast with the positive samples. During sampling, we uniformly sample from the feature library and also need to exclude the current positive sample. For example, for each generated sample, 1000 features are randomly selected as the negative sample set .

[0109] The target domain feature library is the core component of neighbor contrast learning, and its design directly affects the alignment effect between the generated image and the target image. The feature library needs to be updated in real time to reflect the changes in the target domain distribution while maintaining the stability of historical features.

[0110] Traditional neighbor contrast loss only uses the features of the last layer of the discriminator, ignoring the complementarity of multi-scale information, and uses a unified temperature value, making it difficult to adapt to the difficulty differences of different samples, resulting in simple samples dominating the training and insufficient learning of difficult samples. Therefore, the present invention adjusts the traditional neighbor contrast loss function, calculates the neighbor contrast loss using the multi-scale fusion features output by the improved discriminator, and introduces a temperature scaling factor to achieve adaptive adjustment of the temperature parameter.

[0111]

[0112] To facilitate the calculation of neighbor contrast loss, the features after multi-scale fusion and channel attention weighting need to be projected into a compact vector. Therefore, the present invention adds global average pooling and a fully connected layer after the feature fusion layer of the discriminator. The global average pooling here is different from that in the SE Block. Its purpose is to compress spatial information into a global descriptor, generating a fixed-length feature vector. Its input is the feature map after the fusion layer, with a size of 16×16 and 128 channels. A fully connected layer is added after global average pooling to keep the 128-dimensional vector at 128 dimensions through linear transformation and perform layer normalization LayerNorm to replace L2 normalization and retain the relative ratio between channels. Finally, the projected feature vector is used to calculate the similarity between the generated image and the target domain features.​The discriminator extracts feature maps of different scales in multiple convolutional layers (such as Conv1, Conv2, Conv4), then downsamples the shallow features (Conv1) and middle features (Conv2) to the same spatial size as the deep features (Conv4) through adaptive average pooling, concatenates the features of the three levels along the channel dimension to form fused features, performs dimensionality reduction projection on the fused features, and finally performs global average pooling on the compressed features and normalizes them through a fully connected layer. The finally generated feature vector is used to calculate the neighbor contrast loss.

[0113] Let the multi-scale fused feature be , and the improved neighbor contrast loss function is:

[0114] ;

[0115] where, represents the th neighbor positive sample feature, is the index of the neighbor positive sample feature, represents the th negative sample feature, is the index of the negative sample feature, is the number of near neighbors, is the number of negative samples.

[0116] Traditional neighbor contrast loss functions generally randomly select negative samples, but the randomly selected negative samples may contain a large number of noise samples that are irrelevant to the generated features, resulting in low contrast learning efficiency. Therefore, the present invention introduces similarity in the process of selecting negative samples. As Figure 4 shows, the present invention calculates the cosine similarity matrix of the generated image features , positive sample features and all negative sample features in the feature library, and selects the top 10% negative samples with the highest similarity as hard negative samples for each generated sample. For example, if the capacity of the feature library N = 10000, then 1000 hard negative samples are selected for each generated sample. The weight of the hard negative samples in the contrast loss is increased by 2 times to strengthen the model's ability to distinguish easily confused samples. In contrast learning, the diversity of negative samples directly affects the model's ability to distinguish feature differences.

[0117] The cosine similarity formula is:

[0118] .

[0119] In addition, traditional methods only rely on negative samples of the current batch, which may lead to limited sample coverage and fail to fully capture the multimodal distribution of the target domain. Therefore, the cross-batch negative sample pool significantly improves the diversity and representativeness of negative samples by dynamically maintaining a historical sample queue and combining new and old feature mixed sampling. After optimization, 1000 images are randomly selected from the target domain feature library at the initial stage of model training, features are extracted and filled into the queue to avoid the cold start problem. Its data structure adopts a circular queue with the first-in-first-out (FIFO) principle, and the capacity is fixed at 65536, supporting efficient insertion and elimination. After each training batch, the features of the current target domain real images are updated to the library with a momentum coefficient m = 0.999. When inserting new features, if the queue is not full, the new features are inserted sequentially; if the queue is full, the oldest 1% of the features (about 655) are replaced. In addition, features with too low similarity (such as sim < 0.2) or expired timestamps need to be removed regularly (every 10k steps).

[0120] In the neighbor contrast loss function, the temperature parameter is used to control the smoothness of the similarity distribution. When takes a fixed value (such as 0.07), two problems may occur: one is is too small, then the similarity difference is amplified, and the model overly focuses on difficult samples, resulting in unstable training; the other is is too large, which will flatten the similarity distribution. At this time, easy samples dominate the learning and the convergence becomes slow. It can be seen that it is difficult for a fixed temperature parameter to adapt to the difficulty differences of different samples. Therefore, the present invention realizes the adaptive adjustment of the temperature parameter by dynamically predicting the temperature scaling factor, and substitutes the adjusted temperature parameter into the calculation of the neighbor contrast loss, as Figure 4 shown.

[0121] First, the present invention adds a small fully connected network. The input is the mean value (128 dimensions) of the multi-scale fusion features, and the output is the temperature scaling factor. Its network structure is as follows:

[0122] Input layer: 128 nodes (corresponding to the feature dimension).

[0123] Hidden layer: 64 nodes, using ReLU activation.

[0124] Output layer: 1 node, using the Sigmoid function to constrain the output range to (0, 1).

[0125] In the dynamic calculation of the temperature, the base temperature is set to 0.07, and the temperature scaling factor is generated by the fully connected network. The calculation formula for the final temperature parameter is:

[0126] ;

[0127] Among them, is the temperature scaling factor.

[0128] Through the above formula, it is ensured that , adapting to the feature distributions of different batches and avoiding training instability caused by extreme values. For simple samples with high similarity, the temperature will automatically increase to smooth the differences and prevent overfitting; while for difficult samples with low similarity, the temperature will automatically decrease to strengthen the differential learning.

[0129] The total loss function of the improved CycleGAN model of the present invention includes adversarial loss, cycle consistency loss, and neighbor contrast loss. The adversarial loss is used to drive the generator to generate images with high realism. The cycle consistency loss is used to ensure that when the source image is transformed through the target image and then regenerated back to the source image, the image content remains consistent. By minimizing the difference between the source image, the target image, and then back to the source image, the model ensures the consistency of the image content. The neighbor contrast loss is used to optimize the local texture and structure in the generated image, making the relationship between adjacent pixels more consistent, thereby improving the detail quality of the generated image.

[0130] The expression of the total loss function is as follows:

[0131] ;

[0132] Among them, represents the total loss function; represents the generator from the source domain to the target domain ; represents the discriminator in the target domain , which is used to distinguish real images and generated images, forcing the generator to generate images close to the distribution of the target domain ; represents the generator from the target domain to the source domain , which is used to implement the inverse mapping and ensure the consistency of the mappings of and through the cycle consistency loss; represents the weight coefficient of the cycle consistency loss, represents the weight coefficient of the neighbor contrast loss, and The values of

[0133] The formula of the adversarial loss function is:

[0134] ;

[0135] Among them, represents the real image (source domain image) data distribution, represents the image under the real image data distribution is the discriminator's discrimination result for the input image ; is the noise distribution, represents the noise vector under the noise distribution is the generated image output by the generator under the noise vector, is the discriminator's discrimination result for the image ;

[0136] The cycle consistency loss needs to calculate the forward cycle consistency and the backward cycle consistency separately, and its formula is:

[0137] ;

[0138] where represents the generator outputs the generated image according to the input image , specifically, the generator converts the input source domain image to the target domain and then outputs the image; represents the generator outputs the image according to the image , specifically, the generator converts the image from the target domain back to the source domain and then obtains the image; represents the L1 norm of the image and the image , specifically, it represents taking the absolute value of the difference between each element of the image and the image and then summing them up; represents the target domain image data distribution, represents the target domain image under the target domain image data distribution represents the generator outputs the generated image according to the target domain image , specifically, the generator converts the target domain image to the source domain and then outputs the image; represents the generator outputs the image according to the image The output image, specifically referring to the generator The image From the source domain Converted back to the target domain The resulting image after that; Represents the image And the target domain image The L1 norm of, specifically representing the image And the target domain image Taking the absolute value of the element-by-element difference and then summing them up.

[0139] The present invention adjusts the loss weight in stages. In the warm-up stage, the number of iteration steps is less than the first number of steps, and the contrast loss weight , only optimizing the adversarial loss and the cycle consistency loss to avoid the interference of the early contrast loss on the global mapping learning of the generator. In the embodiment of the present invention, the first number of steps is preferably 10k steps.

[0140] Then, in the contrast learning stage, the number of iteration steps is greater than the first number of steps and less than the second number of steps, Linearly increasing from 0 to a preset weight threshold. In the embodiment of the present invention, the second number of steps is preferably 50k steps, and the preset weight threshold is preferably 0.8.

[0141] The contrast learning stage The formula for is:

[0142] .

[0143] Finally, in the fine-tuning stage, the number of iteration steps is greater than the second number of steps, fixing , reducing the learning rate to n times the initial value, and finely adjusting the generation details, Indicating the current number of steps, where 0 < n < 1, preferably 0.1.

[0144] Step C: Use the preprocessed source image and target image to train the image generation model, enabling the image generation model to have the ability to map the source image to the target image, and obtaining the trained image generation model.

[0145] The traditional CycleGAN training process mainly relies on the alternating optimization of the adversarial loss and the cycle consistency loss. However, after introducing the neighbor contrast loss, it is necessary to redesign the training strategy to balance multi-objective learning, dynamic feature library management, and computational efficiency. Therefore, the present invention adopts a staged training strategy, which is specifically divided into the following three stages:

[0146] 1. Warm-up stage (0 - 10k steps): Only use the adversarial loss and the cycle consistency loss, freeze the dynamic feature library, and linearly increase the learning rate from 0 to the initial value (generator 0.0001, discriminator 0.0002) to initially establish the global seasonal mapping ability.

[0147] 2. Contrastive learning stage (10k - 50k steps): Gradually introduce the neighbor contrastive loss, whose weight linearly increases from 0 to 0.8, activate the update of the feature bank (inject momentum features per batch), initiate the screening of difficult negative samples (Top 10% similarity), and use FAISS (Facebook AI Similarity Search, an open-source similarity search tool by Facebook AI) to accelerate the nearest neighbor retrieval, and refine and optimize the feature alignment.

[0148] 3. Fine-tuning stage (after 50k steps): Fix the contrastive loss weight at 0.8, decay the learning rate to 10% of the initial value, freeze the discriminator parameters, and only optimize the generator. Refine the image details through a low learning rate to avoid overfitting.

[0149] The specific training steps are as follows:

[0150] Step 1. Network initialization. After preparing the training data, initialize the network structure of the image generation model, including initializing the parameters of the generator and discriminator, such as weights and biases.

[0151] There are many ways to initialize, such as random initialization or initializing with the parameters of a pre-trained model.

[0152] Step 2. Use the prepared training data (pre-processed source images and target images) to train the image generation model. During the training process, the generator attempts to generate realistic images, while the discriminator attempts to distinguish between real images and generated images. Training is usually an iterative process, and the performance of the generator and discriminator is continuously optimized through multiple iterations.

[0153] Step 3. After each iterative training, calculate the total model loss based on the features output by the generator and discriminator to measure the performance of the model.

[0154] Step 4. According to the total model loss, use the optimization algorithm Adam to optimize the model parameters. By adjusting the weights and biases, minimize the loss value. If the generator loss is large, the optimization algorithm will adjust the parameters of the generator to make the generated images closer to real images.

[0155] Step 5: Use an independent test dataset to evaluate the performance of the trained model. Judge the effectiveness of the model by comparing metrics such as the similarity between the generated images and the real images and evaluating the quality of the generated images. In this embodiment, the FID (Frechet Inception Distance) is used to measure the distribution distance between the generated images and the real images; the KID (Kernel Inception Distance) is used to seek a more robust distribution similarity metric; meanwhile, a user study is adopted to manually evaluate the visual quality and domain consistency of the generated images.

[0156] Step 6: If the model performs well in the test, save the trained model for subsequent applications. The saved model can be used for actual image generation tasks.

[0157] After the above improvements and training, the improved CycleGAN model after training is obtained, which is used as the image generation model of the present invention to generate generated images in the style of the target image.

[0158] Step D: Input the image to be processed into the trained image generation model to obtain a generated image with the style of the target image.

[0159] Experiment 1:

[0160] This experiment takes the task of seasonal conversion of landscape images as an example to verify the effect of the method of the present invention:

[0161] First, the experimental parameters for the task of seasonal conversion of landscape images are set as follows:

[0162] Source domain (summer landscape): Collect 10,000 summer landscape images containing scenes such as mountains, forests, lakes, etc. The image sources include public datasets (such as Flickr, Unsplash) and self-collected outdoor photography works. Target domain (winter landscape): There are also 10,000 winter landscape images, covering typical winter elements such as snow scenes, frozen lakes, and rime, ensuring a match with the source domain scenes. It is required that the image resolution is not lower than 512×512 pixels to avoid low resolution affecting the generation of details. The image format is unified as JPEG, and the color mode is RGB. The images should avoid containing dynamic interference objects such as people and vehicles, and ensure that the content is mainly natural landscapes.

[0163] In Experiment 1, the structural parameters of the image generation model are configured as shown in Table 1:

[0164] Table 1 Configuration of the structural parameters of the image generation model

[0165]

[0166] The image generation model proposed by the present invention is trained based on the source domain and target domain for the seasonal conversion task of landscape images, and the process is as follows:

[0167] 1. Image preprocessing: In the downloaded source domain images and target domain images, in addition to excluding black and white photos, bilinear interpolation is also used to scale the images to 256×256 pixels, maintaining the aspect ratio (scaling the short side and cropping the long side in the center). The pixel values are linearly mapped from [0, 255] to [-1, 1] to be consistent with the training settings of the network. The formula is:

[0168] ;

[0169] where, is the original pixel value, is the normalized pixel value.

[0170] 2. Input the preprocessed images into the image generation model, and use the basic adversarial loss and cycle consistency loss of CycleGAN to preliminarily train the generator and discriminator, so that the model can learn the basic conversion mapping from summer images to winter images.

[0171] 3. After the preliminary training is completed, add the neighbor contrast loss to optimize the diversity of the generated images. Through the neighbor contrast loss, the generated winter images have diverse features, avoiding the singularity of the generated images and enhancing the detail diversity, such as different snow coverage effects and hue changes.

[0172] 4. Dynamically adjust the weight of the neighbor contrast loss: Gradually increase the weight of the neighbor contrast loss, so that the generated winter landscape images achieve a better balance between details and diversity, ensuring the realism and artistic sense of the generated images.

[0173] In this experiment, (following the default settings of CycleGAN), (which needs to be fine-tuned through experiments).

[0174] 5. When the generator training is completed and these loss functions are optimized and adjusted, we can use the trained generator to perform the image conversion task. Input the new summer landscape images into the generator to generate the corresponding winter landscape images.

[0175] To demonstrate the performance of the method of the present invention through data comparison, the experiments of the present invention also used traditional CycleGAN models, CUT (Contrastive Learning for Unpaired Image-to-Image Translation) models, and DRIT++ (Diverse image-to-image translation via disentangled representations) models to perform the seasonal conversion task of landscape images according to the same source domain and target domain respectively, generating winter landscape images.

[0176] The present invention introduces FID and KID metrics to measure image generation performance, and also introduces user ratings. Users give ratings based on the details and diversity of the finally generated images.

[0177] The comparison of the results of different model methods performing the seasonal conversion task of landscape images is shown in Table 2:

[0178] Table 2 Comparison of the results of different model methods performing the seasonal conversion task of landscape images

[0179]

[0180] In Table 2, NCL (optimized) represents the image generation model of the present invention, which is an improved CycleGAN model introducing Neighbor Contrastive Loss (NCL).

[0181] It can be seen from the data in Table 2 that the FID and KID metrics of the method of the present invention are significantly lower than those of other models. Therefore, the images generated by the present invention are more similar to the target images, and the quality of the generated images is better. In terms of the details and diversity of the generated images, the method of the present invention is significantly higher than other models. Therefore, the model of the present invention has good global feature alignment ability in unstructured scenarios (natural landscapes), including multi-scale feature alignment (distant view mountains + near-view snow cover), natural light gradual change (such as lake surface reflection), and can achieve a better balance between the diversity and quality of the generated images.

[0182] Experiment 2:

[0183] This experiment takes the facial expression conversion task as an example to verify the effect of the method of the present invention:

[0184] First, the experimental parameters of the facial expression conversion task are set as follows:

[0185] In the source domain data (neutral expressions), 10,000 natural neutral-expression face images of different genders, ages, and ethnicities were collected. The data sources include publicly available face datasets (such as CelebA, FFHQ) and self-collected face images to ensure sample diversity. In the target domain data (smiling expressions), 10,000 smiling-expression face images of the corresponding population were collected. The degree of smiling covers different types (smile, chuckle, laugh) to ensure that the target domain samples are diverse and match the source domain face images. It should be noted that in these collected image data, the people should avoid wearing glasses, masks, occluders, and other factors that affect expression recognition. The images must contain a complete face and be at a frontal or near-frontal angle.

[0186] In Experiment 2, the structural parameter configuration of the image generation model is shown in Table 3:

[0187] Table 3 Structural Parameter Configuration of the Image Generation Model

[0188]

[0189] The image generation model is trained using the source domain face images and the target domain face images, and the trained image generation model is used to generate smiling-expression face images. At the same time, to highlight the performance of the method of the present invention, the StarGAN v2 model and the GANimation model are used to perform facial expression conversion tasks respectively based on the same source domain and target domain to generate smiling-expression face images, and the FID metric, expression accuracy rate, and identity retention rate are selected to compare the results of different models.

[0190] The comparison of the results of different model methods performing facial expression conversion tasks is shown in Table 4:

[0191] Table 4 Comparison of the Results of Different Model Methods Performing Facial Expression Conversion Tasks

[0192]

[0193] From the data in Table 4, it can be seen that the FID metric of the method of the present invention is significantly lower than that of other models. Therefore, the images generated by the present invention are more similar to the target images, and the quality of the generated images is better. The uploaded face images of the method of the present invention have a higher expression accuracy rate and identity retention rate, proving that the method of the present invention is more excellent in the local detail control ability in structured data (faces) (including local micro-expression control (corner of the mouth curvature, eye wrinkles), identity information retention (skin color, moles, hairstyle)).

[0194] In the embodiments of the present invention, Experiment 1 and Experiment 2 respectively verified the superiority of the same technical framework in the tasks of "global style transfer" and "local feature editing" through differential scenario design and targeted technical adjustments. The combination of the two experiments can fully demonstrate the cross-domain generality of the solution and reflect the multi-scenario support performance of the method of the present invention.

[0195] Embodiment 2

[0196] Embodiment 1 and Embodiment 2 are based on the same inventive concept. This embodiment introduces a high-quality image generation device, as Figure 5 shown, including an image preprocessing module, a model training module, and an image generation module.

[0197] The image preprocessing module is used to obtain a set of source images and target images, and perform image preprocessing; it is also used to obtain the image to be processed.

[0198] The model training module is used to train an image generation model using the preprocessed source images and target images to obtain a trained image generation model.

[0199] The image generation module is used to input the image to be processed into the trained image generation model to obtain a generated image with the style of the target image.

[0200] In the embodiments of the present invention, the image generation model adopts an improved CycleGAN model; the discriminator of the improved CycleGAN model uses a multi-scale feature fusion module and a channel attention module to distinguish between real and fake images; the improved CycleGAN model calculates the total loss of the model through adversarial loss, cycle consistency loss, and neighbor contrast loss.

[0201] For the specific function implementation of each of the above modules, refer to the relevant content in the method of Embodiment 1, which will not be elaborated here.

[0202] In summary of the above embodiments, the present invention proposes a high-quality image generation method and device, which improves the discriminator structure of CycleGAN, designs a multi-scale feature fusion module and a channel attention mechanism, significantly improves the discriminator's ability to capture semantic features at different levels, and ensures that the generated image not only has global style consistency, but also has more authenticity in the expression of local detail textures.

[0203] The present invention introduces a neighbor contrast loss function into the traditional CycleGAN framework. By imposing fine-grained similarity constraints on the generated images and target domain images in the feature space, the diversity and local detail fidelity of the generated images are effectively improved. The present invention constructs a dynamically updated feature library, and realizes more accurate feature alignment by performing dynamic neighborhood matching and difficult sample screening between the generated images and the target domain images. At the same time, a dynamic temperature parameter prediction strategy and a difficult negative sample screening mechanism are proposed to strengthen the effectiveness and pertinence of the contrast loss training.

[0204] In addition, the present invention also innovatively proposes a phased training strategy, including a warm-up stage, a contrast learning stage, and a fine-tuning stage, which effectively optimizes the balance relationship among the adversarial loss, the cycle consistency loss, and the neighbor contrast loss. Through these innovative means, the present invention exhibits better visual quality and diversity in image conversion application scenarios such as season conversion and face expression transfer, effectively avoiding the problems of single generated image patterns and insufficient detail performance.

[0205] The images generated by the present invention can present diverse style features, such as colorful landscape pictures and season pictures, rich details and micro expressions, different roof decorations and window styles. At the same time, the generated images show diversity in details, making the converted images more artistic and realistic.

[0206] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0207] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more flows or multiple flows and / or blocks Figure 1 one or more blocks or multiple blocks.

[0208] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0209] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0210] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.

Claims

1. A method for generating high-quality images, characterized in that: include: Get the image to be processed; According to the image to be processed, a pre-trained image generation model is used to obtain a generated image with the target image style; The image generation model adopts an improved CycleGAN model; the discriminator of the improved CycleGAN model includes a multi-scale feature fusion module and a channel attention module; the improved CycleGAN model calculates the total model loss through adversarial loss, cycle consistency loss and neighbor contrast loss; The multi-scale feature fusion module is used to: Through multiple convolutional layers, shallow features, middle features, and deep features of the generated image are extracted respectively; Adaptive average pooling is used to downsample shallow and mid-layer features to the same spatial size as deep features, completing the spatial alignment of shallow and deep features. The feature fusion layer is used to splice the shallow features, middle features, and deep features after spatial alignment along the channel dimension to obtain the fused features; Reduce the fused features to the dimension of the generated image to obtain the reduced-dimensional fused features; Perform global average pooling on the fusion features after dimension reduction, and then normalize them through the fully connected layer to obtain multi-scale fusion features; According to the multi-scale fusion features, the improved neighbor contrast loss function is obtained, which is expressed as follows: ; in, is the neighbor comparison loss function, is a multi-scale fusion feature. Indicates the Neighbor positive sample features, is the index of the neighbor positive sample feature, Indicates the negative sample features, is the index of the negative sample feature, is the temperature parameter, is the number of neighbors, is the number of negative samples.

2. The high-quality image generation method according to claim 1, characterized in that The channel attention module is used to: Perform global average pooling on the feature map of each channel output by the multi-scale feature fusion module to obtain the feature map description value of each channel and form a channel description vector; The feature map description value of each channel The expression is as follows: ; in, is the space height, is the space width, Represents the input feature map Middle Channels in space The activation value at is the index in the height direction, is the index in the width direction; Channel description vector Input the first fully connected layer and activate the channel description vector through the ReLU activation function Perform compression and activation to obtain intermediate variables; The intermediate variable is restored to the channel number C of the output feature map of the multi-scale feature fusion module through the second fully connected layer, and then the output is constrained to the [0,1] interval through the Sigmoid function to obtain the channel weight vector , the expression is as follows: ; in, 、 is a learnable parameter, is the ReLU activation function, is the Sigmoid function; Multiply the channel weight vector and the output feature map of the multi-scale feature fusion module channel by channel to obtain the weighted feature map; The weighted feature map is integrated into the discriminator to distinguish true from false images.

3. The high-quality image generation method according to claim 1, wherein: The multi-scale fusion features are input into a fully connected network consisting of an input layer, a hidden layer, and an output layer. The fully connected network outputs the temperature scaling factor ; Using the temperature scaling factor Calculating temperature parameters , the calculation formula is as follows: 。 4. The high-quality image generation method according to claim 1, wherein: The improved CycleGAN model calculates the total model loss through adversarial loss, cycle consistency loss and neighbor contrast loss, which is expressed as follows: ; in, represents the total loss of the model, represents the adversarial loss function, represents the cycle consistency loss function, represents the neighbor contrast loss function, Indicates that the source domain To the target domain The generator of Indicates the target domain The discriminator, Indicates that from the target domain To the source domain The generator of represents the weight coefficient of cycle consistency loss, Represents the weight coefficient of the neighbor contrast loss.

5. The high-quality image generation method according to claim 1, wherein: According to the requirements of the image generation task, a set of source images and target images are obtained and image preprocessing is performed; The preprocessed source image and target image are used to train the image generation model to obtain a trained image generation model.

6. The high-quality image generation method according to claim 5, characterized in that: The image generation model is trained using the preprocessed source image and target image to obtain a trained image generation model, including: When the number of iterations is less than the first step, the weight of the neighbor contrast loss is 0. As the number of iterations increases, the learning rate increases linearly from 0 to the initial value. Based on the preprocessed source and target images, the image generation model is trained using adversarial loss and cycle consistency loss. When the number of iteration steps is greater than the first step and less than the second step, as the number of iteration steps increases, the weight of the neighbor contrast loss increases linearly from 0 to the preset weight threshold. Based on the preprocessed source image and target image, the image generation model is trained using adversarial loss, cycle consistency loss, and neighbor contrast loss. When the number of iteration steps is greater than the second number of steps, the weight of the neighbor contrast loss is fixed to the preset weight threshold, the learning rate decays to n times the initial value, the discriminator parameters are frozen, and the generator is trained using the adversarial loss, cycle consistency loss, and neighbor contrast loss based on the preprocessed source and target images; where 0 <n<1; When the total loss of the model is less than the preset loss threshold, the iteration ends and the trained image generation model is output; Among them, the first step number and the second step number are preset values.

7. The high-quality image generation method according to claim 6, characterized in that: When the number of iteration steps is greater than the first step and less than the second step, negative samples are selected in batches through the dynamic target domain feature library, and difficult negative sample screening is started; The data structure of the dynamic target domain feature library adopts a circular queue. After each iteration, the features of the current target domain real image are updated to the target domain feature library with the momentum coefficient m. When inserting new features, if the queue of the target domain feature library is not full, the new features are inserted sequentially. If the queue is full, replace the oldest feature in the queue; Hard negative samples refer to the top 10% of negative samples with the highest similarity to the generated image.

8. A high-quality image generation device, characterized in that: include: An image preprocessing module, used to obtain images to be processed; An image generation module is used to obtain a generated image with the target image style based on the image to be processed using a pre-trained image generation model; The image generation model adopts an improved CycleGAN model; the discriminator of the improved CycleGAN model includes a multi-scale feature fusion module and a channel attention module; the improved CycleGAN model calculates the total model loss through adversarial loss, cycle consistency loss and neighbor contrast loss; The multi-scale feature fusion module is used to: Through multiple convolutional layers, shallow features, middle features, and deep features of the generated image are extracted respectively; Adaptive average pooling is used to downsample shallow and mid-layer features to the same spatial size as deep features, completing the spatial alignment of shallow and deep features. The feature fusion layer is used to splice the shallow features, middle features, and deep features after spatial alignment along the channel dimension to obtain the fused features; Reduce the fused features to the dimension of the generated image to obtain the reduced-dimensional fused features; Perform global average pooling on the fusion features after dimension reduction, and then normalize them through the fully connected layer to obtain multi-scale fusion features; According to the multi-scale fusion features, the improved neighbor contrast loss function is obtained, which is expressed as follows: ; in, is the neighbor comparison loss function, is a multi-scale fusion feature. Indicates the Neighbor positive sample features, is the index of the neighbor positive sample feature, Indicates the negative sample features, is the index of the negative sample feature, is the temperature parameter, is the number of neighbors, is the number of negative samples.

Citation Information

Patent Citations

  • Image texture synthesis method and system for multi-scale channel attention network

    CN113689517A

  • High mountain accumulated snow environment simulation method and system based on SCyclGAN

    CN116958468A