Cycle-GAN optimization method for improving target position consistency in image style migration process

By introducing the dynamic detection module for target position difference and L2 norm regularization in Cycle-GAN, the problem of target position offset of generated image is solved, and the structural consistency and realism of generated images are improved, which is suitable for medical image synthesis, art style migration and autonomous driving data enhancement.

CN120259473APending Publication Date: 2025-07-04CHINA JILIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510408113.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Cycle-GAN is prone to overfitting during training, resulting in inaccurate target positions of generated images. The prior art cannot effectively control the consistency of target positions of generated images, affecting the generation quality.

Method used

A dynamic detection module for target position difference is introduced, and the horizontal coordinate difference between the generated image and the original image is calculated through the target detection network, and a dynamic regularization loss is constructed based on the L2 norm of the generator, which is integrated into the model loss function and optimized the generator parameters.

Benefits of technology

It effectively suppresses the overfitting phenomenon of Cycle-GAN, ensures the consistency of the target position of the generated image, improves the quality and availability of the generated image, and is suitable for medical image synthesis, art style migration and data enhancement in autonomous driving scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005341798340000031
    Figure BDA0005341798340000031
  • Figure BDA0005341798340000041
    Figure BDA0005341798340000041
  • Figure BDA0005341798340000061
    Figure BDA0005341798340000061
Patent Text Reader

Abstract

The invention discloses a Cycle-GAN optimization method for improving target position consistency in an image style migration process, and the method comprises the steps: carrying out the forward propagation of a randomly selected source domain input image through a currently optimized Cycle-GAN network generator in a model optimization process, and outputting a generated image; respectively performing target detection on an originally input source domain input image and a generated image through a target detection network, extracting and screening a target with the maximum area, calculating and normalizing a central point abscissa absolute difference between the two, and calculating a target position difference; generating dynamic regularization loss and fusing the dynamic regularization loss into a model loss function by combining the L2 norm of the generator model parameter and the target position difference, and updating a current model total loss function; and continuing to optimize the model parameters by using the current model loss function. The device is simple in structure, easy to implement, wide in application range and high in universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital image processing, and particularly to a Cycle-GAN optimization method for improving the consistency of target positions during the process of image style transfer. Background Art

[0002] Generative adversarial networks (GANs) are widely used in tasks such as image generation and image conversion. Among them, Cycle-GAN (Cycle Generative Adversarial Network), as an unsupervised learning model, is particularly suitable for image conversion of unlabeled data. However, Cycle-GAN may encounter overfitting problems during the training process, resulting in problems such as inaccurate target positions in the generated images. Existing technologies usually reduce the complexity of the model by applying L1 or L2 norm regularization to the generator model to prevent overfitting. However, these methods do not impose specific constraints on the spatial structure of the generated images, which may cause the target positions of the generated images to shift. In addition, there are methods that determine the training stop time by monitoring the loss change of the validation set to avoid overfitting, but this method cannot optimize dynamically according to the quality of the generated images and may lead to incomplete utilization of data samples during the training process. Due to insufficient attention to the consistency of target positions in the generated images in the existing technologies, and target position offset is an important factor affecting the generation quality. Summary of the Invention

[0003] The present invention provides a Cycle-GAN optimization method for improving the consistency of target positions during the process of image style transfer. By using this method, the consistency of target positions in the generated images can be effectively controlled, thereby improving the generation effect.

[0004] The purpose of the present invention is to improve the problem of target position offset in the generated images caused by overfitting during the training of the Cycle-GAN model, and to propose an optimization method based on target position difference and parameter regularization to ensure the consistency of target positions in the generated images and effectively suppress the overfitting phenomenon of the model.

[0005] The method includes the following steps: after the end of each training cycle, randomly sample test set images and input them into the current generator for image generation; respectively perform target detection on the original image and the generated image through a target detection network, select the target with the largest area, calculate the difference in the abscissa of the target center point and normalize it; form a dynamic regularization term according to the L2 norm of the generator model and the target position difference, and incorporate this regularization term into the original model loss function.

[0006] A Cycle-GAN optimization method for improving the consistency of target positions during the process of image style transfer includes the following steps:

[0007] (1)Perform unsupervised pre-training on the Cycle-GAN network;

[0008] (2)Randomly sample n images from the source domain test set and input them as source domain input images into the Cycle-GAN generator in the current training state for forward propagation to output the generated images corresponding to the target domain;

[0009] (3)Perform object detection on the source domain input images and the target domain generated images respectively through the object detection network, and obtain the object position difference between the two objects;

[0010] (4)Combine the L2 norm of the Cycle-GAN generator model parameters in the current training state with the object position difference to generate a dynamic regularization loss and fuse it into the model loss function, and inversely optimize the total model loss function;

[0011] (5)Use the optimized total model loss function to update the model parameters;

[0012] (6)Evaluate the generated images of the current obtained generator using the peak signal-to-noise ratio and structural similarity index. If the requirements are not met, return to step (2); if the requirements are met, obtain the trained Cycle-GAN network.

[0013] The method of the present invention is a Cycle-GAN overfitting optimization method based on object position difference and parameter regularization. It calculates the normalized difference of the maximum object center point abscissa between the original image and the generated image through the object detection network, and constructs a dynamic regularization loss in combination with the generator L2 norm; finally, high-quality generated images with overfitting suppression and consistent object positions are obtained.

[0014] The present invention introduces an object position difference dynamic detection module, which accurately locates the object position through the object detection network and performs real-time calculation and constraint on its offset. Based on the combination of object position offset and L2 norm, a dynamically adjusted regularization term is designed, enabling the generator to adaptively adjust the regularization intensity during training and further suppressing the overfitting phenomenon. This method can better control the structural consistency in the generated images by quantifying the object position difference, improving the quality and usability of the generated images.

[0015] Furthermore, the Cycle-GAN network consists of a generator, a discriminator, and an object position difference dynamic detection module. By fusing the dynamic regularization loss into the model loss function through the object position difference dynamic detection module, it can suppress the position offset caused by the overfitting of Cycle-GAN.

[0016] In step (1), for the unsupervised pre-training of the Cycle-GAN network, existing methods can be adopted. The loss function in the pre-training stage consists of the generative adversarial loss LAdv and the cycle consistency loss L Cyc It consists of. During pre-training, the source domain image set X and the target domain image set Y are input into the Cycle-GAN network for unsupervised pre-training. Through the generative adversarial loss L Adv and the cycle consistency loss L Cyc the parameters of the generator G and the discriminator D are optimized until the model reaches a preliminary convergence state, providing an initialization parameter basis for subsequent dynamic regularization optimization.

[0017] In step (2), after the pre-training is completed, dynamic regularization optimization training is performed on the preliminarily converged model. Random sampling is carried out from the source domain test set. For example, 100 images with a size of 256×256 are collected and input into the generator in the current training state for forward propagation to output the generated images corresponding to the target domain. The generator adopts the current parameter configuration to reflect the latest state of the model.

[0018] In step (3), when performing object detection on the source domain input image and the target domain generated image, the object with the largest area is extracted and selected as the object corresponding to the source domain input image and the target domain generated image.

[0019] Furthermore, the object detection network is the YOLO object detection network. During actual detection, first, all object bounding boxes of the original input source domain input image and the generated image are obtained respectively; for each image, the single object with the largest bounding box area is selected as the corresponding object, and the objects corresponding to the source domain input image and the target domain generated image are obtained respectively.

[0020] Furthermore, in step (3), the object position difference Δx is:

[0021]

[0022] x1 is the abscissa of the center point of the object selected in the source domain input image, x2 is the abscissa of the center point of the object selected in the generated image, n is the number of source domain input images, and W is the image width.

[0023] In step (4), the model loss function includes the generative adversarial loss, the cycle consistency loss, and the optimization loss including the L2 norm of the generator model parameters and the object position difference information.

[0024] The total model loss function L is:

[0025] L = L Adv + L Cyc + λL Reg

[0026] L Adv is the generative adversarial loss of Cycle-GAN; L Cycis the cycle consistency loss of Cycle - GAN; λ is the weight controlling the dynamic regularization term, and L Reg is the dynamic regularization loss.

[0027] During the optimization process, the present invention adopts a dynamic adjustment mechanism: λ is optimized in real - time according to the mean value of Δx: if Δx > 0.1, λ is increased by 0.04 - 0.06 to enhance regularization; if Δx < 0.05, λ is decreased by 0.01 - 0.03 to avoid underfitting; when Δx is between 0.05 and 0.1, λ remains unchanged.

[0028] The optimized loss is obtained by multiplying the weight term by the dynamic regularization loss, where the dynamic regularization loss is obtained by multiplying the L2 norm of the generator model parameters by the target position difference information; specifically, the dynamic regularization loss L Reg is:

[0029]

[0030] Δx is the target position difference, is the L2 norm of the generator model parameters.

[0031] In steps (4) and (5), the total loss function L is optimized through backpropagation, and the parameters of the generator and discriminator are updated to complete the training iteration. The dynamic regularization term adjusts the regularization strength in real - time through the cross - domain target position difference and parameter norm, inhibits overfitting and maintains the consistency of the target position of the generated images.

[0032] In step (6), the generated images of the current obtained generator are evaluated using the peak signal - to - noise ratio and structural similarity index. If the requirements are not met, return to step (2); if the requirements are met, obtain the trained Cycle - GAN network.

[0033] In the present invention, further, the generator G is composed of an encoder and a decoder, and there is no limit on the specific number of layers.

[0034] Further, the discriminator D is composed of an encoder, and there is no limit on the specific number of layers.

[0035] Further, the generative adversarial loss L Adv has the following expression:

[0036] L Adv =-logD(Y)-log(1 - D(X′))

[0037] where D(Y) and D(X) are the probability predictions of the discriminator D for the input target - domain image Y and the image X generated by the generator respectively ′ respectively.

[0038] Further, the cycle consistency loss LCyc The expression is as follows:

[0039] L Cyc = X - F(X ′ )

[0040] where F is an auxiliary generator with the same structure as the generator G.

[0041] The optimization method described in the present invention is applicable to image conversion tasks such as medical image synthesis, artistic style transfer, and data augmentation for autonomous driving scenarios.

[0042] Compared with the prior art, the present invention has the following beneficial technical effects:

[0043] By introducing a target position difference dynamic detection module and a parameter regularization strategy, the present invention improves the generation quality and stability of Cycle-GAN in image conversion tasks. This method uses an object detection network to quantify the maximum horizontal offset of the target center point between the generated image and the original image in real time, and constructs a dynamic regularization term by combining the L2 norm of the generator parameters, which suppresses the problem of target position offset caused by overfitting in traditional Cycle-GAN, and at the same time avoids the underfitting risk of the static regularization strategy. By integrating the dynamic regularization term into the loss function, this method realizes a double improvement in the structural consistency and realism of the generated image in high-precision scenarios such as medical image synthesis and autonomous driving data augmentation, without the need for additional labeled data, is compatible with the unsupervised training framework, and the dynamic optimization of the training process further improves the data utilization efficiency and model convergence effect.

[0044] The structure of the present invention is simple, easy to implement, has a wide application range, and strong versatility. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a Cycle-GAN optimization method for improving target position consistency in the process of image style transfer of the present invention.

[0046] Figure 2 It is a flowchart of the implementation of the target position difference dynamic monitoring module in the example of the present invention.

[0047] Figure 3 It is a schematic diagram of target position detection in the example of the present invention.

[0048] Figure 4 It is a schematic diagram of target position difference calculation in the example of the present invention.

[0049] Figure 5 It is a comparison diagram for optimizing the overfitting phenomenon in the example of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The present invention will be described in detail below in conjunction with the accompanying drawings of the specification, but the present invention is not limited thereto.

[0051] As Figure 1 shown, a Cycle-GAN optimization method for improving the consistency of target positions during the image style transfer process includes the following steps:

[0052] (1) Input the source domain image set X and the target domain image set Y into the Cycle-GAN network for unsupervised pre-training;

[0053] First, unify the sizes of the source domain image set X and the target domain image set Y to 256×256 pixels, and perform normalization processing and random flipping augmentation on the images to enhance data diversity. The loss function in the pre-training stage consists of the generative adversarial loss L Adv and the cycle consistency loss L Cyc . The optimizer selects the Adam algorithm (learning rate = 0.0002, β1 = 0.5, β2 = 0.999), the batch size is 1, the pre-training lasts for 100 rounds, and the model checkpoint is saved every 10 rounds to monitor the convergence state, providing a stable initial parameter basis for subsequent dynamic regularization optimization.

[0054] (2) Use the generator G obtained by the pre-training in step (1) to obtain the generated images corresponding to the target domain;

[0055] After the pre-training is completed, randomly sample n images from the source domain test set, input them into the generator G in the current training state for forward propagation, and generate target domain images. The generator adopts the parameter configuration of the current round, and the generated images and the corresponding original source domain images are saved to the specified path for subsequent dynamic detection of target position differences.

[0056] (3) Use the generated images of the target domain in step (2) to perform dynamic detection of cross-domain target position differences;

[0057] Perform object detection on the source domain input images and the target domain generated images respectively through the object detection network, extract and screen the object with the largest area. In this embodiment, the YOLO object detection network is used to obtain all the object bounding boxes of the input image and the generated image; for each image, select the single object with the largest bounding box area, and calculate the absolute difference of the abscissas of the center points of the two and normalize it. The formula is:

[0058]

[0059] where x1 is the abscissa of the center point of the largest object in the source domain image, x2 is the abscissa of the center point of the largest object in the target domain generated image, n is the number of images in the test set, and W is the image width.

[0060] (4) Utilize the target difference in step (3) to construct a dynamic regularization loss;

[0061] Calculate the L2 norm of the generator model parameters:

[0062]

[0063] θ G,i is the i-th parameter of the generator model G. Combine it with the cross-domain target position normalized difference Δx to form a dynamic regularization loss term:

[0064]

[0065] And incorporate the dynamic regularization loss L Reg into the Cycle-GAN loss function:

[0066] L = L Adv + L Cyc + λL Reg

[0067] where L is the optimized Cycle-GAN loss function, L Adv is the generative adversarial loss of Cycle-GAN, L Cyc is the cycle consistency loss of Cycle-GAN, and λ is the weight controlling the dynamic regularization term.

[0068] (5) Utilize the optimized Cycle-GAN loss function L in step (4) to optimize the generator;

[0069] Update the generator and discriminator parameters based on the total loss L. The learning rate decays to 0.5 times the original value every 50 epochs. The dynamic adjustment mechanism optimizes λ in real time according to the mean value of Δx: if Δx > 0.1, λ is increased by 0.05 to enhance regularization; if Δx < 0.05, λ is decreased by 0.02 to avoid underfitting; when Δx is between 0.05 and 0.1, λ remains unchanged.

[0070] Figure 3 is one of the generated images in the target domain obtained by the above method. The single target with the largest bounding box area is in the red frame; Figure 4 are the target regions corresponding to the source domain input image and the target domain generated image, and the cross-domain target position difference information is obtained; Figure 5 is the overfitting phenomenon optimization comparison diagram in the example of the present invention. From top to bottom, they are: (1) the source domain image, (2) the image obtained by restoring the source domain image using the optimized network of the present invention, (3) the restored image obtained by using a method for high-frequency enhancement of a superlens image based on discrete wavelet transform with the publication number CN117522703A; (4) the target domain image used during training. Figure 5It can be seen that due to overfitting during the optimization process, the restored image obtained by using the original optimization method has an obvious position shift relative to the original image; while the restored image obtained by the method of the present invention maintains good position consistency and retains complete key information.

[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the patent and not to limit them. Without departing from the principle of this patent, those of ordinary skill in the art can also make several modifications and improvements, which should also be regarded as the protection scope of the present invention.

Claims

1. A Cycle-GAN optimization method for improving the consistency of target positions during the process of image style transfer, characterized in that, It includes the following steps: (1) Perform unsupervised pre-training on the Cycle-GAN network; (2) Randomly sample n images from the source domain test set and input them into the Cycle-GAN generator in the current training state for forward propagation to output the generated images corresponding to the target domain; (3) Perform object detection on the source domain input images and the target domain generated images respectively through the object detection network, and obtain the object position difference between the two objects; (4) Combine the L2 norm of the model parameters of the Cycle-GAN generator in the current training state with the object position difference to generate a dynamic regularization loss and fuse it into the model loss function, and inversely optimize the total model loss function; (5) Update the model parameters using the optimized total model loss function; (6) Evaluate the generated images of the current obtained generator using the peak signal-to-noise ratio and structural similarity index. If the requirements are not met, return to step (2); if the requirements are met, obtain the trained Cycle-GAN network.

2. The Cycle-GAN optimization method for improving the consistency of target positions during image style transfer according to claim 1, characterized in that In step (3), when performing object detection on the source domain input images and the target domain generated images, extract and select the object with the largest area as the object corresponding to the source domain input images and the target domain generated images.

3. The Cycle-GAN optimization method for improving the consistency of target positions during the image style transfer process according to claim 1, characterized in that, In step (3), the object position difference Δx is: x1 is the abscissa of the center point of the selected object in the source domain input image, x2 is the abscissa of the center point of the selected object in the generated image, n is the number of source domain input images, and W is the image width.

4. The Cycle-GAN optimization method for improving the consistency of target positions during the image style transfer according to claim 3, wherein In step (4), the dynamic regularization loss L Reg is as follows: Δx is the target position difference, which is the L2 norm of the generator model parameters.

5. The Cycle-GAN optimization method for improving the consistency of target positions during image style transfer according to claim 4, characterized in that, In step (4), the total model loss function L is: L = L Adv + L Cyc + λL Reg L Adv is the generative adversarial loss of Cycle - GAN; L Cyc is the cycle consistency loss of Cycle - GAN; λ is the weight controlling the dynamic regularization term, and L Reg is the dynamic regularization loss.

6. The Cycle-GAN optimization method for improving the consistency of target positions during image style transfer according to claim 5, characterized in that, During the optimization process, λ is optimized in real time according to the mean value of Δx: if Δx > 0.1, λ is increased by 0.04 - 0.06 to enhance regularization; if Δx < 0.05, λ is decreased by 0.01 - 0.03 to avoid underfitting; when Δx is between 0.05 and 0.1, λ remains unchanged.

7. The Cycle-GAN optimization method for improving the consistency of target positions during image style transfer according to claim 1, characterized in that, The object detection network is the YOLO object detection network, and all object bounding boxes of the original input source domain input images and the generated images are obtained respectively; for each image, select a single object with the largest bounding box area as the corresponding object.

8. The Cycle-GAN optimization method for improving the consistency of target positions during image style transfer according to claim 1, characterized in that, In step (6), evaluate the generated images of the current obtained generator using the peak signal-to-noise ratio and structural similarity index. If the requirements are not met, return to step (2); if the requirements are met, obtain the trained Cycle-GAN network.

Citation Information

Patent Citations

  • Superlens image high-frequency enhancement method based on discrete wavelet transform

    CN117522703A