Industrial product surface defect sample generation method based on cyclic generative adversarial network

By applying a method based on circular generation adversarial network in the detection of surface defects of industrial products, using technical means such as multi-scale feature extraction and dual-capsule discriminator to generate high-quality and diverse defect samples, the problem of difficulty in generating defect samples in the existing technology is solved, and the adaptability of the detection model is significantly improved.

CN120071076APending Publication Date: 2025-05-30WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510015111.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to generate high-quality and diverse defect samples in the detection of surface defects of industrial products, resulting in poor generalization capabilities of the detection model.

Method used

Using a method based on a circular generation adversarial network (CycleGAN), samples with diverse defect characteristics are generated through multi-scale feature extraction, jump connection and dual capsule discriminator combined with perceived loss function.

Benefits of technology

It significantly improves the adaptability of the industrial defect detection model and the diversity and authenticity of generated samples, and solves the problems of insufficient sample diversity and low quality of generated samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071076A_ABST
    Figure CN120071076A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and industrial visual inspection, in particular to an industrial product surface defect sample generation method based on a cyclic generative adversarial network, and the method comprises the steps: building a training data set through collecting industrial product defect-free samples, inputting the defect-free samples into a multi-scale dense convolution U-Net enhanced generator, and obtaining a training data set; a defect sample having a random defect type is generated. In combination with a circulating double-capsule mixed discriminator, authenticity and consistency discrimination is carried out on samples through a Markov discriminator and a capsule network, and randomization of positions, shapes and number of defect samples is realized through adversarial training optimization generation network and discrimination network parameters. According to the method, the diversity and authenticity of the defect samples are effectively improved, the problems of single sample type and low generation quality in the prior art are solved, and an efficient sample generation solution is provided for industrial defect detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and industrial vision inspection, and particularly relates to a method for generating industrial product surface defect samples based on a cycle generative adversarial network. Background Art

[0002] Machine vision is widely used in the detection of industrial product surface defects. The quality and diversity of defect samples directly affect the performance and adaptability of the detection model. However, obtaining high-quality and sufficiently diverse defect samples is usually costly and difficult. Therefore, defect sample generation technology is crucial for promoting intelligent manufacturing.

[0003] Traditional defect detection methods include rule template matching, etc., with poor adaptability and difficulty in dealing with complex industrial scenarios. Although deep learning methods have certain advantages, they are restricted by the insufficient number of samples, resulting in poor generalization ability. Existing GAN technologies can generate defect samples, but have the following defects:

[0004] Insufficient sample diversity, with a single defect form generated; lack of effective capture of image details and spatial features, resulting in low-quality generated samples; the complexity of the industrial environment makes it difficult for the generation model to generalize. Summary of the Invention

[0005] In view of the many problems existing in the above-mentioned prior art, the present invention provides a method for generating industrial product surface defect samples based on a cycle generative adversarial network. Based on the cycle generative adversarial network (CycleGAN), the present invention generates samples with diverse defect features through multi-scale feature extraction and an efficient discriminator. The introduction of a perceptual loss function and a double capsule discriminator ensures the high fidelity and diversity of the generated samples, significantly improving the adaptability of the industrial defect detection model.

[0006] A method for generating industrial product surface defect samples based on a cycle generative adversarial network includes the following steps:

[0007] Collect industrial product surface defect samples and defect-free samples, and construct a training data set using the collected samples;

[0008] Input the defect-free samples into the generation network of the present invention, capture image details at different scales using the multi-scale feature extraction method, and retain low-level feature information through skip connections to generate defect samples with random defect types;

[0009] Input the generated defect samples and real defect samples into the discriminator network, and comprehensively discriminate the authenticity and integrity of the defect samples by combining the evaluation of local features and the consistency analysis of global features;

[0010] Re - input the generated defective samples into the generation network for feature extraction and reconstruction to generate defect - free samples after resetting;

[0011] Based on the adversarial training of the generation network and the discrimination network, optimize the generation network and the discrimination network to obtain a defective sample generation model.

[0012] Preferably, the generation network realizes feature extraction through a multi - scale convolution module. Convolution kernels of different scales process the input image in parallel to capture multi - level feature information.

[0013] Preferably, the implementation method of the skip connection includes connecting the down - sampling feature layer and the corresponding up - sampling feature layer in the generation network to fuse low - level and high - level feature information.

[0014] Preferably, the discrimination network adopts a dual - branch structure, including a local discrimination branch and a global discrimination branch. The local discrimination branch uses a Markov discriminator to process detailed features, and the global discrimination branch uses a capsule network to analyze the spatial relationship between features.

[0015] Preferably, the Markov discriminator performs a sliding window analysis on the local area of the image to detect the authenticity of the local area in the sample.

[0016] Preferably, the capsule network captures the spatial hierarchical relationship between input features through a dynamic routing mechanism and analyzes the overall structural relationship of the input features.

[0017] Preferably, the loss functions used in adversarial training include adversarial loss, cycle consistency loss, identity loss, and perceptual loss. Each loss function is jointly used to optimize the diversity and authenticity of the generated samples.

[0018] Preferably, the perceptual loss is based on the high - level features extracted by the pre - trained model and is used to compare the feature differences between the generated samples and the target samples to optimize the texture and detail quality of the generated images.

[0019] Preferably, the cycle consistency loss constrains the consistency of the generated samples and the reconstructed samples in terms of content structure.

[0020] Preferably, the optimization of the generation network and the discrimination network is achieved by dynamically adjusting the learning rate.

[0021] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:

[0022] The present invention carefully analyzes multi - scale image features and combines dense connections to ensure the retention of detail information and the authenticity of defective samples;

[0023] The present invention combines a Markov discriminator with a capsule network, deeply learns and evaluates the complex relationship between local defects and the global background, and effectively solves the problem that traditional discriminators are insufficient in capturing subtle features;

[0024] The present invention improves the image quality control ability and optimizes the generated texture and details through advanced feature similarity;

[0025] The present invention randomizes the positions, shapes, and quantities of defect samples, enhancing the diversity of the generated samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic diagram of the structure of the multi-scale perception enhanced capsule cycle generative adversarial network in an embodiment of the present invention;

[0027] Figure 2 It is a schematic diagram A of the generated defect samples by MSP-ECCycleGAN (Markov discriminator) in the DAGM2007 dataset in an embodiment of the present invention;

[0028] Figure 3 It is a schematic diagram B of the generated defect samples by MSP-ECCycleGAN in the DAGM2007 dataset in an embodiment of the present invention;

[0029] Figure 4 It is a schematic flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.

[0031] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0032] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0033] In the realistic context where the advancement of Industry 4.0 makes defect-free samples easy to obtain while high-value defective samples are scarce, aiming at the dual problems of lack of sample diversity and difficulty in obtaining high-quality defective samples in the field of industrial product surface defect detection, this method proposes a method for generating industrial product surface defect samples based on a cyclic generative adversarial network on the basis of the cyclic generative adversarial network. It is a model that can generate defective samples using defect-free samples, namely the Multi-Scale Perceptual Enhanced Capsule CycleGAN (MSP-ECCycleGAN). Among them, the multi-scale dense convolutional U-Net enhances the generator, improving the model's ability to accurately capture and carefully analyze the input image features at multiple scale levels. Further, the internal application of dense connections in the generator not only optimizes the transmission and efficient utilization of features within the network, but also greatly improves the model's ability to capture and encode complex and abstract defect features in industrial images. Relying on the skip connection mechanism inside the U-Net architecture, it ensures the effective retention and restoration of lower-level detail information during the complex image synthesis process, effectively avoiding the common problem of gradient disappearance in deep network training and improving the clarity and authenticity of the synthesized images.

[0034] On this basis, the cyclic dual-capsule hybrid discriminator combines the advantages of the Markov discriminator (PatchGAN) and the capsule network with a dynamic routing mechanism, deeply capturing and analyzing the spatial hierarchical structure and subtle features in the image. This strategy effectively learns and evaluates the local defect features and global background features in industrial images, significantly improving the model's performance in maintaining the detail features of defective samples and effectively solving the challenges of traditional discriminators in capturing the minute feature differences and their spatial relationships in industrial product surface defect samples. The specific network model structure is as Figure 1 shown.

[0035] In addition, a perceptual loss function is newly introduced in the model, further refining the ability to control image quality and effectively solving the problem that the generated image details may be too smooth or lacking caused by relying solely on the pixel-level loss function. By learning the similarity between images at a higher level, it promotes the improvement of the quality of the generated images, especially in terms of texture and detail performance. Using an adversarial training framework composed of perceptual loss, cyclic consistency loss, identity mapping loss, and least squares loss, the model realizes efficient random variations in the position, shape, and quantity of different types of defects, enhancing the diversity of the generated samples of industrial product surface defects.

[0036] The framework includes two generators G P2N and G N2P , and two discriminators D P2N and DN2P . The generator is optimized based on the U-Net network structure, and a perceptual loss function is added as a constraint condition to enhance the feature and semantic information of the generated images. In the discriminator, in this section, a capsule network is introduced based on PatchGAN to learn more detailed global spatial features. In addition, the method also modifies the Sigmoid cross-entropy loss function to the least squares method to solve the problem of gradient disappearance during the training process, avoiding mode collapse and unstable training.

[0037] During the generation process of the CycleGAN model, the loss function includes two parts: the adversarial loss function and the reconstruction loss function. The former aims to make the data distribution of the generated images as similar as possible to the data distribution of the target domain, generating more realistic images. The latter is mainly used to prevent the confusion and misalignment of the mapping relationship between the source domain and the target domain. To generate more realistic industrial product surface defect images, on this basis, in order to improve the performance of the generation model in terms of detail reproduction and sample authenticity when processing industrial product surface defect images with diverse and complex textures, the MSP-ECCycleGAN introduces a perceptual loss function and a capsule loss function, and adds an identity mapping loss to better learn the industrial product surface defect features. The overall loss function of the model is shown in Equation (1).

[0038]

[0039] where G* and D* represent the optimal cases of the generator and discriminator respectively, means minimizing the generator and maximizing the discriminator. X and Y represent the source domain and the target domain respectively, and L represents calculating the loss function Loss for (G P2N , D P2N , G N2P , D N2P , X, Y). In this generative adversarial network, the two-player zero-sum game strategy is also adopted, which belongs to a typical minimax duel problem. That is, finally, the generator and the discriminator compete with each other to reach the Nash equilibrium, as shown in Equation (2).

[0040] L MSP-ECCycleGAN = L GAN (G P2N , D P2N , X, Y)

[0041] + L GAN (G N2P , D N2P , Y, X)

[0042] + λ 1 L cycle (G P2N , G N2P , X, Y)

[0043] + λ2 L identity (G P2N ,G N2P ,X,Y)

[0044] +λ 3 L perceptual (G P2N ,G N2P ,X,Y)(2)

[0045] where L GAN represents the adversarial loss, L cycle represents the cycle loss, L identity represents the identity loss, and L perceptual represents the perceptual loss. λ 1 is the weight of the cycle loss, and λ 2 is the weight of the identity loss, and λ 3 is the weight of the perceptual loss.

[0046] The cycle generative adversarial model uses the logarithmic cross-entropy loss function as the adversarial loss function. This loss function is more suitable for dealing with logical classification problems, but it will cause the problem of vanishing gradients during training, affecting the convergence and optimization of the model. Therefore, in this section, the least squares method is used as the adversarial loss function, and combined with the margin loss of the capsule network to improve the stability and convergence of training, and ensure the authenticity and diversity of the generated samples. In the model training loop, the loss functions of the two parts where defective industrial samples and non-defective industrial samples generate each other are the same. Therefore, the following formula introductions are all based on the generator G P2N and the discriminator D P2N Taking the generation of defective samples from non-defective samples as an example, the adversarial loss function is shown in formula (3). Where x is the non-defective sample and y is the defective sample.

[0047]

[0048] Among them, represents all data in the form of a data distribution, and D P2N-1 is the discriminator branch using the PatchGAN structure, and D P2N-2 is the discriminator branch constructed based on the capsule network. The first line on the right side of the equal sign represents the loss function of the generator, and its optimization goal is to make the discriminator's judgment value D P2N-1 (G(x)) approach 1. The second and third lines are the loss functions corresponding to the discriminator, and its optimization goal is to make the discriminator's evaluation value D P2N-1 (y) approach 1, and the judgment value D P2N-1(G(x)) approaches 0. The fourth and fifth lines are the margin loss functions, and λ is a hyperparameter representing the relative importance of the margin loss in the improved adversarial loss. The capsule discriminator only needs to determine whether the input image is a real image or a generated fake image, which is specifically defined by Equation (4).

[0049] L M (k, v) = T K max(0, m + - v k ) 2 + λ(1 - T K )max(0, v k - m - ) 2 (4)

[0050] where v k refers to the vector output by the discriminative layer of the capsule network. k = 0 represents real data, while k = 1 represents fictional data generated by the generator. If it is desired that the discriminator or generator determines the current data as real, then set T K = 1; conversely, if it is desired to make the discriminator recognize the current data as generated forged data, then set T K = 0, that is, T K is equivalent to the function of assigning labels. m + and m - serve as the judgment thresholds for evaluating the authenticity of the input image respectively. When the norm of the vector exceeds m + , the judgment result is 0; if the norm of the vector is lower than m + , then the square of the difference between the two norms is returned as the result; conversely, if the norm of the vector is less than m - , then the result is 0; and when the norm of the vector is greater than m - , the square value of the difference between the two norms is also returned.

[0051] To further strengthen the control of image quality at the visual perception level, the model adds a perceptual loss function during the training process to enhance the features and semantic information of the generated images. Different from the pixel-level direct comparison, the perceptual loss function optimizes the entire training process of the model by evaluating the high-level feature differences between the generated image and the target image, making the generated image more real and clearer, and better simulating the defect characteristics in the real world. The formula is shown in (5), where φ is the pre-trained VGG19 feature extractor.

[0052]

[0053] Identity loss and cycle loss ensure that the generated defective images are consistent with the input defect-free images in terms of content and structure, prompting the model to effectively generate industrial images with target defect features by constraining the generator under the premise of the same background. The formulas are shown in (6) and (7).

[0054]

[0055] MSP-ECCycleGAN completes enhanced training in Class5-Class10 of the DAGM2007 dataset to evaluate the diversity of generated defective samples. The training is carried out for 100 iterations in total, with a batch size of 1. The initial learning rate is set to 0.0002, and the learning rate starts to decay linearly from the 50th iteration. For the hyperparameters in the loss function, the most suitable values are selected through experiments. In this section, λ 1 is set to 5, λ 2 is set to 0.5, λ 3 is set to 0.5, m in the capsule network + is set to 0.9, m - is set to 0.1, λ is set to 0.5, and the parameter β 1 is set to 0.9, β 2 is set to 0.999 Adam optimizer for subsequent gradient calculation. Other experimental environment parameters are the same as those in the previous chapter. The generated defective samples during training are as shown in Figure 2 and Figure 3 shown.

[0056] By observing the analysis of the generation situation of MSP-ECCycleGAN during training, it is found that the model demonstrates a powerful ability to generate a rich variety of defect types and quantities against the background of defect-free real samples. Especially in datasets such as Class7, Class8, and Class9 where background information interference is relatively frequent, even when the defect features are not very prominent, MSP-ECCycleGAN can still generate defective samples with obvious features and a relatively large number of types. This shows that while the model effectively grasps the complex relationship between local defect features and global background features, it can also efficiently refine and retain the delicate features of defective samples. In datasets like Class6 and Class10 where defect features are obvious, the texture details of the generated defective samples are well preserved and show diverse features. This indicates that the model can accurately capture and meticulously analyze defect features at multiple scale levels, improving the model's encoding ability for complex and abstract defect features in industrial images and further enhancing the quality control of image generation.

[0057] To deeply explore the comprehensive impact of the new design on the quality of generated defect images, ablation experiments were designed and completed in Class6 - Class10 of DAGM2007, and objective image quality evaluation metrics and the number of training parameters were used to evaluate the performance of MSP - ECCycleGAN, reflecting its effects and contributions in various aspects. The specific experimental results are recorded in Table 1. Among them, ING, IND, MSPG, ECD, and PLF represent the initial generator, initial discriminator, multi - scale dense convolutional U - Net enhanced generator, cyclic double - capsule hybrid discriminator, and perceptual loss function respectively.

[0058] Table 1 Comparison of image quality evaluation metrics in the ablation experiment of MSP - ECCycleGAN

[0059]

[0060] By comparing the data in the table, it is found that after introducing the multi - scale dense convolutional U - Net enhanced generator and the cyclic double - capsule hybrid discriminator, compared with the traditional generative adversarial network, the image quality evaluation metrics PSNR, SSIM, and IS are increased from 26.23, 0.58, and 1.86 to 32.15, 0.64, and 2.12 respectively, and FID is decreased from 128.26 to 95.66, indicating that the model performance has been improved. On this basis, by integrating the perceptual loss function, PSNR, SSIM, and IS are further increased from 32.15, 0.64, and 2.12 to 33.85, 0.66, and 2.17 respectively, and FID is decreased from 95.66 to 93.63.

[0061] It shows that the perceptual loss function can further optimize the defect feature performance and the depth of semantic information of the generated samples. The entire experimental results prove that the new design can effectively improve the model's ability to accurately retain and restore low - level detail information during the synthesis of industrial product surface defect images. Especially in terms of the randomness of the type, location, shape, and quantity of industrial image defects, it demonstrates more excellent adaptability and flexibility.

[0062] The specific implementation steps include:

[0063] (1) Collect industrial product surface defect samples and defect - free samples to construct a training dataset

[0064] (2) Input the real defect-free samples into the multi-scale dense convolutional U-Net enhanced generator. First, the network passes the input image through the convolutional layer, and uses the convolutional operation to expand the channels. The size of the convolutional kernel of this layer is set, and the step value is set to 1. Subsequently, the feature maps are collected and reconstructed through four rounds of downsampling and four rounds of upsampling. Finally, the generated defective samples are output through the activation layer. To strengthen the extraction of local detail features of the defective samples and improve the training efficiency and accuracy of the network, in the initial three-layer stage of downsampling, a multi-scale fusion strategy is adopted to accumulate the outputs of each layer and its previous layer, and then send them to the next processing layer. The residual module is combined to further ensure the integrity of the sample features during downsampling. In the transformer layer, the dense connection convolutional module is used to replace the residual network, greatly reducing the number of parameters and the amount of calculation. During the upsampling process, the upsampling module uses skip connections to splice the downsampled feature maps of the corresponding scales, and on the basis of the nearest interpolation upsampling, the feature fusion is achieved through the residual module and then input into the next layer, and finally the generated defective samples are output.

[0065] (3) Input the generated defective samples and the real defective samples into the cyclic double capsule hybrid discriminator at the same time. At this time, after the input image is downsampled through three rounds of feature extraction, it is divided into two branches for result output. The former adopts the original discriminator structure of PatchGAN, which is responsible for judging the local authenticity of the image, and the latter uses the capsule network to realize sample discrimination, which is responsible for judging the global consistency of the image, and the final discrimination result is obtained.

[0066] (4) Input the generated defective samples into the multi-scale dense convolutional U-Net enhanced generator again, and repeat step (2) to obtain the reset defect-free samples.

[0067] (5) Input the reset defect-free samples into the cyclic double capsule hybrid discriminator again, and repeat step (3) to obtain the final discrimination result.

[0068] (6) Calculate and update the loss of G P2N Loss of G P2N Loss:

[0069] G P2N Loss = L GAN (G P2N , D P2N , X, Y) + λ 1 *L cycle (G P2N , G N2P , X, Y)

[0070] +λ 2 *L identity (G P2N , X) + λ 3 *L perceptual(G P2N ,X,Y)(1)

[0071] All the formulas are as follows:

[0072]

[0073] L M (k,v) = T K max(0,m + -v k ) 2 +λ(1 - T K )max(0,v k -m - ) 2 (3)

[0074]

[0075]

[0076] (7) Update the parameters of G to minimize the loss G P2N Loss P2N Loss

[0077] (8) Calculate and update the loss of D P2N Loss of D P2N Loss: D P2N Loss = L GAN (D P2N ,G P2N ,X,Y)

[0078] (9) Update the parameters of D to minimize the D P2N Loss P2N Loss

[0079] (10) Then input the real defective samples into the generator, repeat steps (2)-(5), and update the parameters of D P2N to minimize the loss D P2N Loss

[0080] (11) Calculate and update the loss of G N2P Loss of G N2P Loss:

[0081] G N2P Loss = L GAN (G N2P ,D N2P ,Y,X)+λ 1 *L cycle (G N2P ,G P2N ,Y,X)

[0082] +λ2 *L identity (G N2P ,Y)+λ 3 *L perceptual (G N2P ,Y,X)

[0083] (12) Update G N2P 's parameters to minimize the loss G N2P Loss

[0084] (13) Calculate and update D N2P 's loss D N2P Loss: D N2P Loss = L GAN (D N2P ,G N2P ,Y,X)

[0085] (14) Update D N2P 's parameters to minimize the loss D N2P Loss

[0086] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.

[0087] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for generating surface defect samples of industrial products based on a cyclic generative adversarial network, characterized in that: The following steps are involved: Collect surface defect samples and non-defect samples of industrial products, and use the collected samples to build a training data set; Inputting defect-free samples into the generative network of the present invention, using a multi-scale feature extraction method to capture image details of different scales, retaining low-level feature information through skip connections, and generating defective samples with random defect types; The generated defect samples and real defect samples are input into the identification network, and the authenticity and integrity of the defect samples are comprehensively judged by combining the evaluation of local features and the consistency analysis of global features; The generated defective samples are re-input into the generation network for feature extraction and reconstruction to generate reset defect-free samples; Based on the adversarial training of the generation network and the identification network, the generation network and the identification network are optimized to obtain the defect sample generation model.

2. The method according to claim 1, characterized in that The generative network realizes feature extraction through a multi-scale convolution module, and convolution kernels of different scales process the input image in parallel to capture multi-level feature information.

3. The method according to claim 1, characterized in that The implementation method of the jump connection includes connecting the down-sampling feature layer in the generation network with the corresponding up-sampling feature layer to fuse the low-level and high-level feature information.

4. The method according to claim 1, characterized in that The identification network adopts a dual-branch structure, including a local discriminant branch and a global discriminant branch. The local discriminant branch uses a Markov discriminator to process detail features, and the global discriminant branch uses a capsule network to analyze the spatial relationship between features.

5. The method according to claim 4, characterized in that The Markov discriminator performs sliding window analysis on local regions of the image to detect the authenticity of the local regions in the sample.

6. The method according to claim 4, characterized in that The capsule network captures the spatial hierarchical relationship between input features through a dynamic routing mechanism and analyzes the overall structural relationship of the input features.

7. The method according to claim 1, characterized in that The loss functions used in adversarial training include adversarial loss, cycle consistency loss, identity loss, and perceptual loss. The loss functions are used together to optimize the diversity and authenticity of generated samples.

8. The method according to claim 7, characterized in that The perceptual loss is based on high-level features extracted by the pre-trained model and is used to compare the feature differences between the generated samples and the target samples to optimize the texture and detail quality of the generated images.

9. The method according to claim 7, characterized in that: The cycle consistency loss constrains the consistency of the content structure between the generated samples and the reconstructed samples.

10. The method according to claim 1, characterized in that The optimization of the generator network and the discriminator network is achieved by dynamically adjusting the learning rate.