Generative adversarial network architecture search method based on multi-population coevolution
By jointly searching the network structures of the generator and discriminator through a multi-population co-evolution method, combined with channel attention and Wasserstein distance adaptive training, the problems of low efficiency and poor effect of generative adversarial network architecture search in existing technologies are solved, and more efficient network optimization and generation quality improvement are achieved.
Patent Information
- Application Number
- CN202511172427.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing generative adversarial network architecture search methods fail to fully utilize the adversarial coupling relationship between the generator and the discriminator, population diversity is difficult to maintain, and there is a lack of flexible training scheduling strategies, resulting in low search efficiency and poor results.
A multi-population co-evolution method is used to jointly search the network structure of the generator and the discriminator, the feature expression ability is enhanced through the channel attention mechanism, and the Wasserstein distance is used for adaptive training. The update rhythm is dynamically adjusted, and the gene population is optimized by combining gene similarity partitioning and crossover mutation operations.
It improves the overall collaborative performance and generation quality of the generative adversarial network, avoids falling into local optimality, improves search efficiency and convergence stability, enhances the robustness of adversarial training, and is suitable for image generation and feature synthesis tasks.
Smart Images

Figure CN120671780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep neural network technology, and in particular to a generative adversarial network architecture search method based on multi-population co-evolution. Background Art
[0002] In recent years, Generative Adversarial Networks (GANs) have garnered widespread attention for their impressive performance in image generation, data augmentation, and style transfer. GANs consist of a generator and a discriminator, which use adversarial training to achieve close approximation to the data distribution. The generator aims to produce realistic forged samples to deceive the discriminator, while the discriminator strives to distinguish between real and forged samples. These two networks continuously optimize in a dynamic game, collectively improving network performance.
[0003] However, the performance of GANs is highly dependent on the design of their network structure. Different generator and discriminator architectures significantly impact training stability and generation quality. Due to a lack of unified design guidelines, traditional GAN architectures often rely on manual experience, resulting in low efficiency and poor adaptability to diverse task scenarios. To address this issue, Neural Architecture Search (NAS) has been introduced to GAN architecture design in recent years. This method uses automated search strategies to optimize network structure, improving network performance and generalization.
[0004] Although the existing NAS-based GAN architecture search method has alleviated the burden of manual design to a certain extent, the following problems still exist: First, most methods optimize the generator and discriminator architectures separately and independently, and fail to fully utilize the adversarial coupling relationship between the two, resulting in the searched structure being unable to coordinately improve the overall adversarial performance; second, during the population evolution process, population diversity is difficult to maintain and it is easy to fall into local optimality, affecting the search efficiency and effect; third, the current architecture search method relies on a fixed training process for network structure evaluation and lacks a flexible and dynamic training scheduling strategy, which increases the search overhead and reduces the search quality. Summary of the Invention
[0005] Purpose of the invention: The technical problem to be solved by the present invention is to address the deficiencies of the existing technology and provide a method for searching a generative adversarial network architecture based on multi-population co-evolution, comprising the following steps: Step 1: Define the search space for the generator and discriminator and use binary encoding; Step 2, randomly generate the gene groups of generator and discriminator; Step 3: activate all candidate operations and perform initial training; Step 4: Use a collaborative architecture search method to search for the architecture parameters and network weights of the generative adversarial network; Step 5, using a multi-population collaborative optimization method to optimize and update the gene groups of the generator and the discriminator; Step 6: Repeat steps 3 and 4, alternately optimizing the architectures of the generator and discriminator until the network performance cannot be improved any further. Step 7: Use the searched generative adversarial network to generate unseen samples in zero-shot learning.
[0006] In step 1, the generator G is represented as a directed acyclic graph , the discriminator D is represented as a directed acyclic graph ,in, and The sets of nodes representing the generator and discriminator respectively, and They represent the connection operations of the generator and the discriminator respectively, and S represents the channel attention module common to the generator and the discriminator; In step 1, the node sets of the generator and discriminator include: input nodes, intermediate nodes, and output nodes.
[0007] In step 1, the connection operation between the generator and the discriminator and Represents the operation of connecting the nodes in the generator and the discriminator. The specific operation used will be searched from the candidate operation set E through model training, and the state of the edge is represented by one-hot encoding.
[0008] In step 1, the channel attention module S is designed to be placed before each intermediate node to process the i-th intermediate node Predecessor node Output, intermediate node Input The formula is: , in, Represents a connected node To Node Operation, Represents the jth intermediate node Input, Represents a splicing operation; The channel attention module S constructs an attention matrix by calculating the correlation weights between different node features to measure the importance distribution between channels, and introduces a learnable aggregation parameter to compress the feature dimension, thereby extracting a more discriminative feature representation. Finally, the compressed feature is used as the input of the node. The mathematical expression is: , in, represents the input feature matrix, represents the real number space, c is the number of channels, f is the feature dimension, is the channel attention coefficient of the i-th channel, is the learnable aggregation weight of the i-th channel.
[0009] In step 2, randomly generate the gene groups of the generator and discriminator: the genes of the generator and discriminator are composed of a 25×5 matrix with values 0 and 1, where the 25 rows represent the 25 connecting edges of the generator or discriminator network, and the 5 columns represent the 5 candidate operations on the connecting edges. 0 indicates inactivation and 1 indicates activation. A random candidate operation is effective on each edge.
[0010] Step 4 includes: Step 4.1, evaluate and select the optimal discriminator or generator gene, and fix the optimal discriminator or generator gene as the architecture parameter; Step 4.2: Randomly select a gene A1 from the unfixed generator or discriminator gene group. Gene A1 represents the architecture parameters of the generator or discriminator that are not fixed during the current batch training. Step 4.3: Train the network weights of the generator and discriminator using an adaptive training method based on Wasserstein distance; Step 4.4: Repeat steps 4.2 to 4.3 until one round of data training is completed.
[0011] In step 4.1, the method for evaluating the optimal discriminator and generator is: For the generator population , directly use the currently fixed discriminator for evaluation, the formula is: , in represents the currently fixed discriminator, Indicates the generator represented by the i-th generator gene The fitness of Represents the probability distribution of noise The expected value under the condition of ; i ranges from 1 to n; For the discriminator population , using Wasserstein distance as its performance evaluation value: , in, Indicates the discriminator represented by the i-th discriminator gene The fitness of Represents the probability distribution of real data The expected value under Represents the currently pinned generator.
[0012] In step 4.3, the Wasserstein distance is used as the basis for discriminator training. By monitoring the Wasserstein distance between the discriminator and the generated samples, the discriminator update is adaptively terminated. The Wasserstein distance calculation formula is: , The optimization goal of the generator architecture search is expressed as: , , , in, Represents the optimal generator architecture parameters currently evaluated, Indicates the generator represented by the i-th generator gene The architectural parameters of represents the generator weight after training, represents the discriminator weight after training, represents the fixed discriminator architecture parameters, represents the constraints, represents the generator weight, represents the discriminator weight; represents the fitness function of the generator; The optimization objective of the discriminator architecture search is stated as: , , , in, represents the optimal discriminator architecture parameters currently evaluated, Indicates the generator represented by the i-th generator gene The architectural parameters of represents the fixed generator architecture parameters, represents the fitness function of the discriminator.
[0013] Step 5 includes: Step 5.1, use fitness to evaluate and rank the architectures and filter out low-fitness genes; Step 5.2, select the unassigned gene with the best fitness as the leader L of a subpopulation; Step 5.3, traverse the remaining candidate genes g in the set to meet the similarity Higher than the set value The genes of are considered to be in the same subpopulation as the leader L; Step 5.4: Repeat steps 5.2 to 5.3 until all genes are assigned to subpopulations; Step 5.5, select the top ranking Ranked top in the subpopulation The gene of is the parent; Step 5.6, use crossover, mutation and random injection operations to generate offspring individuals; In step 5.1, use dynamic threshold Filter the genes and calculate the threshold using the following formula : , in, is the fitness function of the i-th gene, , is the set fitness threshold parameter; In step 5.3, the similarity calculation formula is: , in, is the element in row i and column j of the gene matrix of leader L, is the element in row i and column j in the gene matrix of the candidate gene g.
[0014] Beneficial effects: The present invention provides a method for searching the architecture of a generative adversarial network based on multi-population co-evolution, which has many beneficial effects. First, by jointly searching the network structure of the generator and the discriminator, the dynamic dependency between the two in the adversarial training process is fully utilized, which helps to improve the collaborative performance and generation quality of the overall network. Secondly, the introduction of a multi-population co-evolution mechanism based on gene similarity division effectively enhances population diversity, avoids falling into local optimality, and improves search efficiency and convergence stability. At the same time, the channel attention mechanism is integrated into the network structure, so that the model can automatically focus on more discriminative feature channels, thereby enhancing the feature expression ability in the generation and discrimination process. In addition, the present invention adopts an adaptive training strategy based on the Wasserstein distance to dynamically adjust the update rhythm of the generator and the discriminator, effectively alleviating the problems of training instability and mode collapse, and improving the robustness of adversarial training. Overall, the method has a flexible structural design and a wide range of applications. It can significantly improve the architectural optimization quality of the generative adversarial network and has good application prospects and practical value in multiple tasks such as image generation and feature synthesis. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is the overall framework diagram of the present invention.
[0016] Figure 2 It is the search space graph in the present invention.
[0017] Figure 3 This is a flow chart of the multi-population evolution strategy in the present invention.
[0018] Figure 4 This is a schematic diagram of the experimental analysis results when the similarity threshold parameter is set to 0.6.
[0019] Figure 5 This is a schematic diagram of the experimental analysis results when the similarity threshold parameter is set to 0.8.
[0020] Figure 6 This is a schematic diagram of the experimental analysis results when the similarity threshold parameter is set to 0.9.
[0021] Figure 7 This is a comparison chart of the experimental results of this discovery and other methods. DETAILED DESCRIPTION
[0022] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more apparent.
[0023] like Figure 1 As shown, an embodiment of the present invention provides a method for searching a generative adversarial network architecture based on multi-population co-evolution, comprising the following steps: Step 1, such as Figure 2 As shown in Figure 2, the search space for the generator and discriminator is defined and binary encoding is used. The generator and discriminator are represented as directed acyclic graphs and ,In this invention, the generator and the discriminator use the same search space structure, such as Figure 1 As shown. Among them, and The sets of nodes representing the generator and discriminator respectively, and They represent the connection operations of the generator and the discriminator respectively, and S represents the channel attention module common to the generator and the discriminator.
[0024] Step 1.1, as Figure 2 As shown, the node sets of the generator and discriminator include: input nodes, intermediate nodes and output nodes. In this example, the generator network Contains three input nodes: noise vector z, semantic vector , and the concatenation vector ; Discriminator network It also contains three input nodes: visual feature vector x, semantic vector , and the concatenation vector ; Both the generator and the discriminator contain 4 intermediate nodes and one output node.
[0025] In step 1.2, the connection operation between the nodes of the generator and the discriminator will be searched from the candidate set. In this example, the candidate operation set E includes: , Among them, FC represents the fully connected layer, ReLU and LeakyReLU are activation functions, Dropout is a regularization technique, None indicates that there is no connection between nodes, and one-hot encoding is used to represent the state of the edge.
[0026] Step 1.3, the channel attention module S is designed to be placed before each intermediate node to process the intermediate nodes Predecessor node Output, intermediate node The input formula is: , in, represents the channel attention module, Represents a connected node To Node Operation, Representative Node Input, represents the splicing operation; the channel attention module S constructs an attention matrix by calculating the correlation weights between different node features to measure the importance distribution between channels, and on this basis introduces a learnable aggregation parameter to compress the feature dimension, thereby extracting a more discriminative feature representation. Finally, the compressed feature is used as the input of the node. The mathematical expression is: , in, represents the input feature matrix, represents the real number space, c is the number of channels, f is the feature dimension, is the channel attention coefficient of the i-th channel, is the learnable aggregation weight of the i-th channel.
[0027] Step 2: Randomly generate the gene groups for the generator and discriminator. In this example, the genes for the generator and discriminator consist of a 25×5 matrix with values of 0 and 1. The 25 rows represent the 25 edges in the generator or discriminator network, and the 5 columns represent the five candidate operations on the edges. 0 indicates inactive, and 1 indicates active. For each edge, a random candidate operation is applied to generate a random gene. The gene matrix is normalized to avoid isolated nodes, dead-end nodes, or missing inputs. Ultimately, 50 different genes are randomly generated and stored.
[0028] Step 3: Activate all candidate operations and perform initial training. Specifically, use a gene matrix with all values set to 1 to set the architecture of the generator and discriminator and perform training to ensure fairness in the initial search. In this example, the initial training lasts for 5 epochs.
[0029] Step 4, such as Figure 1 As shown, collaborative search generates the architecture parameters and network weights of the adversarial network, including: Step 4.1: Evaluate and select the optimal discriminator or generator and fix its architectural parameters. This includes the following steps: Step 4.1.1, for the generator population , directly use the currently fixed discriminator To evaluate, the formula is: , in, Indicates the generator represented by the i-th generator gene The fitness of Represents the probability distribution of noise Expected value under Step 4.1.2, for the discriminator population , using Wasserstein distance as its performance evaluation value: , in, Indicates the discriminator represented by the i-th discriminator gene The fitness of Represents the probability distribution of real data The expected value under Represents the currently fixed generator; Step 4.2: Randomly select a gene from the group of unfixed generator or discriminator genes. This gene represents the architecture parameters of the generator or discriminator that have never been fixed during the training of this batch. Step 4.3: Train the network weights of the generator and discriminator using an adaptive training method based on Wasserstein distance, including: Step 4.3.1, optimize the discriminator architecture. The optimization objective of the discriminator architecture search is expressed as: , , , in, represents the optimal discriminator architecture parameters currently evaluated, Indicates the generator represented by the i-th generator gene The architectural parameters of represents the discriminator weight after training, represents the generator weight after training, represents the fixed generator architecture parameters, st represents the constraints, Represents the generator weight represents the discriminator weight; In step 4.3.2, the Wasserstein distance is used as the basis for discriminator training. By monitoring the Wasserstein distance between the real sample and the generated sample, the discriminator update is adaptively terminated. The Wasserstein distance formula is as follows: , Step 4.3.3, optimize the generator architecture. The optimization goal of the generator architecture search is expressed as: , , , in, Represents the optimal generator architecture parameters currently evaluated, Indicates the generator represented by the i-th generator gene The architectural parameters of represents the fixed discriminator architecture parameters; Step 4.4, loop through steps 3.2 and 3.3 until one round of training is completed; Step 5, such as Figure 3 As shown, a multi-population collaborative optimization method is used to optimize and update the gene groups of the generator and discriminator, including: Step 5.1, use fitness to evaluate and sort the architecture and filter out low fitness genes, set dynamic threshold to filter genes, threshold Calculated by the following formula: , in, is the fitness function of the i-th gene, , is the threshold value set. In this example, The value is taken as 0.5; Step 5.2, select the best individual as the leader of a subpopulation; Step 5.3, traverse the remaining genes in the set to meet the similarity Higher than the set value The gene g is considered to be in the same subpopulation as the leader L, and the similarity is calculated as follows: , in, is the element in row i and column j of the gene matrix of leader L, is the element in row i and column j of the gene matrix of the candidate gene g. In this example, the similarity threshold parameter The value of is obtained through experimental analysis, such as Figure 4 As shown, The value is taken as 0.8; Step 5.4, repeat steps 5.2 and 5.3 until all genes are assigned to subpopulations; Step 5.5, select the top ranking Ranked top in the subpopulation The gene of the parent is the number of children and subpopulation capacity The parameter values are obtained through experimental analysis. In this example, the maximum total number of stored genotypes is K=100, the maximum storage capacity of selected genes is K / 2=50, and the subpopulation capacity is is set to: , in, Indicates rounding down, such as Figure 4 、 Figure 5 、 Figure 6 As shown, the horizontal axis in the figure is the number of subpopulations The vertical axis is the accuracy of the searched model verified in the generative zero-shot learning task. This embodiment evaluates different similarity threshold parameters. ( Figure 4 The median value is 0.6, Figure 5 The median value is 0.8, Figure 6 The median value is 0.9) and the number of subpopulations The performance of the models searched under the same threshold is shown. The results show that as the similarity threshold increases, the optimal subpopulation size tends to decrease. This is because a higher threshold leads to an increase in the genetic similarity between individuals in the subpopulation, so that a smaller number of genotypes can represent local diversity without reducing performance. When the similarity threshold is set to 90%, the performance drops significantly, which is most likely due to the reduction in diversity between subpopulations, because a higher threshold increases redundancy within the population. This loss of diversity ultimately weakens the effectiveness of multi-population search and reduces the quality of the generated architecture. Therefore, the final setting , ; Step 5.6, use crossover, mutation and random injection operations to generate offspring individuals; Step 6: Repeat steps 3 and 4, alternately optimizing the architectures of the generator and discriminator until the network performance cannot be improved any further. In step 7, the searched GAN is used to generate unseen samples in zero-shot learning. A GAN model is constructed using the searched optimal generator and discriminator architecture. During training, the model uses its semantic vectors to generate synthetic features for seen categories. During inference, the generator synthesizes features based on the semantic vectors of unseen categories. The synthesized unseen category features are combined with seen category data to train a classifier, which is then evaluated using a test set that includes both seen and unseen data.
[0030] like Figure 7 As shown in the figure, to verify the performance, the present invention conducted experiments on the public dataset CUB, which contains 11,788 images covering 200 bird species, each of which is associated with an attribute vector. The figure shows that the method of the present invention performs well in terms of training stability, while other architecture search methods frequently oscillate during training. This is attributed to its adaptive adversarial training component, which enhances the robustness of convergence. Although the method of the present invention converges slowly in the early stages of training, it ultimately achieves higher performance, which is 4.95% and 3.59% higher than the baseline model and other architecture search methods, respectively. These results confirm the effectiveness of co-evolution in promoting better synergy between the generator and the discriminator.
[0031] The present invention provides a method for searching for a generative adversarial network architecture based on multi-population coevolution. There are many methods and approaches for implementing this technical solution. The above is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention. Any components not specified in this embodiment may be implemented using existing technologies.
Claims
1. A generative adversarial network architecture search method based on multi-population co-evolution, characterized by: The following steps are involved: Step 1: Define the search space for the generator and discriminator and use binary encoding; Step 2, randomly generate the gene groups of generator and discriminator; Step 3: activate all candidate operations and perform initial training; Step 4: Use a collaborative architecture search method to search for the architecture parameters and network weights of the generative adversarial network; Step 5, using a multi-population collaborative optimization method to optimize and update the gene groups of the generator and the discriminator; Step 6: Repeat steps 3 and 4, alternately optimizing the architectures of the generator and discriminator until the network performance cannot be improved any further. Step 7: Use the searched generative adversarial network to generate unseen samples in zero-shot learning.
2. The method according to claim 1, characterized in that In step 1, the generator G is represented as a directed acyclic graph , the discriminator D is represented as a directed acyclic graph ,in, and The sets of nodes representing the generator and discriminator respectively, and They represent the connection operations of the generator and the discriminator respectively, and S represents the channel attention module common to the generator and the discriminator.
3. The method according to claim 2, characterized in that In step 1, the node sets of the generator and discriminator include: input nodes, intermediate nodes, and output nodes.
4. The method according to claim 3, characterized in that In step 1, the connection operation between the generator and the discriminator and Represents the operation of connecting the nodes in the generator and the discriminator. The specific operation used will be searched from the candidate operation set E through model training, and the state of the edge is represented by one-hot encoding.
5. The method according to claim 4, characterized in that In step 1, the channel attention module S is designed to be placed before each intermediate node to process the i-th intermediate node Predecessor node Output, intermediate node Input The formula is: , in, Represents a connected node To Node Operation, Represents the jth intermediate node Input, Represents a splicing operation; The channel attention module S constructs an attention matrix by calculating the correlation weights between different node features to measure the importance distribution between channels, and introduces a learnable aggregation parameter to compress the feature dimension, thereby extracting a more discriminative feature representation. Finally, the compressed feature is used as the input of the node. The mathematical expression is: , in, represents the input feature matrix, represents the real number space, c is the number of channels, f is the feature dimension, is the channel attention coefficient of the i-th channel, is the learnable aggregation weight of the i-th channel.
6. The method according to claim 5, characterized in that In step 2, randomly generate the gene groups of the generator and discriminator: the genes of the generator and discriminator are composed of a 25×5 matrix with values 0 and 1, where the 25 rows represent the 25 connecting edges of the generator or discriminator network, and the 5 columns represent the 5 candidate operations on the connecting edges. 0 indicates inactivation and 1 indicates activation. A random candidate operation is effective on each edge.
7. The method according to claim 6, characterized in that Step 4 includes: Step 4.1, evaluate and select the optimal discriminator or generator gene, and fix the optimal discriminator or generator gene as the architecture parameter; Step 4.2: Randomly select a gene A1 from the unfixed generator or discriminator gene group. Gene A1 represents the architecture parameters of the generator or discriminator that are not fixed during the current batch training. Step 4.3: Train the network weights of the generator and discriminator using an adaptive training method based on Wasserstein distance; Step 4.4: Repeat steps 4.2 to 4.3 until one round of data training is completed.
8. The method according to claim 7, characterized in that In step 4.1, the method for evaluating the optimal discriminator and generator is: For the generator population , directly use the currently fixed discriminator for evaluation, the formula is: , in represents the currently fixed discriminator, Indicates the generator represented by the i-th generator gene The fitness of Represents the probability distribution of noise The expected value under the condition of ; i ranges from 1 to n; For the discriminator population , using Wasserstein distance as its performance evaluation value: , in, Indicates the discriminator represented by the i-th discriminator gene The fitness of Represents the probability distribution of real data The expected value under Represents the currently pinned generator.
9. The method according to claim 8, characterized in that In step 4.3, the Wasserstein distance is used as the basis for discriminator training. By monitoring the Wasserstein distance between the discriminator and the generated samples, the discriminator update is adaptively terminated. The Wasserstein distance calculation formula is: , The optimization goal of the generator architecture search is expressed as: , , , in, Represents the optimal generator architecture parameters currently evaluated, Indicates the generator represented by the i-th generator gene The architectural parameters of represents the generator weight after training, represents the discriminator weight after training, represents the fixed discriminator architecture parameters, represents the constraints, represents the generator weight, represents the discriminator weight; represents the fitness function of the generator; The optimization objective of the discriminator architecture search is stated as: , , , in, represents the optimal discriminator architecture parameters currently evaluated, Indicates the generator represented by the i-th generator gene The architectural parameters of represents the fixed generator architecture parameters, represents the fitness function of the discriminator.
10. The method according to claim 9, characterized in that Step 5 includes: Step 5.1, use fitness to evaluate and rank the architectures and filter out low-fitness genes; Step 5.2, select the unassigned gene with the best fitness as the leader L of a subpopulation; Step 5.3, traverse the remaining candidate genes g in the set to meet the similarity Higher than the set value The genes of are considered to be in the same subpopulation as the leader L; Step 5.4: Repeat steps 5.2 to 5.3 until all genes are assigned to subpopulations; Step 5.5, select the top ranking Ranked top in the subpopulation The gene of is the parent; Step 5.6, use crossover, mutation and random injection operations to generate offspring individuals; In step 5.1, use dynamic threshold Filter the genes and calculate the threshold using the following formula : , in, is the fitness function of the i-th gene, , is the set fitness threshold parameter; In step 5.3, the similarity calculation formula is: , in, is the element in row i and column j of the gene matrix of leader L, is the element in row i and column j in the gene matrix of the candidate gene g.
Citation Information
Patent Citations
Fan output scene generation and reduction method based on conditional generative adversarial network
CN115828441A
Ranking learning method and system based on evolution condition generative adversarial network and application
CN116245146A
Synthetic aperture radar detection performance evaluation method and device and terminal equipment
CN118625320A
Generative adversarial network architecture search method and system based on GA-PSO hybrid algorithm
CN120124685A
Method and device for generative adversarial network training
US20190130221A1