Generative adversarial network architecture search method and system based on GA-PSO hybrid algorithm
Global and local optimization in generative adversarial network architecture search is solved through the GA-PSO hybrid algorithm, and the problems of inefficiency and training instability in the existing methods are solved, achieving high-quality and diverse generation samples.
Patent Information
- Application Number
- CN202510580529.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-07
AI Technical Summary
Existing methods have inefficiency, training instability, and multi-objective collaboration problems in Generative Adversarial Network (GAN) architecture search, resulting in insufficient quality and diversity of generated samples.
Using a method based on GA-PSO hybrid algorithm, a high-performance generative adversarial network architecture is constructed through global exploration of genetic algorithms and local optimization collaborative search of particle swarm optimization. The method includes building a hypernet architecture, multi-objective optimization using the NSGA-II algorithm, and tuning architecture parameters through particle swarm optimization.
It effectively breaks through the limitations of traditional methods that are prone to falling into small model traps, improves search efficiency, obtains high-quality and diverse generation samples, and provides an efficient and robust automated design solution.
Smart Images

Figure CN120124685A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for searching network architectures, specifically a method and system for searching generative adversarial network architectures based on a GA-PSO hybrid algorithm, belonging to the technical field of deep neural networks. Background Art
[0002] Since the generative adversarial network (GAN) was proposed in 2014, with the adversarial training mechanism of the generator and discriminator, it has shown great potential in fields such as image generation and super-resolution reconstruction. However, its practical application still faces multiple challenges. The training process of GAN is highly unstable and is easily affected by problems such as mode collapse and vanishing gradients, resulting in fluctuations in the quality of generated samples and insufficient diversity. Moreover, the network architecture design has long relied on expert experience. Traditional manually designed models need to adjust parameters through a large number of trials and errors, making it difficult to meet the requirements of complex tasks. In addition, the network is extremely sensitive to architecture details, and subtle structural differences may lead to a significant degradation in generation performance. In recent years, neural architecture search (NAS) technology has provided new ideas for automatically designing networks, but there are still key bottlenecks in its application in GAN: Differentiable search methods based on gradient optimization are prone to failure due to the non-convexity of the adversarial training objective and the dynamic imbalance between the generator and discriminator; Reinforcement learning methods need to independently train and validate each candidate architecture, and since a single training of GAN takes hundreds of rounds, the search cost increases sharply; Evolutionary algorithms can globally explore the architecture space, but tend to select models with simple structures and few parameters, leading to the "small model trap". Simple architectures dominate the population due to fast training convergence, while complex architectures are prematurely eliminated due to the need for long-term tuning, further exacerbating the training instability. In existing research, multi-objective optimization attempts to balance generation quality and distribution matching, but under the weight sharing mechanism, the gradient update of the supernetwork is dominated by simple architectures, and the parameters of complex models are difficult to fully optimize, forming a self-reinforcing resource tilt. In addition, the traditional Pareto front screening does not effectively combine local fine-grained search, the sparsity of the solution set distribution is insufficient, and the trade-off between the number of parameters and generation quality emphasizes computational efficiency, resulting in the suppression of the potential of deep models. These problems highlight the systematic defects of existing methods in terms of efficiency, stability, and multi-objective coordination, and it is urgent to break through the core bottleneck of automated architecture search. Summary of the Invention
[0003] Object of the Invention: Aiming at the above problems, the object of the present invention is to provide a method and system for searching generative adversarial network architectures based on a GA-PSO hybrid algorithm, which synergistically searches for high-performance network architectures through the global exploration of genetic algorithms and the local optimization of particle swarm optimization, constructs a supernetwork, and evaluates subnetworks based on weight sharing to find the generative adversarial network architecture with the optimal performance.
[0004] Technical Solution: On the one hand, the present invention proposes a method for searching generative adversarial network architectures based on a GA-PSO hybrid algorithm, including the following steps: Step 1, construct the supernetwork architecture of the generative adversarial network, including a fully connected layer, an upsampling layer, a downsampling layer, and multiple convolutional modules. The multiple convolutional modules constitute the search space, and the candidate operations in the search space are defined; Step 2, use the supernetwork architecture as an individual in the population, and perform global search using the NSGA-II algorithm to obtain the non-dominated layer; Step 3, perform hierarchical screening on the population through non-dominated sorting and crowding distance to obtain elite offspring; Step 4, use the particle swarm optimization algorithm to perform local parameter adjustment on the top n 0 elite offspring, where ; Step 5, determine whether the iterative termination condition is satisfied. If not, jump to Step 2; otherwise, terminate the search and output the optimal supernetwork architecture that meets the convergence condition.
[0005] Furthermore, Step 2 includes: Step 201, randomly generate an initial population P 0 composed of supernetwork architectures, and the individual encoding corresponding to each supernetwork architecture is an integer matrix; Step 202, calculate the two-objective metrics: IS and FID for each supernetwork architecture through the pre-trained Inception v3 network. IS represents maximizing the image quality, and FID represents minimizing the gap with the real distribution; Step 203, calculate the objective function value of each supernetwork architecture, and the calculation formula is: , , , In the formula, represents the supernetwork architecture parameters, represents the optimal supernetwork architecture parameters, is a multi-objective optimization function, represents minimizing the distribution difference between the generated samples and the real samples, represents maximizing the diversity and quality of the generated samples, is the weight of the supernetwork generator, is the weight of the optimal supernetwork generator, represents the calculation of the expectation of the random noise variable z, and calculates the statistical expectation on the input noise distribution of the generator G(z), that is, the expected performance measurement of the generator generating fake samples through the noise vector, is the adversarial loss of the generator; represents the weight of the supernetwork discriminator, is the weight of the optimal supernetwork discriminator, Denotes the expected value calculation for the real data sample 𝑥, calculating the expected value of the discriminator D(x)'s response to samples under the real data distribution. Is the discriminative loss of the discriminator D for real data. The discriminator D needs to maximize the probability D(x) that the real sample x is judged to be real, and the generator tries to make the generated sample The probability of being judged real by the discriminator D Approach 0; Step 204, screening solutions through the elitist retention strategy, screening out the top n 1 Individuals directly enter the next generation, and the recombined and crossed offspring and the mutated individuals generate a new population, where ; Step 205, dividing the new population into different non-dominated levels through non-dominated sorting.
[0006] Furthermore, in step 205, dividing the population into different non-dominated levels through non-dominated sorting includes: If the IS of individual A ≥ the IS of individual B, and the FID of individual A ≤ the FID of individual B, then individual A dominates individual B. Individuals not dominated by any individual belong to the first layer. Remove the dominated individuals layer by layer until all individuals are stratified to obtain different non-dominated levels.
[0007] Furthermore, step 3 includes: Select individuals from the first non-dominated layer according to the preset number of selections. When the number of individuals in the first non-dominated layer is insufficient, it is supplemented by the next non-dominated layer, and so on; when individuals must be selected from the same non-dominated layer, give priority to selecting individuals with a large crowding distance; use the selected individuals as elite offspring.
[0008] Furthermore, before calculating the crowding distance, standardize the IS and FID values to the [0,1] interval, set the objective space as a two-dimensional grid, and take the sum of the interval distances between adjacent solutions of each network structure in the two objective dimensions of IS and FID to calculate the crowding distance of network structure i. The formula is: , In the formula, Denotes the objective Dimension, the objective value of the next adjacent individual After sorting, Denotes the objective Dimension, the objective value of the previous adjacent individual After sorting.
[0009] Furthermore, step 4 includes: Step 401, the first n 0The elite offspring is used as the initial particle swarm. The corresponding integer matrix is flattened into a vector, and each dimension ranges from [0, 6], which serves as the initial position of the particle. The velocity vector is set to random floating-point numbers in the interval [-1, 1]. Step 402: Calculate the fitness function, which fuses FID and IS with boundary constraints. The fitness function is as follows: , In the formula, represents the boundary constraint penalty term, represents the weight coefficient used to adjust the penalty intensity, represents the dimension parameter of particle k; Step 403: Linearly decrease the inertia weight from to , and calculate the inertia weight at the current iteration. The formula is: , In the formula, t is the current iteration number, T max is the maximum number of iterations, is the initial value of the inertia weight, is the final value of the inertia weight; Step 404: Update the velocity. The formula is: , In the formula, is the learning factor, r 1 , r 2 are random numbers in [0, 1], is the best architecture parameter generated by this particle in history, is the current global best generator architecture parameter, continues the original movement direction of the particle, w is the inertia weight, is the th iteration of the velocity vector of particle , is the th iteration of the position vector of particle i, drives the particle to approach its own optimal solution, guides the particle to aggregate towards the global optimal solution; Step 405: After the velocity is updated, obtain the floating-point position increment. Round the floating-point result to an integer according to the integer encoding of the convolution type and number of channels of the architecture parameter, and use the np.clip function to forcibly constrain the value range of the architecture parameter and limit it within the predefined integer range. The formula is: , In the formula, Obtain the enumeration values of the convolution type and the number of channels of the architecture parameters, which respectively represent the minimum and maximum operation encodings; Step 406, update the position, and the formula is: , and clip it to ; which respectively represent the minimum and maximum values of the particle position parameters, is the particle velocity vector at the t-th iteration; Step 407, stop the iteration when the maximum number of iterations is reached or the solution converges, and obtain the optimal architecture solution; otherwise, return to Step 402.
[0010] On the other hand, the present invention proposes a generative adversarial network architecture search system based on the GA-PSO hybrid algorithm, including: A construction module that constructs a generative adversarial network supernet, with a cell-level configuration of 3 3 convolutions, 5 5 dilated convolutions, skip connection candidate operations, allowing a maximum of 3 input edges for the connection topology between nodes, and supporting dynamic selection of the normalization layer and activation function; A training module that only activates one bilinear upsampling or 3 3 depthwise separable convolution operation paths each time, and performs adversarial training with a discriminator of a fixed structure, adopting a non-repetitive uniform sampling strategy, and each path needs to be fairly iterated in the same adversarial environment; An optimization module that generates an initial population based on the genetic algorithm, explores a wide solution space through crossover and mutation; particle swarm optimization tunes the results of the genetic algorithm, adopts dynamic inertia weight and gradient sensitivity analysis, and breaks through the local optimum; the mixing ratio is set to n 3 genetic algorithm search and (1 - n 3 ) particle swarm optimization, where 65% < n 3 < 75%, and global particle swarm optimization reinforcement is performed every preset number of generations, and the architecture parameters are ensured to meet the preset constraints through discretized position update.
[0011] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: By integrating the global exploration ability of the genetic algorithm and the local optimization ability of particle swarm optimization, adopting the NSGA-II multi-objective optimization framework, generating an initial Pareto front solution set and dynamically adjusting the inertia weight and learning factor, the present invention effectively breaks through the limitation that traditional evolutionary algorithms are prone to falling into the small model trap; During the search process, by constructing a supernet architecture containing multi-granularity modules, including 9 searchable convolution modules, supporting dynamic selection of the optimal operator from 5 candidate convolution operations, and combining the weight sharing mechanism to directly inherit the subnet weights, the search efficiency is effectively improved; To solve the problem of computing resource consumption, the present invention designs a hierarchical coding strategy to decompose architecture parameters into discrete topological variables and continuous hyperparameters, maintains the diversity of the solution set through crowding distance calculation, and adopts a non-replacement sampling strategy to eliminate evaluation bias. This collaborative optimization mechanism enables the final optimal architecture obtained, providing an efficient and robust solution for the automated design of generative adversarial networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is the flowchart of the GA-PSO hybrid algorithm in the present invention; Figure 2 is the architecture diagram of the hypernetwork convolutional module of the generative adversarial network in the present invention; Figure 3 is the basic framework diagram of the hypernetwork training of the generative adversarial network in the present invention; Figure 4 is the schematic diagram of the architecture search space of the generative adversarial network in the present invention; Figure 5 is the trend diagram of the IS score of architecture optimization in the present invention; Figure 6 is the trend diagram of the FID score of architecture optimization in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0014] Embodiment 1 A method for searching the architecture of a generative adversarial network based on the GA-PSO hybrid algorithm described in this embodiment has a flowchart as Figure 1 shown, and includes the following steps: Step 1, construct the hypernetwork architecture of the generative adversarial network, including a fully connected layer, an upsampling layer, a downsampling layer, and multiple convolutional modules. The multiple convolutional modules constitute the search space, and the candidate operations in the search space are defined; In this example, the generative adversarial network includes a generator and a discriminator. The network structure of the generator is not fixed and needs to perform an optimization search in the search space. The hypernetwork in the generative adversarial network is trained using the single-path sampling strategy method, and only one path in the hypernetwork generator is activated and trained adversarially with the discriminator with a fixed structure.
[0015] In one example, the backbone architecture hypernetwork of the generator includes 1 fully connected layer, 4 downsampling layers, 4 upsampling layers, and 9 convolutional modules. Among them, the fully connected layer, downsampling layer, and upsampling layer are fixed layers and do not participate in the search. The convolutional modules are search layers, and the search process is to select an optimal operation from each edge of the convolutional module as the operation of the final optimal subnet.
[0016] Each downsampling layer contains 1 global max pooling operation; each upsampling layer contains 1 transposed convolution operation; each convolution module contains 3 nodes and 2 edges, where the 3 nodes represent 3 groups of feature maps, and the 2 edges represent 2 groups of convolution operations. Each group of convolution operations contains 2 ordinary convolution operations and 3 depthwise separable convolution operations.
[0017] The edges connecting different intermediate nodes perform certain operations to process the data of the input nodes and finally output the results by the output nodes. The candidate operations include: Convolution, Convolution, Depthwise separable convolution, Depthwise separable convolution and Depthwise separable convolution operations to achieve effective feature extraction, information fusion and resolution adjustment to improve the performance and performance of the model.
[0018] To more clearly express the network architecture of the generative adversarial network, in the present invention, a matrix is used to represent the architecture of a generator. This matrix represents 7 types of parameters in 3 network layers, including convolution kernel size, number of channels, normalization method, etc. The value range is limited to [0, 6]. The 0 within the interval represents convolution, 1 represents convolution, 2 represents depthwise separable convolution, 3 represents depthwise separable convolution, 4 represents depthwise separable convolution, 5 represents hybrid convolution, and 6 represents skip connection.
[0019] In one example, as Figure 2 shown in the architecture diagram of the supernet convolution module, the supernet contains multiple convolution modules, and each module contains multiple candidate operations, such as convolution, convolution, depthwise separable convolution, etc. The search space adopts a multi-granularity modular design, covering dual configurable dimensions at the cell level and the network level. At the cell level, each module supports dynamically selecting operators from a candidate set containing operations such as , , dilated convolution, skip connection, etc., and allows configuring the connection topology between nodes, with a maximum of 3 input edges, and can also adaptively select the normalization layer and ReLU / Swish activation function; at the network level, the global architecture parameters include the number of subnets of the generator, whether to insert a downsampling pooling layer, and the dynamic scaling strategy of the number of channels in each layer, such as . In addition, through the weight sharing mechanism of the dynamic supernet, all candidate architectures share the same set of differentiable and relaxed weight parameters to achieve end-to-end efficient joint optimization.
[0020] As Figure 4 shown in the generative adversarial network architecture search space, the input noise first enters the fully connected layer, and then enters the generator module composed of Unit 1, Unit 2, and Unit 3. Figure 4 In Figure 4 , the numbers 0, 1, 2, 3, and 4 represent the five nodes contained in each unit. There are corresponding convolution operations between each node, including multiple operations such as upsampling operation, normalization operation, and downsampling operation. The nodes are connected through these operations. The generator network architecture unit includes an upsampling operation and a normalization operation, and the discriminator network architecture unit includes a downsampling operation and a normalization operation. The parameters of the discriminator supernet remain fixed and do not participate in the architecture search.
[0021] The supernet is an underlying architecture library containing multiple candidate operations. It allows all subnets to share the same set of differentiable parameters through a weight sharing mechanism to reduce the search cost; the subnet is an instance architecture generated by dynamically selecting specific operation combinations from the candidate set of the supernet. The supernet is the parent framework, and the subnet is the optimized configuration searched by it. The architecture includes multiple convolution modules, and each module contains multiple candidate operations, such as ordinary convolution, depthwise separable convolution, etc. The feature maps generated by these operations are fused by element-wise addition to finally obtain a new mixed feature map. During the training process, through learning and optimization, the supernet can automatically select and adjust the weights of these candidate operations, thereby finally determining the optimal operation combination and network structure.
[0022] To ensure fairness and reduce bias, a non-replacement sampling strategy is adopted to ensure that under the same discriminator adversarial training, different supernetwork operations are trained for the same duration. After a certain number of trainings, a commonality-based strategy is adopted to identify potential good operations or eliminate poorly performing operations, improving the training efficiency and reducing the negative impact of bad paths on good paths. Since the supernetwork training method adopts a single-path sampling strategy, the subnet can directly inherit the weights from the supernet without retraining, thus obtaining good performance. Therefore, the present invention uses the trained supernet as an auxiliary performance evaluator, where the subnet evaluation process does not require retraining but directly inherits the weights of the supernet to evaluate and search for potential powerful architectures.
[0023] Step 2: Using the supernet architecture as an individual in the population, the NSGA-II algorithm is used for global search to obtain the non-dominated layer.
[0024] Further, Step 2 includes: Step 201: Randomly generate 100 initial populations P 0 composed of supernet architectures, and the individual encoding corresponding to each supernet architecture is an integer matrix, and the matrix values are strictly limited to the interval [0, 6]. The 0 within the interval represents Convolution, where 1 represents Convolution, where 2 represents Depthwise separable convolution, where 3 represents Depthwise separable convolution, where 4 represents Depthwise separable convolution, where 5 represents Hybrid convolution, where 6 represents skip connection; Step 202, each supernet architecture calculates two-objective metrics: IS and FID through a pre-trained Inception v3 network. IS represents maximizing image quality, and FID represents minimizing the gap from the true distribution; The calculation formula of IS is as follows: , In the formula, represents taking the expectation over all samples x under the distribution of the generator , represents a divergence calculation function that measures the difference between two distributions and . The smaller the value, the closer the distributions are. represents the conditional probability, that is, the probability that the generated sample x belongs to class y. represents the marginal probability, that is, the class distribution of all generated samples; The calculation formula of FID is as follows: , In the formula, is the mean of the two feature distributions, is the covariance matrix of the two feature distributions, and Tr is the sum of the diagonal elements of the matrix, which is used to quantify the difference of the covariance matrix; Step 203, calculate the objective function value of each supernet architecture. The calculation formula is: , , , In the formula, represents the supernet architecture parameters, represents the optimal supernet architecture parameters, is a multi-objective optimization function, represents minimizing the distribution difference between the generated samples and the real samples, represents maximizing the diversity and quality of the generated samples, is the weight of the supernet generator, is the weight of the optimal supernet generator, represents the expected calculation of the random noise variable z, and calculates the statistical expectation on the input noise distribution of the generator G(z), that is, the expected performance measure of the generator to generate fake samples through the noise vector, is the adversarial loss of the generator G; represents the weight of the supernet discriminator, is the weight of the optimal supernet discriminator, represents the expected calculation of the real data sample 𝑥, and calculates the expected response of the discriminator D(x) to the sample under the real data distribution. is the discriminant loss of the discriminator D on the real data. The discriminator D needs to maximize the probability D(x) that the real sample x is judged as real. G tries to make the generated sample The probability of being judged as true by the discriminator D Approaching 0; Step 204, screen the solution through the elite retention strategy, select the top 10% individuals according to the objective function value and directly enter the next generation, recombining the crossover offspring and the mutated individuals to generate a new population; Step 205 , dividing the new population into different non-dominated layers by non-dominated sorting.
[0025] Further, in step 205, dividing the population into different non-dominated layers by non-dominated sorting includes: If the IS of individual A ≥ the IS of individual B, and the FID of individual A ≤ the FID of individual B, then individual A dominates individual B. Individuals that are not dominated by any individual belong to the first layer. The dominated individuals are removed layer by layer until all individuals are stratified and different non-dominated layers are obtained.
[0026] Step 3: Perform stratified screening on the population through non-dominated sorting and crowding distance to obtain elite offspring.
[0027] Further, step 3 includes: According to the preset number of screening, individuals are selected from the first non-dominated layer. When the number of individuals in the first non-dominated layer is insufficient, they are supplemented by the next non-dominated layer, and so on. When individuals must be selected from the same non-dominated layer, individuals with high crowding degree are given priority. The screened individuals are used as elite offspring.
[0028] When selecting elite offspring, the principle of frontier level priority and same-layer sparse priority is maintained. Individuals are selected from the first layer of the non-dominated layer. When insufficient, they are supplemented by the second layer, and so on. When selection must be made from the same layer, subnets with high congestion are retained first. In order to prevent falling into the local optimum, if the frontier does not improve for three consecutive generations, the congestion weight is increased to force the search area to expand. At the same time, the IS mean and FID variance of each generation of frontiers are recorded. If the change rate of both is less than 1% for more than 5 generations, early stopping is triggered.
[0029] Further, before calculating the crowding degree, the IS and FID values are normalized to the interval [0, 1] to avoid the interference of dimensional differences on distance calculation. The target space is set as a two-dimensional grid, and the sum of the interval distances between adjacent solutions of each network structure in the two target dimensions of IS and FID is taken to calculate the crowding degree of network structure i. The formula is: , In the formula, represents the target value of the next adjacent individual of the sorted individual in the dimension, and represents the target value of the previous adjacent individual of the sorted individual in the dimension.
[0030] Aiming at the global exploration limitation and premature convergence problems of the genetic algorithm (GA) architecture, in this example, particle swarm optimization (PSO) is introduced for local optimization. Although GA can achieve wide-area search through crossover and mutation, its ability to finely adjust the elite architecture is insufficient and it is prone to falling into suboptimal solutions. While PSO, relying on the redirection movement mechanism of individual and group experience, can accurately adjust parameters within a small range and effectively break through the local optimum. This complementary optimization mechanism can especially solve the "small model trap", that is, the global search of GA tends to small models with low cost, and PSO enhances the competitiveness of complex high-performance architectures through neighborhood perturbation.
[0031] Step 4, use the particle swarm optimization algorithm to perform local parameter adjustment on the first n 0 elite offspring, where .
[0032] In one example, the first 10% of the elite offspring can be selected for local parameter adjustment.
[0033] Further, Step 4 includes: Step 401, take the first 10% of the elite offspring as the initial particle swarm, flatten the corresponding integer matrix into a 21-dimensional vector, with each dimension ranging from [0, 6], as the initial position of the particle, and set the velocity vector as a random floating-point number in the interval [-1, 1]; Step 402, calculate the fitness function, where FID and IS are fused with boundary constraints; the fitness function is: , In the formula, represents the boundary constraint penalty term, represents the weight coefficient used to adjust the penalty intensity, represents the dimensional parameter of particle k; The penalty for out-of-bounds particles linearly increases to the final value with the number of iterations, prompting the architecture to move towards the region that satisfies the fitness function. For the out-of-bounds particle positions a square penalty is imposed to force the particles to converge to legal discrete points. Corresponding to the sum of the 21-dimensional vectors after flattening the architecture matrix, where each dimension corresponds to the operation encoding of the network layer, such as the convolution type and the number of channels; Step 403, set the inertia weight to linearly decrease from to , and calculate the inertia weight at the current iteration. The formula is: , where t is the current iteration number, T max is the maximum number of iterations, is the initial value of the inertia weight, is the final value of the inertia weight; In one example , , the velocity of the particle swarm at the t-th iteration in dimension d, with the inertia weight linearly decreasing from 0.9 to 0.4 to balance the degree of dependence on the historical velocity; Step 404, update the velocity. The formula is: , where represents the velocity of the particle at the t-th iteration, is the learning factor to enhance the weight of the subnet historical experience to prevent premature convergence. r 1 and r 2 are random numbers in the range [0, 1] to introduce exploration randomness, is the best architecture parameter generated by the history of this particle, is the current global best generator architecture parameter, continues the original movement direction of the particle. w is the inertia weight, is the velocity vector of particle i at the t-th iteration, drives the particle towards its own optimal solution, guides the particle to aggregate towards the global optimal solution; Step 405, after updating the velocity, obtain the floating-point position increment. Round the floating-point result to an integer according to the integer encoding of the enumeration values of the architecture parameters such as convolution type and number of channels, and use the np.clip function to forcefully constrain the value range of the architecture parameters and limit them within the predefined integer range. The formula is: , where x takes the enumeration values of the convolution type and number of channels of the architecture parameters, respectively represent the minimum and maximum operation codes; Step 406, update the position, and the formula is: , and clip it to , respectively represent the minimum and maximum values of the particle position parameters, is the particle at the t-th iteration 's velocity vector; Step 407, stop the iteration when the maximum iteration number is reached or the solution converges, and obtain the optimal architecture solution; otherwise, return to Step 402.
[0034] During the particle swarm optimization process, calculate the sensitivity gradient of each dimension of the current optimal particle to FID every 5 generations , and correct the velocity according to the negative gradient. The calculation formula of the sensitivity gradient is: , in order to prevent falling into local optima, the sensitivity gradient is used to correct the velocity.
[0035] In the solution set after PSO optimization, Flatten the integer encoding matrix into a 21-dimensional vector, dynamically adjust the weight of the crossover region after normalizing the IS and FID values, and generate a Pareto optimal architecture set through a hybrid strategy to ensure multi-objective balance.
[0036] Step 5, determine whether the iteration termination condition is satisfied. If not, jump to Step 2; otherwise, terminate the search and output the optimal supernetwork architecture that meets the convergence condition.
[0037] Round the floating-point value corrected by the particle velocity and clip it to the integer encoding range to maintain the architecture validity; calculate the sensitivity gradient of the optimal particle to FID every 5 generations, and apply a velocity increment in the negative direction to break through the local optimum. Control the search through a composite termination condition: terminate when the maximum iteration is reached, the standard deviation of the fitness of consecutive generations is less than the threshold, or the time consumption exceeds 20% of the GA stage. Finally, output the convergent solution set through the elite retention strategy, and ensure that the solution set meets the video memory or latency deployment requirements through boundary clipping.
[0038] The hybrid strategy of the genetic algorithm and the particle swarm optimization algorithm adopts a hierarchical hybrid optimization mechanism to achieve the dynamic balance of global exploration and local optimization through the GA-PSO collaborative strategy in each generation of evolution. The core implementation plan includes a three-stage optimization chain: the GA global search stage uses the improved NSGA-II algorithm to perform multi-objective optimization on the integer matrix architecture encoding, and uses non-dominated sorting and crowding distance calculation to ensure the diversity of the solution set. The top 10% of the elite architectures selected will be input into the PSO for neighborhood refinement adjustment, and a two-layer guidance mechanism is constructed through dynamic inertia weight and learning factor. Gradient sensitivity analysis is introduced in the PSO stage, and the sensitivity gradient of the optimal particle to FID is calculated every 5 generations And apply velocity correction in the negative direction to effectively break through the local optimal trap. The discretization process ensures parameter compliance. After performing floating-point operations on the 21-dimensional vector parameters, round them off, and use the np.clip function to strictly limit the range to [0, 6], maintaining the physical validity of the architecture encoding.
[0039] Perform hypernetwork training, as Figure 3 shown in the left half. The hypernetwork generator and the discriminator with a fixed structure are trained adversarially. In each round of training, randomly activate a single path, and only optimize the weights of the currently activated path through gradient updates. Other unactivated paths do not participate in the optimization. During the training process, evaluate the contribution of each operation to the generation performance in real time, discard bad operations. If the FID value of using bilinear upsampling in a certain layer is continuously higher than that of nearest neighbor interpolation, remove this operation from the candidate set.
[0040] Perform architecture search: as Figure 3 shown in the right half. Based on the trained hypernetwork, search for the Pareto optimal architecture under multi-objective constraints through a hybrid optimization algorithm.
[0041] The termination conditions comprehensively consider the maximum number of iterations, the standard deviation of the particle swarm fitness, and the computing resource limit, and finally output the set of Pareto front optimal architectures after local reinforcement. This stage significantly improves the quality of a single solution through gradient-assisted search in the constrained space while maintaining the diversity of the solution set.
[0042] Evaluation and selection: Evaluate the optimized architectures according to IS and FID, and select the optimal architecture for training. In the system, the Pareto front selection achieves a balance between performance and efficiency through multi-objective collaborative optimization. First, the algorithm clearly defines three major optimization goals: generation quality, model efficiency, and training cost. Dynamic path selection and weight sharing cover diverse subnet structures in a single hypernetwork, supporting flexible switching between lightweight and multi-convolution stacking and channel expansion hybrid modes, providing hardware-friendly underlying module support for multi-objective Pareto front search. To generate the Pareto front, use the NSGA-II evolutionary algorithm: stratify the candidate architectures according to the dominance relationship through non-dominated sorting. Solutions with low FID and low number of parameters are in the upper layer, and non-dominated levels are preferentially retained. At the same time, introduce a crowding distance comparison mechanism to screen individuals with sparse distribution in the objective space, avoid the loss of population diversity caused by the aggregation of the solution set, and ensure that the front covers different trade-off directions. The dynamic maintenance strategy further optimizes the search process. The elitist retention mechanism directly inherits the previous generation's Pareto optimal solutions to prevent the loss of high-quality architectures. Constraint optimization sets hard thresholds for video memory occupancy or latency, and only allows network architectures that meet the constraint conditions to participate in crossover and mutation, thereby guiding the search for excellent network architectures in a targeted manner.
[0043] The performance verification experiments of the present invention were carried out on two benchmark datasets, CIFAR-10 and STL-10. Among them, the CIFAR-10 dataset consists of 60,000 natural images of ten categories with pixels. According to the standard division, 50,000 images are used as the training set and 10,000 images are used as the test set; the STL-10 dataset contains 13,000 images of ten categories with the original resolution of pixels, which are uniformly scaled to Figure 5 pixels in the experiment to balance the computational efficiency and model performance. Through the proposed automated architecture search method, the system successfully obtained a network structure with excellent generation ability, and used known images and this network structure to generate new images. Figure 6 Specifically shows the performance evolution curve of this architecture during the training process on the CIFAR-10 dataset. The generated images are significantly superior to traditional methods in terms of visual fidelity and detail richness, and can synthesize high-quality samples with natural texture features and semantic consistency.
[0044] The performance evaluation of the present invention is based on two core quantitative indicators in the generative adversarial network: IS and FID. As Figure 5 shown, IS extracts the feature distribution of the generated images through a pre-trained Inception-v3 network, and calculates its inter-class separability and diversity. The larger the value, the better the generated samples perform in terms of visual quality and category diversity. Ideally, high definition and cross-category semantic distinguishability should be achieved simultaneously. Figure 6 The FID in
[0045] compares the statistical distribution differences between the generated images and the real images in the feature space. The smaller the value, the higher the degree of alignment of their deep features, and the more the generated results approximate the real data at the probability distribution level. Experiments show that the architecture search method proposed by the present invention can complete efficient search within several hours through dynamic gradient guidance and multi-objective collaborative optimization, and the obtained architecture shows fast convergence characteristics during independent re-training.
[0046] Example 2 The generative adversarial network architecture search system based on the GA-PSO hybrid algorithm described in this embodiment includes: A construction module for constructing a generative adversarial network supernet, with cell-level support configuration Convolution, Atrous convolution, skip connection candidate operations. The connection topology between nodes allows a maximum of 3 input edges and supports dynamic selection of normalization layers and ReLU / Swish activation functions; Training module, only activating one bilinear upsampling or depthwise separable convolution operation path each time, and performing adversarial training with a discriminator of a fixed structure. An unbiased uniform sampling strategy is adopted, and each path needs to be fairly iterated in the same adversarial environment; Optimization module, generating an initial population based on the NSGA-II genetic algorithm and exploring a wide solution space through crossover and mutation; Particle swarm optimization is used to optimize the GA results, adopting dynamic inertia weights and gradient sensitivity analysis to break through local optima; The mixing ratio is set to 70% GA search and 30% PSO optimization, and global PSO reinforcement is performed every 5 generations. The architecture parameters are ensured to meet the preset constraints by discretizing the position update.
Claims
1. A generative adversarial network architecture search method based on a GA-PSO hybrid algorithm, characterized in that: The steps include: Step 1: construct a supernet architecture of a generative adversarial network, including a fully connected layer, an upsampling layer, a downsampling layer, and multiple convolutional modules. Multiple convolutional modules constitute a search space and define candidate operations in the search space. Step 2, using the supernet architecture as an individual in the population, and using the NSGA-II algorithm to perform a global search to obtain the non-dominated layer; Step 3, stratify and screen the population through non-dominated sorting and crowding distance to obtain elite offspring; Step 4: Use the particle swarm optimization algorithm to adjust the local parameters of the first n0 elite offspring, where ; Step 5: determine whether the iteration termination condition is met. If not, jump to step 2. Otherwise, the search is terminated and the optimal supernet architecture that meets the convergence conditions is output.
2. The method for searching a generative adversarial network architecture based on a GA-PSO hybrid algorithm according to claim 1, characterized in that: Step 2 includes: Step 201, randomly generate an initial population P0 consisting of a supernet structure, and the individuals corresponding to each supernet structure are encoded as an integer matrix; Step 202, each supernet architecture calculates dual-objective indicators: IS and FID through the pre-trained Inception v3 network, where IS represents maximizing image quality and FID represents minimizing the gap with the true distribution; Step 203, calculate the objective function value of each supernet architecture, and the calculation formula is: , , , In the formula, represents the supernet architecture parameters, represents the optimal architecture parameters of the supernet, is a multi-objective optimization function. It means minimizing the distribution difference between generated samples and real samples. Represents maximizing the diversity and quality of generated samples, are the weights of the supernet generator, is the weight of the optimal supernet generator, represents the expected calculation of the random noise variable z, and calculates the statistical expectation on the input noise distribution of the generator G(z), that is, the expected performance measure of the generator to generate fake samples through the noise vector, is the adversarial loss of the generator; represents the weight of the supernet discriminator, is the weight of the optimal supernet discriminator, Represents the real data sample x Calculate the expected response of the discriminator D(x) to the sample under the real data distribution. is the discrimination loss of the discriminator D on the real data. The discriminator D needs to maximize the real sample x The probability D(x) is judged to be true, and the generator tries to generate samples The probability of being judged as true by the discriminator D Approaching 0; Step 204, select the solution through the elite retention strategy, select the first n1 individuals according to the objective function value and directly enter the next generation, recombining the crossover offspring and the mutated individuals to generate a new population, where ; Step 205 , dividing the new population into different non-dominated layers by non-dominated sorting.
3. The method for generating adversarial network architecture search based on GA-PSO hybrid algorithm according to claim 2 is characterized in that: In step 205, dividing the population into different non-dominated layers by non-dominated sorting includes: If the IS of individual A ≥ the IS of individual B, and the FID of individual A ≤ the FID of individual B, then individual A dominates individual B. Individuals that are not dominated by any individual belong to the first layer. The dominated individuals are removed layer by layer until all individuals are stratified and different non-dominated layers are obtained.
4. The method for generating adversarial network architecture search based on GA-PSO hybrid algorithm according to claim 3 is characterized in that: Step 3 includes: According to the preset number of screening, individuals are selected from the first non-dominated layer. When the number of individuals in the first non-dominated layer is insufficient, they are supplemented by the next non-dominated layer, and so on. When individuals must be selected from the same non-dominated layer, individuals with high crowding degree are given priority. The screened individuals are used as elite offspring.
5. The method for searching a generative adversarial network architecture based on a GA-PSO hybrid algorithm according to claim 4 is characterized in that: Before calculating the crowding, the IS and FID values are standardized to the [0,1] interval, the target space is set to a two-dimensional grid, and the sum of the interval distances of adjacent solutions of each network structure in the two target dimensions of IS and FID is taken to calculate the crowding of network structure i. The formula is: , In the formula, Indicates that the target The target value of the next adjacent individual of individual i after sorting in dimension, Represents the target value of the previous adjacent individual of the sorted individual i in the target k dimension.
6. The method for searching a generative adversarial network architecture based on a GA-PSO hybrid algorithm according to claim 5, characterized in that: Step 4 includes: Step 401, take the first n0 elite offspring as the initial particle group, flatten the corresponding integer matrix into a vector, each dimension range [0,6], as the initial position of the particle, and set the velocity vector to a random floating point number in the interval [-1,1]; Step 402, calculate the fitness function, which integrates FID and IS with boundary constraints; the fitness function is: , In the formula, represents the boundary constraint penalty term, Represents the weight coefficient, which is used to adjust the penalty intensity. represents the dimension parameter of particle k; Step 403: Set the inertia weight from Linearly decreasing to , calculate the inertia weight under the current number of iterations, the formula is: , Where t is the current iteration number, T max is the maximum number of iterations, is the initial value of inertia weight, is the final value of inertia weight; Step 404, update the speed, the formula is: , In the formula, is the learning factor, r1 and r2 are random numbers in [0,1], The best architecture parameters generated for this particle history, is the current global optimal generator architecture parameter, Continue the original moving direction of the particle, w is the inertia weight, It is The velocity vector of particle i at the iteration, It is The position vector of particle i at the iteration, Drive particles closer to their own optimal solution, Guide particles to gather toward the optimal solution of the group; Step 405, after the speed is updated, obtain the floating point position Increment, according to the enumeration value of the architecture parameter convolution type and the number of channels, the floating point result is rounded to the integer encoding, and the value range of the architecture parameter is constrained by the np.clip function to limit it to the predefined integer range. The formula is: , In the formula, Take the convolution type and number of channels of the architecture parameters. Respectively represent the minimum and maximum operation codes; Step 406, for location update, the formula is: , and crop to ; Respectively represent the minimum and maximum values of the particle position parameters, is the particle at the tth iteration The velocity vector of Step 407, when the maximum number of iterations is reached or the solution converges, the iteration is stopped to obtain the optimal architecture solution; otherwise, the process returns to step 402.
7. A generative adversarial network architecture search system based on a GA-PSO hybrid algorithm, characterized in that: include: Building blocks for constructing generative adversarial network supernets, with cell-level support for configuration convolution, Atrous convolution and skip connection candidate operations, the node connection topology allows up to 3 input edges, and supports dynamic selection of normalization layers and activation functions; Training module, only one bilinear upsampling or The depth-separable convolution operation path is trained adversarially with the fixed-structure discriminator, and a non-repeated uniform sampling strategy is adopted. Each path needs to be iterated fairly in the same adversarial environment. Optimization module, which generates an initial population based on the genetic algorithm and explores the wide-area solution space through crossover and mutation; particle swarm optimization is used to optimize the results of the genetic algorithm, adopting dynamic inertia weight and gradient sensitivity analysis to break through the local optimum; the mixing ratio is set as n3 GA search and (1 - n3) particle swarm optimization, where 65% < n3 < 75%, and global particle swarm optimization reinforcement is performed every preset number of generations to ensure that the architecture parameters meet the preset constraints through discretized position update.
Citation Information
Patent Citations
Multi-objective deep convolution generative adversarial network model and learning method thereof
CN108171266A
Hybrid particle swarm optimization algorithm in combination with genetic algorithm
CN108399451A
Industrial data quality prediction method based on improved PSO-GA and SVM
CN115271237A
Neural architecture searching method based on diffusion evolutionary algorithm
CN119886226A
Cloud task scheduling method based on phagocytosis-based hybrid particle swarm optimization and genetic algorithm
US20210133534A1
Cited By
Generative adversarial network architecture search method based on multi-population coevolution
CN120671780A
Black box confrontation audio generation method and system
CN121393470A
A black-box adversarial audio generation method and system
CN121393470B
Generative adversarial network architecture search method for text-guided image generation task
CN122133725A