A Neural Architecture Search Method Based on Diffusion Evolution Algorithm

By introducing diffusion evolution algorithms into neural network architecture search, building a hypernet and alternately optimizing network weights and subnet coding, the problem of insufficient population diversity and easy to fall into local optimality in the existing methods is solved, and a more comprehensive and efficient neural network architecture search is achieved.

CN119886226BActive Publication Date: 2025-05-30NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510370248.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-05-30
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The existing neural network architecture search methods have problems such as limited population diversity, low architectural performance, and easy to fall into local optimality.

Method used

Using a neural architecture search method based on diffusion evolution algorithm, the neural network architecture with the best performance is found by building a hypernet and alternately optimizing network weights and subnet coding, using the denoising mechanism of the diffusion model and the global search capability of the evolution algorithm.

Benefits of technology

It effectively enhances population diversity, explores more potential excellent architectures, avoids local optimal solutions, improves the comprehensiveness and effectiveness of searches, and increases the probability of finding the optimal neural network architecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886226B_ABST
    Figure CN119886226B_ABST
Patent Text Reader

Abstract

The present invention discloses a neural architecture search method based on a diffusion evolution algorithm. The method first constructs a supernetwork, which includes subnets, and each subnet assigns weights to each operation on each connection edge through continuous coding; then obtains pictures of different categories to construct a data set; finally, alternately optimizes the supernetwork weights and subnet encodings until the optimization stops when the constraint conditions are met, obtaining the optimal subnet encoding, and further obtaining the architecture of the neural network. The present invention innovatively combines the iterative denoising mechanism of the diffusion model with the global search strategy of the evolutionary algorithm. By combining the denoising mechanism of the diffusion model with the global search ability of the evolutionary algorithm, it constructs a supernetwork and alternately optimizes the network weights and subnet encodings; uses adaptive noise scheduling and density estimation to enhance population diversity and avoid local optima, and can better find the optimal neural network architecture suitable for the task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to deep neural network technology, and particularly to a neural architecture search method based on a diffusion evolution algorithm. Background Art

[0002] Deep learning models, especially convolutional neural networks, have achieved remarkable results in computer vision tasks such as image classification and object detection. However, manually designing an effective network architecture requires a large amount of professional knowledge and repeated trials. To solve this problem, neural network architecture search emerged, aiming to automatically discover the optimal network structure and reduce the dependence on manual design.

[0003] Existing neural network architecture search methods are mainly divided into three categories: reinforcement learning-based, gradient-based, and evolutionary algorithm-based. Among them, reinforcement learning-based methods can obtain good search results, but have high computational resource requirements; gradient-based methods such as differentiable architecture search can quickly find potential architectures, but are prone to obtaining sub-optimal architectures with redundancy and many skip connections and are easily trapped in local optima; evolutionary algorithm-based methods have advantages in exploring the search space, but consume too much computational resources when evaluating candidate networks.

[0004] In recent years, researchers have also proposed various improvement methods, but these methods still have problems such as limited population diversity and low performance of the searched architectures in practical applications. At the same time, although the differentiable architecture search framework improves the search efficiency, the gradient descent method is prone to falling into local optimal solutions, and the simultaneous optimization of the network architecture weights of all candidate operations will cause interference between operations. Therefore, a new search method is needed to solve the defects existing in the current neural network architecture search process. Summary of the Invention

[0005] Object of the Invention: Aiming at the above problems, the object of the present invention is to provide a neural architecture search method based on a diffusion evolution algorithm. By combining the denoising mechanism of the diffusion model with the global search ability of the evolutionary algorithm, a super network is constructed and the network weights and subnet encodings are alternately optimized to find the neural network architecture with the optimal performance.

[0006] Technical Solution: A neural architecture search method based on a diffusion evolution algorithm of the present invention includes the following steps:

[0007] Step 1, construct a super network, and the super network includes subnets, and each subnet assigns weights to each operation on each connection edge through continuous encoding, and the encoding form is matrix, where M represents the number of connection edges in the super network, and N represents N candidate operations set on each connection edge;

[0008] Step 2: Obtain images of different categories to construct a dataset, and divide it into a training dataset and a validation dataset according to a ratio;

[0009] Step 3: Alternately optimize the hypernetwork weights and subnet encodings until the optimization stops when the constraint conditions are met, obtaining the optimal subnet encoding, and then obtaining the neural network architecture; among them, use the training dataset and its labels to train and optimize the hypernetwork weights, and use the differential evolution algorithm to optimize the subnet encoding.

[0010] Furthermore, optimizing the subnet encoding using the differential evolution algorithm includes the following steps:

[0011] Step 301: Sample and initialize the population individuals from the normal distribution. The population is represented as , where represents the identity matrix of dimension represents the encoding dimension, T represents the total number of iterations, represents an individual in the population. Each individual includes a subnet encoding and a scaling factor. The encoding dimension of each individual satisfies: , t represents the current iteration number, represents the set of real numbers. The encoding method of the i-th individual at the t-th time step is represented as , Q represents the transpose, represents the encoding method of the M-th connection edge of the i-th individual at the t-th time step, represents the value corresponding to the N-th candidate operation at the N-th position in the encoding information of the i-th individual at the t-th time step, represents the lower limit of the value corresponding to the N-th candidate operation in the encoding of the i-th individual at the t-th time step, represents the upper limit of the value corresponding to the N-th candidate operation in the encoding of the i-th individual at the t-th time step; represents the scaling factor of the i-th individual at the t-th time step, is a random number between 0 and 1;

[0012] Step 302: Update the population individuals using the reverse denoising process. The update formula is:

[0013] ,

[0014] In the formula, represents the cosine schedule, represents the noise variance, represents the noise injection;

[0015] Step 303: Use the validation dataset to calculate the fitness value of each individual. The calculation formula is:

[0016] ,

[0017] In the formula, B represents the number of batches in the validation dataset, represents the true label of the b-th sample corresponding to the c-th category, represents that the model is based on the parameter for the b-th sample in the validation dataset predicted as the probability of category c, and C represents the total number of categories;

[0018] Step 304: Compare the fitness values of the i-th individual at the t-th time step and the updated individual, and save the higher one into the next-generation population, which is expressed as:

[0019] ,

[0020] In the formula, represents the individual saved into the next generation, including the subnet encoding and the scaling factor, represents the i-th individual at the t-th time step, represents the individual corresponding to the currently traversed individual obtained by processing with the diffusion evolution operator, represents the individual the subnet encoding within the fitness value of, represents the individual the subnet encoding within the fitness;

[0021] Step 305: Convert the fitness value of the individual into a selection probability through the probability representation function, and use the selection probability to guide the update of the population individuals;

[0022] Step 306: Update the current iteration time step , and jump to Step 302 until the maximum number of iterations is reached and the iteration stops.

[0023] Furthermore, the assignment rule of noise injection is:

[0024] ,

[0025] In the formula, represents the population mean.

[0026] Furthermore, the setting rule of the noise injection ratio is:

[0027] In the early stage, that is, when, set the noise injection ratio to be greater than or equal to ;

[0028] In the later stage, that is, when, set the noise injection ratio to be less than or equal to .

[0029] Further, the probability representation function types include uniform mapping, kernel density estimation, or Gaussian mixture model;

[0030] Among them, the probability representation function in the case of uniform mapping is expressed as:

[0031] ,

[0032] The probability representation function in the case of kernel density estimation is expressed as:

[0033] ,

[0034] In the formula, h represents the bandwidth, represents the j-th individual in the population, represents the square of the Euclidean norm;

[0035] The probability representation function in the case of Gaussian mixture model is expressed as:

[0036] ,

[0037] In the formula, represents the weight of the k-th Gaussian component in the mixture model, represents the covariance matrix, and the parameter , represents the set of individual indices belonging to the k-th group, is the number of elements in this set.

[0038] Further, the constraint condition in step 3 is expressed as:

[0039] obeys the constraint ,

[0040] In the formula, represents the subnet encoding, represents the subnet encoding corresponding to the individual with the highest fitness, represents the supernet weight, represents the optimal weight obtained after the supernet is trained under the given weight coefficient , represents the validation dataset, represents the value of the weight W when the loss function value is the lowest, represents minimizing the cross-entropy loss function in the classification task, represents the training dataset.

[0041] Further, the encoding matrix is denoted as , where represents the weight coefficient of the candidate operation n on the connection edge m, and the formula is:

[0042] ,

[0043] In the formula, represents the number of the connection edge, represents the number of the candidate operation, is the basic variable for calculating the weight coefficient of the candidate operation n on the connection edge m. It is a part of the individual encoding, and the individual encoding is used to describe the architecture information of the subnet;

[0044] After the optimization is completed, let , and perform discretization processing on the obtained subnet encoding, that is, only keep the one with the largest encoding value obtained after optimization on each connection edge. The discretized encoding method is converted into 0-1 encoding. The encoding corresponding to the operation retained on each connection edge is 1, and the rest are 0. The formula is expressed as:

[0045] .

[0046] Beneficial effects: Compared with the prior art, the significant advantages of the present invention are:

[0047] 1. The present invention introduces a diffusion evolution algorithm. By analogy, the evolution process is regarded as a denoising process, and the individuals are continuously optimized in the iteration, effectively enhancing the population diversity, enabling more potential excellent architectures to be explored when searching for the neural network architecture, avoiding being trapped in local optimal solutions due to insufficient population diversity, and improving the comprehensiveness and effectiveness of the search;

[0048] 2. The present invention utilizes the iterative denoising mechanism of the diffusion model to be able to explore the parameter space more fully;

[0049] 3. By adjusting operations such as noise variance and injecting noise, the present invention enables a wide or fine search of the search space at different stages during the process of optimizing the subnet encoding, improving the probability of finding the optimal neural network architecture, making the search results more reliable and practical; the reverse denoising process and the adaptive noise injection mechanism in it enable the individuals to evolve towards the direction of better solutions during the evolution process and can jump out of the local optimal region when necessary. This mechanism effectively reduces the risk of the algorithm being trapped in local optimal solutions, ensuring that the search process can continuously approach the global optimal solution and improving the search efficiency and the performance of the finally obtained neural network architecture;

[0050] 4. Based on the characteristics of the diffusion evolution algorithm, a fitness function is constructed in the present invention, and the influence of different probability distribution representations on the search performance is deeply studied, providing a theoretical basis and practical guidance for optimizing the algorithm performance, enabling the algorithm to be adjusted according to different task requirements and data characteristics, and further improving the search effect. Description of the Drawings

[0051] Figure 1 It is a flowchart of a neural architecture search method based on a diffusion evolution algorithm;

[0052] Figure 2 It is a flowchart for optimizing subnet encoding using a diffusion evolution algorithm;

[0053] Figure 3 It is a schematic diagram of the structure of the supernet in the example;

[0054] Figure 4 It is an example diagram of subnet encoding during the optimization process;

[0055] Figure 5 It is a schematic diagram of the accuracy of the supernet obtained through training;

[0056] Figure 6 It is a schematic diagram for comparing the method of the present invention with other methods. Detailed implementation manners

[0057] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0058] A neural architecture search method based on a diffusion evolution algorithm described in this embodiment has a flowchart as Figure 1 shown, and this method includes the following steps:

[0059] Step 1, construct a supernet, which includes subnets, and each subnet assigns weights to each operation on each connection edge through continuous encoding, and the encoding form is matrix, where M represents the number of connection edges in the supernet, and N represents N candidate operations set on each connection edge;

[0060] Step 2, obtain pictures of different categories to construct a data set, and divide it into a training data set and a validation data set according to a ratio;

[0061] Step 3, alternately optimize the supernet weights and subnet encoding, and stop the optimization until the constraint conditions are met to obtain the optimal subnet encoding, and then obtain the neural network architecture; among them, the supernet weights are trained and optimized using the training data set and its labels, and the subnet encoding is optimized using the diffusion evolution algorithm.

[0062] A supernet is a hybrid network that integrates all candidate operations and topological connection methods within the search space; a subnet is a component of the supernet, and its feature is that there is only one candidate operation on all connection edges. In this example, a supernet is first constructed, which contains M connection edges, and N candidate operations are set on each connection edge, such as dynamic separable convolution, dilated convolution, Max pooling, skip connections, and null operations are all candidate operations. There are subnets within the supernet. The input information no longer undergoes a single operation in the supernet but will go through multiple operations. For example, the input information will simultaneously go through dynamic separable convolution, dilated convolution and other operations, and then add the feature maps obtained by these operations element by element to generate new hybrid feature map information. When training the network weights using the gradient descent method, the supernet is crucial as it enables the training of the parameters of all candidate operations within the search space, aiming to provide trained network weights for subsequent evaluation of the subnet performance.

[0063] To represent and optimize the subnets, in this example, a matrix encoding form is adopted. Continuous values within the range of 0 to 1 are used to assign weights to each operation on each connection edge. Through this encoding method, the subnets can be discretized into conventional neural networks according to the weight magnitudes, specifically by removing candidate operations and connection edges with smaller weights. During the evolutionary computing process, the subnets are regarded as individuals within the population. When the encoding presents different values, multiple different neural networks can be obtained. The value of N theoretically has no limit, and any operations such as convolution and pooling can be included. The value of M is related to the internal nodes set in the supernet. If the number of internal nodes in the supernet is H, then . In this example, the diffusion evolution algorithm is used to optimize the encoding of the subnets, aiming to accurately find the subnet architecture with the best performance during the search process.

[0064] Specifically, the supernet adopts a cell-based search space, consisting of normal cells and downsampling cells. The normal cells maintain the feature map size unchanged, and the downsampling cells can use convolution operations with a stride of 2 to halve the output feature map size. Each cell contains intermediate nodes, then the number of connection edges , the encoding dimension . This supernet structure design provides a basis for the diversification of neural network architectures. Through the combination of different cells and the selection of connection edges and operations, a rich variety of subnet architectures can be generated. Each subnet represents the architecture weights through continuous encoding, and the encoding form is denoted as matrix , where represents the weight coefficient of candidate operation n on connection edge m, and the formula is:

[0065] ,

[0066] In the formula, represents the number of the connection edge, represents the number of the candidate operation, The basic variables for calculating the weight coefficient of candidate operation n on connection edge m.

[0067] The set of candidate operations includes, but is not limited to, convolution operations, such as depthwise separable convolution, depthwise separable convolution, dilated convolution, dilated convolution, pooling operations such as max pooling, average pooling, as well as skip connections and no-op operations. These encoding methods can accurately represent different subnet architectures and can be discretized according to the weight magnitudes during subsequent optimization to determine the final neural network architecture.

[0068] In step 2, images of different categories are obtained to construct a data set, which is divided into a training data set and a validation data set according to a ratio such as 5:5. Taking medical image classification as an example, the images in the training data are labeled with different disease categories, such as normal, diseased A, diseased B, etc. The supernetwork needs to be trained first before the weights of its internal operations can be used. In this example, the Adam optimizer is used, and the training data set is used to train the network weights W of the supernetwork. The initial learning rate of the Adam optimizer is set to 0.025, and the weight decay is set to The network weights W are continuously adjusted to gradually reduce the difference between the prediction results of the Adam optimizer for the training data and the true labels, achieving the optimization of the supernetwork performance. The commonly used cross-entropy loss function in classification tasks is used, and its expression is:

[0069] ,

[0070] where, represents the predicted probability of the Adam optimizer for class c, C represents the total number of classes, represents the target output or prediction target in the neural network task, represents the input sample x and its corresponding true label y.

[0071] Furthermore, as Figure 2 shown, optimizing the subnet encoding using the differential evolution algorithm includes the following steps:

[0072] Step 301, sample and initialize the population individuals from a normal distribution. The population is represented as , where, is the dimensional identity matrix, represents the encoding dimension, T represents the total number of iterations, represents the individuals in the population. Each individual includes a subnet encoding and a scaling factor, and the encoding dimension of each individual satisfies: , where \(t\) represents the current iteration number, represents the set of real numbers, and the coding method of the \(i\)-th individual at the \(t\)-th time step is expressed as , \(Q\) represents the transpose, represents the coding method of the \(M\)-th connection edge of the \(i\)-th individual at the \(t\)-th time step, represents the value corresponding to the \(N\)-th candidate operation at the \(N\)-th position in the coding information of the \(i\)-th individual at the \(t\)-th time step, represents the lower limit of the value corresponding to the \(N\)-th candidate operation in the coding of the \(i\)-th individual at the \(t\)-th time step, represents the upper limit of the value corresponding to the \(N\)-th candidate operation in the coding of the \(i\)-th individual at the \(t\)-th time step, represents the scaling factor of the \(i\)-th individual at the \(t\)-th time step, is a random number between 0 and 1;

[0073] Step 302, update the population individuals using the reverse denoising process, and the update formula is:

[0074] ,

[0075] In the formula, represents the cosine schedule, represents the noise variance, represents the noise injection;

[0076] Step 303, calculate the fitness value of each individual using the validation dataset, and the calculation formula is:

[0077] ,

[0078] In the formula, \(B\) represents the number of batches in the validation dataset, represents the true label of the \(b\)-th sample corresponding to the \(c\)-th category, represents based on the parameter for the \(b\)-th sample in the validation dataset the probability of predicting as the \(c\)-th category, and \(C\) represents the total number of categories;

[0079] Step 304, compare the fitness values of the \(i\)-th individual at the \(t\)-th time step and the updated individual, and save the higher one into the next-generation population, which is expressed as:

[0080] ,

[0081] In the formula, represents the individual saved into the next generation, including the subnet coding and the scaling factor, represents the \(i\)-th individual in the \(t\)-th time step, represents the individual being traversed currently The individuals obtained by processing with the corresponding diffusion evolution operator represent the individual subnet encoding within the fitness value, represent the individual subnet encoding within the fitness;

[0082] Step 305, convert the fitness value of the individual into a selection probability through a probability representation function, and use the selection probability to guide the update of the population individuals;

[0083] Step 306, update the current iteration time step , jump to Step 302, and stop the iteration until the maximum number of iterations is reached.

[0084] Furthermore, the assignment rule of noise injection is:

[0085] ,

[0086] wherein, represents the population mean.

[0087] In the above assignment rule, the intensity of noise injection is dynamically adjusted according to the dispersion degree of the population individuals, and the dispersion degree can be judged by calculating the distance between the individual and the population mean. is a parameter that controls the overall intensity of the noise, and e is a random noise sampled from the standard normal distribution. When the population individuals are relatively concentrated, that is , in order to increase the population diversity and let the algorithm jump out of the local optimum, the random noise will be adjusted so that it does not exceed ; when the individual distribution is relatively dispersed, the random noise is saved as , and the noise injection is further dynamically adjusted according to the distribution of the population individuals, regulating the noise from the perspective of the macroscopic time step, and e combines the individual state of the population for more delicate noise adjustment. The two complement each other and jointly serve the optimization process of the algorithm.

[0088] Furthermore, the setting rule of the noise injection ratio is:

[0089] In the early stage, that is , set the noise injection ratio to be greater than or equal to ;

[0090] In the later stage, that is , set the noise injection ratio to be less than or equal to .

[0091] In this example, judge the stage where the current time step t is located. If , the noise variance can be calculated , at this time, set the noise injection ratio to be greater than or equal to , specifically, in the early stage, if the noise variance , then through to ensure , so the noise injection ratio is greater than or equal to , enabling the algorithm to more efficiently search for potential solutions during the large - space search phase; if set , the noise injection ratio is less than or equal to , specifically, through to make , the noise injection ratio is less than or equal to .

[0092] After updating the population individuals, it is necessary to calculate the fitness value of each individual and convert the selection probability, and use the validation dataset to evaluate the performance of each individual, that is, the candidate architecture. The model processes the input data of the current batch to obtain the prediction result, calculates the difference between the prediction result and the true label through the loss function, and takes the negative value of this difference as the fitness value of the current batch of individuals. After completing all batch calculations, if a valid fitness value is obtained, the average of the fitness values of multiple batches is calculated to obtain the final fitness value of each individual; if a valid fitness value is not obtained, an empty tensor is returned. In practical applications, such as in image classification tasks, the model classifies and predicts the images in the validation set, calculates the difference between the prediction result and the true label through the cross - entropy loss function, and uses its negative value as the basis for calculating the fitness value. Generally, the fitness function is defined as the negative validation loss, and the formula is , where, represents the fitness value of the i - th individual, B is the number of batches in the validation set, L is the loss function, and are the input data and label of the b - th batch respectively, are the parameters of the model in the b - th batch. Compare the fitness values of the i - th individual at the t - th time step and the mutated individual, and save the higher one into the next - generation population.

[0093] The individual fitness value can, according to different application scenarios and requirements, convert the individual fitness value into a selection probability through different selection probability representation functions, so as to conform to the data type of the diffusion evolution algorithm and the principle of individual update of the diffusion evolution algorithm, and adopt the direct mapping method as the fitness mapping function, directly using the value calculated by the individual fitness value through different selection probability representation functions as its selection probability, and further guiding the optimization process of the population encoding through the probability reasoning unique to the diffusion evolution algorithm, determining which individuals are more likely to participate in the next - generation evolution to facilitate individual update through the diffusion evolution algorithm.

[0094] Furthermore, the probability representation function types include uniform mapping, kernel density estimation, or Gaussian mixture model;

[0095] Among them, the probability representation function in the case of uniform mapping is expressed as:

[0096] ,

[0097] The probability representation function in the case of kernel density estimation is expressed as:

[0098] ,

[0099] In the formula, h represents the bandwidth, represents the square of the Euclidean norm;

[0100] The probability representation function in the case of Gaussian mixture model is expressed as:

[0101] ,

[0102] In the formula, represents the weight of the k-th Gaussian component in the mixture model, represents the covariance matrix, and the parameter , represents the set of individual indices belonging to the K-th group, is the number of elements in this set.

[0103] When the probability representation function type is uniform mapping, in the process of updating individuals, it is assumed that the probability of an individual being selected in the population follows a uniform distribution, that is, the probability of each individual being selected to participate in the next-generation evolution is . When selecting individuals for update, individuals are randomly selected from the current population according to this probability and then participate in the individual update process.

[0104] When the probability representation function type is kernel density estimation, the selection probability function , in the formula, the bandwidth , and the dimension . When updating the individual , this probability representation function is used to measure the similarity between the individual and other individuals , and then affects the probability of the individual being selected to participate in the update. For example, when calculating the individual , the selection probability of is determined by the selection probability function , so that individuals with a higher similarity to have a greater probability of participating in the calculation of

[0105] ,

[0106] In the formula, is an individual-related noise term selected based on probability, which is related to e sampled from a normal distribution, but the selection process is regulated by If calculates that an individual is highly similar to the current individual then the noise term of the corresponding individual is more likely to be selected into , carries noise information related to the selected individual and affects the direction and degree of individual update.

[0107] When the probability representation function type is a Gaussian mixture model, its selection probability function is:

[0108] ,

[0109] In the formula, the number of components , the covariance matrix is a diagonal matrix. When updating the individual at is used to calculate the probability that the individual belongs to different Gaussian distributions, so as to determine the probability that the individual is selected to participate in the update. For example, when calculating , according to probability, different Gaussian components are selected, and the corresponding parameters , will affect the selection of the calculated individual and the calculation of the noise term, where , the diagonal elements , represents the set of individual indices belonging to the Kth group, is the number of elements in this set. At this time, the individual update formula can be expressed as:

[0110] ,

[0111] Among them, is a noise term related to the selection based on probability, which is related to e sampled from a normal distribution, but the selection process is regulated by If calculates that the individual is more likely to belong to a certain Gaussian component, then the noise term corresponding to this component will be selected into , incorporates the considerations of the Gaussian mixture model for the individual state and probability distribution, and plays an adjustment role in individual update.

[0112] In this example, an alternating optimization method for network weights and subnet encoding is adopted until the error rate of the network architecture within the population converges. For example, the hypernetwork weight update period can be set to perform a complete training every 5 generations of evolution. At each update, the training dataset is used to train the hypernetwork weights, update the network weights W, and provide a more accurate evaluation basis for the evolution of subnet encoding. The subnet encoding evolution period is set to perform 100 reverse denoising iterations per generation. By continuously updating the subnet encoding, a better neural network architecture is explored. The training dataset is evenly divided into two parts, denoted as and , represents the validation dataset, represents the training dataset, where represents the data used to train the hypernetwork weights, represents the data for evaluating the subnet performance during the optimization of subnet encoding. At the same time, an early stopping mechanism is set. When the decrease in the validation error rate on the validation set is less than for 10 consecutive generations, the optimization process is terminated.

[0113] Furthermore, the constraint condition in step 3 is expressed as:

[0114] obeys the constraint ,

[0115] In the formula, represents the subnet encoding, represents the subnet encoding corresponding to the individual with the highest fitness, represents the hypernetwork weights, represents the optimal weights obtained after training the hypernetwork given the weight coefficient , represents the validation dataset, represents the value of the weight W when the loss function value is the lowest, represents minimizing the cross-entropy loss function in the classification task, represents the training dataset.

[0116] After the optimization is completed, let , and perform discretization processing on the obtained subnet encoding, that is, only the largest encoded value obtained after optimization is retained on each connection edge. The discretized encoding method is converted to 0-1 encoding. The encoding corresponding to the retained operation on each connection edge is 1, and the rest is 0. The formula is expressed as:

[0117] .

[0118] The superiority of the neural architecture search method based on the diffusion evolution algorithm described in the present invention is further illustrated by the following example.

[0119] Construct a supernet as Figure 3 shown. This supernet contains 4 nodes representing feature maps, and there is an order relationship between the nodes. Feature maps can only be passed from a pre-order node to a post-order node after being processed by the pre-order node. X 1 , X 2 , X 3 , X 4 represent different feature maps respectively. The 3 connection edges of the supernet represent different candidate operation processes. For example, represents convolution, represents convolution, represents skip connection, Figure 4 is an example diagram of subnet encoding during the optimization process. In the upper half of Figure 4 there are multiple two-dimensional tables, and each table represents the encoding situation of an individual subnet. The horizontal table headers of the tables are candidate operations , , , indicating that these three alternative operations exist on the connection edges of the supernet. Figure 4 The lower half of is a directed graph corresponding to the upper half table. The directed edges between the nodes in the graph represent connection relationships, and edges of different colors such as blue, orange, and green represent different connection paths. These paths are actually determined by the encoding weights in the upper table. Each directed graph represents a subnet architecture, showing the flow direction of data or information in the subnet. Taking the input data on node 0 in the supernet as an example, the feature maps obtained by processing the input information through the three candidate operations are weighted and summed according to the encoding weights, and the obtained feature map information is stored on node 1. For subsequent nodes such as node 2 and node 3, they obtain feature map information from multiple pre-order nodes. At this time, multiple feature map information needs to be concatenated in the channel dimension. In practical applications, the supernet adopts a cell-based search space, which is composed of normal cells and downsampling cells. Each cell contains connection edges, and the number of candidate operations on each connection edge in the supernet is 8, and the encoding dimension is . By alternately optimizing the network weights and subnet encoding through the method described in the present invention, the operation with the largest weight on each connection edge is retained, and the final neural network architecture suitable for a specific task is obtained. The final optimal subnet is specifically represented as:

[0120] , where represents a dynamically separable convolution with a convolution kernel size of , represents a dynamically separable convolution with a convolution kernel size of , Indicates dilated convolution with a convolution kernel size of , Indicates dilated convolution with a convolution kernel size of , Indicates max pooling with a size of , Indicates skip connection, Indicates average pooling with a size of .

[0121] The performance of the neural network architecture optimized by the method of the present invention is tested on the CIFAR-100 public dataset, which has 100 classes and 600 images in each class. On the GPU 4090, a neural network architecture with an accuracy as high as can be searched by the method of the present invention, which is higher than the accuracy of the existing differentiable architecture search algorithms. Specifically, first, a supernet is constructed, and then an operation is sampled from each connection edge in the supernet to obtain a subnet. The sampling step is repeated times to obtain the initial population; the accuracy of the finally trained supernet is as shown in Figure 5 .

[0122] In the process of alternately optimizing the weights of the supernet and the encoding of the population individuals, the present invention compares the average validation accuracy of all individuals in each generation with other methods. As shown in Figure 6 , compared with the premature convergence problem caused by the small model trap in other methods, the method of the present invention can converge normally to a lower validation error rate. It further demonstrates that due to the excellent global search ability of the diffusion evolution operator of the present invention, the problem of local optimality in the search process of existing methods can be alleviated.

Claims

1. A neural architecture search method based on diffusion evolution algorithm, characterized in that: The steps include: Step 1: Build a supernet, which includes subnets, each of which assigns a weight to each operation on each edge through continuous encoding. The encoding form is an M×N matrix, where M represents the number of edges in the supernet and N represents the N candidate operations set on each edge. Step 2: Obtain pictures of different categories to construct a data set, and divide it into a training data set and a validation data set in proportion; Step 3, alternately optimize the supernet weights and subnet codes until the constraints are met, stop the optimization, and obtain the optimal subnet code, and then obtain the neural network architecture; The training data set and its labels are used to train and optimize the supernet weights, and the diffusion evolution algorithm is used to optimize the subnet coding; The optimization of subnet coding using the diffusion evolution algorithm includes the following steps: Step 301, sample and initialize individuals of the population from a normal distribution, and the population is represented as ,in, express dimensional identity matrix, represents the encoding dimension, T represents the total number of iterations, Represents individuals in the population. Each individual includes a subnet code and a scaling factor. The coding dimension of each individual satisfies: , t represents the current iteration number, represents a set of real numbers, and the encoding of the i-th individual at the t-th time step is expressed as , Q represents transpose, represents the encoding method of the Mth connecting edge of the i-th individual at the t-th time step, Indicates that in the encoded information of the i-th individual at the t-th time step, the N-th position corresponds to the value of the N-th candidate operation, represents the lower limit of the value of the Nth candidate operation in the ith individual code at the tth time step, It represents the upper limit of the value of the Nth candidate operation in the ith individual code at the tth time step; represents the scaling factor of the ith individual at the tth time step, is a random number between 0 and 1; Step 302, using the reverse denoising process to update the population individuals, the update formula is: , In the formula, represents cosine scheduling, represents the noise variance, represents noise injection; Step 303, using the validation data set to calculate the fitness value of each individual, the calculation formula is: , Where B represents the number of batches in the validation dataset, Indicates the true label of the b-th sample corresponding to the c-th category, Representation model based on parameters For the bth sample in the validation dataset The probability of predicting category c, C represents the total number of categories; Step 304, compare the fitness values ​​of the i-th individual at the t-th time step and the updated individual, and save the higher one to the next generation population, expressed as: , In the formula, Represents the individuals saved to the next generation, including subnet codes and scaling factors, represents the i-th individual in the t-th time step, Indicates the individual currently traversed The corresponding individuals processed by the diffusion evolution operator are: Represents an individual Subnet code within The fitness value of Represents an individual Subnet code within Adaptability; Step 305, converting the fitness value of the individual into a selection probability through a probability representation function, and using the selection probability to guide the update of the individuals in the population; Step 306, update the current iteration time step , jump to step 302, and stop iterating when the maximum number of iterations is reached; The assignment rule for noise injection is: , In the formula, represents the population mean.

2. According to claim 1, a neural architecture search method based on diffusion evolution algorithm is characterized in that: The setting rule of noise injection ratio is: In the early stages When the noise injection ratio is set to be greater than or equal to ; In the later stages When the noise injection ratio is set to be less than or equal to .

3. The neural architecture search method based on diffusion evolution algorithm according to claim 2 is characterized in that: Probabilistic representation function types include uniform mapping, kernel density estimation, or Gaussian mixture models; Among them, the probability representation function during uniform mapping is expressed as: , The probability representation function for kernel density estimation is expressed as: , In the formula, h represents the bandwidth, represents the jth individual in the population, represents the square of the Euclidean norm; The probability representation function of the Gaussian mixture model is expressed as: , In the formula, represents the weight of the kth Gaussian component in the mixture model, represents the covariance matrix, and the parameters , represents the set of individual indexes belonging to the Kth group, Indicates the number of elements in this collection.

4. The neural architecture search method based on diffusion evolution algorithm according to claim 3 is characterized in that: The constraints in step 3 are expressed as: Obey the constraints , In the formula, Indicates the subnet code, Indicates the subnet code corresponding to the individual with the highest fitness, represents the supernet weight, Indicates that at a given weight coefficient In the case of, the optimal weight obtained by the supernet after training, represents the validation dataset, Indicates the value of weight W when the loss function value is the lowest, Represents the minimization of the cross entropy loss function in the classification task, Represents the training dataset.

5. The neural architecture search method based on diffusion evolution algorithm according to claim 4 is characterized in that: The encoding matrix is ​​recorded as ,in Represents the weight coefficient of candidate operation n on the connecting edge m, and the formula is: , In the formula, Indicates the number of the connecting edge, The number of the candidate operation. The basic variable used to calculate the weight coefficient of candidate operation n on the connection edge m; after the optimization is completed, let , the obtained subnet code is discretized, that is, only the code with the largest value after optimization is retained on each connection edge, and the discretized code is converted to 0-1 code. The code corresponding to the operation retained on each connection edge is 1, and the rest are 0. The formula is expressed as: 。

Citation Information

Patent Citations

  • Neural architecture searching method and system based on differential evolution algorithm

    CN118196600A

  • Lightweight potential diffusion model design method and system based on evolutionary neural architecture search

    CN119559286A