Multi-objective Neural Architecture Search Method and System Based on Gradient Similarity Supernet
The gradient similarity-based super-net method addresses high computational costs and gradient conflicts in NAS by weight sharing and conflict resolution, achieving efficient and balanced multi-objective optimization of neural architectures.
Patent Information
- Application Number
- CN202411700143.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-11-26
AI Technical Summary
The existing neural architecture search methods consume high computing resources in multi-objective optimization, and traditional random sampling and gradient conflicts lead to unstable training process, making it difficult to achieve multi-objective balanced optimization.
The hypernet was built for one-time training, and the gradient similarity hypernet method was used to deal with gradient conflicts through gradient norm importance sampling and PCGrad projection method, and non-dominant sorting was performed in combination with the NSGA-III algorithm to optimize the architecture search process.
It significantly reduces the computational cost, improves resource utilization and search efficiency, achieves balanced optimization of multi-objective neural architecture, and improves the adaptability and performance of the model under different tasks.
Smart Images

Figure CN119180305B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automated machine learning, and specifically to a multi-objective neural architecture search method and system based on gradient similarity supernetworks. Background Art
[0002] Neural networks have achieved breakthrough progress in multiple fields (such as computer vision, natural language processing, healthcare, and autonomous driving), becoming the core tools driving the development of the forefront of intelligent technologies. Their architecture design directly determines the advantages and disadvantages of the model in aspects such as feature extraction, representation, and generalization ability. Therefore, a reasonable network structure is crucial. Traditional neural network design often relies on expert experience and repeated trials. This method is not only time-consuming and laborious but also difficult to adapt to the rapidly changing application requirements. For this reason, automated neural architecture search (NAS) has emerged, providing a new way to optimize the neural network structure algorithmically, enabling the model to automatically adapt to different tasks and improving the efficiency and performance of architecture design. NAS technology reduces the dependence on expert knowledge, opening up a new path for building more accurate and adaptable deep learning models, thus promoting the in-depth development of intelligent applications in various fields.
[0003] Although NAS has demonstrated great application potential in optimization and automated design, the high cost of its computing resources has become the main obstacle to further promotion. NAS usually requires complete training and evaluation on a large number of candidate architectures. Especially in scenarios with large-scale datasets and complex network structures, this evaluation process requires a large amount of computing resources and time. As the search space increases, the search complexity and computing resource requirements of NAS increase exponentially, posing a huge challenge to practical applications. To address these issues, researchers have proposed methods such as surrogate models, pruning, and weight sharing in order to reduce the overhead of candidate architecture evaluation. However, these methods have not completely solved the high-cost problem of NAS. Especially when the multi-objective optimization requirements increase, the resource consumption of NAS becomes more obvious, becoming a bottleneck restricting its application.
[0004] In this context, one-time training based on supernets has gradually become a resource-efficient solution. Supernets cover all candidate subnets in one training, so that each subnet can share weights, avoiding the waste of resources from scratch training. However, there are still some key problems in the supernet training process. First, the sampling strategy of the subnet directly affects the training effect, and traditional random sampling may cause instability in the training process; second, the optimization directions of different subnets may conflict with each other, especially in multi-objective optimization tasks, where gradient conflicts between different objectives often make it difficult to find a performance-balanced architecture. In addition, the existing supernet architectures are mostly based on a single performance indicator, and fail to fully consider the needs of multi-objective balanced optimization. Therefore, in response to the above challenges, the present invention proposes a multi-objective neural architecture search method based on gradient similarity. By optimizing the sampling strategy and improving the gradient conflict handling mechanism, multi-objective balanced optimization is achieved, thereby greatly improving the resource utilization efficiency and overall performance of neural architecture search. Summary of the invention
[0005] Purpose of the invention: The technical problem to be solved by the present invention is to provide a multi-objective neural architecture search method and system based on gradient similarity supernet in view of the shortcomings of the prior art. The method comprises the following steps:
[0006] Step 1: Build a supernet containing two or more subnets, and generate weights shared by all subnets by training the supernet; during the supernet training process, calculate the total gradient of each subnet on the optimization target;
[0007] Step 2: Get the gradient norm of each subnetwork, adopt the importance sampling strategy of the gradient norm, and dynamically adjust the sampling probability of the subnetwork according to the size of the gradient norm;
[0008] Step 3: PCGrad (Projected Conflicting Gradient) projection method is used to deal with the gradient conflict problem between different subnetworks in the multi-objective optimization process. By calculating the gradient similarity between different subnetworks, in the case of gradient conflict, the subnetwork gradient with a smaller gradient norm is projected to the orthogonal direction of the larger gradient, and finally the weighted average is performed to obtain the final gradient update direction;
[0009] Step 4: Initialize a population. Each architecture in the population inherits the shared weights directly from the supernet and is trained on the target dataset to obtain performance on the target.
[0010] Step 5: Use the non-dominated sorting genetic algorithm NSGA-III (Non-dominated Sorting Genetic Algorithm III) to perform non-dominated sorting on the sampled architectures, select the top N architectures from the architectures in the Pareto frontier, and form the next generation of parents;
[0011] Step 6: Repeat Step 4 to Step 5 until a predetermined stopping condition is met, and output the Pareto front architectures that are optimal in each target balance performance. Then select the architecture with the best performance from them, and inherit the weights from the supernet for separate training.
[0012] In Step 1, first construct a supernet containing multiple subnets, and generate weights that can be shared by all subnets by training the supernet. During the training process of the supernet, calculate the gradients of each subnet on various targets and the total gradient on multiple targets. The reason for constructing a supernet is that when evaluating the architecture performance, there is no need to train from scratch on the target dataset. Instead, the performance can be directly obtained by training after inheriting the weights of the supernet, which can significantly reduce the training time and resource consumption.
[0013] Step 1 specifically includes:
[0014] Step 1.1, design the loss function of the subnet:
[0015] During the training process of the supernet S, define the following objectives: accuracy, FLOPs (floating point operations per second), and number of parameters Params; for each subnet, design the following multi-objective loss function L:
[0016]
[0017] where K is the number of optimized objectives, is the loss function of the k-th objective, is the weight of the k-th objective, which is used to balance the contribution of each objective to the total gradient.
[0018] Step 1.2, calculate the gradients of each subnet on various targets and the total multi-objective gradient:
[0019] For the subnets sampled during the training process of the supernet, perform forward propagation to calculate the multi-objective loss function L, and then calculate the gradients of the subnet on each target through backpropagation respectively:
[0020]
[0021] where, represents the gradient of the subnet on the k-th objective, is the gradient of the loss function ;
[0022] Weight and sum the gradients on each target according to the weighting method in the loss function to obtain the total gradient of the subnet on multiple targets:
[0023]
[0024] where, is the total gradient of multi-objective optimization. The total gradient reflects the comprehensive contribution of this subnet to all optimization objectives.
[0025] Step 1.1 specifically includes:
[0026] Step 1.1.1, define the weight adjustment index: In the process of multi-objective optimization, assume that the current loss value of each objective is , and define as the relative training speed of the k-th objective at time step t-1, and the formula is:
[0027]
[0028] where represents the time step, and the value range is , where represents the maximum number of iterations for hypernetwork training; is the loss value at the previous time step; by comparing the loss changes between the current and the previous time steps, the learning speed of the current objective can be reflected.
[0029] Step 1.1.2, dynamic adjustment of weights: Calculate the weight of the k-th objective at time step :
[0030]
[0031] where K represents the total number of objectives, is the relative training speed of the i-th objective at time step t-1, T is the temperature coefficient, and exp is the natural exponential function. T is used to adjust the relative difference degree of weights between each objective. When T is small, the weight difference is large, and when T is large, the weights of each objective tend to be more uniform. Through this formula, the objectives with larger weights are usually those with slower convergence speed in the previous time step, so as to ensure that each objective can learn at a similar speed. The dynamically adjusted weights can ensure that in the process of multi-objective optimization, the attention to different objectives is continuously adjusted as the training progresses, giving priority to those objectives with slower convergence or greater importance, thereby improving the overall optimization effect and stability.
[0032] In Step 2, first, the gradient norm of each subnet needs to be obtained, and then according to the calculated magnitude of the gradient norm, the sampling probability of the subnet is dynamically adjusted. Ensure that the training focuses on the paths and data samples that have a greater impact on weight updates, so as to improve the training efficiency and effectiveness of the hypernetwork. Step 2 specifically includes:
[0033] Step 2.1, calculate the gradient norm of the subnet:
[0034] For the subnets sampled during the supernet training process, calculate the total gradient Calculate the gradient norm of:
[0035]
[0036] where is the total gradient of the \(i\)-th subnet on multiple objectives, is the gradient norm of the \(i\)-th subnet, is the \(j\)-th component of the total gradient vector of the \(i\)-th subnet ; The magnitude of the gradient norm reflects the influence of the subnet on the weight update of the entire supernet. Subnets with larger gradient norms indicate that their update directions are more important for the multi-objective optimization task, so these subnets will be given priority during the sampling process.
[0037] Step 2.2, calculate the sampling probability of subnets:
[0038] Let the gradient norm of each subnet \(i\) be , then the sampling probability is proportional to :
[0039]
[0040] where \(M\) represents the number of subnets sampled in a single sampling, is the probability that the \(i\)-th subnet is sampled. This dynamically adjusted sampling probability can ensure that the training process focuses on those paths and data samples that contribute more to weight updates, thereby improving the efficiency and effectiveness of supernet training.
[0041] In Step 3, in existing evolutionary neural architecture search, most of them adopt the crossover method based on the adjacency matrix. First, flatten the adjacency matrix into a one-dimensional vector. After the node operations are one-hot encoded, they are also connected into a one-dimensional vector, and then the two one-dimensional vectors are concatenated as the representation of an architecture.
[0042] Step 3 specifically includes:
[0043] Step 3.1, detect gradient conflicts: During the supernet training process, different subnets may generate mutually conflicting gradients in the multi-objective optimization task. To solve this gradient conflict problem, it is first necessary to detect whether there are gradient conflicts between subnets. For two different subnets and , define their total gradients on multiple objectives as and , respectively. By calculating the cosine similarity of the gradients , determine whether there are gradient conflicts between them:
[0044]
[0045] Among them, the value ranges from -1 to 1. If the cosine similarity <0, it indicates that there is a gradient conflict between the two subnets;
[0046] Step 3.2: When a gradient conflict is detected between subnets, the PCGrad projection method is used to handle the conflict; its main idea is to project the gradient of the subnet with a smaller gradient norm onto the orthogonal direction of the gradient of the subnet with a larger gradient norm to minimize the conflict to the greatest extent.
[0047] Step 3.3: Weight update and training iteration: According to the total gradient G on the multi-objectives calculated in each round, the shared weights of the supernet are updated. This update process follows the conventional gradient descent method, but because the gradients of each subnet have been processed for conflicts during the update process, the update of the weights is more stable and can better balance between multiple objectives.
[0048] Step 3.2 includes:
[0049] Step 3.2.1, Select the dominant gradient: When a gradient conflict is detected between subnets, compare the gradient norms of the two conflicting subnets. If > , then select as the dominant gradient, as the gradient to be projected;
[0050] Step 3.2.2, Gradient projection: Calculate the projection component of the gradient to be projected on the dominant gradient , and project the gradient onto the orthogonal direction:
[0051]
[0052] Among them, is the gradient after the projection of the gradient ;
[0053] Step 3.2.3, Perform weighted averaging on the projected gradient:
[0054]
[0055] Among them, represents the direction of gradient update during the training of the supernet.
[0056] For subnets with similar gradients, that is, When performing weighted averaging of gradients, gradient information can be shared among similar subnets to obtain a more comprehensive gradient optimization direction.
[0057] For subnets with gradient conflicts, that is when, weighted averaging of gradients is also performed, which can prevent a certain gradient from completely dominating the training direction of the supernet, thereby avoiding the wrong direction at the beginning of the supernet training.
[0058] By using the PCGrad projection method to handle gradient conflicts between subnets, it not only effectively reduces the gradient instability caused by multi-objective optimization, but also achieves a better balance between different tasks in multi-objective optimization by dynamically selecting the dominant gradient and projection strategy. At the same time, this method can prevent the problem of ineffective optimization caused by gradient cancellation, thereby improving the overall training efficiency and effect of the supernet.
[0059] Step 4 includes:
[0060] Initialize a population , with a size of , and each individual represents an architecture; each individual in the population is represented as architecture , ; the structure and operation type of architecture are architectures obtained by randomly sampling from the supernet; the construction process of each architecture includes randomly selecting parameters of each layer from the supernet, such as the number of channels, the size of the convolutional kernel, the network depth, etc.
[0061] After initializing the parent population, perform crossover and mutation on the parent population to obtain N new architectures , and merge the parent population with the newly generated architectures to obtain population ; for each architecture in , inherit the shared weights from the supernet S, and perform fine-tuning on the training dataset to obtain the performance of each architecture on multiple objectives, specifically expressed as:
[0062]
[0063] Among them, the elite population library initially stores the architectures of the parent population; represents the performance of the i-th architecture on multiple objectives, specifically expressed as:
[0064]
[0065] Among them represents the classification accuracy of the i-th architecture on the training dataset; Denote the floating - point operation count of the \(i\) - th architecture ; Denote the number of parameters of the \(i\) - th architecture ;
[0066] In step 5, the newly evaluated architecture is merged with the existing architectures in the current population. The merged population contains various possible architectures, providing a rich set of candidate solutions for the subsequent environmental selection. Step 5 specifically includes:
[0067] Step 5.1, perform NSGA - III non - dominated sorting on the architectures in the merged population :
[0068] Non - dominated sorting divides the architecture individuals into more than two levels. The first level consists of non - dominated architectures, indicating that the architectures are not completely dominated by other individuals in all objectives; the second level contains architectures that are only dominated by the first level, and so on. The result of non - dominated sorting is used to stratify the entire population, and the architectures in the higher levels enter the next generation preferentially:
[0069]
[0070] Among them, the objective function vector represents the performance of the \(i\) - th architecture on different optimization objectives;
[0071] Step 5.2, select reference points: After completing non - dominated sorting, NSGA - III uses pre - defined reference points to guide the selection of individuals. For each architecture , calculate the Euclidean distance between and each reference point, and assign to the reference point closest to
[0072] Step 5.3, select and generate the individuals of the next generation: Through the association of non - dominated sorting and reference points, start from the optimal non - dominated level, i.e., the first level, and select individuals layer by layer until \(N\) individuals are filled; the \(N\) selected individuals are used as the parent population of the next generation , and is stored in the elite population library ; Generate a sub - population from through tournament selection, crossover, and mutation operations.
[0073] In step 6, continue to repeat steps 4 to 5 for the new parent population until the set number of generations is reached. The finally updated elite population library includes the selected architectures that perform better on multiple objectives. From the elite population library Select the architecture with the optimal performance from them, and inherit the weights from the supernet for separate training.
[0074] The present invention also provides a multi-objective neural architecture search system based on a gradient similarity supernet, including:
[0075] A shared weight inheritance module based on the supernet: used to inherit the shared weights from the supernet, ensuring that the architecture does not have to be trained from scratch on the target dataset, but can be directly fine-tuned by inheriting the supernet weights, greatly reducing the training time and resource consumption;
[0076] An importance sampling module based on gradient norm: during the training process of the supernet, the importance sampling module based on gradient norm is used to calculate the gradient norms of each subnet, and importance sampling is performed based on these norms. This sampling strategy ensures that subnets that contribute more to the update of the supernet weights are preferentially selected for training, thereby improving the training efficiency and performance;
[0077] A projection module based on gradient similarity: solves the gradient conflict problem between different subnets during the multi-objective optimization process through the PCGrad projection method; processes the conflicting gradients by calculating the gradient similarity between subnets, ensuring stable and efficient weight updates in multi-objective optimization.
[0078] By introducing a supernet method based on gradient similarity, the present invention effectively improves the resource utilization rate and search efficiency of multi-objective neural architecture search. With the supernet shared weight mechanism, each architecture can directly inherit the supernet weights without having to train from scratch, greatly reducing the computational consumption and accelerating the architecture optimization process. In terms of subnet sampling, an importance sampling strategy based on gradient norm is adopted, making the sampling focus on the paths and data that have a greater impact on optimization, further improving the training efficiency. In addition, the present invention proposes a gradient conflict handling method based on the PCGrad projection method, which dynamically selects the dominant direction and projects conflicts through gradient similarity analysis, effectively alleviating the gradient conflict problem in multi-objective optimization and maintaining the stability of the optimization process. Compared with the previous neural architecture search methods that only focus on a single performance metric, the present invention can achieve balanced optimization under multiple objectives such as accuracy, computational amount, and number of parameters. Through the multi-objective non-dominated sorting algorithm NSGA-III, the present invention realizes effective exploration of multiple performance objectives, promoting new progress in the field of multi-objective neural architecture search. Compared with the existing methods, the present invention has better resource efficiency and comprehensive performance.
[0079] The beneficial effects of the present invention are as follows: By constructing a supernet and adopting a shared weight mechanism during the supernet training process, subsequent architectures do not need to be trained from scratch. The design of shared weights enables each architecture to directly inherit the adapted parameters from the supernet, greatly reducing the computational resources and time costs required for repeated training and effectively improving the efficiency of architecture search. At the same time, during the supernet training process, an importance sampling strategy based on gradient norm is introduced. Dynamically adjust the sampling probability of subnets, and preferentially select those architectures and data that have a greater impact on weight update. This strategy ensures that resources and computational efforts are concentrated on subnets that make significant contributions to the final performance, avoiding the instability caused by random sampling and improving the convergence speed of the training effect. In addition, when dealing with multi-objective optimization, the present invention uses the PCGrad projection method to solve the problem of gradient conflict between subnets. Calculate the gradient similarity to achieve gradient projection and conflict resolution. In the environmental selection part, the NSGA-III algorithm is adopted to achieve non-dominated sorting and optimization during the search process. This algorithm not only supports multi-objective sorting but also balances multiple optimization objectives such as accuracy, floating-point operation count, and number of parameters through the non-dominated sorting mechanism, ensuring that an optimal architecture with balanced performance and reasonable resource utilization can be found under different task requirements, thereby further improving the practicality and adaptability of the system. The multi-objective neural architecture search system of the present invention significantly reduces the computational cost on the premise of ensuring the search effect, and has a significant improvement in the balance of multi-objective optimization, providing an efficient and flexible solution for neural network architecture design. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 It is a framework diagram of the method of the present invention.
[0081] Figure 2 It is a description diagram of the supernet training of the present invention.
[0082] Figure 3 It is a description diagram of the gradient projection of the present invention.
[0083] Figure 4 It is a population update flowchart of the present invention.
[0084] Figure 5 It is a curve diagram of the model training accuracy of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0085] The following further specific descriptions of the present invention are made in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0086] In the embodiments of the present invention, the CIFAR-10 dataset is adopted. To ensure the fairness of the experiment and the reliability of the results, the dataset is divided into two parts: the supernet training dataset, which is used to train the supernet to obtain shared weights; and the population performance evaluation dataset, which is used to evaluate the performance of each architecture in the population and determine the optimal architecture during the evolution process. The image data comes from the real world, including 60,000 color images with a size of 32x32, and there are 10 categories in total, namely airplane, car, bird, cat, deer, dog, frog, horse, ship, and truck. Each category has 6,000 pictures. Among them, 50,000 are used for supernet training and 10,000 are used for population evaluation. Figure 1 It is the overall algorithm framework diagram. The implementation process and details of the multi-objective neural architecture search method based on gradient similarity supernet proposed in the embodiments of the present invention will be introduced in detail below. The method of this embodiment specifically includes the following steps:
[0087] Step 1: Supernet construction and training. The supernet constructed in the present invention contains multiple subnets. Each subnet has adjustable numbers of convolutional layers and parameters. The number of convolutional layers is set to 3 to 5 layers, and the convolutional kernel sizes are 3×3, 5×5, and 7×7. The number of channels in the convolutional layers is set to 16, 32, and 64 respectively to adapt to different computing requirements. The subnet design also includes a pooling layer for feature dimensionality reduction; the last layer is a fully connected layer for classification. During the training process of the supernet, the batch size is 64, the number of training epochs is 100, the optimizer is selected as Adam, and the initial learning rate is set to 0.001.
[0088] Step 1.1: Calculate the loss of the subnet. During the supernet training process, the dynamic weight adjustment method is adopted. Let the current loss value of each target be . To measure the training speed of each target and the need for weight adjustment, define as the relative training speed of target k at time step t-1, and its formula is as follows:
[0089]
[0090] where represents the time step, and its value range is , where represents the maximum number of iterations 100 for supernet training; is the loss value at the previous time step. By comparing the loss changes at the current and previous time steps, the learning speed of the current target can be reflected.
[0091] According to the training speed index , calculate the weight of the kth target at time step , and the weight adjustment formula is as follows:
[0092]
[0093] Among them, K represents the total number of targets, which is set to 4; T is the temperature coefficient, initially set to 0.1. This dynamic weight adjustment method makes the training process focus on the targets with slower convergence, thus ensuring the balanced optimization of multiple targets.
[0094] Step 1.2: As Figure 2 shown, for the subnets sampled during the supernet training process, perform forward propagation to calculate the multi-objective loss function L, and then calculate the gradient of the subnet on each target through backpropagation , and weight and sum the gradients on each target according to the weighting method in the loss function to obtain the total gradient of the subnet on multiple targets . And calculate its total gradient of the gradient norm. The calculation formula of the gradient norm is as follows:
[0095]
[0096] Among them is the total gradient of the i-th subnet on multiple targets, is the gradient norm of the i-th subnet, is the j-th component of the total gradient vector of subnet i, and K represents the dimension of the gradient vector (i.e., the total number of optimization targets).
[0097] Step 2: For the subnets sampled during the supernet training process, based on the importance of the subnet gradient norm, adopt a probability sampling strategy to select subnets for training. Take the gradient norm of each subnet as its sampling importance index. Let the gradient norm of each subnet i be , then its sampling probability is proportional to :
[0098]
[0099] Among them, M represents the number of subnets sampled each time, which is set to 16, is the probability that subnet i is sampled. The subnet with the highest probability is used to update the weights of the supernet. The specific process is as Figure 3 shown.
[0100] Step 3: As Figure 3 shown, it describes several situations of the gradients on the multi-objective optimization task that may occur when the supernet weights are updated during the supernet training process. At the same time, gradient projection processing is performed on the conflicting gradients.
[0101] Step 3.1: First, it is necessary to detect whether there is a gradient conflict between subnets. For two different subnets and , define their total gradients and on multiple objectives. By calculating the cosine similarity of the gradients, determine whether there is a conflict between them:
[0102]
[0103] where, The value of ranges from -1 to 1. If the cosine similarity < 0, it indicates that there is a gradient conflict between the two subnets.
[0104] Step 3.2: As shown in "Gradient Conflict" in Figure 3 . When a gradient conflict is detected between subnets, compare the gradient norms of the conflicting subnets. Assume > , then select as the dominant gradient and as the gradient to be projected. Calculate the projection component of the gradient to be projected on the dominant gradient , and project the gradient onto the orthogonal direction of :
[0105]
[0106] And perform a weighted average on the projected gradient:
[0107]
[0108] where, represents the direction of gradient update during the training of the supernet.
[0109] Step 3.3: According to the total gradient on multiple objectives calculated in each round, update the shared weights of the supernet. This update process follows the conventional gradient descent method.
[0110] Step 4: As shown in Figure 4 , initialize a population with a size of , set to 50. Each individual represents an architecture. Each individual in the population is represented as architecture , . After the initialization of the parent population, perform crossover and mutation on the parent population to obtain new architectures , and merge the parent population with the newly generated architectures to obtain For each architecture in an individual inherits shared weights from the supernet to obtain the performance of each architecture on multiple objectives, specifically expressed as:
[0111]
[0112] Among them, the elite population library initially stores the architectures of the parent population; is expressed as the performance of the th architecture on multiple objectives, specifically expressed as:
[0113]
[0114] represents the th architecture's classification accuracy on the training dataset; represents the th architecture's floating-point operation count; represents the th architecture's
[0115] parameter count.
[0116] Step 5: Use the NSGA-III multi-objective evolutionary algorithm to perform non-dominated sorting on the merged architectures. Select the top N architectures from the non-dominated sorted architectures to form the next generation of parents. In Step 5.1, perform NSGA-III non-dominated sorting on the architectures obtained by merging. Non-dominated sorting divides the architecture individuals into multiple levels. The first level consists of non-dominated architectures, indicating that they are not completely dominated by other individuals in all objectives. The second level contains architectures that are only dominated by the first level, and so on. The result of non-dominated sorting is used to stratify the entire population, and architectures at higher levels have priority to enter the next generation.
[0117]
[0118] Among them, the objective function vector represents the th architecture's performance on different optimization objectives.
[0119] In Step 5.2, after completing non-dominated sorting, to ensure the diversity of the population, NSGA-III uses predefined reference points to guide the selection of individuals. For each architecture , calculate the Euclidean distance between it and each reference point. Each individual is assigned to the reference point closest to it.
[0120] Step 5.3, through the association of non-dominated sorting and reference points, starting from the optimal non-dominated level (i.e., non-dominated level 1), select individuals layer by layer until N individuals (set to 50) are filled. The N selected individuals are used as the parent population of the next generation. , and store them in the elite population library . From , generate a sub-population through tournament selection, crossover, and mutation operations .
[0121] Repeat steps 4 to 5 until the stopping condition is reached, and then output a set of architectures that perform well on multiple objectives. Finally, select the architecture with the highest performance in the elite population library as the final optimal architecture. And select the architecture with the best performance from it, and inherit the weights from the supernet for separate training. As Figure 5 shown, it shows the performance of the architecture with the best performance selected from the elite population library after independent training for 100 epochs on CIFAR-10. This model was selected as a candidate representing high performance and low parameter count during the architecture search process. It has 3.8 MB of parameters and reached a Top-1 accuracy (i.e., the highest accuracy) of 97.52% at the end of training. As can be seen from the figure, the accuracy of the model increased rapidly within the first 20 training epochs, showing good initial convergence ability, and continued to grow steadily during the subsequent training process, eventually approaching a saturated high accuracy.
[0122] The present invention provides a multi-objective neural architecture search method and system based on a gradient similarity supernet. There are many methods and ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.
Claims
1. A multi-objective neural architecture search method based on gradient similarity supernet, characterized in that It includes the following steps: Step 1: Construct a supernet containing more than two subnets, and generate weights shared by all subnets by training the supernet; during the training process of the supernet, calculate the total gradient of each subnet on the optimization objective; Step 2: Obtain the gradient norm of each subnet, adopt the importance sampling strategy of the gradient norm, and dynamically adjust the sampling probability of the subnet according to the magnitude of the gradient norm; Step 3: Adopt the PCGrad projection method to handle the gradient conflict problem between different subnets in the multi-objective optimization process, calculate the gradient similarity between different subnets, and perform weighted average to obtain the final gradient update direction; Step 4: Use the CIFAR-10 dataset and divide the dataset into two parts: Supernet training dataset: Used to train the supernet to obtain shared weights; Population performance evaluation dataset: Used to evaluate the performance of each architecture in the population and determine the optimal architecture during the evolution process; Initialize a population, and each architecture in the population directly inherits the shared weights from the supernet and is trained on the target dataset to obtain the performance on the target; Step 5: Use the non-dominated sorting genetic algorithm NSGA-III to perform non-dominated sorting on the sampled architectures, select the top N architectures from the architectures on the Pareto front to form the next generation of parents; Step 6: Repeat Step 4 to Step 5 until the predetermined stop condition is met, output the Pareto front architecture with the best balanced performance on each target, select the architecture with the best performance from it, and inherit the weights from the supernet for separate training; Step 1 includes: Step 1.1, design the loss function of the subnet: During the training process of the supernet S, define the following objectives: accuracy Accuray, floating-point operation count FLOPs, number of parameters Params; for each subnet, design the following multi-objective loss function L: Among them, K is the number of optimized targets, and L k is the loss function of the k-th target, and λ k is the weight of the k-th target; Step 1.2, calculate the gradient of each target of the subnet and the multi-objective total gradient: For the subnets sampled during the training process of the supernet, perform forward propagation to calculate the multi-objective loss function L, and then calculate the gradient of the subnet on each target through backpropagation: Among them, g k represents the gradient of the subnet on the k-th target, which is the gradient of the loss function L k ; Weight the gradients on each target according to the weighting method in the loss function and sum them up to obtain the total gradient of the subnet on multiple targets: where G is the total gradient of the multi-objective optimization.
2. The method according to claim 1, wherein Step 1.1 specifically includes: Step 1.1.1, Define the weight adjustment index: In the multi-objective optimization process, assume that the current loss value of each objective \(L\) k is \(L\) k (t - 1), and define \(r\) k (t - 1) as the relative training speed of the \(k\)-th objective at time step \(t - 1\), and the formula is: where t represents the time step, and its value range is t = 1, 2, 3,..., T max , where T max represents the maximum number of iterations for hypernetwork training; L k (t - 2) is the loss value at the previous time step; Step 1.1.2, dynamic adjustment of weights: calculate the weight of the k-th target at time step t: where K represents the total number of targets, and r i (t - 1) is the relative training speed of the i-th target at time step t - 1, T is the temperature coefficient, and exp is the natural exponential function.
3. The method according to claim 2, wherein Step 2 includes: Step 2.1, calculate the gradient norm of the subnet: For the subnets sampled during the training process of the supernet, calculate the gradient norm of the total gradient G; where G i is the total gradient of the i-th subnet on multiple targets, and ‖G i ‖ is the gradient norm of the i-th subnet, and G ij is the j-th component of the total gradient vector G i of the i-th subnet; Step 2.2, calculate the sampling probability of the subnet: Let the gradient norm of each subnet \(i\) be \(\|G\) i \|. Then the sampling probability is proportional to \(\|G\) i \|:\) where M represents the number of single sampling subnets, and p i is the probability that the i-th subnet is sampled.
4. The method according to claim 3, wherein Step 3 includes: Step 3.1, Detect gradient conflict: For two different subnets S i and S j , define their total gradients on multiple objectives as G i and G j respectively. By calculating the cosine similarity cos(θ ij ), determine whether there is a gradient conflict between them: Among them, if the cosine similarity cos(θ ij ) < 0, it indicates that there is a gradient conflict between the two subnets; Step 3.2: When detecting a gradient conflict between subnets, adopt the PCGrad projection method to handle the conflict; Step 3.3: Weight update and training iteration: Update the shared weights of the supernet according to the total gradient G on the multi-objectives calculated in each round.
5. The method according to claim 4, wherein Step 3.2 includes: Step 3.2.1, select the dominant gradient: When a gradient conflict is detected between subnets, compare the gradient norms of the two conflicting subnets. If ‖G i ‖ > ||G j ||, then select ‖G i ‖ as the dominant gradient and ||G j || as the gradient to be projected; Step 3.2.2, gradient projection: Calculate the gradient G to be projected j On the dominant gradient G i The projection component, project the gradient G j Project onto the orthogonal direction of G i : Among them, G j ′ is the gradient G j after projection; Step 3.2.3, perform weighted average on the projected gradients: Among them, G s represents the direction of gradient update during hypernetwork training.
6. The method according to claim 5, wherein Step 4 includes: Initialize a population P0 with size N, where each individual represents an architecture; each individual in the population is represented as the architecture Arch i , i ∈ [1, N]; the architecture Arch i 's structure and operation type are architectures obtained by randomly sampling from the super network; After the initialization of the parent population, the parent population is subjected to crossover and mutation to obtain N new architectures Q0, and the parent population is merged with the newly generated architectures to obtain a population R0 = P0 ∪ Q0; for each architecture Arch in R0 i , inherit the shared weights W from the supernetwork S S , and perform fine-tuning on the training dataset Data train to obtain the performance of each architecture on multiple objectives, specifically expressed as: Among them, the initial content stored in the elite population library Archive is the architecture of the parent population; X i represents the performance of the i-th architecture on multiple objectives, specifically expressed as: Among them, Acc i represents the classification accuracy of the i-th architecture Arch i on the training dataset; Flops i represents the floating-point computation count of the i-th architecture Arch i ; Pramas i represents the number of parameters of the i-th architecture Arch i 。 7. The method according to claim 6, wherein Step 5 includes: Step 5.1, perform NSGA-III non-dominated sorting on the architectures in the merged population R0: Non-dominated sorting divides the architecture individuals into more than two levels. The first level consists of non-dominated architectures, indicating that the architectures are not completely dominated by other individuals in all objectives; the second level contains architectures that are only dominated by the first level. The result of non-dominated sorting is used to stratify the entire population, and the architectures in the higher levels enter the next generation first: F(Arch i ) = [Acc i , Flops i , Pramas i where the objective function vector F(Arch i ) represents the performance of the i-th architecture Arch i on different optimization objectives; Step 5.2, Select reference points: After non-dominated sorting is completed, NSGA-III uses pre-defined reference points to guide the selection of individuals. For each architecture Arch i , calculate the Euclidean distance between Arch i and each reference point, and assign Arch i to the reference point closest to Arch i ; Step 5.3, select to generate the next generation of individuals: Through non-dominated sorting and the association with reference points, start from the optimal non-dominated level, i.e., the first level, and select individuals layer by layer until N individuals are filled; the N selected individuals are used as the parent population P of the next generation t+1 , store P t+1 in the elite population library Archive; generate the sub-population Q t+1 through tournament selection, crossover, and mutation operations from P t+1 .
8. The method according to claim 7, wherein In Step 6, continue to repeat Steps 4 to 5 for the new parent population until the set number of generations is reached. The finally updated elite population library Archive includes the selected architectures that perform better in multiple objectives. Select the architecture with the best performance from the elite population library Archive and inherit the weights from the supernet for separate training.
9. A multi-objective neural architecture search system based on a gradient similarity super network implemented by the method according to any one of claims 1 to 8, characterized in that, Including: Shared weight inheritance module based on the supernet: used to inherit the shared weights from the supernet, ensuring that the architecture does not need to be trained from scratch on the target dataset, but can be directly fine-tuned by inheriting the supernet weights, greatly reducing the training time and resource consumption; Importance sampling module based on gradient norm: during the training process of the supernet, the importance sampling module based on gradient norm is used to calculate the gradient norms of each subnet and perform importance sampling based on these norms; Projection module based on gradient similarity: solve the gradient conflict problem between different subnets during the multi-objective optimization process through the PCGrad projection method; handle the conflicting gradients by calculating the gradient similarity between subnets to ensure stable and efficient weight updates in multi-objective optimization.