A Multi-Objective Neural Architecture Search Method Based on a Semi-Supervised Performance Predictor
Through the semi-supervised performance predictor and multi-objective selection method, the problem that predictors in the prior art cannot guarantee high confidence is solved, the efficiency and accuracy of neural architecture search is improved, and efficient search of high-precision neural network structures is realized.
Patent Information
- Application Number
- CN202211157727.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-09-22
AI Technical Summary
Existing performance predictors cannot guarantee high confidence in excellent individuals in the predicted population, resulting in the inability to effectively search for high-precision neural networks.
A multi-objective neural architecture search method based on a semi-supervised performance predictor is adopted. By integrating KNN regression model and truncation operation, predictive confidence is constructed to measure prediction accuracy, and environmental selection is combined with genetic algorithms, and candidate network structures with high confidence and high precision are preferred.
It improves the efficiency of neural architecture search, reduces prediction errors, and increases the probability of searching for high-precision neural network structures.
Smart Images

Figure CN115620046B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of semi-supervised learning, and more particularly, to a multi-objective neural architecture search method based on a semi-supervised performance predictor. Background Art
[0002] Deep neural networks have achieved great success in various practical applications such as image classification, natural language processing, and object detection. This is mainly due to the powerful feature extraction ability of neural networks with deep structures, which can directly learn meaningful features from raw data with almost no explicit feature engineering. This enables researchers to focus on the design of neural architectures. However, designing neural structures relies heavily on the prior knowledge and experience of researchers. Currently, promising convolutional neural network (CNN) models are all manually designed by researchers with rich knowledge of neural networks and image processing. In practice, most developers do not have such knowledge. In addition, neural network architectures are usually problem-specific, and different problems mean different architectures. Neural architecture search (NAS) aims to automate the architecture design of neural networks and is considered a promising method to address the above challenges.
[0003] NAS automates the search for neural architectures with limited resources to achieve the best possible performance with minimal human intervention. Early work on NAS used reinforcement learning methods, and the neural structures searched achieved state-of-the-art classification accuracy in image classification tasks. Subsequently, large-scale evolutionary work verified the feasibility of this concept again, obtaining similar results by using evolutionary computation. The key technology behind NAS involves using search strategies to find the best neural structure by comparing the performance of a large number of candidate neural structures. Therefore, the performance ranking of candidate neural structures is very important. Previous NAS usually fully trains candidate neural structures and then obtains the ranking of candidate neural structures based on their performance on the validation set. However, this method is very time-consuming because there are too many candidate neural structures to train. Most researchers find it unacceptable for such a large resource consumption. Therefore, neural architecture search has gradually turned to being efficient and lightweight.
[0004] A common method for accelerating neural architecture search is the performance predictor. It only needs to train a small part of the neural network. By using these neural networks with true accuracy as training data to train the performance predictor, it can predict the performance of other neural networks, avoiding the need to train a large number of models from scratch during the search process. However, the performance predictor trained with only a small amount of data under supervised learning is prone to overfitting, so the accuracy prediction of the searched neural network is inaccurate. Semi-supervised learning uses high-quality unlabeled data as training data to train the performance predictor, greatly alleviating the overfitting phenomenon of supervised learning. However, the performance of the model constructed by semi-supervised learning is largely affected by the base learner, and its performance will also weaken with the addition of unlabeled data. In addition, existing performance predictors ensure the prediction accuracy by screening high-confidence individuals. The convergence of evolutionary algorithms is largely ensured by screening excellent individuals in the population. The performance predictor cannot guarantee the high confidence of excellent individuals in the predicted population, so it cannot guarantee the search for high-precision neural networks.
[0005] In the existing technology, a Chinese invention patent discloses a method for determining the architecture of a task neural network configured to perform a specific machine learning task. The method includes: obtaining data that specifies a current set of candidate architectures for the task neural network; for each candidate architecture in the current set: using a performance prediction neural network with multiple performance prediction parameters to process the data specifying the candidate architecture, the performance prediction neural network being configured to process the data specifying the candidate architecture according to the current values of the performance prediction parameters to generate a performance prediction, the performance prediction characterizing how well a neural network with the candidate architecture will perform after training on a specific machine learning task; and generating an updated set of candidate architectures by selecting one or more candidate architectures from the current set based on the performance predictions of the candidate architectures in the current set, which cannot guarantee the high confidence of excellent individuals in the predicted population and cannot guarantee the search for high-precision neural networks. Summary of the Invention
[0006] To solve the technical defect that the existing performance predictor cannot guarantee the high confidence of excellent individuals in the predicted population and cannot guarantee the search for high-precision neural networks, the present invention provides a multi-objective neural architecture search method based on a semi-supervised performance predictor.
[0007] To achieve the above invention purposes, the technical solutions adopted are as follows:
[0008] A multi-objective neural architecture search method based on a semi-supervised performance predictor includes the following steps:
[0009] S1: Encode the search space, map the operations in the search space to integers to form samples;
[0010] S2: Randomly sample N neural network architectures from the samples and obtain their accuracies as the initial sample P. Before entering the genetic algorithm routine, train two ensemble KNN regression models as the initial semi-supervised predictors;
[0011] S3: The initial sample P generates a candidate population P1, and its accuracy is predicted by the semi-supervised predictor and its confidence is calculated. The candidate population P1 is used as the real initial population P0, and a P0 training performance predictor is obtained;
[0012] S4: The parental population P t Undergo crossover and mutation to obtain a crossover and mutation sub-population P''. The P0 training performance predictor combines the parental population P t and the crossover and mutation sub-population P'' and performs non-dominated sorting;
[0013] S5: According to the non-dominated sorting result, apply the multi-objective selection method to preferentially select individuals with high confidence from the individuals with a low domination rank and add them to the new generation population P t+1 , until the number of the new generation population P t+1 is equal to the population size N, and the search is completed here.
[0014] In the above solution, to improve the efficiency of neural architecture search, a new type of semi-supervised predictor is proposed, and the prediction error of this predictor is effectively reduced through ensemble learning and truncation operations; for this performance predictor, a prediction confidence is constructed to measure the accuracy of the prediction accuracy of the predictor, providing a selection direction for the environmental selection of the genetic algorithm. During the evolution process, the environmental selection problem is transformed into a multi-objective selection problem of the prediction confidence and prediction accuracy of the performance predictor, making the search of the genetic algorithm proceed towards candidate network architectures with high prediction accuracy and high confidence, and improving the probability of searching for a neural network architecture with high accuracy.
[0015] Preferably, in step S1, the search space is the space composed of all encoded architectures, and the search space includes five functional layers, namely 1×1 convolution, 3×3 convolution, 3×3 average pooling, skip connection, and zero.
[0016] Preferably, in step S2, training the ensemble KNN regression model includes the following steps:
[0017] S21: Set the training data size as M, the total number of sub-models of the ensemble model as N, the data dimension as D, and at the same time set the basic model parameters: the number of neighbors k, and the distance calculation parameters as p1 and p2 respectively;
[0018] S22: Take m data where m < M and d dimensions where d < D to form a small training sample. Train the sub-model according to the set number of neighbors k and distance calculation parameters p1 and p2 of the KNN ensemble model, and finally train N sub-regression models;
[0019] S23: N sub-models respectively make predictions on the unlabeled data, and take the average of the prediction results as the prediction result of the ensemble model.
[0020] In the above solution, for the convenience of calculation, two ensemble KNN models EnKNN1 and EnKNN2 are used as the basic models of the semi-supervised regression algorithm, and it is set that the size of the training data set is equal to the size of the prediction data set. The number of neighbors of the ensemble model is set to 3, and the distance calculation methods are set as follows: p1 is the Euclidean distance and p2 is the Minkowski distance under the value p2. The prediction error of the semi-supervised performance predictor is inevitable. At the same time, as the number of unlabeled samples in the training samples increases continuously, the error of the performance predictor will accumulate continuously. Although an ensemble KNN model is constructed to replace the classical KNN algorithm, effectively reducing the prediction error of the samples. However, semi-supervised regression is a process of continuously adding high-quality prediction candidate samples to the training samples, and its prediction error will increase with the addition of the prediction samples. Therefore, the semi-supervised predictor is truncated to avoid the increase of the prediction error. That is, it is set that after predicting N candidate samples each time, the constructed semi-supervised predictor will stop predicting and then retrain.
[0021] Preferably, in step S3, it includes the following steps:
[0022] S31: Randomly sample N neural architectures and encode them, and train to obtain their accuracies;
[0023] S32: Use the N encoded neural architectures as training samples to train two regression models EnKNN1 and EnKNN2;
[0024] S33: Generate N candidate samples, and EnKNN1 and EnKNN2 predict the accuracies of the candidate samples and predict the improved accuracies:
[0025] S34: If the maximum prediction improvement of the samples predicted by EnKNN1 is greater than 0, then take the sample with the maximum prediction improvement and its prediction accuracy as the new training samples of EnKNN2; if the maximum prediction improvement of the samples predicted by EnKNN1 is less than or equal to 0, then no training samples are taken as prediction samples; EnKNN2 performs the same operation; at the same time, delete the predicted samples from the candidate samples;
[0026] S35: If the training samples of EnKNN1 increase, then retrain EnKNN1 with the new training samples; if the training samples of EnKNN2 increase, then retrain EnKNN2 with the new training samples;
[0027] S36: When the training samples of EnKNN1 increase, use EnKNN2 to calculate the confidence of the newly added training samples of EnKNN1, and at the same time use EnKNN1 to calculate the confidence of the newly added training samples of EnKNN2;
[0028] S37: Loop S34, S35, and S36. If P candidate samples are all used as training samples and added to EnKNN1 and ENKNN2 respectively, then the P samples are all predicted, and this performance predictor is truncated and no longer predicts; if the prediction improvements of all the predicted samples of EnKNN1 and EnKNN2 are less than 0, the prediction error of this group of samples is large, and this performance predictor is forcibly truncated;
[0029] S38: Output the prediction results of EnKNN1 and EnKNN2 for the candidate samples and the confidence of the prediction results in step S36;
[0030] S39: Use the neural network with real accuracy obtained by original sampling to train the truncated EnKNN1 and EnKNN2.
[0031] Preferably, in step S33, the accuracy of predicting candidate samples is calculated by the following formula:
[0032] .
[0033] Preferably, in step S36, after retraining the model for the predicted samples, the average prediction deviation of all labeled samples. The calculation formula is as follows:
[0034]
[0035] where represents the number of labeled samples, f(x) represents the true label of x, and y represents the current predicted value.
[0036] In the above solution, the error of the performance predictor is reduced to a certain extent by the ensemble model, but it still exists widely. A confidence is needed to depict the prediction error of the performance predictor. The prediction confidence describes the accuracy of the prediction by the predictor, which has a different meaning from the prediction improvement of the classical semi-supervised regression algorithm. At the same time, the training samples of the semi-supervised performance predictor proposed by the present invention include labeled samples and unlabeled samples. The prediction results of samples without true labels are inaccurate. Therefore, the present invention constructs a prediction confidence.
[0037] Preferably, in step S4, for the parental population P tPerform pairwise crossover to obtain a crossover population P', then perform a mutation operation on the crossover sub-population P' to generate a crossover and mutated sub-population P'', use a semi-supervised performance predictor to predict the accuracy of the crossover and mutated sub-population P'' and calculate the confidence level corresponding to the accuracy. After that, use P0 to train the performance predictor to mix the parent population P t and the crossover and mutated sub-population P'' into a population P p , and use the predicted accuracy and predicted confidence level of the population P p as two objectives to be optimized, and perform non-dominated sorting.
[0038] Preferably, in step S5, according to the non-dominated sorting result, add individuals with a lower domination level to the new generation population P t+1 , when the number of the new generation population P t+1 plus the number of individuals in a certain layer of P is greater than the population size N, use the predicted confidence level as the selection direction, and preferentially select individuals with a high confidence level to join the new generation population Pt+1 until the number of the new generation population Pt+1 is equal to the population size N.
[0039] In the above solution, multi-objective optimization involves the analysis of the domination relationship between solution vectors and the Pareto front. A multi-objective optimization problem with m objectives and n decision variables can be described as:
[0040]
[0041] where, Ω⊆ℝ n is the decision space, x∈{x1,x2,……xn} is the feasible region of the decision variables, and n is the dimension of the variables. ℝ m is the objective space, m is the number of objective function values, and fi(x) is the value of the i-th objective function. When the number of objectives is 2 - 3, it is called a general multi-objective optimization problem. When the number of objectives is 4 or more, it is called a high-dimensional multi-objective optimization problem. h u ( x ) and g v ( x ) are inequality constraints and equality constraints respectively. The solutions that satisfy the constraint conditions are called feasible solutions. The absolute optimal solution needs to simultaneously optimize multiple objectives while satisfying the constraints. However, due to the mutual exclusivity of the decision variables among multiple objectives, it is difficult to obtain the absolute optimal solution. Generally, an optimal solution set is obtained, and this set of optimal solution sets is generally called the Pareto front.
[0042] Preferably, during the two processes where the multi-objective selection method loops within the predefined termination conditions, when the algorithm ends, output the final multi-objective optimal solution set. The final multi-objective optimal solution set contains individuals with both high prediction accuracy and high confidence level.
[0043] Pareto domination: Suppose p and q are any two different individuals in the population NP. When the following conditions are met:
[0044] (1) For all sub-goals, fk(p) ≤ fk(q) (k = 1, 2, 3, ……, m).
[0045] (2) ∃l ∈ {1, 2, 3, ……, m}, such that fl(p) < fl(q). Here, m is the number of sub-goals.
[0046] It is said that p dominates q, denoted as p ≻ q.
[0047] Pareto optimal solution: The optimal solution of multi-objective optimization is usually called the Pareto optimal solution. When there is no other individual in the objective space that dominates x, x is called the Pareto optimal solution.
[0048] Pareto front: The Pareto front (PF) is the projection of the Pareto optimal solution set in the objective space.
[0049] The genetic algorithm retains high-quality individuals in the candidate population and the parent population through environmental selection. In other words, the genetic algorithm continuously explores the region of high-quality individuals in the hope of finding the best individual. The evolutionary algorithm based on performance predictors (SAEA) is no exception. However, most current environmental selection methods of SAEA are based on confidence selection. It retains the individuals with high prediction confidence for the next-generation population, ensuring the accuracy of the prediction. However, the predictor cannot guarantee that the predicted individuals have both high prediction confidence and high accuracy. This indicates that SAEA cannot guarantee to explore the region where high-precision individuals are located. In SAEA, the predicted individuals need to make a trade-off between high prediction confidence and high prediction accuracy. The SAEA algorithm needs to explore more regions where real high-quality individuals are located, and at the same time, it needs to find more individuals with high prediction confidence and high quality. Regarding the selection of prediction confidence and prediction accuracy as a multi-objective optimization problem, the Pareto domination relationship in the evolutionary process is used to screen out the neural network structures with high prediction confidence and high prediction accuracy. For the selection of key individuals, in order to reduce the prediction error of the proposed accuracy predictor, the present invention gives priority to the individuals with high prediction confidence.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] A multi-objective neural architecture search method based on a semi-supervised performance predictor is provided by the present invention. To improve the efficiency of neural architecture search, a novel semi-supervised predictor is proposed, and the prediction error of this predictor is effectively reduced through ensemble learning and truncation operations. For this performance predictor, prediction confidence is constructed to measure the accuracy of the prediction accuracy of the predictor, providing a selection direction for the environmental selection of the genetic algorithm. During the evolution process, the environmental selection problem is transformed into a multi-objective selection problem of the prediction confidence and prediction accuracy of the performance predictor, enabling the search of the genetic algorithm to be directed towards candidate network structures with high prediction accuracy and high confidence, and increasing the probability of searching for a high-precision neural network structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is the flowchart of the method of the present invention;
[0053] Figure 2 is the flowchart of the neural architecture search method of the present invention;
[0054] Figure 3 is the schematic diagram of the ensemble model proposed by the present invention;
[0055] Figure 4 is the flowchart of the semi-supervised performance predictor proposed by the present invention;
[0056] Figure 5 is the search space diagram of NAS-Bench 201 of the present invention;
[0057] Figure 6 is the convergence curve diagram of the pareto front of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0059] The present invention will be further described below with reference to the drawings and embodiments.
[0060] Embodiment 1
[0061] As Figures 1 to 4 shown, a multi-objective neural architecture search method based on a semi-supervised performance predictor includes the following steps:
[0062] S1: Encode the search space, map the operations in the search space to integers, and form samples;
[0063] S2: Randomly sample N neural network structures in the samples and obtain their accuracies as the initial sample P. Before entering the genetic algorithm routine, train two ensemble KNN regression models as the initial semi-supervised predictor;
[0064] S3: The initial sample P generates a candidate population P1, and its accuracy is predicted and its confidence is calculated through a semi-supervised predictor. The candidate population P1 is used as the true initial population P0 to obtain a P0 training performance predictor.
[0065] S4: The parent population P t undergoes crossover and mutation to obtain a crossover and mutation sub-population P''. The P0 training performance predictor mixes the parent population P t and the crossover and mutation sub-population P'' and performs non-dominated sorting.
[0066] S5: According to the non-dominated sorting result, individuals with a lower domination rank are preferentially selected using the multi-objective selection method to add individuals with high confidence to the new generation population P t+1 , until the number of the new generation population P t+1 is equal to the population size N, and the search is completed here.
[0067] In the above solution, to improve the efficiency of neural architecture search, a new type of semi-supervised predictor is proposed, and the prediction error of this predictor is effectively reduced through ensemble learning and truncation operations; a prediction confidence is constructed for this performance predictor to measure the accuracy of the predictor's prediction accuracy, providing a selection direction for the environmental selection of the genetic algorithm. During the evolution process, the environmental selection problem is transformed into a multi-objective selection problem of the prediction confidence and prediction accuracy of the performance predictor, enabling the search of the genetic algorithm to be directed towards candidate network structures with high prediction accuracy and high confidence, and increasing the probability of searching for a high-precision neural network structure.
[0068] Preferably, in step S1, the search space is the space composed of all encoded architectures, and the search space includes five functional layers: 1×1 convolution, 3×3 convolution, 3×3 average pooling, skip connection, and zero.
[0069] Preferably, as Figure 3 shown, in step S2, training the ensemble KNN regression model includes the following steps:
[0070] S21: Set the training data size as M, the total number of sub-models of the ensemble model as N, the data dimension as D, and at the same time set the basic model parameters: the number of neighbors k, and the distance calculation parameters as p1 and p2 respectively.
[0071] S22: Take m data (m < M) and d dimensions (d < D) to form a small training sample. Train the sub-model according to the set number of neighbors k and distance calculation parameters p1 and p2 of the KNN ensemble model, and finally train N sub-regression models.
[0072] S23: The N sub-models respectively predict the unlabeled data, and take the average of the prediction results as the prediction result of the ensemble model.
[0073] In the above solution, for the convenience of calculation, two integrated KNN models, EnKNN1 and EnKNN2, are used as the basic models of the semi-supervised regression algorithm, and the size of the training data set is set to be equal to the size of the prediction data set. The number of neighbors of the integrated model is set to 3, and the distance calculation methods are set as follows: p1 is the Euclidean distance and p2 is the Minkowski distance under the value p2. The prediction error of the semi-supervised performance predictor is inevitable. At the same time, as the number of unlabeled samples in the training samples increases, the error of the performance predictor will accumulate continuously. Although an integrated KNN model is constructed to replace the classical KNN algorithm, effectively reducing the prediction error of the samples. However, semi-supervised regression is a process of continuously adding high-quality prediction candidate samples to the training samples, and its prediction error will increase with the addition of prediction samples. Therefore, the semi-supervised predictor is truncated to avoid the increase of prediction error. That is, after predicting N candidate samples each time, the constructed semi-supervised predictor will stop predicting and then retrain.
[0074] Preferably, in step S3, as Figure 4 shown, it includes the following steps:
[0075] S31: Randomly sample N neural architectures and encode them, and train to obtain their accuracies;
[0076] S32: Use the N encoded neural architectures as training samples to train two regression models, EnKNN1 and EnKNN2;
[0077] S33: Generate N candidate samples, and EnKNN1 and EnKNN2 predict the accuracies of the candidate samples and predict the improved accuracies:
[0078] S34: If the maximum prediction improvement of the samples predicted by EnKNN1 is greater than 0, then use the sample with the maximum prediction improvement and its prediction accuracy as the new training sample of EnKNN2; if the maximum prediction improvement of the samples predicted by EnKNN1 is less than or equal to 0, then no training sample is used as a prediction sample; EnKNN2 performs the same operation; at the same time, the predicted samples are deleted from the candidate samples;
[0079] S35: If the training samples of EnKNN1 increase, retrain EnKNN1 with the new training samples; if the training samples of EnKNN2 increase, retrain EnKNN2 with the new training samples;
[0080] S36: When the training samples of EnKNN1 increase, use EnKNN2 to calculate the confidence of the new training samples of EnKNN1, and at the same time use EnKNN1 to calculate the confidence of the new training samples of EnKNN2;
[0081] S37: Loop through S34, S35, and S36. If all P candidate samples have been added as training samples to EnKNN1 and ENKNN2 respectively, then predictions are obtained for all P samples, and this performance predictor is truncated and no longer makes predictions. If the prediction improvements for all prediction samples in EnKNN1 and ENKNN2 are less than 0, then the prediction error for this set of samples is large, and this performance predictor is forcibly truncated.
[0082] S38: Output the prediction results of EnKNN1 and EnKNN2 for the candidate samples and the confidence levels of the prediction results in step S36.
[0083] S39: Train the truncated EnKNN1 and EnKNN2 using the neural network with true accuracy obtained from the original sampling.
[0084] Preferably, in step S33, the accuracy of predicting the candidate samples is calculated by the following formula:
[0085] .
[0086] Preferably, in step S36, after retraining the model for the prediction samples, the average prediction deviation of all labeled samples. The calculation formula is as follows:
[0087]
[0088] Where represents the number of labeled samples, f(x) represents the true label of x, and y represents the current predicted value.
[0089] In the above solution, the error of the performance predictor is reduced to a certain extent by the ensemble model, but it still widely exists. A confidence level is needed to depict the prediction error of the performance predictor. The prediction confidence describes the accuracy of the prediction by the predictor, which has a different meaning from the prediction improvement of the classical semi-supervised regression algorithm. At the same time, the training samples of the semi-supervised performance predictor proposed in the present invention include labeled samples and unlabeled samples. The prediction results of samples without true labels are inaccurate. Therefore, the present invention constructs a prediction confidence.
[0090] Preferably, in step S4, pair-wise crossover is performed on the parent population P t to obtain a crossover population P'. Then, mutation operations are performed on the crossover sub-population P' to generate a crossover and mutation sub-population P''. The accuracy of the crossover and mutation sub-population P'' is predicted using the semi-supervised performance predictor and the confidence level corresponding to the accuracy is calculated. Then, the performance predictor is trained with P0 to mix the parent population P t and the crossover and mutation sub-population P'' into a population P p . The prediction accuracy and prediction confidence level of the population P p are used as two objectives to be optimized, and non-dominated sorting is performed.
[0091] Preferably, in step S5, according to the non-dominated sorting result, the individuals with a lower domination level are added to the new generation population P t+1 , when the sum of the number of the new generation population P t+1 and the number of individuals in a certain layer of P is greater than the population size N, the prediction confidence is used as the selection direction, and the individuals with high confidence are preferentially selected to be added to the new generation population Pt+1 until the number of the new generation population Pt+1 is equal to the population size N.
[0092] In the above solution, multi-objective optimization involves the domination relationship between solution vectors and the analysis of the Pareto front. A multi-objective optimization problem with m objectives and n decision variables can be described as:
[0093]
[0094] where, Ω⊆ℝ n is the decision space, x∈{x1,x2,……xn} is the feasible region of decision variables, and n is the dimension of variables. ℝ m is the objective space, m is the number of objective function values, and fi(x) is the i-th objective function value. When the number of objectives is 2 - 3, it is called a general multi-objective optimization problem, and when the number of objectives is 4 or more, it is called a high-dimensional multi-objective optimization problem. h u ( x ) and g v ( x ) are the inequality constraint and the equality constraint respectively, and the solutions that satisfy the constraint conditions are called feasible solutions. The absolute optimal solution needs to optimize multiple objectives simultaneously while satisfying the constraints. However, due to the mutual exclusivity of decision variables among multiple objectives, it is difficult to obtain the absolute optimal solution. Generally, an optimal solution set is obtained, and this set of optimal solution sets is generally called the Pareto front.
[0095] Preferably, during the two processes that the multi-objective selection method loops within the predefined termination conditions, when the algorithm ends, the final multi-objective optimal solution set is output, and the individuals with both high prediction accuracy and high confidence are in the final multi-objective optimal solution set.
[0096] Embodiment 2
[0097] Such as Figure 5As shown in the figure, the present invention is based on the test set NAS-Bench 201. NAS201 is a benchmark for the image classification scenario and is one of the most popular NAS benchmark tests. The cell-based search space in NASBench-201 is represented as a DAG, where nodes represent feature maps related to operation transformations and the sum of edges. Each DAG is generated by 4 nodes and 5 related operations: 1×1 convolution, 3×3 convolution, 3×3 average pooling, skip connection, and no operation. The specific structure of the neural network is as Figure 5 shown.
[0098] For the convenience of the search of the genetic algorithm and the training of the semi-supervised performance predictor, the input of the present invention is the encoded neural network structure, and an integer encoding scheme is adopted to encode this search space. The 5 operations of 1×1 convolution, 3×3 convolution, 3×3 average pooling, skip connection, and zero are mapped to the integer space of [0-4], and different prediction effects are obtained. The parameters and attributes of these operations are not considered in the encoding scheme, avoiding the preference caused by artificial setting, combining the advantages of semi-supervised learning and evolutionary algorithms, and being able to improve the performance of the performance predictor from two aspects of improving the quality of the initial samples and enhancing the performance of the basic model, realizing efficient and accurate neural architecture performance prediction.
[0099] Example 3
[0100] Table 1 shows the performance comparison with other NAS algorithms; for three datasets, the training details of each candidate architecture are provided: the validation and test accuracies of pre-training with 200 epochs on CIFAR-10, CIFAR-100, and ImageNet-16-120. Other variables in this structure are fixed. The present invention uses different random seeds to conduct 5 independent trials for each method and reports the mean and standard deviation in the table. Among them, the two algorithms TSNAS-35 and TSNAS-50 are both algorithms proposed by the present invention. 35 means that the number of integrated sub-models is 35, and similarly 50 means 50 sub-models are integrated. The bold in the table represents the best results of all algorithms. It can be seen that compared with other NAS algorithms, the semi-supervised performance predictor proposed by the present invention has the best performance on three different image datasets.
[0101] Table 1 Performance comparison with other NAS algorithms
[0102]
[0103] To more intuitively reflect the prediction effect of the model, the present invention samples at 5 epochs of 1, 5, 10, 15, and 20 during the search process, and visualizes the pareto front of this algorithm. As Figure 6As shown, from the Pareto front curve of 1 - 20 rounds, it can be seen that the proposed algorithm has a clear convergence curve, indicating that the multi - objective selection method of the present invention is effective.
[0104] Obviously, the above - mentioned embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. A multi-objective neural architecture search method based on a semi-supervised performance predictor, characterized in that Applied to image classification, the dataset to be processed is specifically an image dataset, including the following steps: S1: Encode the search space, map the operations in the search space to integers, and form samples; S2: Randomly sample N neural network structures in the samples and obtain their accuracies as the initial sample P. Before entering the genetic algorithm routine, train two ensemble KNN regression models as the initial semi-supervised predictors; S3: The initial sample P generates a candidate population P1, and its accuracy is predicted and its confidence is calculated through the semi-supervised predictor. The candidate population P1 is used as the real initial population P0, and a P0 training performance predictor is obtained; In step S3, the following steps are included: S31: Randomly sample N neural architectures and encode them, and train to obtain their accuracies; S32: Use the N encoded neural architectures as training samples to train two regression models, EnKNN1 and EnKNN2; S33: Generate N candidate samples, and EnKNN1 and EnKNN2 predict the accuracies of the candidate samples and predict the improved accuracies: S34: If the maximum predicted improvement of the sample predicted by EnKNN1 is greater than 0, then the sample with the maximum predicted improvement and its predicted accuracy are used as the new training samples of EnKNN2; If the maximum predicted improvement of the sample predicted by EnKNN1 is less than or equal to 0, then no training samples are used as predicted samples; EnKNN2 performs the same operation; At the same time, the predicted samples are deleted from the candidate samples; S35: If the training samples of EnKNN1 increase, retrain EnKNN1 with the new training samples; If the training samples of EnKNN2 increase, retrain EnKNN2 with the new training samples; S36: When the training samples of EnKNN1 increase, use EnKNN2 to calculate the confidence of the new training samples of EnKNN1, and at the same time use EnKNN1 to calculate the confidence of the new training samples of EnKNN2; S37: Loop S34, S35, S36. If P candidate samples are used as training samples and added to EnKNN1 and ENKNN2 respectively, then the P samples are all predicted, and this performance predictor is truncated and no longer predicted; If the predicted improvements of all the predicted samples of EnKNN1 and ENKNN2 are less than 0, then the prediction errors of this group of samples are large, and this performance predictor is forcibly truncated; S38: Output the prediction results of EnKNN1 and EnKNN2 for the candidate samples and the confidence of the prediction results in step S36; S39: Train the truncated EnKNN1 and EnKNN2 with the neural networks with real accuracies obtained by the original sampling; S4: Take the parental population P t Perform crossover and mutation to obtain the crossover and mutation sub-population P'', and the P0 training performance predictor takes the parental population P t Mix the parental population P and the crossover and mutation sub-population P'' and perform non-dominated sorting; S5: According to the non-dominated sorting result, individuals with a lower domination level are preferentially selected using the multi-objective selection method, and individuals with a high confidence level are added to the new generation population P. t+1 , until the number of the new generation population P t+1 is equal to the population size N, and the search is completed here.
2. The multi-objective neural architecture search method based on a semi-supervised performance predictor according to claim 1, characterized in that In step S1, the search space is the space composed of all encoded architectures.
3. The multi-objective neural architecture search method based on a semi-supervised performance predictor according to claim 2, wherein, The search space includes five functional layers, which are 1×1 convolution, 3×3 convolution, 3×3 average pooling, skip connection, and zero.
4. A multi-objective neural architecture search method based on a semi-supervised performance predictor according to claim 2, characterized in that In step S2, training the ensemble KNN regression model includes the following steps: S21: Set the training data size to M, the total number of sub-models of the ensemble model to N, the data dimension to D, and at the same time set the basic model parameters: the number of neighbors k, and the distance calculation parameters are p1 and p2 respectively; S22: Select m data where m < M, and select d dimensions where d < D to form a small training sample. Train a sub-model according to the number of neighbors k and the distance calculation parameters p1, p2 of the set KNN ensemble model. Finally, train N sub-regression models; S23: The N sub-regression models respectively make predictions on the unlabeled data, and take the average of the prediction results as the prediction result of the ensemble model.
5. A multi-objective neural architecture search method based on a semi-supervised performance predictor according to claim 1, characterized in that In step S33, the accuracy of predicting candidate samples is calculated by the following formula: 。 6. A multi-objective neural architecture search method based on a semi-supervised performance predictor according to claim 5, characterized in that In step S36, after retraining the model for the prediction samples, the average prediction deviation of all labeled samples is calculated as follows: Among them represents the number of labeled samples, f(x) represents the true label of x, and y represents the current predicted value.
7. A multi-objective neural architecture search method based on a semi-supervised performance predictor according to claim 1, characterized in that, In step S4, for the parental population P t perform pairwise crossover to obtain a crossover population P', then perform mutation operations on the crossover sub-population P' to generate a crossover and mutated sub-population P'', use the semi-supervised performance predictor to predict the accuracy of the crossover and mutated sub-population P'' and calculate the confidence level corresponding to the accuracy. After that, use P0 to train the performance predictor to mix the parental population P t and the crossover and mutated sub-population P'' into a population P p , take the prediction accuracy and prediction confidence level of the population P p as the two objectives to be optimized, and perform non-dominated sorting.
8. A multi-objective neural architecture search method based on a semi-supervised performance predictor according to claim 7, characterized in that, In step S5, according to the non-dominated sorting result, the individuals with a lower domination level are added to the new generation population P t+1 , when the sum of the number of the new generation population P t+1 and the number of individuals in a certain layer of P is greater than the population size N, the prediction confidence is used as the selection direction, and the individuals with high confidence are preferentially selected to be added to the new generation population P t+1 , until the number of the new generation population P t+1 is equal to the population size N.
9. A multi-objective neural architecture search method based on a semi-supervised performance predictor according to claim 8, characterized in that, During the two processes looped by the multi-objective selection method within the predefined termination conditions, when the algorithm ends, the final optimal solution set of multiple objectives is output. The individuals in the final optimal solution set of multiple objectives have both high prediction accuracy and high confidence.
Citation Information
Patent Citations
Deep reinforcement learning robust training method and device based on neuron coverage rate
CN113298255A
Optical fiber communication high-order chromatic dispersion prediction calculation method
CN114039659A