A similarity agent-assisted evolutionary neural architecture search method and system
By introducing a similarity proxy model based on evolutionary neural architecture search, using graph convolution networks and multi-layer perceptrons to extract architectural features, predicting the similarity between candidate architecture and benchmark architecture, solving the problem of time-consuming and high computational cost in the existing methods, and achieving efficient neural network architecture search.
Patent Information
- Application Number
- CN202510170315.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Existing evolutionary neural architecture search methods are time-consuming and computationally expensive to evaluate a large number of candidate architectures, resulting in inefficiency.
Using a similarity proxy model method, the architectural features are extracted through graph convolution networks and multi-layer perceptrons, and the proxy model is constructed using the twin neural network framework to predict the similarity between the candidate architecture and the benchmark architecture as the fitness, replacing the traditional real performance evaluation.
It significantly improves the efficiency of evolutionary neural architecture search, reduces the demand for computing resources, achieves stable and accurate fitness prediction, and can quickly generate high-performance neural network architectures that meet specific needs.
Smart Images

Figure CN119623515B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automated machine learning, and in particular to a similarity agent-assisted evolutionary neural architecture search method and system. Background Art
[0002] As a powerful information processing model, neural networks have demonstrated excellent performance and broad application prospects in computer tasks such as image classification, object detection, and natural language processing. However, the performance of neural networks depends largely on their architectural design, including key factors such as the number of layers, layer types, number of nodes, and activation functions. Reasonable network architecture design is crucial to improving model performance. In the early stages of research, the design of network architecture mainly relied on manual adjustment, which was highly dependent on the experience and knowledge accumulation of domain experts, greatly restricting the development speed of neural networks. In this context, neural architecture search, as an automated design paradigm for network architecture, has attracted widespread attention. Compared with traditional manual feature engineering, neural architecture search gets rid of the tedious manual design and adjustment process through automated search, greatly improving efficiency. The network architecture discovered through neural architecture search not only has excellent performance, but also shows strong adaptability to different tasks and data sets, thus significantly promoting the universality and flexibility of artificial intelligence technology in multiple fields.
[0003] Evolutionary algorithm is an optimization method that simulates the principles of natural selection and genetics. It uses mechanisms such as crossover, mutation, and environmental selection to iteratively optimize the population. Given the complexity of neural network architectures, neural architecture search can essentially be viewed as a type of non-convex optimization problem. Evolutionary algorithms have unique advantages in solving complex optimization problems because of their robustness to local minima and the fact that they do not require gradient information. However, evolutionary neural architecture search usually requires the evaluation of a large number of candidate architectures, and the complete evaluation of each architecture often requires time-consuming training, which can take weeks or even months. As the size of the search space continues to expand, this computational burden becomes increasingly severe, becoming the core bottleneck restricting the efficiency of neural architecture search.
[0004] In order to accelerate the evaluation of architecture performance, agent-assisted evolutionary neural architecture search has gradually become a research focus as an efficient solution. This method uses limited data to train the agent model to quickly and reliably predict the performance of the candidate architecture, thereby completing the fitness estimation in milliseconds. Existing agent models usually use low-cost approximate regression models to predict the exact performance of the candidate architecture according to the characteristics of the target task. However, due to the inconsistency of the distribution of the data to be predicted, the reliability and stability of these models are often limited. In fact, in evolutionary neural architecture search, the core goal of fitness evaluation is not to accurately determine the absolute performance of the candidate architecture, but to quickly screen out a subset of candidate architectures with potential from a huge search space. Based on this, the similarity evaluation agent model replaces the traditional fitness evaluation method by predicting the similarity between architectures, thereby realizing the priority sorting of candidate architectures. This type of method does not need to directly predict the exact performance value, but discovers high-potential architectures through relative comparison, significantly suppresses the evaluation error caused by the inconsistency of data distribution, and effectively improves the robustness of fitness prediction. Therefore, a similarity-based agent model is proposed to efficiently replace the fitness evaluation link in evolutionary neural architecture search, which has important theoretical significance and practical value. Summary of the invention
[0005] Purpose of the invention: The technical problem to be solved by the present invention is to provide an evolutionary neural architecture search method and system based on similarity agent assistance in view of the deficiencies in the prior art.
[0006] The method comprises the following steps:
[0007] Step 1: Initialize an architecture population, perform evolution, obtain the real performance of all individuals through real evaluation, and select the architecture with the best performance as the initial benchmark architecture; at the same time, save all individuals and their corresponding performance data as a training set;
[0008] Step 2: Design a similarity-based graph convolutional network transmission and aggregation strategy, build a graph neural network variant as a feature extractor, map the architecture in the search space to the feature space, and use it to learn and extract the feature representation of the architecture;
[0009] Step 3: Use the twin neural network framework to build a proxy model, use the architecture embedding features obtained by the feature extractor to calculate the similarity between architectures, and train the proxy model through a joint loss function;
[0010] Step 4: Generate candidate architectures based on the architectures in the current population, and use the proxy model to predict the similarity between the candidate architecture and the benchmark architecture as the individual fitness; retain high-potential architectures based on the fitness value, and perform real performance evaluation on the high-potential architectures, and add the evaluation results to the training set; at the same time, merge the current population with the high-performance architectures predicted and selected by the proxy model, and update the population through the environmental selection strategy;
[0011] Step 5, repeat steps 3 and 4 until the population performance converges, and finally output the global optimal architecture; the global optimal architecture can adapt to different search spaces and task requirements, especially in scenarios that require fast iteration and high-performance models, such as large-scale image recognition, real-time video processing, medical image analysis and other applications; this method can quickly generate a high-performance neural network architecture that meets specific needs, and provides strong technical support for promoting the application of artificial intelligence in practical problems.
[0012] In step 1, the reason for first evolving the population for several generations is that the proxy model based on similarity evaluation predicts fitness by measuring the similarity between the candidate architecture and the benchmark architecture. If the initial performance of the benchmark architecture is poor, it may cause the search process to fall into an inefficient state and significantly delay the optimization process. Therefore, by evolving the population for several generations, the adverse effects of the initialization of the benchmark architecture caused by the poor performance of the initial architecture can be avoided, thereby accelerating the search convergence.
[0013] In step 1, N architectures are selected from the search space as the initial population through a random sampling strategy. , N represents the population size, and its value range depends on the scale of the search space and the computing resource constraints. For smaller search spaces, N is generally set to dozens; for tasks involving more complex network structures or larger data sets, N is usually set to hundreds; then, the actual performance value of each architecture is obtained through real evaluation, and all evaluated architectures and performance values are recorded in the architecture pool , expressed as:
[0014] ,
[0015] Among them, the real evaluation refers to the complete training and verification of the neural network corresponding to each architecture on the target data set (such as the CIFAR-10 data set), and the classification accuracy of the neural network corresponding to each architecture on the verification set is recorded as the performance value; represents the record pair of the i-th architecture, Indicates architecture, Indicates The performance value corresponding to each architecture is is the total number of schemas in the schema pool; Select the one with the highest performance value The current population , and set the architecture with the best performance as the initial benchmark architecture , used for similarity evaluation and fitness prediction of subsequent proxy models.
[0016] In step 2, each architecture is represented as a graph structure data using a directed acyclic graph The form is expressed as , where the vertex Represents the layer nodes of the neural network corresponding to the architecture, each layer node corresponds to an operation (such as convolution, mean pooling operation, etc.), edge Represents the connection relationship between levels; the node feature matrix of the graph and the adjacency matrix The input is processed into a graph neural network variant consisting of a graph convolutional network and a multi-layer perceptron to realize feature extraction of the architecture, wherein the graph convolutional network is used to extract node features, and the multi-layer perceptron is used to extract structural features;
[0017] Step 2 specifically includes the following steps:
[0018] Step 2.1, capture node features based on graph convolutional network:
[0019] The traditional graph convolutional network is based on the node feature matrix of the graph structure. and the adjacency matrix As input, static propagation and aggregation rules are used to update node features. The mathematical form is:
[0020] ,
[0021] in Represents the traditional graph convolutional network The node feature matrix extracted by the layer, the initial feature matrix ; Indicates The learnable weight matrix of the layer; is a specific nonlinear item-by-item activation function; is the adjacency matrix with self-loops added; is a diagonal matrix whose elements are defined as ; is the degree matrix Middle Line The elements of the column, is the adjacency matrix Middle Line Elements of a column; Represents the adjacency matrix Normalize it, through the dimension matrix The effect of the difference in the size of the neighborhood of the balancing node on the feature propagation essentially reflects the propagation mechanism of the feature state, which can be regarded as a propagation matrix;
[0022] For the task of comparing architectural similarity, excessive attention to the interaction of irrelevant nodes may introduce noise and affect the accuracy of feature extraction. Therefore, in order to effectively learn node features, the static transmission and aggregation rules of the traditional graph convolutional network are improved by introducing a similarity measurement method to capture the correlation between node features, so that the feature aggregation process can adaptively distinguish the importance of domain information. The update formula of the graph convolutional network based on the similarity aggregation strategy is:
[0023] ,
[0024] in Is a node In the The node feature vector of the layer, is a nonlinear activation function, Representation Node In the The learnable weight matrix of the layer, Indicates Layer Slave Node To Node The transmission matrix of Through the Node Features With Node Features Cosine similarity of Multiply Node and Node adjacency Calculated, Representing node features The Euclidean norm of , Represents matrix transpose; neighbor set Representation Node In the The layers are retained during the polymerization process Neighbor index values, defined as }, Indicates participating nodes Feature aggregation index value, is the similarity threshold, according to all Calculate the mean and standard deviation of the distribution and only keep those with similarity above the threshold The neighbor nodes are aggregated and calculated;
[0025] Step 2.2, capture structural features based on multi-layer perceptron:
[0026] Traditional graph neural networks often face the problem of graph data invariance, that is, they are insensitive to structural positions in various graph operations. They usually only focus on the aggregation of node features and ignore the impact of graph structure on feature extraction. While extracting node features, the present invention introduces a multi-layer perceptron to capture the overall structural characteristics of the network architecture;
[0027] The adjacency matrix Flattened to a one-dimensional vector , and input to the multilayer perceptron to capture the structural features. Flatten represents the operation of flattening the matrix into a one-dimensional vector by row priority. The feature propagation formula of the multilayer perceptron is:
[0028] ,
[0029] in, Indicates The structural feature vector of the hidden layer output, and the initial input ; It is The weight matrix of the layer is used to implement linear transformation, For the The bias term of the layer;
[0030] In order to achieve the joint representation of node features and structural features, the output of the graph convolutional network and the multi-layer perceptron are linearly combined, and the formula is:
[0031] ,
[0032] in, It is The node feature vector matrix of the layer, It is The structural feature vector matrix of the layer, is a learnable parameter used to dynamically balance the contribution of node features and structural features to embedding learning. It is The eigenvector matrix of the layer;
[0033] Finally, the output of the feature extractor Flattened to a one-dimensional vector , That is, the captured architectural feature vector, is the total number of layers of the feature extractor. In practical applications, The value of is generally determined according to the task and data scale. For simple tasks of smaller scale, or , often used for complex tasks .
[0034] Step 3 includes:
[0035] Step 3.1, divide similarity labels and create training data sets:
[0036] For the schema pool The architecture in , generates sample pairs in pairs, and the total is Sample pairs, , which constitutes the training dataset of the proxy model ;
[0037] in, There are two samples and Similarity labels for sample pairs , poor performance Defined as , then the similarity The calculation method is:
[0038] when hour, ,
[0039] when hour, ,
[0040] Where e represents a natural constant; if Greater than the threshold, indicating that two samples The characteristics of The characteristics are not similar;
[0041] If two samples The characteristics of Represents a positive sample pair; if two samples If the characteristics of represents a negative sample pair;
[0042] Step 3.2, build the proxy model:
[0043] Using Siamese Neural Network as a Proxy Model , proxy model The method comprises two branch sub-networks whose structures and parameters are completely shared, so as to ensure that consistent feature extraction is performed on the two input architecture samples; the branch sub-network is composed of the feature extractor defined in step 2, including a graph convolutional network for extracting node features and a multi-layer perceptron for extracting structural features;
[0044] Pair the samples Two samples in and Input two branch sub-networks respectively to get Samples The eigenvector of and Samples The eigenvector of ,calculate and Cosine similarity of As the predicted fitness value, it represents the architecture and Architecture The similarity is:
[0045] ;
[0046] Step 3.3, train the proxy model:
[0047] Using the training dataset , based on the joint loss function Calculate the loss value, update the trainable parameters of the proxy model through the back-propagation algorithm, and repeat step 3.3 until the proxy model converges; It is a learnable parameter used to adjust the contribution ratio of regression loss and similarity loss in the total loss;
[0048] Among them, the similarity loss , represents the total number of sample pairs, To adjust the threshold of the distance between negative sample pairs, it is used to ensure that the feature vectors between negative sample pairs have sufficient discrimination; represents the maximum value function; the similarity loss function promotes the proxy model to learn discriminative feature embedding by optimizing the closeness of similar sample pairs and the separation of dissimilar sample pairs;
[0049] Considering that the main purpose of the proxy model is to replace the fitness evaluation link, it is hoped that similar architectures in the feature space should also show corresponding superiority in performance. Therefore, it is necessary to enhance the coupling between feature similarity and performance similarity. The present invention introduces a performance-based regression loss function in the training of the proxy model. , regression loss ,in Indicates that the similarity with the current architecture is higher than the threshold Before A set of pre-trained samples, , is the similarity threshold, according to all Calculate the mean and standard deviation of the distribution settings; It is The design of the regression loss function can effectively improve the correlation between feature similarity and performance similarity, so that the model can perform better in performance prediction tasks.
[0050] Finally, the paired training datasets Importing Proxy Models , the rational model calculates the joint loss function The loss value is calculated and the trainable parameters of the model are updated using the back-propagation algorithm. This process is repeated until the model converges.
[0051] Step 4 includes:
[0052] Step 4.1, using a simple path-based crossover mutation operator to generate candidate architectures;
[0053] Step 4.2, update the population based on the proxy model prediction.
[0054] An efficient and reasonable architecture generation strategy should maximize the exploration of diverse network architectures under limited computing resources, thereby increasing the probability of discovering high-performance architectures. In this strategy, the crossover mutation operator is a key technical means, which injects diversity and innovation into the population by integrating the characteristic elements of different candidate architectures and introducing a random mutation mechanism. This operator draws on the principles of gene crossover and mutation in biological genetics, aiming to generate new architectures with potential high performance;
[0055] The present invention converts the neural network architecture into graph structure data, and proposes a crossover operator based on a simple path according to the decoupling characteristics between node types and structural connections in the graph structure.
[0056] Step 4.1 includes:
[0057] Step 4.1.1, based on the fitness values of all architecture individuals in the population, i.e., the true performance values, a probability allocation strategy is adopted, in which high-performance architectures are given higher selection probabilities, while low-performance architectures are given lower selection probabilities. Two architectures are randomly selected from the current population as parent architectures through a roulette strategy, and are represented by directed acyclic graphs respectively, and the nodes are numbered according to topological sorting;
[0058] Step 4.1.2, perform depth-first traversal on the directed acyclic graph of the parent architecture respectively to generate all simple paths from the input node to the output node, thereby forming an independent path set; then, based on the random selection strategy, randomly sample several paths from the independent path sets of the two parent architectures to form the path set of the child architecture; this process introduces randomness to ensure that the generated new architecture has unique characteristics, which not only enhances the diversity of the architecture, but also helps to avoid the search from falling into the local optimal solution;
[0059] Step 4.1.3, merge the nodes in the child architecture path set according to the node number, that is, treat the nodes with the same number as the same node, unify the connection relationship between their in-edges and out-edges, and form a complete topological structure, thereby obtaining a directed acyclic graph representation of the child architecture; the directed acyclic graph of the child architecture has a completely new topological structure and may contain connection patterns and node combinations that do not appear in the parent architecture;
[0060] Step 4.1.4, extract the operation type sequence from the two parent architectures respectively: traverse each node in order of node number and record the operation type corresponding to the node to form two parent architecture operation type sequences; then, initialize an empty sequence with the same size as the parent architecture operation type sequence as the child architecture operation type sequence, which is used to store the operation type of each operation of the child architecture; fill the first and last positions of the empty sequence with input and output operations respectively; for the remaining positions in the empty sequence, traverse in turn, and randomly select an operation type of the corresponding position in the parent operation type sequence at each position, and fill the operation type into the operation type sequence of the child architecture; this ensures that the child inherits the beneficial features of the parent and has new potential innovations;
[0061] Step 4.1.5, based on the directed acyclic graph representation of the child architecture and the corresponding operation type sequence, a complete child architecture is formed; the newly formed child architecture is subjected to constraint check, and if the constraints defined in the search space are not met, such as the in-and-out degree of the node does not meet the preset limit, the current architecture is discarded; if the constraints are met, the current architecture is retained; this process ensures that the generated architecture meets the basic requirements of the search space, while improving the effectiveness and reliability of the generated results;
[0062] Step 4.1.6, repeat steps 4.1.2 to 4.1.5 until the predefined quantity is generated A set of candidate architectures .
[0063] Step 4.2 includes:
[0064] Step 4.2.1: Set the candidate architecture and the benchmark architecture Input into the trained proxy model together , generate a set of feature vectors for candidate architectures and the feature vector of the baseline architecture ;
[0065] Step 4.2.2, calculate the similarity between the candidate architecture and the benchmark architecture based on the cosine similarity between the feature vectors:
[0066] ,
[0067] in, It is The feature vector of the sample and the feature vector of the baseline architecture The cosine similarity between ; for each candidate architecture, Indicates Architecture The individual prediction performance value of
[0068] Step 4.2.3, sort the candidate architectures in descending order according to their individual prediction performance values, and select the top candidate architectures, conduct real evaluation, and obtain a set of architectures with known accuracy }, the architecture collection } added to the training set and updated to ;
[0069] Step 4.2.4: Collect the high-performance architectures predicted by the proxy model } and the current population Merge to get a new architecture set }; According to the actual performance value of the architecture , select before The best performing architecture, update the population .
[0070] The present invention also provides an evolutionary neural architecture search system based on similarity agent assistance, comprising:
[0071] Feature extractor based on graph neural network variants: used to extract the latent features of neural network architectures and map complex neural network architecture representations into a low-dimensional, compact feature space, thereby efficiently representing the differences and similarities between architectures;
[0072] Proxy model based on similarity evaluation strategy: used to learn the similarity between architectures; by predicting the similarity value between the candidate architecture and the benchmark architecture as the fitness indicator of the candidate architecture, it can screen out architectures with potential high performance;
[0073] Architecture generator based on simple path decomposition: used to efficiently generate a large number of candidate new architectures, maximize the diversity of the search space, fully explore potential high-performance neural network architectures, and thus accelerate the evolutionary search process.
[0074] The present invention also provides an electronic device, characterized in that it includes a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the described method.
[0075] The method of the present invention constructs a new proxy model, which significantly improves the efficiency of evolutionary neural architecture search by directly predicting the performance of candidate architectures without fully training them. The proxy model based on similarity evaluation uses existing architecture performance data to quickly estimate the potential performance of new architectures by measuring the feature similarity between new architectures and known high-performance architectures, thereby achieving stable and accurate fitness prediction for the proxy-assisted evolutionary search process. In addition, the present invention regards the neural network architecture as a graph structure, introduces graph neural networks to learn feature representation, and proposes a novel transmission and aggregation mechanism based on similarity clustering to optimize the embedding expression ability of graph neural networks. This mechanism can adaptively enhance the importance of neighborhood information in architectural features, thereby effectively improving the accuracy and generalization of architectural representation. In order to further enhance the coupling between feature similarity and performance similarity, the present invention designs a joint loss function to map architectures with similar performance to adjacent positions in the feature space. Through this strategy, the proxy model significantly improves the prediction efficiency and accuracy of candidate architecture performance, bringing important breakthroughs to the field of evolutionary neural architecture search from a new perspective of similarity evaluation. Compared with traditional agent-assisted algorithms, the evolutionary neural architecture search method proposed in this invention shows higher prediction accuracy and better search efficiency, injecting new technical impetus into the field of neural architecture search.
[0076] The beneficial effects of the present invention are as follows: the present invention proposes an evolutionary neural architecture search method based on similarity agent assistance, aiming to optimize the efficiency and performance of evolutionary neural architecture search. The method models the neural network architecture as graph structure data, and adopts a graph neural network variant as a feature extractor to extract node information and topological structure features of the architecture. In order to further improve the feature extraction capability, the present invention innovatively designs a transmission and aggregation mechanism based on the similarity clustering principle, so that the graph neural network can adaptively identify and strengthen the importance of key information in the field, thereby optimizing the expressive ability of architecture feature embedding. Based on the above characteristics, the present invention constructs a proxy model based on similarity evaluation, which quickly and reliably predicts architecture performance from the perspective of the similarity between the candidate architecture and the benchmark architecture. Under the constraint of limited computing resources, the proxy model can effectively predict the potential performance of large-scale candidate architectures, thereby replacing the real performance evaluation link with high computational cost in evolutionary neural architecture search. In addition, the present invention designs a joint loss function for the proxy model, which not only effectively improves the training efficiency of the model by strengthening the coupling relationship between performance similarity and feature similarity, but also significantly improves the accuracy and stability of performance prediction. The innovative method of the present invention enhances the adaptability between evolutionary algorithms and neural architecture search tasks, significantly reduces the demand for computing resources while ensuring high efficiency, and lays a solid foundation for the widespread promotion of neural architecture search technology in practical applications. Compared with existing methods, the present invention shows excellent practical value and broad application potential, bringing breakthrough progress to the field of evolutionary neural architecture search. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 It is the overall framework diagram of the present invention.
[0078] Figure 2 It is a schematic diagram of the feature extractor structure of the present invention.
[0079] Figure 3 It is a schematic diagram of the proxy model structure of the present invention.
[0080] Figure 4 This is the effect comparison of feature extractors Figure 1 .
[0081] Figure 5 This is the effect comparison of feature extractors Figure 2 .
[0082] Figure 6 This is the effect diagram of the feature extractor.
[0083] Figure 7 It is the effect of the proxy model Figure 1 .
[0084] Figure 8 It is the effect of the proxy model Figure 2 .
[0085] Fig. 9 yes Figure 8 A partial enlargement of the results. DETAILED DESCRIPTION
[0086] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0087] like Figure 1 As shown, an embodiment of the present invention provides an evolutionary neural architecture search method based on similarity agent assistance, comprising the following steps:
[0088] Step 1: Randomly sample in the search space of NAS-Bench-101, a benchmark dataset for evaluating neural architecture search methods. The architecture is used as the initial population , Represents the population size, and its value range depends on the scale of the search space and the computing resource constraints. For a smaller search space, it is generally set to For tasks involving more complex network structures or larger data sets, it is usually necessary to set Then, the performance value of each architecture is obtained through real evaluation. All evaluated architectures and their performance values are recorded in the architecture pool, which is recorded as:
[0089] ,
[0090] Among them, the real evaluation refers to the complete training and verification of the neural network corresponding to each architecture on the target dataset (such as the CIFAR-10 dataset), and recording its classification accuracy on the verification set as the performance value; Indicates The record pairs of the architecture, Indicates architecture, Indicates The performance value corresponding to each architecture is is the total number of evaluated architectures in the architecture pool;
[0091] From the schema pool Select the one with the highest performance value The current population , and the architecture with the best performance is set as the initial benchmark architecture, denoted as .
[0092] Step 2: Build Figure 2The architecture feature extractor shown in the figure consists of a graph convolutional network for extracting node features and a multi-layer perceptron for extracting structural features. By representing the neural network architecture as graph structured data, a directed acyclic graph is used. The form is expressed as , where the vertex Represents the layer nodes of the neural network corresponding to the architecture, each layer node corresponds to an operation (such as convolution, mean pooling operation, etc.), edge Represents the connection relationship between levels; the node feature matrix of the graph and the adjacency matrix The input is processed into a graph neural network variant consisting of a graph convolutional network for extracting node features and a multi-layer perceptron for extracting structural features to achieve architectural feature extraction.
[0093] Step 2.1: To effectively learn node features, this paper introduces a similarity measurement method to improve the static transmission and aggregation rules of traditional graph convolutional networks to capture the correlation between neighborhood features. The update formula of the graph convolutional network based on the similarity aggregation strategy is:
[0094] ,
[0095] in Is a node In the The node feature vector of the layer, is a nonlinear activation function, Representation Node In the The learnable weight matrix of the layer, Indicates Layer Slave Node To Node The transmission matrix of Node Features With Node Features Cosine similarity of Multiply Node and Node adjacency calculate, Representing node features The Euclidean norm of , Represents matrix transpose; neighbor set Representation Node In the The layers are retained during the polymerization process Neighbor index values, defined as }, Participating Node Feature aggregation index value, which is is the similarity threshold, according to all Calculate the mean and standard deviation of the distribution and only keep those with similarity above the threshold The neighbor nodes of the node are aggregated and calculated.
[0096] Step 2.2:
[0097] The adjacency matrix Flattened to a one-dimensional vector , and input to the multilayer perceptron to capture the structural features. Flatten represents the operation of flattening the matrix into a one-dimensional vector by row priority. The feature propagation formula of the multilayer perceptron is:
[0098] ,
[0099] in, Represents the multilayer perceptron The structural feature vector of the hidden layer output, and the initial input ; It is The weight matrix of the layer is used to achieve linear transformation of features. It is Layer bias
[0100] In order to achieve the joint representation of node features and structural features, the output of the graph convolutional network and the multi-layer perceptron are linearly combined, and the formula is:
[0101] ,
[0102] in, It is The node feature vector matrix of the layer, It is The structural feature vector matrix of the layer, is a learnable parameter used to dynamically balance the contribution of node features and structural features to embedding learning. It is The eigenvector matrix of the layer.
[0103] Finally, the output of the feature extractor Flattened to a one-dimensional vector , is the captured architectural feature vector, is the total number of layers of the feature extractor. In practical applications, The value of is generally determined according to the task and data scale. For simple tasks of smaller scale, or , often used for complex tasks .
[0104] Step 3: Based on the architecture pool Generate paired training datasets, annotate similarity labels, and perform supervised training on the proxy model.
[0105] Step 3.1: From the schema pool In the two-by-two combination samples, we can get Sample pairs are used as training data sets for the proxy model ;
[0106] in, There are two samples and Similarity labels for sample pairs , poor performance Defined as , then the similarity The calculation method is:
[0107] when hour, ,
[0108] when hour, ,
[0109] Where e represents a natural constant; if Greater than the threshold, indicating that two samples The characteristics of The characteristics are not similar;
[0110] If two samples The characteristics of Represents a positive sample pair; if two samples If the characteristics of represents a negative sample pair;
[0111] Step 3.2: Based on the feature extractor, build Figure 3 The proxy model shown The proxy model is based on a twin neural network and consists of two identical branch sub-networks. The sub-networks are composed of the feature extractors in step 2, and the two sub-networks share weights and parameters to ensure that the same features are extracted from the two input samples.
[0112] Step 3.3, train the proxy model Using the training dataset , based on the designed joint loss function Calculate the loss value, update the trainable parameters of the proxy model through the back-propagation algorithm, and repeat this process until the model converges;
[0113] Among them, the similarity loss ,for Sample pairs, is a similarity label, when two samples are identified as similar architectures represents a positive sample pair, otherwise represents a negative sample pair. Represents the cosine similarity of the feature vectors calculated by the feature extractor. To adjust the threshold of the distance between negative sample pairs, it is used to ensure that the feature vectors between negative sample pairs have sufficient discrimination;
[0114] Regression Loss ,in Indicates that the similarity with the current architecture is higher than the threshold Before A set of pre-trained samples, is the similarity threshold, according to all Compute the mean and standard deviation of a distribution. It is The actual performance value of the samples.
[0115] Step 4: Use a simple path-based crossover mutation operator to generate a large number of candidate architectures, use the proxy model to predict the performance of the candidate architectures, and screen out potentially high-performance architectures to update the population.
[0116] Step 4.1: Figure 4 As shown, based on the current population, two architectures are selected as parent architectures using a roulette strategy, and are represented by directed acyclic graphs, and the nodes are numbered according to topological sorting. Subsequently, the directed acyclic graphs of the parent architectures are traversed in depth first to generate all simple paths from the input nodes to the output nodes, thereby forming an independent path set;
[0117] Secondly, a random selection strategy is adopted to randomly sample several paths from the independent path sets of the two parents to form the path set of the offspring architecture; and the nodes in the offspring architecture path set are merged according to the node numbers to obtain a directed acyclic graph representation of the offspring architecture.
[0118] Next, extract the operation type sequences from the two parent architectures respectively, that is, traverse each node in order according to the node number and record the operation type corresponding to the node to form two parent architecture operation type sequences; then, initialize an empty sequence with the same size as the parent architecture operation type sequence as the child architecture operation type sequence, which is used to store the operation type of each operation of the child architecture; fill the first and last positions of the empty sequence with input and output operations respectively; for the remaining positions in the empty sequence, traverse in turn, and randomly select an operation type of the corresponding position in the parent operation type sequence at each position, and fill it into the operation type sequence of the child architecture;
[0119] Finally, based on the DAG representation of the child architecture and its corresponding sequence of operation types, a complete child architecture can be formed. The newly formed architecture is checked for constraints. If the constraints of the search space are not met, such as the in-and-out degree of the node does not meet the preset limit, the current architecture is discarded. If it is met, it is retained. Repeat the above steps until a predefined number of Candidate architectures .
[0120] Step 4.2: Update the population. Based on the current training set Medium performance value Optimal Architecture , setting it as the current optimal baseline architecture , the proxy model predicts the candidate architecture and The similarity is used as the individual prediction performance value, so as to screen and retain a part of the architecture with potential high performance value.
[0121] Step 5, repeat steps 3 and 4 until the population performance converges, and finally output the global optimal architecture; the global optimal architecture can adapt to different search spaces and task requirements, especially in scenarios that require fast iteration and high-performance models, such as large-scale image recognition, real-time video processing, medical image analysis and other applications. This method can quickly generate a high-performance neural network architecture that meets specific needs, providing strong technical support for promoting the application of artificial intelligence in practical problems.
[0122] NAS-Bench-101 is a popular benchmark dataset for neural architecture search algorithms, providing a compact and diverse search space. The dataset contains 423,624 unique deep neural network architectures, which are used for image classification tasks on the CIFAR-10 dataset. These architectures are fully trained and verified, and the classification accuracy of the architectures on the training set, validation set, and test set is provided as the performance result of the architecture. The present invention conducts experiments based on the NAS-Bench-101 dataset to verify the effectiveness of the proposed method.
[0123] In the performance verification of the feature extractor, the present invention conducted a comparative experiment with the One-Hot encoding representation method and the traditional graph neural network representation method. For all the architectures in the NAS-Bench-101 dataset, the above three representation methods were used to extract their features, and the extracted high-dimensional feature space was reduced by t-SNE, and the distribution of the architecture in the feature space was visualized, such as Figure 5 , Figure 6 , Figure 7 As shown, the color scale on the right represents the ranking of the architecture based on its classification accuracy on the CIFAR-10 dataset. The redder the architecture, the higher its ranking and the better its performance. Figure 5 is the feature space of the One-Hot representation method, Figure 6 is the feature space of the traditional graph neural network representation method, Figure 7 This is the feature space obtained by the feature extractor proposed in the present invention.
[0124] The results show that the clustering performance of the feature extractor of the present invention in the feature space is significantly better than the previous two methods, and its feature space distribution shows a high degree of regularity. Specifically, the architectures with higher performance rankings are concentrated on the left side of the feature space and include the architectures with the best performance in the search space; observing from right to left, the performance ranking of the architecture gradually increases, and architectures with similar performance are significantly mapped to adjacent positions, while architectures with large performance differences maintain a greater distance. The feature extractor of the present invention can effectively map architectures with similar performance to the neighborhood of the feature space, thereby enhancing the distinguishing ability of the feature representation. This feature learning method provides strong support for the subsequent screening of architectures with excellent performance by guiding the proxy model through the similarity between architectures.
[0125] In testing the proxy model of the present invention, the proxy model is used to predict the top 1000 architectures with the best performance in the entire search space of NAS-Bench-101, and the actual performance of these architectures on the validation set and test set is visualized through scatter plots, such as Figure 8 As shown. The blue dots represent all architectures in the search space, and the red dots represent the 1000 architectures selected by the proxy model. For further observation, Fig. 9 A partial zoom of the left image is provided, focusing on the architectures whose performance exceeds 90% on both the validation set and the test set.
[0126] It can be found that the classification accuracy of the architecture selected by the proxy model on the validation set of the image classification task of the CIFAR-10 dataset is concentrated above 93.5%, and the performance on the test set is also concentrated above 93%. This result not only covers almost all the top-ranked architectures in the NAS-Bench-101 search space, but also accurately identifies the global optimal architecture, which strongly proves that the proxy model has shown excellent performance prediction ability and screening accuracy on both the validation set and the test set of the image classification task, and further verifies that the similarity proxy-assisted evolutionary neural architecture search method proposed in the present invention can effectively search for high-performance neural network architectures, and when the searched architecture is used for image classification tasks, it has extremely high benefits in terms of accuracy, stability and generalization ability.
[0127] This method can adapt to different search spaces and task requirements. By dynamically evaluating the similarity and performance potential of architectural features, it effectively avoids the problems of inefficient random search and over-reliance on manual experience in traditional manual architecture design, showing wide applicability. Especially in scenarios that require fast iteration and high-performance models, such as large-scale image recognition, real-time video processing, and medical image analysis, this method can quickly generate high-performance neural network architectures that meet specific needs, providing strong technical support for promoting the application of artificial intelligence in practical problems.
[0128] The similarity agent-assisted evolutionary neural architecture search method provided by the present invention can effectively identify efficient neural network architectures. This efficient architecture search and optimization process significantly reduces the time cost of manual parameter adjustment and architecture design, while providing a more generalized and efficient architecture. In practical applications, this search method is not only applicable to image classification tasks, but can also be extended to multiple deep learning fields. For example, in the search of large models, this method can quickly locate the model architecture with the best performance, significantly reduce the search time and computational cost, and provide efficient design ideas for training large language models (such as ChatGPT, Transformer, etc.); in the construction of diffusion models, this search method can help discover more efficient denoising modules and diffusion process architectures, thereby further improving the quality and stability of generated tasks. In addition, the search method is also applicable to complex tasks including target detection, speech recognition, natural language processing, etc., and the performance of the model architecture is further improved by searching the parameters of the architecture. The method proposed in the present invention not only has a strong generalization ability in specific tasks, but also has great potential in designing efficient deep learning models, helping to promote the development of artificial intelligence models in a more efficient and intelligent direction.
[0129] The present invention provides a method and system for searching an evolutionary neural architecture based on similarity agent assistance. There are many methods and ways to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention. All components not specified in this embodiment can be implemented by existing technologies.
Claims
1. A similarity agent-assisted evolutionary neural architecture search method, characterized in that: The following steps are involved: Step 1: Initialize an architecture population, perform evolution, obtain the real performance of all individuals through real evaluation, and select the architecture with the best performance as the initial benchmark architecture; At the same time, all individuals and their corresponding performance data are saved as a training set; Step 2: Design a similarity-based graph convolutional network transmission and aggregation strategy, build a graph neural network variant as a feature extractor, map the architecture in the search space to the feature space, and use it to learn and extract the feature representation of the architecture; Step 3: Use the twin neural network framework to build a proxy model, use the architecture embedding features obtained by the feature extractor to calculate the similarity between architectures, and train the proxy model through a joint loss function; Step 4: Generate candidate architectures based on the architectures in the current population, and use the proxy model to predict the similarity between the candidate architecture and the benchmark architecture as the individual fitness; retain high-potential architectures based on the fitness value, and perform real performance evaluation on the high-potential architectures, and add the evaluation results to the training set; at the same time, merge the current population with the high-performance architectures predicted and selected by the proxy model, and update the population through the environmental selection strategy; Step 5: Repeat steps 3 and 4 until the population performance converges, and finally output the global optimal architecture; Applying the globally optimal architecture to an image recognition task; In step 1, N architectures are selected from the search space as the initial population P0 through a random sampling strategy, where N represents the population size. Subsequently, the true performance value of each architecture is obtained through real evaluation, and all evaluated architectures and performance values are recorded in the architecture pool D. train , expressed as: Among them, the real evaluation refers to the complete training and verification of the neural network corresponding to each architecture on the target data set, and the classification accuracy of the neural network corresponding to each architecture on the verification set is recorded as the performance value; (Arc i ,Y i ) represents the record pair of the i-th architecture, Arc i represents the i-th architecture, Y i represents the performance value corresponding to the i-th architecture, Num is the total number of architectures in the architecture pool; train Select the N architectures with the highest performance values as the current population The architecture with the best performance is set as the initial benchmark architecture Arc best ; In step 2, each architecture is represented as graph structure data in the form of a directed acyclic graph G represented as G={V,E}, where the vertex V represents the hierarchical node of the neural network corresponding to the architecture, each hierarchical node corresponds to an operation, and the edge E represents the connection relationship between the hierarchies; the node feature matrix X and the adjacency matrix A of the graph are input into a graph neural network variant composed of a graph convolutional network and a multilayer perceptron for processing to realize feature extraction of the architecture, wherein the graph convolutional network is used to extract node features, and the multilayer perceptron is used to extract structural features; Step 2 specifically includes the following steps: Step 2.1, capture node features based on graph convolutional network: By introducing the similarity measurement method, the static transmission and aggregation rules of the graph convolutional network are improved to capture the correlation between node features. The update formula of the graph convolutional network based on the similarity aggregation strategy is: in is the node feature vector of node i in layer l, σ is the nonlinear activation function, W i (l) represents the learnable weight matrix of node i in layer l, represents the transmission matrix from node j to node i in layer l, matrix Through the i-th node feature h i With the j-th node feature h j Cosine similarity of Multiply the adjacency relationship A between node i and node j ij Calculated, ∥h i ∥2 represents node feature h i The Euclidean norm of , Represents matrix transpose; neighbor set It represents the k neighbor index values retained by node i in layer l during the aggregation process, defined as i k represents the kth index value of the feature aggregation of node i, θ∈[0,1] is the similarity threshold, according to all The mean and standard deviation of the distribution are calculated, and only neighbor nodes with similarity higher than the threshold θ are retained for aggregation calculation; Step 2.2, capture structural features based on multi-layer perceptron: Flatten the adjacency matrix A into a one-dimensional vector x A = flatten(A) and input to the multilayer perceptron to capture the structural features. Flatten represents the operation of flattening the matrix into a one-dimensional vector by row priority. The feature propagation formula of the multilayer perceptron is: in, Represents the structural feature vector of the output of the lth hidden layer, and the initial input is the weight matrix of the lth layer, used to implement linear transformation, b (l) is the bias term of the lth layer; In order to achieve the joint representation of node features and structural features, the output of the graph convolutional network and the multi-layer perceptron are linearly combined, and the formula is: Among them, H (l) is the node feature vector matrix of layer l, is the structural feature vector matrix of the lth layer, β∈[0,1] is a learnable parameter used to dynamically balance the contribution of node features and structural features to embedding learning, is the eigenvector matrix of the lth layer; Finally, the output of the feature extractor Flattened to a one-dimensional vector Emb is the captured architectural feature vector, and L is the total number of layers of the feature extractor; Step 3 includes: Step 3.1, divide similarity labels and create training data sets: For architecture pool D train The architecture in , generates sample pairs in pairs, and a total of Sum sample pairs are obtained, Sum = (Num-1) × (Num) / 2, which constitutes the training data set of the proxy model Among them, y ij There are two sample Arc i , and Arc j Similarity labels, for the sample pair {(Arc i ,Y i ),(Arc j ,Y j )}, the performance difference ΔY is defined as ΔY = Y i -Y j , then the calculation method of similarity is: When ΔY ≥ 0, When ΔY<0, Where e represents a natural constant; if similarity is greater than the threshold, it means that the two sample Arc i ,Arc j The characteristics of the two samples are similar; otherwise, it means that the i ,Arc j The characteristics are not similar; If two samples Arc i ,Arc j The characteristics of y are similar, then ij =1 indicates a positive sample pair; if two samples Arc i ,Arc j The characteristics of y are not similar, then ij =0 indicates a negative sample pair; Step 3.2, build the proxy model: A twin neural network is used as the proxy model Model, and the proxy model Model includes two branch sub-networks whose structures and parameters are completely shared; the branch sub-network is composed of the feature extractor defined in step 2, including a graph convolutional network for extracting node features and a multi-layer perceptron for extracting structural features; The sample pair {(Arc i ,Y i ),(Arc j ,Y j )} in two samples (Arc i ,Y i ) and (Arc j ,Y j ) are input into the two branch sub-networks respectively, and the i-th sample (Arc i ,Y i )’s eigenvector Emb i and the jth sample (Arc j ,Y j )’s eigenvector Emb j , calculate Emb i and Emb j The cosine similarity D ij As the predicted fitness value, it represents the architecture Arc i and Architecture Arc j The similarity: Step 3.3, train the proxy model: Using the training data set Data, based on the joint loss function Calculate the loss value, update the trainable parameters of the proxy model through the back-propagation algorithm, and repeat step 3.3 until the proxy model converges; λ∈[0,1] is a learnable parameter; Among them, the similarity loss M represents the total number of sample pairs, margin is the threshold for adjusting the distance between negative sample pairs; max represents the maximum value function; Regression Loss in represents the index value set of the first k pre-trained samples whose similarity with the current architecture is higher than the threshold τ, τ is the similarity threshold; Y i is the true performance value of the i-th sample.
2. The method according to claim 1, characterized in that Step 4 includes: Step 4.1, using a simple path-based crossover mutation operator to generate candidate architectures; Step 4.2, update the population based on the proxy model prediction.
3. The method according to claim 2, characterized in that Step 4.1 includes: Step 4.1.1, randomly select two architectures from the current population as parent architectures through the roulette strategy, represent them with directed acyclic graphs, and number the nodes according to topological sorting; Step 4.1.2, perform depth-first traversal on the directed acyclic graph of the parent architecture respectively to generate all simple paths from the input node to the output node, thereby forming an independent path set; then, based on the random selection strategy, randomly sample paths from the independent path sets of the two parent architectures respectively to form the path set of the child architecture; Step 4.1.3, merge the nodes in the child architecture path set according to the node number, regard the nodes with the same number as the same node, unify the connection relationship between the incoming edge and the outgoing edge, form a complete topological structure, and thus obtain the directed acyclic graph representation of the child architecture; Step 4.1.4, extract the operation type sequence from the two parent architectures respectively: traverse each node in order of node number and record the operation type corresponding to the node to form two parent architecture operation type sequences; then, initialize an empty sequence with the same size as the parent architecture operation type sequence as the child architecture operation type sequence, which is used to store the operation type of each operation of the child architecture; fill the first and last positions of the empty sequence with input and output operations respectively; for the remaining positions in the empty sequence, traverse in turn, and randomly select an operation type of the corresponding position in the parent operation type sequence at each position, and fill the operation type into the operation type sequence of the child architecture; Step 4.1.5, based on the directed acyclic graph representation of the child architecture and the corresponding operation type sequence, a complete child architecture is formed; the constraints of the newly formed child architecture are checked, and if the constraints defined in the search space are not satisfied, the current architecture is discarded; if they are satisfied, the current architecture is retained; Step 4.1.6, repeat steps 4.1.2 to 4.1.5 until a set of candidate architectures with a predefined number S is generated.
4. The method according to claim 3, characterized in that: Step 4.2 includes: Step 4.2.1: Combine the candidate architecture set Inds and the benchmark architecture Arc best Input them together into the trained proxy model Model to generate a set of feature vectors of candidate architectures and the feature vector Emb of the baseline architecture best ; Step 4.2.2, calculate the similarity between the candidate architecture and the benchmark architecture based on the cosine similarity between the feature vectors: Among them, cosine(Emb i ,Emb best ) is the feature vector Emb of the i-th sample i and the feature vector Emb of the baseline architecture best cosine similarity between; for each candidate architecture, cosine(Emb i ,Emb best ) represents the i-th architecture Arc i The individual prediction performance value of Step 4.2.3: Sort the candidate architectures in descending order according to their individual prediction performance values, select the top K candidate architectures, conduct real evaluation, and obtain a set of architectures with known accuracy. The schema collection Add to the training set and update to Step 4.2.4: Collect the high-performance architectures predicted by the proxy model With the current population Merge to get a new architecture set According to the actual performance value Y of the architecture i , select the top N best performing architectures and update the population 5. The similarity agent-assisted evolutionary neural architecture search system implemented according to the method of any one of claims 1 to 4, characterized in that: include: Feature extractor based on graph neural network variants: used to extract the latent features of neural network architectures and map complex neural network architecture representations into a low-dimensional, compact feature space, thereby efficiently representing the differences and similarities between architectures; Proxy model based on similarity evaluation strategy: used to learn the similarity between architectures; by predicting the similarity value between the candidate architecture and the benchmark architecture as the fitness indicator of the candidate architecture, it can screen out architectures with potential high performance; Architecture generator based on simple path decomposition: used to efficiently generate a large number of candidate new architectures, maximize the diversity of the search space, fully explore potential high-performance neural network architectures, and thus accelerate the evolutionary search process.
6. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Neural network architecture evaluation method based on attribute graph optimization
CN110232434A
Evolutionary neural architecture search method and system based on performance level agent assistance
CN118917353A