Particle swarm neural architecture search method and system based on ternary comparison agent assistance

By using the ternary comparison agent-assisted particle swarm algorithm in neural network architecture search, the problem of expensive training data of the proxy model and insufficient prediction accuracy is solved, and efficient and accurate neural network architecture search is achieved.

CN119204082BActive Publication Date: 2025-05-16NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411682871.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-05-16
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

In the existing neural network architecture search methods, training of proxy models requires a large amount of real performance data, and the prediction accuracy is difficult to achieve absolute accuracy, resulting in the impact of search direction.

Method used

The particle swarm neural architecture search method based on ternary comparison agent assisted is adopted. By initializing particle swarm, the architecture performance is truly evaluated, the triple sample is constructed, the agent model is trained, and the particle swarm algorithm and agent model are used for double comparison and relaxation operations to optimize architecture search.

Benefits of technology

It significantly reduces the time cost of architecture evaluation, improves search efficiency and accuracy, reduces dependence on real evaluations, and ensures that excellent architectures are effectively retained in the architecture pool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119204082B_ABST
    Figure CN119204082B_ABST
Patent Text Reader

Abstract

The present invention provides a particle swarm neural architecture search method and system assisted by a ternary comparison agent. The method uses a graph neural network to effectively extract the operation information and topological structure of the network architecture, and performs an evolutionary search of the particle swarm. In order to solve the problem of high cost of architecture evaluation in the current field of neural network structure search, the present invention proposes a dual comparison agent model based on triples. By constructing a training data set in the form of triples, this method can achieve data enhancement, significantly increase the number of training samples compared to traditional methods, and effectively solve the problem of insufficient training samples for agent models in traditional methods. In addition, by judging the pros and cons of two architectures through a dual comparison mechanism, this method can more effectively ensure that excellent architectures are retained in the architecture pool. Therefore, the agent model can replace the real evaluation process to a greater extent, thereby significantly reducing the time cost of architecture evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automated machine learning, and in particular to a particle swarm neural architecture search method and system based on ternary comparison agent assistance. Background Art

[0002] Deep learning has played a key role in the field of machine learning. It has achieved success in various areas of computer vision, including but not limited to image classification, object detection, boundary detection, semantic segmentation, and pose estimation. These achievements are primarily attributed to deep neural networks, which automate feature engineering. The architecture of a deep neural network is typically tailored for a specific task, and its associated weights are acquired through a learning process. To achieve optimal performance on the target task, both the network structure and weights must be optimized. However, designing an effective deep neural network architecture is a complex process that requires a deep understanding of both deep learning and the specific problem, as well as lengthy tuning of numerous hyperparameters. These task-specific network architectures often do not generalize well to other application domains. For example, a network architecture designed for image classification may perform poorly on object detection tasks. Furthermore, network architectures must be designed for diverse deployment environments within limited computing resources (such as latency, memory, and floating-point operations). Manually designing network architectures is inefficient and fails to explore all possible options.

[0003] Due to the inefficiency of manually designed network architectures, neural network architecture search has begun to attract widespread attention. As an automated machine learning method, neural network architecture search focuses on automatically searching for effective neural network architectures for specific tasks on a given dataset. While neural network architecture search has shown great potential for automated neural network design, the significant cost of architecture search remains a major obstacle to its development. Evaluating the performance of candidate architectures is the primary factor affecting search cost, as each candidate requires full training and validation of the network architecture. Early researchers employed low-fidelity estimation methods such as early stopping, reducing input image resolution, using subsets of the full training set, and reducing the number of network channels. However, these methods can lead to inaccurate evaluation of candidate networks, especially for complex and large network architectures. Later, weight inheritance methods were proposed, allowing network architectures to skip weight training. However, these methods rely heavily on supernet design, incur additional optimization overhead, and have limited search spaces, making them unsuitable for specific or complex operations. Therefore, surrogate-assisted neural network architecture search has emerged as a promising alternative. These methods use surrogate models to predict the performance of neural networks, thereby reducing the time cost of actual performance evaluation.

[0004] Methods that use surrogate models as performance predictors for neural network architecture search often suffer from the following drawbacks: Proxy model training requires a large amount of training data, which is inherently expensive, as each piece of data requires the true performance of the neural network architecture; Proxy models' prediction accuracy cannot be absolutely accurate, leading to the possibility of predicting high-performance architectures as low-performance ones, which often affects the model's search direction. Therefore, it is crucial to propose a particle swarm neural architecture search method assisted by a ternary comparison surrogate. Summary of the Invention

[0005] Purpose of the invention: The technical problem to be solved by the present invention is to address the deficiencies of the existing technology and provide a particle swarm neural architecture search method and system based on ternary comparison agent assistance. The method includes the following steps:

[0006] Step 1: Initialize a particle population in the search space. Each particle represents a neural network architecture. Use the particle swarm algorithm to evolve G generations and truly evaluate all architecture individuals in the population to obtain true performance. The performance data of all truly evaluated architectures is recorded.

[0007] Step 2: The architectures that have been evaluated are grouped into triplet samples. Each triplet sample contains three architectures. Based on the actual performance difference, each pair of combinations in the triplet is labeled with a performance difference. This serves as the training dataset for the proxy model, and the proxy model is trained.

[0008] Step 3: Continue to generate architecture individuals through the update formula of the particle swarm algorithm, and use the proxy model to double compare the current architecture with the local optimal architecture of the current particle and take relaxation operations to obtain the pros and cons relationship between the current architecture and the local optimal architecture of the current particle. When the local optimal architecture is replaced, the current architecture and the global optimal architecture of the particle swarm are put into the proxy model for double comparison, and relaxation operations are taken to determine whether to replace the global optimal architecture.

[0009] Step 4: Whenever a local optimal architecture is replaced, the actual performance of the updated local optimal architecture is evaluated and added to the architecture pool. When the number of different architectures in the architecture pool reaches the set threshold, the proxy model is updated with the latest architecture performance data to optimize the model's performance to adapt to different search stages.

[0010] Step 5: Repeat steps 3 to 4 until the stopping condition is met and output the global optimal architecture.

[0011] In step 1, initialize a particle population , N i Represents the i-th neural network architecture, n is the number of particles in the population; and randomly initialized Position and velocity, N i The position is represented by , For the The particle in Position on the dimension; The speed is , For the The particle in Speed ​​in dimension;

[0012] In step 1, It is the entity of the particle swarm in the particle swarm algorithm. Finally, the state update formula in the particle swarm algorithm is used to update the position and speed of each particle. The speed of each particle determines the direction of particle architecture evolution, and the position is the architecture representation of the particle. Generation, here set 2≤G≤3, in this evolutionary G generation, the calculation of the fitness value of each particle is equivalent to the real evaluation of the network architecture represented by each particle. In this process, the performance indicators of each architecture on the task dataset are evaluated and saved in the architecture pool. Where N is the total number of architecture-performance pairs in the architecture pool Archive, (Arch i ,y i ) represents an architecture-performance pair, Arch i is the i-th architecture in the record, y i is the performance corresponding to the i-th architecture. A task dataset refers to a set of data specifically designed for evaluating and training neural network architectures. This dataset is used to test and measure the performance of neural network architectures when processing specific tasks. For example, if the task is image recognition, the dataset may contain various labeled images for training and evaluating neural networks to recognize objects in images. Common image recognition task datasets include CIFAR-10, ImageNet, COCO, etc. The reason why the initial evolutionary number G is set to 2≤G≤3 is that when G is too large, it means that the number of architectures that need to be truly evaluated will increase, and the true evaluation of the architecture is an expensive operation, which goes against the original intention of using agent-assisted optimization; when G is too small, it will result in too little data being provided to the agent model for training, which is likely to lead to insufficient training of the agent model, the learned features cannot fit the entire search space well, and there is a risk of overfitting of the agent model training.

[0013] In step 2, you first need to pool the architecture The architecture with performance indicators is rearranged and combined to obtain a list of triple samples, which is not only conducive to the training of the proxy model, but also solves the biggest bottleneck problem of using proxy-assisted methods to replace the real evaluation process: the training of the proxy model requires a large amount of training data, which is an expensive resource because each data requires the real performance of the neural network architecture. 3 The triplet training dataset is N, where N is the number of architectures in the architecture pool Archive. After obtaining sufficient labeled training data samples, supervised training of the proxy model is performed. Step 2 specifically includes:

[0014] Step 2.1, create a triplet training dataset:

[0015] For the architecture in the architecture pool Archive, the order of magnitude of the non-repeated combination is N 3 The triplet training dataset, a single triplet , the corresponding triplet label is , in order to ensure that similar triples will not be generated, the above satisfy ;

[0016] Step 2.2, design the structure of the proxy model:

[0017] The design of a proxy model's structure directly determines how well it extracts information from the input data. Therefore, designing a model structure that aligns with the application scenario is essential. Graph neural networks can process graph-structured data and naturally capture the connectivity and topology between nodes, which is very useful for understanding the characteristics of neural network architectures.

[0018] The proxy model includes an architecture feature extractor and a performance difference predictor, which are coupled to each other to jointly train and complete the tasks of architecture feature extraction and performance difference prediction;

[0019] The architecture feature extractor includes a graph convolution layer, and the architecture feature extractor extracts the architecture features of the neural network based on the transmission and aggregation of information between nodes to obtain an architecture representation;

[0020] The performance difference predictor includes a fully connected layer, and the performance difference predictor completes the performance difference prediction of a pair of input neural network architectures based on the learning of architecture representation;

[0021] In the graph convolution layer, node information is a unique feature vector for each node. The information transfer operation between nodes is called aggregation operation. The aggregation operation formula is:

[0022] ,

[0023] in, is the activation function; is the adjacency matrix of a graph containing self-loops; It is the degree matrix of all nodes. The degree matrix is ​​a matrix in graph theory, which describes the degree of each node in the graph. The degree represents the number of edges connected to a node. is a degree matrix containing self-loops, is the trainable weight matrix of the l-th layer architecture feature extractor; represents the node feature vector of the l-th layer architectural feature extractor, and , Represents the input of the entire architecture feature extractor It is the initial node feature vector X. The node feature information extracted by the graph convolution layer is then integrated through a global pooling layer to obtain the representation of the architecture diagram in the latent space.

[0024] Next, the architecture representation will be input into the performance difference predictor to predict the performance difference between the two architectures;

[0025] Step 2.3, design a loss function to guide the training process of the proxy model:

[0026] Designing a loss function for the surrogate model is crucial because it is directly related to the goals and effectiveness of model training. A specifically designed loss function ensures that the direction of model learning is closely aligned with the task requirements, thereby optimizing the model's performance on a specific task.

[0027] The loss function of the designed proxy model is expressed as:

[0028] ,

[0029] ,

[0030] ,

[0031] in, is the cross entropy loss function, is the batch size of input data, is the performance difference predicted by the surrogate model, It is The actual performance of each architecture is poor; is the designed symbolic loss function, where the function is a hinge function with a lower limit of 0, and the function is a sign function; and It is an adaptive weight coefficient that controls the importance between two different loss functions; the total loss function is used Guide the training of the surrogate model until convergence.

[0032] The design of the sign loss function is mainly to ensure that the proxy model can ensure the correct sign of the performance difference when making double comparisons, which corresponds to the relaxation operation; The cross-entropy loss function is designed to ensure high accuracy when comparing poorly performing models even with the introduction of a second comparison layer within the anchor architecture. This high requirement placed on the model allows for better prediction accuracy during the relaxation phase of the model. and The two adaptive weight coefficients are designed to ensure that the cross entropy loss and the sign loss are unbalanced in magnitude, helping to update the model gradient more stably.

[0033] In step 3, the particle swarm algorithm (PSO) is used as the neural network architecture search strategy because, as an evolutionary computational method, it effectively addresses global optimization challenges and is insensitive to local optima. This makes it well-suited for addressing the architectural design challenges of deep learning models, which are often non-differentiable and non-convex. Due to its simplicity and fast convergence, the PSO has shown great potential in addressing complex optimization problems, gradually surpassing traditional genetic algorithms. In the context of neural network architecture search, certain parameters, such as the number of filters in a convolutional layer, have a wide linear adjustable range. These parameters are called ordinals. The PSO, through its smooth parameter update mechanism, can intuitively handle these ordinal parameters, effectively searching for optimal values ​​over a wide range.

[0034] Step 3 specifically includes:

[0035] Step 3.1, update the state of particles in the population through the update formula of the particle swarm algorithm;

[0036] Step 3.2: Use the proxy model to predict the performance difference between the newly generated architecture and the local optimal architecture, as well as the performance difference between the newly generated architecture and the global optimal architecture.

[0037] Step 3.3: Based on the double comparison and relaxation operation, it is determined whether the newly generated architecture needs to be truly evaluated.

[0038] In step 3.4, based on the results of the real evaluation of the newly generated architecture, it is further determined whether the local optimal architecture and the global optimal architecture need to be replaced.

[0039] In step 3.1, the update formula of the particle swarm algorithm is:

[0040] ,

[0041] ,

[0042] in, is the inertia factor; Indicates the number of current iterations; , is the acceleration factor; 、 A random number between 0 and 1; For the Daidi The particle in The speed of the dimension, For the Daidi The particle in Position on dimension; pbest id It is The particle in The position of the individual extreme point of dimension, gbest d It is the position of all extreme points of the entire population in the dth dimension.

[0043] In step 3.2, traverse all individuals in the current population, for any individual N i , N i and N i The corresponding local optimal architecture One-hot encoding is performed separately, and the resulting encoding structures are ,Will Directly connected to form the proxy model input , input the proxy model ip The output is sent to the proxy model to obtain the predicted value of the performance difference between the corresponding newly generated architecture and the local optimal architecture. ip ; The global optimal architecture gbest of the current iteration algebra i Perform one-hot encoding and the resulting encoding is ,Will Directly connected to form the proxy model input , input the proxy model ig The output is sent to the proxy model to get the predicted value of the performance difference between the corresponding newly generated architecture and the global optimal architecture. ig ; Get pbest directly from the architecture pool Archive i The real performance of ip with gbest i The real performance of ig , and get the actual performance difference y ipg =y ip -y ig;Finally, according to output ip 、output ig 、y ipg To directly determine whether the current particle individual needs to replace the local optimal architecture with the global optimal architecture.

[0044] In step 3.3, through output ip 、output ig 、y ipg Calculated , , where sign1 is the difference sign obtained by relaxing the performance difference between the newly generated architecture of the current particle and the local optimal architecture of the current particle. If the sign is positive, it means that the performance of the newly generated architecture of the current particle is higher than that of the local optimal architecture of the current particle. If the sign is negative, it means that the performance of the newly generated architecture of the current particle is lower than that of the local optimal architecture of the current particle. sign2 is based on the current global optimal architecture as the anchor point, and with the help of the known real performance difference y ipg , thereby indirectly comparing the performance of the newly generated architecture of the current particle with the performance of the local optimal architecture of the current particle. Sign2 is also a difference sign. If the sign is positive, it means that the performance of the newly generated architecture of the current particle is higher than the performance of the local optimal architecture of the current particle. If the sign is negative, it means that the performance of the newly generated architecture of the current particle is lower than the performance of the local optimal architecture of the current particle. The sign obtained based on the double comparison of sign1 and sign2 is used to make a true judgment, that is, as long as one of sign1 and sign2 is positive, the newly generated architecture of the current particle is truly evaluated to obtain its true performance. Finally, it is judged whether it is necessary to update the local optimal architecture of the current particle and the global optimal architecture of the current population.

[0045] In step 3.4, when the performance y of the real evaluation of the newly generated architecture i Greater than the actual evaluation performance of the local optimal architecture of the particle corresponding to the newly generated architecture obtained from the architecture pool Archive When N i Replace the local optimal architecture pbest of the particle corresponding to the newly generated architecture i When y i Greater than the performance y of obtaining the true evaluation of the global optimal architecture in the architecture pool Archive gbest , then use N i Replaces the global optimal architecture gbest.

[0046] In step 4, when the newly generated architecture is truly evaluated, the performance y of the true evaluation is obtained. i After that, the architecture performance is compared The deduplicated data is stored in the architecture pool Archive. When the architecture performance in the architecture pool Archive is i ,y i When the number of ) reaches the threshold T (T is usually 60), repeat step 3 and use the architecture performance pair (Arch i ,y i ) is made into a training data set of triplets, so as to retrain the proxy model to ensure the prediction accuracy of the proxy model.

[0047] In step 4, when a locally optimal architecture is replaced, a realistic evaluation of the new architecture is performed. This not only accumulates real-world performance metrics for excellent architectures but also ensures the correct evolutionary direction of the evolutionary algorithm. Accumulating real-world performance metrics for excellent architectures is primarily used to retrain the proxy model. The proxy model that undergoes retraining during the architecture search process is called an online proxy model. Online proxy models can update parameters based on the latest data, more accurately reflecting the current search status and trends. Offline proxy models are typically built based on historical data and may not reflect the latest changes in a timely manner. Online proxy models offer greater flexibility, allowing for the addition of new data points or adjustment of model structures as needed during the search process. Offline proxy models typically have all parameters and structures already determined before the search begins. Therefore, updating proxy models during the architecture search process is essential.

[0048] In step 5, after the particle swarm iterates to the target generation, it will Retrain using the task dataset and use it as the optimal architecture finally searched by the algorithm.

[0049] The present invention also provides a particle swarm neural architecture search system based on ternary comparison agent assistance, comprising:

[0050] Evolution module based on particle swarm algorithm: used to generate new neural network architectures and promote particle swarm iteration;

[0051] Graph Neural Network-based Architecture Feature Extraction Module: This module extracts the topological structure and operational information of the network architecture, thereby generating a code that fully describes the network structure.

[0052] Triplet-based proxy model training module: used to generate training data samples for proxy models. Triplet-based data augmentation allows proxy models to learn richer features.

[0053] Neural network architecture performance difference prediction module based on double comparison and relaxation operations: Use the prediction results of the proxy model to replace the real evaluation, thereby reducing the time cost of architecture evaluation.

[0054] This invention proposes an innovative triple-based dual comparison proxy model to address the high cost problem faced by architecture evaluation in the current field of neural network structure search. This method achieves effective data enhancement by constructing a training data set in the form of triples. Compared with traditional methods, it significantly increases the number of training samples, fundamentally solving the problem of lack of training samples for proxy models. By implementing a dual comparison mechanism, the model can more accurately evaluate the pros and cons between two architectures, ensuring that excellent architectures are effectively retained in the architecture pool. This mechanism not only improves the accuracy of the evaluation, but also greatly reduces the dependence on the real evaluation process, enabling the proxy model to replace the traditional evaluation method to a greater extent, thereby significantly shortening the time cost required for architecture evaluation.

[0055] The present invention proposes an innovative algorithm combining particle swarm optimization (PSO) with a ternary comparison proxy for neural network architecture search. Leveraging a ternary comparison proxy model, this algorithm significantly improves search efficiency and accuracy. In this approach, neural network architectures are modeled as directed acyclic graphs (DAGs), and a graph neural network (GNN) is used as a feature extractor. Through the information propagation mechanism of the GNN, this method comprehensively captures and integrates the topological structure and operational information of the network architecture, converting this information into feature encodings that accurately reflect the architecture's performance. These encodings are then input into a performance differential predictor to predict the architecture's performance. The tight integration of the encoder and predictor ensures consistency and accuracy between the feature representation and performance predictions. To generate training samples for the proxy model, the present invention employs a triplet strategy, which not only reduces costs but also ensures a sufficient sample supply, thereby improving the quality and efficiency of model training. Furthermore, the present invention combines the PSO algorithm with a dual comparison prediction mechanism, using the proxy model's predictions to evaluate the particle's fitness and compare them with the local optimal architecture. This combination not only enhances the applicability of evolutionary algorithms for neural architecture search but also further improves prediction accuracy by introducing relaxation operations when predicting performance differentials using the proxy model. Overall, this invention provides an efficient solution to the key problem of neural network architecture evaluation and promotes the advancement of neural network architecture search technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 Flow chart of the method of the present invention.

[0057] Figure 2 This is a diagram of the proxy model structure of the present invention.

[0058] Figure 3 This is the search space structure diagram of the present invention.

[0059] Figure 4This is a double comparison prediction error diagram of the surrogate model of the present invention. DETAILED DESCRIPTION

[0060] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more apparent.

[0061] like Figure 1 As shown, in a specific embodiment of the present invention, a novel neural architecture search method is proposed, which utilizes a ternary comparison agent to enhance the particle swarm optimization algorithm. Figure 1 The complete flowchart of this method is shown, starting from the initialization stage of the population until the output of the final result. The method of this embodiment is evaluated on the standard dataset NAS-Bench-201 in the field of neural architecture search as an example. Next, the execution steps and key details of the ternary comparison agent assisted particle swarm neural architecture search method proposed in the present invention will be described in detail. The method of the present invention specifically includes the following steps:

[0062] Step 1: Initialize a particle population P with N=30 particles and an empty architecture pool Archive. Evaluate the architectures represented by all particles in P on the task dataset and obtain their true performance values. All initialized architectures are added to the architecture pool Archive in the form of architecture performance pairs. ,in, represents the i-th architecture, represents the performance value corresponding to the i-th architecture, and N is the total number of architecture-performance pairs in the archive. The population P is then evolved for G generations, 2≤G≤3, and each newly generated architecture is evaluated and recorded in the archive. The initial evolution generation number G is set to 2≤G≤3. If the value of G is too high, the number of architectures that need to be actually evaluated will increase, which will lead to higher evaluation costs and violate the original intention of using proxy models to assist optimization and reduce costs. Conversely, if the value of G is too low, the number of neural network architectures with true accuracy obtained is insufficient, which means that the proxy model has received insufficient training data. This may lead to insufficient model training and inability to accurately capture the characteristics of the entire search space, thereby increasing the risk of overfitting the proxy model.

[0063] Step 2: Rearrange and combine the architectures with performance indicators in the architecture pool Archive to obtain a triplet sample list. The architecture pool Archive stores the architecture performance pairs that have been deduplicated. The purpose of deduplication is to ensure that the proxy model does not have duplicate data for training, otherwise it will increase the risk of overfitting the model. The creation of triplet lists of training samples solves the biggest bottleneck problem of using proxy-assisted methods to replace the real evaluation process: the training of the proxy model requires a large amount of training data, which is an expensive resource because each data requires the real performance of the neural network architecture. Because the triplet sample combination method can obtain N 3 The reason the proxy model in this invention can use triplet data as training data is because of its unique predictive metric, performance difference. In the particle swarm algorithm, the elements that need to be compared are the current particle, the local optimal architecture of the current particle, and the global optimal architecture. These three elements correspond exactly to the three architectures in a single triplet. Therefore, the proxy model based on triplet training data can be integrated into the particle swarm algorithm with high compatibility. After obtaining sufficient labeled training data samples, the proxy model is trained in a supervised manner.

[0064] Step 2.1: Divide the architecture in the Archive into training data samples in the form of triples. Single triple , the corresponding triplet label is , in order to ensure that similar triples will not be generated, the above i, j, k are strictly required to meet ; ,Triple list It is a batch data set for training the proxy model, and size is the number of samples. The three architectures in Triple are all encapsulated using the Data data class in the torch_geometric library. list It is a list of graphic data converted by the to_data_list method in the torch_geometric library.

[0065] Step 2.2: Build Figure 2 The proxy model shown in the figure consists of an architectural feature extractor and a performance predictor. The architectural feature extractor is composed of three layers of graph convolutional layers plus a global max pooling layer, while the performance loss predictor is composed of three layers of fully connected layers. The feature extractor and performance predictor are coupled, allowing a single loss function to guide the training of both models simultaneously. This also facilitates better information transfer between the two models.

[0066] Step 2.3: Using Triple list As a surrogate model, the supervised training dataset is used. As the loss function to guide the directional training of the proxy model. Among them, L1 is the cross entropy loss function, L2 is the designed symbol loss function, and It is an adaptive weight coefficient that controls the importance between two different loss functions. The design of the sign loss function L2 is mainly to ensure that the proxy model can ensure the correctness of the sign of the performance difference when making double comparisons, which corresponds to the relaxation operation. The design of the cross entropy loss function L1 is mainly to ensure that the comparison between performance differences can still maintain high accuracy when introducing the second comparison of the anchor architecture. The high requirements of this loss on the model can enable the relaxation operation of the model in the prediction stage to achieve better prediction accuracy. and The two adaptive weight coefficients are designed to ensure that the cross entropy loss and the sign loss are unbalanced in magnitude, helping to update the model gradient more stably.

[0067] Step 3: If Figure 1 As shown in the figure, the state update formula of the particle swarm algorithm is used to update the state of the particles, and then the environment selection is performed on the population. The process of environment selection requires comparing the performance of the newly generated architecture of each particle in this generation with the local optimal architecture corresponding to each particle.

[0068] Step 3.1: The reason why the particle swarm algorithm is suitable for architecture search of neural networks is that it is an evolutionary algorithm that is good at finding the global optimal architecture rather than just staying at the local optimum. This algorithm is particularly effective for complex, nonlinear problems in deep learning that are difficult to solve with traditional methods. Compared with traditional genetic algorithms, the particle swarm algorithm is simpler and can find solutions faster. In addition, the three most important elements in the particle swarm algorithm are the current particle state, the local optimal architecture of the current particle, and the global optimal architecture. These three elements are combined with the ternary comparison of the present invention, which greatly improves the fit between the search strategy and the evaluation strategy in the neural architecture search.

[0069] Step 3.2: Use the proxy model to predict the performance difference between the newly generated architecture and the local optimal architecture and the performance difference between the newly generated architecture and the global optimal architecture. Traverse all particles in the current population, for any particle N i , N i and pbest i The corresponding network architectures are one-hot encoded, and the resulting encoding structures are , directly connect the two to form the proxy model input , input the proxy model into The predicted value of the performance difference between the corresponding newly generated architecture and the local optimal architecture is obtained by feeding it into the proxy model. ; gbest i The corresponding network architectures are one-hot encoded, and the resulting encoding is ,Will Direct connection constitutes the proxy model input , input the proxy model into The predicted value of the performance difference between the corresponding newly generated architecture and the global optimal architecture is obtained by feeding it into the proxy model ; From the architecture pool Get pbest directly i with gbest i Real performance and , and get the actual performance difference between the two ; Finally, according to The three are used to directly determine whether the current particle individual needs to replace the local optimal architecture with the global optimal architecture.

[0070] Step 3.3: The three calculations can be obtained , ,in It is the sign of the difference between the performance difference between the newly generated architecture of the current particle and the local optimal architecture of the current particle after the relaxation operation. If the sign is positive, it means that the performance of the newly generated architecture of the current particle is higher than the performance of the local optimal architecture of the current particle. If the sign is negative, it means that the performance of the newly generated architecture of the current particle is lower than the performance of the local optimal architecture of the current particle. and The difference is that it uses the current global optimal architecture as an anchor point and uses known real information , thereby indirectly comparing the performance of the newly generated architecture of the current particle with the performance of the local optimal architecture of the current particle; based on the above and The symbol obtained by the double comparison is judged as true if it is true, that is, and As long as there is a positive sign in , the newly generated architecture of the current particle is truly evaluated to obtain its true performance. Finally, it is determined whether the local optimal architecture of the current particle and the global optimal architecture of the current population need to be updated.

[0071] Step 3.4: When the performance of the newly generated architecture is evaluated Larger than the architecture pool The performance of the actual evaluation of the local optimal architecture of the particle obtained in When N i Alternative to pbest i ;when Greater than the performance of obtaining the true evaluation of the global optimal architecture from the architecture pool archive , then use N iReplaces the global optimal architecture gbest.

[0072] Step 4: When a locally optimal architecture is replaced, a true evaluation of that architecture is performed. Without this evaluation, the online proxy model will degenerate into an offline proxy model. This is because no new labeled architectures are generated during the algorithm's execution. This results in the proxy model lacking new training data samples, which in turn leads to biased predictions. Because offline proxy models cannot adapt well to every stage of the search process, evaluating good architectures during the search process is essential. A true evaluation of an architecture involves training the current architecture with the task dataset and then testing it to obtain the architecture's true performance metrics for the target task.

[0073] Step 5: After the particle swarm iterates to the target generation, it will calculate the current global optimal architecture. The task dataset is retrained and used as the optimal architecture finally searched by the algorithm.

[0074] like Figure 3 As shown in , NAS-Bench-201 defines a neural architecture search framework, where each architecture consists of a fixed skeleton and a series of stacked search units. This process can be viewed as searching for the optimal cell configuration. In this framework, each cell is a densely connected directed acyclic network, such as Figure 3 As shown in the bottom part of . In this graph, nodes represent aggregations of feature maps, while edges represent operations that transform feature maps from one node to another. The size of the search space depends on the number of nodes in the directed acyclic graph and the size of the set of operations to choose from. In NAS-Bench-201, to construct the operation set, this method selected 4 nodes and 5 typical operations, resulting in a search space containing 15,625 different architectures. For each architecture, multiple training runs were performed on three different datasets to ensure the reliability and diversity of the results. The training process of each architecture is recorded in detail at each training iteration, including training logs and performance metrics. These metrics include but are not limited to training accuracy, test accuracy, training loss, and test loss. In addition, the number of parameters and floating-point operations performed for each architecture are also recorded in detail to facilitate analysis and comparison.

[0075] like Figure 4As shown in the figure, a proxy model based on dual comparison was used to conduct a proxy model prediction error experiment on 70% of the 15,625 different architectures in the NAS-Bench-201 space. The experiment trained the proxy model for 100 iterations, tested the proxy model after each training iteration, and recorded the prediction error rate of each proxy model test. The proxy model converged after 50 training generations, and the prediction error rate was stable below 0.02 after 80 training generations. The experiment proves that the proxy model can maintain the prediction error rate after dual comparison prediction of poor architectural performance.

[0076] The present invention provides a particle swarm neural architecture search method and system based on a ternary comparison agent. There are many methods and approaches to implement this technical solution. The above is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements and modifications without departing from the principles of the present invention. These improvements and modifications should also be considered within the scope of protection of the present invention. All components not specified in this embodiment can be implemented using existing technologies.

Claims

1. A particle swarm neural architecture search method based on ternary comparison agent assistance, characterized in that: The following steps are involved: Step 1: Initialize a particle population in the search space. Each particle represents a neural network architecture. Use the particle swarm algorithm to evolve G generations and truly evaluate all architecture individuals in the population to obtain true performance, and record all the truly evaluated architecture performance data. Step 2: Make the architectures that have been actually evaluated into triple samples, and label the performance difference of each pair of combinations in each triple according to the actual performance difference, so as to serve as the training data set of the proxy model, and train the proxy model; Step 3: Continue to generate architecture individuals through the update formula of the particle swarm algorithm, and use the proxy model to double compare the current architecture with the local optimal architecture of the current particle and take relaxation operations to obtain the relationship between the current architecture and the local optimal architecture of the current particle. When the local optimal architecture is replaced, the current architecture and the global optimal architecture of the particle swarm are put into the proxy model for double comparison, and relaxation operations are taken to determine whether to replace the global optimal architecture. Step 4: Whenever a local optimal architecture is replaced, the actual performance of the updated local optimal architecture is evaluated, and the updated local optimal architecture is added to the architecture pool. When the number of different architectures in the architecture pool reaches the set threshold, the proxy model is updated using the latest architecture performance data to optimize the model's performance to adapt to different search stages. Step 5, repeat steps 3 to 4 until the stopping condition is met, and output the global optimal architecture; Step 3 includes: Step 3.1, update the state of particles in the population through the update formula of the particle swarm algorithm; Step 3.2, using the proxy model to predict the performance difference between the newly generated architecture and the local optimal architecture, as well as the performance difference between the newly generated architecture and the global optimal architecture; Step 3.3, based on the double comparison and relaxation operation, determine whether a real evaluation of the newly generated architecture is required; Step 3.4, further determine whether it is necessary to replace the local optimal architecture and the global optimal architecture based on the results of the real evaluation of the newly generated architecture; In step 3.1, the update formula of the particle swarm algorithm is: Among them, ω is the inertia factor; k represents the number of current iterations; c1 and c2 are acceleration coefficients; rand1 and rand2 are random numbers between 0 and 1; is the velocity of the i-th particle of the k-th generation in the d-th dimension, is the position of the i-th particle of the k-th generation in the d-th dimension; pbest id is the position of the individual extreme point of the ith particle in the dth dimension, gbest d It is the position of all extreme points of the entire population in the dth dimension.

2. The method according to claim 1, characterized in that In step 1, initialize a particle population P = [N1, N2, ..., N n ],N i represents the i-th neural network architecture, n is the number of particles in the population; and N is randomly initialized i The position and velocity, N i The position is represented by x id is the position of the ith particle in the dth dimension; N i The speed is v id is the speed of the ith particle in the dth dimension; then evolve for G generations, evaluate the performance indicators of each architecture on the task dataset, and save them in the architecture pool Where N is the total number of architecture-performance pairs in the architecture pool Archive, (Arch i ,y i ) represents an architecture-performance pair, Arch i is the i-th architecture in the record, y i is the performance corresponding to the i-th architecture.

3. The method according to claim 2, characterized in that Step 2 includes: Step 2.1, create a triplet training data set: For the architecture pool Archive, the number of combinations without duplication is N. 3 The triple training data set, a single triple Triple = [Arch i ,Arch j ,Arch k ], the corresponding triple label is Label = [y i -y j ,y i -y k ,(y i -y j )-(y i -y k )], where i, j, k satisfy i≠j, k>j; Step 2.2, design the structure of the proxy model: The proxy model includes an architecture feature extractor and a performance difference predictor, and the architecture feature extractor and the performance difference predictor are coupled to each other to jointly train and complete the tasks of architecture feature extraction and performance difference prediction; The architecture feature extractor includes a graph convolution layer, and the architecture feature extractor extracts the architecture features of the neural network based on the transmission and aggregation of information between nodes to obtain an architecture representation; The performance difference predictor includes a fully connected layer, and the performance difference predictor completes the performance difference prediction of a pair of input neural network architectures based on the learning of the architecture representation; In the graph convolution layer, the node information is a feature vector unique to each node. The information transfer operation between nodes is called the aggregation operation. The aggregation operation formula is: Where σ is the activation function; is the adjacency matrix of the graph with self-loops; D is the degree matrix of all nodes. The degree matrix is ​​a matrix in graph theory that describes the degree of each node in the graph. The degree represents the number of edges connected to a node. is a degree matrix containing self-loops; W (l) is the trainable weight matrix of the l-th layer architecture feature extractor; H (l) represents the node feature vector of the l-th layer architecture feature extractor, and H (0) =X,H (0) =X represents the input H of the entire architecture feature extractor (0) It is the initial node feature vector X. The node feature information extracted by the graph convolution layer is then integrated through a global pooling layer to obtain the representation of the architecture diagram in the latent space. Next, the architecture representation will be input into the performance difference predictor to predict the performance difference between the two architectures; Step 2.3, design a loss function to guide the training process of the proxy model: The loss function of the designed proxy model is expressed as: L=αL1+βL2, Among them, L1 is the cross entropy loss function, m is the batch size of the input data, and predict i is the performance difference predicted by the surrogate model, target i is the actual performance difference corresponding to the i-th architecture; L2 is the designed symbolic loss function, function ψ(x) is the hinge function with a lower limit of 0, and function sign(x) is the sign function; α and β are adaptive weight coefficients; the total loss function L is used to guide the training of the proxy model until convergence.

4. The method according to claim 3, characterized in that In step 3.2, traverse all individuals in the current population, for any individual N i , N i and N i The corresponding local optimal architecture pbest i One-hot encoding is performed respectively, and the resulting encoding structures are Will Directly connected to form the proxy model input Enter the proxy model into input ip Send it to the proxy model to get the predicted value output of the performance difference between the corresponding newly generated architecture and the local optimal architecture ip ; The global optimal architecture gbest of the current iteration i Perform one-hot encoding, and the resulting encoding is Will Directly connected to form the proxy model input Enter the proxy model into input ig Send it to the proxy model to get the predicted value output of the performance difference between the corresponding newly generated architecture and the global optimal architecture ig ; Get pbest directly from the architecture pool Archive i The real performance ip With gbest i The real performance ig , and get the actual performance difference y ipg =y ip -y ig ;Finally, according to output ip 、output ig ,y ipg To directly determine whether the current particle individual needs to replace the local optimal architecture with the global optimal architecture.

5. The method according to claim 4, characterized in that In step 3.3, through output ip 、output ig ,y ipg Calculate sign1 = sign (output ip ),sign2=sign(output ig -y ipg ), where sign1 is the difference sign obtained by relaxing the performance difference between the newly generated architecture of the current particle and the local optimal architecture of the current particle. If the sign is positive, it means that the performance of the newly generated architecture of the current particle is higher than that of the local optimal architecture of the current particle. If the sign is negative, it means that the performance of the newly generated architecture of the current particle is lower than that of the local optimal architecture of the current particle. Sign2 takes the current global optimal architecture as the anchor point and uses the known real performance difference y ipg , thereby indirectly comparing the performance of the newly generated architecture of the current particle with the performance of the local optimal architecture of the current particle. Sign2 is also a difference sign. If the sign is positive, it means that the performance of the newly generated architecture of the current particle is higher than the performance of the local optimal architecture of the current particle. If the sign is negative, it means that the performance of the newly generated architecture of the current particle is lower than the performance of the local optimal architecture of the current particle. The sign obtained based on the double comparison of sign1 and sign2 is used to make a true judgment, that is, as long as one of sign1 and sign2 is positive, the newly generated architecture of the current particle is truly evaluated to obtain its true performance. Finally, it is judged whether it is necessary to update the local optimal architecture of the current particle and the global optimal architecture of the current population.

6. The method according to claim 5, characterized in that In step 3.4, when the real evaluation performance y of the newly generated architecture i Greater than the actual evaluation performance of the local optimal architecture of the particle corresponding to the newly generated architecture obtained from the architecture pool Archive When N i Replace the local optimal architecture pbest of the particle corresponding to the newly generated architecture i When y i Greater than the performance y of obtaining the true evaluation of the global optimal architecture in the architecture pool Archive gbest , then use N i Replaces the global optimal architecture gbest.

7. The method according to claim 6, characterized in that In step 4, when the newly generated architecture is truly evaluated, the performance y of the true evaluation is obtained. i After that, the architecture performance (Arch i ,y i ) is deduplicated and stored in the architecture pool Archive. When the architecture performance in the architecture pool Archive is (Arch i ,y i ) reaches the threshold T, then repeat step 3 and use the architecture performance in the architecture pool Archive to find the architecture performance (Arch i ,y i ) is made into a triplet training data set, so as to retrain the proxy model to ensure the prediction accuracy of the proxy model.

8. The particle swarm neural architecture search based on ternary comparison agent assistance implemented according to the method of any one of claims 1 to 7, characterized in that: include: Evolution module based on particle swarm algorithm: used to generate new neural network architectures and promote the iteration of particle swarms; Architecture feature extraction module based on graph neural network: used to extract the topological structure information and operation information of the network architecture, thereby generating a code that can fully describe the network structure information; Triplet-based proxy model training module: used to generate training data samples for proxy models. Triplet-based data enhancement allows the proxy model to learn richer features. Neural network architecture performance difference prediction module based on double comparison and relaxation operation: Use the prediction results of the proxy model to replace the real evaluation, thereby reducing the time cost of architecture evaluation.

Citation Information

Patent Citations

  • Intelligent identification method for same-opening welding seam image

    CN117237592A

  • Automated variational inference using stochastic models with irregular beliefs

    WO2023249068A1