A neural network architecture search method, device, electronic device and storage medium

By combining genetic algorithms and classification models, the neural network architecture search method is solved, and the problem of complex neural network design is time-consuming and insufficient training data of proxy model is insufficient, and deep neural networks with excellent performance is achieved efficiently, reducing hardware requirements and improving model accuracy.

CN116108384BActive Publication Date: 2025-07-11NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211671424.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-07-11
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

The prior art has problems such as time-consuming and complex in neural network design, relying on manual intervention, difficulty in tuning hyperparameters, and insufficient training data of proxy models, resulting in inefficient neural architecture search.

Method used

The classification model is used to preprocess the classification model through the training data set, and the neural network architecture search is used to use the genetic algorithm, and the trained classification model is used to evaluate the fitness of individual networks in the network population to obtain deep neural networks with excellent performance.

Benefits of technology

Obtaining deep neural networks with excellent performance in a short period of time reduces hardware requirements and improves the efficiency of neural network architecture search and the accuracy of classification models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116108384B_ABST
    Figure CN116108384B_ABST
Patent Text Reader

Abstract

The present invention discloses a neural network architecture search method, device, electronic device and storage medium, belonging to the technical field of automated machine learning. The method includes: obtaining an initial data set; preprocessing the initial data set to obtain a training data set; training a classification model using the training data set; performing neural network architecture search according to a genetic algorithm, and evaluating the fitness of individual networks in a network population using the trained classification model; and obtaining a deep neural network with excellent performance according to the fitness evaluation result. This method can save the time for model training in the neural network architecture search process, reduce the requirements of users for hardware, and obtain a deep neural network with excellent performance in a short time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a neural network architecture search method, device, electronic device and storage medium, and belongs to the technical field of automated machine learning. Background Art

[0002] Deep learning evolved from multi-layer neural networks is the current mainstream method for big data processing and analysis. A neural network is the basic module of deep learning, and its essence is a network model reflecting the feature mapping relationship. Deep neural networks are often designed by professionals based on professional knowledge and past experience. Therefore, it is a time-consuming, complex and error-prone task to search or design excellent structures manually. Moreover, deep neural networks face a large number of hyperparameter selections. The tuning of these hyperparameters requires repeated attempts and continuous trial and error. There is a lack of effective theoretical methods, and in many cases, it is a very skillful and uncertain task.

[0003] In view of the various deficiencies of the above-mentioned manual design of deep neural networks, scholars at home and abroad have proposed many explorations and improvements. In recent years, neural architecture search technology has attracted extensive attention in the industrial and academic fields, that is, by methods such as reinforcement learning, evolutionary algorithms and gradient algorithms, an excellent neural network structure is automatically searched, so as to reduce human intervention in the design of neural networks. During the neural architecture search process, each network structure needs to be fully trained and evaluated on the corresponding dataset, which makes the neural architecture search not only require a large number of computing facilities (such as GPUs, etc.), but also require a large amount of time overhead. Therefore, a method of accelerating with a surrogate model is introduced. It is proposed to use the network structure as the input information of a surrogate model, and use the obtained predicted value as the evaluation of the network structure to select an excellent network structure, thereby greatly improving the speed of neural architecture search. However, most surrogate models use regression machine learning models, and the training of such surrogate models often has the problem of insufficient data.

[0004] The genetic algorithm uses the natural selection mechanism and genetic laws in the biological evolution process to solve optimization problems, effectively combines random search with parallel enhanced nearest neighbor search, and has characteristics such as generality, parallelism and global optimality. Moreover, evolutionary computing uses the natural evolution mechanism to represent complex phenomena and can quickly and effectively solve complex problems. Evolutionary computing has been applied in the field of automated deep learning. Summary of the Invention

[0005] The purpose of the present invention is to provide a neural network architecture search method, device, electronic device and storage medium, which can obtain an excellent neural network within a relatively small time budget based on the assistance of a classification model.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a neural network architecture search method, including:

[0008] Obtain an initial data set;

[0009] Preprocess the initial data set to obtain a training data set;

[0010] Use the training data set to train a pre-constructed classification model;

[0011] Perform neural network architecture search according to a genetic algorithm, and use the trained classification model to evaluate the fitness of individual networks in a network population;

[0012] According to the fitness evaluation result, obtain a deep neural network with excellent performance.

[0013] Combined with the first aspect, further, the initial data set includes network structure information and network structure prediction performance.

[0014] Combined with the first aspect, further, obtaining the initial data set includes:

[0015] Obtain an initial sample and encode the initial sample;

[0016] Use the accuracy of the initial sample as its corresponding label;

[0017] The initial data set is jointly constituted by the encoding and label of the initial sample.

[0018] Combined with the first aspect, further, the accuracy of the initial sample is the accuracy rate when the neural network classifies an image, that is, the ratio of the number of correctly classified images to the total number of images to be classified.

[0019] Combined with the first aspect, further, encoding the initial sample includes:

[0020] Use a One-hot vector to represent the node type of the initial sample;

[0021] Use an adjacency matrix to represent the connection relationship between nodes;

[0022] Flatten the upper triangular matrix of the adjacency matrix into a one-dimensional vector, and use the one-dimensional vector as the encoding of the initial sample.

[0023] Combined with the first aspect, further, preprocessing the initial data set to obtain a training data set includes:

[0024] Select any two encodings in the initial data set for pairing;

[0025] According to the two paired codes, assign a digital label 1 to one of the codes and a digital label 0 to the other code;

[0026] Construct a training data set in the way of pairwise selection from the four codes.

[0027] Combined with the first aspect, further, evaluating the fitness of individual networks in the network population using the trained classification model includes:

[0028] Initialize the fitness to 0;

[0029] Input the training data set into the classification model;

[0030] If the output result of the classification model is 1, add 1 point to the fitness of the network on the coding side corresponding to the digital label 1;

[0031] If the output result of the classification model is 0, add 1 point to the fitness of the network on the coding side corresponding to the digital label 0.

[0032] In the second aspect, the present invention provides a neural network architecture search device, including:

[0033] Data acquisition module: used to acquire the initial data set;

[0034] Preprocessing module: used to preprocess the initial data set to obtain the training data set;

[0035] Model training module: used to train the pre-constructed classification model using the training data set;

[0036] Search module: used to perform neural network architecture search according to the genetic algorithm, and evaluate the fitness of individuals in the population using the trained classification model to obtain a deep neural network with excellent performance.

[0037] In the third aspect, the present invention provides an electronic device, including a processor and a storage medium;

[0038] The storage medium is used to store instructions;

[0039] The processor is used to operate according to the instructions to execute the steps of the method according to any one of the first aspects.

[0040] In the fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of the first aspects are implemented.

[0041] Compared with the prior art, the beneficial effects of the present invention are:

[0042] The present invention first trains a classification model, and uses the classification model to evaluate the fitness of an individual corresponding to a neural network during the neural network architecture search process by means of a genetic algorithm, so as to obtain a deep neural network with excellent performance, thereby saving the time for model training during the neural network architecture search process, reducing the requirements of users for hardware, and being able to obtain a deep neural network with excellent performance within a very short time budget. During the process of using the genetic algorithm to search for a network, the classification model can be further trained, which helps to improve the accuracy of the classification model, and thus helps to obtain a deep neural network with excellent performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flowchart of a neural network architecture search method provided by an embodiment of the present invention;

[0044] Figure 2 is an example diagram of the encoding of a single individual provided by an embodiment of the present invention;

[0045] Figure 3 is an example diagram of the encoding pairing method of two individuals provided by an embodiment of the present invention;

[0046] Figure 4 is an example diagram of the pairing method when obtaining a training data set provided by an embodiment of the present invention;

[0047] Figure 5 is an example diagram of obtaining paired samples when evaluating fitness provided by an embodiment of the present invention;

[0048] Figure 6 is an example diagram of a crossover operation provided by an embodiment of the present invention;

[0049] Figure 7 is an example diagram of a mutation operation provided by an embodiment of the present invention;

[0050] Figure 8 is an example diagram of the search result on the data set NAS - Bench - 101 provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The technical solutions of this patent will be further described in detail below in conjunction with the specific embodiments.

[0052] The embodiments of this patent are described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain this patent and should not be construed as a limitation of this patent. Without conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.

[0053] Embodiment 1:

[0054] Figure 1 This is a flowchart of a neural network architecture search method provided in the first embodiment of the present invention. This flowchart only shows the logical order of the method in this embodiment. On the premise of non - conflict, in other possible embodiments of the present invention, the steps shown or described may be completed in a different Figure 1 order than that shown.

[0055] The neural network architecture search method provided in this embodiment can be applied to a terminal and can be executed by a neural network architecture search device. This device can be implemented in a software - and / or - hardware manner and can be integrated in the terminal. For example: any tablet computer or computer device with communication functions. Refer to Figure 1 This embodiment of the method specifically includes the following steps:

[0056] Step 1: Obtain an initial data set;

[0057] The initial data set includes network structure information and network structure prediction performance.

[0058] Obtaining the initial data set includes the following steps:

[0059] Step A: Obtain an initial sample and encode the initial sample;

[0060] Encoding the initial sample includes the following steps:

[0061] Step ①: Use a One - hot vector to represent the node type of the initial sample;

[0062] Step ②: Use an adjacency matrix to represent the connection relationship between nodes;

[0063] Step ③: Flatten the upper triangular matrix of the adjacency matrix into a one - dimensional vector and use the one - dimensional vector as the encoding of the initial sample.

[0064] Step B: Use the accuracy of the initial sample as its corresponding label;

[0065] Step C: jointly form the initial data set from the encoding and label of the initial sample;

[0066] Among them, the accuracy of the initial sample is the accuracy rate when the neural network classifies images, that is, the ratio of the number of correctly classified images to the total number of images to be classified.

[0067] In this embodiment, a public platform is first established. Users train their own models on the public platform, and the platform collects the data obtained by the users during model training (including model structure, data set used by the model, number of parameters, computational complexity, accuracy, and training time consumption, etc.). In this embodiment, the classification model is a support vector machine for binary classification tasks, and the data set of this example uses an existing data set: NAS-Bench-101.

[0068] Step 2: Preprocess the initial data set to obtain a training data set;

[0069] Preprocessing the initial data set to obtain a training data set includes the following steps:

[0070] Step a: Select any two encodings in the initial data set for pairing;

[0071] Step b: According to the two paired encodings, assign the digital label 1 to one encoding and the digital label 0 to the other encoding;

[0072] Step c: Construct a training data set according to the method of pairwise selection from four encodings;

[0073] The individual encoding of the genetic algorithm is as Figure 2 shown. This example provides a specific encoding method: using a One-hot vector to represent the node type and an adjacency matrix to represent the connection method between nodes. Finally, take the upper triangular matrix and flatten it into a one-dimensional vector. Use this encoding method for all the obtained samples and use their accuracy as the corresponding labels to obtain the initial data set.

[0074] Step 3: Use the training data set to train the pre-constructed classification model;

[0075] In this embodiment, when training the binary classification SVM: first preprocess the initial data set, as Figure 3 shown, randomly select two individual encodings for pairing, with labels 1 and 0. 1 means the left network is better, and 0 means the right network is better. Use this as the input and output corresponding labels of the classification model; then adopt the method of pairwise selection of individuals from four individuals, as Figure 4 shown, to construct a training data set; finally, use the training data set to train the binary classification SVM.

[0076] Step 4: Search for the neural network architecture according to the genetic algorithm, and use the trained classification model to evaluate the fitness of individual networks in the network population;

[0077] Neural architecture search refers to given a set of candidate neural network structures called the search space, and using a certain strategy to search for the optimal network structure from it. The quality of a neural network structure, that is, its performance, is measured by certain metrics such as accuracy and speed, which is called performance evaluation. A classification model refers to using a classification model to evaluate the performance of a neural network structure during the neural architecture search process.

[0078] Evaluating the fitness of individual networks in a network population using a trained classification model includes the following steps:

[0079] Step Ⅰ: Initialize the fitness to 0;

[0080] Step Ⅱ: Input the training data set into the classification model;

[0081] Step Ⅲ: If the output result of the classification model is 1, add 1 point to the fitness of the encoded-side network corresponding to the digital label 1;

[0082] Step Ⅳ: If the output result of the classification model is 0, add 1 point to the fitness of the encoded-side network corresponding to the digital label 0;

[0083] The evaluation method adopted in this embodiment is: Initialize the fitness to 0. As Figure 5 shown, select all paired encodings, use the paired encodings as the features to be predicted and input them into the binary SVM, and the output result is 1 or 0. Each time the classification model outputs, accumulate scores for the corresponding individuals. If the classification model outputs 1, add 1 point to the network on the left side of the encoding. If the classification model outputs 0, add 1 point to the network on the right side of the encoding.

[0084] Step Five: According to the fitness evaluation result, obtain a deep neural network with excellent performance;

[0085] In this embodiment, experiments are carried out by combining the genetic algorithm and the classification model. Start the iterative process of searching using the genetic algorithm, and use the binary SVM to obtain the performance of individuals as fitness; Train the deep neural networks corresponding to the top n individuals in terms of fitness until convergence, make a new original data set from these individuals according to Step One, obtain a new training data set according to Step Two, and further train the binary SVM using the new training data set; Retain the top several individuals as elites in the network population, use the selection strategy to select the parents of the offspring, and generate new individuals by mutation or crossover with a certain probability and add them to the new network population until the number of individuals in the new network population reaches the set value. Figure 6 is the flowchart of the crossover operation. Figure 7It is a flowchart of the mutation operation; thereafter, the next iteration process is carried out, and so on until the specified number of iterations or the running time budget is reached; in the last iteration of training, the neural network corresponding to the individual with the highest training fitness is trained to obtain its performance, and the result is output; Figure 8 It is the result of searching on the dataset NAS - Bench - 101.

[0086] The neural network architecture search method provided in this embodiment first trains a classification model, and uses the classification model to evaluate the fitness of the individuals corresponding to the neural networks during the neural network architecture search process by means of a genetic algorithm, so as to obtain a deep neural network with excellent performance, thereby saving the time for model training in the neural network architecture search process, reducing the requirements of users for hardware, and being able to obtain a deep neural network with excellent performance within a very short time budget. During the process of using the genetic algorithm to search for the network, the classification model can be further trained, which helps to improve the accuracy of the classification model, and thus helps to obtain a deep neural network with excellent performance.

[0087] Embodiment 2:

[0088] This embodiment provides a neural network architecture search device, including:

[0089] A data acquisition module: used to acquire an initial data set;

[0090] A preprocessing module: used to preprocess the initial data set to obtain a training data set;

[0091] A model training module: used to train a pre - constructed classification model using the training data set;

[0092] A search module: used to perform neural network architecture search according to the genetic algorithm, and use the trained classification model to evaluate the fitness of individuals in the population to obtain a deep neural network with excellent performance.

[0093] The neural network architecture search device provided by the embodiments of the present invention can execute the neural network architecture search method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0094] Embodiment 3:

[0095] This embodiment provides an electronic device, including a processor and a storage medium;

[0096] The storage medium is used to store instructions;

[0097] The processor is used to operate according to the instructions to execute the steps of the method in Embodiment 1.

[0098] Embodiment 4:

[0099] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method in Embodiment 1 are implemented.

[0100] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0101] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0102] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0104] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.

Claims

1. A neural network architecture search method, characterized in that, Including: Obtain an initial data set; Preprocess the initial data set to obtain a training data set; Use the training data set to train a pre-constructed classification model; Conduct a neural network architecture search according to a genetic algorithm, and use the trained classification model to evaluate the fitness of individual networks in the network population; Obtain a neural network with excellent performance according to the fitness evaluation result; Obtaining the initial data set includes: Obtain an initial sample and encode the initial sample; Use the accuracy of the initial sample as its corresponding label; The initial data set is jointly composed of the encoding and label of the initial sample; The accuracy of the initial sample is the accuracy rate when the neural network classifies an image, that is, the ratio of the number of correctly classified images to the total number of images to be classified; Preprocessing the initial data set to obtain a training data set includes: Select any two encodings in the initial data set for pairing; According to the two paired encodings, assign the digital label 1 to one encoding and the digital label 0 to the other encoding; Construct a training data set in the way of pairwise selection from four encodings; Using the trained classification model to evaluate the fitness of individual networks in the network population includes: Initialize the fitness to 0; Input the training data set into the classification model; If the output result of the classification model is 1, add 1 point to the fitness of the network on the side of the encoding corresponding to the digital label 1; If the output result of the classification model is 0, add 1 point to the fitness of the network on the side of the encoding corresponding to the digital label 0.

2. The neural network architecture search method according to claim 1, wherein The initial data set includes network structure information and network structure prediction performance.

3. The neural network architecture search method according to claim 1, characterized in that, Encoding the initial sample includes: Use a One-hot vector to represent the node type of the initial sample; Use an adjacency matrix to represent the connection relationship between each node; Flatten the upper triangular matrix of the adjacency matrix into a one-dimensional vector, and use the one-dimensional vector as the encoding of the initial sample.

4. A neural network architecture search device, characterized in that, Including: Data acquisition module: used to obtain an initial data set; Preprocessing module: used to preprocess the initial data set to obtain a training data set; Model training module: used to train a pre-constructed classification model using the training data set; Search module: used to conduct a neural network architecture search according to a genetic algorithm, and use the trained classification model to evaluate the fitness of individuals in the population, and obtain a deep neural network with excellent performance; Obtaining the initial data set includes: Obtain an initial sample and encode the initial sample; Use the accuracy of the initial sample as its corresponding label; The initial data set is jointly composed of the encoding and label of the initial sample; The accuracy of the initial sample is the accuracy rate when the neural network classifies an image, that is, the ratio of the number of correctly classified images to the total number of images to be classified; Preprocessing the initial data set to obtain a training data set includes: Select any two encodings in the initial data set for pairing; According to the two paired encodings, assign the digital label 1 to one encoding and the digital label 0 to the other encoding; Construct a training data set in the way of pairwise selection from four encodings; Evaluating the fitness of individual networks in a network population using a trained classification model includes: Initializing the fitness to 0; Inputting the training data set into the classification model; If the output result of the classification model is 1, adding 1 point to the fitness of the encoded-side network corresponding to the digital label 1; If the output result of the classification model is 0, adding 1 point to the fitness of the encoded-side network corresponding to the digital label 0.

5. An electronic device, characterized in that, Including a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 3.