Neural network architecture search method and device, storage medium and electronic equipment
By determining the search space, sampling and optimization matrix in neural network architecture search, the problems of high resource consumption and limited search performance in the existing technology are solved, low-cost and high-efficiency neural network architecture search are realized, and high-quality neural network model is constructed.
Patent Information
- Application Number
- CN202510519904.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing neural network architecture search technology consumes too much resources in constructing a differentiable supernet and training process, and the search performance is limited, making it difficult to achieve low-cost and high-efficiency architectural search.
By determining the search space of the neural network architecture, initializing the feature matrix and adjacency matrix, sampling multiple times and inputting a pre-trained proxy model for evaluation, adjusting the evaluation function and optimizing the matrix until the preset conditions are met, building the target neural network.
The search neural network architecture is realized at low cost and high sampling efficiency, which improves the search performance and can effectively build high-quality neural network models.
Smart Images

Figure CN120031073A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of neural network technology, and in particular to a neural network architecture search method, device, storage medium, and electronic device. Background Art
[0002] Neural Architecture Search (NAS) is a technology that designs neural network architectures in an automated manner without human intervention. It aims to search for the optimal network structure through algorithms to improve the performance of neural network models on specific tasks.
[0003] In NAS technology, a differentiable proxy model is used to evaluate neural network architectures. If the modeling of a differentiable proxy model is viewed as an approximate representation, or proxy representation, of the original architecture space, then architecture differentiable optimization can be performed directly based on the proxy model. However, a bottleneck in existing techniques is that the neural network architecture space itself is a non-differentiable discrete space, making differentiable optimization impossible.
[0004] To address these issues, existing NAS technologies achieve architectural space relaxation and differentiable optimization by constructing differentiable supernetworks. However, this approach consumes excessive costs in constructing and training the differentiable supernetworks, relies heavily on significant computing power, resources, and time, and ultimately offers limited search performance.
[0005] Therefore, how to achieve neural network architecture search with low search cost and high sampling efficiency is an urgent problem to be solved. Summary of the Invention
[0006] This specification provides a neural network architecture search method, device, storage medium and electronic device to at least partially solve the above-mentioned problems existing in the prior art.
[0007] This manual adopts the following technical solutions: This specification provides a neural network architecture search method, including: Determine a search space for a neural network architecture based on a target task, wherein the search space includes a set of candidate connection relationships and a set of candidate operations for the network layer; Determining matrix sizes and row / column label information of a first probability variable matrix and a second probability variable matrix respectively according to the search space, and initializing values of elements in the first probability variable matrix and the second probability variable matrix respectively, wherein the first probability variable matrix is a probability distribution representation of a feature matrix of a neural network architecture, and the second probability variable matrix is a probability distribution representation of an adjacency matrix of the neural network architecture; Sampling the first probability variable matrix and the second probability variable matrix multiple times to obtain multiple matrix pairs; Inputting each matrix pair into a pre-trained proxy model, respectively, to obtain an evaluation value of each matrix pair output by the proxy model according to the pre-trained evaluation function; Adjusting an evaluation function included in the proxy model according to the evaluation value, and optimizing the first probability variable matrix and the second probability variable matrix according to the gradient of the adjusted evaluation function; Repeating the sampling and optimization process based on the optimized first probability variable matrix and the second probability variable matrix until a preset termination condition is reached; A target matrix pair is determined among all sampled matrix pairs, and a target neural network is constructed based on the target matrix pair.
[0008] Optionally, determine the search space of the neural network architecture based on the target task, specifically including: Determine the target task to which the target neural network to be constructed belongs; Acquire several standard neural network models for the target task, where the standard neural network models are artificially constructed based on prior knowledge; The search space is determined based on the network layers and atomic operations included in each standard neural network model.
[0009] Optionally, determining the matrix scale and row / column label information of the first probability variable matrix and the second probability variable matrix respectively according to the search space specifically includes: The matrix scale of the first probability variable matrix is determined based on the number of network layers adopted in the candidate connection relationship set and the number of candidate operations included in the candidate operation set, and the matrix scale of the second probability variable matrix is determined based on the number of network layers adopted in the candidate connection relationship set.
[0010] Optionally, pre-train the proxy model, specifically including: Obtaining several sample neural network models under the target task; For each sample neural network model, perform a capability test on the sample neural network model to obtain a first labeled evaluation value of the sample neural network model; Determine the sample feature matrix and sample adjacency matrix of the sample neural network model; Inputting the sample feature matrix and the sample adjacency matrix into the proxy model to be trained, and obtaining an evaluation value to be optimized output by the proxy model according to the evaluation function to be optimized; The proxy model is trained according to the difference between the evaluation value to be optimized and the first labeled evaluation value.
[0011] Optionally, performing a capability test on the sample neural network model to obtain a first labeled evaluation value of the sample neural network model specifically includes: Training the sample neural network model and deploying the trained sample neural network model to a hardware platform; Processing a pre-built test set using the sample neural network model on the hardware platform to obtain capability indicators of the sample neural network model, wherein the capability indicators include at least inference accuracy, inference latency, and power consumption; A first labeled evaluation value of the sample neural network model is determined according to the capability indicator.
[0012] Optionally, adjusting the evaluation function included in the proxy model according to the evaluation value specifically includes: For each matrix pair, a capability test is performed on a neural network architecture constructed according to the matrix pair to obtain a second labeled evaluation value of the matrix pair; According to the second labeled evaluation value of each matrix pair and the evaluation value of each matrix pair output by the proxy model, the parameters of the proxy model are adjusted to adjust the evaluation function in the proxy model.
[0013] Optionally, the evaluation function is a function of the first probability variable matrix and the second probability variable matrix; Optimizing the first probability variable matrix and the second probability variable matrix according to the gradient of the adjusted evaluation function specifically includes: Extracting a first sample matrix according to the first probability variable matrix, and extracting a second sample matrix according to the second probability variable matrix; respectively determining a first partial derivative of the evaluation function with respect to the first probability variable matrix and a second partial derivative of the evaluation function with respect to the second probability variable matrix; Substituting the first sample matrix into the first partial derivative to obtain a first forward propagation gradient of the first probability variable matrix, and substituting the second sample matrix into the second partial derivative to obtain a second forward propagation gradient of the second probability variable matrix; The first probability variable matrix is optimized according to the first forward propagation gradient, and the second probability variable matrix is optimized according to the second forward propagation gradient.
[0014] This specification provides a neural network architecture search device, the device comprising: A determination module is used to determine a search space for a neural network architecture according to a target task, wherein the search space includes a set of candidate connection relationships and a set of candidate operations of the network layer; an initialization module, configured to determine the matrix scale and row / column label information of a first probability variable matrix and a second probability variable matrix respectively according to the search space, and to initialize the values of each element in the first probability variable matrix and the second probability variable matrix respectively, wherein the first probability variable matrix is a probability distribution representation of a feature matrix corresponding to the neural network architecture, and the second probability variable matrix is a probability distribution representation of a corresponding adjacency matrix in the neural network architecture; a sampling module, configured to sample the first probability variable matrix and the second probability variable matrix multiple times to obtain multiple matrix pairs; An input module, configured to input each matrix pair into a pre-trained proxy model, and obtain an evaluation value of each matrix pair output by the proxy model according to an evaluation function obtained in the pre-training; an adjustment module, configured to adjust an evaluation function included in the proxy model according to the evaluation value, and optimize the first probability variable matrix and the second probability variable matrix according to a gradient of the adjusted evaluation function; A loop module, configured to repeatedly execute the sampling and optimization process based on the optimized first probability variable matrix and the second probability variable matrix until a preset termination condition is reached; A construction module is used to determine a target matrix pair among all sampled matrix pairs and construct a target neural network based on the target matrix pair.
[0015] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned neural network architecture search method.
[0016] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned neural network architecture search method is implemented.
[0017] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects: In the neural network architecture search method provided in this specification, it can be seen from the above method that a search space for a neural network architecture is determined according to a target task, the search space including a set of candidate connection relationships and a set of candidate operations of a network layer; the matrix scale and row / column label information of a first probability variable matrix and a second probability variable matrix are respectively determined according to the search space, and the values of each element in the first probability variable matrix and the second probability variable matrix are respectively initialized, wherein the first probability variable matrix is a probability distribution representation of a feature matrix of the neural network architecture, and the second probability variable matrix is a probability distribution representation of an adjacency matrix of the neural network architecture; the first probability variable matrix and the second probability variable matrix are sampled multiple times to obtain a plurality of matrix pairs; each matrix pair is input into a pre-trained proxy model to obtain an evaluation value of each matrix pair output by the proxy model according to an evaluation function obtained in the pre-training; the evaluation function included in the proxy model is adjusted according to the evaluation value, and the first probability variable matrix and the second probability variable matrix are optimized according to the gradient of the adjusted evaluation function; the sampling and optimization process is repeated based on the optimized first probability variable matrix and the second probability variable matrix until a preset termination condition is reached; a target matrix pair is determined from all the sampled matrix pairs, and a target neural network is constructed based on the target matrix pair.
[0018] When using the neural network architecture search method provided in this specification to construct a target neural network model under a target task, after determining the search space, a first probability variable matrix and a second probability variable matrix can be constructed respectively, and a matrix pair consisting of a feature matrix and an adjacency matrix can be sampled according to the first probability variable matrix and the second probability variable matrix; the evaluation value of the matrix pair is output by the proxy model, and the proxy model is optimized based on this, and the first probability variable matrix and the second probability variable matrix are updated using the optimized proxy model; the sampling and optimization process is repeated until a preset termination condition is met; finally, the target matrix pair is selected from all the sampled matrix pairs and the target neural network model is constructed. This method uses the distribution variables of the architecture topology and the candidate operation features to realize the continuous relaxation of the original discrete architecture space; through the graph data representation and encoding of the network architecture, and the training of the proxy model based on the graph neural network, the proxy representation of the original space is realized; the reparameterization method is used to realize the differentiable search of the graph topology and the feature matrix based on the gradient respectively; the architecture optimization search is realized end to end in a manner of discrete architecture sampling, proxy model training, and architecture search online and collaborative alternation. In summary, for any architecture space, this method does not require pre-defining the number of data points for pre-training the proxy model. It only needs to determine the total number of sampled architecture points as the predefined search overhead to complete the search of the neural network architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings: Figure 1 A flowchart of a neural network architecture search method in this specification; Figure 2 A schematic diagram of the search space of a neural network architecture provided in this specification; Figure 3 A schematic diagram of a feature matrix and an adjacency matrix provided in this specification; Figure 4 A schematic diagram of a neural network architecture provided in this specification; Figure 5 A schematic diagram of a first probability variable matrix and a second probability variable matrix provided in this specification; Figure 6 A schematic diagram of the process of training a proxy model provided in this manual; Figure 7 A schematic diagram of a process for testing the capability of a sample neural network model provided in this specification; Figure 8 A schematic diagram of a process for optimizing a first probability variable matrix and a second probability variable matrix provided in this specification; Figure 9 A framework flow chart of a neural network architecture search method provided in this specification; Figure 10 A schematic diagram of a neural network architecture search device provided in this specification; Figure 11 The corresponding Figure 1 Schematic diagram of electronic equipment. DETAILED DESCRIPTION
[0020] To make the purpose, technical solutions, and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0021] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0022] Figure 1 This is a flowchart of a neural network architecture search method in this specification, which specifically includes the following steps: S100: Determine a search space for a neural network architecture according to a target task, where the search space includes a set of candidate connection relationships and a set of candidate operations of a network layer.
[0023] All steps in the neural network architecture search method provided in this specification can be implemented by any electronic device with computing capabilities, such as a terminal, server, and other devices.
[0024] This method is mainly used to automatically search for the optimal neural network architecture under the target task. Based on this, in this step, the search space of the neural network architecture can be first determined according to the determined target task.
[0025] Neural network architecture is the structural design of a neural network model, encompassing the number of network layers, connectivity, and the atomic operations within each layer. Simply put, neural network architecture encompasses all the elements that go into building a neural network model. Atomic operations refer to the most fundamental, indivisible operations in neural network calculations, such as element-wise addition, element-wise multiplication, matrix multiplication, convolution, and activation functions.
[0026] Neural network architecture search methods essentially search through a large number of possible neural network architectures to find the optimal one. When searching for a neural network architecture, it's necessary to restrict the search scope to avoid invalid searches and increase search efficiency. This restriction is known as the neural network architecture search space.
[0027] After determining the target task for the target neural network to be built, the search space for the neural network architecture can be further limited based on the target task. In this method, the search space can include a set of candidate connection relationships for network layers and a set of candidate operations. The candidate connection relationships refer to the number of network layers in the neural network architecture and the connection relationships between each network layer; the candidate operations refer to the optional atomic operations in the neural network architecture.
[0028] In one embodiment, it is possible to specifically determine the target task to which the target neural network to be constructed belongs; obtain several standard neural network models under the target task, wherein the standard neural network model is a neural network model artificially constructed based on prior knowledge; and determine the search space based on the network layers and atomic operations included in each standard neural network model.
[0029] Once a target task is determined, several standard neural network models for the target task can be obtained through various channels, such as historical data and online searches. Standard neural network models are artificially constructed based on prior knowledge. In other words, they are relatively mature neural network models with excellent performance in the target task domain. By learning the architecture of standard neural network models, the network layer connections and candidate operations of the neural network architecture suitable for the target task can be determined, which is the appropriate search space.
[0030] Of course, in addition to the methods given in the above embodiments, the search space may also be determined by, for example, manually setting the connection relationship set and candidate operation set contained in the search space, and this specification does not impose any specific restrictions on this.
[0031] Figure 2 This is a search space diagram of a neural network architecture provided in this specification. Figure 2 As shown in Figure 2, in this method, the neural network architecture and search space are represented in the form of a directed acyclic graph. Figure 2 In a specific embodiment, the directed acyclic graph includes 5 nodes, each node represents a network layer in the neural network architecture; the candidate operation set represents the optional atomic operations in each blank node. Among them, In represents the input layer, Out represents the output layer, and the blank nodes numbered 1, 2, and 3 represent the network layers for the candidate operations to be determined. The dotted lines connecting the nodes represent the candidate connection relationships between the nodes, that is, the input-output relationships, and the output of a node is given as output to the next node it points to. In a directed acyclic graph, the connection relationship between nodes is unidirectional, and there will be no situation where the downstream node points to the upstream node. Therefore, the directed acyclic graph also represents the positional relationship between the network layers.
[0032] S102: Determine the matrix scale and row / column label information of the first probability variable matrix and the second probability variable matrix respectively according to the search space, and initialize the value of each element in the first probability variable matrix and the second probability variable matrix respectively, wherein the first probability variable matrix is a probability distribution representation of the feature matrix corresponding to the neural network architecture, and the second probability variable matrix is a probability distribution representation of the corresponding adjacency matrix in the neural network architecture.
[0033] After determining the search space of the neural network architecture in step S100, a first probability variable matrix and a second probability variable matrix can be obtained in this step based on the search space. In this method, the first probability variable matrix is a probability distribution representation of the characteristic matrix of the neural network architecture, and the second probability matrix is a probability distribution representation of the adjacency matrix of the neural network architecture. Among them, the characteristic matrix of the neural network architecture is used to characterize the selection relationship between each network layer and each candidate operation in the neural network architecture; the adjacency matrix of the neural network architecture is used to characterize the input-output relationship between each network layer in the neural network architecture. A pair of characteristic matrices and adjacency matrices can determine a corresponding neural network architecture.
[0034] Figure 3 The instructions provided for this Figure 2 A possible schematic diagram of the feature matrix and adjacency matrix in the search space shown in . Figure 3 As shown in the figure, the left figure represents the characteristic matrix, and the right figure represents the adjacency matrix. Regardless of whether it is a characteristic matrix or an adjacency matrix, the value of each element in the matrix can only be 0 or 1. The characteristic matrix and the adjacency matrix are introduced below.
[0035] Among them, the row label in the feature matrix represents each network layer of the neural network architecture, and the column label represents each candidate operation in the candidate operation set. Then the matrix size of the feature matrix is (number of network layers, number of candidate operations). It should be noted that the input layer, output layer and the corresponding input operations and output operations are also calculated. The value of each element in the feature matrix represents the selection relationship between the network layer represented by the row label of the element and the candidate operation represented by the column label of the element. When the element is 1, it means that the corresponding network layer adopts the corresponding candidate operation. Conversely, when the element is 0, it means that the corresponding network layer does not adopt the corresponding candidate operation. Figure 3 For example, Figure 3 In the feature matrix shown on the left, the output layer In selects the output operation, the network layer 1 selects the conv1×1 (conv1×1-bn-relu) operation, the network layer 2 selects the maxpool3×3 operation, the network layer 3 selects the conv3×3 (conv3×3-bn-relu) operation, and the output layer Out selects the output operation.
[0036] The row labels and column labels in the adjacency matrix represent each network layer of the neural network architecture. The value of each element in the adjacency matrix indicates whether the output of the network layer represented by the row label of the element is the input of the network layer represented by the column label of the element. When the value of the element is 1, it means that the output of the network layer represented by the row label is the input of the network layer represented by the column label. Conversely, when the element is 0, it means that the output of the network layer represented by the row label is not the input of the network layer represented by the column label. Figure 3 For example, Figure 3 In the adjacency matrix shown on the right, the output of the input layer is the input of network layers 1, 2, and 3, the output of network layer 1 is the input of the output layer, the output of network layer 2 is the input of network layer 3 and the output layer, the output of network layer 3 is the input of the output layer, and the output of the output layer is not given to any other network layer.
[0037] It is worth mentioning that since the output of the later network layers will not be used as the input of the earlier network layers, the values of the elements in the lower triangle and diagonal of all adjacency matrices in this application are 0.
[0038] In this application, the combination of a feature matrix and an adjacency matrix is defined as a matrix pair. When a matrix pair is given, a neural network architecture can be uniquely determined. Figure 3 Take the matrix pair shown as an example, Figure 4 A basis for this instruction manual Figure 3 The neural network architecture is constructed by the matrix pair shown in Figure 4 As shown, the candidate operation of network layer 1 is conv1×1, the candidate operation of network layer 2 is maxpool3×3, and the candidate operation of network layer 3 is conv3×3; the output of the input layer flows to network layers 1, 2, and 3, the output of network layer 1 flows to the output layer, the output of network layer 2 flows to network layer 3 and the output layer, and the output of network layer 3 flows to the output layer.
[0039] The process of constructing a matrix can be achieved by determining the size of the matrix, row / column label information, and the value of each element in the matrix. The matrix size and row / column label information of the first probability variable matrix and the second probability variable matrix can be determined by the search space determined in step S100. In one embodiment, the matrix size of the first probability variable matrix can be determined based on the number of network layers used in the candidate connection relationship set and the number of candidate operations included in the candidate operation set, and the matrix size of the second probability variable matrix can be determined based on the number of network layers used in the candidate connection relationship set.
[0040] The feature matrix and adjacency matrix have been introduced above. Since the first probability variable matrix is a probabilistic representation of the feature matrix, and the second probability variable matrix is a probabilistic representation of the adjacency matrix, the matrix size and row / column label information of the first probability variable matrix are the same as those of the feature matrix, and the matrix size and row / column label information of the second probability variable matrix are the same as those of the adjacency matrix.
[0041] On the other hand, the values of each element in the first probability variable matrix and the second probability variable matrix can be initialized directly according to preset rules, such as random assignment, etc., and this specification does not impose specific restrictions on this. Figure 5 The instructions provided for this Figure 3 A schematic diagram of a possible first probability variable matrix and second probability variable matrix under the search space shown. Figure 5 As shown in the figure, the left figure is the first probability variable matrix, and the right figure is the second probability variable matrix. In the first probability variable matrix, the value of each element represents the probability that the network layer corresponding to the element's row label selects the candidate operation corresponding to the element's column label; in the second probability variable matrix, the value of each element represents the probability that the output of the network layer corresponding to the element's row label is the input of the network layer corresponding to the element's column label. Among them, the values of all elements in the lower triangle and diagonal of the second probability variable matrix are 0.
[0042] The first probability variable matrix and the second probability variable matrix corresponding to the search space can be obtained in the above manner.
[0043] S104: Sampling the first probability variable matrix and the second probability variable matrix multiple times to obtain multiple matrix pairs.
[0044] In this step, the first probability variable matrix and the second probability variable matrix obtained in step S102 may be sampled multiple times to obtain multiple matrix pairs. Each sampling step samples the first probability variable matrix to obtain a characteristic matrix, and samples the second probability variable matrix to obtain an adjacency matrix. The characteristic matrix and the adjacency matrix obtained by sampling constitute a matrix pair.
[0045] In this method, α represents the first probability variable matrix, β represents the second probability variable matrix; A o Represents the feature matrix obtained by sampling, A t Represents the adjacency matrix obtained by sampling; N represents the number of network layers, and M represents the number of candidate operations. o According to α sampling, the matrix size is N×M, A t According to β sampling, the matrix size is N×N.
[0046] There are many ways to sample the first probability variable matrix and the second probability variable matrix. This specification provides a specific method for reference. In a specific embodiment, α and β can be sampled using different formulas to obtain A o With A t .
[0047] Since each network layer only selects one candidate operation, corresponding to A o In the middle, it is shown as A o Each row of has only one element with a value of 1, and the values of the remaining elements are 0. Therefore, when sampling α, a candidate operation can be extracted for each network layer through polynomial sampling to obtain Ao , the specific formula is as follows: ——Formula① Where i is the row index and Cat (Categorical) is the multinomial distribution sampling. Formula ① means that for a certain network layer, its distribution vector is subjected to softmax, then the multinomial distribution probability is applied, and then sampling is performed based on the probability.
[0048] For each edge in the search space, that is, each element in the feature matrix, sampling can be performed according to the following formula: ——Formula② Among them, (h, k) is the (row, column) index, σ represents the sigmoid activation function, and β (>h>,>k>) The value of activation is [0, 1]. The meaning of formula ② is that for a certain edge A t (>h>,>k>) , its distribution probability is activated by sigmoid and used as binomial distribution probability, and then sampled according to the probability.
[0049] In this step, multiple matrix pairs with probabilistic values of α and β can be obtained through multiple samplings. Different matrix pairs can represent different neural network architectures.
[0050] S106: Inputting each matrix pair into a pre-trained proxy model respectively to obtain an evaluation value of each matrix pair output by the proxy model according to the pre-trained evaluation function.
[0051] After sampling and obtaining a number of matrix pairs in step S104, these matrix pairs can be input into the proxy model in this step to obtain the evaluation value of each matrix pair output by the evaluation function included in the proxy model. The proxy model operates by taking a matrix pair as input, processing the matrix pair using the evaluation function, and calculating the output evaluation value. The evaluation value represents the overall capability of the neural network architecture corresponding to the matrix pair; a higher evaluation value indicates a more capable neural network architecture.
[0052] In this method, the evaluation function can be recorded as function f, and the input of f is the feature matrix A o With the adjacency matrix A t . Due to A o With A t They are obtained by sampling according to α and β respectively, so A o As a value of α, A t Based on this, we can regard α and β as variables, and f as a function of variables α and β, that is, the evaluation function f(α, β).
[0053] In this method, the proxy model can be trained in a variety of different ways. This specification provides a specific embodiment for reference. Specifically, several sample neural network models under the target task can be obtained; For each sample neural network model, a capability test is performed on the sample neural network model to obtain a first labeled evaluation value of the sample neural network model; a sample feature matrix and a sample adjacency matrix of the sample neural network model are determined; the sample feature matrix and the sample adjacency matrix are input into a proxy model to be trained to obtain an evaluation value to be optimized output by the proxy model according to an evaluation function to be optimized; and the proxy model is trained according to a difference between the evaluation value to be optimized and the first labeled evaluation value.
[0054] In this method, the proxy model can be pre-trained through supervised training. Figure 6 This is a schematic diagram of the process of training the agent model provided in this manual. Figure 6 As shown, during the training process, several sample neural network models for performing target tasks can be obtained through various methods such as historical data and manual construction. By performing a capability test on the sample neural network model, the first labeled evaluation value of the sample neural network model is obtained as the training label. According to the architecture of the sample neural network model, the sample feature matrix and the sample adjacency matrix of the sample neural network model can be uniquely determined, and then the sample feature matrix and the sample adjacency matrix can be input into the proxy model to obtain the evaluation value to be optimized output by the proxy model. The first labeled evaluation value is used as the training label, and the L1 loss is determined according to the difference between the first labeled evaluation value and the evaluation value to be optimized, the proxy model is trained, and the parameters of the proxy model are adjusted, that is, the evaluation function is adjusted. Among them, θ represents the weight of the proxy model.
[0055] Furthermore, the capability test of the sample neural network model can be conducted in a variety of different ways according to the specific requirements of the neural network model. This specification provides a specific method for reference. In a specific embodiment, the sample neural network model can be trained, and the trained sample neural network model can be deployed to a hardware platform; the sample neural network model is used on the hardware platform to process a pre-built test set to obtain a capability index of the sample neural network model, and the capability index includes at least inference accuracy, inference delay and power consumption; and the first labeled evaluation value of the sample neural network model is determined based on the capability index.
[0056] First of all, the sample neural network model being tested should be a trained neural network model that can perform the target task normally. Figure 7 This is a flow chart of the capability test of a sample neural network model provided in this manual. Figure 7As shown in the figure, after the trained sample neural network model is deployed on the hardware platform, it can be used to process a pre-built test set, and the capability indicators of the sample neural network model are determined based on the processing results. The capability indicators include hardware-related indicators and model-related indicators. Hardware-related indicators may include indicators such as inference latency and power consumption, while model-related indicators may include indicators such as inference accuracy.
[0057] S108: Adjusting the evaluation function included in the proxy model according to the evaluation value, and optimizing the first probability variable matrix and the second probability variable matrix according to the gradient of the adjusted evaluation function.
[0058] After obtaining the evaluation value of each sampled matrix pair in step S106, the sampled evaluation value can be used to optimize the proxy model in this step, and the optimized proxy model can be used to optimize the first probability variable matrix and the second probability variable matrix.
[0059] The method for optimizing the proxy model is similar to the method for training the proxy model, and can also be achieved through supervised training. Specifically, for each matrix pair, the ability of the neural network architecture constructed based on the matrix pair can be tested to obtain the second labeled evaluation value of the matrix pair; based on the second labeled evaluation value of each matrix pair and the evaluation value of each matrix pair output by the proxy model, the parameters of the proxy model are adjusted to adjust the evaluation function in the proxy model. The method for testing the ability of the neural network architecture corresponding to the matrix pair can also be the same as the method for testing the ability of the sample neural network model when pre-training the proxy model.
[0060] Updating the proxy model parameters can actually be considered an update to the evaluation function f fitted by the proxy model, that is, f(α, β) has changed. Therefore, based on the updated proxy model, the first probability variable matrix α and the second probability variable matrix β can be further optimized. There are also multiple ways to update α and β. This specification provides a specific method for reference. In one specific embodiment, a first sample matrix can be extracted based on the first probability variable matrix, and a second sample matrix can be extracted based on the second probability variable matrix; the first partial derivative of the evaluation function with respect to the first probability variable matrix and the second partial derivative of the evaluation function with respect to the second probability variable matrix are determined respectively; the first sample matrix is substituted into the first partial derivative to obtain the first forward propagation gradient of the first probability variable matrix, and the second sample matrix is substituted into the second partial derivative to obtain the second forward propagation gradient of the second probability variable matrix; the first probability variable matrix is optimized based on the first forward propagation gradient, and the second probability variable matrix is optimized based on the second forward propagation gradient.
[0061] Figure 8 This is a schematic diagram of a process for optimizing the first probability variable matrix and the second probability variable matrix provided in this specification. Figure 8 As shown, first, the first sample matrix S can be extracted according to the first probability variable matrix o , and extract the second sample matrix S according to the second probability variable matrix t Among them, extract S o With S t There are also many specific ways to extract S. o When , Gumbel-SoftMax can be used to extract samples from a classification distribution with class probability. The specific formula is as follows: ——Formula③ In extracting S t When , we can use the reparameterization technique, the specific formula is as follows: ——Formula④ In the above formulas ③ and ④, (i, j) represents S o (row, column) index, (h, k) represents S t The (row, column) index of ; g follows the Gumbel distribution and is calculated by g = -log(-log(u)); u follows the (0, 1) distribution and is used to introduce noise; τ is the temperature coefficient, which will continue to decay to 0 as the iteration proceeds to approach the original discrete distribution, thus approximating the discrete sampling shown in Formulas ① and ②.
[0062] The above reparameterization technique not only achieves sampling, but also makes the entire propagation path differentiable with respect to (α, β). Combined with the design of the sampling process in step S104, the value range of (α, β) is a real number within (-∞, +∞). Therefore, the optimization search of (α, β) can be implemented based on conventional optimizers such as SGD and ADAM, thereby realizing the learning of neural network architecture distribution.
[0063] While extracting the first sample matrix and the second sample matrix, the partial derivative of the evaluation function f with respect to α can be determined and the partial derivatives with respect to β Then, the first sample matrix and the second sample matrix extracted are substituted into and In the above example, we get the back propagation gradients of α and β. and , and finally the above back-propagation gradient can be used to optimize α and β respectively, changing the values of α and β. Figure 8 The @ in represents matrix multiplication.
[0064] S110: Repeat the sampling and optimization process based on the optimized first probability variable matrix and the second probability variable matrix until a preset termination condition is reached.
[0065] After completing the optimization of the first and second probability variable matrices in the previous steps, the sampling and optimization process described above, i.e., steps S104-S108 of the present method, can be re-executed. This loop can continue for multiple rounds until a preset termination condition is met. The preset termination condition can be set as needed, for example, when the number of loops reaches a specified number, or when the evaluation value of the sampled matrix pair reaches a specified threshold, etc., and this specification does not impose any specific restrictions on this.
[0066] In each cycle, the parameters and evaluation function of the proxy model will change, and the values of the first probability variable matrix and the second probability variable matrix will also change. As a result, the matrix pair obtained in the sampling process of each cycle will also be different.
[0067] In the process of continuous circulation, the optimization algorithm for the first probability variable matrix α and the second probability variable matrix β is as follows: 1. Determine the weight θ of the optimized proxy model, and do not change θ during the optimization of α and β; 2. Determine the interval training rounds oEpochs of α, the interval training rounds tEpochs of β, the total search rounds searchEpochs, the initial Gumbel temperature τ, and the learning rates η1 and η2; 3. Cyclic search round e=0~searchEpochs-1: I. Substitute formula ③ and formula ④ to extract the first sample matrix S o and the second sample matrix S t ; II. S o With S t Input the proxy model and perform forward calculation; III. Calculate the loss and pass partial derivatives and Perform back propagation and get two gradients and ; IV. Update α and β using gradient segments, with update intervals of oEpochs and tEpochs respectively. ; V. Decaying Gumbel temperature τ.
[0068] 4. Output the updated α and β.
[0069] The above process can be packaged as an algorithm with inputs α, β, f, θ; preset parameters are oEpochs, tEpochs, searchEpochs, τ, η1, η2; and output is the optimized α and β, which is implemented by a computing device.
[0070] S112: Determine a target matrix pair among all sampled matrix pairs, and construct a target neural network based on the target matrix pair.
[0071] By continuously executing the sampling and optimization process, a large number of matrix pairs can be sampled and searched while updating the evaluation function of the proxy model and the first and second probability variable matrices. This not only completes the search for the neural network architecture, but also ensures the quality of the searched neural network architecture by continuously optimizing the sampling process. At this point, the optimal target matrix pair can be determined from all the sampled matrix pairs, and the target neural network can be constructed based on the target matrix pair.
[0072] There are various ways to determine the target matrix pair, and this specification provides some specific examples for reference. As mentioned in step S106 of this method, a capability test can be performed on the sample neural network model, and the resulting capability test value can be used as the first labeled evaluation value of the sample neural network model. This capability test method can also be used to select the optimal matrix pair in this step.
[0073] Specifically, since a neural network architecture can be uniquely determined based on each matrix pair, a capability test can be performed on the neural network architecture corresponding to the matrix pair, and the test results can be recorded as the capability test value of the matrix pair. A higher capability test value indicates a stronger performance of the neural network architecture corresponding to the matrix pair. Therefore, the matrix pair with the highest capability test value can be selected from all matrix pairs as the target matrix pair, and the target neural network can be constructed based on the target matrix pair.
[0074] Furthermore, because the number of matrix pairs sampled in this method is relatively large, testing the capabilities of the neural network architectures corresponding to all matrix pairs is a time-consuming and complex process. Therefore, it is better to use the proxy model optimized through multiple iterations in this method to perform a certain degree of screening on the sampled matrix pairs, and then perform capability testing on the matrix pairs that remain after screening to reduce the workload.
[0075] Specifically, all sampled matrix pairs are input into the proxy model, and an evaluation value is obtained for each matrix pair output by the proxy model. Based on the evaluation value, the matrix pairs are screened, retaining those with an evaluation value no less than a specified threshold and discarding the remaining matrix pairs. A capability test is then performed on the retained matrix pairs, and the matrix pair with the highest capability test value is determined as the target matrix pair.
[0076] Compared to performance testing, proxy models can evaluate matrix pairs more quickly. However, while proxy models' data-fitting evaluation method can characterize the performance of the neural network architecture corresponding to a matrix pair to a certain extent, it cannot guarantee 100% accuracy. Therefore, proxy models can be used to perform a preliminary screening of a large number of matrix pairs, retaining those with high evaluation values (i.e., those with better performance). Subsequently, more accurate performance testing is performed to determine the final performance test values of each retained matrix pair, and the matrix pair with the highest performance test value is selected. This method ensures both accuracy in determining the target matrix pair and efficiency of the overall process.
[0077] Finally, the overall process of the neural network architecture search method provided in this manual is summarized. Figure 9 This is a framework flow chart of a neural network architecture search method provided in this specification, such as Figure 9 As shown, the overall implementation framework of this method can be realized according to the following steps: 1. Determine the search space, pre-train the agent model, and obtain the evaluation function f; 2. Determine the maximum number of evaluations C, the number of samples in a single round Q, but the number of round evaluations B; 3. When n≤C: Ⅰ. Sample α and β and generate matrix pairs for batches ; II. Evaluation For each matrix pair in , record the evaluation value to the historical data set; III.n←n+B; 4. Output the matrix pair with the highest historical evaluation value.
[0078] The above process can be packaged into an algorithm with inputs of α, β, and θ; preset parameters of f, C, Q, and B; and output as the matrix pair with the highest historical evaluation value, and implemented through a computing device.
[0079] Furthermore, in the above process, to further improve the quality of the ultimately constructed neural network architecture, the evaluation values output by the proxy model can be used for screening, and the neural network architectures corresponding to the screened matrix pairs can be subjected to capability testing, with the resulting capability test values recorded in the historical dataset. For example, matrix pairs whose evaluation values output by the proxy model are not less than a specified value can be retained, while matrix pairs whose evaluation values are less than the specified value can be filtered out. The neural network architectures corresponding to the retained matrix pairs can then be subjected to capability testing, with the resulting capability test values recorded in the historical dataset.
[0080] When using the neural network architecture search method provided in this specification to construct a target neural network model under a target task, after determining the search space, a first probability variable matrix and a second probability variable matrix can be constructed respectively, and a matrix pair consisting of a feature matrix and an adjacency matrix can be sampled according to the first probability variable matrix and the second probability variable matrix; the evaluation value of the matrix pair is output by the proxy model, and the proxy model is optimized based on this, and the first probability variable matrix and the second probability variable matrix are updated using the optimized proxy model; the sampling and optimization process is repeated until a preset termination condition is met; finally, the target matrix pair is selected from all the sampled matrix pairs and the target neural network model is constructed. This method uses the distribution variables of the architecture topology and the candidate operation features to realize the continuous relaxation of the original discrete architecture space; through the graph data representation and encoding of the network architecture, and the training of the proxy model based on the graph neural network, the proxy representation of the original space is realized; the reparameterization method is used to realize the differentiable search of the graph topology and the feature matrix based on the gradient respectively; the architecture optimization search is realized end to end in a manner of discrete architecture sampling, proxy model training, and architecture search online and collaborative alternation. In summary, for any architecture space, this method does not require pre-defining the number of data points for pre-training the proxy model. It only needs to determine the total number of sampled architecture points as the predefined search overhead to complete the search of the neural network architecture.
[0081] The above is the neural network architecture search method provided in this specification. Based on the same idea, this specification also provides a corresponding neural network architecture search device, such as Figure 10 shown.
[0082] Figure 10 A schematic diagram of a neural network architecture search device provided in this specification, specifically including: A determination module 200 is configured to determine a search space for a neural network architecture based on a target task, wherein the search space includes a set of candidate connection relationships and a set of candidate operations for a network layer; An initialization module 202 is configured to determine the matrix scale and row / column label information of a first probability variable matrix and a second probability variable matrix respectively according to the search space, and initialize the values of each element in the first probability variable matrix and the second probability variable matrix respectively, wherein the first probability variable matrix is a probability distribution representation of a feature matrix corresponding to the neural network architecture, and the second probability variable matrix is a probability distribution representation of an adjacency matrix corresponding to the neural network architecture; A sampling module 204 is configured to sample the first probability variable matrix and the second probability variable matrix multiple times to obtain multiple matrix pairs; An input module 206 is configured to input each matrix pair into a pre-trained proxy model to obtain an evaluation value of each matrix pair output by the proxy model according to the pre-trained evaluation function; An adjustment module 208 is configured to adjust an evaluation function included in the proxy model according to the evaluation value, and optimize the first probability variable matrix and the second probability variable matrix according to a gradient of the adjusted evaluation function; A loop module 210 is configured to repeatedly execute the sampling and optimization process based on the optimized first probability variable matrix and the second probability variable matrix until a preset termination condition is met; The construction module 212 is used to determine a target matrix pair from all sampled matrix pairs and construct a target neural network based on the target matrix pair.
[0083] Optionally, the determination module 200 is specifically used to determine the target task to which the target neural network to be constructed belongs; obtain several standard neural network models under the target task, and the standard neural network model is a neural network model artificially constructed based on prior knowledge; and determine the search space based on the network layers and atomic operations contained in each standard neural network model.
[0084] Optionally, the initialization module 202 is specifically used to determine the matrix scale of the first probability variable matrix based on the number of network layers adopted in the candidate connection relationship set and the number of candidate operations included in the candidate operation set, and to determine the matrix scale of the second probability variable matrix based on the number of network layers adopted in the candidate connection relationship set.
[0085] Optionally, the device also includes a pre-training module 214, which is specifically used to obtain several sample neural network models under the target task; for each sample neural network model, perform a capability test on the sample neural network model to obtain a first labeled evaluation value of the sample neural network model; determine the sample feature matrix and the sample adjacency matrix of the sample neural network model; input the sample feature matrix and the sample adjacency matrix into the proxy model to be trained to obtain the evaluation value to be optimized output by the proxy model according to the evaluation function to be optimized; and train the proxy model based on the difference between the evaluation value to be optimized and the first labeled evaluation value.
[0086] Optionally, the pre-training module 214 is specifically used to train the sample neural network model and deploy the trained sample neural network model to a hardware platform; the sample neural network model is used to process a pre-built test set on the hardware platform to obtain capability indicators of the sample neural network model, and the capability indicators include at least inference accuracy, inference delay and power consumption; and a first labeled evaluation value of the sample neural network model is determined based on the capability indicators.
[0087] Optionally, the adjustment module 208 is specifically used to perform a capability test on the neural network architecture constructed based on each matrix pair to obtain a second labeled evaluation value of the matrix pair; and adjust the parameters of the proxy model according to the second labeled evaluation value of each matrix pair and the evaluation value of each matrix pair output by the proxy model to adjust the evaluation function in the proxy model.
[0088] Optionally, the evaluation function is a function of the first probability variable matrix and the second probability variable matrix; The adjustment module 208 is specifically used to extract a first sample matrix based on the first probability variable matrix, and extract a second sample matrix based on the second probability variable matrix; respectively determine the first partial derivative of the evaluation function with respect to the first probability variable matrix and the second partial derivative of the evaluation function with respect to the second probability variable matrix; substitute the first sample matrix into the first partial derivative to obtain the first forward propagation gradient of the first probability variable matrix, and substitute the second sample matrix into the second partial derivative to obtain the second forward propagation gradient of the second probability variable matrix; optimize the first probability variable matrix according to the first forward propagation gradient, and optimize the second probability variable matrix according to the second forward propagation gradient.
[0089] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 Provided neural network architecture search method.
[0090] This manual also provides Figure 11 The schematic structure diagram of the electronic device shown in FIG. Figure 11 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The neural network architecture search method described above. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0091] Improvements to a technology can be clearly categorized as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with technological advancements, many process flow improvements can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can integrate a digital system onto a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using software called a "logic compiler." This is similar to the software compilers used during program development. Before compilation, the original code must be written in a specific programming language, called a hardware description language (HDL). There are many types of HDL, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0092] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.
[0093] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0094] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0095] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0096] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0097] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0099] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0100] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0101] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0102] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0103] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0104] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0105] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0106] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be encompassed within the scope of the claims of this application.
Claims
1. A neural network architecture search method, characterized in that: include: Determine a search space for a neural network architecture according to a target task, wherein the search space includes a set of candidate connection relationships and a set of candidate operations of a network layer; Determine the matrix scale and row and column label information of the first probability variable matrix and the second probability variable matrix respectively according to the search space, and initialize the value of each element in the first probability variable matrix and the second probability variable matrix respectively, wherein the first probability variable matrix is a probability distribution representation of the feature matrix of the neural network architecture, and the second probability variable matrix is a probability distribution representation of the adjacency matrix of the neural network architecture; Sampling the first probability variable matrix and the second probability variable matrix multiple times to obtain multiple matrix pairs; Inputting each matrix pair into a pre-trained proxy model respectively, and obtaining an evaluation value of each matrix pair output by the proxy model according to an evaluation function obtained by pre-training; Adjusting the evaluation function included in the proxy model according to the evaluation value, and optimizing the first probability variable matrix and the second probability variable matrix according to the gradient of the adjusted evaluation function; Repeating the sampling and optimization process based on the optimized first probability variable matrix and the second probability variable matrix until a preset termination condition is reached; A target matrix pair is determined among all sampled matrix pairs, and a target neural network is constructed based on the target matrix pair.
2. The method according to claim 1, characterized in that Determine the search space of the neural network architecture based on the target task, including: Determine the target task to which the target neural network to be constructed belongs; Acquire several standard neural network models under the target task, wherein the standard neural network model is a neural network model artificially constructed based on prior knowledge; The search space is determined based on the network layers and atomic operations included in each standard neural network model.
3. The method according to claim 1, characterized in that Determining the matrix scale and row / column label information of the first probability variable matrix and the second probability variable matrix respectively according to the search space specifically includes: The matrix scale of the first probability variable matrix is determined according to the number of network layers adopted in the candidate connection relationship set and the number of candidate operations included in the candidate operation set, and the matrix scale of the second probability variable matrix is determined according to the number of network layers adopted in the candidate connection relationship set.
4. The method according to claim 1, characterized in that Pre-trained proxy models, including: Obtaining several sample neural network models under the target task; For each sample neural network model, a capability test is performed on the sample neural network model to obtain a first labeled evaluation value of the sample neural network model; Determine a sample feature matrix and a sample adjacency matrix of the sample neural network model; Inputting the sample feature matrix and the sample adjacency matrix into the proxy model to be trained, and obtaining the evaluation value to be optimized output by the proxy model according to the evaluation function to be optimized; The proxy model is trained according to the difference between the evaluation value to be optimized and the first labeled evaluation value.
5. The method according to claim 4, characterized in that The capability test of the sample neural network model is performed to obtain a first labeled evaluation value of the sample neural network model, specifically including: Training the sample neural network model, and deploying the trained sample neural network model to a hardware platform; Using the sample neural network model to process a pre-built test set on the hardware platform to obtain capability indicators of the sample neural network model, wherein the capability indicators include at least inference accuracy, inference latency, and power consumption; A first annotation evaluation value of the sample neural network model is determined according to the capability indicator.
6. The method according to claim 1, characterized in that Adjusting the evaluation function included in the proxy model according to the evaluation value specifically includes: For each matrix pair, a capability test is performed on a neural network architecture constructed according to the matrix pair to obtain a second labeled evaluation value of the matrix pair; According to the second labeled evaluation value of each matrix pair and the evaluation value of each matrix pair output by the proxy model, the parameters of the proxy model are adjusted to adjust the evaluation function in the proxy model.
7. The method according to claim 1, characterized in that The evaluation function is a function of the first probability variable matrix and the second probability variable matrix; Optimizing the first probability variable matrix and the second probability variable matrix according to the gradient of the adjusted evaluation function specifically includes: Extracting a first sample matrix according to the first probability variable matrix, and extracting a second sample matrix according to the second probability variable matrix; respectively determining a first partial derivative of the evaluation function with respect to the first probability variable matrix and a second partial derivative of the evaluation function with respect to the second probability variable matrix; Substituting the first sample matrix into the first partial derivative to obtain a first forward propagation gradient of the first probability variable matrix, and substituting the second sample matrix into the second partial derivative to obtain a second forward propagation gradient of the second probability variable matrix; The first probability variable matrix is optimized according to the first forward propagation gradient, and the second probability variable matrix is optimized according to the second forward propagation gradient.
8. A neural network architecture search device, characterized in that: include: A determination module, used to determine a search space of a neural network architecture according to a target task, wherein the search space includes a set of candidate connection relationships and a set of candidate operations of a network layer; An initialization module, used to determine the matrix scale and row / column label information of the first probability variable matrix and the second probability variable matrix respectively according to the search space, and initialize the values of each element in the first probability variable matrix and the second probability variable matrix respectively, wherein the first probability variable matrix is a probability distribution representation of the feature matrix corresponding to the neural network architecture, and the second probability variable matrix is a probability distribution representation of the corresponding adjacency matrix in the neural network architecture; A sampling module, used for sampling the first probability variable matrix and the second probability variable matrix multiple times to obtain multiple matrix pairs; An input module, used to input each matrix pair into a pre-trained proxy model, and obtain an evaluation value of each matrix pair output by the proxy model according to an evaluation function obtained by pre-training; An adjustment module, configured to adjust an evaluation function included in the proxy model according to the evaluation value, and optimize the first probability variable matrix and the second probability variable matrix according to a gradient of the adjusted evaluation function; A loop module, used to repeatedly execute the sampling and optimization process based on the optimized first probability variable matrix and the second probability variable matrix until a preset termination condition is reached; A construction module is used to determine a target matrix pair among all sampled matrix pairs, and to construct a target neural network based on the target matrix pair.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Intelligent generation method and system for topological relation of canal system
CN121211629A
Power distribution network operation risk identification method based on optimized proxy neural architecture search
CN122244641A