Differential adaptive neural network architecture search method based on zero-order approximation

Through the differentiable adaptive neural network architecture search method based on zero-order approximation, the problem of insufficient precision and time overhead in the optimization process of DARTS algorithm is solved, and an efficient neural network architecture search adapted to resources of different platforms is realized.

CN120146103APending Publication Date: 2025-06-13TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510264197.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing DARTS automatic architecture search algorithm is not accurate enough in the optimization process, has a large time overhead, and cannot adapt to the computing power limitations of different platforms, resulting in inefficiency, low accuracy and low availability.

Method used

A differentiable adaptive neural network architecture search method based on zero-order approximation is adopted. By constructing a resource-constrained neural network operation set, variable searches are set to adapt to different resource situations, and structural parameters are updated using gradient descent, and neural network operations are selected according to the probability distribution to generate complex neural networks.

Benefits of technology

This method can adapt to resource constraints of different computing platforms and improve the efficiency and accuracy of neural network architecture search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146103A_ABST
    Figure CN120146103A_ABST
Patent Text Reader

Abstract

The invention provides a differential adaptive neural network architecture search method based on zero-order approximation, and belongs to the technical field of artificial neural networks. The resource-constrained neural network operation set is constructed through the selected multiple neural network operations, then the structure parameter set is constructed according to the neural network operation set, variable search suitable for different resource conditions is set, the structure parameters are updated through the gradient descent method, and the structure parameters are optimized. And according to the probability distribution, selecting a corresponding neural network operation in the structure parameter to which each connection edge in the updated structure parameter set belongs to obtain a complex neural network. The differential adaptive neural network architecture search method based on zero-order approximation can adapt to different computing platform resource constraints, and meanwhile, the efficiency of searching a complex neural network and the accuracy of the complex neural network are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial neural networks, and particularly relates to a differentiable adaptive neural network architecture search method based on zero-order approximation. Background Art

[0002] Deep learning is a cutting-edge artificial intelligence technology that can efficiently extract features and analyze content from massive amounts of data. At present, methods in this field have achieved outstanding results in tasks such as image recognition and video understanding. However, the process of designing intelligent models for these tasks requires a large amount of professional knowledge and manual input, which is the biggest obstacle to the large-scale application of deep learning technology in practical problems.

[0003] To solve this problem, leading research experts have begun to consider using the Neural Architecture Search (NAS) method. The DARTS algorithm is a representative differentiable architecture search strategy in NAS. It assigns continuous weights to candidate operations and performs weighted mixing during search, thereby making the search space continuous, forming a differentiable two-layer optimization problem, and using gradient descent for performance optimization. After optimization, DARTS selects the operation with the largest weight from the mixed operations, thereby determining a high-performance neural network framework with a complex topological structure in the rich search space.

[0004] However, the existing DARTS automatic architecture search algorithm is not precise enough in the process of solving the optimization problem, requires a large time overhead, and cannot adapt to the computing power limitations of different platforms. There are still defects of low efficiency, low accuracy, and low usability. Summary of the Invention

[0005] The present invention is made to solve the above problems, and its purpose is to provide a differentiable adaptive neural network architecture search method based on zero-order approximation.

[0006] The present invention provides a differentiable adaptive neural network architecture search method based on zero-order approximation, which has the following features. It is used to obtain a corresponding complex neural network through a training data set, a validation data set, and a neural network operations to obtain a corresponding complex neural network, including the following steps: S10, select b from the predefined a neural network operations as the neural network operation set , which includes a variable-size convolution operation; S20, construct an initial neural network and set the upper limit C of resource consumption U and the lower limit C L , control parameter λ 1 and λ 2 , set the number of iteration rounds n to 1. Among them, the initial neural network includes E groups of neuron groups and E - 1 feature reduction layers, and each group of neuron groups contains S ea neural unit, the neural unit includes J connection edges; S30, respectively corresponding the J connection edges of the neural unit to b neural network operations in the neural network operation set to form a hybrid neural network operation, and obtain the corresponding E*J*b structural parameters denoted as set δ. When the iteration round n≥θ, activate the variable-size search and variable-depth search, and set the K different-size convolution operations to which the variable-size convolution operations in the b neural network operations of each of the J connection edges belong, to obtain the corresponding E groups and each group has J*K structural parameters denoted as set β. At the same time, for the used neural unit, obtain a total of E groups and each group has S e parameters denoted as set γ, construct the structural parameter set {δ,β,γ}, calculate the total resource occupancy C, e = 1, 2, …, E, and set the upper-layer problem loss function:

[0007]

[0008] where θ is a preset iteration threshold, and w is the model parameter set; S40, set the temperature constant τ, the adjustment factor a, and the adjustment factor m, set the minimum value μ and the random unit vector u according to the structural parameter set α, and obtain the fluctuating structural parameter set according to the structural parameter set α, the minimum value μ, and the random unit vector u Integrate all the weight parameter sets of the initial neural network and denote them as the model parameter set w, and set the fluctuating model parameter set according to w According to the training data set, the structural parameter set α, and the fluctuating structural parameter set Randomly select a batch of data each time, and through the lower-layer problem loss function and the gradient descent method, obtain the updated model parameter set w′ and the updated fluctuating model parameter set S50, input the validation data set into the initial neural network, randomly select a batch of data, calculate the approximate gradient of the structural parameter set α through the upper-layer problem loss function, and use the gradient descent method to update and obtain the structural parameter set α′, and use the structural parameter set α′ to update the structural parameter set α; S60, repeat steps S40~S50 until the training data set is traversed, and add 1 to the iteration round n; S70, repeat steps S40 - S60 until the current iteration round n reaches the maximum iteration round q; S80, based on the probability distribution formed by the probability normalization function for the b structural parameters corresponding to each connection edge in the structural parameter set α′, sample the neural network operation corresponding to a certain probability value as the neural network operation of this connection edge, then obtain the complex neural network.

[0009] In the differentiable adaptive neural network architecture search method based on zero-order approximation provided by the present invention, it may further have the following features: Among them, the neural network operation is any operation that preserves the data dimension. The neural network operations include zero operation, skip connection operation, 1×1 convolution operation, size-variable convolution operation, and 3×3 average pooling operation. The loss functions of the training dataset and the validation dataset are both cross-entropy functions. In step S20, the neural unit is a directed acyclic graph composed of a series of data blocks and neural network operations. The data blocks inside the unit have sequential numbers, and the smallest and largest numbers respectively constitute the input and output of the unit.

[0010] In the differentiable adaptive neural network architecture search method based on zero-order approximation provided by the present invention, it may further have the following features: Among them, in step S30, the hybrid neural network operation is a weighted neural network operation. For the edge (i, j) between any i-th and j-th data blocks, the calculation result of its corresponding hybrid neural network operation is: In the formula, represents the structure parameter vector of the (i, j) edge of the unit belonging to the neural unit group e in the structure parameter set α, g(·) represents the probability normalization function combined with the annealing strategy, and g o represents taking the element corresponding to the neural network operation o of the operation result of this function.

[0011] In the differentiable adaptive neural network architecture search method based on zero-order approximation provided by the present invention, it may further have the following features: Among them, the probability normalization function g(·) combined with the annealing strategy is denoted as sparsemax′(·). Its input and output values are both vectors. For a certain vector σ, the probability normalization function sparsemax′(σ) combined with the annealing strategy is:

[0012] In the formula, σ is the input vector, argmin represents the parameter value at which a function obtains the minimum value in its domain, Δ L-1 is an L - 1 dimensional simplex, represents the L-dimensional real vector space, the superscript T is the transpose operation, ||p - σ / (τ * a n / / m )|| 2 represents the squared Euclidean distance between the vector p and σ / (τ * a n / / m ). The integer division symbol / / represents the quotient obtained by dividing two integers and truncating it towards zero. According to the set temperature constant τ and adjustment factors a, m, the input vector σ is selected as required. Through the probability normalization function combined with the annealing strategy, the vector p is calculated, and p is specified as the probability normalization result sparsemax′(σ) of the input vector σ, generating the required neural network weighted probability.

[0013] In the differentiable adaptive neural network architecture search method based on zero-order approximation provided by the present invention, it may further have the following characteristics: Among them, in step S30, the variable-size search includes K different sizes of convolutional kernels. When n is less than the preset iteration threshold θ, the variable-size convolution operation on the j-th side of any neuron in the unit group e only uses the largest convolutional kernel for calculation, otherwise its operation result is: In the formula, represents the structure parameter vector corresponding to the j-th side in the structure parameter matrix belonging to the neuron group e in the β set, and conv k (·) represents convolution with a convolutional kernel size of k*k, and g k represents selecting the k-th element of this normalization function. The depth-variable search is for the S e neurons of the e-th neuron group. The neurons in the group are connected in series, and the output of the previous neuron is used as the input of the next neuron. When n is less than the preset iteration threshold θ, the depth-variable operation result of the unit group e only uses the output of the last neuron as the final output of this unit group and enters the next-level calculation, otherwise its operation result is: In the formula, γ e represents the structure parameter vector belonging to the neuron group e in the γ set, represents the output of the neuron with the sequence number d in the unit group e, is used to represent the number of neurons in the e-th neuron group, and g d represents selecting the d-th element of this normalization function.

[0014] In the differentiable adaptive neural network architecture search method based on zero-order approximation provided by the present invention, it may further have the following characteristics: Among them, in step S30, the total resource occupancy C is a quantitative characteristic of the operation resource usage by a statistical model. For each neural network operation within the e-th neuron group, the expected resource occupancy for a certain topological structure edge (i, j) is: In the formula, represents the resource occupancy of the operation o on the (i, j) edge, represents the structure parameter of the (i, j) structure edge belonging to the e-th neuron group, and the resource occupancy of the variable-size convolution operation is expressed as: In the formula, represents the resource occupancy of the operation with a convolutional kernel of size k, and the resource occupancy of the depth-variable operation of the unit group e is expressed as: In the formula, represents the total resource occupancy of the first d neurons in the e-th neuron group.

[0015] In the differentiable adaptive neural network architecture search method based on zero-order approximation provided by the present invention, it may also have the following features: Among them, in step S40, the minimum value is the vector obtained by concatenating all vectors of the structure parameter set α in the dimension with a width of 1, and the random unit vector u′ is a random vector with the same number of elements as and each element follows a normal distribution respectively. The fluctuating structure parameter set u″ is a set of variables obtained by aligning the dimensions of the random unit vector u with the structure parameter set α, is the element-wise addition.

[0016] In the differentiable adaptive neural network architecture search method based on zero-order approximation provided by the present invention, it may also have the following features: Among them, in step S40, the method for updating the model parameter set w′ and the updated fluctuating model parameter set includes the following sub-steps: A01, set the iteration round t = 1; A02, randomly select a batch of data in the training dataset and input it into the initial neural network, combine the structure parameter set α and the fluctuating structure parameter set and according to the lower-layer problem loss function and the learning rate η w , update the model parameter set w and the fluctuating model parameter set and increment the iteration round t by 1; A03, repeat step A02 until the iteration round t reaches the maximum iteration round T, and then output the final updated model parameter set w′ and the updated fluctuating model parameter set

[0017] In the differentiable adaptive neural network architecture search method based on zero-order approximation provided by the present invention, it may also have the following features: Among them, in step A02, the lower-layer problem loss function is the loss function and the loss function The value of the loss function is the calculation result of the model parameter set w and the structure parameter set α according to the batch data in the training dataset. According to the learning rate η w and the loss function the gradient of the model parameter set w is calculated by the gradient descent method to obtain the updated model parameter set w″, which is used as the model parameter set w at iteration round t + 1. The value of the loss function is the calculation result of the fluctuating model parameter set and the fluctuating structure parameter set according to the batch data in the training dataset. According to the learning rate η w and the loss function for the fluctuating model parameter set Gradient By the gradient descent method, the updated fluctuation model parameter set is calculated As the fluctuation model parameter set at iteration round t + 1

[0018] In the differentiable adaptive neural network architecture search method based on zero - order approximation provided by the present invention, it can also have the following feature: Among them, in step S50, the approximate gradient of the structure parameter set α The calculation formula is: In the formula, Is the upper - layer problem loss function The gradient of the model parameter set w, w′(α) indicates that in this gradient calculation, the updated model parameter set w′ is regarded as related to the structure parameter set α The gradient of the structure parameter set α, w′() indicates that in this gradient calculation, the updated model parameter set w′ is regarded as a fixed value and is independent of the structure parameter set α, and the superscript T is the transpose operation. According to the set learning rate η α And the approximate gradient of the structure parameter set α By the gradient descent method, the structure parameter set α′ is calculated

[0019] Functions and effects of the invention

[0020] According to a differentiable adaptive neural network architecture search method based on zero - order approximation involved in the present invention, since a resource - constrained neural network operation set is constructed by selecting multiple neural network operations, and then a structure parameter set is constructed according to the neural network operation set, a variable search adapting to different resource situations is set, the structure parameters are updated by the gradient descent method, and according to the probability distribution, the neural network operations corresponding to the structure parameters to which each connection edge in the updated structure parameter set belongs are selected, a complex neural network is obtained

[0021] Therefore, the differentiable adaptive neural network architecture search method based on zero - order approximation of the present invention can adapt to the resource constraints of different computing platforms, and at the same time improve the efficiency of searching for complex neural networks and the accuracy of complex neural networks Brief description of the drawings

[0022] Figure 1 Is a flowchart of a differentiable adaptive neural network architecture search method based on zero - order approximation of an embodiment of the present invention

[0023] Figure 2 Is a schematic diagram of the search architecture of a complex neural network model for image recognition in a test example of the present invention

[0024] Figure 3It is a schematic diagram of a certain neuron structure of the complex neural network model for image recognition in the test example of the present invention. Detailed implementation manners

[0025] In order to make the technical means, creative features, achieved purposes and effects realized by the present invention easy to understand, the following embodiments will specifically elaborate on a differentiable adaptive neural network architecture search method based on zero-order approximation of the present invention in conjunction with the accompanying drawings.

[0026] <Embodiment>

[0027] Figure 1 It is a flowchart of a differentiable adaptive neural network architecture search method based on zero-order approximation in the embodiment of the present invention.

[0028] As Figure 1 shown, the present embodiment provides a differentiable adaptive neural network architecture search method based on zero-order approximation, which is used to obtain a corresponding complex neural network through a training data set, a validation data set and a neural network operations. (The loss functions of the training data set and the validation data set are both cross-entropy functions)

[0029] A differentiable adaptive neural network architecture search method based on zero-order approximation in the present embodiment includes the following steps:

[0030] S10. Select b neural network operations from the predefined a neural network operations as the neural network operation set which includes a variable-size convolution operation.

[0031] Among them, the neural network operation is any operation that preserves the data dimension, including zero operation, skip connection operation, 1×1 convolution operation, variable-size convolution operation and 3×3 average pooling operation.

[0032] S20. Construct an initial neural network and set the upper limit C of resource consumption U and the lower limit C L , control parameters λ 1 and λ 2 , and set the number of iterations n to 1.

[0033] Among them, the initial neural network includes E groups of neuron groups and E-1 feature reduction layers. Each group of neuron groups contains S e neurons, and each neuron contains J connection edges.

[0034] A neuron is a directed acyclic graph composed of a series of data blocks and neural network operations. The data blocks inside the unit have sequential numbers, and the smallest and largest numbers respectively constitute the input and output of the unit.

[0035] S30. Establish the upper-layer problem loss function, including the following sub-steps S31 to S33:

[0036] S31. Correspond the J connection edges of the neuron unit to b neural network operations in the neural network operation set to form a hybrid neural network operation, and obtain the corresponding E * J * b structural parameters denoted as set δ.

[0037] Among them, the hybrid neural network operation is a weighted neural network operation. For the edge (i, j) between any i-th and j-th data blocks, the calculation result of its corresponding hybrid neural network operation is:

[0038]

[0039] In the formula, represents the structural parameter vector of the (i, j) edge of the unit belonging to the neuron unit group e in the structural parameter set α, and g(·) represents the probability normalization function combined with the annealing strategy. g o represents taking the element corresponding to the operation result of this function for the neural network operation o.

[0040] Denote the probability normalization function g(·) combined with the annealing strategy as sparsemax′(·). Its input and output values are both vectors. For a certain vector σ, the probability normalization function sparsemax′(σ) combined with the annealing strategy is:

[0041]

[0042] In the formula, σ is the input vector, argmin represents the parameter value at which a function obtains the minimum value in its domain, and Δ L-1 is an L - 1 dimensional simplex, represents the L - dimensional real vector space, the superscript T is the transpose operation, and ||p - σ / (τ * a n / / m )|| 2 represents the squared Euclidean distance between the vector p and σ / (τ * a n / / m ). The integer division symbol / / represents the quotient obtained by dividing two integers and truncating it towards zero.

[0043] In the above formula, according to the set temperature constant τ, adjustment factors a and m, select the input vector σ as required, and calculate the vector p through the probability normalization function combined with the annealing strategy. Specify p as the probability normalization result sparsemax′(σ) of the input vector σ to generate the required neural network weighted probability.

[0044] S32. When the iteration round n ≥ θ, activate the variable-size search and variable-depth search. For the variable-size convolution operations among the b neural network operations of each of the J connection edges, set the K different sizes of convolution operations to which they belong, and obtain the corresponding E groups, with J*K structural parameters in each group, denoted as the set β. At the same time, for the used neurons, obtain a total of E groups, with S e parameters in each group, denoted as the set γ, and construct the set of structural parameters {δ, β, γ}, and calculate the total resource occupancy C, where e = 1, 2,..., E.

[0045] In this step, the variable-size search includes K different sizes of convolution kernels. When n is less than the preset iteration threshold θ, the variable-size convolution operation on the j-th side of any neuron in the unit group e only uses the largest convolution kernel for calculation; otherwise, its operation result is:

[0046]

[0047] In the formula, represents the structural parameter vector corresponding to the j-th side in the structural parameter matrix belonging to the neuron group e in the β set, and conv k (·) represents the convolution with a convolution kernel size of k*k, and g k represents selecting the k-th element of this normalization function.

[0048] The variable-depth search is for the S e neurons in the e-th neuron group. The neurons in the group are connected in series, and the output of the previous neuron is used as the input of the next neuron. When n is less than the preset iteration threshold θ, the variable-depth operation result of the unit group e only uses the output of the last neuron as the final output of this unit group and enters the next-level calculation; otherwise, its operation result is:

[0049]

[0050] In the formula, γ e represents the structural parameter vector belonging to the neuron group e in the γ set, represents the output of the neuron numbered d in the unit group e, is used to represent the number of neurons in the e-th neuron group, and g d represents selecting the d-th element of this normalization function.

[0051] The total resource occupancy C is a quantitative characteristic of the operation resource usage by a statistical model. For each neural network operation within the e-th group of neuron groups, the expected resource occupancy for a certain topological structure edge (i, j) is:

[0052]

[0053] In the formula, The resource occupancy of operation o on the (i, j) edge, represents the structural parameter of the (i, j) structural edge belonging to the e-th group of neural units,

[0054] The resource occupancy of the size-variable convolution operation is expressed as:

[0055]

[0056] In the formula, represents the resource occupancy of the operation of the convolution kernel of size k,

[0057] The resource occupancy of the depth-variable operation of the e-th unit group is expressed as:

[0058]

[0059] In the formula, represents the total resource occupancy of the first d neural units within the e-th group of neural units.

[0060] S33, set the upper-layer problem loss function:

[0061]

[0062] Among them, the upper-layer problem loss function in this embodiment is the cross-entropy function, θ is the preset iteration threshold, w is the set of model parameters, and the upper-layer loss function in the value of is the calculation result of the set of model parameters w and the set of structural parameters α based on the validation data set.

[0063] S40, parameter setting and obtaining the updated set of model parameters w′ and the updated fluctuating set of model parameters includes the following sub-steps S41 to S44:

[0064] S41, set the temperature constant τ, the adjustment factor a, and the adjustment factor m.

[0065] S42, set the minimum value μ and the random unit vector u according to the set of structural parameters α:

[0066] The minimum value is the vector obtained by splicing all the vectors of the set of structural parameters α in the dimension with a width of 1.

[0067] The random unit vector u′ is a random vector with the same number of elements as and each element follows a normal distribution independently.

[0068] S43. Obtain the set of fluctuating structure parameters according to the set of structure parameters α, the minimum value μ, and the random unit vector u u″ is a set of variables obtained by aligning the dimensions of the random unit vector u with the set of structure parameters α, which is the element-wise addition.

[0069] S44. Integrate all the weight parameter sets of the initial neural network and denote them as the model parameter set w. Set the fluctuating model parameter set according to w

[0070] S45. According to the training data set, the set of structure parameters α, and the set of fluctuating structure parameters Randomly select a batch of data each time, and through the lower-layer problem loss function and the gradient descent method, obtain the updated model parameter set w′ and the updated fluctuating model parameter set including the following sub-steps A01 to A03:

[0071] A01. Set the iteration round t = 1.

[0072] A02. Randomly select a batch of data from the training data set and input it into the initial neural network, combining the set of structure parameters α and the set of fluctuating structure parameters and update the model parameter set w and the fluctuating model parameter set according to the lower-layer problem loss function and the learning rate η w , and increment the iteration round t by 1.

[0073] where the learning rate is the minimum value of the learning rate, is the maximum value of the learning rate.

[0074] The lower-layer problem loss function is the loss function and the loss function The lower-layer problem loss function is the cross-entropy function.

[0075] The loss function is the calculation result of the model parameter set w and the set of structure parameters α according to the batch data in the training data set. According to the learning rate η w and the loss function the gradient of the model parameter set w Through the gradient descent method, calculate the updated model parameter set w″ as the model parameter set w at iteration round t + 1:

[0076]

[0077] The loss function is the value of the fluctuating model parameter set according to the batch data in the training data set and the set of fluctuation structure parameters According to the calculation result, and the learning rate η w and the loss function calculate the gradient of the set of fluctuation model parameters Through the gradient descent method, calculate and obtain the updated set of fluctuation model parameters as the set of fluctuation model parameters at the iteration round t + 1

[0078]

[0079] A03. Repeat step A02 until the iteration round t reaches the maximum iteration round T, and then output the final updated set of model parameters w′ and the updated set of fluctuation model parameters

[0080] S50. Input the validation data set into the initial neural network, randomly select a batch of data, calculate the approximate gradient of the set of structure parameters α through the upper-layer problem loss function, and use the gradient descent method to update and obtain the set of structure parameters α′, and use the set of structure parameters α′ to update the set of structure parameters α.

[0081] Among them, the approximate gradient of the set of structure parameters α The calculation formula is:

[0082]

[0083]

[0084] In the formula, is the upper-layer problem loss function is the gradient of the set of model parameters w, and w′(α) means that in the calculation of this gradient, the updated set of model parameters w′ is regarded as related to the set of structure parameters α, is the gradient of the set of structure parameters α, and w′() means that in the calculation of this gradient, the updated set of model parameters w′ is regarded as a fixed value and is independent of the set of structure parameters α. The superscript T is the transpose operation.

[0085] According to the set learning rate η α and the approximate gradient of the set of structure parameters α Through the gradient descent method, calculate and obtain the set of structure parameters α′.

[0086] Specifically, in this embodiment, the gradient descent method is the adaptive moment estimation algorithm with given hyperparameters.

[0087] S60. Repeat steps S40 - S50 until the training data set is traversed, and increment the iteration round n by 1. ​

[0088] S70. Repeat steps S40 - S60 until the current iteration round n reaches the maximum iteration round q.

[0089] S80. Based on the probability distribution formed by the b structural parameters corresponding to each connection edge in the set of structural parameters α′ through the probability normalization function, sample a neural network operation corresponding to a certain probability value as the neural network operation of this connection edge, then a complex neural network is obtained.

[0090] This embodiment also provides a differentiable adaptive neural network architecture search model based on zero - order approximation, which uses the differentiable adaptive neural network architecture search method in this embodiment, including an operation selection module, an initialization module, a dynamic search module, a parameter optimization module, and an architecture generation module.

[0091] The operation selection module is used to construct a set of neural network operations according to the method in step S10.

[0092] The initialization module is used to construct an initial neural network and set the upper limit C of resource consumption according to the method in step S20. U and the lower limit C L , control parameter λ 1 and λ 2 , set the iteration round n to 1.

[0093] The dynamic search module is used to activate size - variable search and depth - variable search and set the upper - layer problem loss function according to the method in step S30.

[0094] The parameter optimization module is used to update the model parameters through the fluctuating structure and gradient descent according to the method in steps S40 - S70.

[0095] The architecture generation module is used to determine the final neural network based on probability distribution sampling according to the method in step S80.

[0096] <Test Example>

[0097] This test example uses a differentiable adaptive neural network architecture search model based on zero - order approximation in the embodiment and conducts tests according to a differentiable adaptive neural network architecture search method in the embodiment.

[0098] Specifically, for the problem scenario of image recognition, experiments are carried out based on three sub - datasets PathMNIST, BloodMNIST, and OrganAMNIST of the MedMNIST dataset. For each of these sub - datasets, all data except the test set is used, equally divided into a training set and a validation set, and a construction is made including E = 3 groups of neuron groups and each group S e= 3 neural units, a total of 2 feature reduction layers. The neural unit contains J = 6 connection edges, and the initial iteration round n is set to 1.

[0099] Set of neural network operations There are b = 5 neural network operations, specifically zero operation, skip connection operation, 1×1 convolution operation, size-variable convolution operation, and 3×3 average pooling operation. Among them, the zero operation means multiplying the data by 0, and the size-variable convolution contains a total of K = 3 types of convolution kernels with widths of 7×7, 5×5, and 3×3.

[0100] Set T = 10, set the learning rate η α = 0.015 and q = 40, set the threshold θ = 20, set the temperature constant τ = 1.5, adjustment factors a = 0.75, m = 5, control parameter λ 1 = λ 2 = 15 / (τ * a n / / m ), and obtain a complex neural network model for image recognition as the model of this test case. Its search architecture is as Figure 2 shown, and the structure of a certain neural unit in the model is as Figure 3 shown.

[0101] Specifically, when setting the control parameter λ 1 = λ 2 = 0, 3 independent experiments are carried out for each selected dataset, and 300 probability samplings are performed based on the results of each search model, obtaining a total of 900 models. Calculate their total number of parameters, and set the three percentile combinations of [0, 20%], [40%, 60%], and [80%, 95%] as the upper resource consumption limit C U and the lower limit C L of three different sizes of models S, M, and L respectively. Then set the control parameters to the required values and perform an architecture search task adapted to different resource amounts.

[0102] Perform model search on the above zero-order approximation-based differentiable adaptive neural network architecture search method / model, that is, the algorithm of the present invention (named ZO-DARTS++), the DARTS algorithm DARTS that uses second-order expansion approximation for the lower-layer problem of double-layer optimization, the ZO-DARTS algorithm based on DARTS and zero-order approximation, and the structure search algorithm MiLeNAS that transforms the DARTS double-layer optimization into a single-layer optimization problem. Compare the consumption time, that is, the search time, of the search models of each algorithm. The image recognition accuracy and search time of the search models are obtained by sampling the original results with probabilities three times after three experiments, and the average value of the results is taken after 300 rounds of retraining of the obtained model structures. Then the search time (seconds) of each algorithm is shown in Table 1 below:

[0103] Table 1 (Search time consumption of each algorithm)

[0104] Algorithm Name PathMNIST BloodMNIST OrganAMNIST DARTS 3470.6 1168.0 1881.4 MiLeNAS 2902.6 946.8 1469.3 ZO-DARTS 1589.8 503.3 735.2 ZO-DARTS++ 2670.6 576.4 1121.7

[0105] The recognition accuracy rate (%) of the model images obtained by the search is shown in Table 2 below:

[0106] Table 2 (Recognition accuracy rate of the model images obtained by the search)

[0107] Algorithm Name PathMNIST BloodMNIST OrganAMNIST DARTS 78.6±3.50 86.7±3.74 92.8±0.86 MiLeNAS 81.1±4.07 86.8±3.76 93.1±0.85 ZO-DARTS 81.6±5.39 85.6±5.82 93.1±1.32 ZO-DARTS++ 80.7±2.83 89.2±1.99 93.3±1.02 ZO-DARTS++S 80.6±3.30 89.2±2.84 92.9±0.40 ZO-DARTS++M 80.8±4.59 89.1±2.46 93.2±0.74 ZO-DARTS++L 80.4±2.41 91.2±1.37 94.2±0.57

[0108] Meanwhile, the average size (in millions of parameters) of each model is shown in Table 3 below:

[0109] Table 3 (Average size of each model)

[0110] Algorithm Name PathMNIST BloodMNIST OrganAMNIST DARTS 0.21±0.1 0.25±0.1 0.22±0.1 MiLeNAS 0.30±0.1 0.21±0.2 0.30±0.1 ZO-DARTS 0.23±0.1 0.24±0.1 0.30±0.1 ZO-DARTS++ 0.45±0.4 0.24±0.1 0.75±0.6 ZO-DARTS++S 0.11±0.0 0.12±0.0 0.21±0.1 ZO-DARTS++M 0.34±0.0 0.21±0.0 0.48±0.1 ZO-DARTS++L 0.82±0.1 0.48±0.1 1.12±0.1

[0111] In the above table, ZO-DARTS++ is the algorithm and corresponding model of the present invention. The suffixes S, M, and L respectively represent the corresponding models obtained by the algorithm of the present invention after setting three different upper and lower limits of model size from small to large.

[0112] It can be seen from the above table that the model obtained by the search of the algorithm of the present invention has a higher accuracy rate for the model for image recognition on the selected data set compared with other algorithms, the required search time consumption is moderate, and it has the characteristic that the model size can be freely adjusted. Among them, the ZO-DARTS++S model with the smallest size has nearly 50% less parameter quantity compared with other methods, and is more excellent than other methods in terms of the comprehensive indexes of performance, efficiency, and size.

[0113] Functions and effects of the embodiment

[0114] According to the zero-order approximation-based differentiable adaptive neural network architecture search method involved in this embodiment, since a set of neural network operations constrained by resources is constructed through a selected plurality of neural network operations, and then a set of structural parameters is constructed according to the set of neural network operations, a variable search adapting to different resource situations is set, the structural parameters are updated by the gradient descent method, and the neural network operations corresponding to the structural parameters to which each connection edge in the updated set of structural parameters belongs are selected according to the probability distribution, a complex neural network is obtained. Therefore, the zero-order approximation-based differentiable adaptive neural network architecture search method of this embodiment can adapt to the resource constraints of different computing platforms, and at the same time improve the efficiency of searching for complex neural networks and the accuracy of complex neural networks.

[0115] Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for searching a differentiable adaptive neural network architecture based on zero-order approximation, characterized in that: It is used to obtain the corresponding complex neural network through the training data set, the verification data set and a neural network operation, including the following steps: S10, select b from the predefined a neural network operations as the neural network operation set It contains a resizable convolution operation; S20, build the initial neural network and set the resource consumption limit C U With lower limit C L , control parameters λ1 and λ2, set the iteration round n to 1, The initial neural network includes E groups of neural units and E-1 feature reduction layers, each of which includes S e A neural unit, wherein the neural unit comprises J connecting edges; S30, respectively corresponding the J connecting edges of the neural unit to the neural network operation set b neural network operations are formed to form a hybrid neural network operation, and the corresponding E*J*b structural parameters are recorded as a set δ. When the iteration round n ≥ θ, activate the variable-size search and variable-depth search, and set the K convolution operations of different sizes subordinate to the variable-size convolution operations in the b neural network operations of the J connecting edges to obtain the corresponding E groups, and each group of J*K structural parameters is recorded as a set β. At the same time, for the neural units used, a total of E groups and each group of S e The parameters are recorded as set γ, and the structural parameter set {δ, β, γ} is constructed to calculate the total resource usage C, e = 1, 2, ..., E, Set the upper problem loss function: Among them, θ is the preset iteration threshold, and w is the model parameter set; S40, setting the temperature constant τ, the adjustment factor a and the adjustment factor m, According to the structural parameter set α, the minimum value μ and the random unit vector u are set. According to the structural parameter set α, the minimum value μ and the random unit vector u, the fluctuating structural parameter set is obtained. All weight parameter sets of the initial neural network are integrated as the model parameter set w, and the fluctuation model parameter set is set according to w According to the training data set, the structural parameter set α and the fluctuation structural parameter set Each time, a batch of data is randomly selected, and the updated model parameter set w′ and the updated fluctuation model parameter set are obtained through the lower-level problem loss function and gradient descent method. S50, inputting the verification data set into the initial neural network, randomly selecting batch data, calculating the approximate gradient of the structure parameter set α through the upper problem loss function, and using the gradient descent method to update the structure parameter set α′, and using the structure parameter set α′ to update the structure parameter set α; S60, repeating steps S40 to S50 until the training data set is traversed, and increasing the iteration round n by 1; S70, repeat steps S40-S60 until the current iteration round n reaches the maximum iteration round q; S80, based on the probability distribution formed by the probability normalization function of the b structural parameters corresponding to each of the connecting edges in the structural parameter set α′, a neural network operation corresponding to a certain probability value is sampled as the neural network operation of the connecting edge, thereby obtaining the complex neural network.

2. The method for searching a differentiable adaptive neural network architecture based on zero-order approximation according to claim 1, characterized in that: in, The neural network operation is any operation that preserves the data dimension. The neural network operations include zero operations, skip connection operations, 1×1 convolution operations, size-variable convolution operations, and 3×3 average pooling operations. The loss functions of the training data set and the validation data set are both cross entropy functions. In step S20, the neural unit is a directed acyclic graph composed of a series of data blocks and neural network operations, and the data blocks inside the unit have sequential numbers, among which the smallest and largest numbers constitute the input and output of the unit respectively.

3. The method for searching a differentiable adaptive neural network architecture based on zero-order approximation according to claim 1, characterized in that: in, In step S30, the hybrid neural network operation is a weighted neural network operation. For any edge (i, j) between data blocks numbered i and j, the corresponding hybrid neural network operation calculation result is: In the formula, represents the structural parameter vector of the (i, j) edge of the unit belonging to the neural unit group e in the structural parameter set α, g(·) represents the probability normalization function combined with the annealing strategy, g o It means taking the element of the neural network operation o corresponding to the result of the function operation.

4. The method for searching a differentiable adaptive neural network architecture based on zero-order approximation according to claim 3, characterized in that: in, The probability normalization function g(·) representing the combined annealing strategy is denoted as sparsemax′(·), whose input and output values ​​are both vectors. For a certain vector σ, the probability normalization function sparsemax′(σ) combined with the annealing strategy is: In the formula, σ is the input vector, argmin represents the parameter value of a function that achieves the minimum value in its domain, Δ L-1 is an L-1 dimensional simplex, represents an L-dimensional real vector space, the superscript T is the transposition operation, ||p-σ / (τ*a n / / m )|| 2 Represents the vector p and σ / (τ*a n / / m ), the integer division symbol / / represents the division of two integers and taking the quotient truncated towards zero. According to the set temperature constant τ and adjustment factors a and m, the input vector σ is selected as required, the vector p is calculated through the probability normalization function combined with the annealing strategy, and p is designated as the probability normalization result sparsemax′(σ) of the input vector σ to generate the required neural network weighted probability.

5. The method for searching a differentiable adaptive neural network architecture based on zero-order approximation according to claim 1, characterized in that: in, In step S30, the resizable search includes K convolution kernels of different sizes. When n is less than the preset iteration threshold θ, the resizable convolution operation on the j-th edge of any neural unit in the unit group e only uses the largest convolution kernel for calculation. Otherwise, the calculation result is: In the formula, represents the structural parameter vector corresponding to the jth edge in the structural parameter matrix of the neural unit group e in the β set, conv k (·) represents the convolution with kernel size k*k, g k Indicates selecting the kth element of the normalized function, The depth variable search targets S of the e-th neural unit group. e neural units, the units in the group are connected in series, and the output of the previous unit is used as the input of the next unit. When n is less than the preset iteration threshold θ, the variable depth operation result of the unit group e only selects the last unit output as the final output of the unit group and enters the next level calculation, otherwise the operation result is: In the formula, γ e represents the structural parameter vector belonging to the neural unit group e in the γ set, represents the output of the neural unit with sequence number d in unit group e, It is used to represent the number of units in the e-th neural unit group, g d Indicates selecting the dth element of the normalization function.

6. The method for searching a differentiable adaptive neural network architecture based on zero-order approximation according to claim 5, characterized in that: in, In step S30, the total resource occupancy number C is a quantitative characteristic of the usage of computing resources by a statistical model. For each neural network operation in the e-th group of neural units, the expected resource occupancy number for a topological structure edge (i, j) is: In the formula, represents the number of resources occupied by operation o on the (i,j) edge, represents the structural parameters of the (i,j)th structural edge belonging to the eth group of neural units, Resizable convolution operation The resource usage is expressed as: In the formula, Indicates the number of resources occupied by the operation of the convolution kernel of size k, The resource usage of the depth-variable operation of unit group e is expressed as: In the formula, Represents the total resource usage of the first d neural units in the e-th neural unit group.

7. The method for searching a differentiable adaptive neural network architecture based on zero-order approximation according to claim 1, characterized in that: in, In step S40, the minimum value is the vector obtained by concatenating all vectors of the structural parameter set α in a dimension with a width of 1, The random unit vector u′ is the number of elements and A random vector that is uniform and each element follows a normal distribution, Wave structure parameter set u″ is the variable set obtained by aligning the dimension of the random unit vector u with the structural parameter set α, It is the addition by element position.

8. The method for searching a differentiable adaptive neural network architecture based on zero-order approximation according to claim 1, Features: in, In step S40, the model parameter set w′ and the fluctuation model parameter set are updated. The method comprises the following sub-steps: A01, set the iteration round t = 1; A02, randomly select batch data from the training data set and input it into the initial neural network, combining the structure parameter set α and the fluctuation structure parameter set And according to the loss function of the lower layer problem and the learning rate η w , update the model parameter set w and the fluctuation model parameter set And increase the iteration round t by 1; A03, repeat step A02 until the iteration round t reaches the maximum iteration round T, then output the final updated model parameter set w′ and the updated fluctuation model parameter set 9. The method for searching for a differentiable adaptive neural network architecture based on zero-order approximation according to claim 8, characterized in that: in, In step A02, the lower layer problem loss function is the loss function And the loss function Loss Function The value of is the calculation result of the model parameter set w and the structure parameter set α based on the batch data in the training data set, based on the learning rate w and the loss function The gradient of the model parameter set w By using the gradient descent method, the updated model parameter set w″ is calculated as the model parameter set w at iteration round t+1. Loss Function The value of is based on the batch data in the training data set for the fluctuation model parameter set and the wave structure parameter set The calculation result is based on the learning rate η w And the loss function The wave model parameter set Gradient By using the gradient descent method, the updated wave model parameter set is calculated. As the set of fluctuation model parameters at iteration round t+1 10. The method for searching a differentiable adaptive neural network architecture based on zero-order approximation according to claim 1, characterized in that: in, In step S50, the approximate gradient of the structural parameter set α The calculation formula is: In the formula, is the loss function of the upper layer problem The gradient of the model parameter set w, w′(α) indicates that the updated model parameter set w′ is considered to be related to the structural parameter set α in the gradient calculation, for The gradient of the structural parameter set α, w′(), indicates that in the gradient calculation, the updated model parameter set w′ is regarded as a constant and has nothing to do with the structural parameter set α. The superscript T is a transposition operation. According to the setting learning rate η α and the approximate gradient of the structural parameter set α The structural parameter set α′ is calculated by the gradient descent method.