A neural network architecture search method, device, electronic device and storage medium

By constructing a neural network architecture search method that does not include jump connections in neural network architecture search and using gradient transmission unit stability algorithm, the problem of the neural network's accuracy decreases when the number of iterations increases, and the optimal architectural accuracy of the neural network is improved.

CN114219964BActive Publication Date: 2025-05-16INSPUR (BEIJING) ELECTRONICS INFORMATION IND CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111676417.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-05-16
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The existing neural network architecture search method can easily lead to a decrease in accuracy when the number of iterations increases. Especially when many jump connections are included, the accuracy of the neural network will be reduced, resulting in collapse.

Method used

A neural network architecture search method is designed. By constructing a neural network containing the structural unit to be searched and the gradient transmission unit, the operation set between the internal nodes in the structural unit to be searched does not contain jump connections, and the optimal operation is determined using the gradient transmission unit stability algorithm to ensure that the gradient of the deep network is effectively transmitted to the shallow network.

Benefits of technology

It effectively avoids the problem that the neural network search algorithm decreases with the increase of iterations, and improves the optimal architectural accuracy of the searched neural network, so that the accuracy of the neural network for image classification increases when the number of iterations increases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114219964B_ABST
    Figure CN114219964B_ABST
Patent Text Reader

Abstract

The present application discloses a neural network architecture search method, device, electronic device and computer-readable storage medium, the method comprising: obtaining a data set for an image classification task; wherein the data set comprises an image and a corresponding category label; constructing a neural network for the image classification task; wherein the neural network comprises a plurality of structural units connected in sequence, each structural unit comprises a structural unit to be searched and a gradient transmission unit, the structural unit to be searched comprises a plurality of internal nodes, and the gradient transmission unit comprises a jump connection operation or a 1×1 convolution operation; defining an operation set between internal nodes in the structural unit to be searched; wherein the operation set does not comprise a jump connection; using the data set to search for the best operation between every two internal nodes of the structural unit to be searched in each structural unit, and determining the structure of the gradient transmission unit. The present application improves the accuracy of the best architecture of the neural network for image classification searched.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image classification, and more specifically, to a neural network architecture search method and device, an electronic device, and a computer-readable storage medium. Background Art

[0002] In the field of deep learning, neural network architectures are constantly evolving, and neural network architecture search, that is, determining the optimal topological structure of a neural network, has become the mainstream method for neural network architecture design. Automatic neural network architecture search (full name: Neural Architecture Search, abbreviated as NAS) has become a current hot research direction.

[0003] In the related art, the Differentiable ARchiTecture Search (DARTS) method uses a gradient descent method to perform architecture search in a differentiable search space. However, as the number of search iterations increases, DARTS tends to prioritize skip connections from the search space during the search process. When a deep neural network contains many skip connections, the accuracy of the neural network will decrease, which will lead to the collapse of the neural network.

[0004] Therefore, how to improve the accuracy of the optimal architecture of the searched neural network is a technical problem that technicians in this field need to solve. Summary of the invention

[0005] The purpose of the present application is to provide a neural network architecture search method, device, electronic device and computer-readable storage medium, which improves the accuracy of the optimal architecture of the neural network for image classification.

[0006] To achieve the above objectives, the present application provides a neural network architecture search method, comprising:

[0007] Obtain a data set for an image classification task; wherein the data set includes images and corresponding category labels;

[0008] Constructing a neural network for an image classification task; wherein the neural network includes a plurality of structural units connected in sequence, each of the structural units includes a structural unit to be searched and a gradient transmission unit, the structural unit to be searched includes a plurality of internal nodes, and the gradient transmission unit includes a jump connection operation or a 1×1 convolution operation;

[0009] Defining an operation set between internal nodes in the structure unit to be searched; wherein the operation set does not include a jump connection;

[0010] The data set is used to search for the best operation between every two internal nodes of the structure unit to be searched in each of the structure units, and the structure of the gradient transmission unit is determined.

[0011] The output of the structure unit to be searched is a separable concatenation of all internal nodes in the structure unit to be searched.

[0012] Wherein, determining the structure of the gradient transmission unit includes:

[0013] If the dimension of the feature map input to the structural unit is the same as the dimension of the feature map output by the structural unit to be searched in the structural unit, then the gradient transmission unit in the structural unit is specifically a jump connection;

[0014] If the dimension of the feature map input to the structural unit is different from the dimension of the feature map output by the structural unit to be searched in the structural unit, the gradient transfer unit in the structural unit is specifically a 1×1 convolution operation.

[0015] The output of the structural unit is the sum of the output of the structural unit to be searched in the structural unit and the output of the gradient transmission unit.

[0016] Among them, the neural network and The structure unit is a reduced-resolution structure unit, N is the number of structure units in the neural network, the stride of the reduced-resolution structure unit is 2, and the strides of the remaining structure units are 1.

[0017] The data set includes a training set and a validation set; using the data set to search for the best operation between every two internal nodes of the structure unit to be searched in each structure unit includes:

[0018] Determining a weight parameter of each operation in the operation set using the training set;

[0019] Determine the architecture parameter corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units by using the verification set;

[0020] The operation with the largest architecture parameter between every two internal nodes is determined as the best operation.

[0021] The step of using the verification set to determine the architecture parameter corresponding to each operation between every two internal nodes of the structure unit to be searched in each structure unit includes:

[0022] Input the image of the verification set into the neural network, calculate the architecture loss corresponding to each operation between every two internal nodes of each structural unit to be searched in each structural unit by using the architecture loss function based on the output of the neural network and the category label corresponding to the image, calculate the architecture loss gradient corresponding to each operation between every two internal nodes of each structural unit to be searched in each structural unit based on the architecture loss, and update the architecture parameters corresponding to each operation between every two internal nodes of each structural unit to be searched in each structural unit based on the architecture loss gradient;

[0023] Among them, the architecture loss function is: Among them, ω * (α) is the optimal weight parameter obtained on the training set, α is the architecture parameter set, L val () is the loss value on the validation set, ω 0-1 is a predefined hyperparameter, M is the total number of all connections to be searched in all structural units to be searched in the neural network, and two internal nodes containing operations to be searched are defined as one connection to be searched. N is the total number of operations of the mth connection to be searched, σ() is the softmax function, α n is the architecture parameter of the nth operation of the mth connection to be searched, O is the operation set, o i,j and o′ i,j is the output of the operation between intermediate node i and intermediate node j.

[0024] To achieve the above objectives, the present application provides a neural network architecture search device, comprising:

[0025] An acquisition module is used to acquire a data set for an image classification task; wherein the data set includes images and corresponding category labels;

[0026] A construction module, used to construct a neural network for an image classification task; wherein the neural network includes a plurality of structure units connected in sequence, each of the structure units includes a structure unit to be searched and a gradient transmission unit, the structure unit to be searched includes a plurality of internal nodes, and the gradient transmission unit includes a jump connection operation or a 1×1 convolution operation;

[0027] A definition module, used to define an operation set between internal nodes in the structure unit to be searched; wherein the operation set does not include a jump connection;

[0028] The search module is used to use the data set to search for the best operation between every two internal nodes of the structure unit to be searched in each of the structure units, and determine the structure of the gradient transmission unit.

[0029] To achieve the above objectives, the present application provides an electronic device, including:

[0030] Memory for storing computer programs;

[0031] A processor is used to implement the steps of the above-mentioned neural network architecture search method when executing the computer program.

[0032] To achieve the above objectives, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned neural network architecture search method are implemented.

[0033] It can be seen from the above scheme that a neural network architecture search method provided by the present application includes: obtaining a data set for an image classification task; wherein the data set contains images and corresponding category labels; constructing a neural network for the image classification task; wherein the neural network includes a plurality of structural units connected in sequence, each of the structural units includes a structural unit to be searched and a gradient transfer unit, the structural unit to be searched includes a plurality of internal nodes, and the gradient transfer unit includes a jump connection operation or a 1×1 convolution operation; defining an operation set between internal nodes in the structural unit to be searched; wherein the operation set does not include a jump connection; using the data set to search for the best operation between every two internal nodes of the structural unit to be searched in each of the structural units, and determining the structure of the gradient transfer unit.

[0034] The present application designs a new neural network for image classification tasks, which includes multiple structural units, each of which includes a structural unit to be searched and a gradient transmission unit. The operation set between the internal nodes in the structural unit to be searched does not include jump connections. Searching for the best operation between every two internal nodes in the structural unit to be searched can prevent jump connections from being searched between internal nodes, thereby avoiding the decrease in accuracy of the neural network search algorithm as the number of search iterations increases. At the same time, the gradient transmission unit is used to stabilize the search process of the algorithm to ensure that the gradient of the deep network is effectively transmitted to the shallow network, further improving the accuracy of the optimal architecture of the searched neural network. It can be seen that the neural network architecture search method provided by the present application increases the accuracy of the neural network for image classification as the number of search iterations increases, thereby improving the accuracy of the optimal architecture of the neural network for image classification. The present application also discloses a neural network architecture search device, an electronic device, and a computer-readable storage medium, which can also achieve the above-mentioned technical effects.

[0035] It should be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. The drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following specific implementation methods, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure. In the drawings:

[0037] Figure 1 is a flowchart of a neural network architecture search method according to an exemplary embodiment;

[0038] Figure 2 is a structural diagram of a neural network for an image classification task according to an exemplary embodiment;

[0039] Figure 3 is a flowchart of another neural network architecture search method according to an exemplary embodiment;

[0040] Figure 4 is a structural diagram of a neural network architecture search device according to an exemplary embodiment;

[0041] Figure 5 The figure is a structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0042] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. In addition, in the embodiments of the present application, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0043] The embodiment of the present application discloses a neural network architecture search method, which improves the accuracy of the optimal architecture of the searched neural network for image classification.

[0044] See also Figure 1 , according to a flowchart of a neural network architecture search method shown in an exemplary embodiment, such as Figure 1 As shown, including:

[0045] S101: Obtain a data set for an image classification task; wherein the data set includes images and corresponding category labels;

[0046] The purpose of this embodiment is to search for the best architecture of a neural network for image classification. In this step, a data set for an image classification task is obtained, which includes images and corresponding category labels. For example, for a digital classification task from 0 to 9, the input data set includes images containing numbers from 0 to 9, and the category label of each image is the number it contains.

[0047] S102: constructing a neural network for an image classification task; wherein the neural network includes a plurality of structure units connected in sequence, each of the structure units includes a structure unit to be searched and a gradient transmission unit, the structure unit to be searched includes a plurality of internal nodes, and the gradient transmission unit includes a jump connection operation or a 1×1 convolution operation;

[0048] S103: defining an operation set between internal nodes in the structure unit to be searched; wherein the operation set does not include a jump connection;

[0049] It should be noted that for a general residual neural network, which consists of N residual blocks, the output of the i+1th residual block is X i+1 , X i+1 =f i+1 (X i , W i+1 )+X i , where X i is the output of the ith residual block, W i+1 is the weight of the i+1th residual block, f i+1 The weight is W i+1 The operation of the i+1th residual block. Assume that the model loss is L. It can be proved that From this formula, we can see that the shallow network always contains the gradient information of the deep network. The skip connection of the residual network alleviates the gradient vanishing phenomenon, making the deep neural network easier to train. Therefore, when the DARTS technology searches for the best operation in the search space containing the skip connection, the skip connection is more likely to be selected, resulting in a "collapse" phenomenon in the search process.

[0050] Therefore, the neural network for the image classification task constructed in this embodiment is as follows: Figure 2As shown, it includes multiple structural units connected in sequence, and each structural unit includes a structural unit to be searched and a gradient transmission unit. The operation set O between the internal nodes in the structural unit to be searched may include operations such as convolution, maximum pooling, such as 3×3 separable convolution, 5×5 separable convolution, 7×7 separable convolution, 9×9 separable convolution, 3×3 convolution, 3×3 dilated convolution, 5×5 dilated convolution, 3×3 average pooling, "zero" operation, etc., which are not specifically limited here. It should be noted that the operation set in this embodiment does not include jump connections.

[0051] It should be noted that this embodiment has no limit on the number of channels operated, which can be 16 channels, 32 channels or other number of channels. The convolution feature map is padded with zeros to maintain the spatial resolution. This embodiment uses the ReLU-Conv-BN sequential operation, and each separable convolution is applied twice.

[0052] Each structural unit to be searched is a directed acyclic graph containing N ordered internal nodes, where the internal nodes represent feature graphs and the edge e from internal node i to internal node j is i,j Represents an operation, and the output of the operation is defined as o i,j (x i ). The value of an internal node j is the sum of the outputs of the edges connected to it, defined as x j =∑ i<j o i,j (x i ). Between internal node i and internal node j, there are several operations, each corresponding to an architectural weight. The output of the structure unit to be searched is the separable concatenation of all internal nodes. The gradient transfer unit branch contains fixed jump connections or 1×1 convolution operations.

[0053] The neural network is composed of N structural units connected in series, and the input node of structural unit k is the output node of structural unit k-1. The output of the structural unit is the sum of the output of the structural unit to be searched and the output of the gradient transmission unit.

[0054] As a preferred embodiment, the neural network and The structure unit is a resolution reduction structure unit, N is the number of structure units in the neural network, the stride of the resolution reduction structure unit is 2, and the stride of the remaining structure units is 1. In a specific implementation, in the neural network and , which is the resolution reduction structural unit of the neural network. In the resolution reduction structural unit, the stride of the operation of connecting the input nodes is 2. Indicates rounding down.

[0055] S104: Using the data set, searching for the best operation between every two internal nodes of the structure unit to be searched in each of the structure units, and determining the structure of the gradient transmission unit.

[0056] The neural network architecture search method provided in this embodiment searches for the best operation between every two internal nodes in the structural unit to be searched, excludes the jump connection operation, and avoids the decrease in accuracy of the neural network search algorithm as the number of search iterations increases, that is, avoids the "collapse" phenomenon in the search process. The gradient transmission unit is used to stabilize the search process of the search method, including the jump connection operation or the 1×1 convolution operation, to ensure that the gradient of the deep network can be effectively transmitted to the shallow network during the gradient back propagation process.

[0057] As a preferred implementation, the structure of the gradient transfer unit is determined, including: if the dimension of the feature map input to the structural unit is the same as the dimension of the feature map output by the structural unit to be searched in the structural unit, the gradient transfer unit in the structural unit is specifically a jump connection; if the dimension of the feature map input to the structural unit is different from the dimension of the feature map output by the structural unit to be searched in the structural unit, the gradient transfer unit in the structural unit is specifically a 1×1 convolution operation. In a specific implementation, when the dimension of the input feature map is the same as the output dimension, the gradient transfer unit is a jump connection. When the dimension of the feature map is different from the output dimension, the gradient transfer unit is a 1×1 convolution operation.

[0058] The embodiment of the present application designs a new neural network for the image classification task, which includes multiple structural units, each of which includes a structural unit to be searched and a gradient transmission unit. The operation set between the internal nodes in the structural unit to be searched does not include a jump connection. The optimal operation between every two internal nodes in the structural unit to be searched can prevent the search for jump connections between the internal nodes, thereby avoiding the decrease in accuracy of the neural network search algorithm as the number of search iterations increases. At the same time, the gradient transmission unit is used to stabilize the search process of the algorithm, ensuring that the gradient of the deep network is effectively transmitted to the shallow network, further improving the accuracy of the optimal architecture of the searched neural network. It can be seen that the neural network architecture search method provided by the embodiment of the present application increases the accuracy of the neural network for image classification as the number of search iterations increases, thereby improving the accuracy of the optimal architecture of the neural network for image classification.

[0059] The embodiment of the present application discloses a neural network architecture search method. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution. Specifically:

[0060] See also Figure 3 , a flowchart of another neural network architecture search method according to an exemplary embodiment is shown, such as Figure 3 As shown, including:

[0061] S201: Obtain a data set for an image classification task; wherein the data set includes a training set and a validation set, and the training set and the validation set contain images and corresponding category labels;

[0062] In this embodiment, the data set is divided into two subsets: a training set and a validation set, both of which contain images and corresponding category labels. The training set is used to train the weight parameters of the neural network, and the validation set is used to train the architecture parameters of the neural network. The weight parameters refer to the weights in the operations of the neural network (for example: 3×3 convolution operations, 5×5 convolution operations, separable convolution operations, dilated convolution operations, etc.). The architecture parameters refer to the parameters representing the importance of the operation to be searched.

[0063] S202: constructing a neural network for an image classification task; wherein the neural network includes a plurality of structure units connected in sequence, each of the structure units includes a structure unit to be searched and a gradient transmission unit, the structure unit to be searched includes a plurality of internal nodes, and the gradient transmission unit includes a jump connection operation or a 1×1 convolution operation;

[0064] In this step, a neural network for the image classification task is constructed, and the weight parameter ω, architecture parameter α, and the total number of search iterations E are initialized.

[0065] S203: defining an operation set between internal nodes in the structure unit to be searched; wherein the operation set does not include a jump connection;

[0066] S204: Determine a weight parameter of each operation in the operation set using the training set;

[0067] S205: Determine, by using the verification set, an architecture parameter corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units;

[0068] As a preferred implementation, the use of the validation set to determine the architecture parameters corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units comprises: inputting the image of the validation set into the neural network, using the architecture loss function to calculate the architecture loss corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units based on the output of the neural network and the category label corresponding to the image, and calculating the architecture loss gradient corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units based on the architecture loss, and updating the architecture parameters corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units based on the architecture loss gradient;

[0069] Among them, the architecture loss function is: Among them, ω * (α) is the optimal weight parameter obtained on the training set, α is the architecture parameter set, L val () is the loss value on the validation set, ω 0-1 is a predefined hyperparameter, M is the total number of all connections to be searched in all structural units to be searched in the neural network, and two internal nodes containing operations to be searched are defined as one connection to be searched. N is the total number of operations of the mth connection to be searched, σ() is the softmax function, α n is the architecture parameter of the nth operation of the mth connection to be searched, O is the operation set, o i,j and o′ i,j is the output of the operation between intermediate node i and intermediate node j.

[0070] In order to make the search space continuous, this embodiment can convert the architecture weight into classification probability: in, Represents the mixed probability. At the end of training, quilt Furthermore, the following architecture parameter constraint function L is defined: 0-1 Classification probability values ​​that drive architecture parameters Approaching 0 or 1, thereby increasing The distinguishability.

[0071] In the specific implementation, the architecture parameters and weight parameters can be jointly learned. The training loss and evaluation loss are defined separately, and the following problems are solved using two-layer optimization:

[0072]

[0073] stω * (α) = argmin ω L train (ω, α); where L train () is the loss value on the training set;

[0074]

[0075]

[0076]

[0077] Starting from the first iteration until the number of iterations reaches E, perform the following operations: Calculate the gradient based on the loss of the training set Update the weight parameter ω and take a set of data (including images and category labels) from the training set. Each sample of this set of data passes through the neural network to obtain the network output value. Use these output values ​​and the corresponding category labels to calculate L train (ω, α). Then, the back propagation algorithm is used to calculate And update ω. Calculate the gradient based on the loss of the validation set Update the architecture parameter α and take a set of data (including images and category labels) from the validation set. Each sample of this set of data passes through the neural network to obtain the network output value. Use these output values ​​and the corresponding category labels to calculate Then, the back propagation algorithm is used to calculate And update α.

[0078] S206: Determine the operation with the largest architecture parameter between every two internal nodes as the optimal operation, and determine the structure of the gradient transmission unit.

[0079] In the specific implementation, for each operation between two internal nodes, according to the corresponding architecture parameter α n The operation with the largest architecture parameter value is selected as the operation between two internal nodes. Furthermore, when the dimension of the input feature map is the same as the output dimension, the gradient transfer unit is a jump connection. When the dimension of the feature map is different from the output dimension, the gradient transfer unit is a 1×1 convolution operation.

[0080] It can be seen that this embodiment constrains the update step of the architecture parameters through the constraint conditions of the architecture parameters, so that the architecture parameters have better distinguishability, which can promote the network search for the optimal architecture with higher accuracy and solve the problem that the optimal architecture is difficult to select.

[0081] A neural network architecture search device provided in an embodiment of the present application is introduced below. The neural network architecture search device described below and the neural network architecture search method described above can be referenced to each other.

[0082] See also Figure 4 , according to an exemplary embodiment, a structural diagram of a neural network architecture search device is shown, such as Figure 4 As shown, including:

[0083] An acquisition module 401 is used to acquire a data set for an image classification task; wherein the data set includes images and corresponding category labels;

[0084] A construction module 402 is used to construct a neural network for an image classification task; wherein the neural network includes a plurality of structure units connected in sequence, each of the structure units includes a structure unit to be searched and a gradient transmission unit, the structure unit to be searched includes a plurality of internal nodes, and the gradient transmission unit includes a jump connection operation or a 1×1 convolution operation;

[0085] A definition module 403 is used to define an operation set between internal nodes in the structure unit to be searched; wherein the operation set does not include a jump connection;

[0086] The search module 404 is used to use the data set to search for the best operation between every two internal nodes of the structure unit to be searched in each of the structure units, and determine the structure of the gradient transmission unit.

[0087] The embodiment of the present application designs a new neural network for the image classification task, which includes multiple structural units, each of which includes a structural unit to be searched and a gradient transmission unit. The operation set between the internal nodes in the structural unit to be searched does not include a jump connection. The optimal operation between every two internal nodes in the structural unit to be searched can prevent the search for jump connections between the internal nodes, thereby avoiding the decrease in accuracy of the neural network search algorithm as the number of search iterations increases. At the same time, the gradient transmission unit is used to stabilize the search process of the algorithm, ensuring that the gradient of the deep network is effectively transmitted to the shallow network, further improving the accuracy of the optimal architecture of the searched neural network. It can be seen that the neural network architecture search device provided in the embodiment of the present application increases the accuracy of the neural network for image classification as the number of search iterations increases, thereby improving the accuracy of the optimal architecture of the neural network for image classification.

[0088] Based on the above embodiment, as a preferred implementation manner, the output of the structure unit to be searched is a separable concatenation of all internal nodes in the structure unit to be searched.

[0089] On the basis of the above embodiments, as a preferred implementation manner, if the dimension of the feature map input to the structural unit is the same as the dimension of the feature map output by the structural unit to be searched in the structural unit, the gradient transmission unit in the structural unit is specifically a jump connection; if the dimension of the feature map input to the structural unit is different from the dimension of the feature map output by the structural unit to be searched in the structural unit, the gradient transmission unit in the structural unit is specifically a 1×1 convolution operation.

[0090] Based on the above embodiment, as a preferred implementation manner, the output of the structure unit is the sum of the output of the structure unit to be searched in the structure unit and the output of the gradient transmission unit.

[0091] Based on the above embodiment, as a preferred implementation, the first and The structure unit is a reduced-resolution structure unit, the stride of the reduced-resolution structure unit is 2, and the strides of the remaining structure units are 1.

[0092] Based on the above embodiment, as a preferred implementation, the data set includes a training set and a validation set; the search module 404 includes:

[0093] A first determining unit, configured to determine a weight parameter of each operation in the operation set by using the training set;

[0094] A second determination unit is used to determine, by using the verification set, an architecture parameter corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units;

[0095] The third determining unit is used to determine the operation with the largest architecture parameter between every two internal nodes as the best operation.

[0096] On the basis of the above embodiment, as a preferred implementation manner, the second determination unit specifically inputs the image of the verification set into the neural network, calculates the architecture loss corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units based on the output of the neural network and the category label corresponding to the image by using the architecture loss function, calculates the architecture loss gradient corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units based on the architecture loss, and updates the unit of the architecture parameter corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units based on the architecture loss gradient;

[0097] Among them, the architecture loss function is: Among them, ω * (α) is the optimal weight parameter obtained on the training set, α is the architecture parameter set, L val () is the loss value on the validation set, ω 0-1 is a predefined hyperparameter, M is the total number of all connections to be searched in all structural units to be searched in the neural network, and two internal nodes containing operations to be searched are defined as one connection to be searched. N is the total number of operations of the mth connection to be searched, σ() is the softmax function, α n is the architecture parameter of the nth operation of the mth connection to be searched, O is the operation set, o i,j and o′ i,j is the output of the operation between intermediate node i and intermediate node j.

[0098] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0099] Based on the hardware implementation of the above program module, and in order to implement the method of the embodiment of the present application, the embodiment of the present application also provides an electronic device, Figure 5 FIG. 1 is a structural diagram of an electronic device according to an exemplary embodiment. Figure 5 As shown, the electronic equipment includes:

[0100] Communication interface 1, capable of exchanging information with other devices such as network devices;

[0101] The processor 2 is connected to the communication interface 1 to realize information exchange with other devices, and is used to execute the neural network architecture search method provided by one or more technical solutions when running a computer program. The computer program is stored in the memory 3.

[0102] Of course, in actual application, the various components in the electronic device are coupled together through the bus system 4. It can be understood that the bus system 4 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 4 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 5 Various buses are labeled as bus system 4.

[0103] The memory 3 in the embodiment of the present application is used to store various types of data to support the operation of the electronic device. Examples of such data include: any computer program used to operate on the electronic device.

[0104] It can be understood that the memory 3 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAM bus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 3 described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memories.

[0105] The method disclosed in the above embodiment of the present application can be applied to the processor 2, or implemented by the processor 2. The processor 2 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 2 or the instruction in the form of software. The above processor 2 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 2 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in the memory 3, and the processor 2 reads the program in the memory 3 and completes the steps of the above method in combination with its hardware.

[0106] When the processor 2 executes the program, the corresponding processes in the various methods of the embodiments of the present application are implemented, which will not be repeated here for the sake of brevity.

[0107] In an exemplary embodiment, the present application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, a memory 3 storing a computer program, and the computer program can be executed by a processor 2 to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.

[0108] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium, which, when executed, executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, disks or optical disks.

[0109] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0110] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A neural network architecture search method, characterized in that: include: Obtain a data set for an image classification task; wherein the data set includes images and corresponding category labels; Constructing a neural network for an image classification task; wherein the neural network includes a plurality of structural units connected in sequence, each of the structural units includes a structural unit to be searched and a gradient transmission unit, the structural unit to be searched includes a plurality of internal nodes, and the gradient transmission unit includes a jump connection operation or a 1×1 convolution operation; Defining an operation set between internal nodes in the structure unit to be searched; wherein the operation set does not include a jump connection; Using the data set, searching for the best operation between every two internal nodes of the structure unit to be searched in each of the structure units, and determining the structure of the gradient transmission unit; The data set includes a training set and a validation set; using the data set to search for the best operation between every two internal nodes of the structure unit to be searched in each structure unit includes: Determining a weight parameter of each operation in the operation set using the training set; Determine the architecture parameter corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units by using the verification set; The operation with the largest architecture parameter between every two internal nodes is determined as the best operation; The step of using the verification set to determine the architecture parameter corresponding to each operation between every two internal nodes of the structure unit to be searched in each structure unit includes: Input the image of the verification set into the neural network, calculate the architecture loss corresponding to each operation between every two internal nodes of each structural unit to be searched in each structural unit by using the architecture loss function based on the output of the neural network and the category label corresponding to the image, calculate the architecture loss gradient corresponding to each operation between every two internal nodes of each structural unit to be searched in each structural unit based on the architecture loss, and update the architecture parameters corresponding to each operation between every two internal nodes of each structural unit to be searched in each structural unit based on the architecture loss gradient; Among them, the architecture loss function is: Among them, ω * (α) is the optimal weight parameter obtained on the training set, α is the architecture parameter set, L val () is the loss value on the validation set, ω 0-1 is a predefined hyperparameter, M is the total number of all connections to be searched in all structural units to be searched in the neural network, and two internal nodes containing operations to be searched are defined as one connection to be searched. N is the total number of operations of the mth connection to be searched, σ() is the softmax function, α n is the architecture parameter of the nth operation of the mth connection to be searched, O is the operation set, o i,j and o′ i,j is the output of the operation between intermediate node i and intermediate node j.

2. The neural network architecture search method according to claim 1, characterized in that: The output of the structure unit to be searched is a separable concatenation of all internal nodes in the structure unit to be searched.

3. The neural network architecture search method according to claim 1, characterized in that: The step of determining the structure of the gradient transmission unit comprises: If the dimension of the feature map input to the structural unit is the same as the dimension of the feature map output by the structural unit to be searched in the structural unit, then the gradient transmission unit in the structural unit is specifically a jump connection; If the dimension of the feature map input to the structural unit is different from the dimension of the feature map output by the structural unit to be searched in the structural unit, the gradient transfer unit in the structural unit is specifically a 1×1 convolution operation.

4. The neural network architecture search method according to claim 2, characterized in that: The output of the structural unit is the sum of the output of the structural unit to be searched and the output of the gradient transmission unit in the structural unit.

5. The neural network architecture search method according to claim 1, characterized in that: The neural network and The structure unit is a reduced-resolution structure unit, N is the number of structure units in the neural network, the stride of the reduced-resolution structure unit is 2, and the strides of the remaining structure units are 1.

6. A neural network architecture search device, characterized in that: include: An acquisition module is used to acquire a data set for an image classification task; wherein the data set includes images and corresponding category labels; A construction module, used to construct a neural network for an image classification task; wherein the neural network includes a plurality of structure units connected in sequence, each of the structure units includes a structure unit to be searched and a gradient transmission unit, the structure unit to be searched includes a plurality of internal nodes, and the gradient transmission unit includes a jump connection operation or a 1×1 convolution operation; A definition module, used to define an operation set between internal nodes in the structure unit to be searched; wherein the operation set does not include a jump connection; A search module, used for searching the optimal operation between every two internal nodes of the structure unit to be searched in each of the structure units by using the data set, and determining the structure of the gradient transmission unit; The data set includes a training set and a validation set; the search module includes: A first determining unit, configured to determine a weight parameter of each operation in the operation set by using the training set; A second determining unit is used to determine, by using the verification set, an architecture parameter corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units; A third determining unit, configured to determine the operation with the largest architecture parameter between every two internal nodes as the best operation; The second determination unit is specifically a unit for inputting the image of the verification set into the neural network, calculating the architecture loss corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units based on the output of the neural network and the category label corresponding to the image by using the architecture loss function, calculating the architecture loss gradient corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units based on the architecture loss, and updating the architecture parameter corresponding to each operation between every two internal nodes of the structure unit to be searched in each of the structure units based on the architecture loss gradient; Among them, the architecture loss function is: Among them, ω * (α) is the optimal weight parameter obtained on the training set, α is the architecture parameter set, L val () is the loss value on the validation set, ω 0-1 is a predefined hyperparameter, M is the total number of all connections to be searched in all structural units to be searched in the neural network, and two internal nodes containing operations to be searched are defined as one connection to be searched. N is the total number of operations of the mth connection to be searched, σ() is the softmax function, α n is the architecture parameter of the nth operation of the mth connection to be searched, O is the operation set, o i,j and o′ i,j is the output of the operation between intermediate node i and intermediate node j.

7. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the neural network architecture search method as claimed in any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the neural network architecture search method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Microstructure searching method based on grouping and layering mechanism

    CN111275186A