Method and device for searching neural network architecture, electronic device, and storage medium
By regularizing the initial loss function, the differences in architecture parameters in the neural network architecture search method are optimized, solving the problem of difficult operator selection in traditional methods and improving the performance of neural networks and image super-resolution effects.
Patent Information
- Application Number
- CN202111579065.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2041-12-22
AI Technical Summary
Traditional neural network architecture search methods suffer from poor performance in the generated neural networks, mainly because the architecture parameters are not very different, making it difficult to select operators.
By regularizing the initial loss function, the target loss function is obtained. Based on the target loss function, the architecture parameters corresponding to the operator are optimized to improve the variability of the architecture parameters, thereby determining the optimal operator among multiple operators.
It improves the performance of neural networks, especially in image super-resolution tasks, enhancing the image quality after super-resolution processing and the generalization performance of neural networks.
Smart Images

Figure CN114298271B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, and particularly relates to a neural network architecture searching method and device, electronic equipment and storage medium. BACKGROUND
[0002] With the development of machine learning technology, machine learning algorithms based on artificial neural networks have achieved high accuracy in image processing, natural language processing and other tasks. Among them, the design of neural networks is an important topic of machine learning, and in order to improve the design efficiency of neural networks, the related technology usually generates neural networks quickly through a neural network architecture searching method, but the traditional neural network architecture searching method has certain defects, thereby affecting the performance of the generated neural network. SUMMARY
[0003] The embodiments of the present application disclose a neural network architecture searching method and device, electronic equipment and storage medium, which can improve the performance of the generated neural network.
[0004] The first aspect of the embodiments of the present application discloses a neural network architecture searching method, comprising:
[0005] determining a plurality of operators corresponding to the first node and the second node, and determining first architecture parameters corresponding to the plurality of operators, respectively, wherein the first node and the second node are any two nodes in at least two nodes obtained from a search space;
[0006] regularizing an initial loss function to obtain a target loss function, and optimizing the first architecture parameters corresponding to the plurality of operators, respectively, according to the target loss function, to obtain second architecture parameters corresponding to the plurality of operators, respectively;
[0007] determining an operator with the maximum second architecture parameter in the plurality of operators as a target operator corresponding to the first node and the second node;
[0008] determining a subnetwork according to the at least two nodes and the target operator corresponding to any two nodes in the at least two nodes, and determining a target neural network according to a plurality of subnetworks.
[0009] The second aspect of the embodiments of the present application discloses a neural network architecture searching device, comprising:
[0010] a first determining unit configured to determine first architecture parameters corresponding to a plurality of operators corresponding to a first node and a second node, respectively, wherein the first node and the second node are any two nodes in at least two nodes obtained from a search space;
[0011] An optimization unit is configured to regularize the initial loss function to obtain a target loss function, and optimize the first architecture parameter corresponding to each of the plurality of operators according to the target loss function to obtain a second architecture parameter corresponding to each of the plurality of operators.
[0012] A second determination unit is configured to determine, from the plurality of operators, an operator with a maximum second architecture parameter as a target operator corresponding to the first node and the second node.
[0013] A third determination unit is configured to determine a sub-network according to the at least two nodes and the target operator corresponding to any two nodes of the at least two nodes, and determine a target neural network according to a plurality of sub-networks.
[0014] A third aspect of an embodiment of the present application discloses an electronic device, comprising:
[0015] A memory storing executable program codes;
[0016] A processor coupled to the memory;
[0017] The processor invokes the executable program codes stored in the memory to execute the method for searching a neural network architecture disclosed in the first aspect of the present application.
[0018] A fourth aspect of an embodiment of the present application discloses a computer readable storage medium storing a computer program, wherein the computer program causes a computer to execute the method for searching a neural network architecture disclosed in the first aspect of the present application.
[0019] A fifth aspect of an embodiment of the present application discloses a computer program product, when the computer program product runs on a computer, causes the computer to execute part or all steps of any one method of the first aspect of the present application.
[0020] A sixth aspect of an embodiment of the present application discloses an application publishing platform, which is configured to publish a computer program product, wherein when the computer program product runs on a computer, causes the computer to execute part or all steps of any one method of the first aspect of the present application.
[0021] Compared with the related art, the embodiments of the present application have the following beneficial effects:
[0022] By the method provided in the embodiments of the present application, at least two nodes obtained from the search space can be determined, and the first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node are determined. Further, the initial loss function can be regularized to obtain a target loss function, and the first architecture parameters corresponding to the plurality of operators are optimized according to the target loss function. The first architecture parameters are optimized by the target loss function after the regularization processing, which can make the plurality of second architecture parameters obtained after the optimization more discrete, improve the difference between each second architecture parameter, and make it more accurate to determine the largest second architecture parameter in the plurality of operators subsequently. Further, the operator with the largest corresponding second architecture parameter can be used as the optimal operator of the first node and the second node, thereby improving the quality of the operators in the subnetwork and further improving the performance of the target neural network built according to the plurality of subnetworks subsequently. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0024] Figure 1 is a flowchart of a network architecture search method in the related art;
[0025] Figure 2 is a structure diagram of a subnetwork disclosed in an embodiment of the present application;
[0026] Figure 3 is a diagram for explaining how to build a subnetwork disclosed in an embodiment of the present application;
[0027] Figure 4 is a flowchart of a neural network architecture search method disclosed in an embodiment of the present application;
[0028] Figure 5 is a flowchart of another neural network architecture search method disclosed in an embodiment of the present application;
[0029] Figure 6 is a flowchart of still another neural network architecture search method disclosed in an embodiment of the present application;
[0030] Figure 7 is a diagram of a U-shaped neural network disclosed in an embodiment of the present application;
[0031] Figure 8 is a structure diagram of a neural network architecture search device disclosed in an embodiment of the present application;
[0032] Figure 9 is a structural schematic diagram of an electronic device disclosed by an embodiment of the present application. DETAILED DESCRIPTION
[0033] The technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0034] It should be noted that the terms "first", "second", "third", and "fourth" in the specification and claims of the present application are used to distinguish different objects, rather than to describe a specific order. The terms "include" and "have" and any variations thereof in the embodiments of the present application are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.
[0035] The embodiments of the present application disclose a neural network architecture search method and device, an electronic device and a storage medium, which can improve the performance of the generated neural network.
[0036] The technical solutions of the present application will be described in detail in conjunction with specific embodiments.
[0037] In order to more clearly illustrate the neural network architecture search method and device, electronic device and storage medium disclosed by the embodiments of the present application, first, the network architecture search (NAS) method in the related art is introduced.
[0038] Please refer to Figure 1 , Figure 1 is a flowchart of the network architecture search method in the related art. As shown in Figure 1 , the NAS method can be roughly divided into three steps, namely: (1) establishing a search space 100; (2) determining a neural network 110 from the search space 100 based on a certain search strategy; and (3) performing performance evaluation on the neural network 110.
[0039] The search space 100 can be regarded as a subset of the network structure space, and a researcher or a developer can artificially delimit a search space 100 in the network structure space according to prior knowledge to facilitate subsequent determination of a neural network in a smaller set. The search space 100 can include various nodes and operators and the like components, where the node represents an input or output feature quantity (for example, a feature map), and the operator represents various computing processes (for example, convolution, pooling, and the like) performed between the nodes.
[0040] After the search space 100 is defined, the neural network 110 can be searched from the search space 100 based on a certain search strategy. The search strategy is how to select appropriate nodes and operators in the search space to construct the neural network 110 according to the nodes and operators.
[0041] The neural network 110 constructed according to the nodes and operators can be further evaluated in performance based on a target data set to verify the performance of the neural network. The search strategy described above aims to find the network with the best performance (such as accuracy) from the search space 100, and therefore the result of the performance evaluation can be fed back to the search strategy to guide the next round of search until a neural network with good performance is finally searched.
[0042] From the above introduction, it can be seen that how to search the neural network 110 from the search space 100 based on a certain search strategy is the key of the NAS method, and a neural network 110 is usually composed of multiple sub-networks (that is, cells, please refer to Figure 2 , Figure 2 The present embodiment discloses a structure schematic diagram of a sub-network, and the sub-network is usually composed of nodes 200 and operators 210, so the sub-network needs to be constructed before the neural network 110 is constructed.
[0043] Please refer to Figure 2 and Figure 3 , Figure 3 The present embodiment discloses a schematic diagram for explaining how to construct the sub-network. As shown in Figure 3 After the nodes 200 are obtained from the search space by the search strategy, the operators 210 between each two nodes usually need to be further determined, but there can be many computing operations (such as Figure 3As shown, there can be 3 different operators between each two nodes, such as the convolution, pooling, etc. described above. In order to determine the optimal operator between two nodes, the related art usually performs continuous relaxation processing on multiple operators to parameterize the multiple operators into architecture parameters, and then optimizes the architecture parameters through a loss function to obtain the optimized architecture parameters corresponding to the multiple operators respectively, and then determines the operator with the largest optimized architecture parameter as the optimal operator. After determining the optimal operator between all nodes, the subnetwork shown in Figure 2
[0044] However, it is found in practice that the architecture parameters optimized through the loss function in the related art usually fall within a concentrated interval, and the difference between the architecture parameters is not large, which makes it difficult to choose the optimal operator. In this regard, the method disclosed in the embodiments of the present application can perform regularization processing on the initial loss function to obtain a target loss function, and optimize the first architecture parameters corresponding to the multiple operators according to the target loss function; wherein the target loss function after regularization processing can make the multiple second architecture parameters obtained after optimization more discrete, improve the difference between the second architecture parameters, and thus facilitate more accurate determination of the second architecture parameter with the largest value in the multiple operators in the subsequent step, and thus determine the optimal operator, thereby solving the problem that the difference between the architecture parameters is not large, making it difficult to choose the operator.
[0045] In an embodiment, the method for searching the neural network architecture disclosed in the embodiments of the present application can be applied to the construction of a neural network for image super-resolution. The neural network for image super-resolution is a neural network for performing the image super-resolution task. Image super-resolution is an image processing algorithm that can maintain the clarity of the image picture even after the image is enlarged by several times, or even tens of times. With the rapid development of electronic devices (such as mobile phones and tablet computers), image super-resolution technology has been widely applied to various electronic devices.
[0046] In addition to solving the problem of operator selection in the construction of the neural network for image super-resolution, the method for searching the neural network architecture disclosed in the embodiments of the present application can also more accurately and reasonably measure the picture quality loss after image super-resolution, and improve the picture quality of the image after image super-resolution processing.
[0047] Based on this, the following describes the method for searching the neural network architecture disclosed in the embodiments of the present application.
[0048] Please refer to Figure 4 , Figure 4 is a flowchart of a method for searching a neural network architecture according to an embodiment of the present application. The method can be applied to various electronic devices that can be used to build a neural network, including but not limited to cloud servers, personal computers, etc. The method can include the following steps:
[0049] 402. Determine the first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node, the first node and the second node being any two nodes of the at least two nodes obtained from the search space.
[0050] In an embodiment of the present application, the electronic device can obtain at least two nodes from the pre-established search space through a certain search strategy. The number and type of nodes obtained can be determined according to the target task (e.g., image super-resolution task, image segmentation task, etc.) corresponding to the neural network to be constructed, or can be set by an open person according to a large amount of development experience, which is not limited here.
[0051] It can be understood that the node represents the characteristic quantity of the input or output, and it is necessary to determine which calculation operation (i.e., operator) needs to be performed between two characteristic quantities. In fact, all operators existing in the search space can become the optimal operator corresponding to the first node and the second node, so the plurality of operators corresponding to the first node and the second node described above can be all operators existing in the search space.
[0052] Determining an optimal operator from all operators existing in the search space is actually a black-box optimization process based on a discrete space, which is difficult to search. For this purpose, the electronic device can perform relaxation processing on the operators existing in the search space according to a normalization function (e.g., softmax function) to obtain weight values corresponding to the plurality of operators, respectively, and use the weight values corresponding to each operator as the first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node, respectively. Thus, all operator parameters existing in the search space are parameterized to convert the black-box optimization problem in the discrete space into a problem of finding the maximum parameter, thereby simplifying the subsequent optimal operator search process.
[0053] 404. Regularize the initial loss function to obtain a target loss function, and optimize the first architecture parameters corresponding to the plurality of operators according to the target loss function to obtain second architecture parameters corresponding to the plurality of operators.
[0054] In an embodiment of the present application, the electronic device can optimize the first architecture parameters by gradient descent to determine the architecture parameters with the smallest loss. The process of gradient descent needs to supervise the optimization through the loss function, so the definition of the loss function will affect the subsequent optimization effect.
[0055] In related technologies, the aforementioned initial loss function is typically used to guide the optimization process of the first architecture parameters. The initial loss function can include various loss functions commonly used in neural network training, including but not limited to L1 loss function, L2 loss function, and hinge loss function. However, in practice, it has been found that optimizing the first architecture parameters through the initial loss function will cause the optimized parameters to fall within a concentrated range, with little difference between the architecture parameters, making it difficult to select operators.
[0056] In this embodiment of the application, the electronic device can perform regularization on the initial loss function to obtain the target loss function. The second architecture parameters optimized by the target loss function are more discrete, which improves the difference between the various second architecture parameters. This makes it easier to determine the largest second architecture parameter among multiple operators more accurately, and thus more accurately determine the optimal operator.
[0057] Optionally, the electronic device may add a regularization term to the initial loss function to regularize it. Optionally, the regularization term satisfies constraints, including but not limited to: differentiability within the domain; the regularization term reaching its maximum value when the architecture parameter is 0.5; and the reciprocal of the regularization term approaching 0 when the architecture parameter is 0.5.
[0058] 406. Among multiple operators, determine the operator with the largest second architecture parameter as the target operator corresponding to the first node and the second node.
[0059] In this embodiment, since the second architecture parameters optimized by the target loss function are more discrete and the differences between the second architecture parameters corresponding to each operator are greater, the electronic device can directly determine the target operator with the largest second architecture parameters from among the multiple operators by comparing the second architecture parameters corresponding to multiple operators.
[0060] Since the second architecture corresponding to the target operator is the largest, the model assigns it a higher weight. Therefore, the target operator can be used as the optimal operator corresponding to the first and second nodes. Optionally, the electronic device can discard all operators except the target operator from among multiple operators, retaining only the target operator.
[0061] 408. Determine a subnetwork based on at least two nodes and the target operators corresponding to any two of the at least two nodes, and determine the target neural network based on multiple subnetworks.
[0062] In this embodiment of the application, after determining the target operator between every two nodes in at least two nodes, the electronic device can determine the sub-network (e.g., based on the at least two nodes and the target operators corresponding to any two of the at least two nodes) according to the at least two nodes. Figure 2 (The subnetwork shown).
[0063] Further, the electronic device can fill the plurality of sub-networks into the network framework to be filled according to a certain filling rule to obtain the target neural network. The network framework can refer to a neural network framework, which defines the positions of modules inside the network, the connection modes between the modules, and model parameters, etc. According to different tasks performed by the network, the network framework can include a feedforward neural network framework, a recurrent network framework, and a symmetric connection network framework, etc.
[0064] In the embodiments of the present application, the electronic device can fill the sub-networks into the network framework as modules (for example, up-sampling modules, down-sampling modules, pooling modules, etc.) inside the network framework, and fill the sub-networks into the network framework according to the module positions, the connection modes between the modules, and the model parameters defined by the network architecture, to obtain the target neural network.
[0065] Implementing the method disclosed in each of the above embodiments can first obtain at least two nodes from the search space, and determine the first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node; further, the initial loss function can be regularized to obtain a target loss function, and the first architecture parameters corresponding to the plurality of operators can be optimized according to the target loss function; wherein the first architecture parameters are optimized by the target loss function after the regularization processing, which can make the plurality of second architecture parameters obtained after the optimization more discrete, improve the difference between each second architecture parameter, and make it more accurate to determine the largest second architecture parameter in the plurality of operators subsequently; further, the operator corresponding to the largest second architecture parameter can be used as the optimal operator of the first node and the second node, thereby improving the quality of the operators inside the sub-network, and further improving the performance of the target neural network built subsequently according to the plurality of sub-networks.
[0066] Please refer to Figure 5 , Figure 5 is a flow diagram of another method for searching a neural network architecture disclosed in the embodiments of the present application. The method can be applied to various electronic devices that can be used to build a neural network, including but not limited to cloud servers, personal computers, etc. The method can include the following steps:
[0067] 502, determining the first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node, the first node and the second node being any two nodes of the at least two nodes obtained from the search space.
[0068] In one embodiment, the electronic device can obtain at least two nodes from the search space according to the target task corresponding to the target neural network, the number and type of the at least two nodes matching the target task, and the target neural network being a neural network to be built.
[0069] Further, the electronic device can relax the plurality of operators corresponding to the first node and the second node according to the normalization function to obtain first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node, the first node and the second node being any two nodes in the at least two nodes.
[0070] Optionally, the normalization function can include a softmax function, a sigmoid function, etc., which is not limited herein.
[0071] Implementing the above method can make the obtained node match a target task corresponding to the target neural network to be constructed, so that the subsequently generated target neural network can better perform the target task.
[0072] In an embodiment, the electronic device can relax the plurality of operators corresponding to the first node and the second node by a sigmoid function to obtain first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node.
[0073] Implementing the above method, the electronic device can relax the plurality of operators corresponding to the first node and the second node by a sigmoid function, so that the obtained first architecture parameters corresponding to the plurality of operators are independent of each other, thereby avoiding the situation that the plurality of operators corresponding to the first node and the second node are dependent on each other, causing the operators to be difficult to be selected.
[0074] 504. Determine a regularization term according to the first architecture parameters corresponding to the plurality of operators, and regularize the initial loss function according to the regularization term to obtain a target loss function.
[0075] In the embodiments of the present application, after determining the first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node, the electronic device can further determine a regularization term according to the first architecture parameters corresponding to the plurality of operators.
[0076] Optionally, the electronic device can determine the regularization term according to the first architecture parameters corresponding to the plurality of operators and the following formula 1, that is:
[0077] Formula 1:
[0078]
[0079] wherein, L re represents the regularization term, N represents the number of the plurality of operators corresponding to the first node and the second node, σ(α i ) represents the first architecture parameter corresponding to the i-th operator.
[0080] Further, the electronic device can regularize the initial loss function according to the regularization term to obtain a target loss function.
[0081] 506、regularize the initial loss function according to the regularization term to obtain a target loss function.
[0082] Optionally, the electronic device calculates a first product of the regularization term and a weight coefficient corresponding to the regularization term, and determines the target loss function according to the first product and the initial loss function.
[0083] The weight coefficient corresponding to the regularization term can be set by a developer according to a large amount of development experience, which is not limited herein.
[0084] In an embodiment, the electronic device can determine the target loss function according to the regularization term, the weight coefficient corresponding to the regularization term, the initial loss function, and the following formula 2, that is:
[0085] Formula 2:
[0086] min w,α I SR +τ·L re
[0087] min w,α L SR +τ·L re represents the target loss function, L SR represents the initial loss function, L re represents the regularization term, τ represents the weight coefficient corresponding to the regularization term, and a represents an architecture parameter and w represents a model parameter.
[0088] Implementing the method disclosed in each of the above embodiments can regularize the initial loss function to obtain a target loss function, and the second architecture parameter optimized by the target loss function will be more discrete, improving the difference between each second architecture parameter, thereby facilitating more accurate determination of the largest second architecture parameter in the plurality of operators in the subsequent, and further determining the optimal operator.
[0089] 508、optimize the first architecture parameter corresponding to each of the plurality of operators according to the target loss function to obtain a second architecture parameter corresponding to each of the plurality of operators.
[0090] In the embodiments of the present application, the electronic device can optimize the first architecture parameter corresponding to each of the plurality of operators according to the target loss function by a target optimization method to obtain a second architecture parameter corresponding to each of the plurality of operators.
[0091] Optionally, the target optimization method can include an Adam optimization method.
[0092] Optionally, the electronic device can also optimize the first architecture parameters and the first model parameters corresponding to the plurality of operators respectively according to a target loss function, to obtain second architecture parameters and second model parameters corresponding to the plurality of operators respectively.
[0093] Optionally, the electronic device can fix the first model parameters (i.e., keep the first model unchanged), and optimize the first architecture parameters corresponding to the plurality of operators respectively according to the target loss function by a target optimization method, to obtain the second architecture parameters corresponding to the plurality of operators respectively.
[0094] Further, the electronic device can fix the second architecture parameters, and optimize the first model parameters corresponding to the plurality of operators respectively according to the target loss function by the target optimization method, to obtain the second model parameters corresponding to the plurality of operators respectively, and alternately optimize until the architecture parameters and the model parameters converge, to determine that the optimization of the architecture parameters and the model parameters is completed.
[0095] Optionally, the electronic device can fix the first model parameters, and determine the second architecture parameters corresponding to the plurality of operators respectively according to the first architecture parameters, the first model parameters and the following formula 3, that is:
[0096] Formula 3:
[0097] s.t.ω*(α)=ar gmin ω L train (ω,α)
[0098] Wherein, ω represents the first model parameters, α represents the first architecture parameters, and ω*(α) represents the second architecture parameters.
[0099] Optionally, the electronic device can fix the second architecture parameters, and determine the second model parameters corresponding to the plurality of operators respectively according to the first model parameters, the second architecture parameters and the following formula 4, that is:
[0100] Formula 4:
[0101]
[0102] Wherein, ω represents the first model parameters, α represents the second architecture parameters, and ω*(α) represents the second model parameters.
[0103] In practice, it is found that since the number of architecture parameters is much smaller than the number of model parameters, if all parameters (i.e., architecture parameters and model parameters) are directly optimized uniformly, the problem of overfitting of architecture parameters or underfitting of model parameters will be caused. To this end, the above method is implemented, and the architecture parameters and the model parameters are converged by the alternately optimization method, thereby avoiding the problem of overfitting of architecture parameters or underfitting of model parameters. The accuracy of subsequently determining the optimal operator is improved.
[0104] 510、determine the operator with the largest second architecture parameter in the plurality of operators as the target operator corresponding to the first node and the second node.
[0105] 512、determine sub-networks according to the at least two nodes and the target operators corresponding to any two nodes in the at least two nodes, and determine the target neural network according to the plurality of sub-networks.
[0106] The method disclosed in the above embodiments can be used to regularize the initial loss function to obtain a target loss function, and optimize the first architecture parameters corresponding to the plurality of operators according to the target loss function; wherein the first architecture parameters are optimized by the target loss function after the regularization, which can make the plurality of second architecture parameters obtained after the optimization more discrete, improve the difference between the second architecture parameters, and make it more accurate to determine the largest second architecture parameter in the plurality of operators subsequently; thereby improving the quality of the operators in the sub-networks, and further improving the performance of the target neural network built according to the plurality of sub-networks subsequently; and the nodes obtained can match the target task corresponding to the target neural network to be built, so that the target neural network generated subsequently can better perform the target task; and the plurality of operators corresponding to the first node and the second node can be relaxed by the sigmoid function, so that the first architecture parameters corresponding to the plurality of operators are independent of each other, thereby avoiding the situation that the plurality of operators corresponding to the first node and the second node are dependent on each other, and the operators are difficult to select; and the initial loss function can be regularized to obtain a target loss function, and the second architecture parameters optimized by the target loss function will be more discrete, which improves the difference between the second architecture parameters, thereby facilitating the subsequent determination of the largest second architecture parameter in the plurality of operators, and further determining the optimal operator.
[0107] Please refer to Figure 6 , Figure 6 is a flow diagram of another method for searching a neural network architecture disclosed in the embodiments of the present application. The method can be applied to various electronic devices that can be used to build a neural network, including but not limited to cloud servers, personal computers, etc. The method can include the following steps:
[0108] 602、determine the first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node, wherein the first node and the second node are any two nodes in the at least two nodes obtained from the search space.
[0109] In the embodiments of the present application, the definition of the search space can be inherited from the definition of the search space by the DARTS algorithm.
[0110] 604. Regularize the initial loss function to obtain the target loss function, and optimize the first architecture parameters corresponding to the multiple operators according to the target loss function to obtain the second architecture parameters corresponding to the multiple operators.
[0111] 606. Among multiple operators, determine the operator with the largest second architecture parameter as the target operator corresponding to the first node and the second node, and determine the sub-network based on at least two nodes and the target operators corresponding to any two of the at least two nodes.
[0112] 608. Obtain the network framework to be filled, and use multiple subnetworks to fill the network framework to obtain the target neural network.
[0113] As mentioned above, the network framework defines the positions of modules within the network, the connection methods between modules, and model parameters. Therefore, after obtaining the network framework to be filled, the electronic device can, based on the module positions, connection methods, and model parameters defined in the network framework, fill one or more sub-networks into the network framework as modules within the network framework to obtain the target neural network.
[0114] In one embodiment, the network framework may include a U-shaped network framework; see [link to relevant documentation]. Figure 7 , Figure 7 This is a schematic diagram of a U-shaped neural network disclosed in an embodiment of this application. The U-shaped network framework may include the same number of downsampling layers 700 and upsampling layers 710, with multiple downsampling layers and multiple upsampling layers connected end to end, and the multiple downsampling layers and multiple upsampling layers are symmetrically arranged; furthermore, the symmetrically arranged upsampling layers and downsampling layers may establish skip connections (for example, the first downsampling layer 720 and the last upsampling layer 730), the output of the last downsampling layer 740 may be connected to the first upsampling layer 750, or the output of the last downsampling layer 740 may be connected to the pooling layer 760, and the pooling layer 700 may be connected to the first upsampling layer 750, which is not limited here.
[0115] Optionally, the electronic device can obtain a target number of sub-networks as next-down layers and a target number of sub-networks as upsampling layers based on the framework information of the U-shaped network framework. The framework information can include the target number of upsampling and downsampling layers that the U-shaped network framework needs to fill, such as 4 or 5. The target number can be set by the designer of the U-shaped network framework or determined by the developers based on the target task corresponding to the target neural network; this is not limited here.
[0116] Further, the electronic device can fill the U-shaped network framework with the target number of down-sampling layers and the target number of up-sampling layers to obtain the target neural network.
[0117] Optionally, the framework information can further include the setting positions of the up-sampling layers and the down-sampling layers, and the connection modes of the up-sampling layers and the down-sampling layers; and the electronic device fills the target number of down-sampling layers and the target number of up-sampling layers into the U-shaped network framework according to the setting positions and the connection modes indicated by the framework information to obtain the target neural network.
[0118] Optionally, the target neural network obtained according to the U-shaped network framework and the plurality of sub-networks can be applied to an image super-resolution processing task to improve the effect of image super-resolution.
[0119] Implementing the above method can fill the U-shaped network framework with the plurality of sub-networks to obtain a target neural network suitable for an image super-resolution task. Since the performance of the target neural network generated by the method disclosed in the embodiments of the present application is higher, the effect of image super-resolution is improved.
[0120] Optionally, the construction standard corresponding to the sub-network can include a peak signal noise ratio (PSNR) and a structural similarity (SSIM), thereby improving the performance of the image super-resolution neural network constructed according to the sub-network.
[0121] In another embodiment, the electronic device can fill the U-shaped network framework with the target number of down-sampling layers and the target number of up-sampling layers to obtain a transition neural network.
[0122] Further, the electronic device can obtain an image processing operator, which is an operator for image super-resolution processing of an image, including but not limited to a PixelShuffle operator (see 770 in Figure 7 Further, the electronic device can connect the image processing operator with a target up-sampling layer in the transition neural network to obtain a target neural network. The target up-sampling layer is an up-sampling layer arranged at the last in the transition neural network.
[0123] Implementing the above method can further add an operator for image super-resolution processing of an image in the constructed target neural network, thereby improving the image super-resolution capability of the target neural network.
[0124] It can be understood that after the target neural network is determined according to the plurality of sub-networks, the target neural network is still in a verification stage, that is, the target neural network is not necessarily suitable for performing the target task (for example, image super-resolution). In this regard, after the target neural network is determined according to the plurality of sub-networks, the electronic device can train the target neural network by using the training data set, and verify the output result of the trained target neural network by using the verification data set, to obtain a verification index of the target neural network, wherein the verification index is used to indicate the accuracy of the output result of the trained target neural network.
[0125] Optionally, the training data set can include one or more images that have not undergone image super-resolution processing, and the verification data set can include one or more images that have undergone image super-resolution processing by other trained neural networks, which is not limited herein.
[0126] Further, if the verification index meets the index requirement, the electronic device can determine that the target neural network is trained.
[0127] Implementing the above method can further train the target neural network after the target neural network is determined, and verify the performance of the trained target neural network, to determine a target neural network with better performance.
[0128] In another embodiment, the electronic device can also obtain one or more second output results of one or more trained target neural networks by using one or more second training data sets, and verify the one or more second output results by using the verification data set, to obtain one or more second verification indexes, and if the one or more second verification indexes meet the index requirement, the electronic device can determine that the target neural network is trained.
[0129] Implementing the above method can train the target neural network by using a plurality of training data sets, to verify whether the target neural network trained by using other training data can also obtain a high-performance network, thereby improving the generalization performance of the target neural network.
[0130] By implementing the method disclosed in the above embodiments, the initial loss function can be regularized to obtain a target loss function, and the first architecture parameters corresponding to the plurality of operators are optimized according to the target loss function; wherein the first architecture parameters are optimized by the target loss function after the regularization processing, which can make the plurality of second architecture parameters obtained after the optimization more discrete, improve the difference between each second architecture parameter, and enable the maximum second architecture parameter to be more accurately determined in the plurality of operators subsequently; thereby improving the quality of the operators within the subnetwork, and further improving the performance of the target neural network built according to the plurality of subnetworks subsequently; and the U-shaped network framework can be filled by the plurality of subnetworks to obtain a target neural network suitable for an image super-resolution task, since the performance of the target neural network generated by the method disclosed in the embodiments of the present application is higher, the image super-resolution effect is improved; and an operator for image super-resolution processing of an image can be further added in the constructed target neural network, thereby improving the image super-resolution capability of the target neural network; and after the target neural network is determined, the target neural network can be further trained, and the performance of the trained target neural network can be verified to determine a target neural network with better performance.
[0131] Please refer to Figure 8 , Figure 8 is a structural schematic diagram of a neural network architecture searching device disclosed in an embodiment of the present application. The device can be applied to various electronic devices that can be used to build a neural network, including but not limited to a cloud server, a personal computer, etc. The device can include a first determining unit 801, an optimization unit 802, a second determining unit 803, and a third determining unit 804, wherein:
[0132] The first determining unit 801 is configured to determine first architecture parameters corresponding to a plurality of operators corresponding to a first node and a second node, the first node and the second node being any two nodes in at least two nodes obtained from a search space;
[0133] The optimization unit 802 is configured to regularize an initial loss function to obtain a target loss function, and optimize the first architecture parameters corresponding to the plurality of operators according to the target loss function, to obtain second architecture parameters corresponding to the plurality of operators;
[0134] The second determining unit 803 is configured to determine an operator with the maximum second architecture parameter as a target operator corresponding to the first node and the second node in the plurality of operators;
[0135] The third determining unit 804 is configured to determine a subnetwork according to the at least two nodes and the target operators corresponding to any two nodes in the at least two nodes, and determine a target neural network according to the plurality of subnetworks.
[0136] By implementing the above device, the first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node can be determined from the at least two nodes obtained from the search space. Further, the initial loss function can be regularized to obtain the target loss function, and the first architecture parameters corresponding to the plurality of operators can be optimized according to the target loss function. The plurality of second architecture parameters obtained after optimization can be more discrete by optimizing the first architecture parameters through the target loss function after regularization, which improves the difference between each second architecture parameter, so that the largest second architecture parameter can be more accurately determined in the plurality of operators subsequently. Further, the operator corresponding to the largest second architecture parameter can be used as the optimal operator of the first node and the second node, thereby improving the quality of the operators in the subnetwork and further improving the performance of the target neural network built according to the plurality of subnetworks subsequently.
[0137] As an optional implementation, the optimization unit 802 is further configured to determine a regularization term according to the first architecture parameters corresponding to the plurality of operators, and to regularize the initial loss function according to the regularization term to obtain the target loss function.
[0138] By implementing the above device, the initial loss function can be regularized to obtain the target loss function, and the second architecture parameters optimized through the target loss function will be more discrete, which improves the difference between each second architecture parameter, thereby facilitating more accurate determination of the largest second architecture parameter in the plurality of operators subsequently, and further determining the optimal operator.
[0139] As an optional implementation, the optimization unit 802 is further configured to calculate a first product of the regularization term and a weight coefficient corresponding to the regularization term, and to determine the target loss function according to the first product and the initial loss function.
[0140] By implementing the above device, the initial loss function can be regularized to obtain the target loss function, and the second architecture parameters optimized through the target loss function will be more discrete, which improves the difference between each second architecture parameter, thereby facilitating more accurate determination of the largest second architecture parameter in the plurality of operators subsequently, and further determining the optimal operator.
[0141] As an optional implementation, the first determination unit 801 is further configured to obtain at least two nodes from the search space, and to relax the plurality of operators corresponding to the first node and the second node according to a normalization function to obtain the first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node.
[0142] The device can make the obtained node match a target task corresponding to the target neural network to be constructed, so that the generated target neural network can better perform the target task.
[0143] As an optional implementation, the normalization function includes a sigmoid function.
[0144] The device can relax the plurality of operators corresponding to the first node and the second node through the sigmoid function, so that the plurality of operators obtained correspond to independent first architecture parameters, thereby avoiding the mutual dependence between the plurality of operators corresponding to the first node and the second node, and the situation that the operators are difficult to select.
[0145] As an optional implementation, the third determination unit 804 is further configured to obtain a network framework to be filled, and fill the network framework to be filled by using the plurality of sub-networks to obtain the target neural network.
[0146] The device can fill the network framework to be filled by using the high-quality sub-networks to obtain the high-performance target neural network, thereby improving the performance of the target neural network.
[0147] As an optional implementation, the network framework includes a U-shaped network framework; the third determination unit 804 is further configured to obtain a target number of sub-networks as down-sampling layers and a target number of sub-networks as up-sampling layers according to framework information of the U-shaped network framework, the framework information including target numbers of the up-sampling layers and the down-sampling layers corresponding to the U-shaped network framework to be filled; and fill the U-shaped network framework by using the target number of down-sampling layers and the target number of up-sampling layers to obtain the target neural network.
[0148] The device can fill the U-shaped network framework by using the plurality of sub-networks to obtain the target neural network suitable for an image super-resolution task, and the performance of the target neural network generated by the method disclosed in the embodiments of the present application is higher, thereby improving the effect of image super-resolution.
[0149] As an optional implementation, the third determination unit 804 is further configured to fill the U-shaped network framework by using the target number of down-sampling layers and the target number of up-sampling layers to obtain a transition neural network; and obtain an image processing operator and connect the image processing operator and a target up-sampling layer in the transition neural network to obtain the target neural network, the image processing operator being used for super-resolution processing of an image, and the target up-sampling layer being an up-sampling layer arranged at the last in the transition neural network.
[0150] The device can further add an operator for image super-resolution processing of the image in the constructed target neural network, so as to improve the image super-resolution capability of the target neural network.
[0151] As an optional implementation, Figure 7 The device also includes a verification unit and a fourth determination unit, which are not shown.
[0152] The verification unit is configured to train the target neural network using a training data set after determining the target neural network according to the plurality of sub-networks, and verify the output result of the trained target neural network using a verification data set to obtain a verification index of the target neural network, the verification index being used to indicate the accuracy of the output result of the trained target neural network.
[0153] The fourth determination unit is configured to determine that the training of the target neural network is completed when the verification index meets an index requirement.
[0154] The device can further train the target neural network after determining the target neural network, and verify the performance of the trained target neural network to determine a target neural network with better performance.
[0155] Please refer to Figure 9 , Figure 9 is a structural schematic diagram of an electronic device disclosed by the embodiments of the present application. As Figure 9 shown, the electronic device can include:
[0156] a memory 901 storing executable program codes;
[0157] a processor 902 coupled with the memory 901;
[0158] The processor 902 calls the executable program codes stored in the memory 901 to execute the search method of the neural network architecture disclosed in each of the above embodiments.
[0159] The embodiments of the present application disclose a computer readable storage medium storing a computer program, wherein the computer program causes a computer to execute the search method of the neural network architecture disclosed in each of the above embodiments.
[0160] The embodiments of the present application also disclose an application publishing platform, wherein the application publishing platform is used to publish a computer program product, and when the computer program product runs on a computer, the computer is caused to execute part or all steps of the method in each of the above method embodiments.
[0161] It should be understood that the term "one embodiment" or "an embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in one embodiment" or "in an embodiment" in various places in the specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It will also be appreciated by those of skill in the art that references to a structure or feature that is
[0162] In various embodiments of the present application, it should be understood that the magnitude of the serial number of the above-mentioned processes does not mean the inevitable sequence of execution, and the execution sequence of the processes should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0163] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e. they can be located in one place, or they can be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0164] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0165] The integrated unit described above, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer accessible memory. Based on such understanding, the technical solutions of the present application, essentially or the part that makes a contribution to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product, which is stored in a memory and includes a number of steps for causing a computer device (which can be a personal computer, a server or a network device, etc., and specifically can be a processor in the computer device) to execute the methods of the embodiments of the present application described above.
[0166] Those skilled in the art can understand that all or part of the steps of various methods in the above embodiments can be completed by instructing the relevant hardware by means of a program, and the program can be stored in a computer readable storage medium, including Read-Only Memory (ROM), Random Access Memory (RAM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), One-time Programmable Read-Only Memory (OTPROM), Electrically-Erasable Programmable Read-Only Memory (EEPROM), Compact Disc Read-Only Memory (CD-ROM) or other optical disk storage, magnetic disk storage, magnetic tape storage, or any other medium that can be used to carry or store data in a computer readable manner.
[0167] The above describes in detail the search method and device of the neural network architecture, the electronic device, and the storage medium disclosed in the embodiments of the present application. The principles and implementation manners of the present application are described by applying specific examples. The above embodiment descriptions are only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges can be changed. In summary, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method for searching a neural network architecture, characterized in that, The method is applied to construction of a neural network for image super-resolution, and the method comprises the following steps: determining first architecture parameters corresponding to a plurality of operators corresponding to a first node and a second node, the first node and the second node being any two nodes in at least two nodes obtained from a search space; determining a regularization term according to the first architecture parameters corresponding to the plurality of operators, performing regularization processing on an initial loss function according to the regularization term to obtain a target loss function, and optimizing the first architecture parameters corresponding to the plurality of operators according to the target loss function to obtain second architecture parameters corresponding to the plurality of operators, the target loss function being used to make the plurality of second architecture parameters obtained after optimization more discrete; determining an operator with the largest second architecture parameter in the plurality of operators as a target operator corresponding to the first node and the second node; determining a subnetwork according to the at least two nodes and the target operators corresponding to any two nodes in the at least two nodes, and determining a target neural network according to a plurality of subnetworks; wherein the determining of the regularization term according to the first architecture parameters corresponding to the plurality of operators comprises: determining the regularization term according to the first architecture parameters corresponding to the plurality of operators and Formula 1, that is: Formula 1: Wherein the first node and the second node correspond to a plurality of operators, and the first node and the second node correspond to a plurality of first architecture parameters and a plurality of second architecture parameters. Wherein, the regularization term is represented by N, which represents the number of the plurality of operators corresponding to the first node and the second node, and the The first architecture parameter corresponding to the i-th operator is represented by 2. The method of claim 1, wherein, the regularization processing on the initial loss function according to the regularization term to obtain the target loss function comprises: calculating a first product of the regularization term and a weight coefficient corresponding to the regularization term; determining the target loss function according to the first product and the initial loss function.
3. The method of claim 1, wherein, the determining of the first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node comprises: obtaining at least two nodes from a search space; performing relaxation processing on the plurality of operators corresponding to the first node and the second node according to a normalization function to obtain the first architecture parameters corresponding to the plurality of operators corresponding to the first node and the second node.
4. The method of claim 3, wherein, The normalization function comprises a sigmoid function.
5. The method of claim 1, wherein, the determining of the target neural network according to the plurality of subnetworks comprises: obtaining a network framework to be filled, and filling the network framework to be filled by using the plurality of subnetworks to obtain the target neural network.
6. The method of claim 5, wherein, The network framework comprises a U-shaped network framework; the filling of the network framework to be filled by using the plurality of subnetworks to obtain the target neural network comprises: obtaining a target number of the subnetworks as down-sampling layers and a target number of the subnetworks as up-sampling layers according to framework information of the U-shaped network framework, the framework information comprising target numbers of the up-sampling layers and the down-sampling layers corresponding to the U-shaped network framework to be filled; filling the U-shaped network framework by using the target number of the down-sampling layers and the target number of the up-sampling layers to obtain the target neural network.
7. The method of claim 6, wherein, the filling of the network framework to be filled by using the target number of the down-sampling layers and the target number of the up-sampling layers to obtain the target neural network comprises: Fill the U-shaped network framework with the target number of down-sampling layers and the target number of up-sampling layers, to obtain a transition neural network; An image processing operator is obtained, and the image processing operator is connected with a target up-sampling layer in the transition neural network to obtain a target neural network, the image processing operator being used for image super-resolution processing of an image, and the target up-sampling layer being an up-sampling layer arranged at the last in the transition neural network.
8. The method according to any one of claims 1 to 7, characterized in that, After the target neural network is determined according to the plurality of sub-networks, the method further comprises: The target neural network is trained by using a training data set, and an output result of the trained target neural network is verified by using a verification data set, to obtain a verification index of the target neural network, the verification index being used for indicating an accuracy rate of the output result of the trained target neural network; If the verification index meets an index requirement, it is determined that the training of the target neural network is completed.
9. An apparatus for searching a neural network architecture, the apparatus comprising: The apparatus is applied to construction of a neural network for image super-resolution, and the apparatus comprises: A first determining unit is configured to determine first architecture parameters corresponding to a plurality of operators corresponding to a first node and a second node, the first node and the second node being any two nodes of at least two nodes obtained from a search space; An optimization unit is configured to determine a regularization term according to the first architecture parameters corresponding to the plurality of operators, to perform regularization processing on an initial loss function according to the regularization term, to obtain a target loss function, and to optimize the first architecture parameters corresponding to the plurality of operators according to the target loss function, to obtain second architecture parameters corresponding to the plurality of operators, the target loss function being used to make the plurality of second architecture parameters more discrete after optimization; A second determining unit is configured to determine an operator with the largest second architecture parameter as a target operator corresponding to the first node and the second node from the plurality of operators; A third determining unit is configured to determine a sub-network according to the at least two nodes and the target operators corresponding to any two nodes of the at least two nodes, and to determine a target neural network according to a plurality of sub-networks; The regularization term is determined according to the first architecture parameters corresponding to the plurality of operators and formula 1, that is, Formula 1: The computer program is executed by the processor to implement the method of any one of claims 1-8. The first node and the second node are connected by a plurality of operators. represents a regularization term, N represents a number of the plurality of operators corresponding to the first node and the second node, and represents a first architecture parameter corresponding to the i-th operator.
10. An electronic device, comprising: The computer program is executed by the processor to implement the method of any one of claims 1-8.
11. A computer readable storage medium storing a computer program, wherein the computer program comprises instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 10. The computer program is executed by the processor to implement the method of any one of claims 1-8.
Citation Information
Patent Citations
Neural network structure obtaining method and device, and storage medium
CN110428046A
Neural network architecture search method and device, image processing method and device and storage medium
CN112561027A