Neural Network Architecture Search Optimization Method, Device and Equipment

Multiple search spaces are generated by transforming the parameters of the neural network model, and using evolutionary algorithm optimization and optimal structure method, the problem of long search time for multi-objective NAS is solved and efficiency is improved.

CN115081622BActive Publication Date: 2025-05-27PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210820860.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-05-27
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

The existing multi-objective neural network structure search (NAS) method has caused the search time to increase exponentially and is inefficient due to the expansion of the search space.

Method used

By transforming the width parameters and input parameters of the neural network model, multiple search spaces are generated, and the neural network model is optimized using evolutionary algorithms to determine the optimal structure, while narrowing the search scope through memory resource limitations.

Benefits of technology

It effectively reduces the use of space, time and resources by neural network models, improves the execution efficiency of NAS, and shortens search time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115081622B_ABST
    Figure CN115081622B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence, and provides a method, apparatus, and device for optimizing neural network architecture search. The method includes obtaining a first neural network model; obtaining multiple different first search spaces by transforming the width parameter and input parameter of the first neural network model; for each first search space, determining a second neural network model corresponding to the first search space according to the width parameter and input parameter corresponding to the first search space; performing performance evaluation on the second neural network model corresponding to each first search space to obtain a first performance evaluation value corresponding to each first search space; determining a second search space from the multiple first search spaces according to the first performance evaluation value; obtaining a third neural network model according to the second search space; and using an evolutionary algorithm to perform search optimization processing on the third neural network model to determine the optimal structure of the third neural network model. The embodiments of this application can improve the execution efficiency of NAS.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a method, device, and equipment for optimizing neural network architecture search. Background Art

[0002] Neural Architecture Search (NAS) is a method for constructing a neural network model. This method can automatically adjust parameters, greatly reducing the time for manual adjustment of hyperparameters. Due to the emergence of edge devices such as MicroController Units (MCUs), people have also included model size and model performance in the optimization metrics of NAS, thus constructing multi-objective NAS. However, since the search space of multi-objective NAS will greatly increase, the search time also increases exponentially. Summary of the Invention

[0003] The purpose of this application is to solve the problems of the prior art to at least a certain extent, and provide a method, device, and equipment for optimizing neural network architecture search, improving the execution efficiency of NAS, and thus reducing the search time.

[0004] The technical solutions of the embodiments of this application are as follows:

[0005] In the first aspect, this application provides a method for optimizing neural network architecture search. The method includes:

[0006] Obtain a first neural network model;

[0007] By transforming the width parameter and input parameter of the first neural network model, obtain multiple different first search spaces;

[0008] For each of the first search spaces, according to the width parameter and the input parameter corresponding to the first search space, determine a second neural network model corresponding to the first search space;

[0009] Perform performance evaluation on the second neural network model corresponding to each first search space, and obtain a first performance evaluation value corresponding to each first search space;

[0010] Determine a second search space from the multiple first search spaces according to the first performance evaluation value. The second search space represents a search space that meets the preset memory resource limit condition;

[0011] Obtain a third neural network model according to the second search space;

[0012] Use an evolutionary algorithm to perform search optimization processing on the third neural network model to determine the optimal structure of the third neural network model.

[0013] According to some embodiments of the present application, the method of using an evolutionary algorithm to perform search optimization processing on the third neural network model to determine the optimal structure of the third neural network model includes:

[0014] Performing optimization processing on the third neural network model to obtain a fourth neural network model;

[0015] Calculating a second performance evaluation value and a third performance evaluation value of the fourth neural network model, where the second performance evaluation value represents the memory resource occupancy, and the third performance evaluation value represents the number of floating-point operations;

[0016] According to the second performance evaluation value and the third performance evaluation value, selecting the fourth neural network model that meets the memory resource limit condition and the preset floating-point operation number limit condition as the parent network;

[0017] Performing evolutionary calculation on the parent network using the evolutionary algorithm to obtain a subnet, where the parent network is connected to the subnet through weight sharing;

[0018] Taking the subnet as the third neural network model, and re-performing optimization processing on the third neural network model;

[0019] When the preset number of iterations is reached, obtaining the optimal structure of the third neural network model according to the subnet corresponding to the maximum number of floating-point operations.

[0020] According to some embodiments of the present application, the step of performing optimization processing on the third neural network model to obtain a fourth neural network model includes:

[0021] Calculating the Euclidean norm corresponding to each network layer of the third neural network model to obtain the norm value corresponding to each network layer;

[0022] Pruning the network layer corresponding to the norm value lower than the preset threshold to obtain a pruned network layer;

[0023] Updating the weights of the pruned network layer, and recalculating the Euclidean norm corresponding to the network layer;

[0024] Calculating the sparsity of the third neural network model according to the norm value corresponding to each network layer, and when the sparsity reaches the preset sparse value, ending the pruning process to obtain the fourth neural network model.

[0025] According to some embodiments of the present application, after performing evolutionary calculation on the parent network using the evolutionary algorithm to obtain a subnet, the method further includes:

[0026] For the sub-network, a scheduling algorithm is used to perform priority sorting on the sub-network to obtain a model priority;

[0027] According to the model priority, the sub-networks that meet the memory resource limit condition and the floating-point operation count limit condition are selected.

[0028] According to some embodiments of the present application, obtaining multiple different first search spaces by transforming the width parameter and input parameter of the first neural network model includes:

[0029] By transforming the width parameter and input parameter of the first neural network model, multiple different third search spaces are obtained;

[0030] Based on sampling processing of multiple third search spaces, multiple first search spaces are obtained.

[0031] According to some embodiments of the present application, obtaining multiple different third search spaces by transforming the width parameter and input parameter of the first neural network model includes:

[0032] Using a width multiplier parameter to transform the width parameter of the first neural network model, where the value range of the width multiplier parameter is between 0 and 1;

[0033] By transforming the resolution of the image, the input parameter of the first neural network model is transformed;

[0034] According to the width parameter and the input parameter, each third search space is obtained.

[0035] According to some embodiments of the present application, calculating the Euclidean norm corresponding to each network layer of the third neural network model to obtain the norm value corresponding to each network layer includes:

[0036] Calculating the Euclidean norm corresponding to each network layer of the third neural network model through the following formula:

[0037]

[0038] where d L2 represents the norm value, C 1i represents a channel of the network layer, C 2i represents another channel of the network layer, and n represents the number of channels of the network layer.

[0039] In a second aspect, the present application provides a neural network structure search optimization device, including:

[0040] A model acquisition module, configured to acquire a first neural network model;

[0041] A parameter processing module, configured to obtain a plurality of different first search spaces by transforming the width parameter and the input parameter of the first neural network model;

[0042] A first processing module, configured to determine, for each of the first search spaces, a second neural network model corresponding to the first search space according to the width parameter and the input parameter corresponding to the first search space;

[0043] A performance evaluation module, configured to perform performance evaluation on the second neural network model corresponding to each first search space, and obtain a first performance evaluation value corresponding to each first search space;

[0044] A second processing module, configured to determine a second search space from the plurality of first search spaces according to the first performance evaluation value, where the second search space represents a search space that satisfies a preset memory resource limit condition;

[0045] A third processing module, configured to obtain a third neural network model according to the second search space;

[0046] An optimization processing module, configured to perform search optimization processing on the third neural network model by using an evolutionary algorithm to determine an optimal structure of the third neural network model.

[0047] In a third aspect, the present application provides a computer device, where the computer device includes a memory and a processor, and computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of any one of the methods described in the first aspect above.

[0048] In a fourth aspect, the present application further provides a computer-readable storage medium, where the storage medium can be read and written by a processor, and the storage medium stores computer instructions. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of any one of the methods described in the first aspect above.

[0049] The technical solution provided by the embodiments of the present application has the following beneficial effects:

[0050] An embodiment of the present application provides a method, apparatus, and device for optimizing neural network architecture search. The method first obtains a first neural network model; by transforming the width parameter and input parameter of the first neural network model, a plurality of different first search spaces are obtained; for each first search space, according to the width parameter and input parameter corresponding to the first search space, a second neural network model corresponding to the first search space is determined; the performance of the second neural network model corresponding to each first search space is evaluated to obtain a first performance evaluation value corresponding to each first search space; according to the first performance evaluation value, a second search space is determined from the plurality of first search spaces, and the second search space represents a search space that meets the preset memory resource limit condition. By obtaining the second search space through the memory resource limit, the search range can be reduced, thereby reducing the occupation of space, time, and resources by the neural network model, and improving the execution efficiency of NAS; a third neural network model is obtained according to the second search space; an evolutionary algorithm is used to perform search optimization processing on the third neural network model to determine the optimal structure of the third neural network model. The embodiment of the present application improves the execution efficiency of NAS, thereby reducing the search time. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a flowchart of a method for optimizing neural network architecture search provided by an embodiment of the present application;

[0052] Figure 2 is Figure 1 a sub-step flowchart of step S700 in

[0053] Figure 3 is Figure 2 a sub-step flowchart of step S710 in

[0054] Figure 4 is a flowchart of a method for optimizing neural network architecture search provided by another embodiment of the present application;

[0055] Figure 5 is Figure 1 a sub-step flowchart of step S200 in

[0056] Figure 6 is Figure 5 a sub-step flowchart of step S210 in

[0057] Figure 7 is an overall flowchart of a method for optimizing neural network architecture search provided by another embodiment of the present application;

[0058] Figure 8 is a structural diagram of an apparatus for optimizing neural network architecture search provided by an embodiment of the present application;

[0059] Figure 9 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0060] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0061] It should be noted that unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0062] The embodiment of the present application provides a neural network structure search optimization method, device and equipment. The method first obtains a first neural network model; obtains a plurality of different first search spaces by transforming the width parameter and input parameter of the first neural network model; for each first search space, determines a second neural network model corresponding to the first search space according to the width parameter and input parameter corresponding to the first search space; performs performance evaluation on the second neural network model corresponding to each first search space to obtain a first performance evaluation value corresponding to each first search space; determines a second search space from the plurality of first search spaces according to the first performance evaluation value, and the second search space represents a search space that meets the preset memory resource limit condition. Obtaining the second search space through the memory resource limit can reduce the search range, thereby reducing the occupancy of the neural network model on space, time and resources, and improving the execution efficiency of NAS; obtains a third neural network model according to the second search space; uses an evolutionary algorithm to perform search optimization processing on the third neural network model to determine the optimal structure of the third neural network model. Through the evolutionary algorithm, a neural network model with a higher accuracy can be obtained. The embodiment of the present application improves the execution efficiency of NAS, thereby reducing the search time.

[0063] It should be noted that the neural network structure search optimization method is applied to a deep neural network model deployed on a microcontroller. In the case of multiple evaluation metrics, it can reduce the search space, thereby reducing resource occupancy; this method can also be applied to a deep neural network model deployed in other controllers, which can reduce the search time of the optimal neural network model and has a wide range of applications.

[0064] Embodiments of the present application can be used in numerous general or specific computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0065] Next, with reference to the accompanying drawings, an optimization method, apparatus, and device for neural network structure search provided by embodiments of the present application will be described.

[0066] See Figure 1 , Figure 1 FIG. shows a schematic flowchart of an optimization method for neural network structure search provided by an embodiment of the present application. The above method includes but is not limited to steps S100, S200, S300, S400, S500, S600, and S700.

[0067] Step S100, obtain a first neural network model.

[0068] It can be understood that the first neural network model can be a deep neural network model, a shallow neural network model, or a lightweight neural network model, which will not be elaborated here. First, obtaining the first neural network model as the initial neural network model is beneficial for subsequent automatic adjustment of the hyperparameters of the initial neural network model.

[0069] Step S200, by transforming the width parameter and input parameter of the first neural network model, obtain multiple different first search spaces.

[0070] Refer to Figure 5 , by transforming the width parameter and input parameter of the first neural network model, obtain multiple different first search spaces, including but not limited to steps S210 and S220.

[0071] Step S210, by transforming the width parameter and input parameter of the first neural network model, obtain multiple different third search spaces.

[0072] It can be understood that for the first neural network model obtained according to step S100, the width parameter and the input parameter of the first neural network model are transformed to form a pair of width parameter and input parameter. One of the parameters or both parameters are transformed to obtain multiple different pairs of width parameter and input parameter. These different pairs of width parameter and input parameter constitute multiple different third search spaces. By obtaining the third search space, it is beneficial to perform subsequent computational processing for narrowing the search space.

[0073] Reference Figure 6 , by transforming the width parameter and the input parameter of the first neural network model, multiple different third search spaces are obtained, including but not limited to the following steps:

[0074] Step S211, use the width multiplier parameter to transform the width parameter of the first neural network model, where the value range of the width multiplier parameter is between 0 and 1.

[0075] It can be understood that the width parameter acts on the number of channels of the model input and output. The value range of the width multiplier parameter is between 0 and 1. Selecting different width multiplier parameters between 0 and 1, denoted as a, makes the number of input channels change from M to aM, and the number of output channels change from N to aN, obtaining different width parameters. The number of channels of the first neural network model also changes accordingly, thus realizing the automatic adjustment of the model hyperparameters. By transforming the width parameter, different hyperparameter groups of the first neural network model can be obtained, which is beneficial to subsequent searching for the optimal structure of the third neural network model among these hyperparameters.

[0076] Step S212, transform the input parameter of the first neural network model by changing the resolution of the image.

[0077] It can be understood that the input parameter of the first neural network model represents the parameter of the image size input to the first neural network model. By changing the resolution of the image (i.e., changing the size of the image, the size of the model input picture can range from 16*16 to 224*22), different input parameters are obtained. The input of the first neural network model is images with different resolutions, which is applicable to the processing of images of different sizes and has wide applicability.

[0078] Step S213, obtain each third search space according to the width parameter and the input parameter.

[0079] It can be understood that, according to step S211 and step S212, by varying the width parameter and the input parameter, pairs of width parameter and input parameter are formed, and each pair of width parameter and input parameter corresponds to a third search space. By varying the width multiplier parameter and the input image size, each third search space is obtained. Exemplarily, a specific combination of the resolution of each input image and the width multiplier parameter forms a search space. If there are 15 choices for the input parameter and 10 choices for the width parameter, there will be 15 * 10 = 150 search spaces. By obtaining the third search space, it is beneficial for subsequent computational processing of narrowing down the search space.

[0080] Step S220: Based on multiple third search spaces, perform sampling processing to obtain multiple first search spaces.

[0081] It can be understood that, based on multiple third search spaces, random sampling is performed, and the pairs of width parameter and input parameter obtained by sampling are used as the first search space. After multiple samplings, multiple first search spaces are obtained; alternatively, the third search space can be sampled according to a preset sampling rule, and the pairs of width parameter and input parameter obtained by sampling are used as the first search space. After multiple samplings, multiple first search spaces are obtained. Exemplarily, the preset sampling rule is to store the third search space row by row, and sample one third search space every 3 search spaces. After multiple samplings, multiple first search spaces can be obtained. By sampling to obtain the first search space, the range of the search space can be reduced, thereby reducing the computational amount.

[0082] Step S300: For each first search space, according to the width parameter and input parameter corresponding to the first search space, determine the second neural network model corresponding to the first search space.

[0083] It can be understood that, according to each first search space, the width parameter and input parameter corresponding to one of the first search spaces are selected. The first neural network model is initialized according to the width parameter and input parameter, and the first neural network model is trained to obtain the second neural network model, where the second neural network model is a neural network model trained according to the hyperparameters corresponding to the first search space. The above processing is performed for each first search space to obtain the second neural network models corresponding to each first search space. By obtaining the second neural network model, it is beneficial for subsequent computational processing of evaluating the performance of the network model.

[0084] Step S400: Evaluate the performance of the second neural network model corresponding to each first search space to obtain the first performance evaluation value corresponding to each first search space.

[0085] It can be understood that, according to step S300, a second neural network model is obtained, and the performance of the second neural network model is evaluated to obtain a first performance evaluation value corresponding to each first search space. Among them, the first performance evaluation value represents the value of the memory occupancy of the second neural network model. Based on the first performance evaluation value, it is beneficial to perform subsequent optimization processing on the search space.

[0086] Step S500: Determine a second search space from multiple first search spaces according to the first performance evaluation value. The second search space represents a search space that meets the preset memory resource limit condition.

[0087] It can be understood that, according to the preset memory resource limit condition, select the first performance evaluation values less than the preset memory resource limit condition from multiple first performance evaluation values, and select the first search space corresponding to the selected first performance evaluation value from the first search spaces to obtain the second search space. Obtaining the second search space through the memory resource limit can reduce the search range, thereby reducing the resource occupancy of the neural network model and improving the execution efficiency of NAS.

[0088] It should be noted that the preset memory resource limit condition can be the minimum memory occupancy, or a value less than the set memory occupancy, and can be set according to different models, which will not be elaborated here. Exemplarily, select the first search space corresponding to the smallest first performance evaluation value from the first search spaces to obtain the second search space. At this time, the second search space also meets the memory occupancy limit condition.

[0089] Step S600: Obtain a third neural network model according to the second search space.

[0090] It can be understood that, according to the above step S500, the second search space is obtained, and the first neural network model is initialized with the width parameter and input parameter corresponding to the second search space to obtain the third neural network model corresponding to the above hyperparameters. By obtaining the third neural network model, it is beneficial to perform subsequent search optimization processing operations on the third neural network model.

[0091] Step S700: Use an evolutionary algorithm to perform search optimization processing on the third neural network model to determine the optimal structure of the third neural network model.

[0092] Reference Figure 2 , using an evolutionary algorithm to perform search optimization processing on the third neural network model to determine the optimal structure of the third neural network model includes, but is not limited to, steps S710, S720, S730, S740, S750, and S760.

[0093] Step S710: Optimize the third neural network model to obtain a fourth neural network model.

[0094] It can be understood that before using the evolutionary algorithm to search the third neural network model, the third neural network model can be optimized through pruning to obtain a fourth neural network model; or the third neural network model can be optimized through other compression model algorithms to obtain a fourth neural network model, which will not be elaborated here. Obtaining the fourth neural network model is beneficial to accelerating the subsequent speed of search and optimization using the evolutionary algorithm.

[0095] Reference Figure 3 , optimizing the third neural network model to obtain a fourth neural network model includes, but is not limited to, the following steps:

[0096] Step S711: Calculate the Euclidean norm corresponding to each network layer of the third neural network model to obtain the norm values corresponding to each network layer.

[0097] It can be understood that each network layer has multiple channels. Calculate the Euclidean norm between the channels of each network layer to obtain the norm values corresponding to each network layer. There are multiple norm values corresponding to each network layer, and the values are determined by the number of channels in the network layer. Calculating the norm values is beneficial to subsequent pruning of each network layer.

[0098] In some embodiments, calculating the Euclidean norm corresponding to each network layer of the third neural network model to obtain the norm values corresponding to each network layer includes:

[0099] Calculate the Euclidean norm corresponding to each network layer of the third neural network model through the following formula:

[0100]

[0101] where d L2 represents the norm value, C 1i represents a channel of the network layer, C 2i represents another channel of the network layer, and n represents the number of channels in the network layer.

[0102] Step S712: Prune the network layers corresponding to the norm values below the preset threshold to obtain the pruned network layers.

[0103] It can be understood that for each of the obtained norm values according to step S711, these norm values can be sorted either in descending order or in ascending order. If sorted in ascending order, compare the sequence number of the start position corresponding to the sorted norm value with a preset threshold. If it is less than the preset threshold, determine the network layer corresponding to this norm value. In the corresponding network layer, set the channel corresponding to the norm value to 0, that is, perform pruning on this network layer to obtain a pruned network layer. Compare the position sequence numbers corresponding to the norm values in the ascending sequence with the preset threshold one by one until the norm value is greater than or equal to the preset threshold, and then stop the comparison operation. If sorted in descending order, compare the sequence number of the last position corresponding to the sorted norm value with the preset threshold. The comparison method is similar to the ascending process and will not be elaborated here. The preset threshold can be 10, or 15, or other values, which will not be elaborated here. Exemplarily, select the first 10 norm values and the corresponding network layer channels, and set the parameters of the above channels to 0. By obtaining the pruned network layer, it is beneficial for subsequent calculation to obtain the fourth neural network.

[0104] It should be noted that during the forward propagation of the third neural network, perform the above pruning operation and calculate the loss function of the third neural network to obtain the value of the loss function, which is beneficial for subsequent backpropagation calculation of the third neural network.

[0105] Step S713: Update the weights of the pruned network layer and recalculate the Euclidean norm corresponding to the network layer.

[0106] It can be understood that use the obtained value of the loss function to update the weights of the pruned network layer to obtain an updated network layer, and then recalculate the Euclidean norms of each channel of each network layer to obtain norm values, and repeat steps S711 to S713. During the training process of the third neural network model, continuously perform the pruning operation in a loop, which is beneficial for the model to reach a suitable sparse value.

[0107] Step S714: Calculate the sparsity of the third neural network model according to the norm values corresponding to each network layer. When the sparsity reaches the preset sparse value, end the pruning process to obtain the fourth neural network model.

[0108] It can be understood that according to the norm values corresponding to each network layer, obtain the pruned network layer. According to the pruned network layer, calculate the sparsity of the third neural network model. When the sparsity reaches the preset sparse value, end the pruning to obtain the fourth neural network model. Among them, the fourth neural network model is the third neural network model that reaches the preset sparse value. The preset sparse value is to reduce the sparsity of the third neural network to 80%, or it can also be 85%, and can be set according to the situation. By obtaining the fourth neural network model, the search optimization speed of the search space can be accelerated.

[0109] Step S720: Calculate the second performance evaluation value and the third performance evaluation value of the fourth neural network model, where the second performance evaluation value represents the memory resource occupancy, and the third performance evaluation value represents the number of floating-point operations.

[0110] It can be understood that the fourth neural network model is obtained according to step S714, and the memory occupancy and the number of floating-point operations per second of the fourth neural network model are calculated to obtain the second performance evaluation value and the third performance evaluation value. Among them, the second performance evaluation value represents the memory resource occupancy, and the third performance evaluation value represents the number of floating-point operations. Obtaining the second performance evaluation value and the third performance evaluation value through calculation is beneficial to subsequent search and optimization of the neural network.

[0111] Step S730: According to the second performance evaluation value and the third performance evaluation value, select the fourth neural network model that meets the memory resource limit condition and the preset floating-point operation number limit condition as the parent network.

[0112] It can be understood that according to the second performance evaluation value and the third performance evaluation value obtained in step S720, select the fourth neural network model whose second performance evaluation value is less than the memory resource limit condition and whose third performance evaluation value is greater than the preset floating-point operation number limit condition as the parent network. When there is only one fourth neural network model obtained according to the pruning algorithm, directly use the fourth neural network model as the parent network, which is beneficial to subsequent optimization processing of the model based on the parent network.

[0113] Step S740: Use an evolutionary algorithm to perform evolutionary calculation on the parent network to obtain a child network, where the parent network is connected to the child network through weight sharing.

[0114] It can be understood that for the parent network obtained according to step S730, a neural network model with a higher number of floating-point operations per second has higher accuracy, that is, the parent network obtained by selecting a value greater than the floating-point operation number limit condition has higher accuracy. Perform evolutionary mutation calculation on the parent network. Each time a mutation occurs, a child network will appear. Multiple evolutions can obtain multiple different child networks. Among them, the parent network is connected to the child network through weight sharing. The evolutionary algorithm can be an age-based evolutionary algorithm or other evolutionary algorithms, which will not be elaborated here. By using an evolutionary algorithm to perform mutation and evolution processing on the parent network, different widths of the network model can be explored, so as to search for a neural network model with higher accuracy.

[0115] Reference Figure 4 , after using an evolutionary algorithm to perform evolutionary calculation on the parent network to obtain a child network, the method further includes:

[0116] Step S770: For the sub-networks, use a scheduling algorithm to perform priority sorting on the sub-networks to obtain the model priorities.

[0117] It can be understood that for the multiple sub-networks generated according to Step S740, in order to speed up the selection of sub-networks with relatively high accuracy and reduce time, a scheduling algorithm is used to perform priority sorting on the sub-networks. The first-come-first-served algorithm can be adopted, where the sub-networks generated earlier have higher priorities and are thus preferentially selected; or the round-robin method can be used to obtain the model priorities. These model priorities can preferentially select the eligible sub-networks, speed up the search, and reduce time.

[0118] Step S780: According to the model priorities, select the sub-networks that meet the memory resource limit condition and the floating-point operation count limit condition.

[0119] It can be understood that according to the model priorities, calculate the memory resource occupancy and the number of floating-point operations per second of the sub-networks within the priorities, and select the sub-networks that meet the memory resource limit condition and the floating-point operation count limit condition. Networks with excellent performance also produce networks with relatively high accuracy. By selecting sub-networks, the search time can be reduced.

[0120] Step S750: Take the sub-network as the third neural network model and re-optimize the third neural network model.

[0121] It can be understood that the parent network and the generated sub-networks can be stored through a queue. The parent network is stored at the head position of the queue, and the generated sub-networks are sequentially added to the queue according to the model priorities. During the process of search and optimization, the network corresponding to the head position of the queue is deleted, and then the selected sub-network is output from the queue. The sub-network is taken as the third neural network model, and pruning processing is re-performed on the third neural network model. Then, the evolutionary algorithm is used to optimize the pruned network model, and this is executed in a loop. The stack can also be used to store the parent network and the generated sub-networks. According to the principle of the stack that the first-in-last-out, the sub-networks with higher model priorities enter the stack last. The sub-network popped out of the stack is taken as the third neural network, and then this is executed in a loop. Through continuous loop execution, a neural network model with relatively high accuracy can be selected.

[0122] Step S760: When the preset number of iterations is reached, obtain the optimal structure of the third neural network model according to the sub-network corresponding to the maximum floating-point operation count.

[0123] It can be understood that in the storage method of the queue in step S750, by deleting the parent network, when the preset number of iterations is reached, there is only one sub-network in the queue, which is the sub-network corresponding to the maximum floating-point operation count. This sub-network is used as the optimal structure of the third neural network model. Here, the preset number of iterations can be 200 or 220, and it can be set according to the network model, which will not be elaborated here. The optimal structure of the obtained third neural network model has low memory occupancy, moderate model sparsity, high operation count per second, and high accuracy.

[0124] Reference Figure 7 , Exemplarily, the first neural network model is a deep neural network model, which can be VGG19 or resnet20, and can be used for image classification processing. By changing the width parameter of the deep neural network model and the size of the input image to the deep neural network model, a first search space is obtained. Step S801 is executed to randomly sample the first search space, and the sampled depth parameter and input parameter are used to initialize the deep neural network model to obtain a second neural network model. Step S802 is executed to calculate the first performance evaluation value for the second neural network model. According to the first performance evaluation value, step S803 is executed to determine the search space that meets the preset memory resource limit condition as the second search space; according to the width parameter and input parameter corresponding to the second search space, a third neural network model is obtained. By sampling and the preset memory resource limit condition, the search range of the search space is reduced, improving the NAS execution efficiency; and the obtained third neural network model can reduce the occupancy of the neural network model on space, time, and resources.

[0125] Exemplarily, according to the third neural network model, step S804 is executed to perform pruning on the third neural network model, calculate the norm values corresponding to each network layer of the third neural network model, and set the parameters of the channels corresponding to each network layer to 0 to implement the pruning operation, and perform cyclic training until the preset sparsity is reached to obtain a fourth neural network model, which is used as the parent network. By pruning the neural network model, the resource occupancy during model training can be reduced.

[0126] Exemplarily, step S805 is executed to perform evolutionary computation on the parent network using an age-based evolutionary algorithm to obtain each sub-network, calculate the memory occupancy and the number of floating-point operations per time of the sub-network, select the sub-network that meets the resource occupancy limit condition and the maximum number of floating-point operations per time and add it to the queue, delete the oldest network in the queue, and execute step S806 to determine whether the number of loops reaches a preset number of iterations. If the preset number of iterations is not reached, return to execute step S804 and step S805. If the preset number of iterations is reached, end the search to obtain the optimal structure of the third neural network model. The optimal structure of the third neural network model has less resource occupancy and higher classification accuracy.

[0127] Reference Figure 8 , an embodiment of the present application provides a neural network structure search optimization device 100. The device 100 includes a model acquisition module 110 for acquiring a first neural network model; a parameter processing module 120 for obtaining a plurality of different first search spaces by transforming the width parameter and the input parameter of the first neural network model; a first processing module 130 for determining, for each first search space, a second neural network model corresponding to the first search space according to the width parameter and the input parameter corresponding to the first search space; a performance evaluation module 140 for performing performance evaluation on the second neural network model corresponding to each first search space to obtain a first performance evaluation value corresponding to each first search space; a second processing module 150 for determining a second search space from the plurality of first search spaces according to the first performance evaluation value. The second search space represents a search space that meets the preset memory resource limit condition. Obtaining the second search space through the memory resource limit can reduce the search range, thereby reducing the occupancy of the neural network model on space, time, and resources, and improving the execution efficiency of NAS; a third processing module 160 for obtaining a third neural network model according to the second search space; an optimization processing module 170 for performing search optimization processing on the third neural network model using an evolutionary algorithm to determine the optimal structure of the third neural network model. A neural network model with higher accuracy can be obtained through the evolutionary algorithm.

[0128] It should be noted that the model acquisition module 110 is connected to the parameter processing module 120, the parameter processing module 120 is connected to the first processing module 130, the first processing module 130 is connected to the performance evaluation module 140, the performance evaluation module 140 is connected to the second processing module 150, the second processing module 150 is connected to the third processing module 160, and the third processing module 160 is connected to the optimization processing module 170. The above neural network structure search optimization method acts on the neural network structure search optimization device 100. The neural network structure search optimization device 100 can not only reduce the occupancy of the neural network model on space, time, and resources, but also improve the execution efficiency of NAS and obtain a neural network model with higher accuracy.

[0129] It should also be noted that the first processing module 130, the second processing module 150, and the third processing module 160 are all central processing units. A central processing unit generally consists of a logical operation unit, a control unit, and a storage unit. Using the central processing unit for calculation saves a large amount of human resources.

[0130] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0131] Figure 9 FIG. 500 shows a computer device provided by an embodiment of the present application. The computer device 500 may be a server or a terminal. The internal structure of the computer device 500 includes, but is not limited to:

[0132] A memory 510 for storing programs;

[0133] A processor 520 for executing the programs stored in the memory 510. When the processor 520 executes the programs stored in the memory 510, the processor 520 is used to execute the above-mentioned neural network structure search optimization method.

[0134] The processor 520 and the memory 510 may be connected by a bus or other means.

[0135] The memory 510, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs, such as the neural network structure search optimization method described in any embodiment of the present invention. The processor 520 realizes the above-mentioned neural network structure search optimization method by running the non-transitory software programs and instructions stored in the memory 510.

[0136] The memory 510 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store the execution of the above-mentioned neural network structure search optimization method. In addition, the memory 510 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 510 may optionally include a memory remotely provided with respect to the processor 520, and these remote memories may be connected to the processor 520 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof.

[0137] The non-transitory software program and instructions required to implement the above neural network architecture search optimization method are stored in the memory 510. When executed by one or more processors 520, the neural network architecture search optimization method provided by any embodiment of the present invention is executed.

[0138] An embodiment of the present application also provides a computer-readable storage medium storing computer-executable instructions for executing the above neural network architecture search optimization method.

[0139] In one embodiment, the storage medium stores computer-executable instructions that are executed by one or more control processors 520. For example, when executed by a processor 520 in the above computer device 500, the one or more processors 520 can be caused to execute the neural network architecture search optimization method provided by any embodiment of the present invention.

[0140] The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0141] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0142] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (items) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0143] Those of ordinary skill in the art can understand that all or some of the steps and systems disclosed in the above methods can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0144] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of this application. These equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.

Claims

1. A method for optimizing neural network architecture search, characterized in that, the method includes: Obtain a first neural network model; wherein, the input of the first neural network model is images of different resolutions; By changing the width parameter and input parameter of the first neural network model, obtain a plurality of different first search spaces; wherein, the input parameter represents a parameter of the image size input into the first neural network model; For each of the first search spaces, determine a second neural network model corresponding to the first search space according to the width parameter and the input parameter corresponding to the first search space; Perform performance evaluation on the second neural network model corresponding to each first search space to obtain a first performance evaluation value corresponding to each first search space; Determine a second search space from the plurality of first search spaces according to the first performance evaluation value, and the second search space represents a search space that meets a preset memory resource limit condition; Obtain a third neural network model according to the second search space; Perform optimization processing on the third neural network model to obtain a fourth neural network model; Calculate the second performance evaluation value and the third performance evaluation value of the fourth neural network model, wherein the second performance evaluation value represents the memory resource occupancy, and the third performance evaluation value represents the floating-point operation count; According to the second performance evaluation value and the third performance evaluation value, select the fourth neural network model that meets the memory resource limit condition and the preset floating-point operation count limit condition as the parent network; Use an evolutionary algorithm to perform evolutionary calculation on the parent network to obtain a subnet, wherein the parent network is connected to the subnet through weight sharing; Use the subnet as the third neural network model, and re-perform optimization processing on the third neural network model; When the preset number of iterations is reached, obtain the optimal structure of the third neural network model according to the subnet corresponding to the maximum floating-point operation count.

2. The method according to claim 1, characterized in that, The performing optimization processing on the third neural network model to obtain a fourth neural network model includes: Calculate the Euclidean norm corresponding to each network layer of the third neural network model to obtain the norm value corresponding to each network layer; Perform pruning processing on the network layer corresponding to the norm value lower than the preset threshold to obtain the pruned network layer; Update the weights of the pruned network layer, and re-calculate the Euclidean norm corresponding to the network layer; According to the norm value corresponding to each network layer, calculate the sparsity of the third neural network model, and when the sparsity reaches the preset sparse value, end the pruning processing to obtain the fourth neural network model.

3. The method according to claim 1, characterized in that, After using the evolutionary algorithm to perform evolutionary calculation on the parent network to obtain a subnet, the method further includes: For the subnet, use a scheduling algorithm to perform priority sorting on the subnet to obtain a model priority; Select the sub-network that meets the memory resource limit condition and the floating-point operation count limit condition according to the model priority.

4. The method according to claim 1, wherein, the obtaining of multiple different first search spaces by transforming the width parameter and the input parameter of the first neural network model includes: obtaining multiple different third search spaces by transforming the width parameter and the input parameter of the first neural network model; performing sampling processing based on the multiple third search spaces to obtain the multiple first search spaces.

5. The method according to claim 4, wherein, the obtaining of multiple different third search spaces by transforming the width parameter and the input parameter of the first neural network model includes: transforming the width parameter of the first neural network model by using a width multiplier parameter, where the value range of the width multiplier parameter is between 0 and 1; transforming the input parameter of the first neural network model by changing the resolution of the image; obtaining each of the third search spaces according to the width parameter and the input parameter.

6. The method according to claim 2, wherein, the calculating of the Euclidean norm corresponding to each network layer of the third neural network model to obtain the norm value corresponding to each network layer includes: calculating the Euclidean norm corresponding to each network layer of the third neural network model by the following formula: Among them, represents the range value, represents a channel of the network layer, represents another channel of the network layer, represents the number of channels of the network layer.

7. A neural network structure search optimization device, wherein, it includes: a model acquisition module, configured to acquire a first neural network model; wherein, the input of the first neural network model is images with different resolutions; a parameter processing module, configured to obtain multiple different first search spaces by transforming the width parameter and the input parameter of the first neural network model; wherein, the input parameter represents a parameter of the size of the image input into the first neural network model; a first processing module, configured to determine, for each of the first search spaces, a second neural network model corresponding to the first search space according to the width parameter and the input parameter corresponding to the first search space; a performance evaluation module, configured to perform performance evaluation on the second neural network model corresponding to each first search space to obtain a first performance evaluation value corresponding to each first search space; a second processing module, configured to determine a second search space from the multiple first search spaces according to the first performance evaluation value, where the second search space represents a search space that meets a preset memory resource limit condition; a third processing module, configured to obtain a third neural network model according to the second search space; an optimization processing module, configured to: perform optimization processing on the third neural network model to obtain a fourth neural network model; calculate a second performance evaluation value and a third performance evaluation value of the fourth neural network model, where the second performance evaluation value represents the memory resource occupancy, and the third performance evaluation value represents the floating-point operation count; Select the fourth neural network model that meets the memory resource limit condition and the preset floating-point operation count limit condition as the parent network according to the second performance evaluation value and the third performance evaluation value; Use an evolutionary algorithm to perform evolutionary computation on the parent network to obtain a child network, where the parent network is connected to the child network through weight sharing; Use the child network as the third neural network model and re-optimize the third neural network model; When the preset number of iterations is reached, obtain the optimal structure of the third neural network model according to the child network corresponding to the maximum floating-point operation count.

8. A computer device, characterized in that, the computer device includes a memory and a processor, and computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by one or more of the processors, one or more of the processors execute the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, the storage medium can be read and written by a processor, and the storage medium stores computer instructions. When the computer-readable instructions are executed by one or more processors, one or more of the processors execute the steps of the method according to any one of claims 1 to 6.