A low-cost automatic search method for neural network architectures for image classification

By designing network blocks based on packet convolution and an improved genetic algorithm, combined with the NTK condition number as fitness, the problem of high search cost for neural network structures in image classification tasks is solved, and fast and low-cost high-precision network structure search is achieved.

CN114299344BActive Publication Date: 2025-06-06JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111669013.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-06-06
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The existing neural network structure search method is costly in image classification tasks and requires a lot of computing resources and time.

Method used

A network block based on packet convolution is designed as the basic unit, combining improved genetic algorithms and three-stage natural selection strategy, using the NTK condition number as individual fitness to quickly search for high-precision and low-parameter network structures.

Benefits of technology

It realizes the search for network structures with superior comprehensive performance in a short time with few computing resources and is suitable for image classification tasks, and the searched network structures perform excellently in classification accuracy and parameter quantity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299344B_ABST
    Figure CN114299344B_ABST
Patent Text Reader

Abstract

The present invention discloses a low-cost automatic search method for a neural network structure for image classification, and belongs to the field of image classification technology. The method designs a network block based on group convolution, and uses the block as a basic unit to construct an extensible network structure. The controllable parameter setting of the block makes the search space of the constructed network structure extensible. Combined with an improved genetic algorithm, a three-stage natural selection strategy is used to better stimulate the exploratory and developmental nature of the search space. At the same time, the conditional number of the non-training indicator NTK is introduced as the individual fitness, and a high-precision and low-parameter network structure is searched at an extremely fast speed, so that when solving practical problems, it is possible to use less computing resources to quickly search for a network structure with superior comprehensive performance. For image classification tasks, experiments have proved that the classification accuracy of the network structure searched by this method is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a low-cost automatic search method for a neural network structure for image classification, and belongs to the technical field of image classification. Background Art

[0002] Deep learning has made great progress in various computer vision tasks. Among them, hand-designed neural network structures are one of the important driving forces in the development of deep learning, such as VGGNet, ResNet, Inception, and DenseNet. Although hand-designed neural network structures can achieve excellent classification performance, the design of the structure requires professional domain knowledge, which only a few experts have. At the same time, since the manual design method requires repeated optimization experiments, it will consume a lot of time and computing resources. This has also prompted a lot of research in the field of neural network structure search (NAS) in recent years to develop automatic design of neural network structures.

[0003] NAS algorithm can automatically design the network structure, which can be used by individuals who are not familiar with professional knowledge, greatly lowering the threshold of network design. The automation of NAS algorithm can reduce manpower and cost, and the network structure searched by NAS algorithm can outperform the manually designed algorithm. However, the search time and computing resource cost of NAS algorithm to find the best network structure are usually very expensive. Most existing NAS algorithms mainly rely on validation data sets to optimize the network structure, which requires a lot of time and intensive computing resources. For example, NASNet uses 500 GPUs and takes 4 days to search for the best network.

[0004] The network structure search problem is usually defined as a single-objective optimization problem, that is, only a single objective is considered at the same time instead of multiple objectives. Most real-world network deployments require not only extremely high classification performance, but also lower computing resources, such as fewer network parameters and less network computational complexity. For this reason, some manually designed network structures have been developed in recent years. While reducing computational consumption, the network can still have high-precision performance, such as MobileNet and MobileNetV2. At the same time, in recent years, some NAS algorithms based on multi-objective optimization have also emerged to make network structures easier to calculate and deploy. For example, NSGA-Net considers the trade-off between the classification accuracy and computational complexity of the network. LEMONADE considers both the classification performance of the network and the number of network parameters.

[0005] However, these methods still require a lot of computing resources and a long search time, but many computer vision tasks have time requirements, such as image classification tasks in many scenarios have real-time requirements. Therefore, how to use less computing resources to quickly search for a network structure with superior comprehensive performance to apply to practical problems in the real world still needs further research. Summary of the invention

[0006] In order to solve the problem of high cost of the current automatic search method for neural network structure in image classification technology, the present invention provides a low-cost automatic search method for neural network structure for image classification, the method comprising:

[0007] Step 1: For the image classification task, determine the main framework of the neural network structure, randomly generate X network structures as the population P, and each individual in the population represents a randomly generated network structure; the main framework of the neural network structure includes a standard convolution layer, unit num Reg Unit modules and a global average pooling layer, each Reg Unit module includes block num group convolution Reg Block; and each Reg Unit module contains a SENet module with a probability of 50%, and the SENet module simulates the attention mechanism through Squeeze-and-Excitation;

[0008] The number of Reg Unit modules unit num, the number of group convolution Reg Block block num, and the width width of the second convolution layer in each branch of the group convolution Reg Block are randomly generated;

[0009] Step 2: Set the three-stage separation point S for the subsequent population evolution stage 1 , S 2 and the maximum number of generations of evolution Max_gen;

[0010] Step 3: Calculate the NTK condition number K of the network structure of each individual in the population P N fitness as an individual;

[0011] Step 4: The population enters the evolution stage, and uses the tournament selection to select individual mutation operations to generate new network structure individuals. Different indicators are selected for environmental selection to eliminate individuals according to the stage of the current evolutionary algebra G.

[0012] Step 5: After reaching the maximum number of evolutionary generations Max_gen, select the fitness K of the individual N The network structure with the smallest value is taken as the searched neural network structure for the image classification task.

[0013] Optionally, each group convolution Reg Block in each network structure contains group branches, and each branch consists of three convolutional layers and one pooling layer, where the pooling layer is in the third layer; the first and fourth convolutional layers use 1×1 kernels to adjust the number of feature maps, the second convolutional layer uses a 3×3 kernel to extract feature maps, and all convolutional layers follow the order of convolution operation, ReLu activation function, and batch normalization layer; the pooling layer in the third layer is used to halve the size of the input data; the input data is image data.

[0014] Optionally, for M×M input data, the number of pooling layers in the third layer of each branch of the group convolution Reg Block cannot be greater than

[0015] Optionally, in step four, different metrics are selected according to the stage to which the current evolutionary generation G belongs for environmental selection to eliminate individuals, including:

[0016] In the first and third stages, that is, when 0 < G ≤ S 1 and S 2 < G ≤ Max_gen, the fitness K N of the individual is selected as the criterion to eliminate individuals;

[0017] In the second stage, that is, when S 1 < G ≤ S 2 , the lifespan of the individual is selected as the criterion to eliminate individuals, and the lifespan of the individual is the number of evolutionary generations experienced by the individual.

[0018] Optionally, the population evolution process includes:

[0019] Randomly select k individuals from the population; from these k individuals, according to the fitness K N value of each individual, select the top t individuals with the best fitness as the parent individuals;

[0020] The t parent individuals generate t offspring individuals through a set of mutation operators; after the offspring individuals are generated, they are evaluated and added to the existing population;

[0021] According to the stage to which the current evolutionary generation belongs, use the corresponding criterion in environmental selection to eliminate individuals; eliminate the t worst individuals according to the current criterion, so that the population size remains unchanged, and the remaining individuals construct a new population and enter the next generation of evolution.

[0022] Optionally, the t parent individuals generate t offspring individuals through a set of mutation operators; after the offspring individuals are generated, they are evaluated and added to the existing population, including:

[0023] Randomly select a mutation position pos within the length of the parent individual ij, which represents the position of the jth Reg Block in the i-th Reg Unit, and the position is determined by the order of the Reg Unit in the network structure and the position order of the Reg Block in the Reg Unit;

[0024] Randomly select a mutation operator to perform mutation of the parent individual, wherein the mutation operator includes an add operator, a remove operator, and a change operator;

[0025] Add operator: at mutation position pos ij Add a Reg Block with random parameter settings;

[0026] Remove operator: remove at mutation position pos ij Reg Block on;

[0027] Change operator: randomly change the mutation position pos ij Parameters of the Reg Block on.

[0028] Optionally, when implementing the add operator, if the length of the parent individual reaches the upper limit, the add operator cannot be implemented, and the only options are to remove the operator or change the operator;

[0029] When implementing the removal operator, when the length of the parent individual reaches the lower limit, the removal operator cannot be performed, and the only options are to add an operator or change the operator.

[0030] The present application also provides an image classification method, which uses the neural network structure searched out by the above method to perform image classification.

[0031] Optionally, the method comprises:

[0032] The image to be classified is input into the neural network structure, and the features of the image to be classified are extracted through the standard convolution layer;

[0033] Further feature extraction is performed through unit num Reg Unit modules, where the output of each group convolution Reg Block in each Reg Unit module is connected by the output features of each branch and the residual connection, and then the feature map is obtained through the SENet module with a probability of 50%, and then the feature map output by the Reg Units is flattened into a feature vector through the global average pooling layer. Finally, a fully connected layer with a softmax layer is set as a classifier to convert the feature vector into the final classification result.

[0034] The beneficial effects of the present invention are:

[0035] By designing a network block based on group convolution, a scalable network structure is constructed with this block as the basic unit. The controllable parameter setting of the block makes the search space of the constructed network structure scalable. Combined with an improved genetic algorithm, a three-stage natural selection strategy is used to better stimulate the exploration and development of the search space. At the same time, the conditional number of the non-training indicator NTK is introduced as the individual fitness to search for a high-precision and low-parameter network structure at an extremely fast speed, so that when solving practical problems, it is possible to use less computing resources to quickly search for a network structure with superior comprehensive performance. For image classification tasks, experiments have proved that the classification accuracy of the searched network structure with superior comprehensive performance is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0037] Figure 1 It is a schematic diagram of the structure of the overall network designed in the low-cost neural network structure search method based on the three-stage evolutionary algorithm disclosed in one embodiment of the present invention and the proposed new network block Reg Block.

[0038] Figure 2 It is a schematic diagram of the selection values ​​of the parameters of the network structure for the image classification problem searched by the low-cost neural network structure search method based on the three-stage evolutionary algorithm disclosed in one embodiment of the present invention.

[0039] Figure 3 It is a schematic diagram of a flexible encoding strategy disclosed in one embodiment of the present invention.

[0040] Figure 4 It is a comparison diagram of the parameters of the group convolution proposed in this application and the standard convolution in the prior art disclosed in one embodiment of the present invention.

[0041] Figure 5A It is a test accuracy comparison chart between the original network structure disclosed in one embodiment of the present invention and the network architecture without the SENet module.

[0042] Figure 5B It is a comparison diagram of the parameter quantities between the original network structure disclosed in one embodiment of the present invention and the network architecture without the SENet module.

[0043] Figure 6is the K in the LoNAS search space on the CIFAR-10 dataset disclosed in one embodiment of the present invention. N Schematic diagram of the negative correlation with network structure test accuracy.

[0044] Figure 7 This is a schematic diagram of the effect of the length of the second stage on the test accuracy under the premise of the same evolutionary length (evolutionary generations are set to 50).

[0045] Figure 8 It is a schematic diagram of adding operators and removing operators in the evolution process disclosed in one embodiment of the present invention. DETAILED DESCRIPTION

[0046] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0047] Embodiment 1:

[0048] This embodiment provides a low-cost neural network structure search method based on a three-stage evolutionary algorithm, the method comprising:

[0049] Step 1. Given a specific set of parameters for Reg Block, the network structure is flexibly encoded; at the same time, the separation point S of the three stages is given 1 , S 2 and the maximum number of evolution generations Max_gen; the Reg Block includes a group convolution and a SENet module, wherein the probability of including the SENet module is 50%;

[0050] The Reg Block contains group branches, each branch consists of three convolutional layers and one pooling layer, where the pooling layer is in the third layer; the first and fourth convolutional layers use 1×1 kernels to adjust the number of feature maps, and the second convolutional layer uses 3×3 kernels to extract feature maps. All convolutional layers follow the order of convolution operation, ReLu activation function, and batch normalization layer; the third pooling layer is used to halve the size of the input data.

[0051] The output of Reg Block is composed of the output features of each branch and the residual connection, plus a SENet module with a probability of 50%; the SENet module simulates the attention mechanism through Squeeze-and-Excitation.

[0052] Step 2. Initialize a population P containing 50 network structure individuals according to the encoding method in step 1;

[0053] The main body of the network structure of each individual includes a standard convolutional layer Conv Unit, unit num Reg Units, and a global average pooling layer, as shown in Figure 1 (a). Each Reg Block structure in the Reg Units is as shown in Figure 1 (b).

[0054] Step 3. Use the CIFAR-10 and CIFAR-100 datasets to calculate the condition number K of the NTK of each network structure N as the fitness of the individual;

[0055] Step 4. The population enters evolution;

[0056] Step 5. Use tournament selection to select individual mutation operations to generate new network structure individuals;

[0057] Step 6. Select different metrics for environmental selection to eliminate individuals according to the current generation number G of evolution;

[0058] Specifically:

[0059] When 0 < G ≤ S 1 , select the fitness K of the individual N as the criterion to eliminate individuals;

[0060] When S 2 < G ≤ Max_gen, select the lifespan of the individual as the criterion to eliminate individuals, and the lifespan of the individual is the number of generations of evolution experienced by the individual;

[0061] Step 7. Return to Step 5 until the maximum number of evolution generations is reached.

[0062] Experiments on the image classification datasets CIFAR-10 and CIFAR-100 can prove that the present invention can search for a network structure that takes into account both classification accuracy and the number of parameters in a very short search time while consuming extremely little computing resources.

[0063] Example 2

[0064] This embodiment provides a low-cost neural network structure search method based on a three-stage evolutionary algorithm. Taking the search for a low-cost neural network structure for image classification tasks as an example, the method includes:

[0065] Step 1. Given a specific parameter set regarding the Reg Block, flexibly encode the network structure; at the same time, given the separation point S of the three stages 1 , S 2and the maximum number of evolution generations Max_gen; the Reg Block includes a group convolution and a SENet module, wherein the probability of including the SENet module is 50%;

[0066] The Reg Block contains group branches, each branch consists of three convolutional layers and one pooling layer, where the pooling layer is in the third layer; the first and fourth convolutional layers use 1×1 kernels to adjust the number of feature maps, and the second convolutional layer uses 3×3 kernels to extract feature maps. All convolutional layers follow the order of convolution operation, ReLu activation function, and batch normalization layer; the third pooling layer is used to halve the size of the input data.

[0067] The output of Reg Block is composed of the output features of each branch and the residual connection, plus a SENet module with a probability of 50%; the SENet module simulates the attention mechanism through Squeeze-and-Excitation.

[0068] Traditional standard convolution can achieve good classification performance, but it also requires more parameters, which is not conducive to designing a high-precision network structure with fewer parameters. Therefore, this application designs a new network block called Reg Block based on ResNet Block. Reg Block consists of group convolution and SENet modules, which can be used to reduce the number of parameters and improve classification performance respectively.

[0069] The topology of Reg Block is as follows Figure 1 (b) In Reg Block, the input features are divided into a certain number of groups, which allows the standard convolution operation to be decomposed into multiple independent convolution branches.

[0070] Compared with the standard convolution operation, the advantage of group convolution is that it can greatly reduce the amount of network computation and the number of parameters without significantly reducing the classification performance. The third layer of the pooling layer in the Reg Block is used to halve the size of the input data. The number of pooling layers cannot be arbitrarily specified and needs to follow the computational constraints. For example, for an M×M input data, the number of pooling layers used to halve the input feature size cannot be greater than Otherwise, the size of the input data will be reduced to less than 1, resulting in errors. Therefore, in the Reg Block, only a part of the pooling layer stride can be set to 2 to halve the number of feature maps, and the other part of the stride is set to 1.

[0071] The output of Reg Block is composed of the output features of each branch and the residual connection, plus a SENet module. The SENet module simulates the attention mechanism through Squeeze-and-Excitation, which can make the network structure pay more attention to the part with the most information in the feature, thereby improving the representation ability of the network structure.

[0072] For the effectiveness of the Reg Block designed by this application, including group convolution and SENet modules, this application conducted two ablation experiments on CIFAR-10. The first one was to verify the effectiveness of group convolution, and the second one was to investigate the effectiveness of SENet modules. The experimental results are shown in Figure 4 As shown in Figure 2, 10 individuals were randomly selected from a final population for the two ablation experiments. These individuals all contained group convolutions and a certain number of SENet modules.

[0073] In the first ablation experiment, the effect of group convolution on the number of parameters of the network structure was verified. First, the number of parameters of each individual was recorded. Then, while keeping other topological structures unchanged, the group convolution of each individual was converted into a standard convolution and the corresponding number of parameters was recorded. The comparison results are shown in Figure 2. Figure 4 As shown in Figure 4, black represents group convolution and gray represents standard convolution. It can be clearly seen from Figure 4 that group convolution has much fewer parameters than standard convolution, and each individual containing group convolution can reduce the number of parameters by about half. Therefore, group convolution can effectively reduce the number of parameters in the network structure.

[0074] In the second ablation experiment, the effectiveness of the SENet module on the network test accuracy and number of parameters was verified. For each individual, 10 independent experiments were performed to obtain the test accuracy and number of parameters of the individual and the individual without all SENet modules. The comparison results of the test accuracy and number of parameters are shown in Figure 2. Figure 5A and Figure 5B As shown in Figure 2, the dotted line and black bar represent the original network structure, and the solid line and gray bar represent the network structure with all SENet modules removed. Figure 5A It is clearly shown that compared with the original network structure, the accuracy performance of the network structure after removing the SENet module is greatly reduced, indicating that the SENet module can improve the test accuracy of the network structure. Figure 5B It shows that compared with the overall number of parameters in the network structure, the addition of the SENet module only brings a small increase in the number of parameters, and has little effect on the number of network parameters. These results show that the SENet module can significantly improve the classification performance of the network structure with only a small increase in the number of parameters.

[0075] Step 2. Initialize a population P containing 50 network structure individuals according to the encoding method in step 1;

[0076] like Figure 1 As shown in (a), the main body of the network structure of each of the 50 network structures includes a standard convolution layer ConvUnit, unit num Reg Units and a global average pooling layer.

[0077] Among them, the standard convolution layer Conv Unit uses a 3×3 kernel to extract the features of the initial input data. When used for image classification tasks, the initial input data is the image to be classified.

[0078] The number of Reg Units unit num is randomly generated; each Reg Unit consists of block num Reg Blocks. Reg Block is randomly generated based on a set of parameters that can be automatically searched, that is, the number of Reg Blocks block num is randomly generated. The number of Reg Blocks in each Reg Unit is also randomly generated, the number of branches group in each RegBlock is randomly generated, and the width width of the second convolutional layer in each branch is randomly generated.

[0079] Thus, a population P initialized by random individuals is obtained, which contains 50 individuals. Each individual represents a randomly generated network structure. The main body of the network structure of all individuals contains a standard convolutional layer Conv Unit, unit num Reg Units and a global average pooling layer.

[0080] A global average pooling layer is placed at the end of each individual network structure to flatten the feature map output by Reg Units into a feature vector. Finally, a fully connected layer with a softmax layer is set as a classifier to convert the feature vector into the final prediction result.

[0081] Step 3. Calculate the condition number K of NTK for each network structure using the CIFAR-10 and CIFAR-100 datasets N fitness as an individual;

[0082] In order to speed up the search process, the present invention introduces NTK to characterize the trainability of the network structure. Higher trainability represents higher classification accuracy performance of the network architecture. NTK can be used to characterize the gradient descent training dynamics of infinite width or finite width deep network architectures. Referring to W.Chen, X.Gong, and Z.Wang, "Neural architecture search onimagenet in four gpu hours: A theoretically inspired perspective," inInternational Conference on Learning Representations, 2020, the condition number K of NTK for each network structure is calculated using the CIFAR-10 and CIFAR-100 datasets. N ;

[0083] Specifically, the eigenvalue λ of NTK between training sets is obtained according to each set of training images and corresponding labels in the CIFAR-10 and CIFAR-100 datasets. k , according to each eigenvalue λ k Get the NTK condition number K of the network structure N , the calculation formula is as follows:

[0084]

[0085] Among them, λ 0 Denotes the eigenvalue λ k The maximum value of λ m Denotes the eigenvalue λ k The minimum value of .

[0086] This application randomly generates 200 network structure individuals and tests their K N The correlation between the accuracy of the network structure test is as follows: Figure 6 As shown. Figure 6 It can be seen that K N It is negatively correlated with the accuracy performance of the network structure.

[0087] Therefore, this application uses K N To evaluate the fitness of individuals. In the evolutionary process, minimize K N It helps to find a network structure with high accuracy performance. N Non-trainable features can directly save a lot of search time and computing resources.

[0088] Calculate the K of each initial individual N value.

[0089] Step 4. The population enters the evolution process. The tournament selection is used to select individuals for mutation operation to generate new individual network structures, and different metrics are selected according to the current generation number G of evolution for environmental selection to eliminate individuals.

[0090] During the evolution process, first, k individuals are randomly selected from the population. From these k individuals, according to the fitness value K N of each individual, the top t individuals with the best fitness are selected as the parent individuals.

[0091] Then, these t parent individuals generate t offspring individuals through a set of mutation operators. After the offspring individuals are generated, they are evaluated and added to the existing population.

[0092] Then, according to the stage to which the current generation number of evolution belongs, corresponding criteria are used in environmental selection to eliminate individuals. According to the current criteria, t individuals with the worst fitness are eliminated, so that the size of the population remains unchanged. The remaining individuals construct a new population and enter the next generation of evolution.

[0093] Specifically:

[0094] In the first stage (0 < G ≤ G 1 ) and the third stage (G 2 < G ≤ Max_gen), the criteria for environmental selection are both based on K N . This helps to retain potential optimal solutions and improve the exploitative ability of the algorithm respectively. In the second stage (G 1 < G ≤ G 2 ), the lifespan of the individual is used as the criterion for environmental selection, ensuring sufficient explorative ability.

[0095] That is:

[0096] When 0 < G ≤ S 1 , the fitness K N of the individual is selected as the criterion to eliminate individuals;

[0097] When S 2 < G ≤ Max_gen, the lifespan of the individual is selected as the criterion to eliminate individuals, and the lifespan of the individual is the number of generations of evolution experienced by the individual;

[0098] Step 5. Return to Step 4 until the maximum number of generations of evolution is reached, and select the individual with the smallest K N as the best network structure found.

[0099] In the traditional evolutionary algorithm, fixed standards are usually used for environmental selection throughout the evolution process. Most of the selected standards can directly reflect the performance of the network structure, such as the test accuracy and number of parameters of the network. Using this method, when the population enters the evolutionary process, individuals with better fitness can be preserved in the population through environmental selection. However, in the subsequent evolutionary process, mutations will occur between these individuals, which will cause most of the offspring to be inherited by these individuals during the evolutionary process. Over time, the algorithm will only focus on these few excellent individuals, which can easily lead to falling into the local optimum, and the algorithm's exploration ability will be greatly reduced.

[0100] Therefore, (E. Real, A. Aggarwal, Y. Huang, and QV Le, "Regularized evolution for image classifier architecture search," in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, 2019, pp. 4780–4789.) proposed an evolutionary algorithm based on individual lifespan to solve this problem. It uses the lifespan of individuals in the population as the criterion for environmental selection. During the evolutionary process, each environmental selection will discard the oldest individual, thereby eliminating individuals with better fitness and longer survival time in the population, increasing the probability of other individuals entering the subsequent evolutionary process, allowing the algorithm to search more space.

[0101] However, the inventors have found through in-depth research that this type of evolution based on individual life span may have the problem of unstable convergence. In the early stages of evolution, the life spans of different individuals in the population are relatively similar. If there are many individuals with good fitness in the population at the beginning, then as the life span of the individuals increases, these individuals will be eliminated one after another in the later evolution process. These individuals are removed as potential optimal solutions in the search space, which will slow down the convergence speed of the population, thereby affecting the effect of population convergence.

[0102] Therefore, the present invention comprehensively considers traditional evolution and evolution based on individual life span, and proposes a new evolutionary algorithm with multi-standard environmental selection. In the first and third stages of evolution, K related to the classification performance of the network structure is selected. N As the criterion for environmental selection, each time the N In the second stage, based on the life span of the individual, individuals with shorter life span are selected and preserved in the population.

[0103] In the first stage, it is ensured that the excellent individuals in the population can enter the later evolutionary process, so that the offspring produced by mutation can inherit from them, improve the overall performance of the population, and ensure that there are enough potential optimal solutions in the population. Then in the second stage, the population is frequently updated to explore more search spaces and increase the diversity of individuals. Finally, in the third stage, excellent individuals are saved in each environmental selection to guide the population to converge to the best optimal solution, which helps to ensure the development of the algorithm.

[0104] In order to verify the effectiveness of the three-stage evolution adopted in this application, this embodiment conducts five independent experiments with different second-stage lengths. The maximum number of generations of each experimental population evolution is the same, and the classification performance of the final population is recorded. By changing the length of the second stage, the lengths of the first and third stages also change accordingly, which helps to study the impact of different lengths of each stage on the final population verification accuracy. The length of the second stage changes from [0-30], Figure 7 The overall accuracy performance of different populations is shown. Figure 7 In the figure, each rectangular box represents the overall verification accuracy of a population, the length of the box represents the deviation of the accuracy between individuals, and the points and dashed lines in the box represent the mean and median of the accuracy. The extension lines at both ends of the box represent the maximum and minimum accuracy in the population. When the length of the second stage is set to 0, the evolutionary algorithm degenerates into a traditional evolutionary algorithm containing fixed standard environment selection. Figure 7 It can be clearly seen that the traditional evolutionary algorithm has the lowest average verification accuracy compared with other three-stage evolutionary algorithms. This shows that the second stage helps to explore more search space and helps the population converge to a network structure with better classification performance. When the length of the second stage increases, the average accuracy of the population shows a trend of increasing first and then decreasing. This can be explained by the fact that the longer second stage causes the population to spend too much time exploring the search space during the entire evolution process, resulting in the population being unable to converge to a better solution in time. At the same time, the length of the third rectangular box and its extension is the shortest, indicating that the difference between individuals is the smallest. This can prove that a third stage with sufficient length can improve exploration, which helps to eliminate individuals with poor fitness and increase the number of optimal solutions. This in turn improves the stability of the evolutionary algorithm during the search process. Therefore, according to the above experimental results, the appropriate length of each stage helps to effectively balance the exploration and development of the algorithm, thereby better searching for the optimal solution.

[0105] During the evolution process, the offspring individuals in the population are generated by the mutation of existing individuals to explore more search space and increase the diversity of individuals. In this application, the mutation operator is only performed in the Reg Unit, and the Conv Unit does not involve mutation due to its specific function. For the mutation operator, first randomly select a mutation position pos within the length of the parent individual ij, which represents the position of the jth Reg Block in the i-th Reg Unit, and the position is determined by the order of the Reg Unit in the network structure and the position order of the Reg Block in the Reg Unit. Then, a mutation operator is randomly selected to perform the mutation of the parent individual. According to the block-based network structure, the designed mutation operator is as follows:

[0106] Add (add a Reg Block with random parameter settings);

[0107] Remove (remove the Reg Block at the selected position);

[0108] Change (randomly change the parameters of the Reg Block at the selected position). More specifically, in the Add operator, a Reg Block with random parameters is generated and inserted at position pos ij After that. In the removal operator, the position pos ij The Reg Block on the .

[0109] In the change operator, a new set of parameters is randomly generated to replace the position pos ij The old parameters of the Reg Block. Figure 8 As shown in the figure, examples of adding and removing operators are shown to better understand the mutation operator. Figure 8 In (a), a new Reg Block is randomly generated and inserted after Reg Block 11. Figure 8 In (b), Reg Block 23 is removed from Reg Unit 2.

[0110] It should be noted that the length of the original parent individual needs to be considered when implementing the add operator and the remove operator. If the length reaches the upper limit, the add operator cannot be implemented and only the other two operators can be selected. When the length of the original individual reaches the lower limit, the remove operator cannot be operated either.

[0111] This application designs a new network block called Reg Block, which combines group convolution and SENet modules to reduce the number of network parameters and improve network classification performance. Based on Reg Block, a flexible encoding strategy is proposed to construct the network structure. By designing network structure constraints, a limited search space can be constructed to find a network structure that takes into account both network classification accuracy and the number of parameters.

[0112] Beneficial effects of this application:

[0113] This application evaluates the fitness of each network structure by analyzing the Neural Tangent Kernel (NTK). NTK can effectively characterize the trainability of the network structure. The number of NTK (K N ) is strongly correlated with the classification accuracy of the network structure. Since the indicator (K N ), which can greatly reduce the search time and save a lot of computing resources.

[0114] This application proposes a three-stage evolutionary algorithm based on multi-criteria environment selection. The criterion for environment selection is based on the number of NTK (K N ) and the life span of the individual. The life span attribute is associated with each individual and represents the number of evolutionary generations that the individual has experienced. In the early stages of the evolutionary process, according to K N By preserving individuals with high fitness to the next generation, a population containing many individuals with high fitness can be formed. In the second stage, older individuals are eliminated according to their life span, so that the population can maintain diversity and avoid premature convergence to the local optimal solution. N The best individuals are retained as the standard to ensure the convergence of the population. The three-stage evolutionary algorithm can well balance the exploration and development in the search process. In addition, this method also designs a simple mutation operator based on a set of Reg Blocks to maintain the evolution of the population.

[0115] In order to verify that the search method provided by this application can search for high-precision, low-parameter network structures in a short time and only requires a small amount of computing resources, the following experiment is conducted by comparing the network structure searched by the method of this application with the existing manually designed network structure, semi-automatic search + manual fine-tuning, and fully automatic search network structure:

[0116] Experiments were conducted on CIFAR-10 and CIFAR-100, and the current mainstream algorithms were compared. The results are shown in Table 1. In Table 1:

[0117] The column below CIFAR-10 and CIFAR-100 represents the corresponding accuracy of the network structure obtained by each method when performing image classification. The higher the accuracy, the better the classification effect.

[0118] Parameters represents the number of parameters of the designed network structure. The fewer the number of parameters, the better the network structure.

[0119] GPU Days represents the search time used by the method. 1 GPU Day means that it takes one day to run on a 1080Ti graphics card. The smaller the value, the less time it takes. GPUs represents the number of graphics cards required. The smaller the value, the less graphics card resources are required. Table 1 shows the comparison results. The results of these algorithms are extracted from the data in their respective seminal papers.

[0120] It should be noted that the CIFAR-10 and CIFAR-100 datasets are public datasets, where the CIFAR-10 dataset consists of 60,000 32x32 color images of 10 classes, with 6,000 images for each class. There are 50,000 training images and 10,000 test images. The dataset is divided into five training batches and one test batch, each with 10,000 images. The test batch contains exactly 1,000 randomly selected images from each class. The training batch contains the remaining images in random order, but some training batches may contain more images from one class than another. Overall, the sum of the five training sets contains exactly 5,000 images from each class. The CIFAR-100 dataset has 100 classes, each containing 600 images. Each class has 500 training images and 100 test images. The 100 classes in CIFAR-100 are divided into 20 superclasses. Each image comes with a "fine" label (the class it belongs to) and a "coarse" label (the superclass it belongs to). For more details, please refer to the introduction on the webpage https: / / www.cnblogs.com / cloud-ken / p / 8456878.html.

[0121] The above existing methods are referenced as follows:

[0122] The ResNet-110 method can be found in “K.He, X.Zhang, S.Ren, and J.Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.”

[0123] The FractalNet method can be found in "G. Larsson, M. Maire, and G. Shakhnarovich. Fractalnet: Ultra-deep neural networks without residuals. arXiv preprint arXiv: 1605.07648, 2016."

[0124] The DenseNet (k = 24) method and DenseNet-B (k = 40) can be referred to the introduction in "G. Huang, Z. Liu, L. Van Der Maaten, and KQ Weinberger. Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4700–4708, 2017.";

[0125] The Wide ResNet method can be found in "S.Zagoruyko and N.Komodakis.Wide residualnetworks.arXiv preprint arXiv:1605.07146,2016.";

[0126] The ResNeXt-29 (8x64d) method can be found in “S. Xie, R. Girshick, P. Doll′ar, Z. Tu, and K. He. Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1492–1500, 2017.”

[0127] The Hierarchical Evolution method can be found in “H.Liu, K.Simonyan, O.Vinyals, C.Fernando, and K. Kavukcuoglu.Hierarchical representations for efficient architecture search. In International Conference on Learning Representations, 2018.”

[0128] For the AmoebaNet-A method, please refer to the introduction in "E.Real, A.Aggarwal, Y.Huang, and Q.V.Le.Regularized evolution for image classifier architecture search.InProceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages4780–4789, 2019.";

[0129] The NASNet-A method can be found in "B.Zoph, V.Vasudevan, J.Shlens, and QVLe.Learning transferable architectures for scalable image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8697–8710, 2018."

[0130] The DARTS method can be found in “H.Liu, K.Simonyan, and Y.Yang.Darts:Differentiablearchitecture search.In International Conference on Learning Representations,2018.”

[0131] The ENAS (macro) method and the ENAS (micro) method can be found in “H. Pham, M. Guan, B. Zoph, Q. Le, and J. Dean. Efficient neural architecture search via parameters sharing. In International Conference on Machine Learning, pages 4095–4104. PMLR, 2018.”;

[0132] The Block-QNN-S method can be found in “Z. Zhong, J. Yan, W. Wu, J. Shao, and C.-L. Liu. Practical block- wise neural network architecture generation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2423–2432, 2018.”

[0133] For the TE-NAS method, please refer to the introduction in "W.Chen, X.Gong, and Z.Wang.Neural architecture search on imagenet in four gpu hours: A theoretically inspired perspective. In International Conference on Learning Representations, 2020.";

[0134] The Large-scale Evolution method can be found in “E. Real, S. Moore, A. Selle, S. Saxena, YL Suematsu, J. Tan, QV Le, and A. Kurakin. Large-scale evolution of image classifiers. In International Conference on Machine Learning, pages 2902–2911. PMLR, 2017.”

[0135] For the AE-CNN method, please refer to the introduction in “Y.Sun, B.Xue, M.Zhang, and GGYen.Completely automatedcnn architecture design based on blocks.IEEE transactions on neural networksand learning systems, 31(4):1242–1254, 2019.”

[0136] For the CNN-GA method, please refer to the introduction in “Y.Sun, B.Xue, M.Zhang, GGYen, and J.Lv.Automaticallydesigning cnn architectures using the genetic algorithm for imageclassification.IEEE transactions on cybernetics, 50(9):3840–3854, 2020.”;

[0137] For the NAS method, please refer to the introduction in "B.Zoph and QVLe.Neural architecture search with reinforcement learning.ArXiv preprint arXiv:1611.01578,2016.";

[0138] The NSGA-Net method can be referred to in “Z. Lu, I. Whalen, V. Boddeti, Y. Dhebar, K. Deb, E. Goodman, and W. Banzhaf. Nsga-net: neural architecture search using multi-objective genetic algorithm. In Proceedings of the Genetic and Evolutionary Computation Conference, pages 419–427, 2019.”.

[0139] The optimal network structure searched by the method proposed in the present invention in Table 1 is denoted as EX-Net.

[0140] Table 1: Comparison results of the proposed method with other algorithms on the CIFAR-10 and CIFAR-100 datasets, in terms of test accuracy (%), number of parameters, number of GPU days searched, and number of GPUs used.

[0141]

[0142] The analysis is as follows:

[0143] 1) Comparison results with manually designed networks

[0144] As can be seen from Table 1, compared with the most advanced network structures designed manually, the network structure EX-Net searched by the method of this application is much better than FractalNet and WideResNet in terms of test accuracy and number of parameters on CIFAR-10 and CIFAR-100. For DenseNet (k = 24), EX-Net shows better test accuracy on CIFAR-10 and CIFAR-100, while the number of parameters obtained by EX-Net on CIFAR-10 and CIFAR-100 is only 6.9% and 15.8% of that of DenseNet (k = 24). The number of parameters in EX-Net is slightly higher than that of ResNet-100, but the test accuracy of EX-Net on both datasets has been greatly improved, by 3.5% and 8.9% respectively. Compared with DenseNet-B (k = 40) and ResNeXt-29 (8x64d), EX-Net has better test accuracy performance on CIFAR-10. On CIFAR-100, although EX-Net's accuracy is slightly lower than theirs, the number of parameters of EX-Net is only 16.8% and 12.5% ​​of the number of parameters of DenseNet-B (k=40) and ResNeXt-29 (8x64d), which is greatly reduced. Compared with ResNeXt-29 (8x64d), EX-Net only uses 1 / 8 of the GPU resources.

[0145] Therefore, compared with the most advanced manually designed network structure, the network structure EX-Net searched by the present invention can achieve higher accuracy performance. At the same time, EX-Net has much fewer parameters than most manually designed network structures.

[0146] 2) Comparison results with semi-automatic NAS algorithms

[0147] As can be seen from Table 1, compared with the semi-automatic NAS algorithm, compared with Hierarchical Evolution, Block-QNN-S and ENAS (macro), the network structure EX-Net searched by the method of this application is completely superior to them in terms of test accuracy and number of parameters, while greatly reducing the search time cost (reduced by 16 to 4500 times). Compared with NASNet-A, EX-Net is slightly worse than it in terms of test accuracy, but the number of parameters of EX-Net is much less than that of NASNet-A. In addition, EX-Net searches 100,000 times faster than NASNet-A, and consumes only 1 / 500 of the GPU resources consumed by NASNet-A. EX-Net has better test accuracy and fewer parameters than AmoebaNet-A. The GPU Days required for EX-Net are only 0.02, which is only 1 / 157,500 of AmoebaNet-A, and the computing resources required by the GPU are only 1 / 450 of AmoebaNet-A. DARTS and ENAS (micro) have slightly better accuracy performance on CIFAR-10 than EX-Net, but EX-Net has much fewer parameters. With the same GPU resource consumption, EX-Net's search time is 75 times and 25 times less than them respectively. In addition, although EX-Ne's accuracy performance is not as good as TE-NAS, the number of parameters of EX-Net and the number of GPU days consumed by EX-Net are only half of TE-NAS.

[0148] Therefore, compared with the semi-automatic NAS algorithm, the network structure EX-Net searched by the present application method is competitive in test accuracy and shows better advantages in the number of parameters. In addition, EX-Net also shows great advantages in search time cost and required computing resource consumption.

[0149] 3) Comparison results with the fully automatic NAS algorithm

[0150] Compared with the fully automatic NAS algorithm, the network structure EX-Net searched by the method of this application shows advantages over Large-scale Evolution and NAS in terms of accuracy performance and number of parameters. In addition, EX-Net only consumes 0.02 GPU Days, which is much lower than Large-scale Evolution and NAS. At the same time, the GPU resources required by EX-Net are 800 times less than NAS. EX-Net is superior to AE-CNN in terms of test accuracy and number of parameters on CIFAR-10 and CIFAR-100. EX-Net has achieved better improvements in search time cost and required GPU resource consumption. Compared with CNN-GA, EX-Net has higher test accuracy on CIFAR-10 and fewer parameters. In addition, EX-Net has better accuracy performance on the more complex CIFAR-100, while the number of parameters is close to CNN-GA. The search time of EX-Net is only about 1 / 1750 of that consumed by CNN-GA. NSGA-Net has slightly better accuracy than EX-NET on CIFAR-10 (97.5% vs. 96.83%), but the number of parameters of EX-Net is only 1 / 13 of that of NSGA-Net (1.9M vs. 26.8M). When using the same computing resources, EX-Net's search time is 200 times less than NSGA-Net.

[0151] Therefore, in the comparison of fully automatic NAS algorithms, the network structure EX-Net searched by the method of this application shows great advantages in all objectives.

[0152] in conclusion

[0153] In summary, the network structure EX-Net searched by the method of this application exceeds most manually designed network structures in test accuracy, while having fewer parameters. EX-Net also shows great advantages over most automatic NAS algorithms in terms of test accuracy and number of parameters. At the same time, it requires fewer GPU resources and reduces the search time by 200 to 1,120,000 times. Compared with the semi-automatic NAS algorithm, considering the difference in search space and the involvement of manual design, the advantage of EX-Net in test accuracy performance is not obvious, but EX-Net has much fewer parameters and greatly reduces the search time cost and computing resource consumption.

[0154] Some steps in the embodiments of the present invention may be implemented using software, and the corresponding software program may be stored in a readable storage medium, such as a CD or a hard disk.

[0155] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A low-cost automatic search method for neural network structures for image classification, It is characterized in that The method comprises: Step 1: For the image classification task, determine the main framework of the neural network structure, randomly generate X network structures as the population P, and each individual in the population represents a randomly generated network structure; the main framework of the neural network structure includes a standard convolution layer, unit num Reg Unit modules and a global average pooling layer, each Reg Unit module includes block num group convolution Reg Block; and each Reg Unit module contains a SENet module with a probability of 50%, and the SENet module simulates the attention mechanism through Squeeze-and-Excitation; The number of Reg Unit modules unit num, the number of group convolution Reg Block block num, the number of branches of group convolution RegBlock group, and the width of the second convolution layer in each branch width are randomly generated; Step 2: Set the three-stage separation point S for the subsequent population evolution stage 1 , S 2 and the maximum number of generations of evolution Max_gen; Step 3: Calculate the NTK condition number K of the network structure of each individual in the population P N As the fitness of an individual; NTK is used to characterize the gradient descent training dynamics of infinite width or finite width deep network architectures; Step 4: The population enters the evolution stage, and the tournament selection is used to select individual mutation operations to generate new network structure individuals. Different indicators are selected for environmental selection to eliminate individuals according to the stage of the current evolutionary algebra G. Step 5: After reaching the maximum number of evolutionary generations Max_gen, select the fitness K of the individual N The network structure with the smallest value is used as the searched neural network structure for image classification tasks; In step 4, different indicators are selected according to the stage of the current evolutionary generation G to perform environmental selection to eliminate individuals, including: In the first and third stages, that is, when 0 < G ≤ S 1 and S 2 < G ≤ Max_gen, the fitness K of the individuals is selected N as the criterion to eliminate individuals; In the second stage, when S 1 <G≤S 2 When selecting an individual, the life span of the individual is used as the criterion to eliminate the individual, and the life span of the individual is the number of evolutionary generations that the individual has experienced; The population evolution process includes: Randomly select k individuals from the population; from these k individuals, according to the fitness K of each individual N The value is large, and the individuals with the best fitness before t are selected as the parent individuals; T parent individuals generate t offspring individuals through a set of mutation operators; after the offspring individuals are generated, they are evaluated and added to the existing population; According to the stage of the current evolutionary generation, the corresponding standard is used to eliminate individuals in the environmental selection; according to the current standard, the t worst individuals are eliminated so that the population size remains unchanged, and the remaining individuals form a new population and enter the next generation of evolution; The group convolution Reg Block in each network structure contains group branches, each branch consists of three convolution layers and one pooling layer, where the pooling layer is in the third layer; the first and fourth convolution layers use 1×1 kernels to adjust the number of feature maps, and the second convolution layer uses 3×3 kernels to extract feature maps. All convolution layers follow the order of convolution operation, ReLu activation function and batch normalization layer; the third pooling layer is used to halve the size of the input data; the input data is image data; For M×M input data, the number of pooling layers in the third layer of each branch of the group convolution Reg Block cannot be greater than The t parent individuals generate t offspring individuals through a set of mutation operators; after the offspring individuals are generated, they are evaluated and added to the existing population, including: Randomly select a mutation position pos within the length of the parent individual ij , which represents the position of the jth RegBlock in the i-th Reg Unit, and the position is determined by the order of the Reg Unit in the network structure and the position order of the Reg Block in the Reg Unit; Randomly select a mutation operator to perform mutation of the parent individual, wherein the mutation operator includes an add operator, a remove operator, and a change operator; Add operator: at mutation position pos ij Add a Reg Block with random parameter settings; Remove operator: remove at mutation position pos ij Reg Block on; Change operator: randomly change the mutation position pos ij Parameters of the Reg Block on ; When implementing the add operator, if the length of the parent individual reaches the upper limit, the add operator cannot be implemented, and the only options are to remove the operator or change the operator; When implementing the removal operator, if the length of the parent individual reaches the lower limit, the removal operator cannot be performed, and the only options are to add an operator or change the operator; Calculate the condition number K of NTK for each network structure using the CIFAR-10 and CIFAR-100 datasets N fitness as an individual; According to each set of training images and corresponding labels in the CIFAR-10 and CIFAR-100 datasets, the eigenvalue λ of NTK between training sets is obtained. k , according to each eigenvalue λ k Get the NTK condition number K of the network structure N , the calculation formula is as follows: Among them, λ 0 Denotes the eigenvalue λ k The maximum value of λ m Denotes the eigenvalue λ k The minimum value of .

2. An image classification method, It is characterized in that The method uses the neural network structure searched out by the method described in claim 1 to perform image classification.

3. The method according to claim 2, It is characterized in that The method comprises: The image to be classified is input into the neural network structure, and the features of the image to be classified are extracted through the standard convolution layer; Further feature extraction is performed through unit num Reg Unit modules, where the output of each group convolution Reg Block in each Reg Unit module is connected by the output features of each branch and the residual connection, and then the feature map is obtained through the SENet module with a probability of 50%, and then the feature map output by the Reg Units is flattened into a feature vector through the global average pooling layer. Finally, a fully connected layer with a softmax layer is set as a classifier to convert the feature vector into the final classification result.

Citation Information

Patent Citations

  • Method for searching convolutional neural network

    CN111242268A

  • Rapid attention neural network architecture search method based on evolutionary method

    CN112465120A