Method and apparatus for determining a neural network
By obtaining multiple initial search spaces, evaluating and selecting candidate neural networks, and using technical means such as Pareto optimal solution and combining regularization layers to optimize neural network combinations, the problem of insufficient performance of neural network combinations in the existing technology is solved, and a higher performance neural network is achieved.
Patent Information
- Application Number
- CN201911090334.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2039-11-08
AI Technical Summary
In terms of how to obtain a high-performance neural network composed of multiple neural networks, it is difficult for the prior art to effectively determine the combination method to improve overall performance.
By obtaining multiple initial search spaces, determining M candidate neural networks and evaluating them, N first target neural networks are selected based on the evaluation results, considering the combination method between candidate neural networks, and using technical means such as Pareto optimal solution and combination regularization layer to optimize the network structure.
A higher performance neural network combination is realized, which improves operating speed and accuracy, meets task requirements, and improves the overall performance of neural networks.
Smart Images

Figure CN112784954B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and more specifically, to a method and apparatus for determining a neural network. Background Art
[0002] A neural network is a type of mathematical computational model that mimics the structure and function of biological neural networks (the central nervous system of animals). A neural network can include multiple layers with different functions, each consisting of parameters and computational formulas. Different layers in a neural network have different names depending on the computational formula or function. For example, a layer that performs convolution calculations is called a convolutional layer, which is often used to extract features from input signals (e.g., images).
[0003] The neural networks used in some application scenarios can be composed of multiple neural networks. For example, the neural network used to perform object detection tasks can be composed of a residual network (ResNet), a multi-level feature extraction model, and a region proposal network (RPN).
[0004] Therefore, how to obtain a neural network composed of multiple neural networks is a technical problem that needs to be solved urgently. Summary of the Invention
[0005] The present application provides a method and related apparatus for determining a neural network, which can obtain a combined neural network with higher performance.
[0006] First aspect, this application provides a method for determining a neural network, which includes: obtaining a plurality of initial search spaces, where each initial search space includes one or more neural networks, the functions of the neural networks in any two of the initial search spaces are different, and the functions of any two neural networks in the same initial search space are the same while their network structures are different; determining M candidate neural networks according to the plurality of initial search spaces, where each candidate neural network includes a plurality of candidate sub-networks, the plurality of candidate sub-networks belong to the plurality of initial search spaces, and any two of the plurality of candidate sub-networks belong to different initial search spaces, and M is a positive integer; evaluating the M candidate neural networks to obtain M evaluation results; according to the M evaluation results, determining N candidate neural networks from the M candidate neural networks, and determining N first target neural networks according to the N candidate neural networks, where the N first target neural networks and the N candidate neural networks are in one-to-one correspondence, each of the N candidate neural networks includes a plurality of candidate sub-networks, each of the N first target neural networks includes a plurality of target sub-networks, and each target sub-network in each first target neural network includes the same blocks as each sub-network in the corresponding candidate neural network, and N is a positive integer less than or equal to M.
[0007] In this method, after sampling candidate neural networks from a plurality of initial search spaces, the entire candidate neural network is evaluated, and then the first target neural network is determined according to the evaluation result and the candidate neural network. After sampling the candidate neural network in this way and determining the first target neural network according to the overall evaluation result of the candidate neural network, compared with the method of separately evaluating the candidate sub-networks and then determining the first target neural network according to the evaluation results of the candidate sub-networks, the combination method between the candidate sub-networks is fully considered, and a first target neural network with better performance can be obtained.
[0008] In some possible implementation manners, the evaluation result of the candidate neural network includes one or more of the following: running speed, accuracy, number of parameters, or floating-point operation count.
[0009] In some possible implementation manners, determining N first target neural networks according to the M evaluation results and the M candidate neural networks includes: according to the M evaluation results, determining the N candidate neural networks whose evaluation results meet the task requirements among the M candidate neural networks as the N first target neural networks.
[0010] For example, determining the N candidate neural networks among the M candidate neural networks whose running speed and / or accuracy meet the preset task requirements as the N first target neural networks.
[0011] In some possible implementation manners, the evaluation results of the candidate neural networks include running speed and accuracy. Among them, determining N first target neural networks according to the M evaluation results and the M candidate neural networks includes: determining the N candidate neural networks from the M candidate neural networks according to the M evaluation results, where the N candidate neural networks are the Pareto optimal solutions of the M candidate neural networks when running speed and accuracy are taken as the objectives; determining the N first target neural networks according to the N candidate neural networks.
[0012] Since the N candidate neural networks obtained according to this implementation manner are the Pareto optimal solutions of the M candidate neural networks, the performance of these N candidate neural networks is better than that of other candidate neural networks, which makes the performance of the N first target neural networks determined according to these N candidate neural networks also better.
[0013] In some possible implementation manners, determining the N first target neural networks according to the N candidate neural networks includes: determining these N candidate neural networks as these N first target neural networks.
[0014] In some possible implementation manners, determining the N first target neural networks according to the N candidate neural networks includes: determining a plurality of target search spaces according to a plurality of candidate sub-networks of the i-th candidate neural network, where the plurality of target search spaces correspond one-to-one to the plurality of candidate sub-networks of the i-th candidate neural network, each target search space in the plurality of target search spaces includes one or more neural networks, and the blocks included in each neural network in each target search space are the same as the blocks included in the candidate sub-network corresponding to each target search space; determining the i-th first target neural network according to the plurality of target search spaces, where a plurality of target sub-networks in the i-th first target neural network belong to the plurality of target search spaces, and any two target sub-networks in the plurality of target sub-networks of the i-th first target neural network belong to different target search spaces.
[0015] That is to say, without changing the blocks, a first target neural network with better performance is re-searched.
[0016] In some possible implementation manners, the method further includes: determining N second target neural networks according to the N first target neural networks, where the i-th second target neural network among the N second target neural networks is obtained by performing one or more of the following processes on the i-th first target neural network: adding a combined regularization layer after a convolutional layer in a target sub-network of the i-th first target neural network, adding a combined regularization layer after a fully-connected layer in the target sub-network of the i-th first target neural network, and normalizing weights of the convolutional layer in the target sub-network of the i-th first target neural network.
[0017] This implementation manner can improve the performance of the second target neural network and enhance the training speed of the second target neural network.
[0018] In some possible implementation manners, the method further includes: evaluating the N second target neural networks to obtain evaluation results of the N second target neural networks. These N evaluation results can be used to select a more suitable second target neural network from the N second target neural networks according to task requirements, thereby improving the completion quality of the task.
[0019] In some possible implementation manners, the evaluating the N second target neural networks to obtain the evaluation results of the N second target neural networks includes: randomly initializing network parameters in the i-th second target neural network; training the i-th second target neural network according to training data; and testing the trained i-th second target neural network according to test data to obtain the evaluation result of the trained i-th second target neural network.
[0020] In some possible implementation manners, the first target neural network is used for object detection, where the multiple initial search spaces include a first initial search space, a second initial search space, a third initial search space, and a fourth initial search space. The first initial search space includes residual networks with different depths, second-generation residual networks (ResNext) with different depths, and / or MobileNet networks with different depths. The second initial search space includes connection paths of features at different levels. The third initial search space includes a region proposal network (RPN) and / or a region proposal by guided anchoring (GA-RPN). The fourth initial search space includes a one-stage detection head network (Retina-head), a fully-connected detection head network, a fully-convolutional detection head network, and / or a cascade network head (Cascade-head).
[0021] In some possible implementation manners, the first target neural network is used for image classification. Among them, the multiple initial search spaces include a first initial search space and a second initial search space. The first initial search space includes residual networks with different depths, ResNext with different depths, and / or dense connection networks (DenseNet) with different widths. The neural networks in the second initial search space include fully connected layers.
[0022] In some possible implementation manners, the first target neural network is used for image segmentation. Among them, the multiple initial search spaces include a first initial search space, a second initial search space, and a third initial search space. The first initial search space includes residual networks with different depths, ResNext with different depths, and / or high-resolution networks with different widths. The second initial search space includes atrous spatial pyramid pooling networks, pyramid pooling networks, and / or networks including dense prediction units. The third initial search space includes U-Net models and / or fully convolutional networks.
[0023] In a second aspect, the present application provides an apparatus for determining a neural network. The apparatus includes: an acquisition module, configured to acquire multiple initial search spaces, where the initial search space includes one or more neural networks, the functions of the neural networks in any two of the initial search spaces are different, and the functions of any two neural networks in the same initial search space are the same and the network structures are different; a determination module, configured to determine M candidate neural networks according to the multiple initial search spaces, where the candidate neural networks include multiple candidate sub-networks, the multiple candidate sub-networks belong to the multiple initial search spaces, and any two of the multiple candidate sub-networks belong to different initial search spaces; an evaluation module, configured to evaluate the M candidate neural networks to obtain M evaluation results, where M is a positive integer; the determination module is further configured to: according to the M evaluation results, determine N candidate neural networks from the M candidate neural networks, and determine N first target neural networks according to the N candidate neural networks, where the N first target neural networks and the N candidate neural networks are in one-to-one correspondence, each of the N candidate neural networks includes multiple candidate sub-networks, each of the N first target neural networks includes multiple target sub-networks, and each target sub-network included in each first target neural network includes the same blocks as each sub-network included in the corresponding candidate neural network, and N is a positive integer less than or equal to M.
[0024] In some possible implementation manners, the evaluation results of the candidate neural networks include one or more of the following: running speed, accuracy, number of parameters, or floating-point operation count.
[0025] In some possible implementation manners, the evaluation result of the candidate neural network includes the running speed and the accuracy. Among them, the determining module is specifically configured to: determine the N candidate neural networks from the M candidate neural networks according to the M evaluation results, where the N candidate neural networks are the Pareto optimal solutions of the M candidate neural networks when the running speed and the accuracy are taken as the targets; determine the N first target neural networks according to the N candidate neural networks.
[0026] In some possible implementation manners, the determining module is specifically configured to: determine a plurality of target search spaces according to a plurality of candidate sub-networks of the i-th candidate neural network, where the plurality of target search spaces correspond to the plurality of candidate sub-networks of the i-th candidate neural network one by one, each target search space in the plurality of target search spaces includes one or more neural networks, and the blocks included in each neural network in each target search space are the same as the blocks included in the candidate sub-network corresponding to each target search space; determine the i-th first target neural network according to the plurality of target search spaces, where the plurality of target sub-networks in the i-th first target neural network belong to the plurality of target search spaces, and any two target sub-networks in the plurality of target sub-networks of the i-th first target neural network belong to different target search spaces.
[0027] In some possible implementation manners, the determining module is further configured to: determine N second target neural networks according to the N first target neural networks, where the i-th second target neural network in the N second target neural networks is obtained by performing one or more of the following processes on the i-th first target neural network: adding a combined regularization layer after the convolutional layer in the target sub-network of the i-th first target neural network, adding a combined regularization layer after the fully connected layer in the target sub-network of the i-th first target neural network, and normalizing the weights of the convolutional layer in the target sub-network of the i-th first target neural network.
[0028] In some possible implementation manners, the evaluation module is further configured to: evaluate the N second target neural networks to obtain the evaluation results of the N second target neural networks.
[0029] In some possible implementation manners, the evaluation module is specifically configured to: randomly initialize the network parameters in the i-th second target neural network; train the i-th second target neural network according to the training data; test the trained i-th second target neural network according to the test data to obtain the evaluation result of the trained i-th second target neural network.
[0030] In some possible implementation manners, the first target neural network is used for object detection. Among them, the multiple initial search spaces include a first initial search space, a second initial search space, a third initial search space, and a fourth initial search space. The first initial search space includes residual networks with different depths, second-generation residual networks with different depths, and / or mobile networks with different depths. The second initial search space includes connection paths of features at different levels. The third initial search space includes general region candidate networks and / or anchor-guided region candidate networks. The fourth initial search space includes one-stage detection head networks, fully-connected detection head networks, fully-convolutional detection head networks, and / or cascaded detection head networks.
[0031] In some possible implementation manners, the first target neural network is used for image classification. Among them, the multiple initial search spaces include a first initial search space and a second initial search space. The first initial search space includes residual networks with different depths, second-generation residual networks with different depths, and / or densely connected networks with different widths. The neural network in the second initial search space includes fully-connected layers.
[0032] In some possible implementation manners, the first target neural network is used for image segmentation. Among them, the multiple initial search spaces include a first initial search space, a second initial search space, and a third initial search space. The first initial search space includes residual networks with different depths, second-generation residual networks with different depths, and / or high-resolution networks with different widths. The second initial search space includes atrous spatial pyramid pooling networks, pyramid pooling networks, and / or networks including dense prediction units. The third initial search space includes U-Net models and / or fully-convolutional networks.
[0033] In a third aspect, a device for determining a neural network is provided. The device includes: a memory for storing a program; a processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method in the first aspect.
[0034] In a fourth aspect, a computer-readable medium is provided. The computer-readable medium stores instructions for a device to execute, and the instructions are used to implement the method in the first aspect.
[0035] In a fifth aspect, a computer program product including instructions is provided. When the computer program product runs on a computer, the computer is caused to execute the method in the first aspect above.
[0036] In a sixth aspect, a chip is provided. The chip includes a processor and a data interface. The processor reads instructions stored on a memory through the data interface and executes the method in the first aspect above.
[0037] Optionally, as an implementation, the chip may further include a memory in which instructions are stored, and the processor is configured to execute the instructions stored on the memory. When the instructions are executed, the processor is configured to execute the method in the first aspect. Description of the Drawings
[0038] Figure 1 is an exemplary flowchart of the method for determining a neural network in the present application;
[0039] Figure 2 is an example diagram of the initial search space of the neural network for performing an object detection task in the present application;
[0040] Figure 3 is an example diagram of the initial search space of the neural network for performing an image classification task in the present application;
[0041] Figure 4 is an example diagram of the initial search space of the neural network for performing an image segmentation task in the present application;
[0042] Figure 5 is another exemplary flowchart of the method for determining a neural network in the present application;
[0043] Figure 6 is an example diagram of the Pareto front of the candidate neural network in the present application;
[0044] Figure 7 is another exemplary flowchart of the method for determining a neural network in the present application;
[0045] Figure 8 is another exemplary flowchart of the method for determining a neural network in the present application;
[0046] Figure 9 is an exemplary structural diagram of the apparatus for determining a neural network according to an embodiment of the present application;
[0047] Figure 10 is an exemplary structural diagram of the apparatus for determining a neural network according to an embodiment of the present application;
[0048] Figure 11 is another example diagram of the Pareto front of the candidate neural network in the present application. Detailed Embodiments
[0049] For ease of understanding, the following gives an explanation of the concepts related to the present application.
[0050] (1) Neural Network
[0051] A neural network can be composed of neural units. A neural unit can refer to an operation unit with x s and intercept 1 as inputs. The output of this operation unit can be:
[0052]
[0053] where s = 1, 2, …… n, n is a natural number greater than 1, W s is the weight of x s , b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer, and the activation function can be the sigmoid function. A neural network is a network formed by connecting multiple such single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.
[0054] (2) Deep neural network
[0055] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with multiple hidden layers. According to the position of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i + 1-th layer.
[0056] Although the DNN seems very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: where, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer only performs such a simple operation on the input vector to obtain the output vector Due to the large number of layers in the DNN, the number of coefficients W and offset vectors is also relatively large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscript corresponds to the third-layer index 2 of the output and the second-layer index 4 of the input.
[0057] In summary, the coefficient from the k-th neuron in the (L - 1)-th layer to the j-th neuron in the L-th layer is defined as
[0058] It should be noted that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0059] (3) Convolutional Neural Network
[0060] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers, and this feature extractor can be regarded as a filter. A convolutional layer refers to the neuron layer in a convolutional neural network that performs convolutional processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can only be connected to some neighboring-layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some neurons arranged in a rectangle. The neurons in the same feature plane share weights, and the shared weight here is the convolutional kernel. Sharing weights can be understood as a way of extracting image information that is independent of position. The convolutional kernel can be initialized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolutional kernel can learn to obtain reasonable weights. Additionally, the direct benefit of sharing weights is to reduce the connections between layers of the convolutional neural network while also reducing the risk of overfitting.
[0061] (4) Loss Function
[0062] During the process of training a deep neural network, since we hope that the output of the deep neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the real target value, and then update the weight vector of each layer of the neural network according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is too high, we adjust the weight vector to make it predict lower, and keep adjusting until the deep neural network can predict the real target value or a value very close to the real target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0063] (5) Backpropagation algorithm
[0064] The neural network can use the backpropagation (BP) algorithm to correct the magnitudes of the parameters in the initial neural network during the training process, so that the reconstruction error loss of the neural network becomes smaller and smaller. Specifically, forward propagating the input signal until the output will generate an error loss, and updating the parameters in the initial neural network by backpropagating the error loss information, so as to make the error loss converge. The backpropagation algorithm is a reverse propagation movement dominated by the error loss, aiming to obtain the optimal parameters of the neural network, such as the weight matrix.
[0065] (6) Pareto solution
[0066] The Pareto solution is also called the non-dominated solution or non-dominated solution. When there are multiple goals, due to the conflict and incomparability between goals, a solution may be the best for one goal and the worst for other goals. These solutions that will necessarily weaken at least one other goal while improving any goal are called non-dominated solutions or Pareto solutions.
[0067] Pareto Optimality is a state of resource allocation. Without making any goal worse, it is impossible to make some goals better. Pareto Optimality is also called Pareto efficiency and Pareto improvement.
[0068] A set of target optimal solutions is called the Pareto optimal set. The surface formed by the optimal set in space is called the Pareto front.
[0069] For example, when the goals are the running speed and accuracy of a neural network, when the running speed of a neural network is better than that of other neural networks, its accuracy may be very poor. When the accuracy of this neural network is better than that of other neural networks, its running speed may be very poor. For a certain neural network, if it is impossible to improve its prediction accuracy without deteriorating its running accuracy, then this neural network can be called a Pareto optimal solution with running accuracy and prediction accuracy as the goals.
[0070] (7) Backbone network
[0071] The backbone network is used to extract the features of the input image to obtain multi-level (multi-scale) features of the image. Commonly used backbone networks include ResNet, ResNext, MobileNet, or DenseNet with different depths. The main difference between different series of backbone networks lies in the different basic units that make up the network. For example, the ResNet series includes ResNet-50, ResNet-101, and ResNet-152, and its basic unit is the bottleneck network block. ResNet-50 contains 16 bottleneck network blocks, ResNet-101 contains 33 bottleneck network blocks, and ResNet-152 contains 50 bottleneck network blocks. The difference between the ResNext series and the ResNet series is that the basic unit is replaced from the bottleneck network block to the bottleneck network block with grouped convolution. The basic unit of the MobileNet series is the depthwise separable convolution. The basic units of the DenseNet series are the dense unit module and the transition network module.
[0072] (8) Multi-level feature extraction network (Neck)
[0073] The multi-level feature extraction network is used to screen and fuse multi-scale features to generate a more compact and expressive feature vector. The multi-level feature extraction network can include a fully convolutional pyramid network with different scale connections, an atrous spatial pyramid pooling (ASPP) network, a pooling pyramid network, or a network including dense prediction units.
[0074] (9) Prediction module
[0075] The prediction module is used to output prediction results related to the application task.
[0076] The prediction module may include a head prediction network for converting features into prediction results that ultimately meet the requirements of the task. For example, in an image classification task, the final output prediction result is a probability vector indicating the probability of the input image belonging to each category; in an object detection task, the prediction result is the coordinates of all candidate object bounding boxes present in the input image and the probability of each candidate object bounding box belonging to each category; in an image segmentation task, the prediction module needs to output a class classification probability map at the pixel level of the image.
[0077] The head prediction network may include Retina-head, a fully connected detection head network, Cascade-head, a U-Net model, or a fully convolutional detection head network.
[0078] When the prediction module is used for the object detection task in computer vision tasks, the prediction module may include a region proposal network (RPN) and a head prediction network.
[0079] RPN is a component module in a two-stage detection network, a fast regression classifier for generating rough object positions and class label information, mainly composed of two branches. The first branch classifies foreground and background for each anchor point, and the second branch calculates the offset of the bounding box relative to the anchor point.
[0080] Generally, a two-layer simple network including a binary classifier and bounding box regression is used to implement RPN. Bounding box regression is a regression model for object detection that searches for a regression window closer to the real window with a smaller loss function value near the object location obtained from the sliding window.
[0081] At this time, the head prediction network is used to further optimize the classification and detection results obtained by RPN, generally implemented through a multi-layer network much more complex than RPN. The combination of RPN and the head prediction network enables the object detection system to quickly remove a large number of invalid image regions and concentrate on carefully detecting more potential image regions, achieving fast and good results.
[0082] The methods and devices of this application can be applied in many fields of artificial intelligence, such as intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, autonomous driving, and other fields.
[0083] Specifically, the methods and devices of this application can be specifically applied to fields that require the use of (deep) neural networks, such as autonomous driving, image classification, image segmentation, object detection, image retrieval, image semantic segmentation, image quality enhancement, image super-resolution, and natural language processing.
[0084] For example, by using the method of the present application to obtain a neural network suitable for album classification, the album classification neural network can be used to classify pictures, so as to label pictures of different categories, which is convenient for users to view and search. In addition, the classification labels of these pictures can also be provided to the album management system for classification management, saving the management time of users, improving the efficiency of album management, and enhancing the user experience.
[0085] For another example, by using the method of the present application to obtain a neural network that can detect targets such as pedestrians, vehicles, traffic signs or lane lines, it can help autonomous vehicles drive more safely on the road.
[0086] For another example, by using the method of the present application to obtain a neural network that can segment objects in an image, so as to understand the content of the currently captured image according to the segmentation result, and give a decision basis for the rendering of the photo-taking effect, thereby providing the best image rendering effect for users.
[0087] Next, the technical solutions in the present application will be described with reference to the accompanying drawings.
[0088] Figure 1 It is an exemplary flowchart of the method for determining a neural network in the present application. The method includes S110 to S140.
[0089] S110, obtain a plurality of initial search spaces. In the plurality of initial search spaces, each initial search space includes one or more neural networks. The functions of the neural networks in any two initial search spaces are different, and the functions of any two neural networks in the same initial search space are the same and the network structures are different.
[0090] Among them, at least one of the plurality of initial search spaces includes a plurality of neural networks.
[0091] In the embodiments of the present application, the network structure of a neural network may include one or more stages, and each stage may include at least one block. Among them, a block may be composed of basic atoms in a convolutional neural network, and these basic atoms include: convolutional layer, pooling layer, fully connected layer or non-linear activation layer, etc. A block may also be referred to as a basic unit or a basic module.
[0092] In a convolutional neural network, features usually exist in a three-dimensional form (length, width and depth). A feature can be regarded as a superposition of multiple two-dimensional features. Among them, each two-dimensional feature of a feature can be called a feature map. Or, a feature map (two-dimensional feature) of a feature can also be called a channel of a feature. The length and width of a feature map can also be called the resolution of the feature map.
[0093] When a neural network includes multiple stages, the number of blocks in different stages can be different. Similarly, the resolutions of the input feature maps and output feature maps processed by different stages can also be different.
[0094] When a segment in a neural network includes multiple blocks, the number of channels of different blocks can be different. It should be understood that the number of channels of a block can also be referred to as the width of the block. Similarly, the resolutions of the input feature maps and output feature maps processed by different blocks can also be different.
[0095] The difference in the network structures of any two neural networks can include: the number of stages included in any two neural networks, the number of blocks in the stages, the number of channels of the blocks, the resolution of the input feature map of the stage, the resolution of the output feature map of the stage, the resolution of the input feature map of the block, and / or the resolution of the output feature map of the block are different.
[0096] Generally, the initial search space is determined according to the target task. That is to say, it is necessary to first determine the target task, and then determine which neural networks with what functions can be combined to form the target neural network for achieving the target task according to the target task, and then construct the initial search space of the neural network with this function.
[0097] Taking the target task of high-level computer vision tasks as an example, the implementation method of determining the initial search space is introduced below.
[0098] The target neural network for solving high-level computer vision tasks can be a convolutional neural network with a unified design paradigm. High-level computer vision tasks include object detection, image segmentation, image classification, etc.
[0099] Since the target neural network for performing the object detection task can include a backbone network, a multi-level feature extraction network, and a prediction network, and the prediction network includes a region proposal network and a head prediction network, the initial search space of the backbone network, the initial search space of the multi-level feature extraction network, the initial search space of the region proposal network, and the initial search space of the head prediction network can be constructed. In addition, the initial search space of the resolution of the input image of the backbone network can also be constructed.
[0100] Such as Figure 2As shown, the initial search space for the resolution of the input image may include 512×512, 800×600, 1333×800, etc.; the initial search space for the backbone network may include ResNet with a depth of 18, 34 (i.e., d = 18, 34…), etc., ResNext with a depth of 18, 34, etc., and MobileNet; the initial search space for the multi-level feature extraction network may include the fusion paths of different scales in the backbone network, such as the Feature Pyramid Network (FPN) with the reduction multiples of the corresponding feature resolution scales in the backbone network relative to the original image being 1, 2, 3, 4 1,2,3,4 and the Feature Pyramid Network (FPN) with reduction multiples of 2, 4, and 5 2,4,5 ; the initial search space for the region candidate network may include the ordinary region candidate network and the anchor-guided region candidate network (region proposal by guided anchoring, GA-RPN); the initial search space for the head prediction network may include the fully-connected detection head (FC detection head), the detection head containing a one-stage detector, the detection head containing a two-stage detector, and the cascaded detection head with the cascading times of 2, 3, etc., where n represents the cascading times.
[0101] Since the target neural network for performing the image classification task may include the backbone network and the head prediction network, the initial search space for the backbone network and the initial search space for the head prediction network can be constructed.
[0102] As Figure 3 shown, the initial search space for the backbone network may include backbone networks for classification such as ResNet, ResNext, and DenseNet; the initial search space for the head prediction network may include FC.
[0103] Since the target neural network for performing the image task may include the backbone network, the multi-level feature extraction network, and the head prediction network, the initial search space for the backbone network, the initial search space for the multi-level feature extraction network, and the initial search space for the head prediction network can be constructed.
[0104] As Figure 4As shown in the figure, the initial search space of the backbone network may include ResNet, ResNext, and the VGG network proposed by the Visual Geometry Group of the University of Oxford; the initial search space of the multi-level feature extraction network may include the ASPP network, the pyramid pooling network, and the upsampling+concate network for merging multi-scale features; the initial search space of the head prediction network may include the U-Net model, the fully convolutional networks (FCN), and the Dense Prediction Cell network (DPC).
[0105] Figures 2 to 4 The "+" in it represents the connection relationship after the neural networks in the search space are sampled.
[0106] S120. Determine M candidate neural networks according to the multiple initial search spaces. The candidate neural networks include multiple candidate sub-networks. The multiple candidate sub-networks belong to the multiple initial search spaces, and any two candidate sub-networks among the multiple candidate sub-networks belong to different initial search spaces. M is a positive integer.
[0107] For example, a neural network can be randomly sampled from each initial search space, and all the sampled neural networks are combined into a complete neural network, which is called a candidate neural network.
[0108] Another example is that a neural network can be randomly sampled from each initial search space, and all the sampled neural networks are combined into a complete neural network. Then, calculate the floating-point operations per second (FLOPS) of this complete neural network. If the FLOPS of this complete neural network meets the task requirements, then determine this complete neural network as a candidate neural network; otherwise, discard this complete neural network and resample.
[0109] For example, when the finally determined target neural network is used on a terminal device with relatively low computing power, generally, the FLOPS of this complete neural network cannot exceed the computing power of this terminal device, otherwise it is meaningless to apply this neural network to execute tasks on this terminal device.
[0110] If the network structure of the complete neural network obtained by each sampling is the same as that of the complete neural network obtained by the previous sampling, the complete neural network obtained by this sampling can be discarded and resampled.
[0111] Optionally, sampling can be performed from a partial search space to obtain candidate neural network models. The candidate neural networks obtained by sampling in this way can only include neural networks in the partial search space.
[0112] Perform multiple samplings according to the multiple initial search spaces. For example, perform at least M samplings to obtain M candidate neural networks.
[0113] S130, evaluate the M candidate neural networks to obtain M evaluation results of the M candidate neural networks.
[0114] For example, initialize the network parameters in each of the M candidate neural networks; input training data into each candidate neural network and train each candidate neural network to obtain M trained candidate neural networks. After obtaining the M trained candidate neural networks, input test data into the M trained candidate neural networks to obtain the evaluation results of the M candidate neural networks.
[0115] Among them, if the candidate sub-network in the candidate neural network has been trained before forming the candidate neural network, when initializing the network parameters in the candidate sub-network, the network parameters obtained by the previous training of the candidate sub-network can be loaded to complete the initialization. This can improve the training efficiency of the candidate neural network and ensure the convergence of the candidate neural network.
[0116] For example, when the candidate sub-network is ResNet trained by the ImageNet dataset, the network parameters obtained by training the ResNet with the ImageNet dataset can be loaded.
[0117] The ImageNet dataset refers to the public dataset used in the ImageNet large scale visual recognition challenge (ILSVRC) competition.
[0118] Of course, the network parameters in the candidate neural network can also be initialized in other ways, such as randomly generating the network parameters in the candidate neural network.
[0119] The evaluation results of the candidate neural network can include one or more of the following: the running speed, accuracy, number of parameters, or floating-point operation amount of the candidate neural network. The accuracy here refers to the accuracy of the task result obtained by the candidate neural network after inputting the test data and performing the corresponding task compared with the expected result.
[0120] Under normal circumstances, the number of training times of the candidate neural network can be less than the normal number of training times of neural networks in the art, the learning rate of each training of the candidate neural network can be less than the normal learning rate of neural networks in the art, and the training duration of the candidate neural network can be less than the normal training duration of neural networks in the art. That is to say, the candidate neural network is trained quickly.
[0121] S140. According to the M evaluation results and the M candidate neural networks, determine N first target neural networks. The first target neural networks include multiple target sub-networks. The N first target neural networks correspond one-to-one to N candidate neural networks among the M candidate neural networks. For the i-th first target neural network among the N first target neural networks, the multiple target sub-networks included therein correspond one-to-one to the multiple candidate sub-networks included in the i-th candidate neural network among the N candidate neural networks. For each target sub-network included in the multiple target sub-networks included in the i-th first target neural network, the blocks included therein are the same as the blocks included in the corresponding candidate sub-network among the multiple candidate sub-networks included in the i-th candidate neural network. N is a positive integer less than or equal to M, and i is a positive integer less than or equal to N.
[0122] Among them, the connection relationship between the target sub-networks in the first target neural network is the same as the connection relationship between the corresponding candidate sub-networks in the candidate sub-networks.
[0123] Among them, the blocks included in each target sub-network being the same as the blocks included in the corresponding candidate sub-network may include: the basic atoms in the blocks included in each target sub-network and the basic atoms in the blocks included in the corresponding candidate sub-network, and the number of these basic atoms and the connection relationship between these basic atoms are the same. For example, the candidate sub-network is a multi-level feature extraction module, and the multi-level feature extraction module is specifically a feature pyramid network, and when the feature pyramid network is fused at scales 2, 3, and 4, the corresponding target sub-network still maintains the fusion at scales 2, 3, and 4. Another example is that the candidate sub-network is a prediction module, and when the prediction module includes a head prediction network with a cascade number of 2, the target sub-network still includes a head prediction network with a cascade number of 2.
[0124] It can be understood that one or more of the stacking times of the blocks, the number of channels of the blocks, the upsampling position, the downsampling position of the feature map, or the convolution kernel size in each target sub-network may be different from those of the blocks in the corresponding candidate sub-network.
[0125] In some possible implementation manners, determining N first target neural networks according to the M evaluation results and the M candidate neural networks may include: determining, according to the M evaluation results, N candidate neural networks among the M candidate neural networks whose evaluation results meet the task requirements as the N first target neural networks.
[0126] For example, determining N candidate neural networks among the M candidate neural networks whose running speed and / or accuracy meet the preset task requirements as the N first target neural networks.
[0127] After sampling candidate neural networks from multiple initial search spaces, evaluating the entire candidate neural networks, and then determining the first target neural networks according to the evaluation results and the candidate neural networks. After sampling to obtain candidate neural networks, determining the first target neural networks according to the overall evaluation results of the candidate neural networks, compared with the method of separately evaluating candidate sub-networks and then determining the first target neural networks according to the evaluation results of the candidate sub-networks, can fully consider the combination ways between candidate sub-networks, and can obtain first target neural networks with better performance, so that when using the first target neural networks to execute tasks, better completion quality can be obtained.
[0128] In some possible implementation manners, the evaluation results of the candidate neural networks may include running speed and accuracy. In this implementation manner, determining N first target neural networks according to the M evaluation results and the M candidate neural networks may include: determining, according to the M evaluation results, N candidate neural networks from the M candidate neural networks, where the N candidate neural networks are the Pareto optimal solutions of the M candidate neural networks when running speed and accuracy are taken as the goals; determining N first target neural networks according to the N candidate neural networks.
[0129] Since the N candidate neural networks obtained according to this implementation manner are the Pareto optimal solutions of the M candidate neural networks, the performance of these N candidate neural networks is better than that of other candidate neural networks, which makes the performance of the N first target neural networks determined according to these N candidate neural networks also better.
[0130] The evaluation results of the candidate neural networks include running speed and prediction accuracy. When the running speed is taken as the abscissa and the prediction accuracy is taken as the ordinate, the spatial position relationship of the M candidate neural networks is as Figure 5 shown. Among them, the dotted line represents the Pareto front of these multiple first candidate neural networks. The first candidate neural networks located on the dotted line are the Pareto optimal solutions, and the set of all first candidate neural networks located on the dotted line is the Pareto optimal set.
[0131] Among them, after the first time determining a new first candidate neural network and its evaluation result according to M initial searches each time based on M initial search spaces, according to the spatial position relationship between the evaluation result and the evaluation results of the previous first candidate neural networks, the Pareto front of the first candidate neural network is re-determined, that is, the Pareto optimal set of the first candidate neural network is updated.
[0132] In this embodiment, when determining N first target neural networks according to the N candidate neural networks, the i-th first target neural network can be determined according to the i-th candidate neural network.
[0133] In some possible implementation manners, determining the i-th first target neural network according to the i-th candidate neural network may include: determining the i-th candidate neural network as the i-th first target neural network.
[0134] An exemplary flowchart of another implementation manner for determining the i-th first target neural network according to the i-th candidate neural network is as Figure 5 shown. The method may include S510 and S520.
[0135] S510, determining a plurality of target search spaces according to a plurality of candidate sub-networks of the i-th candidate neural network, the plurality of target search spaces corresponding one-to-one to the plurality of candidate sub-networks of the i-th candidate neural network, each target search space in the plurality of target search spaces including one or more neural networks, and each neural network included in each target search space having the same blocks as the candidate sub-network corresponding to each target search space.
[0136] Specifically, determining the target search space corresponding to each candidate sub-network according to each candidate sub-network in the plurality of candidate sub-networks, and finally obtaining a plurality of target search spaces. Each target search space may include one or more neural networks, but generally speaking, at least one target search space includes a plurality of neural networks.
[0137] When determining a plurality of target search spaces according to a plurality of candidate sub-networks of the i-th candidate neural network, the target search space corresponding to each candidate sub-network can be determined. For example, determining the target search space based on the structure of the blocks included in each candidate sub-network.
[0138] In some implementation manners, the candidate sub-network can be directly used as the target search space corresponding to the candidate sub-network. At this time, only one neural network is included in the target search space. That is to say, the candidate sub-network remains unchanged and is directly used as a target sub-network, and the target sub-networks corresponding to other candidate sub-networks in the i-th candidate neural network are searched, and then all the target sub-networks are combined into a target neural network.
[0139] In some other implementations, a corresponding target search space may be constructed based on the candidate sub-network. The target search space includes multiple target sub-networks, and the blocks included in each target sub-network in the target search space are the same as those included in the candidate sub-network.
[0140] At this time, the blocks included in each target sub-network are the same as those included in the candidate sub-network, which can be understood as including: the basic atoms in the blocks included in each target sub-network and the basic atoms in the corresponding blocks included in the candidate sub-network, and the number of these basic atoms and the connection relationship between these basic atoms are the same. For example, if the candidate sub-network is a multi-level feature extraction module, specifically a feature pyramid network, and the feature pyramid network fuses at scales 2, 3, and 4, the corresponding target sub-network still maintains the fusion at scales 2, 3, and 4. Another example is that if the candidate sub-network is a prediction module and the prediction module includes a head prediction network with a cascade number of 2, the target sub-network still includes a head prediction network with a cascade number of 2.
[0141] It can be understood that one or more of the stacking times of the blocks, the number of channels of the blocks, the upsampling position, the downsampling position of the feature map, or the convolution kernel size in each target sub-network may be different from those of the corresponding blocks in the candidate sub-network.
[0142] S520. Determine the i-th first target neural network according to the multiple target search spaces. The multiple target sub-networks in the i-th first target neural network belong to the multiple target search spaces, and any two target sub-networks in the multiple target sub-networks of the i-th first target neural network belong to different target search spaces.
[0143] For example, select one target sub-network from each target search space respectively, and then combine all the selected target sub-networks into a complete neural network.
[0144] When selecting a target sub-network from each target search space, a neural network can be randomly selected as the target sub-network; or the number of parameters of each neural network in the target search space can be calculated first, and then the neural network with a smaller number of parameters can be selected as the target sub-network. Of course, other methods can also be used to select the target sub-network. For example, the method of searching for neural networks in the prior art can be used to select the target sub-network, and this embodiment does not limit this.
[0145] After obtaining the complete neural network, in one implementation, the FLOPS of the neural network can be calculated. When the FLOPS of the neural network meets the task requirements, the complete neural network is used as the first target neural network.
[0146] After performing the method shown for each of the N candidate neural networks, N first target neural networks can be obtained. Figure 5 After performing the method shown for each of the N candidate neural networks, N first target neural networks can be obtained.
[0147] In this embodiment, after determining the N first target neural networks, the N first target neural networks can be evaluated to obtain N evaluation results of the N first target neural networks, and the N evaluation results can be saved so that the user can determine which first target neural networks meet the task requirements based on the N evaluation results, and thus determine whether to select which first target neural networks.
[0148] The evaluation result of each first target neural network can include one or more of the following: running speed, accuracy, or number of parameters. The accuracy here refers to the accuracy of the task result obtained after the first target neural network inputs the test data and performs the corresponding task compared with the expected result.
[0149] In one implementation of evaluating the first target neural network, it can include: initializing the network parameters in the first target neural network; inputting training data into the first target neural network to train the first target neural network; inputting test data into the trained first target neural network to obtain the evaluation result of the first target neural network.
[0150] In this embodiment, the number of training times of the first target neural network can be greater than that of the candidate neural network, the learning rate of each training of the first target neural network can be greater than that of the candidate neural network, and the training duration of the first target neural network can be less than the normal training duration of the candidate neural network. In this way, a target neural network with higher accuracy can be trained.
[0151] In this embodiment, after obtaining the N first target neural networks, in the first implementation, a group normalization (GN) layer can be added after each convolutional layer and / or each fully connected layer in each target sub-network of the first target neural network to obtain a second target neural network corresponding to the first target neural network. The performance and training speed of the second target neural network will be improved compared with the first target neural network. Among them, if there is originally a batch normalization (BN) layer in the target sub-network, the BN layer can be replaced with a GN layer.
[0152] For example, the first target neural network is a convolutional neural network for performing computer vision tasks, and the convolutional neural network is a neural network composed of a backbone network module, a multi-level feature extraction module, and a prediction module. The GN layer can be used to replace the BN layer in the backbone network module, and GN layers can be added after each convolutional layer and each fully connected layer in the multi-level feature extraction module and the prediction module, so as to obtain the corresponding second target neural network.
[0153] Since computer vision tasks require large-sized input images and are limited by the video memory capacity of the graphics processing unit (GPU) used for training, a relatively small input batch (i.e., fewer images are input at one time) is usually adopted during the training process. This will result in inaccurate statistics (mean and variance) of the input data estimated by using BN-related strategies, thereby reducing the accuracy of the first target neural network after training. GN is insensitive to the batch size, so it can better estimate the statistics of the input data, thereby improving the performance of the second target neural network and accelerating its training speed.
[0154] In the embodiments of this application, after obtaining N first target neural networks, in the second implementation manner, the weights (weight standardization, WS) of all convolutional layers in each first target neural network can be standardized, so as to obtain the corresponding second target neural network. That is to say, in addition to standardizing the activation function, the weights of the convolutional layers are also standardized to accelerate the training speed and avoid dependence on the input batch size.
[0155] Standardizing the weights of the convolutional layers can also be referred to as normalizing the convolutional layers. For example, the convolutional layers can be normalized through the following formula:
[0156]
[0157] I = C in ×K
[0158] Where, represents the weight matrix of the convolutional layer, * represents the convolution operation, O represents the number of output channels, C in represents the number of input channels, I represents the number of input channels within the convolutional kernel region for each output channel, x represents the input of the convolutional layer, y represents the output of the convolutional layer, represents the weight on the input channel within the jth convolutional kernel region corresponding to the ith output channel; K represents the convolutional kernel size.
[0159] For example, when the first target neural network is a convolutional neural network for performing computer vision tasks, multiple loss functions usually need to be optimized during the training process of the convolutional neural network. For example, when the first target neural network is a convolutional neural network for object detection, it is necessary to optimize the classification loss function of foreground and background in the region proposal network, the bounding box regression loss function, as well as the classification loss function of specific categories and the bounding box regression loss function in the head prediction network. The complexity of these loss functions will hinder the backpropagation of the gradients of the loss functions to the backbone network. By normalizing the weights in the convolutional layer, each loss function can be made smoother, which helps the gradients of the loss functions to backpropagate to the backbone network, thereby improving the performance of the corresponding second target neural network and its training speed.
[0160] In the embodiments of the present application, after obtaining N first target neural networks, in the third implementation manner, it is possible to not only normalize the weights of all convolutional layers in each first target neural network, but also add a combined regularization layer after each convolutional layer and each fully connected layer in each target sub-network of the first target neural network.
[0161] In this embodiment, after obtaining N second target neural networks, the evaluation results of these N second target neural networks can be obtained. The obtaining method can refer to the obtaining method of the evaluation results of the first target neural network, which will not be elaborated here.
[0162] In this embodiment, after obtaining the candidate neural network and the evaluation result of the candidate neural network, the Pareto optimal set of the candidate neural network can be updated according to the evaluation result.
[0163] When the evaluation result of the candidate neural network includes the running speed and the prediction accuracy, a two-dimensional spatial coordinate system is constructed with the running speed as the abscissa and the prediction accuracy as the ordinate. Then, the spatial position relationship of multiple candidate neural networks obtained by repeatedly executing S120 and S130 is as Figure 6 shown. Among them, a point represents the evaluation result of a candidate neural network, the dotted line represents the Pareto front of multiple candidate neural networks, the candidate neural networks located on the dotted line are the Pareto optimal solutions, and the set of all candidate neural network combinations located on the dotted line is the Pareto optimal set.
[0164] After each time a new candidate neural network and its evaluation result are determined, according to the spatial position relationship between the evaluation result and the evaluation results of the previous candidate neural networks, the Pareto front of the candidate neural network is re-determined, that is, the Pareto optimal set of the candidate neural network is updated.
[0165] In some implementations, the evaluation result of a candidate neural network that is a Pareto optimal solution can be considered as an evaluation result that meets the task requirements, and thus the target neural network can be further determined based on this candidate neural network.
[0166] In other implementations, one or more Pareto optimal solutions can be selected from the Pareto optimal set, and the evaluation results of these one or more Pareto optimal solutions are considered as evaluation results that meet the task requirements. For example, when the task requirement is that the running speed of the first target neural network is less than a certain threshold, the evaluation result of the first candidate neural network in the Pareto optimal set whose running speed is less than this threshold is the evaluation result that meets the task requirements.
[0167] For a candidate neural network that meets the task requirements, construct the target search space for each candidate sub-network in this candidate neural network, and search for the target sub-network corresponding to this candidate sub-network from the target search space of each candidate sub-network. Each target sub-network obtained by searching in multiple target search spaces constitutes the first target neural network.
[0168] In this embodiment, the steps in Figure 3 can be executed in parallel for multiple candidate neural networks to obtain multiple target neural networks corresponding to these multiple candidate neural networks. This can save search time and improve search efficiency.
[0169] Next, in combination with Figure 7 an exemplary flowchart of the method for determining a neural network according to the present application is introduced.
[0170] S701, Prepare task data. Specifically, accurate training data and test data.
[0171] S702, Initialize the initial search space and initial search parameters.
[0172] Among them, the implementation manner of initializing the initial search space can refer to the implementation manner of determining the initial search space described above, and will not be elaborated here.
[0173] Among them, the initial search parameters include the training parameters when training each candidate neural network. For example, the initial search reference can include the number of training times, learning rate, and / or training duration, etc. for each candidate neural network.
[0174] S703, Sample candidate neural networks. The implementation manner of this step can refer to the implementation manner of determining candidate neural networks according to multiple initialized search spaces described above, and will not be elaborated here.
[0175] S704, Performance evaluation. The implementation manner of this step can refer to the implementation manner of evaluating candidate neural networks described above, and will not be elaborated here.
[0176] S705, Update the Pareto front. This step can refer to the implementation method of updating the Pareto front described above and will not be elaborated here.
[0177] S706, Determine whether the termination condition is met. If yes, repeat S703; otherwise, execute S707. When the termination condition is met, multiple candidate neural networks can be searched for.
[0178] For example, when the difference between the evaluation results of the current candidate neural network and the previous candidate neural network is less than or equal to a preset threshold, it is determined that the termination condition is met.
[0179] S707, Pareto front screening. That is, select n candidate neural networks from the Pareto front obtained in S705. These n candidate neural networks are E1 to En in sequence. Then, execute S708 to S712 in parallel for these n candidate neural networks.
[0180] For example, select n candidate neural networks with a running speed less than or equal to a preset threshold from the Pareto front obtained in S705.
[0181] Then, for each of the n selected candidate neural networks, execute the Figure 8 method in.
[0182] S808, Initialize the target search space and target search parameters.
[0183] Among them, the implementation method of initializing the target search space can refer to the implementation method of determining the target search space described above and will not be elaborated here.
[0184] Among them, the target search parameters include the training parameters when training each first target neural network. For example, the target search reference can include the number of training times, learning rate, and / or training duration, etc. for each first target neural network.
[0185] S809, Sample the first target neural network. The implementation method of this step can refer to the implementation method of determining the first target neural network according to multiple targeted search spaces described above and will not be elaborated here.
[0186] S810, Performance evaluation. The implementation method of this step can refer to the implementation method of evaluating the first target neural network described above and will not be elaborated here.
[0187] S811, Update the Pareto front. Regard the first target neural network as a candidate neural network, and update the Pareto front of the n candidate neural networks screened in S707 according to the evaluation result of the first target neural network. The specific update method refers to the foregoing content and will not be elaborated here.
[0188] S812, Determine whether the termination condition is met. If yes, repeat S809; otherwise, execute S813.
[0189] For example, when the difference between the current first target neural network and the evaluation result of the first target neural network obtained by the previous execution of S809 is less than or equal to a preset threshold, it is determined that the termination condition is met.
[0190] Taking Figure 6 the Pareto front shown as an example, after the termination condition is met, the finally updated Pareto front is as shown by the solid line in Figure 11 . As shown in Figure 11 , for the target neural network corresponding to the finally updated Pareto front, under the constraint of the same running speed, the prediction accuracy is better.
[0191] S813, Output the first target neural network. In addition, the evaluation results of these n first target neural networks can also be output.
[0192] For example, output the first target neural network corresponding to the Pareto front updated by S811.
[0193] The following introduces the structures and related information of 6 exemplary first target neural networks (E1 to E6) obtained by using the method of this application in combination with Table 1.
[0194] Table 1 Network Structure and Related Information Table of the First Target Neural Network
[0195]
[0196]
[0197] In Table 1, mAP represents the average accuracy of the target detection prediction results. For the backbone network module, the first placeholder is the selection of the convolutional module; the second is the number of basic channels; "-" separates each stage at different resolutions, and the resolution of the latter stage is halved compared to the previous stage; "1" represents a conventional block without changing the channels, and "2" represents that the number of basic channels in this block is doubled. For the network structure (Neck) of the multi-level feature extraction module, P1 - P5 represent the selected feature levels from the backbone network module and "c" represents the number of channels of the output of the Neck; for the RCNN head; "2FC" is two shared fully connected layers; "n" represents the number of cascades of the prediction head network; the time is the processing time after each picture is input into the first target neural network, in milliseconds (ms); the unit of the floating-point operations per second of the backbone network module is Giga (G).
[0198] The following describes the experimental results of the second target neural network obtained by normalizing the convolutional layer weights of the first target neural network and adding a combined regularization layer after each convolutional layer and fully connected layer in the first target neural network, in conjunction with Table 2.
[0199] Table 2 Performance table of neural networks obtained by different training methods
[0200] Training method Epoch Batch Learning rate mAP BN 12 2*8 0.02 24.8 BN 12 8*8 0.20 28.3 GN 12 2*8 0.02 29.4 GN + WS 12 4*8 0.02 30.7
[0201] Among them, the backbone network module of the first target neural network is a ResNet-50 structure, the multi-level feature extraction module is a feature pyramid network, and the head prediction module is two-layer FC. Moreover, different strategies are used to conduct effectiveness analysis experiments on the first target neural network, and the evaluation is carried out on the COCO (common objects in context) dataset. The COCO dataset is constructed by the Microsoft team and is a well-known dataset in the field of object detection; Epoch is the number of training epochs (traversing the training subset once represents one training epoch), Batch Size is the input batch size, and Experiments 1 to 2 follow the training procedures of the standard detection model and are trained for 12 epochs respectively. By comparing Experiments 1, 2, and 3, it can be found that a smaller input batch size will lead to incorrect estimation of the statistics of the input data, resulting in a decrease in accuracy; while using group regularization can alleviate this problem and increase the mAP from 24.8% to 29.4%. According to the comparison between Experiment 3 and Experiment 4, it is found that adding WS can further smooth the training process and increase the mAP by 1.3%. Therefore, the method of training the detection network from scratch can end the training earlier than the method of using the parameters pre-trained on ImageNet as initialization.
[0202] Figure 9 is an exemplary structural diagram of the device for training a neural network in this application. The device 900 includes an acquisition module 910, a determination module 920, and an evaluation module 930. The device 900 can implement the foregoing Figure 1 、 Figure 5 or Figure 7 shown methods.
[0203] For example, the acquisition module 910 is used to execute S110, the determination module 220 is used to execute S120 and S140, and the evaluation module 930 is used to execute S130.
[0204] The device 900 can be deployed in a cloud environment, which is an entity that uses basic resources to provide cloud services to users under the cloud computing model. The cloud environment includes a cloud data center and a cloud service platform. The cloud data center includes a large number of basic resources (including computing resources, storage resources, and network resources) owned by the cloud service provider. The computing resources included in the cloud data center can be a large number of computing devices (such as servers). The device 900 can be a server in the cloud data center for training a neural network. The device 900 can also be a virtual machine created in the cloud data center for training a neural network. The device 900 can also be a software device deployed on a server or a virtual machine in the cloud data center. This software device is used to train a neural network and can be deployed distributively on multiple servers, or distributively on multiple virtual machines, or distributively on virtual machines and servers. For example, the acquisition module 910, the determination module 920, and the evaluation module 930 in the device 900 can be deployed distributively on multiple servers, or distributively on multiple virtual machines, or distributively on virtual machines and servers. Another example is that when the determination module 920 includes multiple sub-modules, these multiple sub-modules can be deployed on multiple servers, or distributively on multiple virtual machines, or distributively on virtual machines and servers.
[0205] The device 900 can be abstracted by the cloud service provider into a cloud service for determining a neural network on the cloud service platform and provided to users. After the user purchases this cloud service on the cloud service platform, the cloud environment uses this cloud service to provide the user with the cloud service for determining a neural network. The user can upload task requirements to the cloud environment through an application program interface (API) or through the web interface provided by the cloud service platform. The device 900 receives the task requirements, determines the neural network for implementing this task, and the finally obtained neural network is returned by the device 900 to the edge device where the user is located.
[0206] When the device 900 is a software device, the device 900 can also be deployed alone on a computing device in any environment.
[0207] This application also provides a device 1000 as shown in Figure 10 The device 1000 includes a processor 1002, a communication interface 1003, and a memory 1004. An example of the device 1000 is a chip. Another example of the device 1000 is a computing device.
[0208] Communication can occur between the processor 1002, the memory 1004, and the communication interface 1003 via a bus. Executable code is stored in the memory 1004, and the processor 1002 reads the executable code in the memory 1004 to execute the corresponding method. The memory 1004 may also include other software modules required for running processes such as an operating system. The operating system can be LINUX TM , UNIX TM , WINDOWS TM , etc.
[0209] For example, the executable code in the memory 1004 is used to implement the Figure 1 method shown, and the processor 1002 reads the executable code in the memory 1004 to execute the Figure 1 method shown.
[0210] Among them, the processor 1002 can be a central processing unit (CPU). The memory 1004 may include volatile memory, such as random access memory (RAM). The memory 1004 may also include non-volatile memory (2non-volatile memory, 2NVM), such as read-only memory (2read-only memory, 2ROM), flash memory, a hard disk drive (HDD), or a solid state disk (SSD).
[0211] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0212] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0213] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0214] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0215] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0216] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0217] As described above, the above are only the specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for determining a neural network, characterized in that, Including: Obtain a plurality of initial search spaces, where the initial search spaces include one or more neural networks, the functions of the neural networks in any two of the initial search spaces are different, and the functions of any two neural networks in the same initial search space are the same and the network structures are different; Determine M candidate neural networks according to the plurality of initial search spaces, where the candidate neural networks include a plurality of candidate sub-networks, the plurality of candidate sub-networks belong to the plurality of initial search spaces, and any two of the plurality of candidate sub-networks belong to different initial search spaces, and M is a positive integer; Evaluate the M candidate neural networks based on input data to obtain M evaluation results, where the input data is related to an image; According to the M evaluation results, determine N candidate neural networks from the M candidate neural networks, and determine N first target neural networks according to the N candidate neural networks, where the first target neural networks are used for object detection, image classification or image segmentation based on the input data, the N first target neural networks and the N candidate neural networks are in one-to-one correspondence, each of the N candidate neural networks includes a plurality of candidate sub-networks, each of the N first target neural networks includes a plurality of target sub-networks, and each target sub-network included in each first target neural network includes the same blocks as each sub-network included in the corresponding candidate neural network. N is a positive integer less than or equal to M.
2. The method according to claim 1, wherein The evaluation results of the candidate neural networks include one or more of the following: running speed, accuracy, number of parameters or floating-point operation count.
3. The method according to claim 2, wherein The evaluation results of the candidate neural networks include running speed and accuracy; Among them, the step of determining N candidate neural networks from the M candidate neural networks according to the M evaluation results and determining N first target neural networks according to the N candidate neural networks includes: Determine the N candidate neural networks from the M candidate neural networks according to the M evaluation results, where the N candidate neural networks are the Pareto optimal solutions of the M candidate neural networks when running speed and accuracy are the goals; Determine the N first target neural networks according to the N candidate neural networks.
4. The method according to claim 3, characterized in that, The step of determining the N first target neural networks according to the N candidate neural networks includes: Determine a plurality of target search spaces according to the plurality of candidate sub-networks of the i-th candidate neural network, the plurality of target search spaces are in one-to-one correspondence with the plurality of candidate sub-networks of the i-th candidate neural network, each of the plurality of target search spaces includes one or more neural networks, and each neural network included in each of the plurality of target search spaces includes the same blocks as the candidate sub-network corresponding to each of the plurality of target search spaces; Determine the $i$-th first target neural network according to the multiple target search spaces, where multiple target sub-networks in the $i$-th first target neural network belong to the multiple target search spaces, and any two target sub-networks in the multiple target sub-networks of the $i$-th first target neural network belong to different target search spaces, and $i$ is a positive integer less than or equal to $N$.
5. The method according to any one of claims 1 to 4, characterized in that The method further includes: Determine $N$ second target neural networks according to the $N$ first target neural networks, where the $i$-th second target neural network in the $N$ second target neural networks is obtained by performing one or more of the following processes on the $i$-th first target neural network: adding a combined regularization layer after the convolutional layer in the target sub-network of the $i$-th first target neural network, adding a combined regularization layer after the fully connected layer in the target sub-network of the $i$-th first target neural network, and normalizing the weights of the convolutional layer in the target sub-network of the $i$-th first target neural network.
6. The method according to claim 5, characterized in that, The method further includes: Evaluate the $N$ second target neural networks to obtain evaluation results of the $N$ second target neural networks.
7. The method according to claim 6, characterized in that, The evaluating the $N$ second target neural networks to obtain the evaluation results of the $N$ second target neural networks includes: Randomly initialize the network parameters in the $i$-th second target neural network; Train the $i$-th second target neural network according to training data; Test the trained $i$-th second target neural network according to test data to obtain the evaluation result of the trained $i$-th second target neural network.
8. The method according to any one of claims 1 to 4, characterized in that, The first target neural network is used for object detection. Among them, the multiple initial search spaces include a first initial search space, a second initial search space, a third initial search space, and a fourth initial search space. The first initial search space includes at least one of residual networks with different depths, second-generation residual networks with different depths, and mobile networks with different depths. The second initial search space includes connection paths of features at different levels. The third initial search space includes at least one of a general region candidate network and an anchor-guided region candidate network. The fourth initial search space includes at least one of a one-stage detection head network, a fully connected detection head network, a fully convolutional detection head network, and a cascaded detection head network.
9. The method according to any one of claims 1 to 4, characterized in that The first target neural network is used for image classification. Among them, the multiple initial search spaces include a first initial search space and a second initial search space. The first initial search space includes at least one of residual networks with different depths, second-generation residual networks with different depths, and dense connection networks with different widths. The neural network in the second initial search space includes a fully connected layer.
10. The method according to any one of claims 1 to 4, characterized in that, The first target neural network is used for image segmentation. Among them, the multiple initial search spaces include a first initial search space, a second initial search space, and a third initial search space. The first initial search space includes at least one of residual networks with different depths, second-generation residual networks with different depths, and high-resolution networks with different widths. The second initial search space includes at least one of a dilated spatial pyramid pooling network, a pyramid pooling network, and a network including a dense prediction unit. The third initial search space includes at least one of a U-Net model and a fully convolutional network.
11. An apparatus for determining a neural network, characterized in that, Comprising: An acquisition module, configured to acquire multiple initial search spaces, where the initial search space includes one or more neural networks, the functions of the neural networks in any two of the initial search spaces are different, and the functions of any two neural networks in the same initial search space are the same and the network structures are different; A determination module, configured to determine M candidate neural networks according to the multiple initial search spaces. The candidate neural networks include multiple candidate sub-networks, the multiple candidate sub-networks belong to the multiple initial search spaces, and any two of the multiple candidate sub-networks belong to different initial search spaces, where M is a positive integer; An evaluation module, configured to evaluate the M candidate neural networks based on input data to obtain M evaluation results, where the input data is related to images; The determination module is further configured to: according to the M evaluation results, determine N candidate neural networks from the M candidate neural networks, and determine N first target neural networks according to the N candidate neural networks. Among them, the first target neural network is used for object detection, image classification, or image segmentation based on the input data. The N first target neural networks correspond to the N candidate neural networks one by one. Each of the N candidate neural networks includes multiple candidate sub-networks, and each of the N first target neural networks includes multiple target sub-networks. Each target sub-network in each first target neural network includes the same blocks as each sub-network in the corresponding candidate neural network. N is a positive integer less than or equal to M.
12. The device according to claim 11, characterized in that, The evaluation results of the candidate neural networks include one or more of the following: running speed, accuracy, number of parameters, or floating-point operation count.
13. The device according to claim 12, wherein The evaluation results of the candidate neural networks include running speed and accuracy; Among them, the determination module is specifically configured to: According to the M evaluation results, determine the N candidate neural networks from the M candidate neural networks. The N candidate neural networks are the Pareto optimal solutions of the M candidate neural networks when running speed and accuracy are the goals; Determine the N first target neural networks according to the N candidate neural networks.
14. The device according to claim 13, characterized in that, The determination module is specifically configured to: Determine a plurality of target search spaces according to a plurality of candidate sub-networks of the i-th candidate neural network, where the plurality of target search spaces correspond one-to-one to the plurality of candidate sub-networks of the i-th candidate neural network, each target search space in the plurality of target search spaces includes one or more neural networks, and the blocks included in each neural network in each target search space are the same as the blocks included in the candidate sub-network corresponding to each target search space; Determine the i-th first target neural network according to the plurality of target search spaces, where the plurality of target sub-networks in the i-th first target neural network belong to the plurality of target search spaces, and any two target sub-networks in the plurality of target sub-networks of the i-th first target neural network belong to different target search spaces.
15. The device according to any one of claims 11 to 14, characterized in that, The determining module is further configured to: Determine N second target neural networks according to the N first target neural networks, where the i-th second target neural network in the N second target neural networks is obtained by performing one or more of the following processes on the i-th first target neural network: adding a combined regularization layer after the convolutional layer in the target sub-network of the i-th first target neural network, adding a combined regularization layer after the fully connected layer in the target sub-network of the i-th first target neural network, and normalizing the weights of the convolutional layer in the target sub-network of the i-th first target neural network.
16. The device according to claim 15, characterized in that, The evaluation module is further configured to: Evaluate the N second target neural networks to obtain evaluation results of the N second target neural networks.
17. The device according to claim 16, characterized in that, Specifically, the evaluation module is configured to: Randomly initialize the network parameters in the i-th second target neural network; Train the i-th second target neural network according to training data; Test the trained i-th second target neural network according to test data to obtain the evaluation result of the trained i-th second target neural network.
18. The device according to any one of claims 11 to 14, characterized in that, The first target neural network is used for object detection, where the plurality of initial search spaces include a first initial search space, a second initial search space, a third initial search space, and a fourth initial search space. The first initial search space includes at least one of residual networks with different depths, second-generation residual networks with different depths, and mobile networks with different depths. The second initial search space includes connection paths of features at different levels. The third initial search space includes at least one of a general region candidate network and an anchor-guided region candidate network. The fourth initial search space includes at least one of a one-stage detection head network, a fully-connected detection head network, a fully convolutional detection head network, and a cascaded detection head network.
19. The device according to any one of claims 11 to 14, characterized in that The first target neural network is used for image classification, where the plurality of initial search spaces include a first initial search space and a second initial search space. The first initial search space includes at least one of residual networks with different depths, second-generation residual networks with different depths, and dense connection networks with different widths. The neural network in the second initial search space includes a fully connected layer.
20. The device according to any one of claims 11 to 14, characterized in that, The first target neural network is used for image segmentation. Among them, the multiple initial search spaces include a first initial search space, a second initial search space, and a third initial search space. The first initial search space includes at least one of residual networks with different depths, second-generation residual networks with different depths, and high-resolution networks with different widths. The second initial search space includes at least one of a dilated spatial pyramid pooling network, a pyramid pooling network, and a network including a dense prediction unit. The third initial search space includes at least one of a U-Net model and a fully convolutional network.
21. An apparatus for determining a neural network, characterized in that, Comprising: a memory for storing a program; a processor for executing the program stored in the memory, and when the program stored in the memory is executed, implementing the method according to any one of claims 1 to 10.
22. A computer-readable storage medium, characterized in that, The computer-readable medium stores instructions for a computing device to execute, and when the computing device executes the instructions, implementing the method according to any one of claims 1 to 10.
23. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, implementing the method according to any one of claims 1 to 10.
24. A chip, characterized in that, Comprising a processor and a data interface, the processor reads the instructions stored on the memory through the data interface and executes the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
A structure searching method and device of a depth neural network
CN109284820A
An image small target detection method based on combination of two-stage detection
CN109598290A