Neural Architecture Search Device, Neural Architecture Search Method, and Program
By constructing a super network with fully connected layers and training candidate layers in parts, the neural architecture search process is optimized, addressing the challenges of time-consuming training and structural gaps between supernet and subnet models.
Patent Information
- Application Number
- JP2024566342
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-05-16
AI Technical Summary
The challenge in neural architecture search (NAS) is the time-consuming process of training large-scale supernet models and the significant structural gap between the supernet and the subnet, which affects efficiency and accuracy.
The proposed solution involves constructing a super network with fully connected layers, training the candidate layers and fully connected layers in parts, and selecting the best-performing parts to optimize the neural architecture search process.
This approach significantly reduces the training time of the supernet and narrows the structural gap between the supernet and the subnet, leading to a more efficient neural architecture search for computer vision tasks.
Smart Images

Figure 2025517163000001_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a neural architecture search device, a neural architecture search method, and a program.
Background Art
[0002] In the past few decades, convolutional neural network (CNN) models have become state-of-the-art solutions for computer vision tasks such as image classification, object detection, and semantic segmentation. The main reason for the success of CNN models is the ability to achieve high accuracy. In real-time applications, the time taken for the execution of a CNN model, generally referred to as the execution time, is also very important.
[0003] While CNN models that achieve high accuracy tend to have several CNN layers, CNN models that achieve high speed (i.e., short execution time) tend to have fewer CNN layers. Therefore, there is a trade-off between accuracy and speed with respect to the number of CNN layers employed in a CNN model. Furthermore, there are several hyperparameters related to CNN layers, such as kernel size, input channels, and output channels. Manually optimizing each hyperparameter for each layer is a time-consuming task that requires a lot of human expertise.
[0004] Recently, an efficient technique for such problems, namely neural architecture search (NAS), has been developed. The NAS technique generally includes three steps. First, as shown in Figure 2, a network consisting of several candidate CNN layers is constructed. A large-scale network with multiple CNN layer candidates is called a SuperNet. As the first step of NAS, the SuperNet is trained on a dataset. Next, as the second step, the SuperNet is intelligently pruned to become a smaller network with fewer CNN layers for the purpose of minimizing the degradation of accuracy. A smaller network with fewer CNN layers is called a SubNet. Finally, in the third step, the SubNet is further trained on the dataset to recover the accuracy.
[0005] The CNN model for the object detection task mainly consists of three blocks, namely, a backbone block, a neck block, and a head block. The main task of the backbone block is to perform shallow-level feature extraction from the input image, the neck block performs deeper-level feature extraction, and the head block performs the task of predicting labels based on the features extracted by the backbone and the neck blocks. The NAS method can be applied to one or more blocks. NAS relaxes the requirement for human expertise in designing CNN models. However, a concern of the NAS method is that it takes time for the training of the SuperNet and the search and training of the optimal SubNet. To address this concern, in Non-Patent Document 1, special architecture parameters that are also trained during the training of the SuperNet are introduced. By using the special architecture parameters, the pruning of the SuperNet can be performed quickly, and by performing the second step quickly, the NAS method can be accelerated. However, the time required to train a large-scale SuperNet is very long, resulting in a large delay in obtaining the final SubNet.
Prior Art Documents
Non-Patent Documents
[0006] Fast Neural Network Adaptation via Parameter Remapping and Architecture Search, Jiemin Fang, Yuzhu Sun, Kangjian Peng, Qian Zhang, Yuan Li, Wenyu Liu, Xinggang Wang, https: / / arxiv.org / abs / 2001.02525
Summary of the Invention
Problems to be Solved by the Invention
[0007] The problem of NAS is how to train the supernet model. The supernet model is composed of multiple layers, and each layer has several candidates for convolutional layers, so the size of the CNN model becomes large. It takes time to train such a large-sized CNN model.
[0008] Another problem is that the structural gap between the supernet and the subnet is large. In the supernet, multiple candidates existing in each layer are trained under the condition that there are multiple parallel layers in the previous and subsequent layers, while in the subnet, there is only one or fewer CNN layers in the previous and subsequent layers.
[0009] One embodiment of the present invention is made in view of such problems, and its purpose is to provide time-efficient neural architecture search for the backbone blocks of computer vision tasks.
Means for Solving the Problems
[0010] To achieve the above object, the neural architecture search device of the present invention includes a construction means for constructing a super network, wherein the target layer for optimizing the super network is replaced by a plurality of candidate layers, and the super network includes a plurality of fully connected layers; a training means for training the super network, wherein the plurality of candidate layers are trained in parts, and the plurality of fully connected layers are trained corresponding to the parts of the plurality of candidate layers; and a selection means for evaluating the trained super network and selecting the parts of the plurality of candidate layers corresponding to the parts of the plurality of fully connected layers with the best performance.
[0011] To achieve the above object, the neural architecture search method of the present invention includes a step of constructing a super network, wherein the target layer for optimizing the super network is replaced by a plurality of candidate layers, and the super network includes a plurality of fully connected layers; a step of training the super network, wherein the plurality of candidate layers are trained in parts, and the plurality of fully connected layers are trained corresponding to the parts of the plurality of candidate layers; and a step of evaluating the trained super network and selecting the parts of the plurality of candidate layers corresponding to the parts of the plurality of fully connected layers with the best performance.
[0012] Also, to achieve the above object, the program causes a computer to function as a neural architecture search device, and the program causes the computer to function as construction means, training means, and selection means.
Advantages of the Invention
[0013] According to an exemplary aspect of the present invention, it is possible to provide a time-efficient neural architecture search for a backbone block of a computer vision task.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
[0015] In the following description, details of the first embodiment according to the present invention will be described with reference to the drawings. Note that the first embodiment is a basic embodiment for subsequent embodiments.
[0016] In the first embodiment, a neural architecture search device and a neural architecture search method will be described with reference to FIGS. 1 to 3.
[0017] (Configuration of Neural Architecture Search Device) Hereinafter, with reference to FIG. 1, the configuration of the neural architecture search device 1 according to the first embodiment will be described. FIG. 1 is a block diagram showing the configuration of the neural architecture search device 1. As shown in FIG. 1, the neural network architecture search device 1 includes a construction unit 11, a training unit 12, and a selection unit 13.
[0018] The neural architecture search device 1 trains a large-scale network (referred to as a supernetwork) including a plurality of candidate network layers that are candidates for optimization. After the training process and the optimization process, the neural architecture search device 1 outputs a pruned network (also referred to as a subnetwork), which is a network smaller than the supernetwork. Further, the neural network search device 1 may train the subnetwork.
[0019] The feature construction unit 11 is an example of the construction means described in the claims. The training unit 12 is an example of the training means described in the claims. The selection unit 13 is an example of the selection means described in the claims.
[0020] The construction unit 11 constructs a supernetwork. Here, as described above, the supernetwork is a neural network including a plurality of candidate network layers. The candidate network layer is a layer that is optimized in the optimization process by the neural architecture search device 1.
[0021] FIG. 2 is a schematic diagram showing the configuration of the supernetwork SN. As shown in FIG. 2, the supernetwork SN includes one or more target layers (TL) and one or more non-target layers (NTL1, NTL2,...). As shown in FIG. 2, the optimization target layer of the supernetwork SN is replaced by a plurality of candidate layers (CL1, CL2, CL3,...).
[0022] Furthermore, as shown in FIG. 2, the super network SN includes a plurality of fully connected layers (FCL1, FCL2, FCL3, ...). Each of these fully connected layers may correspond to each of the candidate layers (CL1, CL2, CL3, ...).
[0023] When the super network SN is trained, each of the candidate layers (CL1, CL2, CL3, ...) included in the target layer (TL) is trained together with the corresponding fully connected layer (FCL1, FCL2, FCL3, ...).
[0024] That is, the training unit 12 trains the super network. Here, a plurality of candidate layers (CL1, CL2, CL3, ...) are trained in parts, and a plurality of fully connected layers (FCL1, FCL2, FCL3, ...) are trained corresponding to the parts of the plurality of candidate layers.
[0025] When the super network SN is trained, a predefined loss function is used to evaluate the loss for the super network SN. The candidate layers (CL1, CL2, CL3, ...) and the fully connected layers (FCL1, FCL2, FCL3, ...) are selected with reference to the evaluation.
[0026] That is, the selection unit 13 evaluates the trained super network and selects the part corresponding to the best performing part of the plurality of fully connected layers (FCL1, FCL2, FCL3, ...) among the plurality of candidate layers (CL1, CL2, CL3, ...).
[0027] (Flow of Neural Architecture Search Method) Hereinafter, with reference to FIG. 3, the flow of the neural architecture search method according to the first embodiment will be described. FIG. 3 is a flowchart showing the flow of the neural architecture search method S1. As shown in FIG. 3, the flow of the neural architecture search method includes steps S11 to S13.
[0028] In step S11, the construction unit 11 of the neural architecture search device 1 constructs a super network SN. The super network is a neural network including a plurality of candidate network layers. The candidate network layer is a layer that is optimized in the optimization process by the neural architecture search device 1. That is, the neural architecture search method S1 includes a step of constructing a super network, and the optimization target layer of the super network is replaced by a plurality of candidate layers (CL1, CL2, CL3,...), and the super network includes a plurality of fully connected layers (FCL1, FCL2, FCL3,...).
[0029] In step S12, the training unit 12 trains the super network as described above. That is, the neural architecture search method S1 includes a step of training the super network by training a plurality of candidate layers (CL1, CL2, CL3,...) for each part and training a plurality of fully connected layers (FCL1, FCL2, FCL3,...) corresponding to the parts of the plurality of candidate layers (CL1, CL2, CL3,...).
[0030] In step S13, the selection unit 13 evaluates the trained super network and selects a part of the plurality of candidate layers (CL1, CL2, CL3,...). That is, the neural architecture search means S1 includes a step of evaluating the trained super network and selecting a part of the plurality of candidate layers corresponding to the part with the best performance among the plurality of fully connected layers.
[0031] (Advantageous effects of the first embodiment) As described above, according to the first embodiment, when training the super network, a plurality of candidate layers are trained part by part, and a plurality of fully connected layers are trained corresponding to the parts of the plurality of candidate layers. By doing so, compared with the case of training all layers simultaneously, it can be trained in a shorter time. As a result, the training time of the super network becomes very efficient. In this way, it is possible to realize a time-efficient neural architecture search for the backbone block of the computer vision task.
[0032] <Second Embodiment> In the following description, the details of the second embodiment of the present invention will be described with reference to the drawings.
[0033] (Configuration of Model Training System Based on Neural Architecture Search) Hereinafter, with reference to FIG. 4, the configuration of the CNN model training system 100 based on neural architecture search according to the second embodiment will be described. FIG. 4 is a block diagram showing the configuration of the model training system 100 based on neural architecture search. The CNN model is a super network. Hereinafter, the "super network" may also be referred to as "SuperNet". In the second embodiment, the model based on neural architecture search trained by the model training system 100 based on neural architecture search is used for at least one of the object detection task and the object classification task.
[0034] As shown in FIG. 4, the model based on neural architecture search includes a training data set 200 for object detection tasks, a super network builder 300, a super network trainer 400 with object detection and classification tasks, and a neural architecture selector 500.
[0035] The training dataset 200 for the object detection task is a dataset provided to train the CNN model of the model training system 100 based on neural architecture search for use in the object detection task. The training dataset for the object detection task 200 includes images and labels. The images are the input, and the labels are the predicted values intended to be generated as output by the supernet and subnet CNN models.
[0036] The supernet builder 300 corresponds to the construction unit 11 in the first embodiment. The supernet trainer 400 with the object detection task and classification task corresponds to the training unit 12 in the first embodiment. Also, the neural architecture selection unit 500 corresponds to the selection unit 13 in the first embodiment.
[0037] As described in the first embodiment, the supernet trainer 400 with the object detection task and classification task trains the supernet. Here, a plurality of candidate layers are trained in parts, and a plurality of fully connected layers are trained corresponding to the parts of the plurality of candidate layers. In this embodiment, more specifically, the case of training a plurality of candidate layers and corresponding fully connected layers one by one will be described.
[0038] The model training system 100 based on neural architecture search further includes a dataset transformer 600, a training dataset 700 for the classification task, and an optimized CNN model 800.
[0039] The training dataset 700 for the classification task includes images and labels. The images are the input, and the labels are the categories of the objects present in each image. The labels are the predicted values intended to be generated as output by the supernet during training based on the classification task.
[0040] The optimized CNN model 800 is obtained by training the supernet trainer 400 with the object detection task and classification task and the selection by the neural architecture selector 500.
[0041] (Data Transformer 600) The dataset transformer 600 is a functional block that functions as a conversion means for converting a training dataset for an object detection task into a training dataset for a classification task. The training dataset 700 for the classification task is a dataset for training the CNN model of the CNN model training system 100 based on the neural architecture search used in the classification task.
[0042] The dataset transformer 600 includes means for receiving an object detection dataset, means for converting the object detection dataset into a classification dataset, and means for providing the classification dataset as an output. Since the means necessary for converting the object detection dataset into a classification dataset are merely engineering operations, they are not described in this specification.
[0043] Thus, the model training system 100 based on neural architecture search includes conversion means for converting an object detection dataset into a classification dataset.
[0044] (Supernet CNN Model and Supernet Builder 300) FIG. 5 schematically shows an example of a supernet CNN model having a candidate search space. The supernet CNN model 900a includes a backbone 901a, a neck 902a, and a head 903a. The details of the backbone 901a are illustrated as blocks (blocks 904a, 905a,... 907a). For example, the backbone 901a includes N blocks.
[0045] Each block may be formed by various types of neural architectures such as Conv 3x3, SW 3x3, MAX, Skip, etc. In FIG. 5, only the details of block 905a are shown, but other blocks may also have this detailed structure. Conventionally, the training of a CNN model has been performed using various types of neural architectures for each block (blocks 904a, 905a,... 907a). However, such training is time-consuming and inefficient. In this embodiment, candidates for replacing each block are prepared.
[0046] FIG. 6 is a block diagram showing a supernet CNN model constructed by a supernet builder 300. The supernet builder 300 includes means for receiving a training dataset of an object detection task 200 and a supernet, and means for constructing the supernet. When the system 100 is executing an iteration other than the first iteration, it receives the supernet from a neural architecture selector 500. Then, the supernet builder 300 uses the pre-constructed supernet and modifies it. When the system 100 performs the first iteration, the supernet builder 300 constructs a supernet from scratch as shown in FIG. 6.
[0047] In FIG. 6, an image 2001 is input to the supernet CNN model 900.
[0048] F CNN layers are arranged in series and are called a fixed layer 904.
[0049] To the output of the last fixed layer 904, N CNN layers are arranged in series. These N CNN layers in series are B i which is called, and in FIG. 6, i is an index where 0 < i ≦ N. In this case, from layer B 1 to layer B N are arranged in series. To the output of layer B N are arranged M parallel FC (fully connected layer) blocks 9081 to 9083 of the same structure, and the construction of the backbone 901 of the supernet is completed.
[0050] Fig. 7 is a block diagram showing the internal structure of the FC block 9080. As shown in Fig. 7, the FC blocks 9081 to 9083 are basically one layer or multiple fully connected layers arranged in series.
[0051] Finally, B. N At the output of the layer, necks 902 and heads 903 are placed as shown in Fig. 6. The necks 902 and heads 903 are designed according to the requirements of the object detection task. Thus, the output of the head 903 is provided to the object detection output 2002. Meanwhile, the outputs of the FC blocks 9081-9083 are provided to the classification outputs 9091-9093, respectively.
[0052] In FIG. 6, an image 2001 is input to a supernet CNN model 900. The image 2001 is first input to a fixed layer 904 of a backbone 901. The fixed layer 904 is divided into multiple sublayers (B 1 1 , B 1 2 , ..., B 1 M ) are connected to multiple sub-layers (B 1 1 , B 1 2 , ..., B 1 M ) correspond to the multiple candidate layers (CL1, CL2, CL3, . . . ) in the first embodiment.
[0053] First, B 1 Start by optimizing the layers. 1 We replace the layer with M parallel CNN layers, also called sublayers, arranged as shown in Figure 6. The M parallel CNN sublayers are i j (0 <j≦M)として与えられる。M個のサブ層は、基本的にハイパーパラメータを変化させたCNN層のいくつかのバリエーションである。M個のサブ層のうち1個が、ニューラルアーキテクチャーセレクタ500によって実行される選択中に勝者として選択され、残りのM-1個のサブ層は脱落する。
[0054] As shown in FIG. 6, the output from the last fixed layer 904 is input to all M parallel sub-layers. The outputs of all M parallel sub-layers are combined, without limitation, by operations such as concatenation and summation, and the output is provided to layer B 2 layer. Layer B 2 to B N layers each have only one CNN layer for each iteration to train the respective M parallel sub-layers. Training will be described in detail later.
[0055] The layer for optimization is also called the target layer. In this case, layer B 1 is the target layer. The target layer may be changed from layer B 1 to layer B N layer. As described above, the output of layer B N is connected to M parallel FC (fully-connected layer) blocks 9081 to 9083. In this way, multiple fully-connected layers are connected to the output of the target layer or a layer deeper than the target layer.
[0056] B N layer is connected to the neck 902, and the neck 902 is connected to the head 903. The output of the head 903 is shown as the object detection output 2002. In this way, in the case of object detection, the neck 902, the head 903, and the N layers of the backbone 901 are used for prediction.
[0057] Also, layer B N is connected to M parallel FC blocks 9081 to 9083. The output of FC 1 block 9081 is shown as the classification output 9091. Similarly, the output of the FC 2 block and the output of FC 3 block 9081 are shown as the classification output 9092 and the classification output 9093. In this way, in the case of object classification, the FC blocks and the N layers of the backbone 901 are used for prediction.
[0058] (Supernet Trainer 400 with Object Detection Task and Classification Task) The super network trainer 400 with object detection task and classification task, also called STOC400, comprises means for receiving a training data set of the object detection task 200, a training data set of the classification task 200, and a training data set of the classification task 700, and means for receiving a super network from the super network builder 300. STOC400 also comprises means for performing training of the super network for the object detection task and the classification task. Finally, STOC400 also comprises means for outputting the trained super network output.
[0059] The basic function of STOC400 is to train all M sub-layers of layer B together with other B k (0 < k ≦ N; i ≠ k;) layers, and to train all FC blocks 9080 in the classification task. i The training of all sub-layers of layer B is performed one by one. Similarly, the training of all FC blocks is performed one by one. Also, the training of sub-layers and FC blocks is performed alternately. That is, first, sub-layer 9051 is trained in the object detection task, and then FC i block 9081 is trained in the classification task. Next, sub-layer 9052 of layer B is trained in the object detection task, and then FC i 1 block 9082 is trained in the classification task. Thus, STOC400 trains the super network for at least one of the object detection task using the object detection data set and the classification task using the classification data set. 1 i 2 2
[0060] (Neural architecture selector 500) The neural architecture selector 500 includes means for receiving, as inputs, a trained supernet and a classification dataset (the training dataset for the classification task 700), means for evaluating a loss function on the supernet, means for pruning the supernet, and finally means for outputting the pruned supernet. The pruned supernet is also called a subnet.
[0061] In this embodiment, the CNN model training system 100 based on neural architecture search executes NAS (Neural Architecture Search) techniques on the backbone 901. However, the CNN model training system 100 based on neural architecture search can be easily extended to the neck 902 and the head 903 with little effort.
[0062] Hereinafter, the CNN model training system 100 based on neural architecture search is also simply referred to as the system 100.
[0063] (Flow of the neural architecture search process executed by the system 100) FIG. 8 is a flowchart showing the flow of the neural architecture search process executed by the system 100 in the second embodiment.
[0064] In step S101, the system 100 constructs a supernet using the supernet builder 300. The process executed by the supernet builder 300 will be described later.
[0065] In step S102, the system 100 trains the supernet using the supernet trainer 400 with object detection tasks and classification tasks. The process executed by the supernet trainer 400 with object detection tasks and classification tasks will be described in detail below.
[0066] In step S103, system 100 performs candidate selection using neural architecture selector 500. By this step, one sublayer is selected from M parallel sublayers, thereby optimizing the target layer. The process executed by neural architecture selector 500 will be described later.
[0067] In step S104, system 100 determines whether all N layers of the supernet have been optimized or covered. If not, the value of parameter i and the pruned supernet from neural architecture selector 500 are provided as input to supernet builder 300 for the next iteration. In this way, system 100 executes the next iteration.
[0068] The processes of steps S101 - 104 are repeatedly executed until it is determined in S104 that all N layers of the supernet have been covered.
[0069] If it is determined in step S104 that all N layers of the supernet have been covered, the process of step S105 is executed. In step S105, system 100 outputs the optimized CNN model 800.
[0070] That is, system 100 performs such iterations N times to cover all N layers of the supernet. A target layer is set in each iteration. For example, in the first iteration, the target layer is B 1 is. In the second iteration, the target layer is B 2 is. Finally, the target layer is B N is. At the end of the Nth iteration, the final pruned supernet from neural architecture selector 500 is given as output.
[0071] Optionally, the pruned supernet may be trained for one or more epochs to improve accuracy using the means by which the STOC400 trains the supernet on the object detection task. The ultimately pruned supernet, also called the optimal subnetwork or optimal CNN model, serves as the output of system 100.
[0072] During supernet optimization, since only one layer is targeted at a time, faster exploration is possible compared to simultaneous optimization of all layers. Also, except for the layer one before and the layer one after the target layer B in a particular iteration, all the other N-2 layers including the sub-layers of B are trained with one layer as input and one layer as output. This is the architecture of a general output subnetwork or output optimal CNN model. i except for the layer one before and the layer one after the target layer B i all the other N-2 layers including the sub-layers of B are trained with one layer as input and one layer as output. This is the architecture of a general output subnetwork or output optimal CNN model.
[0073] (Processing by Supernet Builder 300) When system 100 is performing the first iteration of supernet (supernetwork) construction, the supernet is constructed from scratch as shown in FIG. 6. When system 100 is performing an iteration other than the first iteration, it receives the supernet from neural architecture selector 500. Then, supernet builder 300 modifies the pre-constructed supernet using it.
[0074] FIG. 9 is a flowchart for explaining the processing executed by supernet builder 300.
[0075] In step S301, supernet builder 300 determines whether the current processing is being executed as an initial supernet construction. The initial supernet construction corresponds to the first iteration. If supernet builder 300 determines that the current processing is being executed as an initial supernet construction, the processing of step S302 is executed.
[0076] In step S302, the supernet builder 300 sets the value of parameter i to "1".
[0077] In step S303, the supernet builder 300 constructs a supernet having a backbone 901, a neck 902, and a head 903. Here, the backbone 901 has F fixed layers, N serial layers, M parallel sub-layers in the i-th layer (0 < i ≤ N), and M FC blocks connected to the output of the N-th layer. That is, the supernet builder 300 constructs a supernet CNN model as described with reference to FIG. 6.
[0078] In step S304, the supernet builder 300 initializes the weights for all layers and sub-layers.
[0079] There are several options for weight initialization. For example, random initialization, parameter remapping, Xavier initialization, etc. The weight initialization task can be performed in the STOC400 as needed.
[0080] The constructed and weight-initialized supernet is output in step S308 of FIG. 9.
[0081] If in step S301 the supernet builder 300 determines not to execute the current process as the initial supernet construction, step S305 is executed. In this case, the system 100 executes the second or subsequent iteration of the supernet construction.
[0082] In step S305, the supernet builder 300 receives the pruned supernet from the neural architecture selector 500 and the value of parameter i. The supernet from the previous iteration is also called the pruned supernet. The pruned supernet is the output of the neural architecture selector 500. The supernet received from the neural architecture selector 500 has one layer for each of all N serial layers. Also, from the neural architecture selector 500, information regarding the next B i layers to be optimized is received as the value of parameter i.
[0083] In step S306, the supernet builder 300 replaces the i-th layer with M parallel sub-layers and connects new M FC blocks to the output of the N-th layer of the pruned supernet. In the received supernet, B i layers are replaced with M parallel sub-layers, and the remaining N - 1 layers are maintained as they are.
[0084] In step S307, the supernet builder 300 performs weight initialization only for the newly added sub-layers in the i-th layer. Performs weight initialization for all newly added M sub-layers arranged in the B i layer.
[0085] The modified supernet is given as the output of the supernet builder 300 in step S308 of FIG. 9. The output supernet is provided to the STOC400.
[0086] In this way, the processing of the supernet builder 300 is executed.
[0087] (Processing executed by the supernet trainer 400 with object detection task and classification task) FIG. 10 is a flowchart for explaining the processing executed by the STOC400.
[0088] In step S401, STOC400 sets parameter j to "1".
[0089] In step S402, STOC400 freezes all FC blocks. First, the weights of all FC blocks 9080 are frozen. This means that the weights do not change during the training of the supernet by the object detection task in step S402.
[0090] In step S403, STOC400 trains the supernet using the object detection task. Here, the supernet has only the B i layer among the m sub - layers of the B i j layer, the B k (0 < k ≤ N; i ≠ k;) layers, and the neck 902 and the head 903.
[0091] Here, the supernet is trained with the training dataset 200 having the object detection task. During the training of the supernet, among all M sub - layers of the B 1 layer, only the B 1 1 sub - layer 9051 participates. In other words, during forward and backward propagation, only the B 1 sub - layer 9051 out of the M sub - layers of the B 1 1 layer participates with the other B k (0 < k ≤ N; i ≠ k;) layers of the backbone 901, the neck 902, and the head 903. The remaining M - 1 sub - layers of the B 1 do not participate during training.
[0092] Figure 11 shows the current training in step S403 of Figure 10 when it is executed for the first time. In this case, it shows the sub - layer B 1 1 during the forward propagation in the training phase based on the object detection task of the supernet model in the first iteration. In Figure 11, thick curved arrows are drawn to identify the blocks used in the current training. In Figure 11, the thick curved arrows indicate the fixed layer 904, the B 1 1Sub-layer 9051, B 2 From layer 906 to B N It passes through layer 907, neck 902, head 903, and object detection output 2002.
[0093] In step S404, STOC400 freezes the weights of the entire supernet, while STOC400 unfreezes the weights of the FC 1 block 9081. Then, the supernet is trained for the classification task. During training, B 1 Among the M sub-layers of the B layer, B 1 1 sub-layer 9051, and the FC 1 block 9081, and the other B k (0 < k ≦ N; i ≠ k;) layers only participate.
[0094] In step S405, STOC400 trains the supernet for the classification task. Here, the supernet freezes only B i from the M sub-layers of the B layer, B i j only, freezes the B k (0 < k ≦ N; i ≠ k;) layers, and unfreezes the FC j block.
[0095] Figure 12 shows the current training in step S405 of Figure 10 when it is executed for the first time. In this case, the forward propagation sub-layer B in the training phase based on the classification task of the FC block in the first iteration is shown. In Figure 12, thick curved arrows are drawn to identify the blocks used in the current training. In Figure 12, the thick curved arrows pass through the fixed layer 904, B 1 1 sub-layer 9051, B 1 1 from layer 906 to B 2 layer 907, FC N block 9081, and the classification output 9091. In this case, only the weights of the FC 1 block 9081 are updated during training. 1 block 9081 are updated.
[0096] In this way, the training described with reference to FIGS. 11 and 12 is performed for each sub-layer.
[0097] In step S406, STOC400 determines whether all sub-layers of the i-th layer are covered or whether the value of parameter j is equal to M. If STOC400 determines that not all sub-layers of the i-th layer are covered and the value of parameter j is not equal to M, step S407 is executed. In step S407, STOC400 increments the value of parameter j by only "1". Then, the processes of steps S402 to S406 are executed again.
[0098] FIG. 13 is a diagram showing the current training executed for the second time in step S403 of FIG. 10. In this case, sub-layer B in the forward propagation in the training phase based on the object detection task of the super network model in the first iteration 1 2 is depicted. In FIG. 13, thick curved arrows are drawn to identify the blocks used in the current training. In FIG. 13, the thick curved arrows pass through the fixed layer 904, B 1 2 sub-layer 9052, B 2 layer 906 to B N layer 907, neck 902, head 9903, and object detection output 2002.
[0099] FIG. 14 shows the current training executed for the second time in step S404 of FIG. 10. In this case, sub-layer B in the forward propagation in the training phase based on the classification task of the FC block in the first iteration 1 2 is shown. In FIG. 14, thick curved arrows are drawn to identify the blocks used in the current training. In FIG. 14, the thick curved arrows pass through the fixed layer 904, B 1 2 sub-layer 9052, B 2 layer 906 to B N layer 907, FC 2It passes through block 9082 and classification output 9092. In this case, during training, only the weights of FC 2 block 9081 are updated.
[0100] The processes from step S402 to S406 are repeatedly executed until STOC400 determines that all sub - layers of the i - th layer are covered or the value of parameter j is equal to M. That is, in the case of the object detection task, one sub - layer is involved from the B i layer, and in the case of the classification task, one of the corresponding FC blocks is involved, and the process of training the super - net is repeatedly executed M - 1 times for the other M - 1 sub - layers and M - 1 FC blocks.
[0101] In this way, in the training process by STOC400, multiple candidate layers are trained one by one, and corresponding to one of the multiple candidate layers, multiple fully - connected layers are trained.
[0102] After repeating the training of the super - net by the object detection task and the classification task M times, in step S408, the trained super - net and all FC blocks are output. Thereby, step S102 in FIG. 3 executed to train the super - net partially by the object detection task and partially by the classification task is completed.
[0103] The trained super - net is input to the neural architecture selector 500, and the neural architecture selector 500 executes the selection of sub - layers described as step S103 in FIG. 8.
[0104] In this way, the process of STOC400 is executed.
[0105] (Process executed by neural architecture selector 500) FIG. 15 is a flowchart for explaining the process executed by the neural architecture selector 500.
[0106] In step S501, the neural architecture selector 500 receives the trained supernet from the STOC400. Also, the neural architecture selector 500 receives the training dataset for the classification task 700 from the dataset transformer 600.
[0107] In step S502, the neural architecture selector 500 selects, as the winner of the i-th layer, the sublayer in which the corresponding FC block has the minimum loss / maximum accuracy in the i-th layer.
[0108] First, for the FC 1 Using the output of the FC block 9081 and the ground truth label from the training dataset of the classification task 700, the loss is evaluated using a predefined loss function. For the output FC 1 During the evaluation of the loss in the FC block 9081, only the B k (0 < k ≦ N; i ≠ k;) layers and the FC 1 block 9081, along with the B 1 layers of the B i 1 sub-layers participate.
[0109] The definition of the loss function may vary depending on the purpose of the NAS. If the only purpose is to achieve high accuracy, the definition of the loss function may consist only of the classification loss as represented by the following equation (1).
[0110]
Equation
[0111]
Equation
[0112] In this way, the neural architecture selector 500 selects one candidate layer corresponding to the best-performing one among the plurality of fully-connected layers from among the plurality of candidate layers.
[0113] The more efficiently the sublayer can extract features, the higher the likelihood that the corresponding FC block can accurately classify the input image. Therefore, it can be surely said that the sublayer corresponding to the best-performing FC block is the best choice among the M sublayers compared to the sublayers of the other M - 1 FC blocks.
[0114] In step S503, the neural architecture selector 500 retains only the winning sublayer and deletes the remaining M - 1 sublayers from the supernet. Since the remaining M - 1 sublayers are pruned, at the end of step S503, B 1 The layer is optimized. The pruned supernet having one layer for each of the total N layers and the other fixed layers 904, and having the neck 902 and the head 903, is given as the output to the backbone 901.
[0115] In step S504, the neural architecture selector 500 determines whether the search for all layers in the supernet has been completed. If the neural architecture selector 500 determines that the search for all layers of the supernet has not been completed, the processing of step S505 is executed.
[0116] In step S505, the neural architecture selector 500 outputs the current supernet, which is the pruned supernet. That is, the neural architecture selector 500 includes output means for outputting the pruned supernetwork by the selection process of the neural architecture selector 500.
[0117] By the process of step S505, as shown in step 3, S103 of FIG. 8, the candidate selection process for the B layer in the backbone 901 is completed by the neural architecture selector 500. 1 Also, the neural architecture selector 500 increments the value of parameter i to optimize the next layer. The output supernet is provided to the supernet builder 300. The process of step S305 is executed on the pruned supernet.
[0118] In step S504, when the neural architecture selector 500 determines that the search for all layers of the supernet has been completed, the process of step S506 is executed. In step S506, the neural architecture selector 500 outputs the current supernet. Since the search for all layers in the supernet has been completed, in step S104 of FIG. 8, it is determined that all N layers in the supernet are covered. Then, in step S105 of FIG. 8, the supernet output from the neural architecture selector 500 is output as an optimized CNN model.
[0119] In this way, the process of the neural architecture selector 500 is executed.
[0120] (Advantageous effects of the second embodiment) According to the second embodiment, only one target layer is optimized at a time. Compared with optimizing all layers simultaneously, the search for the neural architecture can be performed in a short time. Therefore, the time efficiency of training the supernet is very good.
[0121] Also, except for the layer immediately before and the layer immediately after the target layer B in a specific iteration, the other N - 2 layers including the sub - layer B i are all trained with one layer as the input and one layer as the output. This is the architecture of a general output sub - network or the optimal CNN model to be output. Therefore, by optimizing one layer in one iteration, the architecture gap between the super - network and the sub - network can be significantly reduced. i Furthermore, the advantage of such a narrow gap may also shorten the training time of the sub - network, thereby further shortening the training time required by the NAS - based CNN model training system 100.
[0122]
[0123] <Third Embodiment> Hereinafter, the third embodiment will be described with reference to the drawings. In the following description, only the differences between the system 100 and the neural architecture search process according to the second embodiment and the system 100 and the neural architecture search process according to the third embodiment will be described.
[0124] Note that elements having the same functions as those described in the second embodiment are assigned the same reference numerals, and the description thereof will be omitted as appropriate.
[0125] In the second embodiment, the optimization of the Bi layer is executed in the order of the B 1 layer, the B 2 layer, ··· the B N layer. However, in order to select the target layer, the order may be randomly changed (random traverse) so that 0 < i ≤ N.
[0126] In this embodiment, when executing the first iteration, any layer of the B i layer can be optimized. The only constraint is that all B i layers must be covered in N iterations.
[0127] Figure 16 shows the supernet CNN model constructed by the supernet builder 300 in this embodiment. The supernet CNN model shown in Figure 16 is constructed in the first iteration of step S101 in Figure 8. In this case, the random traversal is described as "N", which means that the initial value i of the parameter is "N".
[0128] In the example shown in Figure 16, the optimization of the backbone 901 starts from layer B N That is, first, layer B N is replaced with M parallel CNN layers (sub-layers). Therefore, in layer B N , a plurality of sub-layers (B N 1 , B N 2 ,..., B N M ) are illustrated. The M FC blocks 9081 to 9083 are respectively connected to the sub-layers 9071 to 9073.
[0129] On the other hand, in the second embodiment, as shown in Figure 6, first, layer B 1 is replaced by M parallel CNN layers (sub-layers).
[0130] Figure 17 is a flowchart for explaining the processing executed by the supernet builder 300 in this embodiment.
[0131] In step S301 of Figure 17, the supernet builder 300 determines whether the current processing is executed as the initial supernet construction. The initial supernet construction corresponds to the first iteration. When the supernet builder 300 determines that the current processing is executed as the initial supernet construction, the processing of step S302 is executed.
[0132] In step S302 of Figure 17, the supernet builder 300 sets the value of the parameter i to "N".
[0133] On the other hand, in the second embodiment, as shown in step S302 of FIG. 9, the supernet builder 300 sets the value of parameter i to "1".
[0134] FIG. 18 is a flowchart for explaining the process executed by the neural architecture selector 500 in the present embodiment.
[0135] In step S504 of FIG. 18, the neural architecture selector 500 determines whether the search for all layers in the supernet has been completed. If the neural architecture selector 500 determines that the search for all layers of the supernet has not been completed, the process of step S505 is executed.
[0136] In step S505 of FIG. 18, the neural architecture selector 500 outputs the current supernet, which is the pruned supernet. By the process of step S505, the current B in the backbone 901 by the neural architecture selector 500 i The candidate selection procedure for the layer is completed as shown in steps 3 and S103 of FIG. 8. The output supernet is provided to the supernet builder 300. Thereafter, the process of S305 is executed using the pruned supernet.
[0137] Also, the neural architecture selector 500 sets the value of parameter i so as to optimize the next layer. Also, the next layer is selected by random traversal. Assuming that the system 100 executes the first iteration of setting the value of parameter i to "N", in step S505 of FIG. 18, values from 1 to N - 1 are randomly selected.
[0138] On the other hand, in the second embodiment, as shown in step S505 of FIG. 15, the supernet builder 300 sets the value of parameter i to "i + 1".
[0139] (Advantageous Effects of the Second Embodiment) As described above, by randomly selecting the target layer B i the degree of freedom for optimization can be improved.
[0140] <Fourth Embodiment> Hereinafter, the fourth embodiment will be described with reference to the drawings. In the following description, only the differences between the system 100 and the neural architecture search process according to the second and third embodiments and the system 100 and the neural architecture search process according to the fourth embodiment will be described.
[0141] Note that elements having the same functions as those described in the second and third embodiments are denoted by the same reference numerals, and the description thereof will be omitted as appropriate.
[0142] In the second and third embodiments, the M FC blocks 9081 to 9083 are connected to the output of the B N layer 907. However, the M FC blocks may be connected to the output of any B later layer such that i ≤ later. That is, in the present embodiment, the M FC blocks may be connected to any of the B i layers deeper than or equal to the target layer.
[0143] FIG. 19 shows the supernet CNN model constructed by the supernet builder 300 in the present embodiment. The supernet CNN model shown in FIG. 19 is constructed in the first iteration of step S101 in FIG. 8, and the target layer is the B 1 layer.
[0144] Therefore, a plurality of sub-layers (B 1 1 1 1 2 1 M ) is illustrated. The M FC blocks 9081 to 9083 are respectively connected to the sub-layers 9051 to 9053. In this case, the M FC blocks are respectively connected to the output of the target layer. This means that the value of the parameter later is equal to the value of the parameter i.
[0145] On the other hand, in the second and third embodiments, as shown in FIGS. 6 and 16, the M FC blocks are respectively connected to the output of the B N layer regardless of the layer currently selected as the target layer.
[0146] FIG. 20 is a flowchart for explaining the process executed by the supernet builder 300 according to the present embodiment.
[0147] In step S303 of FIG. 20, the supernet builder 300 constructs a supernet having a backbone 901, a neck 902, and a head 903. Here, the backbone 901 has F fixed layers, N serial layers, M parallel sub-layers in the i-th layer (0 < i ≤ N), and M FC blocks connected to the output of the later layer (i ≤ later).
[0148] On the other hand, in the second and third embodiments, as shown in step 303 of FIG. 9 and step 303 of FIG. 17, the supernet builder 300 constructs a supernet having a backbone 901, a neck 902, and a head 903, where the backbone 901 has F fixed layers, N serial layers, M parallel sub-layers in the i-th layer (0 < i ≤ N), and M FC blocks connected to the output of the N-th layer.
[0149] (Advantageous effects of the fourth embodiment) As described above, by connecting the M FC blocks to either a layer equal to or deeper than the target layer, the degree of freedom in optimization can be improved. i
[0150] (Configuration Example by Software) One or more of all the functions of the system 100 can be realized by hardware such as an integrated circuit (IC chip) or, alternatively, by software.
[0151] In the latter case, the system 100 is realized by, for example, a computer that executes the instructions of a program that is software for realizing the aforementioned functions. FIG. 21 shows an example of such a computer (hereinafter referred to as “computer C”). The computer C includes at least one processor C1 and at least one memory C2. The program P for causing the computer C to function as the system 100 is stored in the memory C2. In the computer C, the function of the system 100 is realized when the processor C1 reads and executes the program P from the memory C2.
[0152] As the processor C1, for example, a CPU (Central Processing Unit), GPU (Graphic Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating Point Number Processing Unit), PPU (Physics Processing Unit), a microcontroller, or a combination thereof can be used. The memory C2 can be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.
[0153] Note that the computer C may further include a random access memory (RAM) in which the program P is loaded and various data are temporarily stored when the program P is executed. The computer C may further include a communication interface for transmitting and receiving data to and from other devices. The computer C may further include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.
[0154] The program P can be stored in a non-transitory tangible storage medium M readable by the computer C. As the storage medium M, for example, a tape, a disk, a card, a semiconductor memory, a programmable logic circuit, etc. can be used. The computer C can acquire the program P via the storage medium M. The program P can be transmitted via a transmission medium. The transmission medium is, for example, a communication network, a broadcast wave, etc. The computer C can also acquire the program P via such a transmission medium.
[0155] [Appendix 1] The present invention is not limited to the above-described embodiments, and various modifications can be made by those skilled in the art within the scope described in the claims. For example, the present invention also includes any embodiment derived by appropriately combining the technical means disclosed in the above-described embodiments within its technical scope.
[0156] [Appendix 2] All or part of the embodiments disclosed above can be described as follows. However, it should be noted that the present invention is not limited to the aspects of the following embodiments.
[0157] [Supplementary Explanation] Aspects of the present invention can also be expressed as follows: (Aspect 1) Constructing means for constructing a super network, wherein the target layer for optimizing the super network is replaced by a plurality of candidate layers, and the super network includes a plurality of fully connected layers; Training means for training the super network, wherein the plurality of candidate layers are trained in parts, and the plurality of fully connected layers are trained corresponding to the parts of the plurality of candidate layers; and, Selecting means for evaluating the trained super network and selecting the parts of the plurality of candidate layers corresponding to the parts of the plurality of fully connected layers with the best performance; A neural architecture search device comprising the above.
[0158] (Aspect 2) In the training process by the training means, the plurality of candidate layers are trained one by one, and the plurality of fully connected layers are trained corresponding to the one of the plurality of candidate layers, The selection means selects one of the plurality of candidate layers corresponding to the layer with the best performance among the plurality of fully connected layers, and the neural architecture search device according to Aspect 1.
[0159] (Aspect 3) The plurality of fully connected layers are connected to the output of the target layer or a layer deeper than the target layer, and the neural architecture search device according to Aspect 1.
[0160] (Aspect 4) The training means is For training the super network for at least one of an object detection task using an object detection dataset and A classification task using a classification dataset, and the neural architecture search device according to Aspect 1.
[0161] (Aspect 5) The neural architecture search device according to Aspect 4, further comprising conversion means for converting the object detection dataset into the classification dataset.
[0162] (Aspect 6) The super network includes a backbone block, a neck block, and a head block, The backbone block includes a plurality of CNN layers arranged in series and the plurality of fully connected layers, The target layer is a neural architecture search device according to Aspect 1, selected from the plurality of CNN layers arranged in series.
[0163] (Aspect 7) An output means for outputting a super network pruned by the selection process of the selection means, further comprising a neural architecture search device according to Aspect 1.
[0164] (Aspect 8) A step of constructing a super network, wherein a target layer for optimizing the super network is replaced with a plurality of candidate layers, and the super network includes a plurality of fully connected layers; A step of training the super network, wherein the plurality of candidate layers are trained in parts, and the plurality of fully connected layers are trained corresponding to the parts of the plurality of candidate layers; and A step of evaluating the trained super network and selecting a part of the plurality of candidate layers corresponding to the part of the plurality of fully connected layers with the best performance, A neural architecture search method including.
[0165] (Aspect 9) A program for causing a computer to function as the neural architecture search device according to claim 1, wherein the program causes the computer to function as the construction means, the training means, and the selection means.
Explanation of Signs
[0166] 1 Neural architecture search device 11 Construction unit 12 Training unit 13 Selection Unit 100 Model Training System Based on Neural Architecture Search 200 Training Dataset for Object Detection Task 300 Supernet Builder 400 Supernet Trainer for Object Detection Task and Classification Task 500 Neural Architecture Selector 600 Dataset Transformer 700 Training Dataset for Classification Task 800 Optimized CNN Model 900 Supernet 901 Backbone 902 Neck 903 Head
Claims
1. Constructing means for constructing a super network, wherein the target layer for optimizing the super network is replaced by a plurality of candidate layers, and the super network includes a plurality of fully connected layers; Training means for training the super network, wherein the plurality of candidate layers are trained in parts, and the plurality of fully connected layers are trained corresponding to the parts of the plurality of candidate layers; and Selecting means for evaluating the trained super network and selecting the parts of the plurality of candidate layers corresponding to the parts of the plurality of fully connected layers with the best performance; A neural architecture search device comprising the above.
2. In the training process by the training means, the plurality of candidate layers are trained one by one, and the plurality of fully connected layers are trained corresponding to the one of the plurality of candidate layers, The selecting means selects one of the plurality of candidate layers corresponding to the layer with the best performance among the plurality of fully connected layers. The neural architecture search device according to Claim 1.
3. The plurality of fully connected layers are connected to the output of the target layer or a layer deeper than the target layer. The neural architecture search device according to Claim 1.
4. The training means Trains the super network for at least one of an object detection task using an object detection dataset and A classification task using a classification dataset. The neural architecture search device according to Claim 1.
5. The neural architecture search device according to Claim 4, further comprising conversion means for converting the object detection dataset into the classification dataset.
6. The super network includes a backbone block, a neck block, and a head block, The backbone block includes a plurality of CNN layers arranged in series and the plurality of fully connected layers, The target layer is selected from the plurality of CNN layers arranged in series. The neural architecture search device according to any one of Claims 1 to 5.
7. Output means for outputting a super network pruned by the selection process of the selecting means. The neural architecture search device according to Claim 1.
8. A step of constructing a super network, comprising replacing an optimization target layer of the super network with a plurality of candidate layers, wherein the super network includes a plurality of fully-connected layers; A step of training the super network, comprising training the plurality of candidate layers in parts and training the plurality of fully-connected layers corresponding to the parts of the plurality of candidate layers; and A step of evaluating the trained super network and selecting parts of the plurality of candidate layers corresponding to the parts of the plurality of fully-connected layers with the best performance, A neural architecture search method comprising the above steps.
9. A program for causing a computer to function as the neural architecture search device according to claim 1, wherein the program causes the computer to function as the constructing means, the training means, and the selecting means.
Citation Information
Patent Citations
Machine learning systems and methods for determining home value
US20210110439A1
Technique to perform neural network architecture search with federated learning
US20210374502A1