Learning method, apparatus, program, and model selection method

By evaluating and selecting DNN models based on their performance after initial pruning, followed by additional learning and pruning, the method addresses the challenge of maintaining performance post-pruning, thereby efficiently generating a desired inference model.

JP2025096782APending Publication Date: 2025-06-30DENSO TEN LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023212691
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-30

AI Technical Summary

Technical Problem

The performance of deep neural networks (DNNs) deteriorates rapidly after pruning, making it challenging to generate a desired inference model efficiently, as re-learning may not restore the target performance.

Method used

A learning method that involves evaluating the performance of multiple DNN models after initial pruning, selecting a target model based on this evaluation, and then performing additional learning and second pruning to generate an inference model with sustained performance.

Benefits of technology

This method allows for the selection of a DNN model with high performance after pruning, reducing the likelihood of needing to restart the model selection process and enhancing the efficiency of generating a desired inference model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096782000001_ABST
    Figure 2025096782000001_ABST
Patent Text Reader

Abstract

To efficiently generate a desired inference model.SOLUTION: A learning method for generating an inference model that performs inference based on input data includes: after evaluation learning that learns a learning model composed of a neural network, based on learning data, executing evaluation unit processing (S10) which evaluates the performance of the learning model through first pruning on the learning model, on a plurality of learning models having different structures; selecting (S20) a target model from among the plurality of learning models before the first pruning, based on results of evaluation in the evaluation unit processing; and generating an inference model, after full-scale learning (S21) that learns the target model using the learning data with the number of times of learning more than the evaluation learning, by performing second pruning (S23) on the target model.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning method, apparatus, program, and model selection method.

Background Art

[0002] Various arithmetic processes using a model having a DNN (Deep Neural Network) have been put into practical use. In the design of a DNN, a designer needs to determine design parameters (such as kernel size and number of output channels) for each layer of the DNN (convolutional layer, pooling layer, etc.). In recent years, the DNN has been becoming more multi-layered, and there are many DNNs having more than 100 layers. In the design process of a DNN, optimization of the DNN design is achieved by repeating a process including DNN structure setting and DNN performance evaluation a plurality of times. However, the number of candidates for design parameters in a multi-layer DNN is enormous, and it is difficult for a designer alone to optimize the DNN design.

[0003] As a method for automating the DNN design, a method called NAS (Neural Architecture Search) has been proposed. In NAS, candidate DNN design patterns (search space) are determined in advance. Then, after sampling a model from the search space, a series of processes of evaluating the model through learning of the model (learning for about several iterations to several epochs) in a relatively short time are repeated until a desired target performance is obtained. Incidentally, Patent Document 1 below discloses a method of selecting search space information according to target constraint conditions of target hardware.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Generally, after exploring the DNN structure by NAS, the most performant model is selected, and learning with an increased number of learning iterations is performed on the selected model. Thereafter, when applying the learned model to hardware such as a SoC (System-on-a-chip), pruning is often performed in order to reduce the size and increase the speed of the model. However, the performance of the model may rapidly degrade due to pruning, and the target performance may not be obtained even if the model is re-learned thereafter. In such a case, it is necessary to start over from selecting another model, and the setback is significant. The occurrence of such a setback hinders the efficient generation of the desired inference model (learned model).

[0006] An object of the present invention is to provide a learning method, apparatus, and program, as well as a model selection method, that contribute to the efficient generation of a desired inference model.

Means for Solving the Problem

[0007] The learning method according to the present invention is a learning method for generating an inference model that makes inferences based on input data. After performing evaluation learning in which a learning model composed of a neural network is learned based on learning data, an evaluation unit process for evaluating the performance of the learning model after undergoing first pruning on the learning model is executed for a plurality of learning models having mutually different structures. In this learning method, based on the result of the evaluation in the evaluation unit process, a target model is selected from among the plurality of learning models before the first pruning, and after this main learning in which the target model is learned using the learning data with a greater number of learning iterations than the evaluation learning, the inference model is generated by performing second pruning on the target model.

Effect of the Invention

[0008] Based on the evaluation results in the evaluation unit process, a learning model with relatively high performance as the performance after the first pruning can be selected as the target model. A learning model with relatively high performance as the performance after the first pruning is expected to have relatively high performance even after the second pruning. That is, according to the above learning method, it is possible to select, as the target model, a model that is expected to have high performance even after the second pruning. Conversely, it becomes difficult to select, as the target model, a model whose performance fails to meet the target due to the second pruning. For this reason, after this learning and the second pruning, it is less likely that it becomes necessary to start over from the selection of the target model because the performance of the target model does not meet the target (the occurrence of backtracking is suppressed). That is, it is possible to efficiently (with less computational cost or time cost) generate a desired inference model (a target model that meets the requirements).

Brief Description of Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Modes for Carrying Out the Invention

[0010] Hereinafter, examples of embodiments of the present invention will be specifically described with reference to the drawings. In each of the drawings to be referred to, the same parts are denoted by the same reference numerals, and redundant explanations regarding the same parts are omitted in principle. In this specification, for the sake of simplification of description, the names of information, signals, physical quantities, functional units, circuits, elements, or components, etc. corresponding to the symbols or reference numerals may be omitted or abbreviated by writing the symbols or reference numerals for referring to the information, signals, physical quantities, functional units, circuits, elements, or components, etc.

[0011] Fig. 1 shows the overall configuration of a learning system (machine learning system) according to an embodiment of the present invention. The learning system in Fig. 1 includes a learning device 10 which is a machine learning device and a database 20. The learning device 10 is connected to a communication network including the Internet. The learning device 10 may be configured by one or more computer devices (server devices) connected to the communication network. The learning device 10 may be configured using cloud computing. The learning device 10 includes a controller 11, a memory 12, and a communication unit 13. Note that the learning described in this embodiment is machine learning.

[0012] The controller 11 comprehensively controls the operations of each part in the learning device 10. The controller 11 includes an arithmetic processing unit including a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit) etc. as hardware resources. The controller 11 may realize the following respective functions to be realized by the controller 11 by executing a program recorded in the memory 12, the database 20, or any other arbitrary recording medium (not shown).

[0013] The memory 12 is configured to have a non-volatile memory such as a ROM (Read only memory) or a flash memory, and a volatile memory such as a RAM (Random access memory). In the memory 12, each data referred to by the controller 11 is stored, and various programs to be executed by the controller 11 are also stored.

[0014] The communication unit 13 transmits and receives arbitrary signals to and from a counterpart device different from the learning device 10. The counterpart device for the communication unit 13 includes the database 20 and also includes any computer device connected to the communication network. Note that the controller 11 can transmit and receive arbitrary information to and from the counterpart device (the counterpart device for the communication unit 13) using the communication unit 13, but the description of the communication unit 13 may be omitted hereinafter.

[0015] The database 20 is a large-capacity recording medium that holds learning data 21. The learning data may be read as a learning data set. The database 20 is connected to the learning device 10 either wired or wirelessly. The controller 11 can freely read any data held in the database 20. The database 20 may be a collection of physically separated multiple recording media. In this case, the learning data 21 is held in the multiple recording media. Part or all of the database 20 may be provided in the learning device 10.

[0016] In the controller 11, the model 110 shown in FIG. 2 is constructed. Since machine learning is performed on the model 110, the model 110 can be referred to as a learning model. The model 110 is an algorithm that generates model output data Dout from model input data Din. The model 110 has an NN120 that is a neural network. NN120 is configured using the hardware resources within the controller 11. NN120 may be a deep neural network (DNN). The controller 11 executes machine learning of the model 110 using the learning data 21. Note that the machine learning of the model 110 can also be said to be the machine learning of NN120. It is assumed that the machine learning performed by the controller 11 in this embodiment is supervised machine learning.

[0017] The model 110 that has undergone machine learning is a learned model (inference model) that performs a predetermined inference. The learned model may be, for example, an object detector that performs object detection as an inference. In object detection, the existence region of the image of the object and the type of the object within the two-dimensional image input to the model 110 are estimated (in other words, detected).

[0018] When object detection is performed as an inference, the learning data 21 is the image dataset 21a shown in FIG. 3. The image dataset 21a includes the image data and the correct answer data of a large number of learning images. In the image dataset 21a, the correct answer data is added to each learning image for the learning image. Each learning image in the image dataset 21a is a two-dimensional image including images of various types of objects. The correct answer data for a certain learning image includes information indicating the type of object existing in the learning image and information specifying the position and shape of the region where the image of the object exists in the learning image. The controller 11 can construct a model 110 that performs object detection as an inference by executing machine learning of the model 110 using the image dataset 21a.

[0019] The process in which machine learning is performed is referred to as a machine learning process. The machine learning process in the case of constructing a model 110 that performs object detection by machine learning will be described. For the sake of concretizing the explanation, the noted learning image is referred to as a noted image. In the machine learning process, the controller 11 inputs the image data of the noted image into the model 110 as model input data Din. Then, the model 110 estimates the existence region of the image of the object in the noted image and the type of the object based on the model input data Din, and generates model output data Dout indicating the estimation result. The correct answer data for the noted image indicates the correct answer of the content to be estimated by the model 110.

[0020] In the machine learning process, the controller 11 derives a loss function representing the error between the model output data Dout of the model 110 and the correct answer data for the noted image. Then, the controller 11 adjusts the parameters of the model 110 using the error backpropagation method so that the value of the loss function is reduced. The parameters of the model 110 are the parameters of the NN120 and include the parameters of the weights and biases in the NN120. The above adjustment is repeated until a predetermined learning end condition is satisfied, and when the learning end condition is satisfied, the machine learning process is terminated. The model 110 after going through the machine learning process functions as an object detector (inference model) that performs object detection as an inference.

[0021] Object detection is a type of image recognition. The inference performed by model 110 can be any image recognition (such as image classification or semantic segmentation). When the inference performed by model 110 is image recognition, a convolutional neural network can be used as NN120. The task (inference) performed by model 110 is arbitrary and can be a clustering task, a generation task, a regression task, a reinforcement task, etc. Hereinafter, for the sake of concretizing the description, unless otherwise specified, it is assumed that the learning data 21 is the image dataset 21a and the model 110 is made to function as an object detector through machine learning.

[0022] The database 20 also holds the evaluation data 22. The evaluation data 22 is data for evaluating the performance of the model 110. The evaluation data 22 has the same data configuration as the learning data 21. However, usually, the amount of data of the evaluation data 22 is less than the amount of data of the learning data 21.

[0023] An example of the structure of NN120 is shown in FIG. 4. Any NN120 is a forward-propagation neural network and includes an input layer 121, an intermediate layer 122, and an output layer 123. The input layer 121 receives the model input data Din and outputs the model input data Din to the intermediate layer 122. The intermediate layer 122 is provided between the input layer 121 and the output layer 123. The intermediate layer 122 performs operations on the model input data Din according to the structure and parameters of the intermediate layer 122 and outputs the operation result to the output layer 123. The output layer 123 generates and outputs the model output data Dout based on the operation result output from the intermediate layer 122.

[0024] A plurality of layers are provided in the intermediate layer 122. A target layer, which is an arbitrary layer in the intermediate layer 122, consists of a plurality of nodes. The target layer receives data from the layer arranged in front of itself (specifically, the layer arranged immediately in front of itself), and performs operations on the received data according to its own structure and parameters. Then, the target layer outputs the operation result to the layer arranged behind itself (specifically, the layer arranged immediately behind itself). When the target layer is the first layer provided in the intermediate layer 122, the model input data Din from the input layer 121 is received by the target layer. When the target layer is the last layer provided in the intermediate layer 122, the operation result by the target layer is output to the output layer 123.

[0025] The internal structure of the intermediate layer 122 is various. In the intermediate layer 122 in the NN120 of FIG. 4, a convolutional layer 122a and 122b, ···, a pooling layer 122c, a fully connected layer 122d and 122e are provided from the input layer 121 to the output layer 123. The convolutional layer performs a convolution operation on the feature map supplied to itself, and outputs the result of the convolution operation to the layer arranged behind itself (specifically, the layer arranged immediately behind itself). The convolution operation is spatial filtering using a spatial filter. When a convolutional layer is provided at the first stage of the intermediate layer 122, the feature map supplied to the convolutional layer corresponds to the model input data Din. The pooling layer reduces the resolution of the feature map supplied to itself by spatial downsampling, and outputs the feature map after the resolution reduction to the layer arranged behind itself (specifically, the layer arranged immediately behind itself). The fully connected layer is a layer in which all nodes (neurons) constituting itself are connected to all nodes of the layer arranged in front of itself. The fully connected layer performs a linear transformation using the output data from the previous layer and outputs the result of the linear transformation. In addition, a batch normalization layer or the like may be provided in the intermediate layer 122.

[0026] The controller 11 is configured to be able to change the structure of the NN 120 in various ways. If the structure of the NN 120 changes, the characteristics of the model 110 change. Let the number of types of structures of the NN 120 set by the controller 11 be represented by "n". The number of types of structures that the controller 11 can set is much larger than n, but here we focus on n types of structures. When the NN 120 has the i-th structure, the model 110 is referred to as the i-th model. n represents any integer of 2 or more, and i represents a natural number of n or less. The controller 11 can cause the model 110 to function as the first to n-th models by changing the structure of the NN 120 in n steps. The structure of the NN 120 is also the structure of the model 110. That is, the i-th structure is the structure of the NN 120 that forms the i-th model and can also be said to be the structure of the i-th model.

[0027] The first to n-th structures are different from each other. That is, for all combinations of natural numbers e and f, the design parameters set for the NN 120 of the e-th structure and the design parameters set for the NN 120 of the f-th structure are different from each other. e and f represent any different natural numbers of n or less.

[0028] The design parameters are the parameters that define the structure of the NN 120. Therefore, the design parameters may also be referred to as structure parameters. For example, the number of filters, the number of channels, and the kernel size in the convolutional layer correspond to the design parameters. The number of filters, the number of channels, and the kernel size in the pooling layer also correspond to the design parameters. The number of layers constituting the NN 120 of the e-th structure and the number of layers constituting the NN 120 of the f-th structure may also be different from each other.

[0029] The learning device 10 selects a target model from among the first to n-th models and performs machine learning (main learning described later) used for the learning data 21 on the target model. The target model is typically a model expected to have the maximum inference accuracy. The target model corresponds to the optimal model extracted from among the first to n-th models. The method for selecting the target model will become clear from the following explanation.

[0030] Figure 5 shows how two layers are combined. Each layer of NN120 includes a plurality of nodes. In NN120, each node generates its output data from the input data to itself through forward propagation and supplies the output data to the nodes connected to itself. In Figure 5, layer L1 and layer L2 are any two adjacent layers provided in NN120, and layer L2 is provided after layer L1. In Figure 5, each white circle represents a node. Each node in layer L1 is connected to the corresponding node in layer L2. Each node generates its output data by performing an operation according to the learning parameters set for itself on the input data to itself. The learning parameters include weights and biases.

[0031] A large number of nodes are provided in NN120, and there are also a large number of combinations of connections between nodes. Among the connections between nodes, connections (node connections) that have no or almost no influence on the model output data Dout can be deleted, and the process of performing such deletion is called pruning. Pruning can be realized by setting the weights set for the nodes to zero. Pruning aims to reduce the weight and speed up the model.

[0032] Figure 6 shows an operation flowchart according to a reference example. The operation according to the reference example is not the operation executed by the learning device 10, but is provided for comparison with the operation of the learning device 10. For convenience, the device that performs the operation according to the reference example is referred to as a reference device.

[0033] The reference device first performs a search for a neural network that meets the reference performance by the NAS in step S911. In step S911, a plurality of structures are set as the structure of the neural network. Then, for each set structure, after training a model having a neural network for a relatively short time, the performance of the model is evaluated. In step S912 following step S911, the reference device selects a model that meets the predetermined performance. The reference device trains the selected model with a larger number of training times than in step S911, and then performs pruning through performance evaluation (steps S913 to S915). Pruning results in a deviation from the appropriate value of each training parameter. Therefore, after the reference device retrains the model after pruning, it performs performance evaluation again (steps S916 and S917). Then, it is confirmed whether the model that has undergone pruning and retraining meets the target performance (step S918). At this time, if the target performance is met, there is no problem.

[0034] However, if the target performance is not met (No in step S918), it is necessary to start over from returning to step S912 to select another model, and the backtracking is large. That is, when returning to step S912, the training in step S913 that requires a relatively long time (for example, several days to one month) is performed again for another model, and the time required to obtain the necessary trained model increases significantly. The costs (costs related to calculation or time) associated with the processes in steps S913 and S916 increase as the model or training data is scaled up, and it is practically difficult to repeat the process many times.

[0035] With reference to FIG. 7, the influence of pruning on the calculation amount and performance will be described. The calculation amount here represents the calculation amount of the model during inference. Consider the first and second reference models as two models according to the reference example. Assume a situation where the first and second reference models are individually selected in step S912.

[0036] In FIG. 7, plot 931 represents the amount of calculation and performance in the first reference model before pruning (specifically, the amount of calculation and performance in the first reference model immediately after learning in step S913). Plot 932 represents the amount of calculation and performance in the first reference model after pruning (specifically, the amount of calculation and performance in the first reference model after relearning in step S916). Plot 941 represents the amount of calculation and performance in the second reference model before pruning (specifically, the amount of calculation and performance in the second reference model immediately after learning in step S913). Plot 942 represents the amount of calculation and performance in the second reference model after pruning (specifically, the amount of calculation and performance in the second reference model after relearning in step S916).

[0037] Before pruning, both the first and second reference models satisfy the target performance. In step S915, pruning is performed so that the amount of calculation of each reference model is equal to or less than the target amount of calculation. As a result of pruning, the first reference model does not satisfy the target performance even after relearning. In contrast, the second reference model satisfies the target performance after pruning and relearning.

[0038] If the second reference model could be selected in step S912 in the reference device first, there would be no need to return from step S918 to step S912. However, when the first reference model is selected in step S912 first, the determination in step S918 for the first reference model is negative (No), so it is necessary to return to step S912. In the reference device, since the influence of pruning is not considered at the stages of steps S911 and S912, there is a high risk of the above-mentioned backtracking occurring.

[0039] If the selection of the model could be made considering the influence of pruning before executing the learning that requires a long time, the risk of backtracking could be reduced. Considering these, in the learning device 10 according to the present embodiment, pruning is performed at the stage of searching for the model structure by NAS, the performance of the model after pruning is evaluated, and a model to which the subsequent processing (such as the present learning described later) is applied is selected based on the evaluation result.

[0040] Fig. 8 shows the operation flowchart of the learning device 10. After the learning device 10 is activated, the learning program is executed by the controller 11, and the processes of each step shown in Fig. 8 are executed by the controller 11 starting from the model evaluation process in step S10. In the model evaluation process, the first to nth models are respectively learned. The learning performed in the model evaluation process is particularly referred to as evaluation learning. In the model evaluation process, an evaluation index corresponding to the performance of the model is derived based on the result of the evaluation learning for each model. After step S10, through the selection of the target model (step S20), the learning of the target model is executed in step S21. The learning performed in step S21 is particularly referred to as this learning.

[0041] The evaluation learning is short-time learning implemented for the purpose of selecting a target model from among the first to nth models. The short time here refers to being short in comparison with the execution time of this learning. That is, the performance of each model is evaluated in the model evaluation process, and this learning that requires a long time is performed on the target model selected based on the result of the evaluation. Compared with the number of learning times in the evaluation learning of each of the first to nth models, the number of learning times in the this learning of the target model is large. The number of learning times is represented by, for example, the number of epochs. The number of epochs (for example, several tens) in the this learning of the target model is larger than the number of epochs (for example, 1 or 2) in the evaluation learning of the ith model. While the this learning of the target model is implemented using all of the learning data 21, the evaluation learning of the ith model may be implemented using only a part of the learning data 21.

[0042] Referring to Fig. 9, the flow of the model evaluation process executed in step S10 of Fig. 8 will be described. Fig. 9 is a flowchart of the model evaluation process. The model evaluation process consists of the processes of steps S11 to S18. In the model evaluation process, the controller 11 first substitutes 1 for the variable i managed by itself in step S11. After step S11, it proceeds to step S12.

[0043] In step S12, the controller 11 performs sampling of the model using NAS. NAS is an abbreviation of "Neural Architecture Search".

[0044] NAS itself is well-known. A method of sampling a model by NAS will be briefly described. Prior to the execution of the model evaluation process, a search space is defined by an operator of the learning device 10, and search space information indicating the search space is given to the learning device 10. When each of a plurality of structures that can be adopted as the structure of the NN120 is referred to as a candidate structure, the search space is a set of a plurality of candidate structures. The search space defines from which set of structures the optimal structure of the NN120 is to be searched.

[0045] In NAS, a method for searching for the optimal structure of the NN120 from the search space is called a search strategy. In step S12, the controller 11 samples a model from the search space according to a determined search strategy. As the search strategy, Bayesian optimization using TPE, an evolutionary algorithm using NSGA, or reinforcement learning can be used. TPE is an abbreviation of "Tree-Structured Parzen Estimator". NSGA is an abbreviation of "Non-dominated Sorting Genetic Algorithm".

[0046] The model sampled in step S12 is the i-th model, and the i-th model is set to the model 110. That is, in step S12, the model 110 having the NN120 of the i-th structure is constructed. The loop process from step S12 to step S17 is repeatedly executed. The i-th model in the first loop process is the model 110 as the first model, and the i-th model in the second loop process is the model 110 as the second model. The same applies to the loop processes after the third time. After step S12, the process proceeds to step S13.

[0047] In step S13, the controller 11 initializes the i-th model. After step S13, it proceeds to step S14. In the initialization, a predetermined initial value or a value randomly selected from a predetermined range is set as the learning parameter of the i-th model. Any model has learning parameters. The learning parameters refer to the parameters that are adjusted during learning, and the above-mentioned weights and biases belong to the learning parameters. When the batch normalization layer is included in the NN120 of the i-th model, the scale parameter and the shift parameter in the batch normalization layer also belong to the learning parameters.

[0048] In step S14, the controller 11 trains the i-th model using the training data 21. The training in step S14 is evaluation training. Specifically, in step S14, the controller 11 extracts a mini-batch from the training data 21. The mini-batch is data having a predetermined mini-batch size within the training data 21. For example, when the training data 21 includes image data of 100,000 training images, the mini-batch includes image data of 10 training images therein and also includes the correct data for the 10 training images. In the evaluation training, the controller 11 inputs the image data of the training images included in the mini-batch as the model input data Din to the i-th model. Thereby, the i-th model generates the model output data Dout based on the current learning parameters. The correct data included in the mini-batch corresponds to the correct data for the model output data Dout.

[0049] The controller 11 derives a loss function representing the error between the model output data Dout generated based on the mini-batch and the correct answer data included in the mini-batch. Further, the controller 11 derives the gradient of the loss function of the i-th model by the error backpropagation method, and updates the learning parameters of the i-th model based on the derived gradient. Thereafter, a new mini-batch is extracted. Each time a mini-batch is extracted, the learning parameters of the i-th model are updated. In the learning for evaluation, a series of processes of updating the learning parameters of the i-th model starting from the extraction of the mini-batch are repeated. When the series of processes are performed a predetermined number of times or for a predetermined time with respect to the i-th model, the learning for evaluation of the i-th model ends, and a transition from step S14 to step S15 occurs.

[0050] Incidentally, hereinafter, the i-th model before the learning for evaluation is performed may be referred to as model M A [i], and the i-th model immediately after the learning for evaluation is performed may be referred to as model M B [i]. In step S14, by performing the learning for evaluation on model M A [i], model M B [i] is generated. In other words, by performing the learning for evaluation on the i-th model, the i-th model is updated from model M A [i] to model M B [i].

[0051] In step S15, the controller 11 performs pruning on the i-th model after the learning for evaluation based on the learning parameters of the i-th model after the learning for evaluation. That is, the controller 11 performs pruning on model M B [i] based on the learning parameters of model M B [i].

[0052] In step S15, among the connections between a large number of nodes existing in the NN120 of the i-th model, the connections (connections of nodes) that have no or almost no influence on the model output data Dout of the i-th model (model M B [i]) are specified as unnecessary connections. In step S15, the i-th model (model MB [i])'s unnecessary connections are deleted. The process of deleting unnecessary connections is pruning.

[0053] Add an explanation about the pruning in step S15. In the explanation of the pruning in step S15, the i-th model refers to the i-th model after learning for evaluation (i.e., model M B [i]). The controller 11 according to step S15 sets a threshold value TH A . Then, the controller 11 extracts the weights in the NN120 of the model M B [i] that are less than the threshold value TH A , and identifies the connections (connections between nodes) corresponding to the weights less than the threshold value TH A as unnecessary connections. The connection corresponding to the weight less than the threshold value TH A refers to the connection between the node with the weight less than the threshold value TH A set and the subsequent-stage node connected to the node. The controller 11 first performs initial pruning to identify unnecessary connections with the threshold value TH A set to a predetermined initial small value, and deletes the identified unnecessary connections. The controller 11 evaluates the calculation amount of the i-th model after the initial pruning. Then, the controller 11 gradually increases the threshold value TH A by a predetermined step value. Each time the threshold value TH A increases, the location of the identified unnecessary connections also increases. Each time the threshold value TH A increases, the controller 11 performs stage pruning to newly identify unnecessary connections and delete the newly identified unnecessary connections. Each time the controller 11 deletes a newly identified unnecessary connection, it evaluates the calculation amount of the i-th model after the deletion. The above operations are repeated until the calculation amount of the i-th model becomes less than or equal to the target calculation amount. When the calculation amount of the i-th model becomes less than or equal to the target calculation amount, the pruning in step S15 ends. Therefore, the calculation amount of the i-th model at the end of step S15 is either equal to the target calculation amount or slightly less than the target calculation amount. Note that the deletion of unnecessary connections may be realized by setting zero to the weights in the unnecessary connections (weights less than the threshold value TH A ).

[0054] Any two nodes provided in the i-th model before pruning and connected to each other are referred to as a first target node and a second target node. The second target node is arranged downstream of the first target node. The first target node generates its output data by performing an operation according to the learning parameters set for itself on the input data to itself. Before pruning, the output data of the first target node is input to the second target node. In step S15, the controller 11 uses the weight used when generating the input data from the first target node to the second target node, and sets the magnitude of the weight set for the first target node as the threshold TH A for comparison. If the magnitude of the weight is less than the threshold TH A , it is determined that the connection between the first and second target nodes is an unnecessary connection. At this time, the deletion of the connection between the first and second target nodes may be realized by setting zero for the weight of the first target node when generating the input data from the first target node to the second target node. If the magnitude of the weight is greater than or equal to the threshold TH A , it is not determined that the connection between the first and second target nodes is an unnecessary connection. If the weight set for the first target node is a vector quantity, the magnitude of the weight represents the norm of the weight. The same processing is performed for all combinations of two nodes provided in the i-th model before pruning and connected to each other.

[0055] The calculation amount of the i-th model represents the amount of calculation performed by the i-th model when performing one inference in the i-th model. That is, the calculation amount of the i-th model, in the case of image recognition, represents the amount of calculation performed by the i-th model from when the image data of one input image is input to the i-th model until the model output data Dout based on the input image is output from the i-th model. Since the method of deriving the calculation amount itself is known, detailed description is omitted. The target calculation amount is set in advance, and information indicating the target calculation amount is held in the memory 12.

[0056] Hereinafter, the i-th model immediately after the pruning in step S15 is the model M Cis sometimes referred to as [i]. In step S15, pruning is performed on model M B [i] to generate model M C [i]. In other words, by performing pruning on the i-th model after the evaluation learning, unnecessary connections in the structure of the i-th model are removed, and as a result, the i-th model is changed from model M B [i] to model M C [i]. Model M C [i] is the i-th model after the evaluation learning and pruning in steps S14 and S15. The computational amount of model M C [i] is equal to or less than the target computational amount. After step S15, the process proceeds to step S16.

[0057] In step S16, the controller 11 performs performance evaluation on the i-th model (i.e., model M C [i]) after the evaluation learning and pruning, and derives a performance index P C according to the result of the performance evaluation. The performance index P C derived for the i-th model is particularly referred to as performance index P C [i]. The performance index P C [i] has a value corresponding to the performance of model M C [i]. For any model, the performance of the model is understood to refer to the generalization performance of the model. Here, it is assumed that the higher the performance of model M C [i] is evaluated, the higher the value of the performance index P C [i] becomes. The performance evaluation in step S16 is performed using the evaluation data 22 (see FIG. 1).

[0058] Assume that object detection is performed as an inference as described above. At this time, the evaluation data 22 is a second image data set (not shown) including image data and correct answer data of a plurality of evaluation images. In the second image data set, the correct answer data is added to each evaluation image for the evaluation image. Each evaluation image in the second image data set is a two-dimensional image including images of various types of objects. The correct answer data for a certain evaluation image includes information indicating the type of object existing in the evaluation image and information specifying the position and shape of the region where the object image exists in the evaluation image.

[0059] The controller 11 according to step S16 supplies the image data of the evaluation image to the model M C [i] as model input data Din to cause the model M C [i] to perform inference. The inference performed in step S16 is particularly referred to as evaluation inference C, and the evaluation inference C in the model M C [i] is particularly referred to as evaluation inference C[i]. The controller 11 evaluates the performance of the model M C [i] based on the result of the evaluation inference C[i].

[0060] The performance index P C [i] may be the inference accuracy in the evaluation inference C[i]. At this time, the higher the inference accuracy in the evaluation inference C[i], the higher the value of the performance index P C [i]. Alternatively, the performance index P C [i] may be the inference error in the evaluation inference C[i]. At this time, the performance index P C [i] has a value corresponding to the error between the correct answer data in the evaluation data 22 and the model output data Dout of the model M C [i] in the evaluation inference C[i], and the smaller the error, the higher the value of the performance index P C [i]. In addition, any index representing the performance of the model M C [i] may be adopted as the performance index P C [i].

[0061] Still, after step S15 and before step S16, learning equivalent to the evaluation learning may be performed on the i-th model. That is, after the pruning in step S15, the i-th model obtained by performing the learning equivalent to the evaluation learning on the i-th model again is the model M C [i]. After step S16, proceed to step S17.

[0062] In step S17, the controller 11 determines whether a predetermined search end condition is satisfied. For example, when the variable i reaches a predetermined value, the search end condition is satisfied. Or, for example, when a predetermined evaluation upper limit time has elapsed after the start of the model evaluation process, the search end condition is satisfied. Further, for example, when the performance of the i-th model evaluated in step S16 satisfies the reference performance, the search end condition is satisfied. That the performance of the i-th model evaluated in step S16 satisfies the reference performance means that the performance of the i-th model evaluated in step S16 is equal to or higher than the reference performance. The performance index P C [i] is equal to or higher than the reference performance threshold corresponding to the reference performance, the performance of the i-th model evaluated in step S16 satisfies the reference performance. Further, for example, when the total number of models determined to satisfy the reference performance in the performance evaluation of step S16 reaches a predetermined number (for example, 3), the search end condition may be satisfied. The reference performance and the reference performance threshold are set in advance, and the information indicating them is held in the memory 12.

[0063] If the search end condition is not satisfied in step S17 (No in step S17), add 1 to the variable i in step S18, then return to step S12, and execute the processing after step S12 again. If the search end condition is satisfied in step S17 (Yes in step S17), end the model evaluation process in FIG. 9. The value of the variable i at the time when the search end condition is satisfied corresponds to the value of n.

[0064] In the second and subsequent executions of step S12, the performance index P derived in step S16 CSampling of the model may be performed based on this. Actually, for example, in the first to tenth steps S12, the first to tenth models having the first to tenth structures are sampled by randomly determining the structure of the model. That is, for example, in step S12 when "1 ≦ i ≦ 10" is satisfied, the i-th model having the i-th structure is sampled by randomly determining the structure of the model. Then, in step S12 when "11 ≦ i" is satisfied, based on the structures of the first to (i - 1)-th models and the already-derived performance indicators P C [1] to P C [i - 1], the structure of the i-th model (i.e., the i-th structure) is set according to the search strategy. At this time, in comparison with the first to (i - 1)-th models, the structure in which an improvement in the performance indicator P C is expected is set as the i-th structure for the i-th model.

[0065] In the model evaluation process of FIG. 9, the evaluation unit process composed of the processes of steps S14 to S16 is executed for a plurality of learning models (the first to n-th models) having different structures from each other. In the evaluation unit process, after the evaluation learning for learning the learning model based on the learning data 21, the performance of the learning model is evaluated after pruning the learning model.

[0066] FIG. 10 conceptually shows the search space by NAS. Now, let NN120 be composed of the first to J-th layers (J is an integer of 2 or more). For each of the first to J-th layers, a plurality of candidate structures are defined in the search space information as candidates for the layer structure. The solid line with an arrow in FIG. 10 represents an example of the selection of candidate structures by NAS. In NAS, through repeated sampling of the model, candidate structures for each layer that contribute to improving the accuracy are selected, and finally, the structure of NN120 that can achieve high inference accuracy is specified.

[0067] Referring to FIG. 8, the operation after the model evaluation process will be described. When the model evaluation process in step S10 ends, a transition from step S10 to step S20 occurs. At the stage of transitioning to step S20, for the first to n-th models, the performance indicators P C [1] to PC [n] has been derived.

[0068] In step S20, the controller 11 selects one of the first to nth models as the target model based on the performance indicators P C [1] to P C [n]. At this time, the controller 11 selects one of the performance indicators P C [1] to P C [n], identifies the maximum performance indicator P C , and selects the model corresponding to the maximum performance indicator P C as the target model. Therefore, for example, among the performance indicators P C [1] to P C [n], if the performance indicator P C [n - 1] is the maximum, the (n - 1)th model is selected as the target model. Similarly, for example, among the performance indicators P C [1] to P C [n], if the performance indicator P C [n - 2] is the maximum, the (n - 2)th model is selected as the target model.

[0069] The performance indicator P C [i] represents the performance of the ith model after the evaluation learning and pruning. Therefore, in step S20, the performance of the models after the evaluation learning and pruning is compared among the first to nth models, and the target model is selected based on the comparison result. Specifically, the controller 11 selects, as the target model from among the first to nth models, the model whose performance after the evaluation learning and pruning is the maximum. Note that the pruning in the description "after the evaluation learning and pruning" refers to the pruning in step S15.

[0070] For example, when the performance indicator P C [i] is represented by the inference accuracy in the evaluation inference C[i], the model with the maximum inference accuracy among the inference accuracies of the first to nth models is selected as the target model from among the first to nth models. The inference accuracies of the first to nth models referred to in this selection refer to the inference accuracies in the evaluation inferences C[1] to C[n]. Or, for example, the performance indicator PC When [i] is represented by the inference error in the inference C[i] for evaluation, among the inference errors of the first to n-th models, the model with the minimum inference error is selected as the target model from among the first to n-th models. The inference errors of the first to n-th models referred to in this selection refer to the inference errors in the inferences C[1] to C[n] for evaluation. The same applies when indicators other than inference accuracy and inference error are used.

[0071] Referring to FIG. 11, the influence of pruning on the calculation amount and performance will be described. For convenience of explanation, among the first to n-th models, any two models are referred to as model P and model Q.

[0072] In FIG. 11, plot 611 represents the calculation amount and performance of model P after the learning for evaluation in step S14 and before the pruning in step S15. Plot 612 represents the calculation amount and performance of model P after the pruning in step S15 (that is, the calculation amount and performance of model P at the stage of step S16). Plot 621 represents the calculation amount and performance of model Q after the learning for evaluation in step S14 and before the pruning in step S15. Plot 622 represents the calculation amount and performance of model Q after the pruning in step S15 (that is, the calculation amount and performance of model Q at the stage of step S16).

[0073] Before pruning, both model P and model Q satisfy the reference performance. The reference performance is the performance corresponding to the target performance to be finally achieved through this learning. However, since the number of learning times of the learning for evaluation is less than that of this learning, the reference performance is set lower than the target performance (however, the reference performance may be the same as the target performance). In step S15 of FIG. 9, pruning is performed so that the calculation amount of each of model P and model Q becomes less than or equal to the target calculation amount. After this pruning, the performance of model P falls below the reference performance, while the performance of model Q is maintained above the reference performance. In the example of FIG. 11, model Q is set as the target model prior to model P.

[0074] After step S20, proceed to step S21. In step S21, the controller 11 sets the target model to model 110 and trains the target model using the training data 21. That is, for example, when the target model is the i-th model, a model 110 having an NN120 of the i-th structure is constructed as the target model. The training in step S21 is this main training. The target model to which this main training is applied has the structure before the pruning in step S15. That is, this main training is executed on the target model before pruning. Therefore, when the i-th model is selected as the target model, the main training is executed on the model M A [i] or M B [i]. It can be said that the target model is selected from among the first to n-th models before the pruning in step S15.

[0075] Specifically, in step S21, the controller 11 extracts a mini-batch from the training data 21. The mini-batch is data having a predetermined mini-batch size within the training data 21. For example, when the training data 21 includes image data of 100,000 training images, the mini-batch includes image data of 10 training images therein and also includes correct answer data for the 10 training images. In this main training, the controller 11 inputs the image data of the training images included in the mini-batch as model input data Din to the target model. Thereby, the target model generates model output data Dout based on the current training parameters. The correct answer data included in the mini-batch corresponds to the correct answer data for the model output data Dout.

[0076] The controller 11 derives a loss function representing the error between the model output data Dout generated based on the mini-batch and the correct answer data included in the mini-batch. Further, the controller 11 derives the gradient of the loss function of the target model by the error backpropagation method, and updates the learning parameters of the target model based on the derived gradient. Thereafter, a new mini-batch is extracted. Each time a mini-batch is extracted, the learning parameters of the target model are updated. In this learning, a series of processes starting from the extraction of the mini-batch and updating the learning parameters of the target model are repeated. When the series of processes is performed a predetermined number of times or for a predetermined time for the target model, this learning for the target model ends, and a transition from step S21 to step S22 occurs. Here, as described above, the number of learning times in this learning is larger than the number of learning times in the learning for evaluation. This learning may end when the loss function of the target model converges.

[0077] In the following, the target model before this learning is performed may be referred to as the target model M α or model M α and the target model immediately after this learning is performed may be referred to as the target model M β or model M β In step S21, by performing this learning on the target model M α the target model M β is generated. In other words, by performing this learning on the target model, the target model is updated from model M α to model M β .

[0078] In step S22, the controller 11 performs performance evaluation of the target model after this learning and before pruning (that is, the target model M β ), and derives a performance index P β according to the result of the performance evaluation. The performance index P β has a value according to the performance of the target model M β . Here, the higher the performance of the target model M β is evaluated, the larger the performance index P βLet the value increase. The performance evaluation in step S22 is performed using the evaluation data 22 (see FIG. 1). As described above, it is assumed that the evaluation data 22 is a second image data set (not shown) including the image data and correct answer data of a plurality of evaluation images.

[0079] The controller 11 according to step S22 supplies the image data of the evaluation image to the target model M β as the model input data Din to cause the target model M β to perform inference. The inference performed in step S22 is particularly referred to as the evaluation inference β. The controller 11 evaluates the performance of the target model M β based on the result of the evaluation inference β.

[0080] The performance index P β may be the inference accuracy in the evaluation inference β. At this time, the higher the inference accuracy in the evaluation inference β, the higher the value of the performance index P β . Alternatively, the performance index P β may be the inference error in the evaluation inference β. At this time, the performance index P β has a value corresponding to the error between the correct answer data in the evaluation data 22 and the model output data Dout of the target model M β in the evaluation inference β, and it is assumed that the smaller the error, the higher the value of the performance index P β . In addition, any index representing the performance of the target model M β may be adopted as the performance index P β . After step S22, the process proceeds to step S23.

[0081] In step S23, the controller 11 performs pruning on the target model after the current learning based on the learning parameters of the target model after the current learning. That is, the controller 11 performs pruning on the target model M β based on the learning parameters of the target model M β .

[0082] In step S23, for the target model M βAmong the connections between a large number of nodes existing in the NN120 of, the target model M β connections (connections between nodes) that have little or no impact on the model output data Dout of are identified as unnecessary connections. In step S23, the unnecessary connections in the target model M β are deleted. The process of deleting unnecessary connections is pruning.

[0083] An explanation of the pruning in step S23 will be added. In the explanation of the pruning in step S23, the target model refers to the target model after this learning (that is, the target model M β ). The controller 11 according to step S23 sets a threshold value TH A . Then, the controller 11 extracts the weights in the NN120 of the target model that are less than the threshold value TH A , and identifies the connections (connections between nodes) corresponding to the weights less than the threshold value TH A as unnecessary connections. The connection corresponding to the weight less than the threshold value TH A refers to the connection between the node with the weight less than the threshold value TH A set and the subsequent-stage node connected to the node. The controller 11 first identifies unnecessary connections with the threshold value TH A set to a predetermined initial small value, and performs initial pruning to delete the identified unnecessary connections. The controller 11 evaluates the computational amount of the target model after the initial pruning. Thereafter, the controller 11 gradually increases the threshold value TH A by a predetermined step value. Each time the threshold value TH A increases, the location of the identified unnecessary connection also increases. The controller 11 sets the threshold value TH AAs the degree of increase, perform stage pruning to newly identify unnecessary connections and delete the newly identified unnecessary connections. Each time the controller 11 deletes a newly identified unnecessary connection, it evaluates the computational complexity of the target model after the deletion. The above operations are repeated until the computational complexity of the target model becomes equal to or less than the target computational complexity. When the computational complexity of the target model becomes equal to or less than the target computational complexity, the pruning in step S23 ends. Therefore, the computational complexity of the target model at the end of step S23 is equal to the target computational complexity or slightly lower than the target computational complexity. Note that the deletion of unnecessary connections may be realized by setting zero to the weight (threshold TH A less than the weight) in the unnecessary connection.

[0084] Any two nodes provided in the target model before pruning and connected to each other are referred to as the first target node and the second target node. The second target node is arranged after the first target node. The first target node generates its output data by performing an operation according to the learning parameter set for itself on the input data to itself. Before pruning, the output data of the first target node is input to the second target node. In step S23, the controller 11 compares the magnitude of the weight set for the first target node, which is the weight used when generating the input data from the first target node to the second target node, with the threshold TH A . If the magnitude of the weight is less than the threshold TH A , it is determined that the connection between the first and second target nodes is an unnecessary connection. At this time, the deletion of the connection between the first and second target nodes may be realized by setting zero to the weight of the first target node when generating the input data from the first target node to the second target node. If the magnitude of the weight is greater than or equal to the threshold TH A , it is not determined that the connection between the first and second target nodes is an unnecessary connection. Note that if the weight set for the first target node is a vector quantity, the magnitude of the weight represents the norm of the weight. The same processing is performed for all combinations of two nodes provided in the target model before pruning and connected to each other.

[0085] The amount of computation of the target model represents the amount of computation performed by the target model when making one inference with the target model. That is, the amount of computation of the target model, in the case of image recognition, represents the amount of computation performed by the target model from when the image data of one input image is input into the target model until the model output data Dout based on the input image is output from the target model. Since the method itself for deriving the amount of computation is known, detailed description thereof is omitted. The target amount of computation is set in advance, and information indicating the target amount of computation is held in the memory 12.

[0086] Hereinafter, the target model immediately after the pruning in step S23 is the target model M γ or model M γ and may be referred to as such. In step S23, pruning is performed on the target model M β to generate the target model M γ That is, by performing pruning on the target model after this learning, unnecessary connections in the structure of the target model are deleted, and as a result, the target model is changed from model M β to model M γ The target model M γ is the target model after this learning and pruning in steps S21 and S23. The amount of computation of model M γ is equal to or less than the target amount of computation. After step S23, the process proceeds to step S24.

[0087] In step S24, the controller 11 uses the learning data 21 to relearn the target model after this learning and pruning. The learning in step S24 is referred to as relearning. The target model to which relearning is applied has the structure after the pruning in step S23. That is, the relearning is for the target model M γIt is executed for. The method of re-learning is the same as the method of this learning. The number of learning times in re-learning and the number of learning times in this learning may coincide with each other or may be different from each other. Re-learning may be terminated when the loss function of the target model converges. The appropriate value of the learning parameter changes due to pruning. During re-learning, the learning parameter is re-adjusted so as to be suitable for the structure after pruning.

[0088] In addition, hereinafter, the target model immediately after re-learning is the target model M δ or model M δ and may be referred to as. In step S24, re-learning is performed on the target model M γ so that the target model M δ is generated. In other words, by performing re-learning on the target model, the target model is updated from model M γ to model M δ . After step S24, the process proceeds to step S25.

[0089] In step S25, the controller 11 performs performance evaluation of the target model (that is, the target model M δ ) after this learning, pruning, and re-learning, and derives a performance index P δ according to the result of the performance evaluation. The performance index P δ has a value corresponding to the performance of the target model M δ . Here, it is assumed that the higher the performance of the target model M δ is evaluated, the higher the value of the performance index P δ . The performance evaluation in step S25 is performed using the evaluation data 22 (see FIG. 1). As described above, the evaluation data 22 is assumed to be a second image data set (not shown) including image data and correct answer data of a plurality of evaluation images.

[0090] The controller 11 according to step S25 supplies the image data of the evaluation image to the target model M δ as the model input data Din, so that the target model M δLet it perform inference. The inference performed in step S25 is particularly referred to as evaluation inference δ. The controller 11 evaluates the performance of the target model M δ based on the result of the evaluation inference δ.

[0091] The performance index P δ may be the inference accuracy in the evaluation inference δ. In this case, the higher the inference accuracy in the evaluation inference δ, the higher the value of the performance index P δ . Alternatively, the performance index P δ may be the inference error in the evaluation inference δ. In this case, the performance index P δ has a value corresponding to the error between the correct data in the evaluation data 22 and the model output data Dout of the target model M δ in the evaluation inference δ, and the smaller the error, the higher the value of the performance index P δ . In addition, any index representing the performance of the target model M δ may be adopted as the performance index P δ . After step S25, the process proceeds to step S26.

[0092] In step S26, the controller 11 determines whether the performance of the target model M δ meets the target performance. That the performance of the target model M δ meets the target performance means that the performance of the target model M δ evaluated in step S25 is equal to or higher than the target performance. When the performance index P δ is equal to or higher than the target performance threshold corresponding to the target performance, the performance of the target model M δ meets the target performance. The target performance and the target performance threshold are set in advance, and the information indicating them is held in the memory 12.

[0093] In step S26, when the performance of the target model M δ meets the target performance (Yes in step S26), the operation of FIG. 8 ends. The target model M δis a learned model to be finally generated by the learning device 10, and is an inference model that performs a predetermined inference (for example, inference for object detection) based on the model input data Din. In step S26, the target model M δ When the performance of does not meet the target performance (No in step S26), the process returns from step S26 to step S20 and the processes after step S20 are repeated.

[0094] However, when returning from step S26 to step S20, a model different from the model selected as the target model in the first step S20 is selected as the target model in the second step S20. That is, for example, in the first step S20, among the performance indicators P C [1] to P C [n], consider the case where the performance indicator P C [n - 1] is the largest and thus the (n - 1)th model is selected as the target model. In this case, when returning from step S26 to step S20, in the second step S20, the controller 11 determines the performance indicators P C [1] to P C [n], and identifies the second largest performance indicator P C . Then, the model corresponding to the second largest performance indicator P C is selected as the new target model. For example, if the second largest performance indicator P C is the performance indicator P C [n - 2], then in the second step S20, the (n - 2)th model is selected as the new target model. The same applies to the third and subsequent steps S20.

[0095] The fact that the determination result in step S26 is negative (No) and the process returns to step S20 corresponds to the above-mentioned backtracking. However, in the learning device 10, the occurrence probability of the above-mentioned backtracking is not high in comparison with the reference example (see FIG. 6). This is because in the method according to the present embodiment, the selection of the target model is performed in consideration of the influence of pruning. That is, at the stage of the model evaluation process in step S10, the performance of each model after pruning is evaluated, and a model that maintains appropriate performance even after pruning is selected as the target model. If the number of learning times of the evaluation learning is sufficiently close to the number of learning times of the present learning, the change tendency of the performance of the i-th model when pruning is performed after the evaluation learning and the change tendency of the performance of the i-th model when pruning is performed after the present learning are sufficiently approximated. As the number of learning times of the evaluation learning decreases, although a difference occurs between the former change tendency and the latter change tendency, it is considered that there is a correlation between the former change tendency and the latter change tendency for models having the same structure. Therefore, a model having relatively high performance after pruning through the evaluation learning is likely to have relatively high performance even after pruning through the present learning.

[0096] In this way, the controller 11 executes the evaluation unit process for a plurality of learning models (the first to n-th models) having different structures from each other. In the evaluation unit process, after the evaluation learning in which the learning model is learned based on the learning data 21, the performance of the learning model is evaluated through the first pruning for the learning model (steps S14 to S16). Here, the pruning in step S15 is referred to as the first pruning for convenience. Then, the controller 11 is based on the result of the evaluation in the evaluation unit process (performance index P C [1] to P CBased on [n], a target model is selected from a plurality of learning models (the first to the nth models) before the first pruning (step S20). Then, the controller 11 executes this learning to train the target model with a number of learning times greater than that of the evaluation learning using the learning data 21 (step S21). The target model to which this learning is applied is the target model before any pruning is performed. After this learning, the controller 11 performs second pruning on the target model (step S23), thereby generating an inference model (target model M δ ) that satisfies the target performance. Here, the pruning in step S23 is referred to as the second pruning for convenience.

[0097] By this method, a model that is expected to have high performance even after the second pruning is more likely to be selected as the target model. Conversely, a model whose performance does not meet the target due to the second pruning is less likely to be selected as the target model. For this reason, the possibility of repeatedly performing this learning on a plurality of models is suppressed (the occurrence of backtracking is suppressed). That is, a desired inference model (target model M δ ) that satisfies the target performance can be generated efficiently (with less computational cost or time cost).

[0098] Refer to FIG. 12. The performance change from plot 611a to plot 612a in FIG. 12 shows the performance change of the target model before and after pruning assuming that the above-described model P (refer to FIG. 11) is selected as the target model. The performance change from plot 621a to plot 622a in FIG. 12 shows the performance change of the target model before and after pruning assuming that the above-described model Q (refer to FIG. 11) is selected as the target model. Note that the performance of the target model after pruning refers to the performance evaluated in step S25.

[0099] The impact of the first pruning (the pruning in step S15) on models P and Q is as described above with reference to FIG. 11. In the example of FIG. 11, since the performance of model Q after the first pruning is higher than that of model P after the first pruning, if only models P and Q are considered, model Q is selected as the target model. When the performance of model Q is higher than that of model P after the first pruning, it is highly likely that the performance of model Q is also higher than that of model P after the second pruning. Therefore, selecting model Q as the target model rather than model P makes it less likely for the above-mentioned setback to occur.

[0100] In the evaluation unit process, for each learning model, the controller 11 causes the evaluation learning using the evaluation data 22 and the evaluation inference C in step S16 to be performed on the learning model after the first pruning. Then, based on the results of the evaluation inference C (performance indicators P C [1]~P C [n]) for each learning model, the controller 11 compares the performance of the learning models after the evaluation learning and the first pruning among a plurality of learning models (the first to the nth models). The controller 11 selects a target model based on the comparison result (step S20).

[0101] As a result, based on the performance levels of each learning model after the first pruning, it is possible to select, as the target model, a learning model that is expected to have high performance even after the second pruning in this learning. Therefore, the occurrence of setbacks is suppressed, and a desired inference model can be efficiently generated.

[0102] Specifically, the controller 11 identifies, among a plurality of learning models (first to nth models), the learning model that has the maximum performance after the evaluation learning and the first pruning. Then, the controller 11 may select the identified learning model (the learning model before the first pruning which is the identified learning model) as the target model (step S20). Thereby, even if the second pruning is performed after this learning, a learning model expected to have high performance (ideally, maximum performance) can be selected as the target model. For this reason, the occurrence of going back is suppressed, and a desired inference model can be efficiently generated.

[0103] More specifically, the performance of each learning model may be evaluated by the inference accuracy or the inference error in the evaluation inference C in step S16. The controller 11 identifies, from among a plurality of learning models (first to nth models), the learning model with the maximum inference accuracy or the learning model with the minimum inference error. Then, the controller 11 may select the identified learning model (the learning model before the first pruning which is the identified learning model) as the target model. After the first pruning, the learning model with the maximum inference accuracy or the learning model with the minimum inference error is expected to have high performance even if the second pruning is performed after this learning. As a result, by the above selection method, the occurrence of going back is suppressed, and a desired inference model can be efficiently generated.

[0104] In the evaluation unit process, the controller 11 performs the first pruning for each learning model (that is, for each of the first to nth models) so that the amount of calculation by the learning model after the evaluation learning and the first pruning is equal to or less than the target amount of calculation (step S15). Also, the controller 11 performs the second pruning so that the amount of calculation by the target model after this learning and the second pruning is equal to or less than the target amount of calculation (step S23).

[0105] As a result, the situation when the second pruning is performed after this learning can be simulated by the first pruning. As a result of the simulation, a learning model that is expected to maintain high performance even after pruning can be selected as the target model. A learning model that has high performance after the first pruning is expected to have high performance even when the second pruning is performed after this learning. Therefore, the occurrence of backtracking is suppressed, and a desired inference model can be efficiently generated.

[0106] Hereinafter, in a plurality of embodiments, supplementary matters, applied technologies, modified technologies, etc. for some of the above-described operations or configurations will be described. The matters described above in this embodiment are applied to the following respective embodiments unless otherwise specified and without contradiction. In each embodiment, if there is a matter conflicting with the above-described matters, the description in each embodiment may be given priority. Also, without contradiction, among the plurality of embodiments shown below, the matters described in any one of the embodiments can be applied to any other embodiment (that is, it is also possible to combine any two or more of the plurality of embodiments).

[0107] <<First Embodiment>> The first embodiment will be described. In step S20, the controller 11 may select, as the target model, a model corresponding to the maximum performance index P among the models that satisfy the reference performance among the first to nth models. That is, for example, among the performance indexes P C [1] to P C [n], if only the performance indexes P C [n - 3] to P C [n] are equal to or higher than the reference performance threshold value, only the (n - 3)th to nth models among the first to nth models satisfy the reference performance. In this case, among the performance indexes P C [n - 3] to P C [n] corresponding to the (n - 3)th to nth models, the maximum performance index P C is specified, and the model corresponding to the maximum performance index P C may be selected as the target model. C

[0108] If, in step S20, none of the first to nth models satisfy the reference performance (that is, when all of the performance indicators P C [1] to P C [n] are less than the reference performance threshold), it may be possible to return from step S20 to step S10. When returning to step S10, the controller 11 performs the model evaluation process again.

[0109] <<Second Embodiment>> The second embodiment will be described. In step S15 of FIG. 9, the controller 11 according to the second embodiment performs pruning of the model M B [i] under fixed conditions.

[0110] As a method of performing pruning of the model M B [i] under fixed conditions, there is a method MTD_2A of performing pruning at a fixed pruning rate. In the method MTD_2A, among all the connections between a plurality of nodes existing in the NN120 of the model M B [i], the ratio of unnecessary connections is made to match the fixed pruning rate. That is, for example, if the fixed pruning rate is 10%, among all the connections between a plurality of nodes existing in the NN120 of the model M B [i], unnecessary connections are identified so that 10% of the connections are identified as unnecessary connections. Starting from the ones with smaller corresponding weights in order, they are identified as unnecessary connections, and when the ratio of unnecessary connections reaches the fixed pruning rate, the additional identification of unnecessary connections may be terminated. In step S15, the unnecessary connections in the model M B [i] are deleted.

[0111] As a method of performing pruning of the model M B [i] under fixed conditions, there is a method MTD_2B of performing pruning at a fixed threshold TH CNST . The controller 11 according to the method MTD_2B extracts the weights in the NN120 of the model M B [i] that are less than the fixed threshold TH CNST , and the fixed threshold TH CNSTIdentify the connections corresponding to weights less than the fixed threshold TH as unnecessary connections (connections between nodes). CNST The connection corresponding to a weight less than the fixed threshold TH CNST refers to the connection between a node with a weight less than the fixed threshold TH set and the subsequent node connected to the node. Fixed threshold TH CNST is invariant and is preset in the learning device 10. In step S15, the model M B [i] The unnecessary connections are deleted.

[0112] When using method MTD_2A or MTD_2B, the model M which is the i-th model after pruning C [i] The computational complexity may exceed the target computational complexity or may be less than or equal to the target computational complexity. For model M C [i] In the case where the computational complexity of exceeds the target computational complexity, assume that the model M C [i] is selected as the target model and pruning is performed on the target model through this learning. A conceptual diagram of this case is shown in FIG. 13. In FIG. 13, plot 651 represents the performance of the target model before pruning. Plot 652 represents the performance of the target model when it is assumed that the target model is pruned by approximately the same amount of pruning as in step S15. Plot 653 represents the performance of the target model when it is assumed that the target model is pruned by approximately the same amount of pruning as in step S23. In this case, there is a possibility that the performance of the target model will drop beyond the allowable range due to the pruning in step S23. This is because in step S23 after this learning, pruning is performed so that the computational complexity of the target model becomes less than or equal to the target computational complexity, and as a result, it is expected that the amount of pruning in step S23 will be larger than the amount of pruning in step S15.

[0113] Therefore, when using method MTD_2A or MTD_2B, for each integer i satisfying "1 ≦ i ≦ n", it is preferable to check in step S20 whether the computational complexity of model M C [i] exceeds the target computational complexity. And for each integer i satisfying "1 ≦ i ≦ n", for model M COnly when the computational complexity of [i] is less than or equal to the target computational complexity, it is advisable for the controller 11 to include the i-th model as a candidate for the target model. That is, for model M C When the computational complexity of [i] exceeds the target computational complexity, it is advisable for the controller 11 to exclude the i-th model from the candidates for the target model so that the i-th model will not be selected as the target model.

[0114] <<Third Embodiment>> The third embodiment will be described. The selection of the target model in step S20 may be performed by the following selection method MTD_3A or MTD_3B. The selection methods MTD_3A and MTD_3B can be combined with the first or second embodiment.

[0115] The controller 11 according to the selection method MTD_3A identifies, as candidate models, one or more models that satisfy the reference performance among the first to n-th models after the evaluation learning and pruning in steps S14 and S15. Thereafter, the controller 11 according to the selection method MTD_3A selects, as the target model, the one with the smallest computational complexity from among the identified one or more candidate models. Here, the computational complexity refers to the computational complexity of the model (M C [i]) after the evaluation learning and pruning in steps S14 and S15.

[0116] For example, in the selection method MTD_3A, consider the case where only the performance indicators P C [1] to P C [n], among which only the performance indicators P C [n - 3] to P C [n] are greater than or equal to the reference performance threshold. In this case, among the first to n-th models after the evaluation learning and pruning in steps S14 and S15, only the (n - 3)-th to n-th models satisfy the reference performance. Therefore, the (n - 3)-th to n-th models become the candidate models. And in this case, the controller 11 compares the computational complexities of the (n - 3)-th to n-th models and identifies the smallest computational complexity among them. The computational complexities compared here are the models M C [n - 3] to M CIt is the computational complexity of [n]. For example, if the minimum computational complexity is that of the (n - 3)th model, the (n - 3)th model is selected as the target model; if the minimum computational complexity is that of the (n - 2)th model, the (n - 2)th model is selected as the target model.

[0117] After the evaluation learning and pruning in steps S14 and S15, the controller 11 according to the selection method MTD_3B identifies one or more models among the first to nth models that meet the reference performance and have a computational complexity less than or equal to the target computational complexity. Then, the controller 11 selects, as the target model, the one with the maximum performance or the minimum computational complexity from among the identified one or more models. Alternatively, the controller 11 according to the selection method MTD_3B selects the target model from among the identified one or more models in consideration of both the performance and the computational complexity of each identified model. Here, the computational complexity refers to the computational complexity of the model (M C [i]) after the evaluation learning and pruning in steps S14 and S15.

[0118] Assume the following cases in the selection method MTD_3B. In the assumed cases, among the performance indicators P C [1] to P C [n], only the performance indicators P C [n - 3] to P C [n] are greater than or equal to the reference performance threshold. In addition, in the assumed cases, among the computational complexities of the first to nth models after the evaluation learning and pruning in steps S14 and S15, only the computational complexities of the (n - 2)th to nth models are less than or equal to the target computational complexity. That is, among the computational complexities of the models M C [1] to M C [n], only the computational complexities of the models M C [n - 2] to M C [n] are less than or equal to the target computational complexity. Then, only the (n - 2)th to nth models are identified as candidate models. In this case, the controller 11 selects the target model from among the (n - 2)th to nth models based on a comparison of the performances of the (n - 2)th to nth models. The performances of the (n - 2)th to nth models here are the performance indicators P C [n - 2] to P CIt is represented by [n]. Alternatively, the controller 11 selects a target model from among the (n - 2)th to nth models based on a comparison of the computation amounts of the (n - 2)th to nth models. The computation amounts of the (n - 2)th to nth models here refer to the computation amounts corresponding to models M C [n - 2] to M C [n]. Further alternatively, the controller 11 selects a target model from among the (n - 2)th to nth models based on the performance of the (n - 2)th to nth models and the computation amounts of the (n - 2)th to nth models.

[0119] <<Fourth Embodiment>> The fourth embodiment will be described. The controller 11 may select a target model in consideration of the amount of change in performance before and after the pruning in step S15. In this case, for each of the first to nth models, the performance of the i-th model is also evaluated before the pruning in step S15. Thus, by combining with the performance evaluation in step S16, the amount of change in performance before and after the pruning in step S15 can be obtained. The amount of change in performance is obtained for each of the first to nth models.

[0120] The amount of change in performance in the i-th model is denoted as the performance change amount Δ[i]. The performance change amount Δ[i] is the amount of change in the performance of model M B [i] as seen from the performance of model M C [i] (see FIG. 9). Assuming that the performance of the model always decreases during the pruning in step S15, it is assumed that "Δ[i]>0". Then, the performance change amount Δ[i] is the amount of decrease in the performance of the i-th model due to performing the pruning in step S15 on the i-th model. For example, when the performance of each model is represented by the inference accuracy, the result obtained by subtracting the inference accuracy of model M B [i] from the inference accuracy of model M C [i] is the performance change amount Δ[i].

[0121] In step S20, the controller 11 according to the fourth embodiment selects a target model in consideration of the performance change amounts Δ[1] to Δ[n]. At this time, the larger the performance change amount is, the more difficult it is to be selected as the target model. Specifically, for example, the performance index P C [n - 1] of the (n - 1)-th model and the performance index P C [n] of the n-th model are the same, and the calculation amount of the model M C [n - 1] and the calculation amount of the model M C [n] are the same. In this case, the controller 11 compares the performance change amounts Δ[n - 1] and Δ[n]. Then, if "Δ[n - 1]>Δ[n]", the controller 11 preferentially selects the n-th model as the target model over the (n - 1)-th model, and if "Δ[n - 1]<Δ[n]", the controller 11 may preferentially select the (n - 1)-th model as the target model over the n-th model.

[0122] When a model with a relatively large performance change amount is selected as the target model, there is a concern that the performance of the target model will be significantly reduced by the pruning after this learning. This concern can be particularly prominent when the method shown in the fourth embodiment is combined with the second embodiment. When a model with a large performance change amount between the plots 651 and 652 in FIG. 13 is selected as the target model, the performance of the target model may be significantly reduced to the performance corresponding to the plot 653 by the pruning after this learning. Therefore, the method shown in the fourth embodiment is particularly preferably combined with the second embodiment.

[0123] <<Fifth Embodiment>> The fifth embodiment will be described. The performance evaluation in step S22 of FIG. 8 is executed to estimate how the performance of the target model has changed before and after pruning for the target model. That is, the controller 11 can recognize the influence of pruning on the performance of the target model by comparing the performance index P β derived in step S22 and the performance index P δ derived in step S25. However, for a target model (M δ) The execution of step S22 is not essential for the generation of . Therefore, the process of step S22 may be omitted from the operation of FIG. 8.

[0124] <<Sixth Embodiment>> The sixth embodiment will be described. FIG. 14 shows a functional block diagram of the controller 11. Functional blocks F1 to F6 are provided in the controller 11. By causing the controller 11 to execute a program recorded in the memory 12, the database 20, or any other arbitrary recording medium (not shown), each function of the functional blocks F1 to F6 may be realized. The relationship between each functional block and the flowcharts of FIGS. 8 and 9 will be described.

[0125] Functional block F1 is a model evaluation unit that performs the model evaluation process of step S10. Functional block F2 is a target model selection unit that performs the process of step S20. Functional block F3 is a main learning processing unit that performs the main learning of step S21. Functional block F4 is a pruning processing unit that performs the pruning of step S23. The pruning processing unit F4 may further perform the pruning of step S15. Functional block F5 is a relearning processing unit that performs the relearning of step S24. Functional block F6 is a performance evaluation unit that performs each performance evaluation of steps S22 and S25. The performance evaluation unit F6 may further perform the performance evaluation of step S16.

[0126] <<Seventh Embodiment>> The seventh embodiment will be described. The target model M at the stage when the operation of FIG. 8 is completed δ is a learned model to be finally generated by the learning device 10, and is an inference model that performs a predetermined inference (for example, inference for object detection) based on the model input data Din. The inference model can be applied to any device (hereinafter referred to as a model application device). More specifically, the inference model is mounted on hardware such as an SoC (System-on-a-chip) provided in the model application device. At this time, the target performance and the target calculation amount according to the specifications of the model application device are defined. In order to satisfy the target performance and the target calculation amount defined here, the operations of FIGS. 8 and 9 are executed.

[0127] The model application device may be an in-vehicle device mounted on a vehicle such as an automobile. In this case, for example, the inference model performs object detection as an inference, and the in-vehicle device performs automatic driving or driving support or the like based on the result of the object detection.

[0128] <<Eighth Embodiment>> The eighth embodiment will be described. Refer to FIGS. 15(a) to (c). A performance estimator 412 for estimating the performance of the model after evaluation learning and pruning may be constructed in advance from the structure of the neural network. Then, in the model evaluation process of FIG. 9, the performance of each model may be evaluated using the performance estimator 412 without actually performing the processes of steps S14 and S15. An explanation will be added to this method.

[0129] In order to obtain the performance estimator 412, first, in an estimator generation device 410 shown in FIG. 15(a), a reference neural network and a reference model (both not shown) are set. The reference neural network is a neural network having a structure similar to that of the NN120. The reference model is a model having a structure similar to that of the model 110 and has the reference neural network. The estimator generation device 410 may be the learning device 10 or any other arbitrary computer device.

[0130] In the estimator generation device 410, reference set data including structure data indicating the structure of the reference neural network and a performance index P CREF indicating the performance of the reference model is set. Although the performance estimator 412 is constructed through learning, the reference set data constitutes a data set used for the learning.

[0131] Performance index P CREFThe derivation method will be described. In the estimator generation device 410, learning for evaluation equivalent to that in step S14 is performed on the reference model, and then pruning equivalent to that in step S15 (hereinafter referred to as pre-stage pruning for convenience) is performed. Then, in the estimator generation device 410, by performing performance evaluation of the reference model that has undergone learning for evaluation and pre-stage pruning, the performance index P CREF is derived. The performance evaluation here is equivalent to the performance evaluation in step S16, and the performance index P CREF has a value corresponding to the performance of the reference model that has undergone learning for evaluation and pre-stage pruning.

[0132] A large number of the above-mentioned reference set data are prepared. The structure of the reference neural network is different for each reference set data. Therefore, for any two reference set data, the structure of the reference neural network in one reference set data and the structure of the reference neural network in the other reference set data are different from each other. The structure data in the reference set data indicates the structure of the reference neural network before pre-stage pruning is performed.

[0133] The estimator generation device 410 has a learning model 411 composed of a neural network. After obtaining a plurality of reference set data, which are a sufficient number of reference set data, the estimator generation device 410 uses the plurality of reference set data to train the learning model 411. The training of the learning model 411 is supervised machine learning, and in the training of the learning model 411, the correct data is the performance index P CREF is.

[0134] In the training of the learning model 411, the learning model 411 executes a process of deriving a performance index indicating the performance of the reference model from the structure data in the reference set data. The estimator generation device 410 compares the performance index derived by the learning model 411 with the performance index P CREFBased on the error with [something], the learning parameters of the learning model 411 are updated by the error backpropagation method. The learning model 411 at the time when the learning of the learning model 411 is completed is a performance estimator 412 that estimates the performance index of an arbitrary neural network from the structural data of the neural network. However, as understood from the data used in the learning of the learning model 411, the performance index estimated by the performance estimator 412 represents the performance of the neural network after performing evaluation learning and pre-pruning on the neural network.

[0135] The created performance estimator 412 is provided to the controller 11 of the learning device 10. The controller 11 according to the eighth embodiment does not execute the processes of steps S14 and S15 in the model evaluation process. The controller 11 according to the eighth embodiment, in step S16, inputs the structural data indicating the structure of the i-th model (specifically, the structure of the model M A [i]) to the performance estimator 412. Thereby, in step S16, the performance estimator 412 derives the performance index of the i-th model from the structural data of the i-th model. The performance index derived here estimates the performance index of the i-th model after actually performing the evaluation learning and pruning of steps S14 and S15 on the i-th model. Therefore, the performance index of the i-th model derived by the performance estimator 412 can be used as the performance index P C [i].

[0136] According to the method of this embodiment, without the need to execute evaluation learning and pruning for each model, the performance index P C [i] can be obtained, and thus the target model can be selected. Actually, for example, the operations shown in FIGS. 8 and 9 are repeated, and the structural data and performance index P C [i] of each model obtained therein are accumulated to form reference set data. After accumulating a sufficient amount of reference set data, the method shown in this embodiment may be implemented.

[0137] <<Ninth Embodiment>> A ninth embodiment will be described. Refer to FIGS. 16(a) to 16(c). A performance estimator 432 for estimating the model performance after the present learning, pruning, and re-learning from the structure of the neural network may be constructed in advance. Then, in step S25 of FIG. 8, the performance estimator 432 may be used to perform a performance evaluation of the target model (specifically, the performance evaluation of the target model M δ ). An explanation will be added to this method.

[0138] In order to obtain the performance estimator 432, first, in an estimator generation device 430 shown in FIG. 16(a), a reference neural network and a reference model (both not shown) are set. The reference neural network is a neural network having a structure similar to that of NN120. The reference model is a model having a structure similar to that of the model 110 and includes the reference neural network. The estimator generation device 430 may be the learning device 10 or any other arbitrary computer device.

[0139] In the estimator generation device 430, reference set data including structure data indicating the structure of the reference neural network and a performance index P δREF indicating the performance of the reference model is set. Although the performance estimator 432 is constructed through learning, the reference set data constitutes a data set used for the learning.

[0140] A method for deriving the performance index P δREF will be described. In the estimator generation device 430, the same present learning as that in step S21 is performed on the reference model, and then, the same pruning as that in step S23 (hereinafter, referred to as post-stage pruning for convenience) is performed. Further, the same re-learning as that in step S24 is performed on the reference model after the present learning and the post-stage pruning. Then, in the estimator generation device 430, by performing a performance evaluation of the reference model that has undergone the present learning, post-stage pruning, and re-learning, the performance index P δREF is derived. The performance evaluation here is the same as the performance evaluation in step S25, and the performance index P δREF has a value corresponding to the performance of the reference model that has undergone the present learning, post-stage pruning, and re-learning.

[0141] A number of the above-mentioned reference set data are prepared. For each reference set data, the structure of the reference neural network is different. Therefore, for any two reference set data, the structure of the reference neural network in one reference set data and the structure of the reference neural network in the other reference set data are different from each other. The structure data in the reference set data indicates the structure of the reference neural network before the subsequent pruning is performed.

[0142] The estimator generation device 430 has a learning model 431 composed of a neural network. After obtaining a plurality of reference set data, which are a sufficient number of reference set data, the estimator generation device 430 uses the plurality of reference set data to train the learning model 431. The training of the learning model 431 is supervised machine learning, and in the training of the learning model 431, the correct data is the performance index P δREF is.

[0143] In the training of the learning model 431, the learning model 431 executes a process of deriving a performance index indicating the performance of the reference model from the structure data in the reference set data. The estimator generation device 430 updates the learning parameters of the learning model 431 by the error backpropagation method based on the error between the performance index derived by the learning model 431 and the performance index P δREF and. When the training of the learning model 431 is completed, the learning model 431 is a performance estimator 432 that estimates the performance index of an arbitrary neural network structure data from the neural network. However, as understood from the data used in the training of the learning model 431, the performance index estimated by the performance estimator 432 represents the performance of the neural network after the present training, subsequent pruning, and retraining for the neural network.

[0144] The created performance estimator 432 is provided to the controller 11 of the learning device 10. The controller 11 according to the ninth embodiment has the structure of the target model in step S25 (specifically, the target model M αInput the structure data indicating the structure) into the performance estimator 432. As a result, in step S25, the performance estimator 432 derives the performance index of the target model from the structure data of the target model. The performance index derived here estimates the performance index of the target model after actually performing the main learning, pruning, and re-learning in steps S21, S23, and S24 on the target model. Therefore, the performance index of the target model derived by the performance estimator 432 is used as the performance index P δ and can be used.

[0145] According to the method of this embodiment, the performance of the target model after main learning, pruning, and re-learning can be estimated without actually performing performance evaluation. Actually, for example, the operations shown in FIGS. 8 and 9 are repeatedly performed, and the structure data and performance index P of the target model obtained therein δ are used to accumulate reference set data. After accumulating a sufficient amount of reference set data, the method shown in this embodiment may be implemented.

[0146] <<Example 10>> Example 10 will be described. In Example 10, modified techniques or applied techniques for the above-mentioned various matters are given.

[0147] Although the method of sampling the first to nth models (setting the first to nth structures) using NAS has been described above, the use of NAS is not essential. As long as the first to nth models are n types of models having different structures from each other, the setting method of the first to nth models is arbitrary. The structures of the first to nth models may be determined based on the input operation information given to the learning device 10 by the operator of the learning device 10.

[0148] The method shown in FIGS. 8 and 9 and executed by the learning device 10 can be referred to as a learning method. The learning method includes a model selection method. The model selection method is configured by the processes of steps S10 and S20, and selects a target model that becomes the basis of the inference model based on the result of the model evaluation process. The device that realizes the model selection method can also be referred to as a model selection device, and the model selection device is included in the learning device 10.

[0149] A program for causing a computer device to execute any method described in the embodiments of the present invention, and a non-volatile recording medium storing the program are included in the scope of the embodiments of the present invention. The program for causing a computer device to execute any method described in the embodiments of the present invention may be a main program incorporated into an arbitrary main program or a subprogram called by an arbitrary main program. The learning device 10 is a kind of computer device. Any process in the embodiments of the present invention may be realized by hardware such as a semiconductor integrated circuit, software corresponding to the above program, or a combination of hardware and software.

[0150] The embodiments of the present invention can be appropriately modified in various ways within the scope of the technical idea shown in the claims. The above embodiments are merely examples of the embodiments of the present invention, and the meanings of the terms of the present invention or each constituent element are not limited to those described in the above embodiments. The specific numerical values shown in the above description are merely examples, and as a matter of course, they can be changed to various numerical values.

Explanation of Reference Numerals

[0151] 10 Learning device 11 Controller 12 Memory 13 Communication unit 110 Model 120 NN 20 Database 21 Learning data 21a Image dataset 22 Evaluation data Din Model input data Dout Model output data 121 Input layer 122 Intermediate layer 123 Output layer 122a, 122b Convolutional layer 122c Pooling layer 122d, 122e Fully connected layer F1 Model Evaluation Unit F2 Target Model Selection Unit F3 Main Learning Processing Unit F4 Pruning Processing Unit F5 Re - learning Processing Unit F6 Performance Evaluation Unit 410, 430 Estimator Generation Device 411, 431 Learning Model 412, 432 Performance Estimator

Claims

1. A learning method for generating an inference model that makes inferences based on input data, comprising: After evaluation learning in which a learning model composed of a neural network is learned based on learning data, evaluation unit processing for evaluating the performance of the learning model through first pruning on the learning model is executed for a plurality of learning models having mutually different structures; Based on the result of the evaluation in the evaluation unit processing, a target model is selected from among the plurality of learning models before the first pruning; After main learning in which the target model is learned using the learning data with a larger number of learning times than the evaluation learning, the inference model is generated by performing second pruning on the target model , a learning method.

2. In the evaluation unit processing, for each learning model, evaluation inference is performed on the learning model after the evaluation learning and the first pruning using evaluation data; Based on the result of the evaluation inference for each learning model, the performance of the learning model after the evaluation learning and the first pruning is compared among the plurality of learning models, and the target model is selected based on the comparison result , the learning method according to Claim 1.

3. Among the plurality of learning models, the learning model having the maximum performance after the evaluation learning and the first pruning is selected as the target model , the learning method according to Claim 2.

4. The performance is evaluated by the inference accuracy in the evaluation inference or the inference error in the evaluation inference, The learning model with the maximum inference accuracy or the learning model with the minimum inference error is selected as the target model from among the plurality of learning models , the learning method according to Claim 2.

5. In the evaluation unit processing, for each learning model, the first pruning is performed so that the amount of calculation by the learning model after the evaluation learning and the first pruning is equal to or less than a target amount of calculation, The second pruning is performed so that the amount of calculation by the target model after the main learning and the second pruning is equal to or less than the target amount of calculation , the learning method according to any one of Claims 1 to 4.

6. A model selection method used for generating an inference model that makes inferences based on input data, comprising: After evaluation learning in which a learning model composed of a neural network is learned based on learning data, evaluation unit processing for evaluating the performance of the learning model through pruning of the learning model is performed on a plurality of learning models having different structures from each other, Based on the results of the evaluation for each learning model, a target model that is the source of the inference model is selected from among the plurality of learning models before the pruning. A model selection method.

7. A learning device that generates an inference model for making inferences based on input data, After evaluation learning in which a learning model composed of a neural network is learned based on learning data, evaluation unit processing for evaluating the performance of the learning model through first pruning of the learning model is performed on a plurality of learning models having different structures from each other, Based on the results of the evaluation in the evaluation unit processing, a target model is selected from among the plurality of learning models before the first pruning, After main learning in which the target model is learned using the learning data for a greater number of learning times than the evaluation learning, the inference model is generated by performing second pruning on the target model. A learning device.

8. A program for causing a computer to execute a learning method for generating an inference model for making inferences based on input data, In the learning method, after evaluation learning in which a learning model composed of a neural network is learned based on learning data, evaluation unit processing for evaluating the performance of the learning model through first pruning of the learning model is performed on a plurality of learning models having different structures from each other, Based on the results of the evaluation in the evaluation unit processing, a target model is selected from among the plurality of learning models before the first pruning, After main learning in which the target model is learned using the learning data for a greater number of learning times than the evaluation learning, the inference model is generated by performing second pruning on the target model. A learning program.

Citation Information

Patent Citations

  • JP7111671A