Learning model evaluation apparatus, method, program and learning apparatus

The proposed learning model evaluation apparatus addresses the issue of inappropriate model quality assessment in Zero-Shot NAS by using an evaluation index based on weight parameter values, resulting in improved model serialization and accuracy.

JP2025096780APending Publication Date: 2025-06-30DENSO TEN LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023212685
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-30

AI Technical Summary

Technical Problem

Existing Zero-Shot NAS methods rely on evaluation indices that are highly correlated with the number of DNN parameters, leading to inappropriate determinations of model quality and potential overfitting due to excessive parameters.

Method used

A learning model evaluation apparatus and method that evaluates neural networks based on an evaluation index derived from the values of weight parameters in the convolutional layer, allowing for appropriate assessment of model structure quality without relying on parameter count.

Benefits of technology

Enables accurate evaluation of model structure quality, leading to better serialization of models and the generation of highly accurate learned models with improved generalization performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096780000001_ABST
    Figure 2025096780000001_ABST
Patent Text Reader

Abstract

To properly evaluate the quality of a learning model.SOLUTION: A learning model evaluation apparatus evaluates a learning model composed of a neural network having a convolutional layer. The learning model evaluation apparatus performs evaluation learning that learns the learning model using learning data, and derives an evaluation index based on values of a plurality of weight parameters in the convolutional layer of the learning model subjected to the evaluation learning.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning model evaluation apparatus, method and program, and a learning apparatus.

Background Art

[0002] Various arithmetic processes using a model having a DNN (Deep Neural Network) have been put into practical use. In the design of a DNN, it is necessary for a designer to determine design parameters (such as a kernel size and the number of output channels) for each layer (such as a convolutional layer and a pooling layer) of the DNN. In recent years, the DNN has been becoming more multi-layered, and there are many DNNs having more than 100 layers. In the DNN design process, the optimization of the DNN design is achieved by repeating a process including the DNN structure setting and the DNN performance evaluation a plurality of times. However, the number of candidates for design parameters in a multi-layer DNN is enormous, and it is difficult for only the designer himself / herself to optimize the DNN design.

[0003] As a method for automating the DNN design, a method called NAS (Neural Architecture Search) has been proposed. In NAS, a candidate (search space) for a DNN design pattern is determined in advance. Then, after sampling a model from the search space, a series of processes of evaluating the model through learning of the model (learning from several iterations to about several epochs) in a relatively short time is repeated until a desired target performance is obtained. At this time, as an efficient sampling method for obtaining the target performance with a small number of search times, Bayesian optimization using TPE, an evolutionary algorithm using NSGA, reinforcement learning, etc. are often used. TPE is an abbreviation of "Tree-Structured Parzen Estimator". NSGA is an abbreviation of "Non-dominated Sorting Genetic Algorithm". Note that Patent Document 1 below discloses a method for selecting search space information according to target constraint conditions of target hardware.

[0004] Conventional NAS needs to learn the model to some extent. Therefore, when the DNN or the dataset is large-scale, it takes a considerable amount of time to search until the target performance is obtained. In view of this, Zero-Shot NAS has been proposed as a NAS that does not require learning or can complete the DNN design with a small number of learning times when designing a DNN. According to Zero-Shot NAS, the search time can be significantly reduced.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] In Zero-Shot NAS, an evaluation index is required to determine the quality of the DNN structure. Existing evaluation indices are highly correlated with the number of DNN parameters. And in the evaluation method using existing evaluation indices (the evaluation method by existing Zero-Shot NAS), there is a tendency to determine that a DNN with a larger number of parameters is a better one. Although such a determination may be correct in some cases, in reality, depending on the dataset, learning task, search space, etc., this determination is often not appropriate. Also, although a larger number of parameters allows for more complex operations, if the number of parameters is unnecessarily large, there will be more wasted operations. An excessive number of parameters may also lead to overfitting.

[0007] Therefore, a proper evaluation index that does not depend on the number of DNN parameters is required. If the quality of the DNN structure (in other words, the quality of the model structure) can be evaluated using a proper evaluation index, the superiority and inferiority between models can be correctly evaluated. By performing learning on a more excellent model, it becomes possible to obtain a good trained model.

[0008] An object of the present invention is to provide a learning model evaluation apparatus, method, and program that contribute to appropriate evaluation of a learning model. Another object of the present invention is to provide a learning apparatus that contributes to generation of a good learned model.

Means for Solving the Problems

[0009] A learning model evaluation apparatus according to the present invention is a learning model evaluation apparatus that evaluates a learning model constituted by a neural network having a convolutional layer, performs evaluation learning for learning the learning model using learning data, and derives an evaluation index based on each value of a plurality of weight parameters in the convolutional layer of the learning model after the evaluation learning.

Effects of the Invention

[0010] By using an evaluation index based on each value of the weight parameter, it is possible to appropriately evaluate whether the structure of the learning model is good or bad. If the quality of the structure of the learning model can be appropriately evaluated, a plurality of learning models can be appropriately serialized (a high rank can be given to a truly excellent model). And for example, by performing learning with a higher number of learning times than the evaluation learning for the learning model to which the highest rank is assigned, a highly accurate learned model can be obtained.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Embodiments for Carrying Out the Invention

[0012] Hereinafter, examples of embodiments of the present invention will be specifically described with reference to the drawings. In each of the drawings referred to, the same parts are denoted by the same reference numerals, and redundant explanations regarding the same parts are omitted as a rule. In this specification, for the sake of simplification of description, the names of information, signals, physical quantities, functional units, circuits, elements, or components, etc. corresponding to the symbols or signs indicating them may be omitted or abbreviated by writing the symbols or signs.

[0013] Fig. 1 shows the overall configuration of a learning system (machine learning system) according to an embodiment of the present invention. The learning system in Fig. 1 includes a learning device 10 which is a machine learning device and a database 20. The learning device 10 is connected to a communication network including the Internet. The learning device 10 may be configured by one or more computer devices (server devices) connected to the communication network. The learning device 10 may be configured using cloud computing. The learning device 10 includes a controller 11, a memory 12, and a communication unit 13. Note that the learning described in this embodiment is machine learning.

[0014] The controller 11 comprehensively controls the operations of each part in the learning device 10. The controller 11 includes an arithmetic processing unit including a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit) etc. as hardware resources. The controller 11 may realize the following respective functions to be realized by the controller 11 by executing a program recorded in the memory 12, the database 20 or any other recording medium (not shown).

[0015] The memory 12 is configured to have a non-volatile memory such as a ROM (Read only memory) or a flash memory, and a volatile memory such as a RAM (Random access memory). In addition to storing each data referred to by the controller 11, the memory 12 stores various programs to be executed by the controller 11.

[0016] The communication unit 13 transmits and receives any signal to and from a counterpart device different from the learning device 10. The counterpart device for the communication unit 13 includes the database 20 and any computer device connected to the communication network. Note that the controller 11 can transmit and receive any information to and from the counterpart device (the counterpart device for the communication unit 13) using the communication unit 13, but the description of the communication unit 13 may be omitted hereinafter.

[0017] The database 20 is a large-capacity recording medium that holds the learning data 21. The learning data may be read as a learning data set. The database 20 is connected to the learning device 10 either wired or wirelessly. The controller 11 can freely read any data held in the database 20. The database 20 may be a collection of physically separated multiple recording media. In this case, the learning data 21 is held in the multiple recording media. Part or all of the database 20 may be provided in the learning device 10.

[0018] In the controller 11, the model 110 shown in FIG. 2 is constructed. Since machine learning is performed on the model 110, the model 110 can be referred to as a learning model. The model 110 is an algorithm that generates the model output data Dout from the model input data Din. The model 110 has an NN120 that is a neural network. The NN120 is configured using the hardware resources within the controller 11. The NN120 may be a deep neural network (DNN). The controller 11 executes the machine learning of the model 110 using the learning data 21. Incidentally, the machine learning of the model 110 can also be said to be the machine learning of the NN120. It is assumed that the machine learning performed by the controller 11 in this embodiment is supervised machine learning.

[0019] The model 110 that has undergone machine learning is a learned model (inference model) that performs a predetermined inference. The learned model may be, for example, an object detector that performs object detection as an inference. In object detection, the existence region of the image of the object and the type of the object in the two-dimensional image input to the model 110 are estimated (in other words, detected).

[0020] When object detection is performed as an inference, the learning data 21 is the image dataset 21a shown in FIG. 3. The image dataset 21a includes the image data and the correct data of a large number of learning images. In the image dataset 21a, the correct data is added to each learning image for the learning image. Each learning image in the image dataset 21a is a two-dimensional image including images of various types of objects. The correct data for a certain learning image includes information indicating the type of object existing in the learning image and information specifying the position and shape of the region where the image of the object exists in the learning image. The controller 11 can construct a model 110 that performs object detection as an inference by executing machine learning of the model 110 using the image dataset 21a.

[0021] The process in which machine learning is performed is referred to as the machine learning process. The machine learning process in the case of constructing the model 110 that performs object detection by machine learning will be described. For the sake of concretizing the description, the noted learning image is referred to as the noted image. In the machine learning process, the controller 11 inputs the image data of the noted image as the model input data Din into the model 110. Then, the model 110 estimates the existence region of the image of the object in the noted image and the type of the object based on the model input data Din, and generates model output data Dout indicating the estimation result. The correct data for the noted image indicates the correct answer of the content to be estimated by the model 110.

[0022] In the machine learning process, the controller 11 derives a loss function representing the error between the model output data Dout of the model 110 and the correct data for the noted image. Then, the controller 11 adjusts the parameters of the model 110 using the error backpropagation method so that the value of the loss function is reduced. The parameters of the model 110 are the parameters of the NN120 and include the weight and bias parameters in the NN120. The above adjustment is repeated until a predetermined learning end condition is satisfied, and when the learning end condition is satisfied, the machine learning process is terminated. The model 110 after the machine learning process functions as an object detector (inference model) that performs object detection as an inference.

[0023] Object detection is a type of image recognition. The inference performed by model 110 can be any image recognition (such as image classification or semantic segmentation). When the inference performed by model 110 is image recognition, a convolutional neural network can be used as NN120. The task (inference) performed by model 110 is arbitrary and can be a clustering task, a generation task, a regression task, a reinforcement task, etc. In the following, for the sake of concretizing the explanation, unless otherwise specified, it is assumed that the training data 21 is the image dataset 21a, and through machine learning, model 110 is made to function as an object detector.

[0024] Fig. 4 shows the basic structure of NN120. NN120 is a forward propagation neural network and has an input layer 121, an intermediate layer 122, and an output layer 123. The input layer 121 receives the model input data Din and outputs the model input data Din to the intermediate layer 122. The intermediate layer 122 is provided between the input layer 121 and the output layer 123. The intermediate layer 122 performs operations on the model input data Din according to the structure and parameters of the intermediate layer 122, and outputs the operation result to the output layer 123. The output layer 123 generates and outputs the model output data Dout based on the operation result output from the intermediate layer 122.

[0025] A plurality of layers are provided in the intermediate layer 122. A target layer, which is an arbitrary layer in the intermediate layer 122, receives data from the layer arranged in front of itself (specifically, the layer arranged immediately in front of itself), and performs operations on the received data according to its own structure and parameters. Then, the target layer outputs the operation result to the layer arranged behind itself (specifically, the layer arranged immediately behind itself). When the target layer is the first layer provided in the intermediate layer 122, the model input data Din from the input layer 121 is received by the target layer. When the target layer is the last layer provided in the intermediate layer 122, the operation result by the target layer is output to the output layer 123.

[0026] In the intermediate layer 122, the convolutional layer L CNVand a batch normalization layer L BN One or more blocks BLK having CNV are provided. In each block BLK, a convolutional layer L BN is provided at the subsequent stage of CNV A plurality of convolutional layers L BN may be provided within one block BLK. Also, a plurality of batch normalization layers L CNV and batch normalization layer L BN are provided. However, here, it is assumed that one convolutional layer L

[0027] convolutional layer L CNV performs a convolution operation on the data supplied to itself and outputs the result of the convolution operation to the layer arranged at the subsequent stage of itself (specifically, the layer arranged at the immediate subsequent stage of itself). The batch normalization layer L BN is a layer that performs known batch normalization. As is well known, batch normalization enables faster and more stable learning. In each block BLK, the data obtained by performing batch normalization on the output data of the convolutional layer L CNV is output from the batch normalization layer L BN

[0028] The intermediate layer 122 may also be provided with layers (such as pooling layers and fully connected layers) that are not classified into the convolutional layer L CNV and the batch normalization layer L BN . However, hereinafter, unless particularly necessary, attention is paid only to the convolutional layer L CNV and the batch normalization layer L BN among the layers provided in the intermediate layer 122.

[0029] ​The controller 11 is configured to be able to change the structure of the NN 120 in various ways. If the structure of the NN 120 changes, the characteristics of the model 110 change. Let the number of types of structures of the NN 120 set by the controller 11 be represented by "n". The number of types of structures that the controller 11 can set is much larger than n, but here we focus on n types of structures. When the NN 120 has the i-th structure, the model 110 is referred to as the i-th model. n represents any integer of 2 or more, and i represents a natural number of n or less. The controller 11 can make the model 110 function as the first to n-th models by changing the structure of the NN 120 in n steps. The structure of the NN 120 is also the structure of the model 110. That is, the i-th structure is the structure of the NN 120 that forms the i-th model and can also be said to be the structure of the i-th model.

[0030] The first to n-th structures are different from each other. That is, for all combinations of natural numbers e and f, the design parameters set for the NN 120 of the e-th structure and the design parameters set for the NN 120 of the f-th structure are different from each other. e and f represent any different natural numbers of n or less.

[0031] The design parameters are parameters that define the structure of the NN 120. Therefore, the design parameters may be referred to as structure parameters. For example, the number of filters, the number of channels, and the kernel size in the convolutional layer L CNV correspond to the design parameters. The number of filters, the number of channels, and the kernel size in the pooling layer also correspond to the design parameters. The number of layers constituting the NN 120 of the e-th structure and the number of layers constituting the NN 120 of the f-th structure may also be different from each other. Therefore, for example, the number of convolutional layers L CNV included in the NN 120 of the e-th structure and the number of convolutional layers L CNV included in the NN 120 of the f-th structure may be different from each other. Similarly, for example, the number of batch normalization layers L BN included in the NN 120 of the e-th structure and the number of batch normalization layers L BN included in the NN 120 of the f-th structure may be different from each other. However, the NN 120 has at least the basic structure shown in FIG. 4.

[0032] The learning device 10 selects a target model from among the first to nth models, and performs machine learning (main learning described later) used for the learning data 21 on the target model. The target model is typically a model expected to have the maximum inference accuracy. The target model corresponds to the optimal model extracted from among the first to nth models. The method for selecting the target model will become clear from the following explanation.

[0033] Fig. 5 shows the operation flowchart of the learning device 10. After the learning device 10 is activated, when the evaluation program starts execution in the controller 11, the model evaluation process in step S10 is executed by the controller 11. In the model evaluation process, the first to nth models are each learned. The learning performed in the model evaluation process is particularly referred to as evaluation learning. In the model evaluation process, an evaluation index is derived based on the results of the evaluation learning for each model.

[0034] Thereafter, in step S30, the serialization process is executed by the controller 11. When the serialization program starts execution in the controller 11, the serialization process is realized. The serialization program may be a part of the evaluation program. In the serialization process, the first to nth models are serialized based on the evaluation index derived in step S10, and thereby the target model is selected from among the first to nth models. For this reason, the serialization process and the serialization program may also be referred to as the target model selection process and the target model selection program. After the serialization process, the process proceeds to step S50.

[0035] In step S50, the main learning process is executed by the controller 11. When the main learning program starts execution in the controller 11, the main learning process is realized. In the main learning process, learning of the target model is executed. The learning performed in the main learning process is particularly referred to as main learning.

[0036] The learning for evaluation is a short-term learning carried out for the purpose of selecting a target model (optimal model) from among the first to nth models. The short-term here refers to being short in comparison to the execution time of this learning. That is, after evaluating the quality of the structure of each model using the model evaluation process, long-term learning that requires a long time is performed on the selected target model. The number of learning times in the long-term learning of the target model is more than the number of learning times in the learning for evaluation of each of the first to nth models. The number of learning times is represented by, for example, the number of epochs. The number of epochs (for example, several tens) in the long-term learning of the target model is larger than the number of epochs (for example, 1 or 2) in the learning for evaluation of the ith model. While performing the long-term learning of the target model using all of the learning data 21, the learning for evaluation of the ith model may be performed using only a part of the learning data 21.

[0037] An integrated program that includes all of the evaluation program, serialization program, and this learning program may be configured. The evaluation program, serialization program, and this learning program may be programs for subroutines called from the main program. The evaluation program, serialization program, this learning program, and the main program may be programs recorded in the memory 12, database 20, or any other arbitrary recording medium (not shown). The same applies to the integrated program.

[0038] [Model Evaluation Process] Referring to FIG. 6, the flow of the model evaluation process executed in step S10 of FIG. 5 will be described. FIG. 6 is a flowchart of the model evaluation process. The model evaluation process consists of the processes of steps S11 to S20. In the model evaluation process, the controller 11 first assigns 1 to the variable i that it manages in step S11. After step S11, it proceeds to step S12.

[0039] In step S12, the controller 11 performs sampling of the model using NAS. NAS is an abbreviation for "Neural Architecture Search".

[0040] The NAS itself is well-known. A method of sampling a model by NAS will be briefly described. Prior to the execution of the model evaluation process, a search space is defined by an operator of the learning device 10, and search space information indicating the search space is given to the learning device 10. When each of a plurality of structures that can be adopted as the structure of the NN120 is referred to as a candidate structure, the search space is a set of a plurality of candidate structures. The search space defines from which set of structures the optimal structure of the NN120 is to be searched.

[0041] In NAS, a method for searching for the optimal structure of the NN120 from the search space is called a search strategy. In step S12, the controller 11 samples a model from the search space according to a determined search strategy. As the search strategy, Bayesian optimization using TPE, an evolutionary algorithm using NSGA, or reinforcement learning can be used. TPE is an abbreviation for "Tree-Structured Parzen Estimator". NSGA is an abbreviation for "Non-dominated Sorting Genetic Algorithm".

[0042] The model sampled in step S12 is the i-th model, and the i-th model is set in the model 110. That is, in step S12, a model 110 having the NN120 of the i-th structure is constructed. The loop process from step S12 to step S19 is repeatedly executed. The i-th model in the first loop process is the model 110 as the first model, and the i-th model in the second loop process is the model 110 as the second model. The same applies to the loop processes after the third time. After step S12, the process proceeds to step S13.

[0043] In step S13, the controller 11 initializes the i-th model. After step S13, the process proceeds to step S14. In the initialization, a predetermined initial value or a value randomly selected from a predetermined range is set as the learning parameter of the i-th model. Any model has learning parameters. The learning parameters refer to the parameters that are adjusted during learning. The weight parameters in the convolutional layer L CNV and the scale parameter and the shift parameter in the batch normalization layer L BN belong to the learning parameters.

[0044] In the model evaluation process, each of the first to n-th models is learned based on the learning data 21. The learning performed in the model evaluation process is evaluation learning. The evaluation learning is executed in the processes of steps S14 to S17 (evaluation learning process). In steps S14 to S17, the above-described machine learning process is performed on the i-th model.

[0045] In step S14, the controller 11 samples a mini-batch from the learning data 21. Mini-batch learning is used in the machine learning in the controller 11. The mini-batch is data having a predetermined mini-batch size within the learning data 21. For example, when the learning data 21 includes image data of 100,000 learning images, the mini-batch includes the image data of 10 learning images therein and also includes the correct data for the 10 learning images. In the model evaluation process, the mini-batch sampled in step S14 is referred to as unit training data. The acquired unit training data is temporarily stored in the memory 12 and used for the subsequent processes. After step S14, the process proceeds to step S15.

[0046] In step S15, the controller 11 derives the loss function of the i-th model based on the unit training data. More specifically, in step S15, the controller 11 inputs the image data of the learning image included in the unit training data as the model input data Din into the i-th model (that is, inputs it into the NN120 in the i-th model). Thereby, the i-th model generates model output data Dout based on the current learning parameters. The correct data included in the unit training data corresponds to the correct data for the model output data Dout. The controller 11 derives a loss function representing the error between the model output data Dout based on the unit training data and the correct data included in the unit training data. After step S15, the process proceeds to step S16.

[0047] In step S16, the controller 11 derives the gradient of the loss function of the i-th model by the error backpropagation method, and updates the learning parameters of the i-th model based on the derived gradient. After step S16, the process proceeds to step S17.

[0048] In step S17, the controller 11 determines whether a predetermined first end condition is satisfied. The first end condition corresponds to the learning end condition in the learning for evaluation. For example, when the number of executions of the mini-batch learning executed for the i-th model (the number of executions of the series of processes consisting of steps S14 to S16) reaches a predetermined number, the first end condition is satisfied. Or for example, when a certain period of time has elapsed after the start of the learning for evaluation for the i-th model, the first end condition is satisfied.

[0049] If the first end condition is not satisfied in step S17 (No in step S17), the process returns to step S14 and the series of processes starting from step S14 is executed again. If the first end condition is satisfied in step S17 (Yes in step S17), the process proceeds to step S18.

[0050] In step S18, the controller 11 executes an evaluation index derivation process for deriving the evaluation index EV. The evaluation index EV derived for the i-th model is particularly referred to as the evaluation index EV[i]. The evaluation index EV[i] is an index indicating the goodness or badness (degree of goodness) of the structure of the i-th model. In the evaluation index derivation process, the evaluation index EV[i] is derived based on the learning parameters of the i-th model at the current time (that is, based on the learning parameters of the i-th model after learning for evaluation). The method for deriving the evaluation index EV[i] will be described in detail later. After step S18, the process proceeds to step S19.

[0051] In step S19, the controller 11 determines whether a predetermined second end condition is satisfied. For example, when the variable i reaches a predetermined value, the second end condition is satisfied. Or, for example, when a predetermined evaluation upper limit time has elapsed after the start of the model evaluation process, the second end condition is satisfied.

[0052] If the second end condition is not satisfied in step S19 (No in step S19), in step S20, 1 is added to the variable i, then the process returns to step S12, and the processes after step S12 are executed again. If the second end condition is satisfied in step S19 (Yes in step S19), the model evaluation process in FIG. 6 is terminated. The value of the variable i at the time when the second end condition is satisfied corresponds to the value of n.

[0053] In the step S12 after the second time, sampling of the model may be performed based on the evaluation index EV derived in the step S18. Actually, for example, in the step S12 from the first time to the tenth time, the first to tenth models having the first to tenth structures are sampled by randomly determining the structure of the model. That is, for example, in the step S12 when "1 ≦ i ≦ 10" is satisfied, the i-th model having the i-th structure is sampled by randomly determining the structure of the model. Thereafter, in the step S12 when "11 ≦ i" is satisfied, based on the structures of the first to (i - 1)-th models and the derived evaluation indexes EV[1] to EV[i - 1], the structure of the i-th model (that is, the i-th structure) is set according to the search strategy. At this time, the structure in which an improvement in the evaluation index EV is expected in comparison with the first to (i - 1)-th models is set as the i-th structure for the i-th model.

[0054] [Serialization process] Referring to FIG. 7, the flow of the serialization process executed in step S30 of FIG. 5 will be described. FIG. 7 is a flowchart of the serialization process. The serialization process consists of the processes of steps S31 and S32. In the model evaluation process executed before the serialization process, the evaluation indexes EV[1] to EV[n] for the first to n-th models have been derived. In the serialization process, the controller 11 first performs serialization of the first to n-th models based on the evaluation indexes EV[1] to EV[n] in step S31. In the serialization, the controller 11 assigns a rank to each of the first to n-th models. Among the first to n-th models, the same rank may be assigned to any two or more models, but basically, different ranks are assigned to the first to n-th models. Hereinafter, for the sake of convenience of explanation, it is assumed that different ranks are assigned to the first to n-th models.

[0055] In the serialization of the first to n-th models, the first to n-th models are respectively ranked, and one of the first to n-th ranks is assigned to each of the first to n-th models. The j-th rank has a higher rank than the (j + 1)-th rank (j is a natural number). Therefore, the model to which the first rank is assigned has the highest rank, and the model to which the n-th rank is assigned has the lowest rank.

[0056] In step S32 following step S31, the controller 11 selects a target model from among the first to nth models based on the serialization result. Among the first to nth models, the model assigned the first rank is selected as the target model. Therefore, it can be said that the serialization process is a process for selecting a target model. Note that in the serialization process, it is not essential to assign ranks to all models. In the serialization process, it is sufficient if at least the model assigned the first rank is specified.

[0057] [This learning process] Referring to FIG. 8, the flow of this learning process executed in step S50 of FIG. 5 will be described. FIG. 8 is a flowchart of this learning process. This learning process consists of the processes of steps S51 to S55.

[0058] This learning process starts with the process of step S51. In step S51, the controller 11 sets the target model selected in the serialization process to the model 110. That is, when the target model is the first model, the model 110 having the NN120 of the first structure is constructed as the target model. Similarly, when the target model is the second model, the model 110 having the NN120 of the second structure is constructed as the target model. The same applies when the target model is the third model or the like.

[0059] The target model after evaluation learning may be set to model 110 in step S51. That is, for example, if the first model is the target model, the learning parameters of the first model after evaluation learning may be set to model 110 in step S51. Similarly, for example, if the second model is the target model, the learning parameters of the second model after evaluation learning may be set to model 110 in step S51. The same applies when the target model is the third model or the like. Thereby, starting from a state where learning has progressed to a certain extent, the main learning of the target model is executed. However, the target model may be initialized in step S51. In this case, in the initialization of step S51, a predetermined initial value or a value randomly selected from a predetermined range is set as the learning parameters of the target model. After step S51, the process proceeds to step S52.

[0060] In this main learning process, the target model is learned based on the learning data 21. The learning performed in this main learning process is the main learning. The main learning is realized in the processes of steps S52 to S55 (main learning process). In steps S52 to S55, the above-described machine learning process is performed on the target model.

[0061] In step S52, the controller 11 samples a mini-batch from the learning data 21. As described above, mini-batch learning is used in the machine learning in the controller 11, and the mini-batch is data having a predetermined mini-batch size within the learning data 21. The unit training data in this main learning process is the mini-batch sampled in step S52. The acquired unit training data is temporarily stored in the memory 12 and used for subsequent processing. After step S52, the process proceeds to step S53.

[0062] In step S53, the controller 11 derives the loss function of the target model based on the unit training data. More specifically, in step S53, the controller 11 inputs the image data of the learning image included in the unit training data as the model input data Din into the target model (that is, inputs it into the NN120 in the target model). Thereby, the target model generates the model output data Dout based on the current learning parameters. The correct data included in the unit training data corresponds to the correct data for the model output data Dout. The controller 11 derives the loss function representing the error between the model output data Dout based on the unit training data and the correct data included in the unit training data. After step S53, the process proceeds to step S54.

[0063] In step S54, the controller 11 derives the gradient of the loss function of the target model by the error backpropagation method, and updates the learning parameters of the target model based on the derived gradient. After step S54, the process proceeds to step S55.

[0064] In step S55, the controller 11 determines whether a predetermined learning end condition is satisfied. When the learning in this learning process converges, the learning end condition is satisfied. For example, when the value of the loss function of the target model becomes equal to or less than a predetermined small value, it is determined that the learning has converged. Alternatively, for example, when the number of executions of the mini-batch learning executed on the target model (the number of executions of a series of processes consisting of steps S52 to S54) reaches a predetermined number, the learning end condition may be satisfied. Further alternatively, for example, when a certain time has elapsed after the start of learning for the target model, the learning end condition may be satisfied.

[0065] If the learning end condition is not satisfied in step S55 (No in step S55), the process returns to step S52 and the series of processes starting from step S52 are executed again. If the learning end condition is satisfied in step S55 (Yes in step S55), the main learning process in FIG. 8 is ended. The target model having the learning parameters after this main learning (i.e., the target model after the satisfaction of the learning end condition) is a learned model that should also be referred to as an inference model, and functions as an inference model capable of performing good inference.

[0066] [Convolutional layer] Next, the input and output data of the convolutional layer L CNV will be described. Refer to FIG. 9. For any noted convolutional layer L CNV , the input data to the convolutional layer L CNV is referred to as the input map IN CNV , and the output data from the convolutional layer L CNV is referred to as the output map OUT CNV . The output map OUT CNV represents the result of the convolution operation by the convolutional layer L CNV . In learning or inference, the convolutional layer L CNV extracts the feature amount of the input image to the model 110. The output map OUT CNV is a map (feature map) representing the feature amount extracted by the convolutional layer L CNV . When the noted convolutional layer L CNV is the first layer provided in the intermediate layer 122, the input map IN CNV is the model input data Din. When the noted convolutional layer L CNV is a layer other than the first layer provided in the intermediate layer 122, the input map IN CNV is a feature map generated within the intermediate layer 122.

[0067] The input map IN CNV has a two-dimensional map (two-dimensional image) having W pixels in the width direction and H pixels in the height direction and includes one or more channels. The input map IN CNVWhen having two or more two-dimensional maps, the two or more two-dimensional maps are arranged in the channel direction. Input map IN CNV The number of channels in is represented by CHa. Then, the input map IN CNV can be considered as a three-dimensional map (three-dimensional image) consisting of (W×H×CHa) pixels. W and H have integer values of 2 or more. CHa has an integer value of 1 or more. Hereinafter, it is assumed that “CHa≧2”. CHa represents the number of channels of the input map IN CNV and also represents the number of input channels of the convolutional layer L CNV . The input map IN CNV has a pixel value for each pixel.

[0068] Here, for the sake of concretizing the description, the width direction is regarded as the X-axis direction, the height direction is regarded as the Y-axis direction, and the channel direction is regarded as the Z-axis direction. The X-axis, Y-axis, and Z-axis are orthogonal to each other. The two-dimensional map of the j-th channel in the input map IN CNV is represented by IN[j]. Then, the input map IN CNV consists of two-dimensional maps IN[1] to IN[CHa]. For example, when the input map IN CNV is the image data of a two-dimensional image consisting of an R component, a G component, and a B component, the R component, G component, and B component of the two-dimensional image are set to the two-dimensional maps IN[1], IN[2], and IN[3], respectively.

[0069] Convolutional layer L CNV is provided with K filters FLT. The filter FLT is a convolutional filter. K has an integer value of 1 or more. Among the K filters FLT, the j-th filter FLT is referred to as filter FLT[j] (j is an integer). Therefore, the convolutional layer L CNV is provided with filters FLT[1] to FLT[K].

[0070] Each filter FLT has kernels for the number of channels CHa. Each kernel is a two-dimensional spatial filter having a size of K W pixels in the width direction and a size of K H pixels in the height direction. Each kernel is K in the width direction.W has pixels and K in the height direction H can also be considered to have pixels. K W and K H has an integer value of 1 or more, and K W and K H at least one of them has a value of 2 or more. In each filter FLT, kernels for the number of channels CHa are arranged in the channel direction. Therefore, the size of each filter FLT is (K W ×K H ×CHa). That is, each filter FLT is a three-dimensional spatial filter having a size of (K W ×K H ×CHa). In each filter FLT, a coefficient is set for each pixel. The coefficient in each filter FLT is a weight parameter. For this reason, each filter FLT has (K W ×K H ×CHa) weight parameters. Therefore, the total number of weight parameters set for one convolutional layer L CNV is (K W ×K H ×CHa×K). The weight parameters belong to the learning parameters and are updated during learning.

[0071] Fig. 10 shows a configuration example of one filter FLT. In Fig. 10, it is assumed that (K W , K H , CHa) = (3, 3, 5). In Fig. 10, w1 to w 45 represent 45 weight parameters set for one filter FLT.

[0072] The output map OUT CNV is a two-dimensional map (two-dimensional image) having W pixels in the width direction and H pixels in the height direction and having CHb channels. CHb represents the number of channels of the output map OUT CNV . When "CHb ≧ 2" in the output map OUT CNV , two or more two-dimensional maps are arranged in the channel direction. Then, the output map OUT CNVIt can be considered as a three-dimensional map (three-dimensional image) composed of (W×H×CHb) pixels. Output map OUT CNV has a pixel value for each pixel. Output map OUT CNV The two-dimensional map of the j-th channel in is represented by OUT[j]. K representing the number of filters FLT is the output map OUT CNV is the same as the number of channels CHb of (i.e., the output channels of the convolutional layer L CNV ). Therefore, the output map OUT CNV consists of two-dimensional maps OUT[1] to OUT[K], and the two-dimensional map OUT[K] and the two-dimensional map OUT[CHb] refer to the same thing.

[0073] Controller 10 performs a convolution operation in the convolutional layer L CNV . In the convolution operation, a target position is set on the XY-plane parallel to the X-axis and Y-axis. Input map IN CNV The convolution operation on is performed for each filter FLT, and thus consists of the first to K product-sum operations. The two-dimensional map OUT[j] is derived by the j-th product-sum operation.

[0074] In the first product-sum operation, with the filter FLT[1] placed at the target position in the input map IN CNV , the product of the values of the pixels in the overlapping input map IN CNV and the coefficients (weight parameters) of the pixels in the filter FLT[1] is obtained. Then, a setting operation is performed in which the first bias is added to the sum of the obtained products, and the resulting value is set as the pixel value at the target position in the two-dimensional map OUT[1]. In the first product-sum operation, the setting operation is sequentially executed while shifting the target position on the XY-plane. Thereby, the pixel values of all pixels in the two-dimensional map OUT[1] are obtained.

[0075] If, as shown in FIG. 10, (K W , K H , CHa) = (3, 3, 5), then K W ×K H×CHa = 3×3×5 = 45. Therefore, in the first product-sum operation, 45 products are obtained for one target position, and the value obtained by adding the first bias to the sum of the 45 products is set as the pixel value of the target position in the two-dimensional map OUT[1].

[0076] The second to K-th product-sum operations are the same as the first product-sum operation. That is, in the j-th product-sum operation, with the filter FLT[j] placed at the target position in the input map IN CNV while overlapping each other, the product of the value of the pixel in the input map IN CNV and the coefficient (weight parameter) of the pixel in the filter FLT[j] is obtained. Then, in the j-th product-sum operation, a setting operation is performed in which the value obtained by adding the j-th bias to the sum of the obtained products is set as the pixel value of the target position in the two-dimensional map OUT[j]. In the j-th product-sum operation, this setting operation is sequentially executed while shifting the target position on the XY plane. Thereby, the pixel values of all the pixels in the two-dimensional map OUT[j] are obtained. The first to K-th biases belong to the learning parameters. Here, it is assumed that the stride and padding of the convolutional layer L CNV are set so that the sizes in the width direction and height direction in the input map IN CNV are the same as the sizes in the width direction and height direction in the output map OUT CNV .

[0077] The structure of the convolutional layer L CNV is defined by the design parameters. In the convolutional layer L CNV , K representing the number of filters FLT, CHa representing the number of input channels, the width direction size K W of each kernel, and the height direction size K H of each kernel belong to the design parameters. The design parameters, unlike the learning parameters, are not changed during learning. The design parameters are defined by the structure of the NN120 including the convolutional layer L CNV , and the design parameters are different between the e-th model and the f-th model (as described above, e and f represent different natural numbers less than or equal to n).

[0078] [Batch Normalization Layer] Next, the input and output data of the batch normalization layer L BN will be described. Refer to Fig. 11. For any noted batch normalization layer L BN , the input data to the batch normalization layer L BN is referred to as the input map IN BN , and the output data from the batch normalization layer L BN is referred to as the output map OUT BN . The output map OUT BN is data obtained by performing batch normalization on the input map IN BN . Here, it is assumed that the input map IN BN is an input map for one mini-batch based on the learning data 21 for one mini-batch (i.e., one unit training data). In a certain block BLK, the output map OUT CNV of the convolutional layer L CNV becomes the input map IN BN to the batch normalization layer L BN . In the following, it is considered that the output map OUT CNV shown in Fig. 9 is input as the input map IN BN to the batch normalization layer L BN .

[0079] Then, the input map IN BN has a two-dimensional map (two-dimensional image) with W pixels in the width direction and H pixels in the height direction, and has CHb channels. CHb represents the number of channels of the input map IN BN (the number of input channels of the batch normalization layer L BN ). When "CHb≧2" in the input map IN BN , two or more two-dimensional maps are arranged in the channel direction. Then, it can be considered that the input map IN BN is a three-dimensional map (three-dimensional image) consisting of (W×H×CHb) pixels. The input map IN BN has a pixel value for each pixel. The two-dimensional map of the j-th channel in the input map IN BN is represented by U[j]. Then, the input map IN BNIt consists of two-dimensional maps U[1] to U[CHb].

[0080] Output map OUT BN is the input map IN BN Similar to IN, it has a two-dimensional map (two-dimensional image) with W pixels in the width direction and H pixels in the height direction, and is provided with CHb channels. CHb represents the number of channels of the output map OUT BN (the number of output channels of the batch normalization layer L BN ). When "CHb ≧ 2" in the output map OUT BN , two or more two-dimensional maps are arranged in the channel direction. Then, the output map OUT BN can be considered as a three-dimensional map (three-dimensional image) consisting of (W × H × CHb) pixels. The output map OUT BN has a pixel value for each pixel. The two-dimensional map of the j-th channel in the output map OUT BN is represented by V[j]. Then, the output map OUT BN consists of two-dimensional maps V[1] to V[CHb].

[0081] In the two-dimensional map U[j], the pixel at an arbitrary attention position is called the attention pixel. And, as shown in FIG. 12, the pixel value at the attention pixel on the two-dimensional map U[j] is represented by u[j, k]. Further, on the two-dimensional map V[j], the pixel value of the pixel arranged at the position corresponding to the attention position is represented by v[j, k].

[0082] Batch normalization layer L BN In it, the pixel value v[j, k] is derived according to the following formulas (1) and (2). In the two-dimensional maps U[j] and V[j], the pixel values u[j, k] and v[j, k], and the formulas (1) and (2), j has an integer value satisfying "1 ≦ j ≦ CHb", and k has an integer value less than or equal to the number of pixels of the two-dimensional map U[j]. The number of pixels of the two-dimensional map U[j] is equal to the number of pixels of the two-dimensional map V[j].

[0083]

Equation

[0084] The calculation of the pixel value v[j, k] according to formulas (1) and (2) is executed for each channel. In formula (1), μ j represents the average of the pixel values of all pixels in the two-dimensional map U[j]. In formula (1), σ j 2 represents the variance of the pixel values of all pixels in the two-dimensional map U[j]. ε is a predetermined small amount to avoid the denominator on the right side of formula (1) becoming zero. The pixel value u[j, k] is normalized within the j-th channel by formula (1). Norm_u[j, k] shown on the left side of formula (1) corresponds to the pixel value u[j, k] after normalization.

[0085] Batch normalization layer L BN sets scale parameters and shift parameters for each channel. In formula (2), γ j represents the scale parameter set for the j-th channel, and β j represents the shift parameter set for the j-th channel. By scaling and shifting the pixel value Norm_u[j, k] according to formula (2), the pixel value v[j, k] is derived. The scale parameter and the shift parameter belong to the learning parameters and are updated during learning.

[0086] [Evaluation Index] As described above, in the evaluation method using the existing evaluation index (the evaluation method by the existing Zero-Shot NAS), there is a tendency to determine that a DNN with a larger number of parameters is better, but such a determination is not always appropriate. In the present embodiment, when evaluating the quality of the structure of NN120, an appropriate evaluation index is introduced.

[0087] Hereinafter, in a plurality of embodiments, some specific examples, application technologies, modified technologies, etc. related to the derivation of evaluation indicators will be described. The matters described above in this embodiment are applied to the following respective embodiments unless otherwise specified and without contradiction. In each embodiment, if there are matters conflicting with the above-described matters, the description in each embodiment may be given priority. Also, without contradiction, among the plurality of embodiments shown below, the matters described in any one of the embodiments can be applied to any other embodiment (that is, it is also possible to combine any two or more of the plurality of embodiments).

[0088] <<First Embodiment>> The first embodiment will be described. The controller 11 according to the first embodiment performs singular value decomposition on the weight parameters of the i-th model after evaluation learning, and performs an operation to obtain an evaluation index EV[i] based on the average value of a plurality of singular values obtained by the singular value decomposition. In the evaluation index derivation process in step S18 of FIG. 6, by executing this operation for each of the first to n-th models, evaluation indexes EV for each of the first to n-th models are derived (that is, evaluation indexes EV[1] to EV[n] are derived).

[0089] The weight parameters of the i-th model after evaluation learning refer to the weight parameters of the i-th model at the stage when the evaluation index derivation process is executed after evaluation learning for the i-th model. Specifically, the weight parameters of the i-th model refer to the weight parameters of the convolutional layer L provided in the NN120 of the i-th model CNV . In the first embodiment, the following weight parameters refer to the weight parameters after evaluation learning unless otherwise specified.

[0090] One convolutional layer L that has been noted CNV is referred to as the noted convolutional layer L CNV , and the singular value decomposition related to the noted convolutional layer L CNV will be described. It is assumed that the noted convolutional layer L CNV has the structure shown in FIG. 9. Regarding matters not specifically described here, the known singular value decomposition theorem is followed.

[0091] The noted convolutional layer LCNV Method MTD_A1, which is the first method of singular value decomposition for CNV , will be described. In method MTD_A1, singular value decomposition of the weight parameter is performed for each kernel. The convolutional layer L of interest CNV is provided with K filters FLT, and each filter FLT consists of CHa kernels. Therefore, the convolutional layer L of interest CNV is provided with (K × CHa) kernels. In method MTD_A1, singular value decomposition of the weight parameter is performed individually for each of the (K × CHa) kernels.

[0092] One kernel has (K W × K H ) weight parameters, and in method MTD_A1, the (K W × K H ) weight parameters are regarded as a matrix A of K W rows and K H columns (see Fig. 9). At this time, the matrix A can be decomposed into three matrices U, D, and V T as shown in the following formula (3) and Fig. 13. Here, the matrix D is a diagonal matrix having positive numbers α1 to α r as diagonal components (r represents the rank of the matrix A). The matrix U is a matrix of K W rows and r columns, the matrix V is a matrix of K H rows and r columns, and the matrix V T is the transposed matrix of the matrix V. The positive numbers α1 to α r are the singular values of the matrix A. A = U·D·V T ···(3)

[0093] In method MTD_A1, for each kernel, the singular values α1 to α r of the matrix A related to the kernel are obtained. Therefore, for example, as shown in Fig. 10, assuming (K W , K H , CHa) = (3, 3, 5), three singular values are derived for each kernel, so a total of 15 (= 3 × CHa) singular values are derived for one filter FLT. Then, since a total of K filters FLT are provided in the convolutional layer L of interest CNV (see Fig. 9), in the convolutional layer L of interest CNVA total of (15 × K) singular values will be derived for it.

[0094] Attention convolutional layer L CNV Method MTD_A2, which is the second method of singular value decomposition for, will be described. In method MTD_A2, singular value decomposition of the weight parameters is performed for each filter FLT. Therefore, in method MTD_A2, singular value decomposition of the weight parameters is performed individually for each of the K filters FLT.

[0095] One filter FLT has (K W × K H × K) weight parameters, and in method MTD_A2, the (K W × K H × K) weight parameters are regarded as a matrix A with K W rows and (K H × K) columns, or a matrix A with (K W × K) rows and K H columns (see Figure 9). Then, by performing singular value decomposition on the matrix A related to method MTD_A2, the singular values of the matrix A are obtained. Then, in method MTD_A2, compared with method MTD_A1, the number of singular values of the matrix A becomes K times. However, the number of matrices A defined for the attention convolutional layer L CNV is (1 / K) times that of method MTD_A1 in method MTD_A2. Therefore, the total number of singular values derived for the attention convolutional layer L CNV is the same between methods MTD_A1 and MTD_A2.

[0096] The number of rows and columns of the matrix A having the weight parameters in the filter FLT as components may be different from those shown in methods MTD_A1 and MTD_A2. Anyway, since r singular values are defined per kernel, a total of (r × CHa × K) singular values are derived for the attention convolutional layer L CNV by singular value decomposition.

[0097] Case CS_A1 where only one convolutional layer L CNV is provided in the NN120 of the i-th model, and convolutional layer L CNVCase CS_A2 with m provided, and the method for deriving the evaluation index EV[i] will be explained separately. For the sake of convenience in explanation, in the first embodiment, the symbol “EV A ” is introduced as the symbol related to the evaluation index, and the evaluation index of the i-th model is specifically represented by “EV A [i]”. m represents an integer of 2 or more.

[0098] ---Case CS_A1 (Convolution layer L CNV is 1)--- The controller 11 related to Case CS_A1 obtains the following average value AVE[i]_A1 as the evaluation index EV A [i]. The average value AVE[i]_A1 is the average value of the total (r×CHa×K) singular values derived for the convolution layer L CNV in the i-th model.

[0099] As a variation, the controller 11 related to Case CS_A1 may obtain the evaluation index EV A [i] by substituting the average value AVE[i]_A1 into a predetermined arithmetic expression. For example, the evaluation index EV A [i] may be the sum of the average value AVE[i]_A1 and a predetermined value. In any case, it is assumed that the evaluation index EV A [i] increases as the average value AVE[i]_A1 increases.

[0100] ---Case CS_A2 (Convolution layer L CNV is m)--- In Case CS_A2, the m convolution layers L CNV provided in the NN120 of the i-th model are referred to as the first to m-th convolution layers L CNV . The controller 11 related to Case CS_A2 obtains the average value of the singular values for each convolution layer L CNV , and then takes the average of the total m average values obtained for the first to m-th convolution layers L CNV . Specifically, the controller 11 related to Case CS_A2 obtains the following average value AVE[i]_A2 as the evaluation index EV A [i].

[0101] The average value AVE[i]_A2 is the average of the average values AVE[i,1]_A2 to AVE[i,m]_A2. The average value AVE[i,1]_A2 is the average of the total (r × CHa × K) singular values derived for the first convolutional layer L in the i-th model. Similarly, the average value AVE[i,2]_A2 is the average of the total (r × CHa × K) singular values derived for the second convolutional layer L in the i-th model. The same applies to the average values AVE[i,3]_A2 and so on. That is, for an integer j satisfying “1 ≦ j ≦ m”, the average value AVE[i,j]_A2 is the average of the total (r × CHa × K) singular values derived for the j-th convolutional layer L in the i-th model. CNV For the CNV layer, it is the average of the total (r × CHa × K) singular values derived. Similarly, for the average value AVE[i,2]_A2, it is the average of the total (r × CHa × K) singular values derived for the second convolutional layer L in the i-th model. The same applies to the average values AVE[i,3]_A2 and so on. That is, for an integer j satisfying “1 ≦ j ≦ m”, the average value AVE[i,j]_A2 is the average of the total (r × CHa × K) singular values derived for the j-th convolutional layer L in the i-th model. CNV For the CNV layer, it is the average of the total (r × CHa × K) singular values derived. Similarly, for the average value AVE[i,3]_A2, it is the average of the total (r × CHa × K) singular values derived for the third convolutional layer L in the i-th model. The same applies to the average values AVE[i,4]_A2 and so on. That is, for an integer j satisfying “1 ≦ j ≦ m”, the average value AVE[i,j]_A2 is the average of the total (r × CHa × K) singular values derived for the j-th convolutional layer L in the i-th model. CNV For the CNV layer, it is the average of the total (r × CHa × K) singular values derived. Similarly, for the average value AVE[i,4]_A2, it is the average of the total (r × CHa × K) singular values derived for the fourth convolutional layer L in the i-th model. The same applies to the average values AVE[i,5]_A2 and so on. That is, for an integer j satisfying “1 ≦ j ≦ m”, the average value AVE[i,j]_A2 is the average of the total (r × CHa × K) singular values derived for the j-th convolutional layer L in the i-th model.

[0102] In case CS_A2, since the design parameters may be different for each convolutional layer L CNV one or more of the values of r, CHa, and K may be different for each convolutional layer L CNV For example, in the first convolutional layer L CNV “CHa = 8”, while in the second convolutional layer L CNV “CHa = 16” may be the case.

[0103] The controller 11 according to case CS_A2 may set the average value AVE[i]_A3 of all the singular values obtained for the first to m-th convolutional layers L of the i-th model as the evaluation index EV CNV [i]. If the values of r, CHa, and K are common among the first to m-th convolutional layers L A then “AVE[i]_A3 = AVE[i]_A2” holds, but if not, “AVE[i]_A3 = AVE[i]_A2” may not hold. CNV As a variation, the controller 11 according to case CS_A2 may obtain the evaluation index EV

[0104] [i] by substituting the average value AVE[i]_A2 or AVE[i]_A3 into a predetermined arithmetic expression. For example, for the evaluation index EV A [i], it may be obtained by substituting the average value AVE[i]_A2 or AVE[i]_A3 into a predetermined arithmetic expression. For example, for the evaluation index EV A[i] may be the sum of the average value AVE[i]_A2 or AVE[i]_A3 and a predetermined value. In any case, it is assumed that the evaluation index EV A increases as the average value AVE[i]_A2 or AVE[i]_A3 increases.

[0105] --- Serialization based on the evaluation index --- In the first embodiment, the evaluation index EV A itself is the evaluation index EV[i]. Therefore, the controller 11 according to the first embodiment performs serialization of the first to nth models based on the evaluation index EV A [1] to EV A [n].

[0106] Specifically, the controller 11 assigns a higher rank to the i-th model as the evaluation index EV A [i] is larger. That is, the controller 11 assigns a higher rank to the e-th model than to the f-th model when the evaluation index EV A [e] is larger than the evaluation index EV A [f]. Therefore, among the evaluation indexes EV A [1] to EV A [n], if the evaluation index EV A [z] is the largest, the z-th model is assigned the first rank (z is a natural number not exceeding n), and the z-th model is set as the target model.

[0107] "EV A [i] = AVE[i]_A1", if "AVE[e]_A1 > AVE[f]_A1" holds, the controller 11 assigns a higher rank to the e-th model than to the f-th model. Conversely, in the case of "EV A [i] = AVE[i]_A1", if "AVE[e]_A1 < AVE[f]_A1" holds, the controller 11 assigns a higher rank to the f-th model than to the e-th model.

[0108] Similarly, "EV AWhen [i] = AVE[i]_A2, if "AVE[e]_A2 > AVE[f]_A2" holds, the controller 11 gives a higher rank to the e-th model than the f-th model. Conversely, "EV A When [i] = AVE[i]_A2, if "AVE[e]_A2 < AVE[f]_A2" holds, the controller 11 gives a higher rank to the f-th model than the e-th model. "EV A The same applies when [i] = AVE[i]_A3. As described above, e and f represent any different natural numbers less than or equal to n.

[0109] In the learning of a neural network including a convolutional layer, it is known that the singular values of the weight matrix corresponding to the above matrix A are related to the learning and generalization performance. As a document showing this matter, Non-Patent Document 1: Daisuke Suzuki, "Kyushu University Intensive Lecture Mathematics of Deep Learning and Machine Learning", [online], September 2020, [searched on September 13, 2023], Internet <URL: http: / / ibis.t.u-tokyo.ac.jp / suzuki / lecture / 2020 / intensive2 / Kyusyu_2020_Deep.pdf>, can be cited. It is known that the larger the singular value of the weight matrix, the faster the error between the model output data Dout and the correct data tends to decay.

[0110] Now, consider training a model α with a relatively large singular value of the weight matrix and a model β with a relatively small singular value of the weight matrix. At this time, the above error decays faster (the learning progresses faster) for the former model α than for the latter model β. If there were no time limit for this training, the generalization performance of both models after this training could be made the same. However, in reality, there is a limit to the time available for this training. Considering terminating this training within the limited time, the former model α can obtain higher generalization performance than the latter model β. That is, the former model α after this training is expected to have higher inference accuracy than the latter model β after this training.

[0111] In the first embodiment, an evaluation index EV A [i] is derived based on the average value of the singular values of the weight matrix in the i-th model. Therefore, it is possible to appropriately evaluate the quality of the structure of the i-th model based on the magnitude of the average value of the singular values. This is because a model with a relatively high average value of singular values is considered to correspond to model α, and a model with a relatively low average value of singular values is considered to correspond to model β. Further, the evaluation of good or bad here is an evaluation of whether the learning of the i-th model progresses quickly / slowly, or whether the expected generalization performance is high / low. Based on the result of this evaluation, the first to n-th models can be serialized, and this learning is performed on the model (target model) to which the highest rank is assigned. Thereby, a highly accurate learned model (the target model after this learning having good generalization performance) can be obtained.

[0112] <<Second Embodiment>> The second embodiment will be described. The controller 11 according to the second embodiment performs singular value decomposition on the weight parameters of the i-th model after learning for evaluation, and among the plurality of singular values obtained by the singular value decomposition, the threshold TH B counts the total number Q B [i] of the singular values exceeding it. The total number of singular values obtained by the singular value decomposition of the weight parameters of the i-th model is represented by the symbol "P B [i]", and is referred to as the total number P B [i] of singular values, etc. The controller 11 calculates the ratio R B of the total number Q B [i] of the singular values exceeding the threshold TH B to the total number P B [i] of singular values, and performs an operation to obtain the evaluation index EV[i] based on the ratio R B [i]. In the evaluation index derivation process in step S18 of FIG. 6, by executing this operation for each of the first to n-th models, the evaluation index EV for each of the first to n-th models is derived (that is, the evaluation indexes EV[1] to EV[n] are derived). The threshold TH B is predetermined. That is, prior to the execution of the operation in FIG. 5, the threshold TH B is given to the learning device 10, and the threshold TH Bis retained.

[0113] The weight parameter of the i-th model after the evaluation learning refers to the weight parameter of the i-th model at the stage where the evaluation index derivation process is executed through the evaluation learning for the i-th model. Specifically, the weight parameter of the i-th model refers to the weight parameter of the convolutional layer L provided in the NN120 of the i-th model CNV In the second embodiment, unless otherwise specified, the weight parameter after the evaluation learning is referred to as the weight parameter shown below.

[0114] The singular value decomposition is as shown in the first embodiment. Therefore, for the target convolutional layer L CNV which is the target convolutional layer L CNV (r×CHa×K) singular values are derived by singular value decomposition.

[0115] For the case CS_B1 where only one convolutional layer L CNV is provided in the NN120 of the i-th model and the case CS_B2 where m convolutional layers L CNV are provided in the NN120 of the i-th model, the method for deriving the evaluation index EV[i] will be described separately. However, for convenience of explanation, in the second embodiment, the symbol "EV B " is introduced as the symbol related to the evaluation index, and the evaluation index of the i-th model is specifically represented by "EV B [i]". m represents an integer of 2 or more.

[0116] ---Case CS_B1 (one convolutional layer L CNV )--- In case CS_B1, the total number P B [i] of singular values is (r×CHa×K). Among the P B [i] singular values, the number of singular values exceeding the threshold TH B becomes the value of Q B [i]. The above-mentioned ratio R B [i] satisfies "R B [i]=Q B [i] / P B [i]". The controller 11 uses the ratio R B [i] as the evaluation index EVB Find it as [i].

[0117] For example, in the case of case CS_B1 where “(r, CHa, K) = (3, 5, 8)”, “P B [i]=r×CHa×K = 3×5×8 = 120”. In this case, among the 120 singular values derived for the i-th model, only 50 singular values exceed the threshold TH B When it exceeds, “R B [i]=Q B [i] / P B [i]=50 / 120”.

[0118] As a variation, the controller 11 related to case CS_B1 may obtain the evaluation index EV B [i] by substituting the ratio R B [i] into a predetermined arithmetic expression. For example, the evaluation index EV B [i] may be the sum of the ratio R B [i] and a predetermined value. In any case, it is assumed that the evaluation index EV B [i] increases as the ratio R B [i] increases.

[0119] ---Case CS_B2 (Convolutional layer L CNV is m pieces)--- In case CS_B2, the m convolutional layers L CNV provided in the NN120 of the i-th model are referred to as the first to m-th convolutional layers L CNV .

[0120] In case CS_B2, by performing singular value decomposition for each convolutional layer L CNV , (r×CHa×K) singular values are obtained for each convolutional layer L CNV . However, the values of r, CHa, and K may differ between two or more convolutional layers L CNV . Here, for the sake of simplicity of explanation, it is assumed that the values of r, CHa, and K are common among the first to m-th convolutional layers L CNV . Then, among the first to m-th convolutional layers L CNVBy performing singular value decomposition on, a total of "(r × CHa × K) × m" singular values are obtained. That is, in case CS_B2, the total number of singular values P B [i] is "(r × CHa × K) × m". P B Among the P B [i] singular values, the number of singular values exceeding the threshold TH B becomes the value of Q B [i]. The above-mentioned ratio R B [i] satisfies "R B [i] = Q B [i] / P B [i]". The controller 11 obtains the ratio R B [i] as the evaluation index EV

[0121] For example, in case CS_B2, when "(r, CHa, K) = (3, 5, 8)" and "m = 4", "P B [i] = r × CHa × K × m = 3 × 5 × 8 × 4 = 480". In this case, among the 480 singular values derived for the i-th model, if only 200 singular values exceed the threshold TH B , then "R B [i] = Q B [i] / P B [i] = 200 / 480".

[0122] As a variation, the controller 11 according to case CS_B2 may obtain the evaluation index EV B [i] by substituting the ratio R B [i] into a predetermined arithmetic expression. For example, the evaluation index EV B [i] may be the sum of the ratio R B [i] and a predetermined value. In any case, it is assumed that the evaluation index EV B [i] increases as the ratio R B [i] increases.

[0123] ---Serialization Based on Evaluation Index--- In the second embodiment, the evaluation index EV B [i] itself is the evaluation index EV[i]. Therefore, the controller 11 according to the second embodiment uses the evaluation index EV B[1] to EV B Perform serialization of the first to the nth models based on [n].

[0124] Specifically, the controller 11 uses the evaluation index EV B The larger [i] is, the higher the rank given to the ith model. That is, the controller 11 uses the evaluation index EV B If [e] is the evaluation index EV B and [e] is larger than [f], the controller 11 gives a higher rank to the eth model than to the fth model. Therefore, among the evaluation indices EV B [1] to EV B [n], if the evaluation index EV B [z] is the largest, the zth model is assigned the first rank (z is a natural number less than or equal to n), and the zth model is set as the target model.

[0125] "EV B [i]=R B [i]". In the case where "R B [e]>R B [f]" holds, the controller 11 gives a higher rank to the eth model than to the fth model. Conversely, in the case of "EV B [i]=R B [i]". If "R B [e]<R B [f]" holds, the controller 11 gives a higher rank to the fth model than to the eth model. As described above, e and f represent any different natural numbers less than or equal to n.

[0126] Now, consider co-learning a model α with relatively large singular values of the weight matrix and a model β with relatively small singular values of the weight matrix. At this time, the error of the former model α decays faster (the learning progresses faster) than that of the latter model β. If there were no time limit for co-learning, it would be possible to make the generalization performance of both models after co-learning comparable. However, in reality, there is a limit to the time available for co-learning. When considering terminating co-learning within the limited time, the former model α can obtain higher generalization performance than the latter model β. That is, the former model α after co-learning is expected to have higher inference accuracy than the latter model β after co-learning.

[0127] In the second embodiment, among the total number of singular values of the weight matrix in the i-th model, the ratio R B of the total number of singular values exceeding the threshold TH B [i] is used to derive the evaluation index EV B [i]. Therefore, based on the magnitude of the ratio R B [i], the quality of the structure of the i-th model can be appropriately evaluated. A model with a relatively high ratio R B [i] corresponds to model α, and a model with a relatively low ratio R B [i] is considered to correspond to model β. Also, the evaluation of good or bad here is an evaluation of whether the learning of the i-th model progresses quickly / slowly, or whether the expected generalization performance is high / low. Based on the result of this evaluation, the first to n-th models can be serialized, and co-learning is performed on the model (target model) assigned the highest rank. Thereby, a highly accurate learned model (the target model after co-learning with good generalization performance) can be obtained.

[0128] <<Third Embodiment>> The third embodiment will be described. The controller 11 according to the third embodiment counts the total number Q C of weight parameters exceeding the threshold TH C among the weight parameters of the i-th model after evaluation learning. The total number of weight parameters in the i-th model is denoted by the symbol "P CDenoted by "[i]", the total number P of weight parameters C Such as "[i]". The controller 11 has a threshold value TH C The total number Q of weight parameters exceeding C [i] is the total number P of weight parameters C [i] The ratio R occupying C [i] is derived, and an operation for obtaining the evaluation index EV[i] is performed based on the ratio R. In the evaluation index derivation process in step S18 of FIG. 6, by executing this operation for each of the first to nth models, the evaluation index EV for each of the first to nth models is derived (that is, the evaluation indexes EV[1] to EV[n] are derived). The threshold value TH C Is predetermined. That is, prior to the execution of the operation in FIG. 5, the threshold value TH is given to the learning device 10, and the threshold value TH is held in the memory 12 or the like. C Is given, and the threshold value TH is held in the memory 12 or the like. C Is given, and the threshold value TH is held in the memory 12 or the like. C Is retained.

[0129] The weight parameter of the i-th model after learning for evaluation refers to the weight parameter of the i-th model at the stage where the evaluation index derivation process is executed through evaluation learning for the i-th model. Specifically, the weight parameter of the i-th model refers to the weight parameter of the convolutional layer L provided in the NN120 of the i-th model. CNV In the third embodiment, unless otherwise specified, the following weight parameters refer to the weight parameters after learning for evaluation.

[0130] The convolutional layer L in the NN120 of the i-th model CNV Case CS_C1 where only one is provided, and case CS_C2 where m convolutional layers L are provided in the NN120 of the i-th model will be separately described for the method of deriving the evaluation index EV[i]. However, for the sake of convenience of explanation, in the third embodiment, the symbol "EV" is introduced as the symbol related to the evaluation index, and the evaluation index of the i-th model is specifically represented by "EV". CNV "[i]". m represents an integer of 2 or more. C " is introduced, and the evaluation index of the i-th model is specifically represented by "EV". C [i]". m represents an integer of 2 or more.

[0131] ---Case CS_C1 (Convolutional layer L CNV Is one)--- In case CS_C1, the total number P of weight parameters C [i], as described with reference to FIG. 9, is (K W ×K H ×CHa×K). P C Among the P C [i] weight parameters, the number of weight parameters exceeding the threshold TH C becomes the value of Q C [i]. The above-mentioned ratio R C [i] satisfies "R C [i]=Q C [i] / P C [i]". The controller 11 obtains the ratio R C [i] as the evaluation index EV

[0132] For example, in case CS_C1, when "(K W ×K H ×CHa×K)=(3, 3, 5, 8)", "P C [i]=K W ×K H ×CHa×K = 3×3×5×8 = 360". In this case, among the 360 weight parameters derived for the i-th model, if only 110 weight parameters exceed the threshold TH C , "R C [i]=Q C [i] / P C [i]=110 / 360".

[0133] As a variation, the controller 11 according to case CS_C1 may obtain the evaluation index EV C [i] by substituting the ratio R C [i] into a predetermined arithmetic expression. For example, the evaluation index EV C [i] may be the sum of the ratio R C [i] and a predetermined value. In any case, it is assumed that the evaluation index EV C [i] increases as the ratio R C [i] increases.

[0134] ---Case CS_C2 (Convolution layer L CNV has m pieces)--- In case CS_C2, the m convolutional layers L provided in the NN120 of the i-th model CNV are referred to as the first to the m-th convolutional layers L CNV .

[0135] In case CS_C2, for each convolutional layer L CNV the total number of weight parameters is (K W × K H × CHa × K). However, the values of K W , K H , CHa, and K can be different between the convolutional layers L CNV where the number is 2 or more. For simplicity of explanation here, assume that the values of K W , K H , CHa, and K are common among the first to the m-th convolutional layers L CNV . Then, since the total number of weight parameters in the first to the m-th convolutional layers L CNV is the total number P C [i] of the weight parameters, "P C [i] = (K W × K H × CHa × K) × m". Among the P C [i] weight parameters, the number of weight parameters exceeding the threshold TH C becomes the value of Q C [i]. The above ratio R C [i] satisfies "R C [i] = Q C [i] / P C [i]". The controller 11 obtains the ratio R C [i] as the evaluation index EV C [i].

[0136] For example, in case CS_C2, when "(K W × K H × CHa × K) = (3, 3, 5, 8)" and "m = 4", "P C [i] = K W × K H × CHa × K × m = 3 × 3 × 5 × 8 × 4 = 1440". In this case, among the 1440 weight parameters derived for the i-th model, only 730 weight parametersC When it exceeds, "R C [i]=Q C [i] / P C [i]=730 / 1440".

[0137] As a variation, the controller 11 according to the case CS_C2 may obtain the evaluation index EV C [i] by substituting the ratio R C [i] into a predetermined arithmetic expression. For example, the evaluation index EV C [i] may be the sum of the ratio R C [i] and a predetermined value. In any case, it is assumed that the evaluation index EV C [i] increases as the ratio R C [i] increases.

[0138] --- Serialization Based on Evaluation Index --- In the third embodiment, the evaluation index EV C [i] itself is the evaluation index EV[i]. Therefore, the controller 11 according to the third embodiment performs serialization of the first to nth models based on the evaluation indexes EV C [1] to EV C [n].

[0139] Specifically, the controller 11 assigns a higher rank to the ith model as the evaluation index EV C [i] is larger. That is, the controller 11 assigns a higher rank to the eth model than the fth model when the evaluation index EV C [e] is larger than the evaluation index EV C [f]. Therefore, among the evaluation indexes EV C [1] to EV C [n], if the evaluation index EV C [z] is the largest, the first rank is assigned to the zth model (z is a natural number less than or equal to n), and the zth model is set as the target model.

[0140] "EV C [i]=R C [i]" In the case of, "R C [e]>R CIf " C C f] = R C C e] < R

[0141] In the model after evaluation learning, a filter FLT with a large number of large weight parameters can be considered to extract more significant features in the input map IN CNV than a filter FLT without such parameters (see Fig. 9). Therefore, a model with a relatively high ratio R C B i] is expected to extract more significant features in the input map IN CNV than a model with a relatively low ratio R

[0142] Therefore, in the third embodiment, the quality of the structure of the i-th model is evaluated based on the magnitude of the ratio R C C i]. Since the degree of extraction of significant features can be estimated from the magnitude of the ratio R CNV

[0143] --- Consideration on the model evaluation method of the first to third embodiments --- ​​​​​Here, an explanation will be added to the model evaluation method shown in the first to third embodiments. The model evaluation method shown in the first to third embodiments belongs to a method of deriving an evaluation index EV[i] based on each value of a plurality of weight parameters in the convolutional layer L of the i-th model after evaluation learning. Then, using the evaluation indices EV[1] to EV[n] derived for the first to n-th models, the quality of the structure of each model is evaluated. Based on the evaluation of good or bad, the first to n-th models are serialized. This is different from the method of determining that a DNN with a larger number of parameters is better, such as the existing evaluation method using existing evaluation indices (the evaluation method by existing Zero-Shot NAS). CNV It belongs to a method of deriving an evaluation index EV[i] based on each value of a plurality of weight parameters in the convolutional layer L CNV of the i-th model after evaluation learning. Then, using the evaluation indices EV[1] to EV[n] derived for the first to n-th models, the quality of the structure of each model is evaluated. Based on the evaluation of good or bad, the first to n-th models are serialized. This is different from the method of determining that a DNN with a larger number of parameters is better, such as the existing evaluation method using existing evaluation indices (the evaluation method by existing Zero-Shot NAS).

[0144] By using the evaluation index EV based on each value of the weight parameters, it is possible to appropriately evaluate the quality of the structure of each model. If the quality of the structure of each model can be appropriately evaluated, the first to n-th models can be appropriately serialized (a high rank can be given to a truly excellent model). Then, by performing this learning on the model (target model) to which the highest rank is assigned, a highly accurate learned model (the target model after this learning having good generalization performance) can be obtained.

[0145] <<Fourth Embodiment>> The fourth embodiment will be described. The controller 11 according to the fourth embodiment performs an operation of obtaining an evaluation index EV[i] based on the average value of the scale parameters in the i-th model after evaluation learning. In the evaluation index derivation process in step S18 of FIG. 6, by executing this operation for each of the first to n-th models, the evaluation index EV for each of the first to n-th models is derived (that is, the evaluation indices EV[1] to EV[n] are derived).

[0146] The scale parameter of the i-th model after evaluation learning refers to the scale parameter of the i-th model at the stage when the evaluation index derivation process is executed after evaluation learning for the i-th model. Specifically, the scale parameter of the i-th model is, in detail, the batch normalization layer L provided in the NN120 of the i-th model BN BNrefers to the scale parameter. In the fourth embodiment, unless otherwise specified, the scale parameter shown below refers to the scale parameter after learning for evaluation.

[0147] For the NN120 of the i-th model, the batch normalization layer L BN Case CS_D1 where only one is provided, and for the NN120 of the i-th model, the batch normalization layer L BN Case CS_D2 where m are provided, will be separately described for the method of deriving the evaluation index EV[i]. However, for the sake of convenience in explanation, in the fourth embodiment, the symbol “EV D ” is introduced as the symbol related to the evaluation index, and the evaluation index of the i-th model is specifically represented by “EV D [i]”. m represents an integer of 2 or more.

[0148] ---Case CS_D1 (batch normalization layer L BN is one)--- In case CS_D1, the single batch normalization layer L BN has the structure shown in FIG. 11. At this time, the scale parameter of the j-th channel in the batch normalization layer L BN is γ j (see the above formula (2)). For the batch normalization layer L BN , a total of CHb scale parameters γ1 to γ BN which are the scale parameters for the number of channels CHb of the batch normalization layer L CHb are set.

[0149] The controller 11 related to case CS_D1 obtains the average value of the scale parameters γ1 to γ CHb as the average value AVE[i]_D1, and sets the average value AVE[i]_D1 as the evaluation index EV D [i].

[0150] As a variation, the controller 11 related to case CS_D1 may obtain the evaluation index EV D [i] by substituting the average value AVE[i]_D1 into a predetermined arithmetic expression. For example, the evaluation index EV D[i] may be the sum of the average value AVE[i]_D1 and a predetermined value. In any case, it is assumed that the evaluation index EV D [i] increases as the average value AVE[i]_D1 increases.

[0151] ---Case CS_D2 (Batch normalization layer L BN is m)--- In case CS_D2, the m batch normalization layers L BN provided in the NN120 of the i-th model are referred to as the 1st to m-th batch normalization layers L BN .

[0152] In case CS_D2, the number of scale parameters in each batch normalization layer L BN is CHb. However, the value of CHb may be different between the batch normalization layers L BN where the value is 2 or more. Here, for simplicity of explanation, it is assumed that the value of CHb is common among the 1st to m-th batch normalization layers L BN . Then, the total number of scale parameters in the 1st to m-th batch normalization layers L BN is (CHb × m).

[0153] The controller 11 according to case CS_D2 obtains the average value of the scale parameters for each batch normalization layer L BN , and then takes the average of the total m average values obtained for the 1st to m-th batch normalization layers L BN . Specifically, the controller 11 according to case CS_D2 obtains the following average value AVE[i]_D2 as the evaluation index EV D [i].

[0154] The average value AVE[i]_D2 is the average value of the average values AVE[i,1]_D2 to AVE[i,m]_D2. The average value AVE[i,1]_D2 is the average value of the total CHb scale parameters set for the 1st batch normalization layer L BN in the i-th model. Similarly, the average value AVE[i,2]_D2 is the average value of the scale parameters of the 2nd batch normalization layer L BNIt is the average value of a total of CHb scale parameters set for. The same applies to the average value AVE[i,3]_D2 and the like. That is, for an integer j satisfying “1 ≦ j ≦ m”, the average value AVE[i,j]_D2 is the j-th batch normalization layer L in the i-th model BN It is the average value of a total of CHb scale parameters set for.

[0155] The controller 11 according to the case CS_D2 is the 1st to m-th batch normalization layers L of the i-th model BN The average value AVE[i]_D3 of all scale parameters set for may be set as the evaluation index EV D [i]. If the value of CHb is common among the 1st to m-th batch normalization layers L BN then “AVE[i]_D3 = AVE[i]_D2” holds, but if not, “AVE[i]_D3 = AVE[i]_D2” may not hold.

[0156] As a variation, the controller 11 according to the case CS_D2 may obtain the evaluation index EV D [i] by substituting the average value AVE[i]_D2 or AVE[i]_D3 into a predetermined arithmetic expression. For example, the evaluation index EV D [i] may be the sum of the average value AVE[i]_D2 or AVE[i]_D3 and a predetermined value. In any case, it is assumed that the evaluation index EV D [i] increases as the average value AVE[i]_D2 or AVE[i]_D3 increases.

[0157] --- Serialization Based on Evaluation Index --- In the fourth embodiment, the evaluation index EV D [i] itself is the evaluation index EV[i]. Therefore, the controller 11 according to the fourth embodiment performs serialization of the 1st to n-th models based on the evaluation indexes EV D [1] to EV D [n].

[0158] Specifically, the controller 11 is the evaluation index EV DThe larger [i] is, the higher rank is given to the i-th model. That is, the controller 11 uses the evaluation index EV D when [e] is greater than the evaluation index EV D [f], a higher rank is given to the e-th model than to the f-th model. Therefore, among the evaluation indexes EV D [1] to EV D [n], if the evaluation index EV D [z] is the largest, the first rank is assigned to the z-th model (z is a natural number not exceeding n), and the z-th model is set as the target model.

[0159] When "EV D [i]=AVE[i]_D1", if "AVE[e]_D1>AVE[f]_D1" holds, the controller 11 gives a higher rank to the e-th model than to the f-th model. Conversely, when "EV D [i]=AVE[i]_D1", if "AVE[e]_D1<AVE[f]_D1" holds, the controller 11 gives a higher rank to the f-th model than to the e-th model.

[0160] Similarly, when "EV D [i]=AVE[i]_D2", if "AVE[e]_D2>AVE[f]_D2" holds, the controller 11 gives a higher rank to the e-th model than to the f-th model. Conversely, when "EV D [i]=AVE[i]_D2", if "AVE[e]_D2<AVE[f]_D2" holds, the controller 11 gives a higher rank to the f-th model than to the e-th model. The same applies to the case of "EV D [i]=AVE[i]_D3". As described above, e and f represent any different natural numbers not exceeding n.

[0161] In the batch normalization layer L BN the input map IN BNScaling is performed by multiplying a large scale parameter for significant features in []. For this reason, a model with a relatively large scale parameter is expected to contain more significant features for the input map IN than a model with a relatively small scale parameter. BN In a model with a relatively large scale parameter, the average value of the scale parameter becomes higher than that in a model with a relatively small scale parameter. The significant features mean features useful for identifying the detection target in object detection, for example.

[0162] Therefore, in the fourth embodiment, the quality of the structure of the i-th model is evaluated based on the average value of the scale parameter. Since the degree of inclusion of significant features in the input map IN can be estimated from the magnitude of the average value of the scale parameter, the quality of the structure of the i-th model can be appropriately evaluated. The evaluation of good or bad here is an evaluation of whether the input map IN BN to the batch normalization layer L BN contains many significant features or not. In other words, it is an evaluation of whether significant features can be appropriately extracted in the convolutional layer L BN preceding the batch normalization layer L BN or not. Based on the result of the evaluation, the first to n-th models can be serialized, and the present learning is performed on the model (target model) assigned the highest rank. If a model that can appropriately extract significant features is set as the target model, the generalization performance obtained after the present learning will be improved. That is, a highly accurate learned model (the target model after the present learning having good generalization performance) can be obtained. CNV

[0163] The model evaluation method shown in the fourth embodiment is for the batch normalization layer L of the i-th model after the learning for evaluation BNIt belongs to a method of deriving an evaluation index EV[i] based on each value of a plurality of scale parameters in []. Then, the goodness or badness of the structure of each model is evaluated using the evaluation indices EV[1] to EV[n] derived for the first to nth models. The first to nth models are serialized based on the goodness or badness evaluation. This is different from a method of determining that a DNN with a larger number of parameters is better, such as an evaluation method using existing evaluation indices (evaluation method by existing Zero-Shot NAS).

[0164] By using the evaluation index EV[i] based on each value of the scale parameter, it is possible to appropriately evaluate the goodness or badness of the structure of each model. If the goodness or badness of the structure of each model can be appropriately evaluated, the first to nth models can be appropriately serialized (a high rank can be given to a truly excellent model). Then, by performing main learning on the model (target model) to which the highest rank is assigned, a highly accurate learned model (target model after main learning having good generalization performance) can be obtained.

[0165] <<Fifth Embodiment>> The fifth embodiment will be described. Among the first to fourth embodiments, the evaluation index EV may be obtained by combining any two or more embodiments. That is, the evaluation index EV may be based on two or more evaluation indices among the evaluation indices EV A , EV B , EV C and EV D . A method of combining all of the first to fourth embodiments will be described.

[0166] In the evaluation index derivation process in step S18 of FIG. 6, the controller 11 performs an operation of obtaining the evaluation indices EV A [i], EV B [i], EV C [i] and EV D [i] for the i-th model. This operation is executed for each of the first to nth models. Therefore, for the first model, the evaluation indices EV A [1] to EV D [1] are derived, and for the second model, the evaluation indices EV A[2] to EV D [2] is derived. The same applies to other models. Evaluation index EV A [i] to EV D The derivation method of [i] is as shown in the first to fourth embodiments.

[0167] Various values derived by the controller 11 according to the fifth embodiment are shown in FIG. 14. The range of values that the evaluation index has is the evaluation index EV A to EV D Since they are different between, it is inappropriate to perform serialization by the sum of the evaluation index EV A to EV D Therefore, perform the following operations.

[0168] In the serialization process, the controller 11 uses the evaluation index EV A [1] to EV A [n] to perform a first operation of setting a ranking evaluation value RNK A for each model. The ranking evaluation value RNK A set for the i-th model is particularly referred to as the ranking evaluation value RNK A [i]. In the first operation, the controller 11 sorts the evaluation indices EV A [1] to EV A [n] in descending order, so that among the evaluation indices EV A [1] to EV A [n], for the j-th largest evaluation index EV A the value of j is set for the ranking evaluation value RNK A corresponding to the model. That is, when the evaluation index EV A [i] is the j-th largest evaluation index among the evaluation indices EV A [1] to EV A [n], for the ranking evaluation value RNK A [i] corresponding to the evaluation index EV A [i], the value of j is set. In the case where two or more evaluation indices having the same value are included in the evaluation indices EV A [1] to EV A [n], for each ranking evaluation value RNK of the two or more models corresponding to the two or more evaluation indicesA For this, the same value is set. However, this case is ignored here (the same applies to the second to fourth operations described later).

[0169] In the serialization process, the controller 11 performs a second operation of setting the ranking evaluation value RNK for each model based on the evaluation indicators EV B [1] to EV B [n]. The ranking evaluation value RNK B set for the i-th model is particularly referred to as the ranking evaluation value RNK B and is denoted as RNK B [i]. In the second operation, the controller 11 sorts the evaluation indicators EV B [1] to EV B [n] in descending order, and sets the value of j for the ranking evaluation value RNK B corresponding to the j-th largest evaluation indicator EV B among EV B [1] to EV B [n]. That is, if the evaluation indicator EV B [i] is the j-th largest evaluation indicator among EV B [1] to EV B [n], the value of j is set for the ranking evaluation value RNK B of the i-th model corresponding to the evaluation indicator EV B [i].

[0170] In the serialization process, the controller 11 performs a third operation of setting the ranking evaluation value RNK for each model based on the evaluation indicators EV C [1] to EV C [n]. The ranking evaluation value RNK C set for the i-th model is particularly referred to as the ranking evaluation value RNK C and is denoted as RNK C [i]. In the third operation, the controller 11 sorts the evaluation indicators EV C [1] to EV C [n] in descending order, and sets the value of j for the ranking evaluation value RNK C corresponding to the j-th largest evaluation indicator EV C among EV CSequence evaluation value RNK of the model corresponding to C Set the value of j for it. That is, evaluation index EV C [i] is the evaluation index EV C [1] to EV C [n], if it is the j-th largest evaluation index, the sequence evaluation value RNK of the i-th model corresponding to the evaluation index EV C [i] C Set the value of j for [i].

[0171] In the serialization process, the controller 11 performs a fourth operation of setting the sequence evaluation value RNK for each model based on the evaluation index EV D [1] to EV D [n]. The sequence evaluation value RNK set for the i-th model D Is specifically referred to as the sequence evaluation value RNK D Of [i]. D Is called RNK In the fourth operation, the controller 11 sorts the evaluation index EV D [1] to EV D [n] in descending order, so that among the evaluation index EV D [1] to EV D [n], for the j-th largest evaluation index EV D Set the value of j for the sequence evaluation value RNK of the corresponding model D That is, if the evaluation index EV D [i] is the j-th largest evaluation index among the evaluation index EV D [1] to EV D [n], the sequence evaluation value RNK of the i-th model corresponding to the evaluation index EV D [i] D Set the value of j for [i].

[0172] Furthermore, the controller 11 obtains the comprehensive evaluation value RNK for each of the first to n-th models in the serialization process TOTAL The comprehensive evaluation value RNK derived for the i-th model TOTAL Is specifically referred to as the comprehensive evaluation value RNK TOTAL [i]. The comprehensive evaluation value RNK TOTAL [i] is derived according to the following formula (4). Coefficient k A ~kD has a predetermined positive value. Typically, the coefficient k A ~k D may all be "1". RNK TOTAL [i]=k A ·RNK A [i]+k B ·RNK B [i] +k C ·RNK C [i]+k D ·RNK D [i] ···(4)

[0173] After that, the controller 11 ranks the first to nth models based on the comprehensive evaluation values RNK TOTAL [1]~RNK TOTAL [n]. In the ranking of the first to nth models, the first to nth models are respectively ranked. The controller 11 sorts the comprehensive evaluation values RNK TOTAL [1]~RNK TOTAL [n] in ascending order. Then, the controller 11 assigns the jth rank to the model corresponding to the jth smallest comprehensive evaluation value RNK TOTAL [1] among RNK TOTAL [1]~RNK TOTAL [n]. Therefore, for example, among the comprehensive evaluation values RNK TOTAL [1]~RNK TOTAL [n], if the comprehensive evaluation value RNK TOTAL [z] is the smallest, the first rank is assigned to the zth model corresponding to the comprehensive evaluation value RNK TOTAL [z] (z is a natural number not exceeding n). Among the first to nth models, the model to which the first rank is assigned is selected as the target model.

[0174] By the method shown in this embodiment, multiple types of evaluation indicators (EV A ~EV DBased on , the quality of the model structure is evaluated, and the target model to which this learning is applied is selected through ranking of the quality. By evaluating the quality of the model structure from multiple viewpoints, it is expected that a well-balanced and appropriate model can be selected as the target model.

[0175] Incidentally, for the coefficients k A ~k D , zero may be set for any one or two of the coefficients. For example, "k A =k B =0" and "k C >0" and "k D >0" may hold, and in this case, the serialization process combining the third and fourth embodiments is performed.

[0176] <<Sixth Embodiment>> The sixth embodiment will be described. Fig. 15 shows a functional block diagram of the controller 11. Functional blocks F1 to F3 are provided in the controller 11. By causing the controller 11 to execute a program recorded in the memory 12, the database 20, or any other arbitrary recording medium (not shown), each function of the functional blocks F1 to F3 may be realized. The relationship between each functional block and the flowchart of Fig. 5 will be described.

[0177] The functional block F1 is a model evaluation unit that performs the model evaluation process of step S10. The functional block F2 is a serialization process unit that performs the serialization process of step S30. The functional block F3 is a main learning process unit that performs the main learning process of step S50.

[0178] <<Seventh Embodiment>> The seventh embodiment will be described. In the seventh embodiment, modified techniques, applied techniques, etc. for each of the above matters are cited.

[0179] Although the method of sampling the first to nth models (setting the first to nth structures) using NAS has been described above, the use of NAS is not essential. As long as the first to nth models are n types of models having different structures from each other, the setting method of the first to nth models is arbitrary. Each structure of the first to nth models may be determined based on the input operation information given to the learning device 10 by the operator of the learning device 10.

[0180] The learned model (inference model), which is the target model after this learning, can be applied to an arbitrary device (hereinafter referred to as a model application device). Prior to this application, pruning is performed on the learned model in the learning device 10 or another device, and the learned model after pruning is re-learned using the learning data 21. Then, the learned model after pruning and after re-learning may be installed in the arithmetic processing unit in the model application device. The model application device may be an in-vehicle device installed in a vehicle such as an automobile. At this time, for example, the learned model performs object detection as an inference, and the in-vehicle device performs automatic driving or driving support or the like based on the result of the object detection.

[0181] For the method according to the present invention, an evaluation method using existing evaluation indices may be combined. That is, based on evaluation indices (EV A ~EV D ) that do not depend on the number of parameters and existing evaluation indices that depend on the number of parameters, the target model may be selected by serializing the first to nth models.

[0182] The learning device 10 includes a learning model evaluation device that evaluates the quality of the structures of the first to nth models through the derivation of the evaluation index EV. The learning device 10 may also be referred to as an information processing device. When evaluating the quality of the structures of the first to nth models through the derivation of the evaluation index EV, the information processing device functions as a learning model evaluation device. After this evaluation, when performing the present learning process based on the evaluation result, the information processing device functions as a learning device. The learning device can be considered as a device having a learning model evaluation device. The learning model evaluation device has a model evaluation unit F1 and a serialization processing unit F2, and the learning device can be considered as having a main learning processing unit F3 in addition to the model evaluation unit F1 and the serialization processing unit F2 (see Fig. 15).

[0183] A program for causing a computer device to execute any method described in the embodiments of the present invention, and a non-volatile recording medium recording the program are included in the scope of the embodiments of the present invention. The program for causing a computer device to execute any method described in the embodiments of the present invention may be a subprogram incorporated into an arbitrary main program or called by an arbitrary main program. The learning device 10 is a kind of computer device. Any process in the embodiments of the present invention may be realized by hardware such as a semiconductor integrated circuit, software corresponding to the above program, or a combination of hardware and software.

[0184] In the embodiments of the present invention, various modifications can be appropriately made within the scope of the technical idea shown in the claims. The above embodiments are merely examples of the embodiments of the present invention, and the meanings of the terms of the present invention or each constituent element are not limited to those described in the above embodiments. The specific numerical values shown in the above description are merely examples, and of course, they can be changed to various numerical values.

Explanation of Reference Numerals

[0185] 10 Learning device 11 Controller 12 Memory 13 Communication unit 110 model 120 NN 20 database 21 training data 21a image dataset Din model input data Dout model output data 121 input layer 122 intermediate layer 123 output layer BLK block L CNV convolution layer L BN batch normalization layer F1 model evaluation unit F2 serialization processing unit F3 main training processing unit

Claims

1. A learning model evaluation apparatus for evaluating a learning model constituted by a neural network having a convolutional layer, comprising: Performing evaluation learning for training the learning model using training data, and deriving an evaluation index based on each value of a plurality of weight parameters in the convolutional layer of the learning model after the evaluation learning. A learning model evaluation apparatus.

2. Performing singular value decomposition on the plurality of weight parameters in the learning model after the evaluation learning, and deriving the evaluation index based on an average value of a plurality of singular values obtained by the singular value decomposition. The learning model evaluation apparatus according to Claim 1.

3. Performing singular value decomposition on the plurality of weight parameters in the learning model after the evaluation learning, deriving a ratio of the total number of singular values exceeding a threshold value to the total number of the plurality of singular values obtained by the singular value decomposition, and deriving the evaluation index based on the ratio. The learning model evaluation apparatus according to Claim 1.

4. Deriving a ratio of the total number of weight parameters exceeding a threshold value to the total number of the plurality of weight parameters in the learning model after the evaluation learning, and deriving the evaluation index based on the ratio. The learning model evaluation apparatus according to Claim 1.

5. A learning model evaluation apparatus for evaluating a learning model constituted by a neural network having a convolutional layer and a batch normalization layer arranged after the convolutional layer, comprising: Performing evaluation learning for training the learning model using training data, and deriving an evaluation index based on each value of a plurality of scale parameters in the batch normalization layer of the learning model after the evaluation learning. A learning model evaluation apparatus.

6. Deriving the evaluation index based on an average value of the plurality of scale parameters in the learning model after the evaluation learning. The learning model evaluation apparatus according to Claim 5.

7. A learning apparatus having the learning model evaluation apparatus according to any one of Claims 1 to 6, comprising: Deriving a plurality of evaluation indexes for a plurality of learning models using the learning model evaluation apparatus; Selecting a target model from the plurality of learning models based on the plurality of evaluation indexes, and training the target model using the training data for a number of training times greater than that of the evaluation learning. A learning apparatus.

8. A learning model evaluation method for evaluating a learning model composed of a neural network having a convolutional layer, performing evaluation learning for training the learning model using training data, and deriving an evaluation index based on each value of a plurality of weight parameters in the convolutional layer of the learning model after the evaluation learning , a learning model evaluation method.

9. A learning model evaluation method for evaluating a learning model composed of a neural network having a convolutional layer and a batch normalization layer arranged after the convolutional layer, performing evaluation learning for training the learning model using training data, and deriving an evaluation index based on each value of a plurality of scale parameters in the batch normalization layer of the learning model after the evaluation learning , a learning model evaluation method.

10. A program for causing a computer to execute a learning model evaluation method for evaluating a learning model composed of a neural network having a convolutional layer, wherein in the learning model evaluation method, evaluation learning for training the learning model using training data is performed, and an evaluation index is derived based on each value of a plurality of weight parameters in the convolutional layer of the learning model after the evaluation learning , a learning model evaluation program.

11. A program for causing a computer to execute a learning model evaluation method for evaluating a learning model composed of a neural network having a convolutional layer and a batch normalization layer arranged after the convolutional layer, wherein in the learning model evaluation method, evaluation learning for training the learning model using training data is performed, and an evaluation index is derived based on each value of a plurality of scale parameters in the batch normalization layer of the learning model after the evaluation learning , a learning model evaluation program.

Citation Information

Patent Citations

  • JP7111671A