Machine learning system

The machine learning system addresses the time-consuming nature of evaluating pre-trained models for transfer learning by using a fitness evaluation unit to select suitable models without performing actual transfer learning, thus speeding up the process and reducing computational demands.

JP7696844B2Active Publication Date: 2025-06-23HITACHI HIGH TECH CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022004834
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-17
Publication Date
2025-06-23
Estimated Expiration
2042-01-17

AI Technical Summary

Technical Problem

Existing machine learning techniques for transfer learning are time-consuming when dealing with large combinations of datasets and model structures, as they require actual transfer learning processes to evaluate the effectiveness of pre-trained models.

Method used

A machine learning system that includes a pre-trained model acquisition unit, a transfer learning dataset storage unit, a pre-trained model fitness evaluation unit, and a transfer learning unit, which allows for the selection of a pre-trained model without performing actual transfer learning by evaluating the fitness of the models with respect to the target task dataset.

Benefits of technology

Enables the rapid selection of a suitable pre-trained model for transfer learning, significantly reducing the time required to obtain a new machine learning model, while minimizing computational resources by avoiding extensive parameter updates during the evaluation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696844000001
    Figure 0007696844000001
  • Figure 0007696844000002
    Figure 0007696844000002
  • Figure 0007696844000003
    Figure 0007696844000003
Patent Text Reader

Abstract

To provide a machine learning system and a machine learning method that can select a pre-trained model for use in transition learning in a short time without actually performing the transition learning.SOLUTION: A machine learning system of the present invention includes a pre-trained model acquisition unit that acquires at least one pre-trained model from a pre-trained model storage unit that stores a plurality of pre-trained models obtained by learning tasks of transition sources at respective conditions, a transition learning data set storage unit that stores a data set about the tasks of a transition destination, a pre-trained model fitness evaluation unit that evaluates a fitness for the pre-trained model acquired by the pre-trained model acquisition unit, and a transition learning unit that performs transition learning using the selected pre-trained model and the data set based on an evaluation result by the pre-trained model fitness evaluation unit and outputs a learning result as a trained model.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a machine learning system and a machine learning method.

Background Art

[0002] In machine learning techniques for data processing, particularly in multi-layer neural networks called deep learning, a technique called transfer learning is frequently used to improve performance on a new dataset by using past pre-trained models and training data. For example, Patent Document 1 describes evaluating whether the data used in pre-training was effective for transfer learning after performing transfer learning efficiently.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the technique described in Patent Document 1, the effectiveness of transfer learning is evaluated after actually performing transfer learning. However, when the combination of the dataset and model structure used for transfer learning is enormous, it takes time to perform transfer learning for all pre-trained models. As a result, in the technique described in Patent Document 1, it may take a long time to obtain a new machine learning model by transfer learning.

[0005] An object of the present invention is to provide a machine learning system and a machine learning method capable of selecting a pre-trained model used in transfer learning in a short time without actually performing transfer learning.

Means for Solving the Problems

[0006] In order to solve the above problems, the machine learning system of the present invention includes a pre-trained model acquisition unit that acquires at least one pre-trained model from a pre-trained model storage unit that stores a plurality of pre-trained models obtained by learning source tasks under respective conditions, a transfer learning dataset storage unit that stores a dataset related to a target task, a pre-trained model fitness evaluation unit that evaluates the fitness of each of the pre-trained models acquired by the pre-trained model acquisition unit with respect to the dataset related to the target task, and a transfer learning unit that performs transfer learning using the selected pre-trained model and the dataset based on the evaluation result of the pre-trained model fitness evaluation unit and outputs the learning result as a learned model.

[0007] Further, the machine learning method of the present invention includes a step of acquiring a pre-trained model obtained by learning a source task, a step of reading a dataset related to a target task, a step of evaluating the fitness of the pre-trained model with respect to the dataset related to the target task, and a step of performing transfer learning using the pre-trained model and the dataset based on the evaluation result and outputting the learning result as a learned model.

Advantages of the Invention

[0008] According to the present invention, it is possible to provide a machine learning system and a machine learning method capable of selecting a pre-trained model used in transfer learning in a short time without actually performing transfer learning.

[0009] Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Mode for Carrying Out the Invention

[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Detailed descriptions of overlapping parts will be omitted.

[0012] First, an overview of the method for selecting a pre-trained model according to this embodiment will be described with reference to FIG. 1. In this specification, a task to be subjected to transfer learning by learning with teacher information is called a target task or a downstream task, and a task to be learned by a pre-trained model used for transfer learning is called a source task or a pre-training task. Also, in this specification, teacher information may be referred to as teaching information.

[0013] In the downstream task, an input image 101, a teacher segmentation 102, and a teacher detection result 103 are provided. Here, either the teacher segmentation 102 or the teacher detection result 103 may be used. The teacher segmentation 102 provides information indicating to which class or instance each pixel belongs. Here, segmentation refers to a recognition process called semantic segmentation, instance segmentation, or panoptic segmentation. Also, the teacher detection result 103 provides information indicating the region to which the class of the detection target belongs in the input image 101 and the detected class. In transfer learning using the downstream task, a model that takes the input image 101 as input and outputs the teacher segmentation 102 or the teacher detection result 103 is learned.

[0014] Region segmentation results 104a and 104b are region segmentation results obtained from pre-trained models, specifically from different pre-trained models A and B, respectively. This indicates that they are segmented into regions with different texture differences. That is, in the region segmentation result 104a, the dog and the cat are segmented into regions belonging to the same class, while the tree and the house are segmented into regions belonging to different classes. On the other hand, in the region segmentation result 104b, the tree and the house are segmented into regions belonging to the same class, while the dog and the cat are segmented into regions belonging to different classes, and furthermore, the head and the body are also segmented into different classes. Here, the region segmentation result by the pre-trained model is, like the segmentation task, the result of dividing the input data into multiple regions so that they have the same meaning within the same region.

[0015] In the machine learning system according to this embodiment, a pre-trained model including a function for performing such region segmentation is learned using a pre-training task.

[0016] By comparing such region segmentation results 104a and 104b with the teacher segmentation 102 and the teacher detection result 103, a pre-trained model suitable for the downstream task can be selected. For example, the teacher segmentation 102 is a task of recognizing dogs and cats as different classes respectively. In the case of this task, the pre-trained model B that outputs the region segmentation result 104b in which dogs and cats are segmented as different classes is more suitable for transfer learning than the pre-trained model A that outputs the region segmentation result 104a in which dogs and cats are segmented as the same class. Similarly for the teacher detection result 103, in the case of a task of detecting dogs and cats as different classes, the pre-trained model B that outputs the region segmentation result 104b in which dogs and cats are different classes is more suitable for transfer learning.

[0017] Thus, in this embodiment, by evaluating the degree of fitness between the region segmentation results by each pre-trained model including the function of performing region segmentation and the teaching information of the downstream task, it is possible to select a pre-trained model effective for transfer learning in a short time without actually performing transfer learning.

[0018] In addition, the pre-trained model according to this embodiment has, in addition to the function of performing region segmentation, a function of outputting feature amounts for each position of the input data, and outputs similar feature amounts within regions estimated to have the same meaning. Therefore, the more the pre-trained model has a high relevance (similarity) with teaching information such as the region segmentation results 104a and 104b and the teacher segmentation 102 and the teacher detection result 103, the more it has learned different feature amounts for each given teaching information, has a high ability to identify the given teaching information, and is considered a model suitable for transfer learning.

Example

[0019] An example of a machine learning system that realizes the method for selecting a pre-trained model shown in FIG. 1 will be described. FIG. 2 is a diagram showing a configuration example of the machine learning system according to Embodiment 1 of the present invention, and FIG. 3 is a flowchart showing the behavior of the machine learning system according to Embodiment 1 of the present invention.

[0020] As shown in FIG. 2, the machine learning system according to this embodiment includes a pre-trained model storage unit 202, a pre-trained model acquisition unit 201, a transfer learning dataset storage unit 204, a pre-trained model fitness evaluation unit 203, a pre-trained model selection unit 205, and a transfer learning unit 206. The pre-trained model storage unit 202 stores a plurality of pre-trained models obtained by learning a pre-training task. The pre-trained model acquisition unit 201 acquires at least one pre-trained model from the pre-trained model storage unit 202. The transfer learning dataset storage unit 204 stores a dataset related to a downstream task. The pre-trained model fitness evaluation unit 203 evaluates the fitness of each pre-trained model acquired by the pre-trained model acquisition unit 201 with respect to the dataset related to the downstream task. Note that the pre-trained model to be evaluated by the pre-trained model fitness evaluation unit 203 is determined based on the pre-trained model evaluation target information 200. The pre-trained model selection unit 205 selects, based on the evaluation result of the pre-trained model fitness evaluation unit 203, a pre-trained model to be used for transfer learning from among the plurality of pre-trained models stored in the pre-trained model storage unit 202. The transfer learning unit 206 performs transfer learning using the pre-trained model selected by the pre-trained model selection unit 205 and the dataset related to the downstream task stored in the transfer learning dataset storage unit 204, and outputs the learning result as a learned model 207.

[0021] Note that each unit of the machine learning system can be configured by using hardware such as a circuit device that implements its function, or can be configured by having an arithmetic unit execute software that implements its function.

[0022] Next, the processing of the machine learning system according to this embodiment will be described with reference to FIG. 3. This processing flow is executed at a timing when the operator designates the execution of transfer learning.

[0023] In step S301, the pre-trained model fitness evaluation unit 203 reads, from the transfer learning dataset storage unit 204, a dataset related to the downstream task for which transfer learning is to be performed.

[0024] FIG. 4 is a diagram showing a configuration example of the transfer learning dataset storage unit 204. As shown in FIG. 4, the dataset related to the downstream task stored in the transfer learning dataset storage unit 204 is composed of a learning dataset 401 for learning and an evaluation dataset 404 for performance evaluation. Further, the learning dataset 401 and the evaluation dataset 404 are each composed of input data 402, 405 and teaching information 403, 406. The formats of these data are common between the learning dataset 401 and the evaluation dataset 404, but the contents of the data are different. The input data 402, 405 are, for example, image data such as the input image 101, and the teaching information 403, 406 are, for example, the teacher segmentation 102 and the teacher detection result 103.

[0025] In step S301, the pre-trained model fitness evaluation unit 203 reads the input data 402 and the teaching information 403 related to the learning dataset 401 from the transfer learning dataset storage unit 204.

[0026] In step S302, the pre-trained model acquisition unit 201 acquires one pre-trained model to be evaluated based on the pre-trained model evaluation target information 200. The pre-trained model evaluation target information 200 stores information related to the pre-trained model to be evaluated. The evaluation target may be specified by an operator, or may be automatically specified, for example, by conditions related to the dataset used for pre-training or conditions related to the configuration of the pre-trained model.

[0027] In step S303, the pre-training model fitness evaluation unit 203 reads the calculation procedure and parameters of the pre-training model to be evaluated obtained in step S302 from the pre-training model storage unit 202. Here, the calculation procedure is information composed of the type and order of calculations applied to the input data in the case of a multi-layer neural network. For example, the configuration of the neural network, pre-processing, post-processing, etc. are applicable. Also, the parameters are the parameters used in the calculation procedure, such as the learned weight parameters and hyperparameters of each layer of the neural network, or the parameters of pre-processing and post-processing.

[0028] Figure 5 is a diagram showing a configuration example of the pre-training model storage unit 202. As shown in Figure 5, one or more pre-training models 501 (pre-training model A, pre-training model B,...) stored in the pre-training model storage unit 202 are each composed of calculation procedure information 502, parameter information 503, learning dataset information 504, and learning condition information 505. The calculation procedure information 502 and the parameter information 503 are the calculation procedure and parameters that the pre-training model fitness evaluation unit 203 acquires in step S303. The learning dataset information 504 stores information about the dataset used for the learning of the pre-training model 501. For example, it is information for specifying the dataset, or tag information such as the domain, acquisition conditions, classes, etc. associated with the dataset. The learning condition information 505 is information about the learning conditions of the pre-training model 501, such as learning schedule information such as the learning period, or information about hyperparameters in learning. The learning dataset information 504 and the learning condition information 505 are presented to the operator to determine the pre-training model evaluation target information 200.

[0029] In step S304, the pre-training model fitness evaluation unit 203 evaluates the fitness of the pre-training model read in step S303 with respect to the dataset of the downstream task to be evaluated read in step S301. The fitness in this embodiment is the degree of fitness with respect to the teaching information 403 of the region division results 104a and 104b by the pre-training model. The higher the fitness, the more effective the pre-training model for transfer learning can be considered.

[0030] A specific method for evaluating the fitness will be described. Here, when the assignment by region division by the pre-training model is the i-th class and the assignment by the teaching information 403 is the j-th class, the cost of the assignment is C ij is expressed as. The fitness can be expressed as the negative value obtained by multiplying the sum of the costs when solving the assignment problem between the region division by the pre-training model and the teaching information by -1. This is because the higher the fitness, the better the value, while the assignment problem is generally formulated to minimize the cost. Note that the assignment problem may perform a one-to-one assignment, or a many-to-one assignment may be performed when the number of region divisions by the pre-training model is larger than the number of classes included in the teaching information.

[0031] The assignment cost C ij As a calculation method of, a plurality of methods can be considered. For example, a method using conditional probability can be considered. The conditional probability that the teaching information is class j when the assignment by the pre-training model is class i can be expressed as p(y = j|z = i). y is the class information of the teaching information, and z is the class information divided by the pre-training model. At this time, the assignment cost is C ij = 1 - p(y = j|z = i). In addition to conditional probability, the assignment cost C ij can also be calculated using cross entropy, mutual information, etc. By using these values, it becomes possible to evaluate the probability of correctly obtaining the class of the teaching information 403 when the class of the region division by the pre-training model is determined.

[0032] Also, the assignment cost C ijFor the calculation of [the above], an index called Intersection over Union (IoU) may be used. IoU is a value obtained by dividing the intersection of the region segmentation result by the pre-trained model and the region given by the teaching information by the union set. In this case, the assignment cost C ij is given by = 1 - IoU(y = j|z = i). The second term on the right side is a function that calculates the IoU between the regions where the region segmented by the pre-trained model is class i and the teaching information is class j.

[0033] Also, for the calculation of the assignment cost C ij the similarity between feature quantities may be utilized. Specifically, in addition to the function of performing region segmentation on the pre-trained model, a function for outputting feature quantities for each region is added, and for each region of the class determined by the region segmentation result by the pre-trained model and the teaching information, it is obtained by comparing the averages of the feature quantities. At this time, the assignment cost C ij is given by = D(G(f, y = j), G(f, z = i)). Here, G is a function that obtains the average of the feature quantity f for each region specified from the feature quantity f, and G(f, y = j) and G(f, z = i) are the averages of the feature quantity f for each region where the region segmentation result by the pre-trained model and the teaching information are class i and class j, respectively. D is a function representing the distance between feature quantities, and is a distance function such as the Euclidean distance or the cosine distance.

[0034] Also, when the number of region segmentations by the pre-trained model is larger than the number of classes included in the teaching information, after reducing the number of region segmentations by the pre-trained model by applying clustering to the average feature quantity for each region segmentation by the pre-trained model, the evaluation of the assignment problem may be executed. Furthermore, when instance recognition is required instead of class units such as object detection, instance segmentation, or panoptic segmentation, the evaluation of the assignment problem may be performed in units of instances instead of class units. Also, when the teaching information is given as a detection window as in object detection, the region within the detection window may be treated as the region taught by the teaching information.

[0035] In step S305, the pre-training model fitness evaluation unit 203 checks whether all the evaluation targets included in the pre-training model evaluation target information 200 have been evaluated. If the evaluation is not completed, it returns to step S302 to continue the evaluation of the unevaluated pre-training model. If the evaluation of all the pre-training models of the evaluation targets is completed, it proceeds to step S306.

[0036] In step S306, the pre-training model selection unit 205 selects one or more models with high fitness for each pre-training model evaluated by the pre-training model fitness evaluation unit 203. In this embodiment, it is assumed that the pre-training model selection unit 205 automatically selects the model with high fitness. However, as in the embodiment 2 described later, it may be selected by the operator.

[0037] In step S307, the transfer learning unit 206 performs transfer learning on the pre-training model selected by the pre-training model selection unit 205 using the dataset stored in the transfer learning dataset storage unit 204, and outputs the learning result as the trained model 207.

[0038] According to the machine learning system according to this embodiment, it is possible to select, without actually performing transfer learning, a pre-trained model suitable for the dataset of the downstream task stored in the transfer learning dataset storage unit 204 from the pre-trained model stored in the pre-trained model storage unit 202. In particular, the pre-trained model fitness evaluation unit 203 of this embodiment only needs to perform forward propagation of the input data of the downstream task for each pre-trained model a smaller number of times (for example, once, at most 10 times per image) than during transfer learning. That is, according to this embodiment, compared with the case of actually performing transfer learning in which parameter updates by forward propagation and backward propagation are repeatedly performed many times on the pre-trained model, the amount of calculation can be significantly reduced, and the fitness can be evaluated in a short time. If you want to further reduce the amount of calculation related to the evaluation of the effectiveness of transfer learning, instead of all the input data included in the dataset to be evaluated, a part of the input data is acquired by a method considering randomness or the bias in the number of samples for each class, and the evaluation of the pre-trained model fitness evaluation unit 203 is performed only on a part of the input data.

[0039] The pre-training performed in this embodiment will be described with reference to FIGS. 6 and 7. FIG. 6 is a diagram showing an overview of pre-training, and FIG. 7 is a flowchart showing the processing of pre-training.

[0040] First, in step S701, the pre-training unit 603 reads the pre-training conditions 601. The pre-training conditions 601 are information regarding the conditions of pre-training, and are composed of information such as the pre-training dataset to be used, the calculation procedure of the model to be learned, information regarding parameter conditions, and hyperparameters such as the learning schedule of pre-training.

[0041] In step S702, the pre-training unit 603 acquires the pre-training dataset set in the pre-training conditions 601 from the pre-training dataset storage unit 602. One or more pre-training datasets may be acquired.

[0042] Here, the pre-training dataset storage unit 602 will be described with reference to FIG. 8. FIG. 8 is a diagram showing a configuration example of the pre-training dataset storage unit 602. As shown in FIG. 8, the pre-training dataset storage unit 602 stores pre-training datasets 801A, 801B, ··· and one or more pre-training datasets 801. Each pre-training dataset 801 is associated with input data 802A, 802B and tag information 803A, 803B. The input data 802A, 802B are the input data used in the pre-training task. In the pre-training of this embodiment, since self-supervised learning (unsupervised learning) is used, it is not necessary to have teaching information for each pre-training dataset 801. Also, since teaching information is not used, it is possible to combine and use a plurality of pre-training datasets 801. Note that the tag information 803A, 803B is tag information such as the domain, acquisition conditions, class, etc. associated with the dataset, and is the information stored as learning dataset information 504 in the pre-training model storage unit 202.

[0043] In step S703, the pre-training unit 603 initializes the pre-training model based on the pre-training model information set in the pre-training condition 601. Specifically, random number initialization and the like are performed according to the calculation procedure and parameter conditions set in the pre-training condition 601.

[0044] In step S704, the pre-training unit 603 updates the parameters of the pre-training model initialized in step S703 using the pre-training dataset acquired in step S702. Note that the parameter update is performed by an optimization method such as the stochastic gradient descent method.

[0045] Here, the pre-trained model will be described with reference to FIG. 9. FIG. 9 is a diagram showing a configuration example of the pre-trained model. The pre-trained model used in this embodiment includes a feature extraction unit 902 that extracts features from the input data 901, a region division unit 905 that divides the input data 901 into a plurality of regions so as to have the same meaning within the region using the features extracted by the feature extraction unit 902, and a feature prediction unit 903 that predicts different features for each region divided by the region division unit 905.

[0046] When the input data 901 is input to the pre-trained model, the feature extraction unit 902 converts the input data 901 into features. The feature extraction unit 902 is composed of a neural network and is composed of a configuration such as a multi-layer perceptron, a convolutional neural network, an attention mechanism, or a combination thereof. The feature prediction unit 903 and the region division unit 905 are also composed of neural networks and output predicted features 904 and region division results 906, respectively.

[0047] In the pre-training step S704, learning is performed using a loss function that recommends that the predicted features 904 have different features for each region divided by the region division unit 905. At this time, a learning method called contrastive learning may be combined to apply different noises to the input data so that the same features are obtained at points belonging to the same region division class on the original image and different features are obtained at points belonging to different region division classes on the original image.

[0048] In addition, the region division unit 905 learns region division using a loss function based on clustering or mutual information. When using clustering, clustering is applied to the output for each position of the input data 901 of the region division unit 905, and learning is performed so as to output the center of the cluster to which each output is assigned. In the case of learning based on mutual information, the assignment to the region division class is learned by optimizing the mutual information between the assigned region division class and the corresponding position of the input data 901 for the output for each position of the input data 901 of the region division unit 905.

[0049] Furthermore, the region division unit 905 may be configured to output a predetermined number of vectors. In this case, the similarity between the prediction feature amount 904 or the vector, which is the output of the feature amount extraction unit 902 or the output of the feature amount prediction unit 903, is calculated, and the assignment result to the vector with the highest similarity is treated as the region division result 906.

[0050] Note that a plurality of region division results 906 may be output using a plurality of region division units 905. In this case, for the evaluation of the fitness, a plurality of values may be output, or the value with the highest value may be selected from a plurality of evaluation values.

[0051] Also, the region division result 906 is used for the fitness evaluation performed by the pre-learning model fitness evaluation unit 203. Further, in the transfer learning performed by the transfer learning unit 206, only the feature amount extraction unit 902 among the pre-learning models may be used, or the feature amount prediction unit 903 or the region division unit 905 may be used.

[0052] In step S705, the pre-learning unit 603 determines the learning end condition defined in the pre-learning condition 601. If the end condition is satisfied, the process proceeds to step S706. If not, the process returns to step S704 and the parameter update of the pre-learning model is performed. Here, the learning end condition is, for example, whether or not the number of parameter updates has reached a predetermined number.

[0053] In step S706, the pre-learning unit 603 associates the pre-learning model obtained by pre-learning with the information related to the pre-learning data set acquired in the pre-learning condition 601 and step S702, stores it in the pre-learning model storage unit 202, and ends the pre-learning flow.

Example

[0054] The machine learning system according to Example 2 further includes a display / operation unit, unlike Example 1, in order for the operator to more efficiently perform pre-learning and transfer learning.

[0055] FIG. 10 is a diagram showing a configuration example of the machine learning system according to the second embodiment. As shown in FIG. 10, in this embodiment, an operator can set the pre-learning conditions 601 for the pre-learning unit 603 or set the pre-learning model evaluation target information 200 for the pre-learning model fitness evaluation unit 203 using the display / operation unit 1001. Further, the display / operation unit 1001 can also instruct the pre-learning unit 603, the pre-learning model fitness evaluation unit 203, and the transfer learning unit 206 to start pre-learning, fitness evaluation, and transfer learning, respectively.

[0056] FIG. 11 is an example of a pre-learning setting screen. As shown in FIG. 11, the pre-learning setting screen includes, for example, a pre-learning parameter setting unit 1101, a dataset filter input unit 1102, a dataset selection unit 1103, and a learning start unit 1104. On this screen, it is possible to set the pre-learning conditions 601 and perform operations related to the start of pre-learning.

[0057] Here, the pre-learning parameter setting unit 1101 sets the calculation procedure and hyperparameters used for pre-learning. For example, it sets the network configuration of the neural network, values related to the number of region classes divided by the region division unit 905, values related to the dimension of the feature amount predicted by the feature amount prediction unit 903, the batch size and learning schedule used for pre-learning, the number of parameter update times, and so on. The dataset filter input unit 1102 and the dataset selection unit 1103 select the dataset used for pre-learning using the tag information 803 stored in the pre-learning dataset storage unit 602. In the example of FIG. 11, the dataset including the tag information 803 corresponding to the domain and imaging conditions input by the dataset filter input unit 1102 is displayed on the dataset selection unit 1103. When the operator sets these parameters, selects the dataset used for pre-learning, and then operates the learning start unit 1104, the pre-learning unit 603 executes pre-learning.

[0058] FIG. 12 is an example of a setting screen for fitness evaluation and transfer learning. As shown in FIG. 12, this setting screen is composed of, for example, a transfer learning data selection section 1201, a transfer learning data set tag input section 1202, a pre-trained model display selection section 1203, a fitness evaluation start section 1204, a transfer learning condition setting section 1205, and a transfer learning start section 1206.

[0059] Here, the transfer learning data selection section 1201 is for the operator to select a data set of a downstream task to be used in transfer learning. The transfer learning data set tag input section 1202 is for the operator to input tag information such as the domain and imaging conditions of the data set selected by the transfer learning data selection section 1201. The pre-trained model display selection section 1203 displays, for each pre-trained model, information on the pre-training conditions, the data set used in pre-training, the fitness evaluated by the pre-trained model fitness evaluation section 203, and the evaluation index in the evaluation data set 404 when transfer learning is executed by the transfer learning section 206. Note that the pre-trained model display selection section 1203 may search for and extract a pre-trained model having tag information corresponding to the information input by the transfer learning data set tag input section 1202 from the pre-trained model storage section 202 and display it. The transfer learning condition setting section 1205 is for setting the learning conditions to be used in transfer learning.

[0060] When the operator selects one or more pre-trained models displayed in the pre-trained model display selection section 1203 and operates the fitness evaluation start section 1204, the pre-trained models selected in the pre-trained model display selection section 1203 are sent as pre-trained model evaluation target information 200 to the pre-trained model fitness evaluation section 203, and the pre-trained model fitness evaluation section 203 executes the fitness evaluation. By referring to the evaluation result of the fitness displayed in the pre-trained model display selection section 1203, the operator can efficiently select a model that is expected to have a good performance evaluation value during transfer learning.

[0061] On the one hand, when the operator operates the transfer learning start unit 1206, the pre-trained model selected by the pre-trained model display selection unit 1203 is sent to the transfer learning unit 206 as the model to be used for transfer learning. The transfer learning unit 206 executes transfer learning and displays the evaluation result in the evaluation dataset 404 on the pre-trained model display selection unit 1203.

[0062] FIG. 13 is an example of a result confirmation screen for fitness evaluation. The screen example in FIG. 13 can be transitioned from the screen example in FIG. 12 and is composed of, for example, a pre-trained model display unit 1301, a dataset display unit 1302, a fitness display unit 1303, an instructional information display unit 1304, and a pre-trained model region division result display unit 1305.

[0063] Here, the pre-trained model display unit 1301 displays the pre-trained model for which the operator wants to confirm the fitness. The dataset display unit 1302 displays the dataset of the downstream task for which the operator wants to confirm the fitness. The fitness display unit 1303 displays the fitness evaluated by the pre-trained model displayed in the pre-trained model display unit 1301 and the dataset displayed in the dataset display unit 1302. The instructional information display unit 1304 displays the instructional information corresponding to the input data included in the dataset displayed in the dataset display unit 1302. The pre-trained model region division result display unit 1305 displays the result of region division by the pre-trained model displayed in the pre-trained model display unit 1301 for the input data included in the dataset displayed in the dataset display unit 1302.

[0064] As described above, in the display and operation unit 1001 of this embodiment, the teaching information of the data set and the result of region division by the selected pre-trained model are displayed together, so that the operator can easily visually compare the two. Therefore, it is possible to select a more appropriate pre-trained model than when selecting a pre-trained model only based on the fitness. Also, in the example of the fitness evaluation result confirmation screen shown in FIG. 13, the region division results by a plurality of pre-trained models may be arranged and displayed simultaneously so that different pre-trained models can be compared. Further, the teaching information and the region division results regarding a plurality of input data may be arranged and displayed simultaneously.

Embodiment

[0065] Embodiment 3 applies the machine learning system of the aforementioned Embodiment 1 or the embodiment to image inspection.

[0066] In the case of image inspection, it is necessary to train the inspection neural network for each inspection process. This is because the inspection targets to be confirmed for each inspection process are different, so it is unrealistic to train an inspection neural network that covers all inspection targets, and it is to flexibly expand the inspection neural network when the number of inspection targets increases.

[0067] Therefore, in the case of image inspection, the downstream task of transfer learning is to be performed for each inspection process. When there are a plurality of inspection processes for the same product, the features that the neural network should acquire are often the same between inspection processes. Therefore, in such a case, it is effective to use the images obtained in other inspection processes for transfer learning. However, when the tasks to be learned for each inspection process are different, that is, when segmentation is performed in one process and a detection task is performed in another process, it becomes difficult to train a common neural network. In this embodiment, by using self-teaching learning, it is possible to integrate the input data of each inspection process and train a common neural network.

[0068] Therefore, in this embodiment, the data sets obtained for each product and each inspection process are stored in the pre-training data set storage unit 602. The input data 802 in this embodiment is an inspection image obtained for each inspection process, and the tag information 803 in this embodiment is the product type, the type of inspection process, the imaging conditions of the inspection image, and the like.

[0069] When an operator wants to newly train an inspection neural network, a data set for pre-training is constructed using tag information such as the product type, the inspection process type, or the imaging conditions, which are expected to have similar properties related to the inspection process to be trained, and pre-training is performed.

[0070] Thereafter, transfer learning is performed using the data set of the inspection process to be trained, a trained model 207 is constructed, and after confirming that the desired accuracy is achieved, it is used as a new inspection neural network, so that it is possible to efficiently configure an inspection neural network suitable for each of various inspection processes.

[0071] Note that the present invention is not limited to the above-described embodiments, and various modifications are included. For example, the above-described embodiments have been described in detail to facilitate understanding of the present invention, and are not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment can be replaced with the configuration of another embodiment, and the configuration of another embodiment can be added to the configuration of one embodiment. Also, for a part of the configuration of each embodiment, addition, deletion, or replacement with other configurations is possible.

Explanation of Reference Numerals

[0072] 101... Input image, 102... Teacher segmentation, 103... Teacher detection result, 104a, 104b... Region division result, 202... Pre-training model storage unit, 203... Pre-training model fitness evaluation unit, 204... Transfer learning data set storage unit, 205... Pre-training model selection unit, 206... Transfer learning unit

Claims

1. A pre-trained model acquisition unit that acquires at least one pre-trained model from a pre-trained model storage unit that stores a plurality of pre-trained models obtained by learning source tasks under respective conditions; A transfer learning dataset storage unit that stores a dataset related to a target task; A pre-trained model fitness evaluation unit that evaluates the fitness of each of the pre-trained models acquired by the pre-trained model acquisition unit with respect to the dataset related to the target task; A transfer learning unit that performs transfer learning using the selected pre-trained model and the dataset based on the evaluation result of the pre-trained model fitness evaluation unit, and outputs the learning result as a learned model; comprising: The pre-trained model includes a feature extraction unit that extracts features from input data, a region division unit that divides the input data into a plurality of regions so as to have the same meaning within the same region, and a feature prediction unit that predicts different features for each of the divided regions. A machine learning system.

2. In the machine learning system according to claim 1, The pre-trained model fitness evaluation unit is a machine learning system that evaluates the fitness between the result of region-dividing the input data of the dataset by the region division unit and the teaching information of the dataset.

3. In the machine learning system according to claim 2, The pre-trained model is one obtained by self-supervised learning of the source task.

4. In the machine learning system according to claim 2, The machine learning system further includes a display unit that displays the result of the region division and the teaching information of the dataset together.

5. In the machine learning system according to claim 2, The fitness is calculated by a function representing the relationship between the result of the region division and the teaching information of the dataset. The function is a machine learning system that utilizes at least one of probability density, degree of overlap of divided regions, and similarity of feature amounts for each divided region.

6. In the machine learning system according to Claim 1, The machine learning system further includes a pre-trained model selection unit that selects a pre-trained model to be used for transfer learning from among a plurality of pre-trained models stored in the pre-trained model storage unit.

7. In the machine learning system according to Claim 6, The machine learning system further includes a display unit that displays the fitness for each pre-trained model.

8. In the machine learning system according to Claim 1, The pre-trained model fitness evaluation unit evaluates the fitness by causing the pre-trained model to perform forward propagation a smaller number of times than during transfer learning on the input data of the dataset.

9. In the machine learning system according to Claim 1, The transfer learning dataset storage unit stores a dataset having an image obtained in an inspection process as input data, The pre-trained model storage unit stores a pre-trained model obtained by learning a task related to the inspection process.

Citation Information

Patent Citations

  • Machine learning apparatus

    JP2016191975A

  • Vehicle and server

    JP2021193280A

  • Information processing apparatus, information processing method, and program

    JP2023074114A