Learning method, learning device, and learning program
The learning method addresses inefficiencies and inaccuracies in conventional transfer learning by using synthetic and pseudo-sample generation and semi-supervised learning, enabling effective transfer learning across different model architectures and label spaces.
Patent Information
- Application Number
- JP2024513651
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-04-07
AI Technical Summary
Conventional transfer learning methods are inefficient and inaccurate when the architecture of the target model differs from the source model, and when the label spaces of the source and target tasks do not overlap.
A learning method that involves synthetic sample generation, pre-learning, pseudo-sample generation, and semi-supervised learning, allowing for efficient and accurate transfer learning even when the source and target model architectures are different and their label spaces do not overlap.
Enables efficient and accurate transfer learning by generating synthetic and pseudo-samples, allowing the method to perform well without direct access to the source dataset and across different model architectures and label spaces.
Smart Images

Figure 0007683817000005 
Figure 0007683817000006 
Figure 0007683817000007
Abstract
Description
Technical Field
[0001] The present invention relates to a learning method, a learning device, and a learning program.
Background Art
[0002] With the recent development of deep learning technology, the application of AI has advanced in various industrial fields. Also, for the learning of existing successful deep models (for example, deep neural networks (DNNs)), a large amount of learning data is required. For example, Non-Patent Document 1 describes that the accuracy of a task by a deep model is proportional to the log scale of the learning data size.
[0003] Also, transfer learning is known, which performs learning of a deep model with less learning data or calculation time by diverting a source data set (transfer source data set) different from the target data set or a learned model.
[0004] Transfer learning may require access to a source data set. On the other hand, due to privacy or license issues, it may not be possible to directly access a commercially created source data set.
[0005] Also, transfer learning includes a procedure called fine-tuning (see, for example, Non-Patent Document 2). Fine-tuning is a method of using the weight parameters of a pre-trained DNN for a known task (source task) as initial values during the learning of a new target task.
[0006] Fine-tuning is used as a standard method, for example, in the processing of images, voices, texts, etc. using DNNs.
Prior Art Documents
Non-Patent Documents
[0007]
Non-Patent Document 1
[0008] However, there is a problem in the conventional technology that transfer learning may not be efficiently and accurately performed.
[0009] For example, when the architecture of a target model (e.g., a deep model to be learned in transfer learning) is different from the architecture of a source model (e.g., a deep model from which transfer is made), transfer learning may not be performed by conventional methods. For example, the architecture of the target model is a task-specialized architecture described in Non-Patent Document 3. [Means for Solving the Problems]
[0010] In order to solve the above-described problems and achieve the object, a learning method is a learning method executed by a learning device, and includes a synthetic sample generation step of generating a synthetic sample by inputting a synthetic label imitating a label included in a first dataset into a first generation model learned by the first dataset, a pre-learning step of performing learning of the first classification model based on a label obtained by inputting the synthetic sample into the first classification model and the synthetic label, a pseudo-sample generation step of generating a pseudo-sample by inputting a label obtained by using the learned first classification model and a sample included in a second dataset into the first generation model, and a semi-supervised learning step of performing unsupervised learning of a second classification model using the pseudo-sample and performing supervised learning of the second classification model using the second dataset.
Advantages of the Invention
[0011] According to the present invention, transfer learning can be efficiently and accurately performed.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Embodiments for Carrying Out the Invention
[0013] Hereinafter, embodiments of the learning method, learning apparatus, and learning program according to the present application will be described in detail with reference to the drawings. Note that the present invention is not limited to the embodiments described below.
[0014] [Configuration of the First Embodiment] Figure 1 is a diagram showing a configuration example of a learning apparatus according to the first embodiment. The learning apparatus 10 executes transfer learning.
[0015] Here, the source means the transfer source in transfer learning. For example, the source task, dataset, and model are called the source task, source dataset, and source model, respectively.
[0016] On the other hand, the target means the transfer destination (the object of interest) in transfer learning. For example, the target task, dataset, and model are called the target task, target dataset, and target model, respectively.
[0017] The model in the embodiment is a DNN. Also, the architecture of the source model (source architecture) and the architecture of the target model (target architecture) are different from each other. The architecture is, for example, the number of layers or nodes constituting the DNN.
[0018] The task in the embodiment is a classification task. Both the source task and the target task execute a classification task for estimating scores for each label. On the other hand, in the source task and the target task, the output label spaces are different.
[0019] In the embodiment, it is assumed that the source data and the target data are data related to images. Note that the source data and the target data are not limited to images, and may be data related to, for example, voice and text.
[0020] The learning device 10 can perform transfer learning even when the source architecture and the target architecture are different.
[0021] Also, the learning device 10 can perform transfer learning even when the label space of the source task and the label space of the target task do not overlap.
[0022] Note that when the tasks and the model architectures are common between the transfer source and the transfer destination, the learning device 10 can naturally perform transfer learning.
[0023] The learning device 10 receives an input of a target data set and outputs information about the learned target model. Also, the learning device 10 acquires information about the source model at the timing of performing transfer learning or in advance.
[0024] As shown in FIG. 1, the learning device 10 includes an input / output unit 11, a storage unit 12, and a control unit 13.
[0025] The input / output unit 11 is an interface for inputting and outputting data. For example, the input / output unit 11 may be a communication interface such as a NIC (Network Interface Card) for performing data communication with other devices via a network. Also, the input / output unit 11 may be an interface for connecting input devices such as a mouse and a keyboard, and output devices such as a display.
[0026] The storage unit 12 is a storage device such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or an optical disk. Note that the storage unit 12 may be a semiconductor memory capable of rewriting data, such as a RAM (Random Access Memory), a flash memory, or an NVSRAM (Non Volatile Static Random Access Memory). The storage unit 12 stores an OS (Operating System) and various programs executed by the learning device 10. The storage unit 12 also stores model information 121.
[0027] The model information 121 is information such as parameters for constructing a model. The model information 121 includes information regarding each of a plurality of models. A part of the model information 121 is appropriately updated in the learning process. Also, the updated model information 121 may be output to other devices or the like via the input / output unit 11.
[0028] The control unit 13 controls the entire learning device 10. The control unit 13 is, for example, an electronic circuit such as a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or a GPU (Graphics Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0029] The control unit 13 also has an internal memory for storing programs and control data that define various processing procedures, and executes each process using the internal memory. The control unit 13 functions as various processing units when various programs operate.
[0030] For example, the control unit 13 includes a pre-training unit 131 and a semi-supervised learning unit 132. The pre-training unit 131 includes a generation unit 131a, a calculation unit 131b, and an update unit 131c. The semi-supervised learning unit 132 includes a generation unit 132a, a calculation unit 132b, and an update unit 132c.
[0031] Here, the transfer learning by the learning device 10 will be described. The transfer learning by the learning device 10 includes pre-training by the pre-training unit 131 and semi-supervised learning by the semi-supervised learning unit 132.
[0032] Note that the pre-training and semi-supervised learning by the learning device 10 may be called pseudo pre-training and pseudo semi-supervised learning because, in addition to real data, pseudo-generated data is used as learning data.
[0033] [Pre-training] The pre-training unit 131 performs learning of the classification model used for learning of the generation model. Assume that the generation model in the embodiment is a conditional adversarial generation model (GAN: Generative Adversarial Networks) (see, for example, Reference 1). Note that the generation model may be a variational autoencoder (VAE). Reference 1: Goodfellow, Ian, et al. "Generative adversarial nets." Advances in neural information processing systems 27, 2014. Reference 2: Kingma, Diederik P., and Max Welling. "Auto-encoding variational bayes." arXiv preprint arXiv:1312.6114 (2013).
[0034] FIG. 2 is a diagram for explaining pre-training. The source generation model G s is the source dataset D sAssume that it has been learned by []. Source dataset D s is data that combines sample x s and label y s and is represented as in equation (1). N s is the data size (number of samples) of source dataset D s .
[0035]
Number
[0036] For example, source dataset D s is, for example, a source dataset created commercially and is an example of a dataset that is difficult to access due to privacy or licensing issues.
[0037] Also, target dataset D t is a dataset that combines sample x t and label y t and is represented as in equation (2). N t is the data size (number of samples) of target dataset D t .
[0038]
Number
[0039] Target dataset D t is easier to access compared to source dataset D s , but the data size may not be sufficient to improve the classification model with high precision when used as learning data.
[0040] As shown in FIG. 2, the generation unit 131a of the pre-training unit 131 randomly generates a synthetic label y s by mimicking source dataset D s and inputs it into source generation model G s to obtain synthetic sample G s (ys ) is generated.
[0041] The calculation unit 131b of the pre-training unit 131 is G s (y s ) is input into the source classification model C´ s Based on the classification result obtained, the loss of the source task is calculated as in Equation (3).
[0042] [Semi-supervised learning]
[0043] The update unit 131c of the pre-training unit 131 updates the parameters of the source classification model C´ s so that the loss in Equation (3) is minimized.
[0044] [Semi-supervised learning] FIG. 3 is a diagram for explaining a method of generating samples. The generation unit 132a of the semi-supervised learning unit 132 first inputs a sample x s of the target data set D t into the source classification model C t and generates a pseudo label y s←t .
[0045] Here, the source classification model C s is obtained by replacing the final layer of the source classification model C´ s learned by the pre-training unit 131 with a specific function. For example, the semi-supervised learning unit 132 replaces the final layer of the source classification model C s which is the softmax function with a temperature-softmax function or a function (arg max function) that returns the label with the maximum score. As a result, the tendency of the label (the pseudo label y s←t ) that controls data generation is changed, and finally the nature of the data (the pseudo source data set D s ) generated from the source generation model G s←t ) can be changed. Note that the source classification model C s is the source classification model C´s It may be the same as that.
[0046] In the example of FIG. 3, in the target data set D t the sample x t is combined with the label y t = "Hummer". Also, the target sample x t and the source sample x s are in the same space, such as the image data itself, or features extracted from the image, etc. For example, the sample x t and the sample x s are represented by vectors of the same number of dimensions.
[0047] On the other hand, the target label y t and the source label y s may be different. For example, the label y t and the label y s are vectors in which each element corresponds to the score of the object to be classified, and are represented by vectors with different numbers of dimensions from each other.
[0048] The source classification model C s generates pseudo-labels such as y s←t = ("Jeep": 0.41, "Limousine": 0.11, "MovingVan": 0.06,...).
[0049] Furthermore, the generation unit 132a inputs the pseudo-label y s←t into the source generation model G s to generate a pseudo-source data set D s←t
[0050] In this way, the semi-supervised learning unit 132 connects the source classification model C s and the source generation model G s to generate a pseudo-source data set D s←t
[0051] Subsequently, the semi-supervised learning unit 132 uses the target data set D t and the pseudo-source dataset D s←t are used to train the target classification model C t .
[0052] The target dataset D t contains the sample x t and the label y t . Therefore, the semi-supervised learning unit 132 uses the target dataset D t as the training data for supervised learning. On the other hand, since the pseudo-source dataset D s←t does not contain labels, the semi-supervised learning unit 132 uses the pseudo-source dataset D s←t as the training data for unsupervised learning.
[0053] Thereby, the semi-supervised learning unit 132 realizes semi-supervised learning as shown in FIG. 4. FIG. 4 is a diagram for explaining semi-supervised learning.
[0054] As shown in FIG. 4, the calculation unit 132b of the semi-supervised learning unit 132 calculates the supervised learning loss L t based on the classification result obtained by inputting the target dataset D t into the target classification model C sup .
[0055] Also, the calculation unit 132b calculates the unsupervised learning loss L s←t based on the classification result obtained by inputting the pseudo-source dataset D t into the target classification model C unsup .
[0056] Then, the calculation unit 132b calculates the total loss as shown in equation (4) from the loss L sup and the loss L unsup .
[0057] [Equation]
[0058] The update unit 132c of the semi-supervised learning unit 132 updates the parameters of the target classification model C so that the loss in Equation (4) is minimized. t
[0059] [Processing of the First Embodiment] Using FIG. 5, the process of pre-training will be described. FIG. 5 is a flowchart showing the process of pre-training.
[0060] As shown in FIG. 5, the learning device 10 first randomly generates a label y that mimics the label of the source dataset (step S101). Next, the label y is input into the source generation model G, and G(y) is generated (step S102). Here, the source generation model G is trained using the source dataset. Also, G(y) is, for example, an image. s s s s s s s s
[0061] Furthermore, the learning device 10 inputs G(y) into the source classification model C', and calculates the output result (classification result) of the source task, C'(G(y)) (step S103). s s s s s s
[0062] Subsequently, the learning device 10 calculates the loss L of the source task, L(C'(G(y)),y) (step S104). Then, the learning device 10 updates the parameters of the source classification model C' by the error backpropagation method of the loss L (step S105). sup s s s s sup s
[0063] Here, when the number of repetitions (number of learning steps) from step S101 to S105 is less than the predetermined maximum number of learning steps (step S106, True), the learning device 10 returns to step S101 and repeats the process.
[0064] On the other hand, when the number of learning steps is greater than or equal to the maximum number of learning steps (step S106, False), the learning device 10 ends the process.
[0065] Using FIG. 6, the flow of the process for generating a sample will be described. FIG. 6 is a flowchart showing the flow of the method for generating a sample.
[0066] As shown in FIG. 6, the learning device 10 first randomly selects a sample x t from the target dataset D t (step S201).
[0067] Next, the learning device 10 inputs the sample x t into the source classification model C s to generate a pseudo-label C s (x t ) (step S202). For example, the source classification model C s is a model in which the final layer of the source classification model C´ s is replaced with a specific function.
[0068] Subsequently, the learning device 10 inputs the pseudo-label C s (x t ) into the source generation model G s to generate a pseudo-sample x s←t = G s (C s (x t )) (step S203).
[0069] Then, the learning device 10 adds the pseudo-sample x s←t to the pseudo-source dataset D s←t (step S204). Here, the pseudo-sample x s←tNo label representing the correct answer corresponding thereto is generated.
[0070] Here, when the number of repetitions (number of generation steps) from step S201 to S204 is less than the predetermined maximum number of generation steps (step S205, True), the learning device 10 returns to step S201 and repeats the process.
[0071] On the other hand, when the number of generation steps is greater than or equal to the maximum number of generation steps (step S205, False), the learning device 10 ends the process.
[0072] The flow of semi-supervised learning will be described with reference to FIG. 7. FIG. 7 is a flowchart showing the flow of semi-supervised learning.
[0073] As shown in FIG. 7, the learning device 10 first obtains a pair (x t , y t ) of a sample x t and a label y t from the target data set D t (step S301). Also, the learning device 10 obtains a sample x s←t from the pseudo-source data set D s←t (step S302).
[0074] The learning device 10 calculates the loss L sup (C t (x t ), y t ) of supervised learning (step S303). Also, the learning device 10 calculates the loss L unsup (C t (x s←t )) of unsupervised learning (step S304).
[0075] Furthermore, the learning device 10 calculates the total loss L total = L sup (C t (x t ), y t ) + L unsup (C t (x s←tCalculate (Step S305). Then, the learning device 10 calculates the total loss L total By the error backpropagation method of, the target classification model C t Update the parameters of (Step S306).
[0076] Here, when the number of repetitions (number of learning steps) from Step S301 to S306 is less than the predetermined maximum number of learning steps (Step S307, True), the learning device 10 returns to Step S301 and repeats the process.
[0077] On the other hand, when the number of learning steps is greater than or equal to the maximum number of learning steps (Step S307, False), the learning device 10 ends the process.
[0078] [Effect of the First Embodiment] As described so far, the learning method executed by the learning device 10 includes a synthetic sample generation step, a pre-learning step, a pseudo-sample generation step, and a semi-supervised learning step.
[0079] In the synthetic sample generation step, the learning device 10 generates a synthetic label (e.g., y s ) that imitates the labels included in the first dataset (e.g., the source dataset D s ) and inputs it to the first generation model (e.g., the source generation model G s ) that has been learned by the first dataset to generate a synthetic sample (e.g., G s (y s ).
[0080] In the pre-learning step, the learning device 10 learns the first classification model based on the label obtained by inputting the synthetic sample to the first classification model (e.g., C´ s ) and the synthetic label.
[0081] In the pseudo-sample generation step, the learning device 10 uses the learned first classification model and the samples (e.g., x t ) included in the second dataset (e.g., the target dataset Dt ) and the label obtained using (for example, C s (x t )) are input into the first generation model to generate pseudo-samples (for example, x s←t ).
[0082] In the semi-supervised learning process, the learning device 10 performs unsupervised learning of the second classification model (for example, the target classification model C t ) using the pseudo-samples, and performs supervised learning of the second classification model using the second dataset.
[0083] Note that the synthetic sample generation process and the pre-training process are executed by the pre-training unit 131. Also, the pseudo-sample generation process and the semi-supervised learning process are executed by the semi-supervised learning unit 132.
[0084] Thereby, even when the architectures of the source classification model C' s or the source classification model C s and the target classification model C t are different, transfer learning can be realized. The knowledge of the source dataset is reflected in the pseudo-source dataset. As a result, according to the embodiment, transfer learning can be efficiently and accurately implemented.
[0085] Also, in the synthetic sample generation process, the learning device 10 inputs a synthetic label randomly generated based on the labels included in the first dataset into the first generation model to generate synthetic samples. Thereby, transfer learning can be realized without accessing the source dataset at all.
[0086] The pseudo-sample generation process inputs the labels obtained by inputting the samples included in the second dataset into a model in which the final layer of the first classification model, which is a DNN, is replaced with a function for knowledge distillation (for example, a temperature-softmax function or an arg max function), into the first generation model to generate pseudo-samples. As a result, the information of the labels used for generating the pseudo-samples can be reduced, and the amount of calculation can be suppressed.
[0087] [Experiment] An experiment actually conducted by implementing the above-described embodiment will be described. In the experiment, transfer learning was performed using the above-described embodiment and other methods, and the results were compared.
[0088] The settings of the experiment are as follows. · Dataset (image): Source dataset: ImageNet Target dataset: (9 types) Caltech-256-60, CUB-200-2011, DTD, FGVC-Aircraft, Indoor67, OxfordFlower, OxfordPets, StanfordCars, StanfordDogs ※ Among the target datasets, 90% was used as the learning and dataset, and 10% was used as the validation dataset. · Neural network architecture: Source classification model C s : ResNet-50 Target classification model C t : (6 types) ResNet-50, WRN-50-2, MNASNet1.0, MobileNetV3-L, EfficientNet-B0, EfficientNet-B5
[0089] The procedure of the experiment is as follows. (1) Perform pre-training for 100 epochs using the synthetic dataset D' consisting of synthetic labels and synthetic samples. s (2) Target dataset D tand the pseudo-source dataset D s←t was used to perform semi-supervised learning for 300 epochs. (3) Adopt the accuracy of the model with the highest accuracy (Top-1 accuracy) on the validation dataset on the test dataset. (4) Perform the above five times and report the average value and standard deviation.
[0090] The experimental patterns are as follows. Scratch: Train the target classification model only using the target dataset D t Logit Matching: Application model of the existing method (transfer knowledge by distillation). Soft Target: Application model of the existing method (transfer knowledge by distillation). PP: Only apply the pre-training of the embodiment. P-UDA: P-SSL model using the semi-supervised learning method UDA. PP+P-UDA: Combination of PP and P-UDA. FT: Fine-tuning using the real source dataset D s (difficult to implement in reality). R-UDA: Semi-supervised learning using the real source dataset D s (difficult to implement in reality). Note that the transfer of knowledge by distillation corresponds to replacing C´ s with C s .
[0091] The experimental results are shown below. Figures 8, 9, and 10 are figures showing the results of the experiment.
[0092] Figure 8 shows the results of each experimental pattern for each architecture of the target classification model. As shown in Figure 9, high accuracy is obtained by applying the embodiment (such as PP+P-UDA) in all architectures. Note that the target dataset in this case is StanfordCars.
[0093] Figure 9 shows the results of each experimental pattern for each target dataset. As shown in Figure 9, high accuracy is obtained by applying the embodiments (such as PP+P-UDA, etc.) in all target datasets. Note that the architecture of the target classification model in this case is ResNet-50.
[0094] Figure 10 shows the comparison results between the experimental patterns and the embodiments (PP, P-UDA) when it is assumed that the source dataset can be accessed, such as FT and R-UDA. As shown in Figure 10, the embodiments obtain performance equivalent to that of FT and R-UDA despite the inability to access the source dataset.
[0095] In Figure 9, P-UDA outperforms R-UDA. This is presumably because the distance (FID (Frechet Inception Distance)) between the distributions between the pseudo-source dataset D s is closer to D s←t than the source dataset D t and contains more useful information.
[0096] [System configuration, etc.] Also, each component of each illustrated device is conceptually functional and does not necessarily have to be physically configured as shown in the figure. That is, the specific forms of dispersion and integration of each device are not limited to those shown in the figure, and all or part of them can be functionally or physically dispersed or integrated in any unit according to various loads, usage situations, etc. Furthermore, each processing function performed by each device can be realized in whole or in any part by a CPU (Central Processing Unit) and a program analyzed and executed by the CPU, or can be realized as hardware by wired logic. Note that the program may be executed not only by the CPU but also by other processors such as GPUs.
[0097] In addition, among the processes described in the present embodiment, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be arbitrarily changed unless otherwise specified.
[0098] [Program] As one embodiment, the learning device 10 can be implemented by installing a learning program that executes the above-described processes as package software or online software on a desired computer. For example, by causing the information processing device to execute the above-described learning program, the information processing device can function as the learning device 10. The information processing device mentioned here includes desktop or notebook personal computers. In addition, other information processing devices include mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone System), and further slate terminals such as PDAs (Personal Digital Assistants) are included in this category.
[0099] In addition, the learning device 10 can also be implemented as a server device that uses the terminal device used by the user as a client and provides services related to the above-described processes to the client. For example, the server device is implemented as a server device that provides a learning service that takes a target data set as an input and outputs information on a target classification model that has been transferred and learned. In this case, the server device may be implemented as a Web server, or may be implemented as a cloud that provides services related to the above-described processes through outsourcing.
[0100] FIG. 11 is a diagram showing an example of a computer that executes a learning program. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0101] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM (Random Access Memory) 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System), for example. The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100, for example. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to a display 1130, for example.
[0102] The hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the program that defines each process of the learning device 10 is implemented as a program module 1093 in which executable code by a computer is described. The program module 1093 is stored in the hard disk drive 1090, for example. For example, a program module 1093 for executing the same process as the functional configuration in the learning device 10 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
[0103] Also, the setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads out the program module 1093 and the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as needed, and executes the processing of the above-described embodiment.
[0104] Note that the program module 1093 and the program data 1094 are not limited to being stored in the hard disk drive 1090, and may be stored, for example, in a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and the program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). Then, the program module 1093 and the program data 1094 may be read by the CPU 1020 from another computer via the network interface 1070.
Explanation of Reference Numerals
[0105] 10 Learning device 11 Input / output unit 12 Storage unit 13 Control unit 121 Model information 131 Pre-training unit 131a Generation unit 131b Calculation unit 131c Update unit 132 Semi-supervised learning unit 132a Generation unit 132b Calculation unit 132c Update unit
Claims
1. A learning method executed by a learning device, comprising: a synthetic sample generation step of inputting a synthetic label imitating a label included in a first dataset into a first generation model learned by the first dataset to generate a synthetic sample; a pre-learning step of learning the first classification model based on the label obtained by inputting the synthetic sample into the first classification model and the synthetic label; a pseudo-sample generation step of inputting the label obtained by using the learned first classification model and a sample included in a second dataset into the first generation model to generate a pseudo-sample; a semi-supervised learning step of performing unsupervised learning of a second classification model using the pseudo-sample and performing supervised learning of the second classification model using the second dataset; A learning method characterized by including the above.
2. The synthetic sample generation step according to claim 1, wherein the synthetic sample is generated by inputting the synthetic label randomly generated based on the label included in the first dataset into the first generation model.
3. The pseudo-sample generation step according to claim 1, wherein the label obtained by inputting the sample included in the second dataset into a model in which the final layer of the first classification model, which is a DNN, is replaced with a function for knowledge distillation is input into the first generation model to generate a pseudo-sample.
4. The pseudo-sample generation step according to claim 3, wherein the label obtained by inputting the sample included in the second dataset into a model in which the final layer of the first classification model, which is a DNN, is replaced with a temperature-softmax function or an arg max function is input into the first generation model to generate a pseudo-sample.
5. A pre-learning unit that generates a synthetic sample by inputting a synthetic label imitating a label included in a first dataset into a first generation model learned by the first dataset, and learns the first classification model based on the label obtained by inputting the synthetic sample into the first classification model and the synthetic label; Using the learned first classification model and the labels obtained using the samples included in the second dataset, input the labels into the first generation model to generate pseudo-samples, perform unsupervised learning of the second classification model using the pseudo-samples, and perform semi-supervised learning including performing supervised learning of the second classification model using the second dataset. A learning device characterized by having the above. **Claim 6** On a computer, A synthetic sample generation step of inputting a synthetic label imitating the labels included in the first dataset into the first generation model learned by the first dataset to generate a synthetic sample; A pre-learning step of learning the first classification model based on the labels obtained by inputting the synthetic sample into the first classification model and the synthetic label; A pseudo-sample generation step of inputting the labels obtained using the learned first classification model and the samples included in the second dataset into the first generation model to generate pseudo-samples; An unsupervised learning step of performing unsupervised learning of the second classification model using the pseudo-samples and a semi-supervised learning step of performing supervised learning of the second classification model using the second dataset; A learning program characterized by causing the above to be executed.