Fine-tuning device and fine-tuning method
The fine tuning device addresses the issue of over-learning and performance degradation in deep learning models by initializing and dynamically updating a prior distribution on a feature space, ensuring feature norm and entropy are maintained, thus enhancing generalization performance.
Patent Information
- Application Number
- PCT/JP2023/039737
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-08
AI Technical Summary
Fine tuning of deep learning models with limited training data can lead to over-learning and performance degradation due to the loss of source knowledge and feature localization.
A fine tuning device that initializes a prior distribution on a feature space defined for each class using statistical values of features, and dynamically updates this distribution during optimization to prevent feature localization and maintain feature entropy.
The solution effectively suppresses over-learning and prevents performance degradation by maintaining feature norm and entropy, thereby improving the generalization performance of fine tuned models.
Smart Images

Figure JP2023039737_08052025_PF_FP_ABST
Abstract
Description
Fine tuning device and fine tuning method
[0001] The present invention relates to a fine tuning device and a fine tuning method.
[0002] In recent years, advances in deep learning technology have led to the widespread application of AI in various industrial fields. However, deep learning models require large amounts of training data. In response to this, a technique called fine tuning (FT) has been proposed that trains high-performance models using a small amount of training data (see Non-Patent Documents 1 to 8). In FT, the weight parameters of a deep learning neural network (DNN) model that has been pre-trained on a known source task are used as initial values when training a desired target task.
[0003] Yosinski, Jason, et al. “How transferable are features in deep neural networks?”, Advances in neural information processing systems 27 (2014), [online], 2014, [Retrieved October 17, 2023], Internet<URL:https: / / papers.nips.cc / paper / 2014 / file / 375c71349b295fbe2dcdca9206f20a06-Paper.pdf> You, Kaichao, et al. “Co-Tuning for Transfer Learning”, Advances in Neural Information Processing Systems 33 (2020), pp.17236-17246, [online], 2020, [Retrieved October 17, 2023], Internet<URL:https: / / papers.nips.cc / paper / 2020 / file / c8067ad1937f728f51288b3eb986afaa-Paper.pdf> Liu, Ziquan, et al. “Improved Fine-Tuning by Better Leveraging Pre-Training Data”, Advances in Neural Information Processing Systems 35 (2022), pp.32568-32581, [online], 2022, [Retrieved October 17, 2023], Internet <URL:https: / / proceedings.neurips.cc / paper_files / paper / 2022 / file / d1c88f9790765146ec8fb5d02e5653a0-Paper-Conference.pdf> Radford, Alec, et al. “Learning Transferable Visual Models From Natural Language Supervision”, International conference on machine learning. PMLR, 2021.[online], 2021, [Retrieved October 17, 2023], Internet<URL:http: / / proceedings.mlr.press / v139 / radford21a / radford21a.pdf> Zhong, Yang, and Atsuto Maki. “Regularizing CNN Transfer Learning with Randomized Regression”, Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020. [online], 2020, [Retrieved October 17, 2023], Internet <URL:https: / / openaccess.thecvf.com / content_CVPR_2020 / papers / Zhong_Regularizing_CNN_Transfer_Learning_With_Randomised_Regression_CVPR_2020_paper.pdf> Xuhong, L. I., Yves Grandvalet, and Franck Davoine. “Explicit Inductive Bias for Transfer Learning with Convolutional Networks”, International Conference on Machine Learning. PMLR, 2018. [online], 2013. [Retrieved October 17, 2023], Internet<URL:http: / / proceedings.mlr.press / v80 / li18a / li18a.pdf> Li, Xingjian, et al. "DELTA: DEEP LEARNING TRANSFER USING FEATURE MAP WITH ATTENTION FOR CONVOLUTIONAL NETWORKS", arXiv preprint arXiv:1901.09229 (2019). [online], 2019, [Retrieved October 17, 2023], Internet <URL:https: / / arxiv.org / pdf / 1901.09229.pdf>Chen, Xinyang, et al. "Catastrophic Forgetting Meets Negative Transfer: Batch Spectral Shrinkage for Safe Transfer Learning", Advances in Neural Information Processing Systems 32 (2019). [online], 2019, [Retrieved October 17, 2023], Internet <URL:https: / / papers.nips.cc / paper_files / paper / 2019 / file / c6bff625bdb0393992c9d4db0c6bbe45-Paper.pdf> .
[0004] However, conventional techniques can sometimes cause overfitting during fine tuning. In other words, when there is little training data, there is a risk of overfitting, resulting in poor generalization performance even if accurate inference can be achieved from training examples.
[0005] The present invention has been made in view of the above, and has as its object to prevent performance degradation by suppressing overlearning in fine tuning.
[0006] In order to solve the above-mentioned problems and achieve the object, the fine-tuning device of the present invention is characterized by having an initialization unit that initializes a prior distribution in a feature space defined for each class of data to be processed using predetermined statistics of the features of the class of the data to be processed; an optimization unit that uses the prior distribution to perform an optimization process of a model that outputs the features and classes of the data from input data to be processed; and an update unit that updates the prior distribution using the features output by the model from the input data to be processed during the optimization process.
[0007] According to the present invention, it is possible to prevent performance degradation by suppressing overlearning in fine tuning.
[0008] FIG. 1 is a diagram for explaining an overview of the fine tuning device of this embodiment. FIG. 2 is a diagram for explaining an overview of the fine tuning device of this embodiment. FIG. 3 is a diagram for explaining an overview of the fine tuning device of this embodiment. FIG. 4 is a diagram for explaining an overview of the fine tuning device of this embodiment. FIG. 5 is a schematic diagram illustrating an example of the general configuration of the fine tuning device of this embodiment. FIG. 6 is a flowchart showing a fine tuning processing procedure. FIG. 7 is a diagram for explaining an example. FIG. 8 is a diagram for explaining an example. FIG. 9 is a diagram for explaining an example. FIG. 10 is a diagram showing an example of a computer that executes a fine tuning program.
[0009] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are designated by the same reference numerals.
[0010] 1 to 4 are diagrams for explaining an overview of the fine-tuning device of this embodiment. First, as illustrated in FIG. 1, in FT, weight parameters of a deep neural network (DNN) model that has been pre-trained on a known source task are used as initial values when training a desired target task. Here, the source data set D s , target dataset D t is expressed by the following equation (1). θ s is the source classifier, f θ t is the target classifier.
[0011]
[0012] On the other hand, in FT, if the target dataset is small, overfitting is likely to occur, for example, because the source knowledge is lost when the target data is optimized. Therefore, as shown in Figure 2, random feature regularization is used as an FT method that does not use source information. In random feature regularization, the deep model f θ Feature extractor gφ A regularization term shown in the following equation (2) is added to the loss function so that the output of p(z) approaches noise z generated in the feature space from the prior distribution p(z). Therefore, accuracy is improved by adding perturbation to the gradient of the loss function shown in the following equation (3).
[0013]
[0014] FIG. 3 illustrates an example of a feature space. In FIG. 3, classes are distinguished by the difference in the color intensity of the points. As illustrated in FIG. 3, random feature regularization minimizes the difference between the prior distribution and the feature vector, resulting in localized features. This results in a reduction in the feature norm, as illustrated in FIG. 4(a). This may result in a decrease in the norm of the gradient of the classification loss, as shown in the following equation (4), which may hinder learning.
[0015]
[0016] Furthermore, as shown in Fig. 4B, the feature entropy H(z) is reduced, which may reduce the mutual information between the feature and the label, as shown in the following equation (5), and thus may hinder the performance of the model.
[0017]
[0018] Therefore, the fine-tuning device of this embodiment dynamically updates the prior distribution that generates noise in the feature space in random feature regularization according to the learning situation. Specifically, since the use of only a single prior distribution is a cause of feature localization, first, a prior distribution is defined for each class, and then widely distributed in the feature space to expand the entropy H(z).
[0019] In addition, the system calculates statistics such as the mean and variance of features for each class in the pre-trained model before FT and sets them as the initial values of the prior distribution defined for each class, thereby eliminating the need to search for the prior distribution parameters for each class.
[0020] Since the initial values are not necessarily optimal, the prior distribution for each class is dynamically updated according to the learning situation. For example, the parameters of the prior distribution for each class are updated according to the current features output by the model during learning. For example, if the prior distribution is a normal distribution, the parameters of the prior distribution are updated to minimize the distance between the mean vector of the features of the target data and the mean parameters of the prior distribution for each class, i.e., the squared error.
[0021] Alternatively, the feature sets of each class are separated between classes by updating the parameters of the prior distribution so as to maximize the distance (minimize the similarity) between the classes and expand the entropy H(z).
[0022] In this way, the fine tuning device prevents the feature norm from shrinking and increases the feature entropy as learning progresses, thereby preventing over-learning during fine tuning and preventing performance degradation.
[0023] [Configuration of the Fine-Tuning Device] Fig. 5 is a schematic diagram illustrating the general configuration of the fine-tuning device of this embodiment. As illustrated in Fig. 5, the fine-tuning device 10 of this embodiment is realized by a general-purpose computer such as a personal computer, and includes an input unit 11, an output unit 12, a communication control unit 13, a storage unit 14, and a control unit 15.
[0024] The input unit 11 is realized using input devices such as a keyboard and a mouse, and inputs various instruction information such as a command to start processing to the control unit 15 in response to input operations by an operator. The output unit 12 is realized by a display device such as a liquid crystal display, a printing device such as a printer, etc. For example, the output unit 12 displays the results of the fine tuning process described below.
[0025] The communication control unit 13 is realized by a NIC (Network Interface Card) or the like, and controls communication between the control unit 15 and external devices via telecommunication lines such as a LAN (Local Area Network) or the Internet. For example, the communication control unit 13 controls communication between the control unit 15 and a management device or the like that manages various types of information.
[0026] The storage unit 14 is realized by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 14 stores in advance the processing program that operates the fine-tuning device 10, data used during execution of the processing program, and the like, or temporarily stores the data each time processing is performed. In particular, in this embodiment, the storage unit 14 stores, for example, parameters of a model for the fine-tuning process and parameters of a prior distribution. The storage unit 14 may be configured to communicate with the control unit 15 via the communication control unit 13.
[0027] The control unit 15 is realized using a CPU (Central Processing Unit), NP (Network Processor), FPGA (Field Programmable Gate Array), etc., and executes a processing program stored in memory. As a result, the control unit 15 functions as an acquisition unit 15a, an initialization unit 15b, an optimization unit 15c, and an update unit 15d, as exemplified in FIG. 5, to perform fine-tuning processing. Note that these functional units may be implemented individually or in part in different hardware. The control unit 15 may also include other functional units.
[0028] The acquisition unit 15a acquires a target dataset D to be processed from a management device that manages various types of data, etc., via the input unit 11 or the communication control unit 13. The target dataset D is made up of sets of N input samples and labels. The acquisition unit 15a may store the acquired target dataset D in the storage unit 14.
[0029] Here, we consider a deep learning model f with parameters θ trained on the source dataset. θ is defined as the following equation (6).
[0030]
[0031] Furthermore, although the following description exemplifies a classification task, the present invention is not limited to this. For example, the processing of this embodiment can be applied to object detection, segmentation, and other tasks in which a categorical variable y exists.
[0032] The initialization unit 15b initializes a prior distribution in a feature space defined for each class of data to be processed using predetermined statistics of the features for each class of the data to be processed. Specifically, while noise in the feature space has conventionally been generated from a single distribution, the initialization unit 15b first defines the prior distribution of noise z in the feature interval of the data to be processed using a normal distribution or the like, as shown in the following equation (7). The normal distribution is defined with a mean and a variance as parameters.
[0033]
[0034] Note that by random feature regularization, the feature extractor g φ Feature generation is performed so that each class approaches a different distribution. Therefore, the entropy H(z) is expected to expand.
[0035] Next, the initialization unit 15b initializes the class y i The statistical quantities such as the mean and variance of the features of the training data, which are pre-calculated before training, are set as the initial values of the parameters of the prior distribution. Here, the parameters of the prior distribution are, for example, the mean μ i ・Dispersion σ i 2 In that case, class y i The prior distribution parameter (μ i , σ i 2 ) is the initial value of the feature extractor g φ The initial parameter φ 0 are used to calculate the following equations (8) and (9):
[0036]
[0037] The optimization unit 15c uses a prior distribution defined for each class to generate a model f that outputs the features and classes of the input data to be processed. θSpecifically, the optimization unit 15c performs the optimization process of the supervised learning loss L sup (θ) and the random feature regularization term L rand (φ), the total loss L shown in the following equation (10) is calculated.
[0038]
[0039] Then, the optimization unit 15c uses the backpropagation algorithm to update the parameter θ=[φ, Ψ] of the total loss L. Then, the optimization unit 15c optimizes the model by repeating the same process after the updating unit 15d dynamically updates the parameters of the prior distribution, which will be described below.
[0040] During the optimization process, the update unit 15d updates the prior distribution defined for each class using features output by the model from the input data to be processed.
[0041] Here, the prior distribution is updated during optimization because the initial value of the prior distribution is not necessarily the optimal distribution. For example, the updating unit 15d updates the mean μ of each class as a parameter of the prior distribution. In this case, the updating unit 15d updates f θ At each learning step, the mean μ of each class is calculated by the gradient method. i} (i=1 to Nc). That is, the updating unit 15d updates the random feature regularization term L ada The parameter μ of the prior distribution is updated so as to minimize (μ).
[0042] Specifically, the update unit 15d updates the parameters of the prior distribution so as to minimize the distance between the statistics of the features of the data to be processed and the statistics of the prior distribution defined for each class.
[0043] That is, the update unit 15d calculates the mean vector of the features of the training data for each class and the mean μ of the prior distribution as shown in the following equation (11). i Distance from intra Objective function L including adaThe parameter μ of the prior distribution is updated to minimize (μ). Here, the mean vector of the features of class i of the training data is the moving average at each step. The distance function D is, for example, the cosine distance.
[0044] Alternatively, the updating unit 15d updates the parameters of the prior distribution so as to maximize the distance between the statistics of the classes. That is, the updating unit 15d updates the parameters of the prior distribution so as to maximize the distance between the statistics of the classes. ada (μ) is the average distance between classes l inter (μ) and updates the parameter μ of the prior distribution using backpropagation.
[0045]
[0046] In this case, l intra is minimized, and l inter (μ) is maximized. This adjusts z in accordance with the learning, preventing the norm of z from shrinking and enabling the expansion of H(z).
[0047] The update unit 15d calculates the distance l between the statistics of the features of the data to be processed and the statistics of the prior distribution defined for each class. intra The process of minimizing the distance l between classes inter Alternatively, both the process of maximizing the value of the sigma and the process of maximizing the value of the sigma may be performed.
[0048] The update unit 15d also uses the variance σ as a parameter of the prior distribution. 2 In this case, the mean μ and variance σ 2 Both of these may be updated, or only one of them may be updated.
[0049] [Fine Tuning Process] Next, the fine tuning process performed by the fine tuning device 10 according to this embodiment will be described with reference to Fig. 6. Fig. 6 is a flowchart showing the procedure of the fine tuning process. The flowchart in Fig. 6 starts, for example, when the user performs an operation input to instruct the start of the process.
[0050] First, the acquisition unit 15a acquires a target data set D to be processed. Then, the initialization unit 15b initializes a prior distribution in a feature space defined for each class using predetermined statistics of the features of each class of the data to be processed. Specifically, when the prior distribution is a normal distribution, the parameter mean μ i ・Dispersion σ i 2 class y i The mean and variance of the learning data are calculated before each learning (step S1).
[0051] Next, the optimization unit 15c and the update unit 15d repeat the processes of steps S3 to S10 described below until a predetermined maximum number of learning steps is reached (step S2, False).
[0052] That is, the optimization unit 15c acquires (x, y) from the target dataset D (step S3). The optimization unit 15c also acquires noise z in the feature space generated from a prior distribution defined for each class (step S4).
[0053] Then, the optimization unit 15c generates a model f that outputs the features and classes of the input data to be processed. θ Specifically, the optimization unit 15c optimizes the supervised learning loss L sup (θ) and the random feature regularization term L rand (φ), the total loss L shown in the above equation (10) is calculated (steps S5 to S7).
[0054] Then, the optimization unit 15c updates the parameter θ=[φ, Ψ] of the total loss L by backpropagation (step S8).
[0055] The update unit 15d θ For each learning step, a dynamic loss function L is calculated using the current features of the training data for each class. ada For example, the update unit 15d calculates the mean vector of the features of the training data and the mean μ of the prior distribution for each class. i Distance from intra Loss function L including ada(μ) is calculated. ada (μ) is the average distance between classes, l, as shown in the above formula (11). inter (μ) may be included.
[0056] Then, the update unit 15d calculates the loss function L by backpropagation. ada The parameter μ of the prior distribution is updated so as to minimize (μ) (step S10). intra is minimized, and l inter (μ) is maximized.
[0057] If the number of learning steps has not reached the predetermined maximum number of learning steps (Step S2, True), the optimization unit 15c and the update unit 15d repeat the processes of Steps S3 to S10. If the number of learning steps has reached the predetermined maximum number of learning steps (Step S2, False), the series of fine-tuning processes ends.
[0058] [Effects] As described above, in the fine-tuning device 10 of this embodiment, the initialization unit 15b initializes a prior distribution in a feature space defined for each class of data to be processed using predetermined statistics of the features for each class of the data to be processed. The optimization unit 15c executes an optimization process for a model that outputs the features and classes of input data to be processed using the prior distribution defined for each class. The update unit 15d updates the prior distribution defined for each class during the optimization process using the features output by the model from the input data to be processed.
[0059] Specifically, the update unit 15d updates the prior distribution so as to minimize the distance between the statistics of the features of the data to be processed and the statistics of the prior distribution defined for each class.
[0060] Alternatively, the updating unit 15d updates the prior distribution so as to maximize the distance between the statistics of the classes.
[0061] This prevents feature localization and prevents the reduction of feature norm. Therefore, the reduction of classification loss gradient is prevented, and the progress of learning is not impeded. Also, the reduction of feature entropy is prevented. Therefore, the mutual information between features and labels is improved, and a deterioration in model performance is suppressed. Furthermore, since source information is not used, applicable models are not limited. In this way, the fine-tuning device 10 makes it possible to flexibly and at low cost improve performance degradation due to overfitting.
[0062] [Example] Figures 7 to 9 are diagrams for explaining an example. In this example, ImageNet and CLIP were used as the source dataset of images, and Stanford Cars, a vehicle image classification dataset, was used as the target dataset with the data volume adjusted to 10%, 25%, 50%, 75%, and 100%. The ratio of the training dataset to the validation dataset was 9:1. ResNet-50 was used as the neural network.
[0063] And the target dataset D t De f θ The learning was carried out for 200 epochs. The accuracy of the test data set of the model with the highest accuracy (Top-1 accuracy) in the validation data set was adopted. In addition, each experimental pattern shown below was carried out five times, and the average value and standard deviation were reported.
[0064] FIG. 7 illustrates the results of this example. The experimental patterns (methods) shown in FIG. 7 were as follows: "Finetuning" is an existing technique that uses learned weights to train C_t with D_t. "RandReg-U(0,1)" is random regularization using a uniform distribution. "RandReg-N(0,1)" is random regularization using a standard normal distribution. "L2SP" is regularization that minimizes the distance between source weights and target weights. "DELTA" is regularization that makes the outputs per channel of the source and target feature extractors similar. "BSS" is regularization that maximizes the singular value of the feature map. "Co-Tuning" is regularization that pseudo-solves the source task during FT using soft source labels corresponding to pre-calculated target classes. "UOT" is a regularization method that extracts samples related to the target from the source dataset in advance and solves the target and source tasks simultaneously in FT itself. "AdaRand" is the method of the above embodiment.
[0065] As shown in Figure 7, it was confirmed that the "AdaRand" method of the above embodiment can achieve a performance of 91% or more, which is higher than other conventional methods. In particular, it was confirmed that the standard deviation (Standard Dev.) for ImageNet Classification was also within a good range of ±0.1 to ±0.2. Note that the "AdaRand" method of the above embodiment does not require additional source information, so it can be applied to any pre-trained model.
[0066] 8 illustrates changes in the feature space, similar to FIG. 3. As shown in FIG. 8, the "AdaRand" method of the above embodiment shows that clusters for each class are more aggregated than with "Fine-tuning." It also shows that the degeneration of the feature space can be prevented compared to "RandReg-U(0,1)." Therefore, it can be seen that a feature space that is easier to classify is formed than with conventional methods.
[0067] 9 illustrates the feature norm and feature entropy over the course of a training epoch. As illustrated in FIG. 9A, the "AdaRand" method of the embodiment prevents the feature norm from shrinking. Therefore, the classification loss gradient is prevented from shrinking, and the progress of training is not impeded.
[0068] 9B, the "AdaRand" method of the embodiment can prevent the reduction of feature entropy. Therefore, the mutual information between features and labels can be improved, and the degradation of model performance can be suppressed.
[0069] [Program] A program written in a computer-executable language may be created to execute the processes performed by the fine tuning device 10 according to the above embodiment. In one embodiment, the fine tuning device 10 can be implemented by installing a fine tuning program that executes the fine tuning process described above as package software or online software on a desired computer. For example, by executing the fine tuning program on an information processing device, the information processing device can function as the fine tuning device 10. The information processing device referred to here includes desktop and notebook personal computers. Other examples of information processing devices include mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone Systems), as well as slate terminals such as PDAs (Personal Digital Assistants). The functions of the fine tuning device 10 may also be implemented on a cloud server.
[0070] 10 is a diagram showing an example of a computer that executes a fine-tuning program. The computer 1000 includes, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0071] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1031. The disk drive interface 1040 is connected to a disk drive 1041. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1041. The serial port interface 1050 is connected to a mouse 1051 and a keyboard 1052, for example. The video adapter 1060 is connected to a display 1061, for example.
[0072] Here, the hard disk drive 1031 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. The various pieces of information described in the above embodiments are stored in the hard disk drive 1031 or the memory 1010, for example.
[0073] The fine tuning program is stored in the hard disk drive 1031 as a program module 1093 in which instructions to be executed by the computer 1000 are written. Specifically, the program module 1093 in which each process executed by the fine tuning device 10 described in the above embodiment is written is stored in the hard disk drive 1031.
[0074] Furthermore, data used for information processing by the fine-tuning program is stored as program data 1094, for example, in the hard disk drive 1031. Then, the CPU 1020 reads the program module 1093 and program data 1094 stored in the hard disk drive 1031 into the RAM 1012 as necessary, and executes each of the above-described procedures.
[0075] The program module 1093 and program data 1094 related to the fine-tuning program are not limited to being stored in the hard disk drive 1031, but may be stored in a removable storage medium, for example, and read by the CPU 1020 via the disk drive 1041. Alternatively, the program module 1093 and program data 1094 related to the fine-tuning program may be stored in another computer connected via a network such as a LAN or a WAN (Wide Area Network), and read by the CPU 1020 via the network interface 1070.
[0076] Although the present invention has been described above as an embodiment, the present invention is not limited to the description and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention.
[0077] REFERENCE SIGNS LIST 10 Fine tuning device 11 Input unit 12 Output unit 13 Communication control unit 14 Storage unit 15 Control unit 15a Acquisition unit 15b Initialization unit 15c Optimization unit 15d Update unit
Claims
1. A fine-tuning device comprising: an initialization unit that initializes a prior distribution on a feature space defined for each class of data to be processed using predetermined statistics of the features for each class of the data to be processed; an optimization unit that uses the prior distribution to execute an optimization process of a model that outputs the features and class of the data from the input data to be processed; and an update unit that updates the prior distribution using the features output by the model from the input data to be processed during the optimization process.
2. The fine-tuning device according to claim 1, characterized in that the update unit updates the prior distribution so as to minimize the distance between the statistics of the features of the data to be processed and the statistics of the prior distribution.
3. The fine-tuning device according to claim 1, wherein the update unit updates the prior distribution so as to maximize the distance of the statistics between classes.
4. A fine-tuning method executed by a fine-tuning device, comprising: an initialization step for initializing a prior distribution on a feature space defined for each class of data to be processed using predetermined statistics of the features for each class of the data to be processed; an optimization step for executing an optimization process of a model that outputs the features and class of the data from input data to be processed using the prior distribution; and an update step for updating the prior distribution using the features output by the model from the input data to be processed during the optimization process.
Citation Information
Patent Citations
Method of training deep neural network to classify data
JP2022044564A
Method for generating integrated model, image inspection system, device for generating model for image inspection, program for generating model for image inspection, and image inspection device
JP2022139417A