Learning device, method and program

The learning device and method address the challenge of reduced classification accuracy in domain adaptation by calculating and minimizing distribution loss between estimated and target distributions, thereby improving learning accuracy in defect image classification tasks.

JP7675562B2Active Publication Date: 2025-05-13KK TOSHIBA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021091243
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-31
Publication Date
2025-05-13
Estimated Expiration
2041-05-31

AI Technical Summary

Technical Problem

Existing domain adaptation methods for defect image classification face challenges when the expected distribution of classifications between the source and target domains differs, leading to reduced classification accuracy, especially when pseudo-labels are incorrect.

Method used

A learning device and method that includes a first acquisition unit, a classification unit, a distribution loss calculation unit, and an update unit, which acquires learning data, generates estimation vectors, calculates distribution loss between estimated and target distributions, and updates the model parameters based on this loss, thereby improving learning accuracy.

Benefits of technology

The proposed solution effectively improves learning accuracy by calculating and minimizing distribution loss between estimated and target distributions, even in scenarios where no labels are provided to the training data, thus enhancing classification accuracy over traditional unsupervised learning methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007675562000007
    Figure 0007675562000007
  • Figure 0007675562000008
    Figure 0007675562000008
  • Figure 0007675562000009
    Figure 0007675562000009
Patent Text Reader

Abstract

To improve learning precision.SOLUTION: A learning apparatus according to the present embodiment includes a first acquisition unit, a classification unit, a first generation unit, a distribution loss calculation unit, and an update unit. The first acquisition unit acquires first data for learning. The classification unit inputs the first data for learning to a model, and generates a plurality of estimation vectors that are processing results of the model. The first generation unit generates an estimation distribution from the plurality of estimation vectors. The distribution loss calculation unit calculates a distribution loss between the estimation distribution and a target distribution to be a target in an inference using the model. The update unit updates parameters of the model, based on the distribution loss.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] FIELD An embodiment of the present invention relates to a learning device, a learning method, and a learning program. [Background technology]

[0002] In defect image classification, which uses machine learning to identify defects in products, etc., previously trained models may no longer be usable due to product generation changes. For example, changes in the shape or manufacturing process of a product may result in defects that differ from those encountered during the manufacturing process of the previous generation of products, such as the tendency for dust to be mixed into the new product during the manufacturing process. In other words, the tendency of the data used for learning may change. To address this issue, there is a method called domain adaptation, which combines learning data from the previous generation (source domain) and the new generation (target domain) to improve classification accuracy for the target domain. However, in domain adaptation, if the expected classification distribution (class distribution) between the source domain and the target domain is different, the classification accuracy in the target domain will be low. Although there are domain adaptation methods using pseudo labels, the pseudo labels are not necessarily correct, and therefore erroneous learning may occur, making it difficult to improve classification accuracy. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Shuhan Tan et al., “Class-imbalanced Domain Adaptation: An Empirical Odyssey”, [online], September 19, 2020, [searched on May 12, 2021], Internet<URL : http: / / arxiv.org / abs / 1910.10320> Summary of the Invention [Problem to be solved by the invention]

[0004] The present disclosure has been made to solve the above-mentioned problems, and aims to provide a learning device, method, and program that can improve learning accuracy. [Means for solving the problem]

[0005] The learning device according to this embodiment includes a first acquisition unit, a classification unit, a first generation unit, a distribution loss calculation unit, and an update unit. The first acquisition unit acquires first learning data. The classification unit inputs the first learning data to a model and generates a plurality of estimated vectors that are processing results of the model. The first generation unit generates an estimated distribution from the plurality of estimated vectors. The distribution loss calculation unit calculates a distribution loss between the estimated distribution and a target distribution that is a target in inference using the model. The update unit updates parameters of the model based on the distribution loss. [Brief description of the drawings]

[0006] [Figure 1] FIG. 1 is a block diagram showing a learning device according to a first embodiment. [Diagram 2] 4 is a flowchart showing the operation of the learning device according to the first embodiment. [Diagram 3] FIG. 11 is a block diagram showing a learning device according to a second embodiment. [Figure 4] 10 is a flowchart showing the operation of a learning device according to a second embodiment. [Diagram 5] FIG. 13 is a block diagram showing a learning device according to a third embodiment. [Figure 6] 13 is a flowchart showing the operation of a learning device according to a third embodiment. [Figure 7] FIG. 13 is a conceptual diagram showing the operation of the inference device according to the fourth embodiment. [Figure 8] FIG. 2 is a block diagram showing an example of the hardware configuration of the learning device according to the embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0007] Hereinafter, a learning device, a method, and a program according to the present embodiment will be described in detail with reference to the drawings. In the following embodiments, parts with the same reference numerals perform the same operations, and duplicated descriptions will be omitted as appropriate.

[0008] (First embodiment) A learning device according to a first embodiment will be described with reference to the block diagram of FIG. The learning device 10 according to the first embodiment includes a data acquiring unit 101, a classifying unit 102, an estimated distribution generating unit 103, a distribution loss calculating unit 104, and an updating unit 105.

[0009] The data acquisition unit 101 acquires learning data from an external source. The learning data is assumed to be data to which no instruction label is assigned. The learning data is data of a target domain to be inferred by a trained model.

[0010] The classification unit 102 receives learning data from the data acquisition unit 101, and inputs the learning data to a model to generate an estimated vector, which is a processing result of the model. The estimated vector is a classification result corresponding to the output of the model, which indicates the probability of belonging to each of a plurality of classes. Note that in this embodiment including the first embodiment and the following embodiments, the "model" is assumed to be a machine learning model such as a neural network aimed at a class classification task, but is not limited to this. In other words, the learning process by the learning device 10 according to this embodiment can be similarly applied to machine learning models aimed at other tasks such as image recognition and regression.

[0011] The estimated distribution generating unit 103 receives the estimated vector from the classifying unit 102, and generates an estimated class distribution (also called an estimated distribution) that is estimated for the entire training data using the estimated vector.

[0012] The distribution loss calculation unit 104 calculates a distribution loss, which is a difference between the estimated class distribution from the estimated distribution generation unit 103 and a target class distribution (also called a target distribution) that is a target in inference using a model. The target class distribution is, for example, a distribution of class classification predicted or assumed in class classification for data in a target domain.

[0013] The update unit 105 acquires the distribution loss from the distribution loss calculation unit 104, and updates the model parameters based on the distribution loss. The update unit 105 stops updating the model parameters based on a predetermined condition, thereby generating a trained model.

[0014] Next, the operation of the learning device 10 according to the first embodiment will be described with reference to the flowchart of FIG. In step S201, the data acquisition unit 101 acquires a plurality of pieces of learning data. Specifically, the learning data may be, for example, photographed images of a product.

[0015] In step S202, the classification unit 102 classifies a mini-batch X t Obtain minibatch X t is a subset of data selected from the training data, and multiple mini-batches are generated from the training data. A general mini-batch generation method can be used to generate mini-batches, such as randomly selecting a certain number of data from the training data. Here, one mini-batch is obtained from multiple mini-batches generated by an existing method.

[0016] In step S203, the mini-batch X t The training data x contained in t The model performs classification on the estimated vector f θ (y|x t ) where θ is the weight, bias, and other parameters set in the model. The estimated vector f θ (y|x t ) is the training data x tIt indicates which of multiple class classifications the given value falls into and the percentage of confidence in the corresponding class classification. Specifically, let us assume that the learning data are photographed images of a product, and a product defect classification task is to be performed. In the case where a product has a defect, and the photographed image is used to estimate and classify whether the defect corresponds to a scratch, foreign matter contamination, or dirt, let us assume that the estimated vector obtained is [scratch, foreign matter contamination, dirt] = [0.7, 0.2, 0.1]. In this case, the estimated vector indicates that there is a high possibility that the product defect is a scratch, in other words, that there is a high degree of certainty that the product defect is a scratch. Thus, here, one learning data x t , one estimate vector f θ (y|x t ) is generated.

[0017] In step S204, the mini-batch X t Determine whether classification processing has been completed for all learning data in mini-batch X t If all the learning data included in has been processed, proceed to step S205 and t If there is unprocessed learning data included in the set, the process of step S204 is repeated.

[0018] In step S205, the estimated distribution generating unit 103 generates a mini-batch X t The estimated distribution generation unit 103 generates an estimated class distribution, which is the confidence level for each class classification when the mini-batch X t The estimated class distribution can be calculated by averaging multiple estimated vectors generated from all the training data in t ) can be expressed by equation (1). Note that “^” represents a superscript hat, and q^ represents an estimate of the class distribution.

[0019]

number

[0020] In addition, the estimated class distribution q^(y|X t ) is a set of multiple estimated vectors f θ (y|x t ) but may be calculated using other statistics such as the median of each element of a vector.

[0021] In step S206, the distribution loss calculation unit 104 calculates the estimated class distribution q^(y|X t ) and the target class distribution q(y), is calculated using the Kullback-Leibler divergence (KL (Kullback-Leibler) divergence). Note that any method that can verify the difference between two probability distributions, such as JS (Jensen-Shannon) divergence, can be used instead of the Kullback-Leibler divergence. The target class distribution q(y) may be calculated from data that is manually labeled as the correct answer for data in the same domain as the training data, for example, or if the target distribution or expected distribution for the classification results can be known in advance, that distribution can be used as the target class distribution q(y). Specifically, the distribution loss L using the Kullback-Leibler divergence is dist (X t ) can be expressed by equation (2).

[0022]

number

[0023] Here, D(q^(y|Xt)||q(y)) on the right-hand side of equation (2) is the Kullback-Leibler divergence calculated by equation (3). Note that Y is the set of all classes.

[0024]

number

[0025] In step S207, the update unit 105 determines whether the learning is completed. The learning may be determined to be completed when a predetermined number of epochs have been completed, or when the distribution loss L dist(X t ) is equal to or less than a threshold, it may be determined that learning has ended. If learning has ended, parameter update is discontinued and processing ends. This generates a trained model. On the other hand, if learning has not ended, proceed to step S208.

[0026] In step S208, the update unit 105 updates the distribution loss L dist (X t The model parameters (such as the weights and biases of the neural network) are updated, for example, by gradient descent and backpropagation, so that θ(θ) is minimized. Then, the process returns to step S202 and executes steps S202 to S208 on the unprocessed mini-batch.

[0027] According to the first embodiment described above, by calculating the distribution loss between the estimated class distribution and the target class distribution and learning the model parameters based on the distribution loss, it is possible to perform appropriate learning even if the learning data is not labeled, and to improve the learning accuracy. In this embodiment, it is possible to improve the classification accuracy compared to a trained model trained by general unsupervised learning.

[0028] Second embodiment The second embodiment differs from the first embodiment in that a target class distribution is generated from learning data with instruction labels.

[0029] The learning device according to the second embodiment will be described with reference to the block diagram of FIG. The learning device 30 according to the second embodiment includes a data acquiring unit 101, a classifying unit 102, an estimated distribution generating unit 103, a distribution loss calculating unit 104, an updating unit 105, and a target distribution generating unit 301. The classification unit 102, the estimated distribution generation unit 103, the distribution loss calculation unit 104, and the update unit 105 are similar to those in the first embodiment, and therefore the description thereof will be omitted.

[0030] The data acquiring unit 101 acquires learning data to which an instruction label has been assigned (hereinafter referred to as second learning data) in addition to the learning data acquired in the first embodiment (hereinafter referred to as first learning data). The second learning data is data in the same domain as the first learning data, and is, for example, data to which an instruction label has been manually assigned to a part of the first learning data. The target distribution generating unit 301 receives the second learning data and the teaching labels from the data acquiring unit 101, and generates a target class distribution based on the teaching labels.

[0031] Next, the operation of the learning device according to the second embodiment will be described with reference to FIG. Steps S201 to S208 are the same as those in the first embodiment, and therefore the description will be omitted.

[0032] In step S401, the data acquiring unit 101 acquires a plurality of second learning data with instruction labels. In step S402, the target distribution generating unit 301 generates a target class distribution based on the teaching labels. For example, the target class distribution may be a frequency distribution obtained by dividing the number of cases for each teaching label class by the total number of teaching labels. In the case of the target class distribution, the distribution of class classification may be calculated using other statistics, as in the case of the estimated class distribution. In step S206, the distribution loss calculation unit 104 may calculate the distribution loss between the target class distribution and the estimated class distribution based on the target class distribution calculated in step S402.

[0033] According to the second embodiment described above, by generating a target class distribution from the second training data with instruction labels, it is possible to perform appropriate learning and improve the learning accuracy even if the first training data is not labeled, as in the first embodiment.

[0034] (Third embodiment) The third embodiment differs from the second embodiment in that the entropy loss between estimated vectors is calculated when generating an estimated class distribution.

[0035] A learning device according to the third embodiment will be described with reference to the block diagram of FIG. The learning device 50 according to the third embodiment includes a data acquiring unit 101, a classifying unit 102, an estimated distribution generating unit 103, a distribution loss calculating unit 104, an updating unit 105, a target distribution generating unit 301, and an entropy loss calculating unit 501. The data acquisition unit 101, the classification unit 102, the estimated distribution generation unit 103, the distribution loss calculation unit 104, and the target distribution generation unit 301 are similar to those in the first and second embodiments, and therefore description thereof will be omitted.

[0036] The entropy loss calculation unit 501 receives a plurality of estimated vectors from the classification unit 102 and calculates the entropy loss between the estimated vectors. Entropy represents the bias of the estimated vector. If the estimation is only updated based on the distribution loss, the individual estimated vectors estimated from one piece of learning data may be trained to approach the target class distribution, resulting in an ambiguous estimation. For each piece of learning data, by learning that this class is close to 1 and the others are zero, a confident estimation is performed for each piece of learning data, and learning that approaches the target class distribution as a whole can be performed. Therefore, by adding an entropy loss to penalize ambiguous estimation, the classification accuracy can be improved. The update unit 105 receives the entropy loss from the entropy loss calculation unit 501 and the distribution loss from the distribution loss calculation unit 104, and updates the parameters of the model based on the entropy loss and the distribution loss.

[0037] The operation of the learning device 50 according to the third embodiment will be described with reference to the flowchart of FIG. Steps other than step S207, step S601, and step S602 are the same as those in the flowchart shown in FIG. 4, and therefore will not be described.

[0038] In step S601, the entropy loss calculation unit 501 calculates the entropy loss based on a plurality of estimated vectors. ent (X t ) can be expressed, for example, by equation (4).

[0039]

number

[0040] In step S207, the update unit 105 updates the distribution loss L dist (X t ) and entropy loss L ent (X t ) to determine whether learning has been completed. For example, the distribution loss L dist (X t ) and entropy loss L ent (X t ) and the combined loss L(X t If ) is equal to or smaller than the threshold, learning ends, and if it is greater than the threshold, the process proceeds to step S602.

[0041]

number

[0042] Or, as shown in equation (6), the distribution loss L dist (X t ) and entropy loss L ent (X t ) and the weighted sum of the composite loss L(X t ) where α and β are any real numbers.

[0043]

number

[0044] In step S602, the update unit 105 updates the parameters based on the combined loss of equation (5) or (6) so as to minimize the combined loss. In addition, when the combined loss is calculated by a weighted sum of the entropy loss and the distribution loss, the parameters may be updated according to the weight so that the loss with a larger weight is preferentially minimized.

[0045] According to the third embodiment described above, the entropy loss of a plurality of estimated vectors is calculated, and the model parameters are updated using the distribution loss and the entropy loss. By using the entropy loss, it is possible to penalize ambiguous estimation, and thus it is possible to improve the learning accuracy of the model.

[0046] (Fourth embodiment) In the fourth embodiment, an example is shown in which inference is performed by an inference device including a trained model trained by a learning device according to any one of the first to third embodiments. FIG. 7 shows a conceptual diagram of the operation of an inference device 70 according to the fourth embodiment. As shown in FIG. 7, an inference device 70 includes a trained model 701 trained by the learning device according to the embodiment described above, and a model execution unit 702. When target data 71, which is the subject of inference, is input to the inference device 70, the model execution unit 702 performs inference on the target data 71 using a trained model, and outputs, as the inference result, a classification result 72 indicating the probability of class classification for the target data 71. According to the fourth embodiment described above, by performing inference using the trained model trained in the above-described embodiment, it is possible to perform highly accurate inference (e.g., class classification).

[0047] Next, an example of the hardware configuration of the learning device 10 and the inference device 70 according to the above-described embodiment is shown in a block diagram in FIG.

[0048] The learning device 10 and the inference device 70 include a CPU (Central Processing Unit) 81, a RAM (Random Access Memory) 82, a ROM (Read Only Memory) 83, a storage 84, a display device 85, an input device 86, and a communication device 87, each of which is connected by a bus.

[0049] The CPU 81 is a processor that executes arithmetic processing and control processing according to a program. The CPU 81 executes the processing of each part of the learning device 10 and the inference device 70 described above in cooperation with the programs stored in the ROM 83 and the storage 84, etc., using a predetermined area of ​​the RAM 82 as a working area.

[0050] The RAM 82 is a memory such as a Synchronous Dynamic Random Access Memory (SDRAM). The RAM 82 functions as a work area for the CPU 81. The ROM 83 is a memory that stores programs and various information in a non-rewritable manner.

[0051] The storage 84 is a device that writes and reads data to a magnetic recording medium such as a hard disk drive (HDD), a semiconductor storage medium such as a flash memory, a magnetically recordable storage medium such as a HDD, an optically recordable storage medium, etc. The storage 84 writes and reads data to the storage medium in response to control from the CPU 81.

[0052] The display device 85 is a display device such as an LCD (Liquid Crystal Display), etc. The display device 85 displays various information based on a display signal from the CPU 81. The input device 86 is an input device such as a mouse, a keyboard, etc. The input device 86 receives information input by a user as an instruction signal, and outputs the instruction signal to the CPU 81. The communication device 87 communicates with external devices via a network under the control of the CPU 81 .

[0053] The instructions shown in the processing procedures shown in the above-mentioned embodiments can be executed based on a program, which is software. A general-purpose computer system can store this program in advance and obtain the same effect as the control operation of the learning device and inference device described above by reading this program. The instructions described in the above-mentioned embodiments are recorded as a program that can be executed by a computer on a magnetic disk (flexible disk, hard disk, etc.), an optical disk (CD-ROM, CD-R, CD-RW, DVD-ROM, DVD±R, DVD±RW, Blu-ray (registered trademark) Disc, etc.), a semiconductor memory, or a recording medium similar thereto. The recording medium may be in any storage format as long as it is readable by a computer or an embedded system. If the computer reads the program from this recording medium and causes the CPU to execute the instructions described in the program based on this program, it can realize an operation similar to the control of the learning device and inference device of the above-mentioned embodiments. Of course, when the computer acquires or reads the program, it may acquire or read it through a network. In addition, an OS (operating system), database management software, network, or other MW (middleware) running on a computer may execute some of the processes required to realize this embodiment based on instructions from a program installed on the computer or embedded system from a recording medium. Furthermore, the recording medium in this embodiment is not limited to a medium independent of a computer or an embedded system, but also includes a recording medium that stores or temporarily stores a program downloaded via a LAN, the Internet, or the like. Furthermore, the number of recording media is not limited to one, and cases in which the processing in this embodiment is executed from multiple media are also included in the recording media in this embodiment, and the media may have any configuration.

[0054] The computer or embedded system in this embodiment is for executing each process in this embodiment based on a program stored in a recording medium, and may be configured as any one of a device such as a personal computer or a microcomputer, or a system in which multiple devices are connected to a network. In addition, the computer in this embodiment is not limited to a personal computer but also includes an arithmetic processing device, a microcomputer, etc. included in information processing equipment, and is a general term for equipment or devices that can realize the functions in this embodiment by a program.

[0055] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included in the scope and spirit of the invention, and are included in the scope of the invention and its equivalents described in the claims. [Explanation of symbols]

[0056] 10,30,50 Learning Device 70 Reasoning device 71 Target Data 72 Classification results 81 CPU 82 RAM 83 ROM 84 Storage 85 Display device 86 Input Devices 87 Communication Equipment 101 Data Acquisition Section 102 Classification Department 103 Estimated distribution generator 104 Distribution Loss Calculation Section 105 Update section 301 Target distribution generation unit 501 Entropy Loss Calculation Unit 701 trained models 702 Model Execution Department

Claims

1. A first acquisition unit that acquires first learning data; a classification unit that inputs the first training data into a model and generates a plurality of estimated vectors that are processing results of the model; an entropy loss calculation unit that calculates an entropy loss from the plurality of estimated vectors; a first generation unit that generates an estimated distribution from the plurality of estimated vectors; a distribution loss calculation unit that calculates a distribution loss between the estimated distribution and a target distribution that is a target in inference using the model; an update unit that updates parameters of the model so as to minimize a combined loss of the entropy loss and the distribution loss; A learning device comprising:

2. The learning device according to claim 1 , wherein the update unit updates the parameters by weighting the entropy loss and the distribution loss.

3. a second acquisition unit that acquires second learning data that is in the same domain as the first learning data and a label to be assigned to the second learning data; The learning device according to claim 1 or 2, further comprising: a second generation unit that generates the target distribution based on the label.

4. The learning device according to claim 3 , wherein the second generation unit generates a frequency distribution of the labels as the target distribution.

5. The classification unit generates the plurality of estimated vectors by inputting the first training data to the model in mini-batches; The learning device according to claim 1 , wherein the first generation unit generates an average of the plurality of estimated vectors as the estimated distribution.

6. The learning device according to claim 1 , wherein the distribution loss calculation unit calculates a Kullback-Leibler divergence between the estimated distribution and the target distribution as the distribution loss.

7. The learning device according to claim 1 , wherein the trained model is generated by the update unit interrupting the update of the parameters of the model based on a predetermined condition.

8. A computer comprising: Acquire first learning data; inputting the first training data into a model and generating a plurality of estimated vectors that are a processing result of the model; Calculating an entropy loss from the plurality of estimated vectors; generating an estimated distribution from the plurality of estimated vectors; Calculating a distribution loss between the estimated distribution and a target distribution that is a target in inference using the model; A learning method, comprising: updating parameters of the model so that a combined loss of the entropy loss and the distribution loss is minimized.

9. Computer, A first acquisition means for acquiring first learning data; A classification means for inputting the first training data into a model and generating a plurality of estimated vectors that are a processing result of the model; an entropy loss calculation means for calculating an entropy loss from the plurality of estimated vectors; a first generating means for generating an estimated distribution from the plurality of estimated vectors; A distribution loss calculation means for calculating a distribution loss between the estimated distribution and a target distribution that is a target in inference using the model; a learning program for functioning as an update means for updating parameters of the model so as to minimize a combined loss of the entropy loss and the distribution loss.

Citation Information

Patent Citations

  • Domain adaptation for image classification using class prior probability

    JP2016058079A

  • Learning method, learning program, and learning device

    JP2020115273A

  • Information processing device and information processing method

    JP2020155101A

  • Learning method, learning program, and learning device

    JP2021015425A

  • Estimating system, estimating device, and estimating method

    JP2021081795A