Information processing apparatus, information processing method, and program
By calculating a coefficient based on inference results for unlabeled data, the method stabilizes neural network learning and reduces user burden, achieving high accuracy without separate verification datasets.
Patent Information
- Application Number
- JP2021111255
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-05
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-07-05
AI Technical Summary
Existing semi-supervised learning methods treat unlabeled data equally, leading to unstable learning and increased user burden due to the need for separate verification teacher-labeled datasets.
A method that calculates a coefficient indicating the influence degree of unlabeled data based on inference results, using unlabeled and labeled data to update neural network weights without requiring a separate verification teacher-labeled dataset.
Enables construction of a highly accurate neural network without the need for a separate verification teacher-labeled dataset, stabilizing the learning process and reducing user burden.
Smart Images

Figure 0007700542000013 
Figure 0007700542000014 
Figure 0007700542000015
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] A neural network (hereinafter also referred to as "NN") has high performance in image recognition and the like. To improve the learning accuracy of the NN, it is known that a huge amount of input data and corresponding teacher labels are required. However, teacher labels are often assigned manually. Therefore, the burden of assigning teacher labels to a huge amount of input data is imposed on the user.
[0003] In recent years, in order to solve this problem, research on semi-supervised learning has become active in which teacher labels are assigned to a small amount of the collected input data and the NN is learned without assigning teacher labels to the remaining data. According to semi-supervised learning, the burden on the user can be greatly reduced. Generally, the loss function used in semi-supervised learning is defined as a weighted sum of a loss function corresponding to a dataset with teacher labels (labeled dataset) and a loss function corresponding to a dataset without teacher labels (unlabeled dataset).
[0004] The method described in Non-Patent Document 1 is one of the semi-supervised learning methods in image recognition in particular, and is a method in which two types of data augmentation are performed on an image without a teacher label, and learning is performed based on comparing the two types of images obtained by the two types of data augmentation. This makes it possible to improve the learning accuracy. In the method described in Non-Patent Document 1, a constant that does not depend on the input data without teacher labels (unlabeled data) is multiplied by the loss function corresponding to each unlabeled data, thereby calculating the loss function corresponding to the entire unlabeled data.
[0005] The method described in Non-Patent Document 2 is a type of semi-supervised learning. It calculates a coefficient indicating the degree of influence for each unlabeled data, and calculates a loss function corresponding to the entire unlabeled data based on the weighted sum of the coefficient for each unlabeled data and the loss function. However, the calculation of the coefficient is performed using a verification teacher-labeled dataset prepared separately from the training data.
Prior Art Documents
Non-Patent Documents
[0006]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0007] However, according to the method described in Non-Patent Document 1, the loss functions corresponding to individual unlabeled data are treated equally, and then the update of the NN based on the loss function is performed. Therefore, since the loss functions corresponding to the unlabeled data that interfere with learning are also treated equally, the learning is likely to become unstable, and it is difficult for the accuracy of the constructed NN to be high.
[0008] Also, according to the method described in Non-Patent Document 2, the influence degree of unlabeled data can be determined for each sample. However, since it is necessary to prepare a verification teacher-labeled data set separately from the learning data, the burden on the user may increase.
[0009] Therefore, the present invention has been proposed to solve these problems, and it is desired to provide a technology that enables construction of a high-precision NN without preparing a verification teacher-labeled data set separately from the learning data.
Means for Solving the Problems
[0010] In order to solve the above problems, according to one aspect of the present invention, an input unit that acquires unlabeled data and acquires labeled data with a teacher label, and based on the unlabeled data, the labeled data, and a neural network, an inference unit that outputs an inference result corresponding to the unlabeled data and an inference result corresponding to the labeled data, a coefficient generation unit that calculates a coefficient indicating the influence degree of the unlabeled data based on the inference result corresponding to the unlabeled data, an unlabeled data evaluation unit that outputs an evaluation result corresponding to the unlabeled data based on the inference result corresponding to the unlabeled data and the coefficient, a labeled data evaluation unit that outputs an evaluation result corresponding to the labeled data based on the inference result corresponding to the labeled data and the teacher label, and an update unit that updates the weight parameters of the neural network based on the evaluation result corresponding to the unlabeled data and the evaluation result corresponding to the labeled data, are provided. In the technology according to the embodiment of the present invention, unlabeled data and labeled data are used for learning. Therefore, the technology according to the embodiment of the present invention corresponds to a technology related to semi-supervised machine learning.
[0011] The inference result corresponding to the unlabeled data includes the inference value for each class. The coefficient generation unit may calculate the difference based on the inference value of the first class with the maximum inference value and the inference values of one or more classes different from the first class for each unlabeled data, and calculate the coefficient based on the difference for each unlabeled data.
[0012] The coefficient generation unit may calculate the difference between the inference value of the first class and the inference value of the second class with the second largest inference value for each unlabeled data.
[0013] The coefficient generation unit may calculate the difference between the inference value of the first class and the average value of the inference values of a plurality of classes different from the first class for each unlabeled data.
[0014] The coefficient generation unit may calculate, as the coefficient, a number having a positive correlation with the difference for each unlabeled data.
[0015] The inference result corresponding to the unlabeled data includes the inference value for each class. The coefficient generation unit may specify the predicted class based on the inference value for each unlabeled data, calculate the number of unlabeled data corresponding to the predicted class as the frequency of the predicted class, and calculate the coefficient based on the frequency of the predicted class for each unlabeled data.
[0016] The coefficient generation unit may specify, as the predicted class, the class with the maximum inference value for each unlabeled data.
[0017] The coefficient generation unit may calculate, as the coefficient, a number having a negative correlation with the frequency of the predicted class for each unlabeled data.
[0018] The unlabeled data evaluation unit may output the evaluation result corresponding to the unlabeled data based on multiplying the loss based on the inference result corresponding to the unlabeled data by the coefficient.
[0019] Further, according to another aspect of the present invention, acquiring unlabeled data and acquiring labeled data with teacher labels, and based on the unlabeled data, the labeled data, and a neural network, outputting an inference result corresponding to the unlabeled data and an inference result corresponding to the labeled data; calculating a coefficient indicating the degree of influence of the unlabeled data based on the inference result corresponding to the unlabeled data; outputting an evaluation result corresponding to the unlabeled data based on the inference result corresponding to the unlabeled data and the coefficient; outputting an evaluation result corresponding to the labeled data based on the inference result corresponding to the labeled data and the teacher label; and updating weight parameters of the neural network based on the evaluation result corresponding to the unlabeled data and the evaluation result corresponding to the labeled data. Executed by a computer An information processing method is provided.
[0020] Further, according to another aspect of the present invention, a program is provided for causing a computer to function as an information processing apparatus including: an input unit that acquires unlabeled data and acquires labeled data with teacher labels; a speculation unit that outputs an inference result corresponding to the unlabeled data and an inference result corresponding to the labeled data based on the unlabeled data, the labeled data, and a neural network; a coefficient generation unit that calculates a coefficient indicating the degree of influence of the unlabeled data based on the inference result corresponding to the unlabeled data; an unlabeled data evaluation unit that outputs an evaluation result corresponding to the unlabeled data based on the inference result corresponding to the unlabeled data and the coefficient; a labeled data evaluation unit that outputs an evaluation result corresponding to the labeled data based on the inference result corresponding to the labeled data and the teacher label; and an update unit that updates weight parameters of the neural network based on the evaluation result corresponding to the unlabeled data and the evaluation result corresponding to the labeled data.
Advantages of the Invention
[0021] As described above, according to the present invention, there is provided a technique that enables construction of a highly accurate NN without preparing a verification teacher-labeled data set separately from learning data.
Brief Description of Drawings
[0022]
Figure 1
Figure 2
Figure 3
Embodiments for Carrying Out the Invention
[0023] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the present specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant description is omitted.
[0024] Also, in the present specification and drawings, a plurality of components having substantially the same functional configuration may be distinguished by attaching different numbers after the same reference numeral. However, when it is not necessary to particularly distinguish each of a plurality of components having substantially the same functional configuration, only the same reference numeral is attached. Also, similar components in different embodiments may be distinguished by attaching different alphabets after the same reference numeral. However, when it is not necessary to particularly distinguish each of similar components in different embodiments, only the same reference numeral is attached.
[0025] (0. Outline of Embodiment) An overview of an embodiment of the present invention will be described. In an embodiment of the present invention, an information processing apparatus (hereinafter also referred to as a "learning apparatus") that performs learning of a neural network will be described. In the learning apparatus, learning of the neural network is performed based on learning data (learning stage). Thereafter, in the identification apparatus, an estimated label is output based on the learned neural network and identification data (test data).
[0026] In an embodiment of the present invention, it is mainly assumed that the learning apparatus and the identification apparatus are realized by the same computer. However, the learning apparatus and the identification apparatus may be realized by different computers. In such a case, the learned neural network generated by the learning apparatus is provided to the identification apparatus. For example, the learned neural network may be provided from the learning apparatus to the identification apparatus via a recording medium, or may be provided via communication. Hereinafter, the "learning stage" executed in the learning apparatus will be described.
[0027] (1. First Embodiment) First, a first embodiment of the present invention will be described. In the first embodiment of the present invention, semi-supervised learning is performed by a learning apparatus.
[0028] (Configuration of Learning Apparatus) With reference to FIG. 1, a configuration example of a learning apparatus according to a first embodiment of the present invention will be described. FIG. 1 is a diagram showing a functional configuration example of a learning apparatus 10 according to a first embodiment of the present invention. As shown in FIG. 1, the learning apparatus 10 according to the first embodiment of the present invention includes an input unit 111, an inference unit 122, a coefficient generation unit 131, a label-free data evaluation unit 132, a labeled data evaluation unit 133, and an update unit 134.
[0029] In the first embodiment of the present invention, it is mainly assumed that the inference unit 122 is included in the neural network 120. That is, the inference unit 122 is configured by connecting calculation graphs constructed by neurons in the order of processing, and can be regarded as one neural network as a whole. Hereinafter, the neural network is also referred to as "NN". More specifically, the inference unit 122 may mainly include a convolutional layer and a pooling layer. Hereinafter, it is mainly assumed that a two-dimensional convolutional layer is used as the convolutional layer, but a three-dimensional convolutional layer may be used.
[0030] In addition to the inference unit 122, the input unit 111, the coefficient generation unit 131, the unlabeled data evaluation unit 132, the labeled data evaluation unit 133, the update unit 134, etc. include an arithmetic device such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), and a program stored in a ROM (Read Only Memory) is expanded to the RAM by the arithmetic device and executed, whereby its function can be realized. At this time, a computer-readable recording medium recording the program may also be provided. Alternatively, these blocks may be configured by dedicated hardware or may be configured by a combination of a plurality of hardware. Data necessary for the calculation by the arithmetic device is appropriately stored by a storage unit (not shown).
[0031] The unlabeled data set 101, the labeled data set 102, and the weight parameter 121 are stored by a storage unit (not shown). Such a storage unit may be configured by a memory such as a RAM (Random Access Memory), a hard disk drive, or a flash memory.
[0032] In the initial state, an initial value is set for the weight parameter 121. For example, the initial value set for the weight parameter 121 may be a random value, but it can be any value. For example, the initial value set for the weight parameter 121 may be a learned value obtained in advance by learning.
[0033] (Unlabeled Dataset 101) The unlabeled dataset 101 is composed of a plurality of learning data (input data) to which no teacher labels are respectively associated. Hereinafter, the learning data to which no teacher label is associated is also referred to as "unlabeled data". In the embodiments of the present invention, it is mainly assumed that the unlabeled data is image data (particularly, still image data). However, the type of the unlabeled data is not particularly limited, and data other than still image data can also be used as the unlabeled data. For example, the unlabeled data may be moving image data including a plurality of frames, or may be time-series data or audio data.
[0034] (Labeled Dataset 102) The labeled dataset 102 is composed of a plurality of learning data (input data) and teacher labels respectively associated with the plurality of learning data. Hereinafter, the learning data to which a teacher label is associated is also referred to as "labeled data". Also, the combination of a teacher label and labeled data is also referred to as "labeled data". The teacher label is given manually or by a function (not shown). Similar to the type of the unlabeled data, the type of the labeled data is not particularly limited.
[0035] (Input Unit 111) The input unit 111 sequentially obtains unlabeled data from the unlabeled dataset 101, creates mini - batches based on the obtained unlabeled data, and outputs the created mini - batches to the inference unit 122 of the neural network 120. Further, the input unit 111 sequentially obtains labeled data (a combination of teacher labels and labeled data) from the labeled dataset 102, creates mini - batches based on the obtained labeled data, and outputs the created mini - batches to the inference unit 122 of the neural network 120. The size of the mini - batch is not particularly limited.
[0036] (Inference unit 122) The inference unit 122 obtains an inference result corresponding to the unlabeled data based on the unlabeled data included in the mini - batch output from the input unit 111 and the neural network 120. More specifically, the inference unit 122 obtains, as the inference result corresponding to the unlabeled data, the data output from the neural network 120 based on inputting the unlabeled data to the neural network 120 with the weight parameters 121 set. The inference unit 122 outputs the inference result corresponding to the unlabeled data to the coefficient generation unit 131.
[0037] At this time, the inference unit 122 can output two types of labels based on the framework of semi - supervised learning to the coefficient generation unit 131 as the inference result corresponding to the unlabeled data. Here, the algorithm for obtaining the two types of labels is not limited to a specific algorithm, and an algorithm used in semi - supervised learning may be used.
[0038] For example, the input unit 111 may obtain two types of unlabeled data based on the unlabeled data obtained from the unlabeled dataset 101. As an example, the input unit 111 may obtain two types of unlabeled data by performing two types of data augmentations on the unlabeled data. At this time, the input unit 111 outputs two types of unlabeled data to the inference unit 122, and the inference unit 122 outputs, as two types of labels, the labels corresponding to each of the two types of unlabeled data to the coefficient generation unit 131.
[0039] Alternatively, there may be one type of label-free data output from the input unit 111 to the inference unit 122, and the inference unit 122 may use two types of weight parameters. As an example, the inference unit 122 may obtain, as two types of labels, data obtained by applying all of the weight parameter 121 and data obtained by applying a part of the weight parameter 121 to the label-free data output from the input unit 111. At this time, the inference unit 122 outputs the two types of labels to the coefficient generation unit 131.
[0040] For example, the inference unit 122 outputs, to the coefficient generation unit 131, one of the two types of labels as a pseudo teacher label corresponding to the label-free data and the other as an estimated label corresponding to the label-free data. Note that which of the two types of labels is used as the pseudo teacher label is not limited. For example, the label obtained by weaker data augmentation may be used as the pseudo teacher label. Alternatively, the label obtained by applying all of the weight parameter 121 may be used as the pseudo teacher label.
[0041] Furthermore, the inference unit 122 obtains an inference result corresponding to the labeled data based on the labeled data included in the mini-batch output from the input unit 111 and the neural network 120. More specifically, the inference unit 122 obtains, as an inference result corresponding to the labeled data, the data output from the neural network 120 based on inputting the labeled data to the neural network 120 in which the weight parameter 121 is set. The inference unit 122 outputs the inference result corresponding to the labeled data to the labeled data evaluation unit 133.
[0042] Note that the format of the inference result output from the inference unit 122 is not particularly limited. However, the format of the inference result output from the inference unit 122 is preferably set in accordance with the format of the teacher label. For example, when the teacher label indicates the class of a classification problem and is a one-hot vector having a length corresponding to the number of classes, the format of the inference result output from the inference unit 122 may also be a vector having a length corresponding to the number of classes. At this time, the inference result output from the inference unit 122 may include a value for each class (hereinafter, also referred to as an "inference value").
[0043] As an example, when the inference unit 122 adjusts so that the sum of the inference values of all classes becomes 1, the inference value corresponding to each class may correspond to the probability corresponding to each class. However, the sum of the inference values of all classes does not necessarily have to be adjusted to 1 by the inference unit 122. In any case, the inference value output from the inference unit 122 can be a larger value as the probability of that class is higher.
[0044] (Coefficient generation unit 131) The coefficient generation unit 131 calculates a coefficient indicating the influence degree of the label-free data based on the inference result corresponding to the label-free data output from the inference unit 122. More specifically, the coefficient generation unit 131 specifies a predicted class based on the inference value for each label-free data included in the mini-batch. Assuming the batch size is B, the predicted class based on the inference values of the label-free data x u ={x1 u ,…,x B u} is specified as y u ={y1 u ,…,y B u}. Note that each element of y u may be the number of the predicted class.
[0045] Here, the prediction class may be specified in any manner. As an example, the class with the maximum inference value is considered to be the class with the highest probability. Thus, the coefficient generation unit 131 may specify, for each piece of unlabeled data included in the mini-batch, the class with the maximum inference value as the prediction class. For example, as the inference value used for specifying the prediction class, either of two types of labels may be used, but it is desirable to use a pseudo teacher label.
[0046] Then, the coefficient generation unit 131 calculates the number of pieces of unlabeled data corresponding to the prediction class as the frequency of the prediction class. For example, when the neural network 120 solves a classification problem into N classes (that is, when the inference result corresponding to the unlabeled data includes inference values for N classes), the frequency c of the prediction class i can be expressed as in the following formula (1).
[0047]
Equation
[0048] For each piece of unlabeled data x included in the mini-batch, the coefficient generation unit 131 calculates a coefficient t based on the frequency c of the prediction class. As an example, it is desirable that the coefficient generation unit 131 calculates, for each piece of unlabeled data x included in the mini-batch, a number having a negative correlation with the frequency c of the prediction class as the coefficient t. As a result, the unlabeled data corresponding to the prediction class with a smaller frequency c is treated as having a higher influence degree. u For each piece of unlabeled data x included in the mini-batch, the coefficient generation unit 131 calculates a coefficient t based on the frequency c of the prediction class. As an example, it is desirable that the coefficient generation unit 131 calculates, for each piece of unlabeled data x included in the mini-batch, a number having a negative correlation with the frequency c of the prediction class as the coefficient t. As a result, the unlabeled data corresponding to the prediction class with a smaller frequency c is treated as having a higher influence degree. u For example, assuming that a function f that returns an output value showing a negative correlation with the input value, the function f that returns an output value showing a negative correlation with the input value can be expressed as in the following formula (2).
[0049] For example, assuming that a function f that returns an output value showing a negative correlation with the input value, the function f that returns an output value showing a negative correlation with the input value can be expressed as in the following formula (2).
[0050]
Equation
[0051] For example, as an example of a function f that returns an output value showing a negative correlation with the input value, the following formula (3) can be cited.
[0052] [Number]
[0053] The coefficient generation unit 131 outputs the inference result and coefficient for each unlabeled data included in the mini-batch to the unlabeled data evaluation unit 132.
[0054] (Unlabeled data evaluation unit 132) Based on the inference result and coefficient for each unlabeled data included in the mini-batch, the unlabeled data evaluation unit 132 obtains an evaluation result corresponding to the unlabeled data. More specifically, for each unlabeled data included in the mini-batch, the unlabeled data evaluation unit 132 calculates a loss based on the inference result, and obtains an evaluation result corresponding to the unlabeled data based on the loss and coefficient for each unlabeled data.
[0055] First, for each unlabeled data, the unlabeled data evaluation unit 132 evaluates the estimated label based on the pseudo ground-truth label and calculates the loss.
[0056] Here, the loss function used for calculating the loss is not limited to a specific function, and a loss function similar to the loss function used in a general neural network may be used. For example, the loss function may be the mean squared error based on the difference between the pseudo ground-truth label corresponding to the unlabeled data and the estimated label corresponding to the unlabeled data, or the cross-entropy error based on the difference between the pseudo ground-truth label corresponding to the unlabeled data and the estimated label corresponding to the unlabeled data.
[0057] Next, the unlabeled data evaluation unit 132 obtains an evaluation result corresponding to the unlabeled data based on the loss and coefficient for each unlabeled data. More specifically, the unlabeled data evaluation unit 132 obtains an evaluation result corresponding to the unlabeled data based on multiplying the loss and coefficient for each unlabeled data.
[0058] As an example, if the unlabeled data is x u and the weight parameter 121 is θ, and the coefficient calculated by the coefficient generation unit 131 is t, the evaluation result l u corresponding to the unlabeled data can be calculated by the sum in the mini-batch of the multiplication result of the loss and coefficient for each unlabeled data, as shown in the following formula (4).
[0059]
Equation
[0060] The unlabeled data evaluation unit 132 outputs the evaluation result corresponding to the unlabeled data to the update unit 134.
[0061] (Labeled data evaluation unit 133) The labeled data evaluation unit 133 evaluates the labeled data for each labeled data included in the mini-batch based on the teacher label corresponding to the labeled data, and obtains an evaluation result for each labeled data. More specifically, the labeled data evaluation unit 133 calculates a loss based on the teacher label corresponding to the labeled data and the labeled data, and obtains an evaluation result corresponding to the labeled data based on the loss for each labeled data.
[0062] First, the labeled data evaluation unit 133 evaluates the labeled data for each labeled data based on the teacher label corresponding to the labeled data, and calculates a loss.
[0063] Here, the loss function is not limited to a specific function, and a loss function similar to the loss function used in a general neural network may be used. For example, the loss function may be the mean squared error based on the difference between the teacher label corresponding to the labeled data and the labeled data, or the cross-entropy error based on the difference between the teacher label corresponding to the labeled data and the labeled data.
[0064] Next, the labeled data evaluation unit 133 obtains an evaluation result corresponding to the labeled data based on the loss for each labeled data. More specifically, the labeled data evaluation unit 133 obtains an evaluation result corresponding to the labeled data by summing up the losses in the mini-batch for each labeled data.
[0065] As an example, let the labeled data be x t and the teacher label corresponding to the labeled data be x t and the weight parameter 121 be θ. Then, the evaluation result l s corresponding to the labeled data can be expressed as shown in the following equation (5).
[0066]
Equation
[0067] The labeled data evaluation unit 133 outputs the evaluation result corresponding to the labeled data to the update unit 134.
[0068] (Update Unit 134) The update unit 134 updates the weight parameter 121 based on the evaluation result corresponding to the unlabeled data output from the unlabeled data evaluation unit 132 and the evaluation result corresponding to the labeled data output from the labeled data evaluation unit 133. As a result, the weight parameter 121 can be trained so that the estimated label corresponding to the unlabeled data approaches the pseudo teacher label corresponding to the unlabeled data, and the labeled data approaches the teacher label corresponding to the labeled data.
[0069] For example, the update unit 134 may update the weight parameter 121 based on a weighted sum of the evaluation result corresponding to the labeled data and the evaluation result corresponding to the unlabeled data (hereinafter, also simply referred to as "weighted sum"). More specifically, the update unit 134 may update the weight parameter 121 by the error backpropagation method (backpropagation) based on the weighted sum of the evaluation result corresponding to the labeled data and the evaluation result corresponding to the unlabeled data.
[0070] The weighted sum may be expressed in any way. As an example, as shown in Equation (5), let the evaluation result corresponding to the labeled data be l s and, as shown in Equation (4), let the evaluation result corresponding to the unlabeled data be l u and, when the hyperparameter for taking the weighted sum is λ, the weighted sum L can be expressed as shown in the following Equation (6).
[0071]
Equation
[0072] Note that each time the update of the weight parameter 121 is completed, the update unit 134 determines whether the learning end condition is satisfied. If it is determined that the learning end condition is not satisfied, the next input data (a combination of labeled data and teacher labels, and unlabeled data) is acquired by the input unit 111, and each of the inference unit 122, the coefficient generation unit 131, the unlabeled data evaluation unit 132, the labeled data evaluation unit 133, and the update unit 134 executes its respective process based on the next input data again. On the other hand, if it is determined that the learning end condition is satisfied, the learning is terminated.
[0073] Note that the learning end condition is not particularly limited, and any condition indicating that the learning of the neural network 120 has been performed to a certain extent is acceptable. Specifically, the learning end condition may include the condition that the value of the weighted sum is smaller than a threshold value. Alternatively, the learning end condition may include the condition that the change in the value of the weighted sum is smaller than a threshold value (the condition that the value of the weighted sum has reached a converged state). Alternatively, the learning end condition may include the condition that the update of the weight parameter 121 has been performed a predetermined number of times. Alternatively, when the accuracy (e.g., correct answer rate, etc.) of the neural network 120 is calculated, the learning end condition may include the condition that the accuracy exceeds a predetermined ratio (e.g., 90%, etc.).
[0074] The configuration example of the learning device according to the first embodiment of the present invention has been described above.
[0075] (Operation in the learning stage) Subsequently, with reference to FIG. 2, the operation flow of the "learning stage" executed by the learning device 10 according to the first embodiment of the present invention will be described. FIG. 2 is a flowchart showing an operation example of the learning stage executed by the learning device 10 according to the first embodiment of the present invention.
[0076] First, the input unit 111 creates a mini-batch by acquiring label-free data of the batch size from the label-free dataset 101, and outputs the created mini-batch to the inference unit 122 of the neural network 120 (S101).
[0077] Subsequently, the inference unit 122 obtains an inference result corresponding to the label-free data based on the label-free data included in the mini-batch created by the input unit 111 and the neural network 120 (S102). The inference unit 122 outputs the inference result corresponding to the label-free data to the coefficient generation unit 131.
[0078] At this time, the inference unit 122 can output two types of labels to the coefficient generation unit 131 as the inference result corresponding to the label-free data. For example, the inference unit 122 outputs one of the two types of labels to the coefficient generation unit 131 as a pseudo teacher label corresponding to the label-free data, and the other as an estimated label corresponding to the label-free data.
[0079] Based on the inference result corresponding to the label-free data output from the inference unit 122, the coefficient generation unit 131 calculates a coefficient indicating the influence degree of the label-free data (S103). More specifically, the coefficient generation unit 131 specifies a prediction class based on the inference value for each label-free data included in the mini-batch. As an example, the coefficient generation unit 131 may specify, as the prediction class, the class with the maximum inference value for each label-free data included in the mini-batch. For example, as the inference value used for specifying the prediction class, either of the two types of labels may be used, but it is desirable to use the pseudo teacher label.
[0080] Then, the coefficient generation unit 131 calculates the number of label-free data corresponding to the prediction class as the frequency of the prediction class. The coefficient generation unit 131 calculates a coefficient based on the frequency of the prediction class for each label-free data included in the mini-batch. As an example, it is desirable that the coefficient generation unit 131 calculates, as the coefficient, a number having a negative correlation with the frequency of the prediction class for each label-free data included in the mini-batch. As a result, the label-free data corresponding to the prediction class with a smaller frequency is treated as having a higher influence degree.
[0081] The coefficient generation unit 131 outputs the inference result and the coefficient for each label-free data included in the mini-batch to the label-free data evaluation unit 132.
[0082] The unlabeled data evaluation unit 132 obtains an evaluation result corresponding to the unlabeled data based on the inference result and coefficient for each unlabeled data included in the mini-batch (S104). More specifically, the unlabeled data evaluation unit 132 calculates a loss based on the inference result for each unlabeled data included in the mini-batch, and obtains an evaluation result corresponding to the unlabeled data based on the loss and coefficient for each unlabeled data. The unlabeled data evaluation unit 132 outputs the evaluation result corresponding to the unlabeled data to the update unit 134.
[0083] Subsequently, the input unit 111 creates a mini-batch by obtaining, as a combination of a teacher label and labeled data, labeled data of a batch size from the labeled data set 102, and outputs the created mini-batch to the inference unit 122 of the neural network 120 (S105).
[0084] Subsequently, the inference unit 122 obtains an inference result corresponding to the labeled data based on the labeled data included in the mini-batch created by the input unit 111 and the neural network 120 (S106). The inference unit 122 outputs the inference result corresponding to the labeled data to the labeled data evaluation unit 133.
[0085] The labeled data evaluation unit 133 evaluates the labeled data for each labeled data included in the mini-batch based on the teacher label corresponding to the labeled data, and obtains an evaluation result for each labeled data (S107). More specifically, the labeled data evaluation unit 133 calculates a loss based on the teacher label corresponding to the labeled data and the labeled data, and obtains an evaluation result corresponding to the labeled data based on the loss for each labeled data. The labeled data evaluation unit 133 outputs the evaluation result corresponding to the labeled data to the update unit 134.
[0086] The update unit 134 updates the weight parameter 121 based on the evaluation result corresponding to the label-free data output from the label-free data evaluation unit 132 and the evaluation result corresponding to the labeled data output from the labeled data evaluation unit 133 (S108). As a result, the weight parameter 121 can be trained so that the estimated label corresponding to the label-free data approaches the pseudo ground-truth label corresponding to the label-free data, and the labeled data approaches the ground-truth label corresponding to the labeled data.
[0087] For example, the update unit 134 may update the weight parameter 121 based on the weighted sum of the evaluation result corresponding to the labeled data and the evaluation result corresponding to the label-free data. More specifically, the update unit 134 may update the weight parameter 121 by the error backpropagation method (backpropagation) based on the weighted sum of the evaluation result corresponding to the labeled data and the evaluation result corresponding to the label-free data.
[0088] Each time the update of the weight parameter 121 is completed, the update unit 134 determines whether the learning end condition is satisfied (S109). If it is determined that the learning end condition is not satisfied (\"NO\" in S109), the operation proceeds to S101, the next input data is acquired by the input unit 111, and each of the inference unit 122, the coefficient generation unit 131, the label-free data evaluation unit 132, the labeled data evaluation unit 133, and the update unit 134 executes its respective process based on the next input data again. On the other hand, if it is determined that the learning end condition is satisfied, the learning is terminated.
[0089] The flow of the operation of the \"learning stage\" executed by the learning device 10 according to the first embodiment of the present invention has been described above.
[0090] (Summary of the First Embodiment) As described above, according to the first embodiment of the present invention, based on the inference result corresponding to the label-free data, a pseudo-prediction class for each label-free data is specified. Then, based on the pseudo-prediction class, the degree of influence on the loss is automatically determined for each label-free data.
[0091] As a result, with respect to a phenomenon that may occur in the learning stage (particularly in the initial learning stage), that is, the phenomenon that the inference results concentrate on a specific class, it becomes possible to reduce the degree of influence on the loss of the class where the inference results concentrate. As a result, the effect that stable learning becomes possible can be enjoyed.
[0092] Also, according to the first embodiment of the present invention, the degree of influence on the loss of the label-free data can be determined without depending on the semi-supervised learning algorithm. Further, according to the first embodiment of the present invention, even when there is label-free data with a relatively small degree of influence on the loss, since the degree of influence is used for learning as a coefficient, manual data selection work using a threshold value or the like is unnecessary, and stable learning becomes possible.
[0093] The first embodiment of the present invention has been described above.
[0094] (2. Second Embodiment) Subsequently, the second embodiment of the present invention will be described. Also in the second embodiment of the present invention, semi-supervised learning is performed by the learning device.
[0095] As shown in FIG. 1, the learning device 10 according to the second embodiment of the present invention includes an input unit 111, an inference unit 122, a coefficient generation unit 131, a label-free data evaluation unit 132, a labeled data evaluation unit 133, and an update unit 134, similar to the learning device 10 according to the first embodiment of the present invention.
[0096] The learning device 10 according to the second embodiment of the present invention mainly differs in the function of the coefficient generation unit 131 as compared with the learning device 10 according to the first embodiment of the present invention. Therefore, hereinafter, the function of the coefficient generation unit 131 will be mainly described, and detailed descriptions of the functions of other blocks will be omitted.
[0097] (Coefficient generation unit 131) The coefficient generation unit 131 calculates a coefficient indicating the degree of influence of the unlabeled data based on the inference result corresponding to the unlabeled data output from the inference unit 122. More specifically, the coefficient generation unit 131 identifies the first class for which the inference value is the maximum for each unlabeled data included in the mini-batch. Then, the coefficient generation unit 131 calculates the difference based on the inference value of the first class and the inference values of one or more classes different from the first class. For example, as the inference value used for calculating the difference, either of the two types of labels may be used, but it is desirable to use a pseudo teacher label. The coefficient generation unit 131 calculates a coefficient based on the difference for each unlabeled data.
[0098] As an example, the coefficient generation unit 131 may identify the second class having the second largest inference value. At this time, the coefficient generation unit 131 may calculate the difference between the inference value of the first class and the inference value of the second class for each unlabeled data. For example, when the neural network 120 solves a classification problem into N classes (that is, when the inference result corresponding to the unlabeled data includes inference values for N classes), the difference d i can be expressed as in the following formula (7).
[0099] [Number]
[0100] In formula (7), v represents the inference result corresponding to the unlabeled data and indicates a set of inference values corresponding to each class. argmax is a function that returns the class number corresponding to the inference value passed as an argument as the output value.
[0101] Note that the method for calculating the difference based on the inference value of the first class and the inference values of one or more classes different from the first class is not limited. For example, the coefficient generation unit 131 may calculate the difference between the inference value of the first class and the average value of the inference values of a plurality of classes different from the first class for each unlabeled data. At this time, if the function for taking the average value is ave, the difference d i can be expressed as in the following formula (8).
[0102]
Equation
[0103] The coefficient generation unit 131 calculates a coefficient t based on the difference d for each unlabeled data included in the mini-batch. As an example, it is desirable for the coefficient generation unit 131 to calculate, as the coefficient t, a number having a positive correlation with the difference d for each unlabeled data included in the mini-batch. Thereby, the unlabeled data corresponding to a larger difference is treated as having a higher influence degree.
[0104] For example, if the function that returns an output value showing a positive correlation with the input value is g, the function g that returns an output value showing a positive correlation with the input value can be expressed as in the following formula (9).
[0105]
Equation
[0106] For example, as an example of the function g that returns an output value showing a positive correlation with the input value, the following formula (10) can be cited.
[0107]
Equation
[0108] The coefficient generation unit 131 outputs the inference result and coefficient for each unlabeled data included in the mini-batch to the unlabeled data evaluation unit 132.
[0109] The above described the configuration example of the learning device according to the second embodiment of the present invention.
[0110] (Operation in the learning stage) Subsequently, with reference to FIG. 2, the operation flow of the "learning stage" executed by the learning device 10 according to the second embodiment of the present invention will be described. The operation of the "learning stage" executed by the learning device 10 according to the second embodiment of the present invention is mainly different in the operation of the coefficient generation unit 131 compared to the operation of the "learning stage" executed by the learning device 10 according to the first embodiment of the present invention. Therefore, hereinafter, the operation of the coefficient generation unit 131 will be mainly described, and detailed descriptions of other operations will be omitted.
[0111] Also in the second embodiment of the present invention, similar to the first embodiment of the present invention, S101 to S102 are executed. Subsequently, the coefficient generation unit 131 calculates a coefficient indicating the influence degree of the label-free data based on the inference result corresponding to the label-free data output from the inference unit 122 (S103).
[0112] More specifically, the coefficient generation unit 131 identifies the first class with the maximum inference value for each label-free data included in the mini-batch. Then, the coefficient generation unit 131 calculates a difference based on the inference value of the first class and the inference values of one or more classes different from the first class. For example, as the inference value used for calculating the difference, either of the two types of labels may be used, but it is desirable to use a pseudo teacher label. The coefficient generation unit 131 calculates a coefficient for each label-free data based on the difference.
[0113] As an example, the coefficient generation unit 131 may identify the second class with the second largest inference value. At this time, the coefficient generation unit 131 may calculate the difference between the inference value of the first class and the inference value of the second class for each label-free data. Alternatively, the coefficient generation unit 131 may calculate the difference between the inference value of the first class and the average value of the inference values of a plurality of classes different from the first class for each label-free data.
[0114] The coefficient generation unit 131 calculates a coefficient based on the difference for each piece of unlabeled data included in the mini-batch. As an example, it is desirable for the coefficient generation unit 131 to calculate, as a coefficient, a number having a positive correlation with the difference for each piece of unlabeled data included in the mini-batch. Thereby, among the pieces of unlabeled data corresponding to large differences, those pieces of unlabeled data are treated as having a higher degree of influence.
[0115] The coefficient generation unit 131 outputs the inference result and the coefficient for each piece of unlabeled data included in the mini-batch to the unlabeled data evaluation unit 132.
[0116] Also in the second embodiment of the present invention, S104 to S109 are executed in the same manner as in the first embodiment of the present invention.
[0117] The flow of the operation of the "learning stage" executed by the learning device 10 according to the second embodiment of the present invention has been described above.
[0118] (Summary of the Second Embodiment) As described above, according to the second embodiment of the present invention, based on the inference result corresponding to the unlabeled data, a difference is calculated based on the inference value of the first class for which the inference value is maximized for each piece of unlabeled data and the inference values of one or more classes different from the first class. Then, the degree of influence on the loss is automatically determined for each piece of unlabeled data based on the difference.
[0119] Thereby, the same effects as those of the first embodiment of the present invention can be enjoyed. Furthermore, according to the second embodiment of the present invention, even when the batch size is relatively small, the degree of influence on the loss can be determined with high accuracy. Therefore, according to the second embodiment of the present invention, stable learning is possible even when the batch size is relatively small.
[0120] The second embodiment of the present invention has been described above.
[0121] (3. Hardware Configuration Example) Next, a hardware configuration example of the learning device 10 according to the first embodiment of the present invention will be described. Note that the hardware configuration of the learning device 10 according to the second embodiment of the present invention can also be realized in the same manner as the hardware configuration of the learning device 10 according to the first embodiment of the present invention.
[0122] Hereinafter, as a hardware configuration example of the learning device 10 according to the first embodiment of the present invention, a hardware configuration example of the information processing device 900 will be described. Note that the hardware configuration example of the information processing device 900 described below is merely an example of the hardware configuration of the learning device 10. Therefore, the hardware configuration of the learning device 10 may have unnecessary configurations deleted from the hardware configuration of the information processing device 900 described below, or new configurations may be added.
[0123] FIG. 3 is a diagram showing the hardware configuration of the information processing device 900 as an example of the learning device 10 according to the first embodiment of the present invention. The information processing device 900 includes a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, a RAM (Random Access Memory) 903, a host bus 904, a bridge 905, an external bus 906, an interface 907, an input device 908, an output device 909, a storage device 910, and a communication device 911.
[0124] The CPU 901 functions as an arithmetic processing unit and a control unit, and controls the overall operation within the information processing device 900 according to various programs. Also, the CPU 901 may be a microprocessor. The ROM 902 stores programs, arithmetic parameters, etc. used by the CPU 901. The RAM 903 temporarily stores programs used in the execution of the CPU 901 and parameters that change as appropriate during the execution. These are interconnected by a host bus 904 composed of a CPU bus or the like.
[0125] The host bus 904 is connected to an external bus 906 such as a PCI (Peripheral Component Interconnect / Interface) bus via a bridge 905. Note that it is not necessarily required to configure the host bus 904, the bridge 905, and the external bus 906 separately, and these functions may be implemented on a single bus.
[0126] The input device 908 is composed of input means such as a mouse, a keyboard, a touch panel, buttons, a microphone, switches, and levers for a user to input information, and an input control circuit that generates an input signal based on the input by the user and outputs it to the CPU 901. A user who operates the information processing device 900 can input various data to the information processing device 900 or instruct a processing operation by operating this input device 908.
[0127] The output device 909 includes, for example, display devices such as a CRT (Cathode Ray Tube) display device, a liquid crystal display (LCD) device, an OLED (Organic Light Emitting Diode) device, a lamp, and a voice output device such as a speaker.
[0128] The storage device 910 is a device for storing data. The storage device 910 may include a storage medium, a recording device for recording data on the storage medium, a reading device for reading data from the storage medium, a deleting device for deleting data recorded on the storage medium, and the like. The storage device 910 is composed of, for example, an HDD (Hard Disk Drive). This storage device 910 drives a hard disk and stores programs and various data executed by the CPU 901.
[0129] The communication device 911 is a communication interface composed of, for example, a communication device for connecting to a network. Also, the communication device 911 may support either wireless communication or wired communication.
[0130] The above has described an example of the hardware configuration of the learning device 10 according to the first embodiment of the present invention.
[0131] (4. Summary) As described above in detail with reference to the accompanying drawings, the preferred embodiments of the present invention are not limited to such examples. It is obvious that those having ordinary knowledge in the technical field to which the present invention pertains can come up with various modification examples or correction examples within the scope of the technical idea described in the claims, and these are also naturally understood to belong to the technical scope of the present invention.
[0132] In the first embodiment and the second embodiment of the present invention, the case where the learning data is image data (particularly, still image data) has been mainly described. However, the type of the learning data is not particularly limited. For example, as long as feature amounts corresponding to the type of the learning data are extracted, data other than still image data can also be used as the learning data. For example, the learning data may be moving image data including a plurality of frames, or may be audio data.
[0133] At this time, when the learning data is still image data, it is common to use a 2D convolutional layer as the convolutional layer included in the inference unit 122. On the other hand, if a 3D convolutional layer is used as the convolutional layer included in the inference unit 122, moving image data can be applied as the learning data.
[0134] In the first embodiment of the present invention, as an example of the function f that returns an output value showing a negative correlation with respect to the input value, the following formula (3) has been described. However, the function f that returns an output value showing a negative correlation with respect to the input value is not limited to such an example. For example, as examples of the function f that returns an output value showing a negative correlation with respect to the input value, the following formula (3-A) and formula (3-B) can also be cited.
[0135]
Equation
[0136] In the second embodiment of the present invention, as an example of the function g that returns an output value showing a positive correlation with the input value, the following formula (9) was described. However, the function g that returns an output value showing a positive correlation with the input value is not limited to such an example. For example, as examples of the function g that returns an output value showing a positive correlation with the input value, the following formula (9-A) and formula (9-B) can also be cited.
[0137]
Equation
Explanation of Signs
[0138] 10 Learning device 101 Unlabeled dataset 102 Labeled dataset 111 Input section 120 Neural network 121 Weight parameter 122 Inference section 131 Coefficient generation section 132 Unlabeled data evaluation section 133 Labeled data evaluation section 134 Update section
Claims
1. An input unit that acquires data without a label and acquires labeled data with a teacher label; A speculation unit that outputs an inference result corresponding to the data without a label and an inference result corresponding to the data with a label based on the data without a label, the data with a label, and a neural network; A coefficient generation unit that calculates a coefficient indicating the degree of influence of the data without a label based on the inference result corresponding to the data without a label; A data-without-label evaluation unit that outputs an evaluation result corresponding to the data without a label based on the inference result corresponding to the data without a label and the coefficient; A data-with-label evaluation unit that outputs an evaluation result corresponding to the data with a label based on the inference result corresponding to the data with a label and the teacher label; An update unit that updates the weight parameters of the neural network based on the evaluation result corresponding to the data without a label and the evaluation result corresponding to the data with a label; An information processing apparatus comprising the above.
2. The inference result corresponding to the data without a label includes inference values for each class. The coefficient generation unit calculates a difference based on the inference value of the first class with the maximum inference value for each data without a label and the inference values of one or more classes different from the first class, and calculates the coefficient for each data without a label based on the difference. The information processing apparatus according to Claim 1.
3. The coefficient generation unit calculates a difference between the inference value of the first class and the inference value of the second class with the second largest inference value for each data without a label. The information processing apparatus according to Claim 2.
4. The coefficient generation unit calculates a difference between the inference value of the first class and the average value of the inference values of a plurality of classes different from the first class for each data without a label. The information processing apparatus according to Claim 2.
5. The coefficient generation unit calculates, as the coefficient, a number having a positive correlation with the difference for each data without a label. The information processing apparatus according to any one of Claims 2 to 4.
6. The inference result corresponding to the data without a label includes inference values for each class. The coefficient generation unit identifies a predicted class based on an inference value for each unlabeled data, calculates the number of unlabeled data corresponding to the predicted class as the frequency of the predicted class, and calculates the coefficient based on the frequency of the predicted class for each unlabeled data. The information processing apparatus according to claim 1.
7. The coefficient generation unit identifies, as a predicted class, the class for which the inference value is maximum for each unlabeled data. The information processing apparatus according to claim 6.
8. The coefficient generation unit calculates, as the coefficient, a number having a negative correlation with the frequency of the predicted class for each unlabeled data. The information processing apparatus according to claim 6 or 7.
9. The unlabeled data evaluation unit outputs an evaluation result corresponding to the unlabeled data based on multiplying a loss based on an inference result corresponding to the unlabeled data by the coefficient. The information processing apparatus according to any one of claims 2 to 8.
10. acquiring unlabeled data and acquiring labeled data with teacher labels, outputting an inference result corresponding to the unlabeled data and an inference result corresponding to the labeled data based on the unlabeled data, the labeled data, and a neural network, calculating a coefficient indicating the degree of influence of the unlabeled data based on an inference result corresponding to the unlabeled data, outputting an evaluation result corresponding to the unlabeled data based on the inference result corresponding to the unlabeled data and the coefficient, outputting an evaluation result corresponding to the labeled data based on the inference result corresponding to the labeled data and the teacher label, updating weight parameters of the neural network based on the evaluation result corresponding to the unlabeled data and the evaluation result corresponding to the labeled data, An information processing method executed by a computer, comprising:
11. A computer, an input unit that acquires unlabeled data and acquires labeled data with teacher labels, a speculation unit that outputs an inference result corresponding to the unlabeled data and an inference result corresponding to the labeled data based on the unlabeled data, the labeled data, and a neural network, A coefficient generation unit that calculates a coefficient indicating the degree of influence of the unlabeled data based on the inference result corresponding to the unlabeled data; An unlabeled data evaluation unit that outputs an evaluation result corresponding to the unlabeled data based on the inference result corresponding to the unlabeled data and the coefficient; A labeled data evaluation unit that outputs an evaluation result corresponding to the labeled data based on the inference result corresponding to the labeled data and the teacher label; An update unit that updates the weight parameters of the neural network based on the evaluation result corresponding to the unlabeled data and the evaluation result corresponding to the labeled data; A program for causing an information processing apparatus to function as described above.
Citation Information
Patent Citations
Method for training multi-class classifier
JP2010231768A
Binary classification learning device, binary classification device, method, and program
JP2017126158A
Deep learning system and method for diagnosis of chest conditions from chest radiographs
WO2021091661A1