Learning device, learning method, and program
By adding an asymmetrically distributed margin to the loss function based on class sample ratios, the learning device and method improve separation performance and fairness in classifier models trained on datasets with uneven class sample numbers, effectively addressing the challenges of class imbalance and multi-label classification.
Patent Information
- Application Number
- JP2024510870
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-03-30
AI Technical Summary
Existing methods for training classifier models using datasets with uneven class sample numbers struggle to achieve excellent separation performance and fairness between classes, particularly in multi-label scenarios.
A learning device and method that add a margin to the loss function during training, where the total margin is fixed and asymmetrically distributed across classes based on the ratio of class samples, to enhance separation performance and fairness.
The proposed solution enables the training of classifier models that excel in separation performance and fairness, even with datasets having uneven class sample numbers, by effectively addressing the challenges of class imbalance and multi-label classification.
Smart Images

Figure 0007697588000016 
Figure 0007697588000017 
Figure 0007697588000018
Abstract
Description
Technical Field
[0001] The present invention relates to a learning device, a learning method, and a recording medium.
Background Art
[0002] Non-Patent Documents 1-6 describe adding a margin to a loss function used for angular distance learning for classification.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Non-Patent Document 5
Non-Patent Document 6
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the methods described in Non-Patent Documents 1 and 2, when training a classifier model using a dataset with uneven class sample numbers, it is difficult to realize a classifier model with excellent separation performance between classes and fairness of separation between classes.
[0005] Also, in Non-Patent Documents 3-6, when training a classifier model using a dataset with uneven class sample numbers, although there are contrivances to improve the separation performance between classes, there are the following problems.
[0006] First, Non-Patent Documents 3 and 4 mainly assume the optimization of single-label multi-class classification, and it is difficult to determine and adjust the margin term to maximize any fairness index in the case of multi-labels. That is, since Non-Patent Documents 3 and 4 assume single-label problems, the assumption is different from that in the case of multi-labels, and the fairness index is not necessarily maximized in the case of multi-labels with the margin determined by this method.
[0007] In addition, Non-Patent Documents 5 and 6 mainly assume a segmentation task that classifies foreground and background. Therefore, it is difficult to determine and adjust the margin term so as to maximize any fairness metric in the case of single-label or multi-label multi-class classification tasks. That is, in Non-Patent Documents 5 and 6, since a margin is added with foreground = m (m is the margin term) / background = 0 in the segmentation task, it is difficult to allocate the margin term that maximizes the fairness metric in the case of multiple classes in the first place. Therefore, it is impossible to solve the existing problems only with the combination of non-patent documents.
[0008] An object of the present invention is to provide a learning device, a learning method, and a recording medium that can realize a classifier model excellent in separation performance between classes and fairness of separation between classes even when learning a classifier model using a data set with uneven numbers of samples in single-label or multi-label classes while solving the above-described problems.
Means for Solving the Problems
[0009] According to one aspect of the present invention, there is provided a learning device for learning a classifier model that performs single-label or multi-label multi-class classification on an image, the learning device including: a learning unit that learns the classifier model using, as an input, feature amounts extracted from learning images; and a margin adding unit that adds a margin to a loss function used for the learning. The margin adding unit fixes the total amount of the margin to be added for the single label or the multi label, and adds a class margin in which the total amount of the margin is asymmetrically distributed to a plurality of classes of the single label or the multi label.
[0010] According to another aspect of the present invention, there is provided a learning method for training a classifier model for single-label or multi-label multi-class classification of images, the method comprising: training the classifier model using, as input, feature amounts extracted from training images; and adding a margin to a loss function used in the training, wherein adding the margin comprises fixing a total amount of the margin added to the single label or the multi-label, and adding a class margin in which the total amount of the margin is asymmetrically distributed to a plurality of classes of the single label or the multi-label.
[0011] According to still another aspect of the present invention, there is provided a recording medium having recorded thereon a program for causing a computer to execute a learning method for training a classifier model for single-label or multi-label multi-class classification of images, the method comprising: training the classifier model using, as input, feature amounts extracted from training images; and adding a margin to a loss function used in the training, wherein adding the margin comprises fixing a total amount of the margin added to the single label or the multi-label, and adding a class margin in which the total amount of the margin is asymmetrically distributed to a plurality of classes of the single label or the multi-label.
Advantages of the Invention
[0012] According to the present invention, even when training a classifier model using a data set with uneven sample numbers for single-label or multi-label classes, it is possible to realize a classifier model excellent in separation performance between classes and fairness of separation between classes.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0014] [First Embodiment] The information processing apparatus and information processing method according to the first embodiment of the present invention will be described with reference to FIGS. 1 to 5.
[0015] First, the configuration of the information processing apparatus according to the present embodiment will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the information processing apparatus 1 according to the present embodiment. In the present embodiment, a case will be described in which the information processing apparatus 1 is a learning apparatus that learns a classifier model for performing multi-label multi-class classification on a face image by angular distance learning, which is deep distance learning using an angle. A classifier model for performing multi-label multi-class classification classifies a target face image into a plurality of classes for each of a plurality of labels. The number of labels is not particularly limited as long as it is a plurality of two or more, and the number of classes is not particularly limited as long as it is a plurality of two or more.
[0016] As shown in FIG. 1, the information processing apparatus 1 according to the present embodiment includes a processor 10, a memory 20, a storage 30, an input device 40, an output device 50, and an interface 60. The processor 10, the memory 20, the storage 30, the input device 40, the output device 50, and the interface 60 are connected to a common bus 70.
[0017] The processor 10 is, for example, a processor such as a CPU (Central Processing Unit) or an MPU (Micro-Processing Unit). The processor 10 operates by executing a program stored in the storage 30 or an external program via the interface 60, and functions as a control unit that controls the operation of the entire information processing apparatus 1. Further, the processor 10 executes an external program via the program stored in the storage 30 or the interface 60 to execute various processes as the information processing apparatus 1.
[0018] Specifically, when the information processing apparatus 1 functions as a learning apparatus, the processor 10 functions as an image acquisition unit 102, a feature extraction unit 104, a classifier learning unit 106, and a margin adding unit 108 as described later by executing a program. Note that the information processing apparatus 1 can also function as an estimation apparatus using a learned classifier model learned by functioning as a learning apparatus. In this case, the processor 10 functions as an image acquisition unit 102, a feature extraction unit 104, and an estimation unit 110 as described in the second embodiment by executing a program. The information processing apparatus 1 functioning as a learning apparatus and the information processing apparatus 1 functioning as an estimation apparatus may be the same apparatus as each other, or may be different apparatuses from each other. When functioning as a learning apparatus, the processor 10 does not necessarily have to function as an estimation unit 110. When functioning as an estimation apparatus, the processor 10 does not necessarily have to function as a classifier learning unit 106 and a margin adding unit 108.
[0019] The memory 20 is a main memory device composed of a volatile memory such as a RAM (Random Access Memory). The memory 20 provides a memory area necessary for the operation of the processor 10 and primarily stores programs executed by the processor 10, data referenced by the processor 10, and the like.
[0020] The storage 30 is an auxiliary storage device composed of, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), a ROM (Read Only Memory), or the like. The storage 30 stores programs executed by the processor 10, data referenced by the processor 10, and the like.
[0021] The storage 30 stores a learning database (DB, Database) 302 in which a plurality of face images with a sample number N are stored as learning face images. Note that the learning DB 302 may be stored in an external device such as a server that can be connected via the interface 60.
[0022] The input device 40 is, for example, a keyboard, a mouse, a touch panel, or the like. The input device 40 receives inputs such as instructions and set values from the user. The input device 40 may be a photographing device such as a digital camera. The output device 50 is, for example, a display, a printer, or the like. The output device 50 that is a display displays various screens such as a setting screen and an execution screen of a program executed by the processor 10.
[0023] The information processing apparatus 1 is connected to an external device such as an external storage device and a peripheral device, a network, or the like via the interface 60. The connection standard of the interface 60 is not particularly limited. Also, the connection method of the interface 60 may be a wired method or a wireless method.
[0024] In this way, the information processing apparatus 1 according to the present embodiment is configured. Note that the information processing apparatus 1 may be a general-purpose computer such as a personal computer or a server, or may be a computer designed specifically. Also, part or all of each function of the information processing apparatus 1 can also be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0025] Next, the learning method by the information processing apparatus 1 according to the present embodiment will be further described with reference to FIGS. 2 and 3. FIG. 2 is a flowchart showing the learning method executed by the information processing apparatus 1 according to the present embodiment. FIG. 3 is a schematic diagram showing the learning method executed by the information processing apparatus 1 according to the present embodiment.
[0026] The processor 10 functions as an image acquisition unit 102, a feature extraction unit 104, a classifier learning unit 106, and a margin adding unit 108 by executing a program stored in the storage 30 or an external program via the interface 60. Hereinafter, the case of learning a classifier model for performing C-class classification of A labels for face images will be described. Here, A is an integer of 2 or more, and C is an integer of 2 or more. For example, the classifier model is a model that performs multi-label multi-class classification on face images to determine face attributes. Specifically, the classifier model is, for example, a model that performs 2-class classification of 3 classes for face images. For example, the classifier model classifies the label of "male" into two classes of "being male" and "not being male", the label of "glasses" into two classes of "having glasses" and "not having glasses", and the label of "smiling face" into two classes of "being a smiling face" and "not being a smiling face". For simplicity, in the following mathematical formulas, the case where the number of classes is unified to C for all labels is considered, but C may be different for each label. In that case, C is replaced by C a and becomes two or more different integers for each label a.
[0027] As shown in FIGS. 2 and 3, the image acquisition unit 102 acquires a mini-batch including a plurality of face images with a batch sample number B from a learning database (DB, Database) 302 in which a plurality of face images with a sample number N are stored as learning face images (step S102).
[0028] Next, the feature extraction unit 104 extracts feature amounts for each face image included in the mini-batch acquired by the image acquisition unit 102 (step S104). The feature extraction unit 104 can extract feature amounts from face images using, for example, a pre-trained convolutional neural network (CNN, Convolutional Neural Network). In this case, the feature extraction unit 104 extracts, as the feature amount of the face image, a D-dimensional feature vector that is an intermediate feature amount output by an intermediate layer of the CNN for the input of the face image to the CNN. The feature extraction unit 104 can normalize the intermediate feature amount by L2 normalization. The intermediate layer of the CNN used for extracting the intermediate feature amount is not particularly limited, but is, for example, an intermediate layer of ResNet (see Kaiming He et al., "Deep Residual Learning for Image Recognition", Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770-778).
[0029] Next, the classifier learning unit 106 performs learning of the classifier model by angular distance learning using, as an input, the feature vector that is the feature amount of each face image. Specifically, it is as follows.
[0030] First, the classifier learning unit 106 calculates the cosine similarity between each of the feature vectors of the face images from the first to the B-th and the representative vector of each class for each of the labels from the first to the A-th (step S106). The classifier learning unit 106 can calculate the cosine similarity using fully connected layers FC1 to FC corresponding to each of the labels from the first to the A-th. A can be used to calculate the cosine similarity.
[0031] Next, the classifier learning unit 106 calculates the loss for each of the labels from the first to the A-th using the calculated cosine similarity by means of a Softmax-type loss function (step S108). The classifier learning unit 106 calculates the loss L a for the a-th label using the loss function represented by the following equation (1-1).
[0032]
Equation
[0033] In Equation (1-1), θ a,b,c is the angle formed by the feature vector of the b-th face image and the representative vector of the c-th class, which is the correct class, for the a-th label. a is an integer satisfying 1 ≤ a ≤ A. b is an integer satisfying 1 ≤ b ≤ B. c is an integer satisfying 1 ≤ c ≤ C. c′ is an integer satisfying c′ ≠ c and 1 ≤ c′ ≤ C. The three summation symbols represent, in order from the left side to the right side of the right side of Equation (1-1), the sum from 1 to B for b, the sum from 1 to C for c, and the sum for c′ of classes other than c. s is a hyperparameter of angular distance learning and is set to, for example, s = 10.
[0034] Also, in Equation (1-1), m a,c (γ a ) is the class margin, which is the margin defined for the c-th class of the a-th label. m a,c (γ a ) is set by the margin adding unit 108. The margin adding unit 108 subtracts m a,b,c from the cosine of the angle θ a,c (γ a ) formed by the feature vector extracted as a feature quantity from the face image and the representative vector of the class in the Softmax-type loss function and then adds it. When calculating the loss, the margin adding unit 108 calculates m a,c (γ a) is set and given to the loss function (step S110).
[0035]
Number
[0036] In formula (2), m a is the label margin, which is the total margin determined for the a-th label. However, m a does not necessarily have to be determined for each label, and a common value m for each label can also be used. α a,c (γ a ) distributes m a according to the number of samples of the c-th class of the a-th label, and is calculated by the following formula (3). Note that α a,c (γ a ) sums to 1 when calculated for c.
[0037]
Number
[0038] In formula (3), N a,c indicates the ratio of the number of samples of face images of the c-th class in the a-th label. N a,c″ indicates the ratio of the number of samples of face images of the c″-th class (where c″ is an integer satisfying 1 ≤ c″ ≤ C) in the a-th label. s is a hyperparameter, and the same one as s in formula (1-1) can be used. The summation symbol in the denominator on the right side of formula (3) means the sum for all integers c″ satisfying 1 ≤ c″ ≤ C. γ a is a parameter that determines by adjusting the intensity of the asymmetry between classes of the class margin. The value of γ a can be either a positive value or a negative value, but is a predetermined value that is not -∞ less than 0, for example.
[0039] In the case of two-class classification with C = 2, m a,c (γ a ) can be set by the following formula (2′).
[0040] [Number]
[0041] In formula (2'), α a,c (γ a ) is calculated by the following formula (3').
[0042] [Number]
[0043] Note that the classifier learning unit 106 can also calculate the loss L for the a-th label using a loss function represented by the following formula (1-2) or (1-3) instead of formula (1-1). In this case as well, the margin adding unit 108 can set m a (γ a,c ) in the same manner as above. a ) can be set.
[0044] [Number]
[0045] [Number]
[0046] In the case of formula (1-2), the margin adding unit 108 sets and adds m a,b,c (γ a,c ) so as to add it to the angle θ a formed by the feature vector extracted as a feature amount from the face image and the representative vector of the class in the Softmax-type loss function. In the case of formula (1-3), the margin adding unit 108 sets and adds m a,b,c (γ a,c ) so as to multiply it by the angle θ a formed by the feature vector extracted as a feature amount from the face image and the representative vector of the class in the Softmax-type loss function.
[0047] In addition, instead of the formula (1-1), the classifier learning unit 106 can also calculate the loss L for the a-th label using the loss function represented by the following formula (1-4) in which three class margins m1, m2, and m3 are used. a Formula (1-4) is a combination of formula (1-1), formula (1-2), and formula (1-3). In this case, for each of m1, m2, and m3, the margin adding unit 108 can be set in the same manner as the above m a,c (γ a ).
[0048]
Number
[0049] Next, the classifier learning unit 106 updates the parameters of the fully connected layer so that the loss L calculated for each label is minimized, and performs learning of the classifier model, and optimizes the parameters of the fully connected layers FC1 to FC a (step S112). For example, the classifier learning unit 106 can perform learning of the classifier model so that the loss L of all A labels calculated by the following formula (4) is minimized, and optimize the parameters of the fully connected layers FC1 to FC A . The summation symbol on the right side of formula (4) means the sum from 1 to A for a. A
[0050]
Number
[0051] Note that the processor 10 can repeatedly execute the processes from step S102 to step S112, and perform learning of the classifier model by mini-batch learning using a plurality of mini-batches. In addition, the processor 10 can also perform learning of the classifier model by batch learning that collectively processes a plurality of face images used for learning, or perform learning of the classifier model by online learning that sequentially processes each of the plurality of face images used for learning.
[0052] The margin adding unit 108 adds a label margin m input by the user via the input device 40 or the like. a and parameter γ a The user can set the label margin m a and parameter γ a By manually adjusting the metric, it is possible to optimize a fairness metric such as Balanced Accuracy, which is an evaluation metric for a classifier model. Note that any metric can be used as the fairness metric, and F1 score, Matthews Correlation Coefficient (MCC), etc. can also be used.
[0053] In addition, the classifier training unit 106 calculates the label margin m a and parameter γ a Instead of manually adjusting the label margin m a and parameter γ a can be used as a learnable parameter. In this way, the classifier training unit 106 can learn the label margin m a and parameter γ a It should be noted that the classifier learning unit 106 does not necessarily determine the label margin m a and parameter γ a There is no need to automatically determine both the label margin and the a and parameter γ a At least one of the above can be determined automatically.
[0054] Label margin m a and parameter γ a When automatically determining, the classifier learning unit 106 can add a constraint condition to the loss function to avoid convergence to a trivial solution (see Non-Patent Document 4). Specifically, the classifier learning unit 106, for example, determines the constraint condition L expressed by the following formula (5): m can be added to the loss expressed by equation (4).
[0055]
Mathematics
[0056] λ is a parameter for adjusting the intensity of L m and the larger λ is, the larger label margin m a will be taken. More precisely, λ can be set for each label, and λ of the a-th label can be incorporated into Equation (5) as λ a a.
[0057] Note that m a can be set to m, and m a can also be learned as a learnable parameter common to all labels. In this case, L m is represented by the following Equation (5′).
[0058]
Mathematics
[0059] The classifier learning unit 106 learns the classifier model as described above to generate a learned classifier model (step S114). The classifier learning unit 106 can store the generated classifier model in a storage device such as the storage 30 or an external storage.
[0060] In recent years, the importance of face authentication that does not depend on attributes such as race and gender, that is, fair face authentication, has been increasing. In order to establish fair authentication, there is also a need for a fair classifier for face attribute estimation that estimates attributes from face images. So far, in order to improve the separation performance of the classifier, a margin has been added to the loss function in angular distance learning. However, since it is difficult to prepare a dataset with a completely uniform number of samples as the samples used for learning, learning has to be performed using a non-uniform dataset including majority samples and minority samples. In such a case, simply adding a margin makes it difficult to realize a classifier with excellent separation performance between classes and fairness of separation between classes when training a classifier using a dataset with non-uniform class sample numbers. In particular, when training a classifier model that performs multi-label multi-class classification, learning may be biased towards simple classes with simple labels, hindering fair learning.
[0061] On the other hand, in this embodiment, the total margin of each label is fixed as the label margin m a , and based on the ratio of class samples, the label margin m a is asymmetrically distributed to the class margin m a by the parameter γ a,c that determines the intensity of asymmetry a .
[0062] Figures 4A and 4B are diagrams visually showing the class margins m a,0 , m a,1 set for classes 0 and 1 of label a, respectively. Figure 4A shows the case where the same margin is set for each class, and Figure 4B shows the case where the margins are set asymmetrically for each class according to this embodiment. W0 and W1 are the representative vectors of classes 0 and 1, respectively. x b is the feature vector extracted from sample b of the face image. As shown in Figure 4B, in this embodiment, the label margin m a is fixed, and the class margins m a where the label margin m a,0 is asymmetrically distributed, m a,1is set. In FIG. 4B, when Class 1 is in the minority, a larger class margin is given to Class 1 than to Class 0, which promotes learning in a compact manner within the class.
[0063] Thus, in this embodiment, the label margin m a is asymmetrically distributed to the class margin m a,c (γ a ). As a result, in this embodiment, it becomes possible to train the classifier model so as to maximize an index related to fairness such as Balanced Accuracy that does not explicitly appear in the loss.
[0064] Also, when automatically determining the label margin m a and the parameter γ a in this embodiment, as shown in Equation (5), the component that determines the asymmetry of the class margin is separated from the constraint condition L m . Therefore, in this embodiment, even if the total amount of label margins reaches the upper limit, the asymmetry of the class margin amounts within the label is determined separately, so that learning that does not impair fairness can be realized even in a multi-label format. The method according to this embodiment is different from Non-Patent Document 4 in that the component that determines the asymmetry of the class margin is separated from the constraint condition L m .
[0065] FIG. 5 is a diagram schematically showing the asymmetry of the class margin automatically determined according to this embodiment. In FIG. 5, for each of Labels 1 to 15, the class margins m0 and m1 of two classes, Class 0 and Class 1, are shown together with the label margin m a . As shown in the figure, the class margins m0 and m1 are asymmetrically determined by the asymmetric distribution of the label margin m a .
[0066] From the above, according to this embodiment, even when training a classifier using a dataset with non-uniform class sample numbers, it is possible to realize a classifier excellent in separation performance between classes and fairness of separation between classes.
[0067] [Second Embodiment] An information processing apparatus and an information processing method according to a second embodiment of the present invention will be described with reference to FIGS. 6 and 7. FIG. 6 is a schematic diagram showing an estimation method executed by the information processing apparatus according to the present embodiment. FIG. 7 is a flowchart showing the estimation method executed by the information processing apparatus according to the present embodiment.
[0068] In the present embodiment, a case will be described in which the information processing apparatus 1 shown in FIG. 1 functions as an estimation apparatus that estimates and classifies the class of a face image using the learned classifier model learned according to the first embodiment. Note that the information processing apparatus 1 functioning as a learning apparatus and the information processing apparatus 1 functioning as an estimation apparatus may be the same apparatus as each other, or may be different apparatuses from each other. The information processing apparatus 1 functioning as an estimation apparatus may not have a function as a learning apparatus.
[0069] The processor 10 functions as an image acquisition unit 102, a feature extraction unit 104, and an estimation unit 110 by executing a program stored in the storage 30 or an external program via the interface 60.
[0070] As shown in FIGS. 6 and 7, the image acquisition unit 102 acquires a face image to be estimated (step S202). The image acquisition unit 102 can also acquire a face image to be estimated stored in the storage 30 in advance from the storage 30, or can acquire a face image to be estimated from an external device via the interface 60. Further, the image acquisition unit 102 can also acquire a face image to be estimated by the input device 40 which is a photographing device.
[0071] Next, the feature extraction unit 104 extracts feature amounts from the face image to be estimated acquired by the image acquisition unit 102 using a CNN in the same manner as in the first embodiment (step S204).
[0072] Next, the estimation unit 110 estimates and classifies the class of each label for the face image to be estimated using the learned classifier model learned by the information processing apparatus 1 according to the first embodiment (step S206). That is, the estimation unit 110 uses the learned fully connected layers FC1 to FC A to calculate the cosine similarity. Next, the estimation unit 110 calculates the classification value of each class as a classification score from the cosine similarity using the Softmax-type function as the output layer.
[0073] In this way, the information processing apparatus 1 estimates and classifies the class of each label for the face image to be estimated.
[0074] [Third Embodiment] According to the third embodiment, the learning apparatus in which the information processing apparatus described in the above embodiment functions can also be configured as shown in FIG. 8. FIG. 8 is a block diagram showing the configuration of the learning apparatus according to the present embodiment.
[0075] As shown in FIG. 8, the learning apparatus 1000 according to the present embodiment is a learning apparatus that learns a classifier model for performing single-label or multi-label multi-class classification on images. The learning apparatus 1000 includes a learning unit 1002 that learns a classifier model using, as input, feature amounts extracted from learning images, and a margin adding unit 1004 that adds a margin to the loss function used for learning. The margin adding unit 1004 fixes the total amount of the margin to be added for single-label or multi-label, and adds a class margin in which the total amount of the margin is asymmetrically distributed to a plurality of classes of single-label or multi-label.
[0076] In the learning apparatus 1000 according to the present embodiment, a class margin in which the total amount of the margin is asymmetrically distributed to a plurality of classes is added. Therefore, according to the learning apparatus 1000 according to other embodiments, even when learning a classifier model using a data set in which the number of samples in each class is non-uniform, a classifier model excellent in separation performance between classes and fairness of separation between classes can be realized.
[0077] [Modified Embodiment] The present invention is not limited to the above-described embodiments, and various modifications are possible.
[0078] For example, in the above-described embodiment, the case of performing multi-label multi-class classification on a face image has been described, but the present invention is not limited thereto. The image for which multi-label multi-class classification is performed may be an object image including one or more objects. In this case, multi-label multi-class classification can be performed on one or more objects recognized in the image.
[0079] Further, in the above-described embodiment, the case of learning a classifier model that performs multi-label multi-class classification has been described, but the present invention is not limited thereto. The classifier model to be learned may perform single-label multi-class classification on an image such as a face image.
[0080] Further, in the above-described embodiment, the case of using a Softmax-type loss function as the loss function has been described, but the present invention is not limited thereto. As the loss function, various functions can be selected according to the object of estimation and the like, and a margin can be given to the loss function in the same manner as described above. As the loss function, in addition to the Softmax-type loss function and the cross-entropy error, the mean squared error, the mean absolute error, etc. can also be used.
[0081] A method of recording a program for operating the configuration of the embodiment so as to realize the functions of the above-described embodiment on a recording medium, reading the program recorded on the recording medium as code, and executing the program on a computer is also included in the scope of each embodiment. That is, a computer-readable recording medium is also included in the scope of each embodiment. Further, not only the recording medium on which the above-described program is recorded, but also the program itself is included in each embodiment.
[0082] As the recording medium, for example, a floppy (registered trademark) disk, a hard disk, an optical disk, a magneto-optical disk, a CD-ROM, a magnetic tape, a non-volatile memory card, etc. can be used. Further, not only those that execute processing with a program recorded on the recording medium alone, but also those that operate on an OS and execute processing in cooperation with the functions of other software and expansion boards are included in the scope of each embodiment.
[0083] Some or all of the above-described embodiments can be described as follows in the appended claims, but are not limited thereto.
[0084] (Appendix 1) A learning device that learns a classifier model for performing single-label or multi-label multi-class classification on an image, a learning unit that learns the classifier model by using, as an input, feature amounts extracted from learning images, and a margin adding unit that adds a margin to a loss function used for the learning, wherein the margin adding unit fixes a total amount of margins to be added for the single label or the multi-label, and adds a class margin in which the total amount of margins is asymmetrically distributed to a plurality of classes of the single label or the multi-label A learning device characterized by the above.
[0085] (Appendix 2) The classifier model performs multi-class classification for each label of the multi-label. The learning device according to Appendix 1, characterized by the above.
[0086] (Appendix 3) The learning unit learns the classifier model by angular distance learning. The learning device according to Appendix 1 or 2, characterized by the above.
[0087] (Appendix 4) The margin adding unit adds the class margin based on a ratio of samples of the class. The learning device according to any one of Appendices 1 to 3, characterized by the above.
[0088] (Appendix 5) The margin giving unit gives the class margin calculated by the following formula (1) to the loss function The learning device according to any one of Appendices 1 to 4, characterized in that
[0089] [Number] (In formula (1), m a is the total margin defined for the a-th label. α a,c (γ a ) is calculated by the following formula (2).
[0090] [Number] In formula (2), N a,c indicates the ratio of the number of samples of the c-th class in the a-th label. N a,c″ indicates the ratio of the number of samples of the c''-th class (where c'' is an integer satisfying 1 ≦ c'' ≦ C) in the a-th label. The summation symbol in the denominator on the right side means the sum over all integers satisfying 1 ≦ c'' ≦ C. s is a hyperparameter.)
[0091] (Appendix 6) The learning unit automatically determines at least one of the m a and the γ a The learning device according to Appendix 5, characterized in that
[0092] (Appendix 7) The loss function is a Softmax-type loss function The learning device according to any one of Appendices 1 to 6, characterized in that
[0093] (Appendix 8) The margin adding unit adds the class margin so as to subtract it from the cosine of the angle formed by the feature vector extracted as the feature amount from the image and the representative vector of the class in the Softmax type loss function. The learning device according to appended note 7, characterized by the above.
[0094] (Appended note 9) The margin adding unit adds the class margin so as to add it to the angle formed by the feature vector extracted as the feature amount from the image and the representative vector of the class in the Softmax type loss function. The learning device according to appended note 7, characterized by the above.
[0095] (Appended note 10) The margin adding unit adds the class margin so as to multiply it by the angle formed by the feature vector extracted as the feature amount from the image and the representative vector of the class in the Softmax type loss function. The learning device according to appended note 7, characterized by the above.
[0096] (Appended note 11) Having a feature extraction unit that extracts the feature amount by a convolutional neural network The learning device according to any one of appended notes 1 to 10, characterized by the above.
[0097] (Appended note 12) The image is a face image The learning device according to any one of appended notes 1 to 11, characterized by the above.
[0098] (Appended note 13) An image acquisition unit that acquires an image, An estimation unit that performs the multi-class classification on the image by the classifier model learned by the learning device according to any one of appended notes 1 to 12 An estimation device, characterized by having the above.
[0099] (Appended note 14) A learning method for training a classifier model that performs single-label or multi-label multi-class classification on images, using the feature quantities extracted from the learning images as inputs to train the classifier model, assigning a margin to the loss function used in the training, wherein assigning the margin includes fixing the total amount of the margin assigned to the single label or the multi-label, and assigning a class margin in which the total amount of the margin is asymmetrically distributed to a plurality of classes of the single label or the multi-label characterized by the learning method.
[0100] (Appendix 15) A recording medium recorded with a program for causing a computer to execute a learning method for training a classifier model that performs single-label or multi-label multi-class classification on images, using the feature quantities extracted from the learning images as inputs to train the classifier model, assigning a margin to the loss function used in the training, wherein assigning the margin includes fixing the total amount of the margin assigned to the single label or the multi-label, and assigning a class margin in which the total amount of the margin is asymmetrically distributed to a plurality of classes of the single label or the multi-label is provided.
[0101] Although the present invention has been described with reference to the embodiments above, the present invention is not limited to the above embodiments. Various modifications that can be understood by those skilled in the art within the scope of the present invention can be made to the configuration and details of the present invention.
Explanation of Signs
[0102] 1... Information processing apparatus 10... Processor 20... Memory 30... Storage 40... Input device 50... Output device 60... Interface 70... Common bus 102…Image acquisition unit 104…Feature extraction unit 106…Classifier learning unit 108…Margin addition unit 110…Estimation unit 1000…Learning device 1002…Learning unit 1004…Margin addition unit
Claims
1. A learning device for training a classifier model that performs single-label or multi-label multi-class classification on images, comprising: a learning unit that trains the classifier model using as input feature quantities extracted from learning images; a margin adding unit that adds a margin to a loss function used for the training; wherein the margin adding unit fixes the total amount of margin to be added for the single label or the multi-label, and adds a class margin in which the total amount of margin is asymmetrically distributed to a plurality of classes of the single label or the multi-label A learning device characterized by the above.
2. The classifier model performs multi-class classification for each label of the multi-label. The learning device according to claim 1, characterized by the above.
3. The learning unit trains the classifier model by angular distance learning. The learning device according to claim 1 or 2, characterized by the above.
4. The margin adding unit adds the class margin based on the ratio of samples of the class. The learning device according to any one of claims 1 to 3, characterized by the above.
5. The margin adding unit adds the class margin calculated by the following formula (1) to the loss function. The learning device according to any one of claims 1 to 4, characterized by the above. 【Number 1】 (In formula (1), m a is the total margin amount defined for the a-th said label. α a,c (γ a ) is calculated by the following formula (2). 【Number 2】 In formula (2), N a,c represents the ratio of the number of samples of the c-th class in the a-th label. N a,c″ represents the ratio of the number of samples of the c''-th class (where c'' is an integer satisfying 1 ≦ c'' ≦ C) in the a-th label. The summation symbol in the denominator on the right side means the sum over all integers c'' satisfying 1 ≦ c'' ≦ C. s is a hyperparameter.)
6. The learning unit automatically determines at least one of the m a and the γ a The learning device according to claim 5, characterized by the above.
7. The loss function is a Softmax-type loss function. The learning device according to any one of claims 1 to 6, characterized by the above.
8. An estimation device comprising: an image acquisition unit that acquires an image; an estimation unit that performs the multi-class classification on the image using the classifier model trained by the learning device according to any one of claims 1 to 7. Characterized by the above.
9. A learning method for training a classifier model that performs single-label or multi-label multi-class classification on images, comprising: a processor trains the classifier model using as input feature quantities extracted from learning images; the processor adds a margin to a loss function used for the training; wherein the processor adding the margin fixes the total amount of margin to be added for the single label or the multi-label, and adds a class margin in which the total amount of margin is asymmetrically distributed to a plurality of classes of the single label or the multi-label A learning method characterized by the above.
10. To a computer, A learning method for training a classifier model that performs single-label or multi-label multi-class classification on images, wherein the classifier model is trained using, as input, feature amounts extracted from learning images, a margin is imparted to a loss function used for the learning, and imparting the margin comprises imparting a class margin in which a total amount of the margin imparted for the single label or the multi-label is fixed and the total amount of the margin is asymmetrically distributed to a plurality of classes of the single label or the multi-label. A program for causing the above to be executed.