Image classification apparatus, image classification method, and image classification program
The image classification technique addresses the challenge of low classification accuracy for additional classes by using a pre-trained apparatus that generates accurate additional class prototypes from text and image features, improving performance even with a small number of images.
Patent Information
- Application Number
- JP2023189210
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-19
AI Technical Summary
Existing image classification technologies face challenges in achieving high classification accuracy for additional classes during incremental learning with a small number of images, leading to issues like catastrophic forgetting and overfitting.
The proposed image classification technique involves an apparatus pre-trained from text and images, which includes units for outputting image and text feature amounts, calculating prototypes, generating weights, and determining similarities to improve classification accuracy for additional classes.
This approach enhances the classification accuracy of additional classes by leveraging pre-trained models and calculating additional class prototypes based on text and image features, even with a small number of images.
Smart Images

Figure 2025077197000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image classification technology.
Background Art
[0002] Humans can learn new knowledge through long-term experience and can maintain the old knowledge without forgetting it. On the other hand, the knowledge of a deep neural network (DNN) using a convolutional neural network (CNN) etc. depends on the dataset used for learning, and in order to adapt to changes in the data distribution, it is necessary to relearn the parameters of the DNN for the entire dataset. In DNN, as learning progresses for a new task, the estimation accuracy for the old task decreases. Thus, in DNN, when continuous learning is performed, catastrophic forgetting, in which the learning results of the old task are forgotten during the learning of the new task, cannot be avoided.
[0003] As a method for avoiding catastrophic forgetting, incremental learning or continual learning has been proposed. Incremental learning or continual learning is a learning method in which, when a new task or new data occurs, instead of learning the model from scratch, the currently learned model is improved and learned.
[0004] Also, humans can learn new knowledge from a small number of images. On the other hand, artificial intelligence using deep learning using a convolutional neural network etc. depends on big data (a large number of images) used for learning. It is known that when artificial intelligence using deep learning is learned with a small number of images, it falls into overfitting where the local performance is good but the generalization performance is poor.
[0005] As a method for avoiding overfitting, few shot learning has been proposed. Few shot learning is a learning method that uses big data in a basic task to learn basic knowledge, and then uses the basic knowledge to learn new knowledge from a small number of images of a new task.
[0006] There is few shot class incremental learning as a method for solving the problems of both continuous learning and few shot learning, and there is a technique that uses an averaged feature vector as a weight vector (Patent Document 1). Also, there is a technique for matching the feature vector of text and the feature vector of an image (Non-Patent Document 1).
Prior Art Documents
Non-Patent Documents
[0007]
Patent Document 1
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] In the prior art, there was a problem that the classification accuracy of images of additional classes did not become sufficiently high for additional learning with a small number of images.
[0009] The present invention has been made in view of such a situation, and an object thereof is to provide an image classification technique capable of improving the classification accuracy of images of additional classes for additional learning with a small number of images.
Means for Solving the Problems
[0010] To solve the above problems, an image classification apparatus according to an aspect of the present invention has been pre-trained from text and images, and includes an image feature amount output unit that takes an image as an input and outputs an image feature amount, an image prototype output unit that calculates the image feature amount for each class and outputs an image prototype for each class, a text feature amount output unit that has been pre-trained from text and images, takes a text describing a class as an input, and outputs a text feature amount, a basic class image prototype selection unit that selects an image prototype of a basic class in the vicinity of the text feature amount of the additional class, a weight generation unit that generates a weight of the additional class from the image feature amount of the additional class, the image prototype of the basic class, and the text feature amount of the additional class, holds the image prototype of the basic class output by the image prototype output unit as a weight of the basic class, further holds the weight of the additional class generated by the weight generation unit, takes the image feature amount output from the image feature amount output unit as an input, and calculates a similarity, and a classification unit that takes the similarity as an input and determines the classification of the image. Another aspect of the present invention is also an image classification apparatus. This apparatus has been pre-trained from text and images, and includes an image feature amount output unit that takes an image as an input and outputs an image feature amount, an image prototype output unit that calculates the image feature amount for each class and outputs an image prototype for each class, a text feature amount output unit that has been pre-trained from text and images, takes a text describing a class as an input, and outputs a text feature amount, a basic class text feature amount selection unit that selects a text feature amount of a basic class in the vicinity of the text feature amount of the additional class, a weight generation unit that generates a weight of the additional class from the image feature amount of the additional class, the text feature amount of the basic class, and the text feature amount of the additional class, holds the image prototype of the basic class output from the image prototype output unit as a weight of the basic class, further holds the weight of the additional class generated by the weight generation unit, takes the image feature amount output from the image feature amount output unit as an input, and calculates a similarity, and a classification unit that takes the similarity as an input and determines the classification of the image. Yet another aspect of the present invention is also an image classification device. This device has been pre-trained from text and images, and includes an image feature amount output unit that takes an image as input and outputs an image feature amount, an image prototype output unit that calculates the image feature amount for each class and outputs an image prototype for each class, a text feature amount output unit that has been pre-trained from text and images, takes a text that describes a class as input, and outputs a text feature amount, a basic class prototype selection unit that selects an image prototype of a basic class and a text feature amount of the basic class that are in the vicinity of the text feature amount of the additional class, a weight generation unit that generates a weight for the additional class from the image feature amount of the additional class, the image prototype of the basic class, the text feature amount of the basic class, and the text feature amount of the additional class, a similarity calculation unit that holds the image prototype of the basic class output from the image prototype output unit as a weight of the basic class, further holds the weight of the additional class generated by the weight generation unit, takes the image feature amount output from the image feature amount output unit as input, and calculates a similarity, and a classification unit that takes the similarity as input and determines the classification of the image.
[0011] Another aspect of the present invention is an image classification method. This method includes an image feature amount output step that has been pre-trained from text and images, takes an image as input, and outputs an image feature amount; an image prototype output step that calculates the image feature amount for each class and outputs an image prototype for each class; a text feature amount output step that has been pre-trained from text and images, takes text explaining a class as input, and outputs a text feature amount; a basic class image prototype selection step that selects an image prototype of a basic class near the text feature amount of an additional class; a weight generation step that generates a weight for the additional class from the image feature amount of the additional class, the image prototype of the basic class, and the text feature amount of the additional class; a step of holding the image prototype of the basic class output by the image prototype output step as a weight of the basic class, further holding the weight of the additional class generated by the weight generation step, taking the image feature amount output from the image feature amount output step as input, and calculating a similarity; and a classification step that takes the similarity as input and determines the classification of the image.
[0012] Note that any combination of the above components, as well as conversions of the expression of the present invention between a method, an apparatus, a system, a recording medium, a computer program, etc., are also effective as aspects of the present invention.
Effects of the Invention
[0013] According to the present invention, it is possible to provide an image classification technique capable of improving the classification accuracy of images of additional classes for additional learning with a small number of images.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Embodiments for Carrying Out the Invention
[0015] First, the basic training dataset and the additional training dataset will be described.
[0016] The basic training dataset includes a large number of basic classes (for example, about 100 to 1000 classes), and each basic class is a supervised dataset composed of a large number of training images (for example, 500 to 3000 images). It is assumed that the basic training dataset has a sufficient amount of data to learn a general classification task alone. Here, the number of basic classes is 60, and the number of training images for each basic class is 500.
[0017] On the other hand, the additional training dataset includes a small number of additional classes (for example, about 1 to 10 classes), and each additional class is a supervised dataset composed of a small number of training images (for example, about 1 to 10 images). Also, here it is assumed to be a small number of images, but it can also be a large number of images if the number of additional classes is small. Here, the number of additional classes is 5, and the number of training images for each additional class is 5.
[0018] The class names of the basic classes and the additional classes are represented by text (sentences).
[0019] In class incremental learning, the classes learned first are called basic classes (known classes), and the classes added later (unknown classes) are called additional classes. When learning the basic classes, the basic dataset is used to learn the basic classes. When learning the additional classes, the additional dataset is used to add the additional classes to the classes learned so far.
[0020] In the classification (inference) stage, it is possible to classify the basic classes and all the additional classes simultaneously.
[0021] There are two methods during additional learning: one is to re-train the image feature extraction unit, and the other is not to re-train the image feature extraction unit. When not re-training the image feature extraction unit, it is relatively easy to maintain the accuracy of the basic classes, but there are challenges in improving the accuracy of the additional classes. Furthermore, when the number of images of the additional classes is small, it becomes more difficult to improve the accuracy of the additional classes.
[0022] An object of the image classification device 100 according to an embodiment of the present invention is to improve the class classification accuracy of additional classes without re-training the image feature extraction unit when the number of images of the additional classes is small.
[0023] (Example) FIG. 1 is a configuration diagram of an image classification device 100 according to an example. The image classification device 100 includes an image feature amount output unit 10, an image prototype generation unit 20, a similarity calculation unit 30, a classification unit 40, a text feature amount output unit 50, a text prototype generation unit 60, an additional class prototype generation unit 70, and a weight generation unit 80.
[0024] The image feature amount output unit 10 is a pre-trained neural network model learned from text and images. It takes an image as input and outputs an image feature amount.
[0025] The image feature amount output unit 10 is composed of a deep neural network and calculates an image feature vector (image feature amount) of the input image. The image feature amount output unit 10 outputs the image feature vector to the image prototype generation unit 20 and the weight generation unit 80.
[0026] Here, it is assumed that the weight parameters of the image feature amount output unit 10 are learned before the data set of the basic classes using an image data set with a sufficient amount of data such as ImageNet. Also, it is assumed that the weight parameters of the image feature amount output unit 10 are fixed after being fine-tuned with the data set of the basic classes.
[0027] Here, ResNet-18 is used as the image feature output unit 10. The dimensionality of the feature vector of ResNet-18 is 512. The image feature output unit 10 is not limited to ResNet-18, and may also be a deep neural network such as ViT-B / 32 (the dimensionality of the image features is 512), ViT-L / 14 (the dimensionality of the image features is 768) of ViT (Vision Transformer), RN50 (the dimensionality of the image features is 1024), RN101 (the dimensionality of the image features is 512), RN50x4 (the dimensionality of the image features is 640), etc. which are extended models of ResNet.
[0028] The image prototype generation unit 20 calculates (for example, obtains the average value, median value, representative value, etc.) the image features for each class and outputs the image prototype for each class.
[0029] The similarity calculation unit 30 is composed of a fully connected layer, and the fully connected layer has weight vectors for multiple classes. That is, the weight vectors exist for each class.
[0030] The similarity calculation unit 30 calculates the cosine similarity, which is the similarity between the input image features and the weight vectors of multiple classes. The similarity calculation unit 30 outputs the cosine similarities of multiple classes to the classification unit 40.
[0031] The weight vectors of the fully connected layer include both the basic classes and the additional classes. Here, the weight vector of each class is the representative vector of the image features of each class, and is also called a prototype. Here, the prototype is the average of the image features.
[0032] It is assumed that the weight vectors of the basic classes of the similarity calculation unit 30 are pre-calculated and fixed using the dataset of the basic classes.
[0033] The image feature amount output unit 10 takes as input all the training images of the basic classes included in the basic training dataset, and outputs the image feature amounts of all the training images of the basic classes. The image prototype generation unit 20 takes as input the image feature amounts of all the training images of the basic classes, averages the image feature amounts for each class, and outputs the respective weight vectors (image prototypes) of the basic classes.
[0034] The similarity calculation unit 30 receives the weight vectors of the basic classes from the image prototype generation unit 20. On the other hand, the weight vectors of the additional classes of the similarity calculation unit 30 are input from the weight generation unit 80. The specific addition method will be described later.
[0035] The similarity calculation unit 30 holds the image prototypes of the basic classes output from the image prototype generation unit 20 as the weights of the basic classes, holds the additional class prototypes output from the weight generation unit 80 as the weights of the additional classes, takes as input the image feature amounts output from the image feature amount output unit 10, and calculates the similarity.
[0036] The classification unit 40 selects the class having the maximum similarity from the similarities calculated by the similarity calculation unit 30. The similarity is, for example, the cosine similarity, and the cosine similarity indicates the similarity for each class. The classification unit 40 selects the class having the maximum cosine similarity.
[0037] The text feature amount output unit 50 is a pre-trained neural network model learned from text and images. It takes as input text that describes a class and outputs the text feature amount.
[0038] The text feature amount output unit 50 is composed of a deep neural network and calculates a text feature vector (text feature amount) from the input text related to the class name.
[0039] The text feature amount output unit 50 outputs the text feature vector to the text prototype generation unit 60.
[0040] The text prototype generation unit 60 calculates text feature amounts for each class (for example, calculates the average value, median value, representative value, etc.) and outputs the text prototype of each class.
[0041] Here, when there is only one piece of text related to a class, since the text feature amount becomes the text prototype as it is, the configuration of the text prototype generation unit 60 may be omitted. In that case, note that the text feature amount output from the text feature amount output unit 50 is directly input to the additional class prototype generation unit 70 as the document prototype as it is.
[0042] As the neural network model of the text feature amount output unit 50, a Transformer model, which is a text encoder learned in a common feature space of an image and text as shown in Non-Patent Document 1, is used. Therefore, the text feature amount output by the text feature amount output unit 50 can be matched with the image feature amount.
[0043] In Non-Patent Document 1, an image encoder and a text encoder are learned so that the image feature vector and the text feature vector of a certain class can be matched in the feature space.
[0044] It is said that the text encoder and the image encoder are not good at detailed classification of classes but are good at conceptual classification.
[0045] An image encoder is generally a very large-scale neural network and is trained with a large-scale dataset. On the other hand, the text encoder does not have to be as large-scale a neural network as the image encoder.
[0046] Here, the conceptual classification ability of the text encoder, which requires a small implementation scale, is used for few-shot learning to improve the class classification accuracy.
[0047] Here, it is assumed that the dimension of the feature vector output by the image feature amount output unit 10 is the same 512 as the dimension of the feature vector output by the text feature amount output unit 50. By matching the number of dimensions, it becomes possible to directly match without reducing the number of dimensions, etc., thereby improving the matching accuracy.
[0048] The prototype calculated based on the feature vector output by the image feature amount output unit 10 is called an "image prototype", and the prototype calculated based on the feature vector output by the text feature amount output unit 50 is called a "text prototype".
[0049] The additional class prototype generation unit 70 generates an additional class prototype using the image prototype of the basic class and the text prototype of the basic class.
[0050] More specifically, the additional class prototype generation unit 70 calculates the difference between the text prototype of the basic class near the text prototype of the additional class and the image prototype of the same basic class as the text prototype of the basic class near the text prototype of the additional class. Further, the additional class prototype generation unit 70 adds this difference to the text feature amount of the additional class to generate an additional class prototype. Here, there may be one or more neighborhoods. Therefore, this difference may be the difference from only one neighborhood or the average of the differences from a plurality of neighborhoods. Note that the neighborhood can be rephrased as within a predetermined distance.
[0051] The additional class prototype generation unit 70 outputs the additional class prototype to the weight generation unit 80.
[0052] The weight generation unit 80 takes the image feature amount of the additional class and the additional class prototype as inputs, and calculates the additional class weight vector of the similarity calculation unit 30 by calculating the image feature amount of the additional class and the additional class prototype.
[0053] FIG. 2 is a flowchart for explaining the processing procedure of additional learning by the image classification apparatus 100.
[0054] The text feature quantity output unit 50 receives a text that describes the basic class. If there is no appropriate text that describes the basic class, the class name of the basic class may be used as the text. The text feature quantity output unit 50 calculates the text feature quantity of the basic class from the input text, and the text prototype generation unit 60 generates the text prototype of the basic class from the text feature quantity of the basic class, and outputs the text prototype of the basic class to the additional class prototype generation unit 70 (S10).
[0055] The image feature quantity output unit 10 takes the image of the basic class as input and outputs the image feature quantity. The image prototype generation unit 20 takes the image feature quantity as input and outputs the image prototype (weight vector) of the basic class to the additional class prototype generation unit 70 (S12).
[0056] The additional class prototype generation unit 70 holds the text prototype of the basic class and the image prototype (weight vector) of the basic class (S14).
[0057] Hereinafter, a processing procedure for additional learning to enable classification of additional classes after the additional classes are given will be described. No major processing such as optimization is required for additional learning, and additional learning can be repeated.
[0058] Here, for the sake of simplicity, it is assumed that additional classes are added one by one, but a plurality of them may be added at once.
[0059] As data of the additional class, a text that describes the additional class j and K images are given. Here, K is an arbitrary integer of 1 or more. If there is no appropriate text that describes the additional class, the class name of the additional class may be used as the text. The text is input to the text feature quantity output unit 50, and the K images are input to the image feature quantity output unit 10.
[0060] When K images are input, the image feature quantity output unit 10 calculates K image feature quantities FVn_Img(j,k) of the additional class j and provides them to the weight generation unit 80 (S16). Here, k = 0, 1, 2, ···, K - 1.
[0061] When a sentence explaining the additional class is input, the sentence feature quantity output unit 50 calculates a sentence prototype PVn_Com(j) of the additional class j and provides it to the additional class prototype generation unit 70 (S18).
[0062] The additional class prototype generation unit 70 selects M sentence prototypes of the basic classes that are in the vicinity of the sentence prototype of the additional class j (for example, M = 3) (S20). Here, M is an arbitrary integer of 1 or more.
[0063] For the selected M basic classes i, the additional class prototype generation unit 70 calculates a movement vector MVb(i) from the sentence prototype PVb_Com(i) to the image prototype PVb_Img(i) as follows (S22). Here, i = 1, 2, ···, M. MVb(i)=PVb_Img(i)-PVb_Com(i)
[0064] The additional class prototype generation unit 70 calculates an average movement vector MVb_ave of the basic classes of the movement vectors of the M basic classes as follows (S24). MVb_ave=ΣMVb(i) / M
[0065] The additional class prototype generation unit 70 adds the average movement vector of the basic classes to the sentence prototype of the additional class j, calculates an additional class prototype PPVn_Img(j) of the additional class j as follows, and provides it to the weight generation unit 80 (S26). PPVn_Img(j)=PVn_Com(j)+MVb_ave
[0066] FIG. 3 is a diagram for explaining an example of calculation of the additional class prototype.
[0067] First, select the text prototypes (black circles) of the basic classes B, C, and E that are in the vicinity of the text prototype of the additional class j. Next, calculate the movement vectors MVb(B), MVb(C), and MVb(E) from the text prototypes of the basic classes B, C, and E to the image prototypes. Average MVb(B), MVb(C), and MVb(E) to calculate the average movement vector MVb_ave. Add the movement vector to the text prototype of the additional class j to calculate the additional class prototype of the additional class j.
[0068] FIG. 4 is a diagram for explaining an operation example of the weight generation unit 80.
[0069] The weight generation unit 80 averages the K image feature amounts of the additional class j input from the image feature amount output unit 10 and the additional class prototype of the additional class j input from the additional class prototype generation unit 70, and calculates the image prototype PVn_Img(j) of the additional class j as follows, and outputs the image prototype of the additional class j to the similarity calculation unit 30 (S28). PVn_Img(j)=(ΣFVn_Img(j,k)+PPVn_Img(j)) / (K + 1)
[0070] The similarity calculation unit 30 adds the image prototype of the additional class j as the weight vector of the additional class j to the fully connected layer (S30).
[0071] As a result, the image classification device 100 can classify the additional class j in addition to the basic classes.
[0072] As another example, for example, if it is known that the accuracy of the text prototype of the additional class j output by the text feature amount output unit 50 is high, as shown in the following formula, the proportion (α in the following formula) of the additional class prototype can be increased to calculate the weighted average image prototype of the additional class j. For example, set α to be greater than 1, such as α = 1.2. PVn_Img(j)=ΣFVn_Img(j,k) / K+α×PPVn_Img(j)
[0073] Also, here, the average of the motion vectors of neighboring classes is used, but for example, statistical quantities other than the average such as the median, maximum value, minimum value, etc. of the motion vectors of neighboring classes may be used.
[0074] Next, the classification (inference) processing procedure by the image classification apparatus 100 after the additional classes can be classified will be described.
[0075] FIG. 5 is a flowchart for explaining the classification processing procedure by the image classification apparatus 100.
[0076] The input image is input to the image feature amount output unit 10. The image feature amount output unit 10 calculates the image feature amount of the input image and gives it to the similarity calculation unit 30 (S40).
[0077] The similarity calculation unit 30 calculates the similarity between the input image feature amount and the weight vectors of all classes, and gives the similarities of all classes to the classification unit 40 (S42).
[0078] The classification unit 40 selects the class having the maximum similarity from among the similarities of all classes (S44). Thereby, the class of the input image is determined.
[0079] As described above, the image classification apparatus 100 according to the embodiment calculates an additional class prototype by using an additional class sentence prototype, an average movement vector calculated from a basic class sentence prototype and a basic class image prototype. Further, an image prototype of the additional class is calculated as a weight vector of the similarity calculation unit 30 by averaging the additional class prototype and K image feature amounts of the additional class. Thereby, even when the image data is a small number of additional classes, by using the highly accurate basic class image prototype calculated from a large number of images, a highly accurate additional class image prototype can be obtained. In addition, the image classification apparatus 100 can classify the additional class with high accuracy in addition to the basic class by weighted-averaging the additional class image prototypes according to the accuracy of the additional class sentence prototypes output by the sentence prototype generation unit 60.
[0080] In the above embodiment, the additional class prototype generation unit 70 selects, as a sentence prototype in the vicinity of the sentence prototype of the additional class j, a basic class that has been learned with a sufficient amount of data. However, an additional class input before the additional class j may be selected as a sentence prototype in the vicinity.
[0081] (Modification Example 1) Here, a simpler configuration of the image classification apparatus 100 of the embodiment will be described. Specifically, a modification example in the case where there is only a class name (label) in the additional class and there is no image data will be described.
[0082] FIG. 6 is a diagram for explaining the image classification apparatus 100 of Modification Example 1. The difference from the embodiment is that there is no weight generation unit 80 and there is no processing flow related to the weight generation unit 80. Thereby, even when there is no image data of the additional class, the additional class can be learned by using the class name of the additional class as sentence data.
[0083] When the feature space of the text feature quantity output unit 50 and the feature space of the image feature quantity output unit 10 are learned to be similar feature spaces, even when there is no image data of the additional class, the additional class can be learned by using the text prototype generated from the class name of the additional class.
[0084] The similar feature space means, for example, that the feature spaces of the text feature quantity output unit 50 and the image feature quantity output unit 10 have the same number of dimensions. When the text feature quantity output by the text feature quantity output unit 50 and the image feature quantity output by the image feature quantity output unit 10 in a certain class are mapped to the same feature space, the distance between the text feature quantity and the image feature quantity becomes close.
[0085] Even for an additional class that has only a class name and no image, by using the additional class prototype, the image classification device 100 can classify the additional class j in addition to the basic class.
[0086] In Modification 1, there is no weight generation unit 80. The additional class prototype generation unit 70 outputs the additional class prototype of the additional class j to the similarity calculation unit 30 as the weight vector of the additional class j. As described above, the text feature quantity and the image feature quantity have the same number of dimensions, and when the text feature quantity and the image feature quantity in a certain class are mapped to the same feature space, the distance becomes close. Therefore, the relationship between the text prototype and the image prototype of the additional class j generated from the text feature quantity and the image feature quantity respectively is the same. Thus, the additional class prototype generation unit 70 can generate an additional class prototype using the text prototype of the additional class j and the image prototype of the basic class.
[0087] FIG. 7 is a flowchart for explaining the additional learning processing procedure by the image classification device 100 of Modification 1. The difference is that steps S16 and S28 in FIG. 2 are omitted, and step S30 is replaced by step S32. Since the other processing procedures are the same as those in FIG. 2, the description of the common processing procedures is omitted, and only the differences are described.
[0088] In step S32, the similarity calculation unit 30 adds the additional class prototype of the additional class j to the fully connected layer as the weight vector of the additional class j.
[0089] (Modification 2) Here, a more simplified configuration of Modification 1 will be described. FIG. 8 is a configuration diagram of the image classification device 100 of Modification 2. The difference from Modification 1 is that there is no additional class prototype generation unit 70. As a result, the text prototype output from the text prototype generation unit 60 (when there is only one text related to the class (for example, only the class name), the text feature amount output from the text feature amount output unit 50) is used as the weight vector of the similarity calculation unit 30.
[0090] Modification 2 can be applied when the feature space of the text feature amount output unit 50 and the feature space of the image feature amount output unit 10 are learned to be close feature spaces, and the text prototype output from the text prototype generation unit 60 can be directly used as the weight vector of the similarity calculation unit 30. Therefore, the configuration can be simplified and the processing load of the image classification device 100 can be reduced.
[0091] Also, even when there is no image data of the additional class, additional learning can be performed only with the class name of the additional class. Furthermore, since it is not necessary to consider parts such as the background that have no relation to the class name in the image, the accuracy of the weight vector can be improved.
[0092] FIG. 9 is a flowchart for explaining the additional learning processing procedure by the image classification device 100 of Modification 2. Steps S10, S12, S14, S16, S20, S22, S24, S26, and S28 in the additional learning processing procedure of the embodiment in FIG. 2 are omitted, step S18 is replaced with step S19, and step S30 is replaced with step S34.
[0093] When the text of the additional class is input, the text feature quantity output unit 50 outputs the text feature quantity to the text prototype generation unit 60. The text prototype generation unit 60 calculates the text prototype PVn_Com(j) of the additional class j from the text feature quantity and provides it to the similarity calculation unit 30 (S19).
[0094] The similarity calculation unit 30 adds the text prototype of the additional class j as the weight vector of the additional class j to the fully connected layer (S34).
[0095] (Modification Example 3) Here, a simple configuration of the embodiment will be described. FIG. 10 is a diagram for explaining the image classification device 100 of Modification Example 3. The difference from the embodiment of FIG. 1 is that there is no additional class prototype generation unit 70.
[0096] The weight generation unit 80 averages the text prototype output from the text prototype generation unit 60 (when there is only one text related to the class (for example, only the class name), the text feature quantity output from the text feature quantity output unit 50) and the image feature quantity output from the image feature quantity output unit 10 to generate the weight vector of the similarity calculation unit 30.
[0097] The feature space of the text feature quantity output unit 50 and the feature space of the image feature quantity output unit 10 are learned to be close feature spaces. When there is image data of the additional class, by using the text prototype of class j output using a text that is high-precision but conceptual information and the image feature quantity of the additional class j output using an image that is specific information, the weight generation unit 80 can generate an appropriate weight vector of the additional class for the similarity calculation unit 30.
[0098] FIG. 11 is a flowchart for explaining the processing procedure of additional learning by the image classification apparatus 100 according to Modification 3. Steps S10, S12, S14, S20, S22, S24, and S26 of the processing procedure of additional learning in the embodiment of FIG. 2 are omitted, and the difference is that step S28 is replaced by step S29. Since the rest is the same as the processing procedure of FIG. 2, the explanation of the common processing procedure is omitted, and only the differences are explained.
[0099] In step S29, the weight generation unit 80 averages the K image feature amounts of the additional class j input from the image feature amount output unit 10 and the text prototype (when there is only one text related to the class, the text feature amount output from the text feature amount output unit 50) input from the text prototype generation unit 60 to calculate the image prototype PVn_Img(j) of the additional class j, and outputs the image prototype of the additional class j to the similarity calculation unit 30.
[0100] FIG. 12 is a diagram for explaining an operation example of the weight generation unit 80 according to Modification 3.
[0101] The weight generation unit 80 averages the K image feature amounts FVn_Img(j,k) of the additional class j input from the image feature amount output unit 10 and the text prototype PVn_Com(j) of the additional class j input from the text prototype generation unit 60 to calculate the image prototype PVn_Img(j) of the additional class j as follows. PVn_Img(j)=(ΣFVn_Img(j,k)+PVn_Com(j)) / (K + 1)
[0102] (Modification 4) Here, a modification example of the additional class prototype generation unit 70 will be described. Assume that a class higher than a given class is defined for the dataset used in the learning of the basic class.
[0103] FIG. 13 is a diagram for explaining the image classification apparatus 100 of Modification 4. The difference from the embodiment is that the additional class prototype generation unit 70 is given upper class information. In the embodiment, the sentence prototypes of the basic classes in the vicinity of the sentence prototype of the additional class are used for calculating the additional class prototype. In Modification 4, the upper class is defined in advance, and the sentence prototypes of the basic classes belonging to the same upper class as the additional class are used for calculating the additional class prototype.
[0104] When comparing images of different classes with the same upper class, they have similar features. Therefore, when the deviation between the image prototype and the sentence prototype is large, the accuracy can be improved by calculating the additional class prototype using the image prototype of another class belonging to the same upper class rather than the correlation based on the sentence prototype.
[0105] Note that Modification 1 may be applied to Modification 4 to omit the weight generation unit 80.
[0106] FIG. 14 is a flowchart for explaining the additional learning processing procedure by the image classification apparatus 100 of Modification 4. The difference is that step S20 in the additional learning processing procedure of the embodiment in FIG. 2 is replaced with step S21, and the rest is the same as the processing procedure in FIG. 2. Therefore, the common processing procedures will be omitted and only the differences will be explained.
[0107] In step S21, the additional class prototype generation unit 70 selects the sentence prototypes of M basic classes belonging to the upper class of the additional class j.
[0108] FIG. 15 is a diagram for explaining an example of the upper class. The upper class (Superclass) is a superordinate concept of the classes (Classes). FIG. 15 shows an example of the CIFAR100 dataset. For example, if the upper class is aquatic mammals and the additional class is a dolphin, the basic classes belonging to the same upper class are beaver, otter, seal, and whale. For example, aquatic mammals have similar features as images, such as having a tail fin.
[0109] FIG. 16 is a diagram for explaining the calculation of the prototype of the additional class by the additional class prototype generation unit 70.
[0110] Let the basic classes, which are the same upper-class as the additional class j, be class A', class B', class B', class C', and class D'. Let the movement vectors from the sentence prototype to the image prototype be MVA', MVB', MVC', and MVD' in the order of the above basic classes, respectively. The additional class prototype generation unit 70 adds the average movement vector, which is the average of MVA', MVB', MVC', and MVD', to the sentence prototype of the additional class j to calculate the prototype of the additional class j.
[0111] For example, assume that the additional class j is a beaver, and that the dolphin, otter, seal, and whale included in the upper-class aquatic mammals are included in the basic classes. Then, the prototypes of the dolphin, otter, seal, and whale become A', B', C', and D', respectively.
[0112] As another example regarding the upper-class, as shown in FIG. 17, the upper-class may be defined using the classification levels of organisms generally defined, such as order, family, genus, and species.
[0113] When using classes with the same upper-class, it can be expected that the closer the upper-class is to the lower classification level, the closer the prototypes are located.
[0114] Also, here, the average of the movement vectors of classes with the same upper-class is used, but for example, the median, maximum value, minimum value, etc. of the movement vectors may be used.
[0115] (Modification Example 5) FIG. 18 is a configuration diagram of the image classification device 100 according to Modification Example 5. The configuration of Modification Example 5 is different from that of the embodiment in that the additional class prototype generation unit 70 is replaced by the basic class sentence prototype selection unit 72, and the operation of the weight generation unit 80 is different.
[0116] The basic class sentence prototype selection unit 72 selects a basic class sentence prototype in the vicinity of the additional class sentence prototype, and provides the basic class sentence prototype and the additional class sentence prototype to the weight generation unit 80.
[0117] The weight generation unit 80 averages the K image feature amounts of the additional class input from the image feature amount output unit 10, the L basic class sentence prototypes input from the basic class sentence prototype selection unit 72, and the 1 additional class sentence prototype input from the basic class sentence prototype selection unit 72 to calculate an image prototype of the additional class, which is used as the weight vector of the additional class for the similarity calculation unit 30. Here, L is an arbitrary integer greater than or equal to 1.
[0118] FIG. 19 is a flowchart for explaining the additional learning processing procedure by the image classification apparatus 100 of Modification 5. In Modification 5, steps S10, S12, and S14 of the additional learning processing procedure of the embodiment in FIG. 2 are replaced with steps S11 and S15 in FIG. 19, and steps S20, S22, S24, S26, and S28 in FIG. 2 are replaced with steps S23 and S27 in FIG. 19, which is different. Since the rest is the same as the processing procedure in FIG. 2, the description of the common processing procedure is omitted, and only the differences are described.
[0119] In step S11, the sentence prototype generation unit 60 outputs a basic class sentence prototype to the basic class sentence prototype selection unit 72.
[0120] In step S15, the additional class prototype generation unit 70 holds the basic class sentence prototype.
[0121] In step S23, the basic class sentence prototype selection unit 72 selects the sentence prototype of the additional class j and L basic class sentence prototypes in the vicinity of the sentence prototype of the additional class j.
[0122] In step S27, the weight generation unit 80 averages the K image feature amounts of the additional class j input from the image feature amount output unit 10, the L text prototypes of the basic classes input from the basic class text prototype selection unit 72, and the text prototype of one additional class, calculates the image prototype PVn_Img(j) of the additional class j, and outputs it to the similarity calculation unit 30 as the weight vector of the additional class j. FIG. 20 shows an operation example of the weight generation unit 80 in Modification 5. As an example, reference numerals 200a, 200b, 200c, 200d, and 200e indicate five image feature amounts of the additional class j, reference numerals 210a, 210b, and 210c indicate text prototypes of three basic classes, and reference numeral 220 indicates the text prototype of one additional class. Reference numeral 230 indicates the image prototype of the additional class j calculated by averaging the five image feature amounts of the additional class j, the three text prototypes of the basic classes, and the text prototype of one additional class.
[0123] It is known that an image prototype calculated using the image feature amounts of a small number of images of a certain class has insufficient ability (generalization performance) to adapt to unknown image data. Therefore, by using the text prototype of the additional class j and the text prototypes of neighboring basic classes, the generalization performance can be improved.
[0124] (Modification 6) FIG. 21 is a configuration diagram of the image classification apparatus 100 in Modification 6. The configuration of Modification 6 is different in that the basic class text prototype selection unit 72 in Modification 5 is replaced by a basic class image prototype selection unit 74, and the operation of the weight generation unit 80 is different.
[0125] The basic class image prototype selection unit 74 selects L image prototypes of basic classes in the vicinity of the text prototype of the additional class, and gives the selected image prototypes of the basic classes and the text prototype of the additional class to the weight generation unit 80.
[0126] The weight generation unit 80 calculates the image prototype of the additional class by averaging the K image feature amounts of the additional class input from the image feature amount output unit 10, the L image prototypes of the basic classes input from the basic class image prototype selection unit 74, and the text prototype of one additional class input from the basic class image prototype selection unit 74, and uses it as the weight vector of the additional class of the similarity calculation unit 30.
[0127] In Modification 6, step S23 in FIG. 19 is replaced as follows. The basic class image prototype selection unit 74 selects the text prototype of the additional class j and the L image prototypes of the basic classes in the vicinity of the text prototype of the additional class j.
[0128] In Modification 6, step S27 in FIG. 19 is replaced as follows. The weight generation unit 80 calculates the image prototype PVn_Img(j) of the additional class j by averaging the K image feature amounts of the additional class j input from the image feature amount output unit 10, the L image prototypes of the basic classes input from the basic class image prototype selection unit 74, and the text prototype of one additional class.
[0129] It is known that the generalization performance of image prototypes calculated using a small number of image feature amounts is insufficient. Therefore, the generalization performance can be improved by using the image prototypes of the basic classes in the vicinity of the additional class j.
[0130] (Modification 7) FIG. 22 is a configuration diagram of the image classification device 100 according to Modification 7. The configuration of Modification 7 is different in that the basic class text prototype selection unit 72 of Modification 5 and the basic class image prototype selection unit 74 of Modification 6 are replaced by the basic class prototype selection unit 76, and the operation of the weight generation unit 80 is different.
[0131] The basic class prototype selection unit 76 selects a basic class text prototype and a basic class image prototype near the additional class text prototype, and provides the basic class text prototype, the basic class image prototype, and the additional class text prototype to the weight generation unit 80.
[0132] The weight generation unit 80 averages the K image feature amounts of the additional class input from the image feature amount output unit 10, the L basic class text prototypes selected by the basic class prototype selection unit 76, the M basic class image prototypes, and the one additional class text prototype input from the basic class prototype selection unit 76 to calculate the additional class image prototype, which is used as the weight vector of the additional class for the similarity calculation unit 30.
[0133] In Modification 7, step S23 in FIG. 19 is replaced as follows. The basic class prototype selection unit 76 selects the text prototype of the additional class j, the L basic class text prototypes near the text prototype of the additional class j, and the M basic class image prototypes near the text prototype of the additional class j.
[0134] In Modification 7, step S27 in FIG. 19 is replaced as follows. The weight generation unit 80 averages the K image feature amounts of the additional class input from the image feature amount output unit 10, the one additional class text prototype input from the basic class prototype selection unit 76, the L basic class text prototypes selected by the basic class prototype selection unit 76, and the M basic class image prototypes selected by the basic class prototype selection unit 76 to calculate the additional class j image prototype PVn_Img(j).
[0135] It is known that the image prototype calculated using the image feature amounts of a small number of images has insufficient generalization performance. Therefore, the generalization performance can be improved by using the basic class image prototypes near the text prototype of the additional class j and the basic class text prototypes.
[0136] Of course, the various processes of the image classification apparatus 100 described above can be realized not only as an apparatus using hardware such as a CPU and a memory, but also by firmware stored in a ROM (Read Only Memory), a flash memory, etc., or software such as a computer. It is also possible to record the firmware program and the software program on a computer-readable recording medium and provide them, or to transmit and receive them to and from a server through a wired or wireless network, or to transmit and receive them as data broadcast of terrestrial or satellite digital broadcast.
[0137] As described above, the present invention has been described based on the embodiments. It is understood by those skilled in the art that the embodiments are examples, and various modifications are possible in the combination of each component and each processing process, and such modifications are also within the scope of the present invention.
Explanation of Reference Numerals
[0138] 10 Image feature quantity output unit, 20 Image prototype generation unit, 30 Similarity calculation unit, 40 Classification unit, 50 Text feature quantity output unit, 60 Text prototype generation unit, 70 Additional class prototype generation unit, 72 Basic class text prototype selection unit, 74 Basic class image prototype selection unit, 76 Basic class prototype selection unit, 80 Weight generation unit, 100 Image classification apparatus.
Claims
1. An image feature output unit that has been pre-trained from text and images, receives an image as input, and outputs an image feature; an image prototype output unit that calculates the image feature amount for each class and outputs an image prototype for each class; A text feature output unit that has been pre-trained from text and images, receives a text describing a class as input, and outputs text features; a base class image prototype selector for selecting image prototypes of the base class that are in the vicinity of the text features of the additional class; a weight generation unit that generates a weight for the additional class from the image feature of the additional class, the image prototype of the basic class, and the text feature of the additional class; a similarity calculation unit that holds the image prototypes of the basic classes output by the image prototype output unit as weights of the basic classes, and further holds the weights of the additional classes generated by the weight generation unit, and receives the image features output by the image feature output unit to calculate a similarity; An image classification device comprising a classification unit that receives the similarity as an input and determines a classification of the image.
2. An image feature output unit that has been pre-trained from text and images, receives an image as input, and outputs an image feature; an image prototype output unit that calculates the image feature amount for each class and outputs an image prototype for each class; A text feature output unit that has been pre-trained from text and images, receives a text describing a class as input, and outputs text features; a basic class sentence feature selection unit for selecting sentence features of a basic class that are in the vicinity of the sentence features of the additional class; a weight generation unit that generates a weight for the additional class from the image feature of the additional class, the text feature of the basic class, and the text feature of the additional class; a similarity calculation unit that holds the image prototypes of the basic classes output from the image prototype output unit as weights of the basic classes, and further holds the weights of the additional classes generated by the weight generation unit, and receives the image features output from the image feature output unit to calculate a similarity; An image classification device comprising a classification unit that receives the similarity as an input and determines a classification of the image.
3. An image feature output unit that has been pre-trained from text and images, receives an image as input, and outputs an image feature; an image prototype output unit that calculates the image feature amount for each class and outputs an image prototype for each class; A text feature output unit that has been pre-trained from text and images, receives a text describing a class as input, and outputs text features; a base class prototype selection unit for selecting image prototypes of the base class and text features of the base class that are in the vicinity of the text features of the additional class; a weight generating unit that generates a weight for the additional class from the image feature of the additional class, the image prototype of the basic class, the text feature of the basic class, and the text feature of the additional class; a similarity calculation unit that holds the image prototypes of the basic classes output from the image prototype output unit as weights of the basic classes, and further holds the weights of the additional classes generated by the weight generation unit, and receives the image features output from the image feature output unit to calculate a similarity; An image classification device comprising a classification unit that receives the similarity as an input and determines a classification of the image.
4. An image feature output step in which the system has been pre-trained from text and images, an image is input, and an image feature is output; an image prototype output step of calculating the image feature amount for each class and outputting an image prototype of each class; A text feature output step in which the system has been pre-trained from text and images, a text describing a class is input, and text feature values are output; a base class image prototype selection step of selecting image prototypes of the base class that are in the vicinity of the text features of the additional class; a weight generation step of generating weights for the additional classes from the image features of the additional classes, the image prototypes of the basic classes, and the text features of the additional classes; a similarity calculation step of holding the image prototypes of the basic classes outputted in the image prototype output step as weights of the basic classes, further holding the weights of the additional classes generated in the weight generation step, and calculating a similarity using the image features outputted in the image feature output step as input; a classification step for determining a classification of the image using the similarity as an input.
5. An image feature output step in which the system has been pre-trained from text and images, an image is input, and an image feature is output; an image prototype output step of calculating the image feature amount for each class and outputting an image prototype of each class; A text feature output step in which the system has been pre-trained from text and images, a text describing a class is input, and text feature values are output; a base class image prototype selection step of selecting image prototypes of the base class that are in the vicinity of the text features of the additional class; a weight generation step of generating weights for the additional classes from the image features of the additional classes, the image prototypes of the basic classes, and the text features of the additional classes; a similarity calculation step of holding the image prototypes of the basic classes outputted in the image prototype output step as weights of the basic classes, further holding the weights of the additional classes generated in the weight generation step, and calculating a similarity using the image features outputted in the image feature output step as input; a classification step of inputting the similarity and determining a classification of the image,
Citation Information
Patent Citations
Image classification apparatus, image classification method, and image classification program
JP2024129938A