Image classification apparatus, image classification method, and image classification program
The image classification technique addresses the challenge of low classification accuracy for additional classes by using pre-trained models and similarity calculations to generate accurate prototypes and weights, improving performance even with a small number of images.
Patent Information
- Application Number
- JP2023189208
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-19
AI Technical Summary
Existing image classification techniques face challenges in achieving high classification accuracy for additional classes during incremental learning with a small number of images, leading to overfitting and poor generalization performance.
The proposed image classification technique involves pre-training from text and images, utilizing an image feature output unit, image prototype generation, text feature output, similarity calculation, and classification units to improve classification accuracy for additional classes by leveraging pre-trained models and similarity calculations.
This approach enhances the classification accuracy of additional classes without retraining the image feature extraction part, even with a small number of images, by effectively combining image and text features to generate accurate prototypes and weights for similarity calculations.
Smart Images

Figure 2025077195000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image classification technology.
Background Art
[0002] Humans can learn new knowledge through long-term experience and maintain their old knowledge without forgetting it. On the other hand, the knowledge of deep neural networks (DNNs) such as convolutional neural networks (CNNs) depends on the dataset used for learning, and in order to adapt to changes in the data distribution, it is necessary to relearn the parameters of the DNN for the entire dataset. In DNNs, as learning progresses for new tasks, the estimation accuracy for old tasks decreases. Thus, in DNNs, catastrophic forgetting, where the learning results of old tasks are forgotten during the learning of new tasks, cannot be avoided when continuous learning is performed.
[0003] As a method for avoiding catastrophic forgetting, incremental learning or continual learning has been proposed. Incremental learning or continual learning is a learning method in which, when new tasks or new data occur, instead of learning the model from scratch, the currently learned model is improved and learned.
[0004] Also, humans can learn new knowledge from a small number of images. On the other hand, artificial intelligence using deep learning such as convolutional neural networks depends on big data (a large number of images) used for learning. It is known that when artificial intelligence using deep learning is learned with a small number of images, although the local performance is good, it falls into overfitting with poor generalization performance.
[0005] As a method for avoiding overfitting, few shot learning has been proposed. Few shot learning is a learning method in which basic knowledge is learned using big data in a basic task, and new knowledge is learned from a small number of images of a new task using the basic knowledge.
[0006] There is few shot class incremental learning as a method for solving the problems of both continuous learning and few shot learning, and there is a technique that uses an averaged feature vector as a weight vector (Patent Document 1). Also, there is a technique for matching the feature vector of text and the feature vector of an image (Non-Patent Document 1).
Prior Art Documents
Non-Patent Documents
[0007]
Patent Document 1
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] In the prior art, there was a problem that the classification accuracy of images of additional classes did not become sufficiently high for additional learning with a small number of images.
[0009] The present invention has been made in view of such a situation, and an object thereof is to provide an image classification technique capable of improving the classification accuracy of images of additional classes for additional learning with a small number of images.
Means for Solving the Problems
[0010] In order to solve the above problems, an image classification apparatus according to an aspect of the present invention has been pre-trained from text and images, and includes an image feature amount output unit that takes an image as an input and outputs an image feature amount, an image prototype generation unit that calculates the image feature amount for each class and outputs an image prototype for each class, a text feature amount output unit that has been pre-trained from text and images, takes a text describing a class as an input, and outputs a text feature amount, a similarity calculation unit that holds the image prototype of the basic class as the weight of the basic class and the text feature amount of the additional class as the weight of the additional class, takes the image feature amount output from the image feature amount output unit as an input, and calculates a similarity, and a classification unit that takes the similarity as an input and determines the classification of the image.
[0011] Another aspect of the present invention is an image classification method. This method includes an image feature amount output step that has been pre-trained from text and images, takes an image as an input, and outputs an image feature amount, an image prototype generation step that calculates the image feature amount for each class and outputs an image prototype for each class, a text feature amount output step that has been pre-trained from text and images, takes a text describing a class as an input, and outputs a text feature amount, a similarity calculation step that holds the image prototype of the basic class as the weight of the basic class and the text feature amount of the additional class as the weight of the additional class, takes the image feature amount output from the image feature amount output step as an input, and calculates a similarity, and a classification step that takes the similarity as an input and determines the classification of the image.
[0012] In addition, any combination of the above components, and those obtained by converting the expression of the present invention among a method, an apparatus, a system, a recording medium, a computer program, etc. are also effective as aspects of the present invention.
Advantages of the Invention
[0013] According to the present invention, it is possible to provide an image classification technique capable of improving the classification accuracy of images of additional classes for additional learning with a small number of images.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Embodiments for Carrying Out the Invention
[0015] First, the basic training dataset and the additional training dataset will be described.
[0016] The basic training dataset includes a large number of basic classes (for example, about 100 to 1000 classes), and each basic class is a supervised dataset composed of a large number of training images (for example, 500 to 3000 images). It is assumed that the basic training dataset has a sufficient amount of data to independently learn a general classification task. Here, the number of basic classes is 60, and the number of training images for each basic class is 500.
[0017] On the other hand, the additional training dataset includes a small number of additional classes (for example, about 1 to 10 classes), and each additional class is a supervised dataset composed of a small number of training images (for example, about 1 to 10 images). Also, here it is assumed to be a small number of images, but it can be a large number of images as long as it is a small number of classes. Here, the number of additional classes is 5, and the number of training images for each additional class is 5.
[0018] The class names of the basic classes and the additional classes are expressed in text (sentences).
[0019] In incremental learning, the first class to be learned is called the base class (known class), and the class added later (unknown class) is called the additional class. During base class learning, the base class is learned using the base dataset. During additional class learning, the additional dataset is used to add the additional class to the classes learned so far.
[0020] In the classification (inference) stage, the base class and all additional classes can be classified simultaneously.
[0021] There are two ways during incremental learning: one is to retrain the image feature extraction part, and the other is not to retrain the image feature extraction part. When not retraining the image feature extraction part, it is relatively easy to maintain the accuracy of the base class, but there are challenges in improving the accuracy of the additional class. Furthermore, when the number of images of the additional class is small, it becomes more difficult to improve the accuracy of the additional class.
[0022] The object of the image classification device 100 according to the embodiment of the present invention is to improve the class classification accuracy of the additional class without retraining the image feature extraction part when the number of images of the additional class is small.
[0023] (Example) FIG. 1 is a configuration diagram of an image classification device 100 according to an example. The image classification device 100 includes an image feature amount output unit 10, an image prototype generation unit 20, a similarity calculation unit 30, a classification unit 40, a text feature amount output unit 50, a text prototype generation unit 60, an additional class prototype generation unit 70, and a weight generation unit 80.
[0024] The image feature amount output unit 10 is a pre-trained neural network model learned from text and images. It takes an image as input and outputs an image feature amount.
[0025] The image feature amount output unit 10 is composed of a deep neural network and calculates an image feature vector (image feature amount) of the input image. The image feature amount output unit 10 outputs the image feature vector to the image prototype generation unit 20 and the weight generation unit 80.
[0026] Here, it is assumed that the weight parameters of the image feature amount output unit 10 are learned before the data set of the basic class in an image data set with a sufficient amount of data such as ImageNet. Also, it is assumed that the weight parameters of the image feature amount output unit 10 are fixed after being fine-tuned with the data set of the basic class.
[0027] Here, ResNet-18 is used as the image feature amount output unit 10. The dimensionality of the feature vector of ResNet-18 is 512. The image feature amount output unit 10 is not limited to ResNet-18, and may also be a deep neural network such as ViT-B / 32 (the dimensionality of the image feature amount is 512), ViT-L / 14 (the dimensionality of the image feature amount is 768) of ViT (Vision Transformer), RN50 (the dimensionality of the image feature amount is 1024), RN101 (the dimensionality of the image feature amount is 512), RN50x4 (the dimensionality of the image feature amount is 640) of the extended model of ResNet, etc.
[0028] The image prototype generation unit 20 calculates the image feature amount for each class (for example, calculates the average value, median value, representative value, etc.) and outputs the image prototype for each class.
[0029] The similarity calculation unit 30 is composed of a fully connected layer, and the fully connected layer has weight vectors for a plurality of classes. That is, the weight vectors exist for each class.
[0030] The similarity calculation unit 30 calculates the cosine similarity, which is the similarity between the input image feature amount and the weight vectors of a plurality of classes. The similarity calculation unit 30 outputs the cosine similarities of a plurality of classes to the classification unit 40.
[0031] The weight vectors of the fully connected layer include both the basic class and the additional class. Here, the weight vector of each class is a representative vector of the image feature amounts of each class, and is also called a prototype. Here, the prototype is the average of the image feature amounts.
[0032] It is assumed that the weight vector of the basic class of the similarity calculation unit 30 is calculated in advance with the dataset of the basic class and is fixed.
[0033] The image feature amount output unit 10 takes as input all the training images of the basic class included in the basic training dataset, and outputs the image feature amounts of all the training images of the basic class. The image prototype generation unit 20 takes as input the image feature amounts of all the training images of the basic class, averages the image feature amounts for each class, and outputs the respective weight vectors (image prototypes) of the basic class.
[0034] The similarity calculation unit 30 receives the weight vector of the basic class from the image prototype generation unit 20. On the other hand, the weight vector of the additional class of the similarity calculation unit 30 is input from the weight generation unit 80. The specific additional method will be described later.
[0035] The similarity calculation unit 30 holds the image prototype of the basic class output from the image prototype generation unit 20 as the weight of the basic class, holds the additional class prototype output from the weight generation unit 80 as the weight of the additional class, takes as input the image feature amount output from the image feature amount output unit 10, and calculates the similarity.
[0036] The classification unit 40 selects the class having the maximum similarity from the similarities calculated by the similarity calculation unit 30. The similarity is, for example, the cosine similarity, and the cosine similarity indicates the similarity for each class. The classification unit 40 selects the class having the maximum cosine similarity.
[0037] The text feature amount output unit 50 is a pre-trained neural network model learned from text and images. It takes as input text that describes a class and outputs a text feature amount.
[0038] The text feature quantity output unit 50 is composed of a deep neural network, and calculates a text feature vector (text feature quantity) from the text related to the input class name.
[0039] The text feature quantity output unit 50 outputs the text feature vector to the text prototype generation unit 60.
[0040] The text prototype generation unit 60 calculates the text feature quantity for each class (for example, obtains the average value, median value, representative value, etc.) and outputs the text prototype of each class.
[0041] Here, when there is only one piece of text related to a class, the text feature quantity directly becomes the text prototype, so the configuration of the text prototype generation unit 60 may be omitted. In that case, note that the text feature quantity output from the text feature quantity output unit 50 is directly input to the additional class prototype generation unit 70 as the document prototype as it is.
[0042] As the neural network model of the text feature quantity output unit 50, a Transformer model, which is a text encoder learned in a common feature space of images and texts as shown in Non-Patent Document 1, is used. Therefore, the text feature quantity output by the text feature quantity output unit 50 can be matched with the image feature quantity.
[0043] In Non-Patent Document 1, an image encoder and a text encoder are learned so that the image feature vector and the text feature vector of a certain class can be matched in the feature space.
[0044] The text encoder and the image encoder are not good at detailed classification of classes, but are considered to be good at conceptual classification.
[0045] An image encoder is generally a very large neural network and is trained with a large dataset. On the other hand, a text encoder does not necessarily need to be as large a neural network as the image encoder.
[0046] Here, the conceptual classification ability of a text encoder that can be implemented on a small scale is utilized for few-shot learning to improve the classification accuracy.
[0047] Here, it is assumed that the dimension of the feature vector output by the image feature output unit 10 is the same 512 as the dimension of the feature vector output by the text feature output unit 50. By matching the number of dimensions, it becomes possible to directly match without reducing the number of dimensions, etc., thus improving the matching accuracy.
[0048] The prototype calculated based on the feature vector output by the image feature output unit 10 is called an "image prototype", and the prototype calculated based on the feature vector output by the text feature output unit 50 is called a "text prototype".
[0049] The additional class prototype generation unit 70 generates additional class prototypes using the image prototypes of the basic classes and the text prototypes of the basic classes.
[0050] More specifically, the additional class prototype generation unit 70 calculates the difference between the text prototype of the basic class near the text prototype of the additional class and the image prototype of the same basic class as the text prototype of the basic class near the text prototype of the additional class. Further, the additional class prototype generation unit 70 adds this difference to the text feature amount of the additional class to generate an additional class prototype. Here, there may be one or more neighborhoods. Therefore, this difference may be the difference with only one neighborhood or the average of the differences with a plurality of neighborhoods. Note that the neighborhood can be paraphrased as within a predetermined distance.
[0051] The additional class prototype generation unit 70 outputs the additional class prototype to the weight generation unit 80.
[0052] The weight generation unit 80 takes as input the image feature amount of the additional class and the additional class prototype, and calculates the weight vector of the additional class of the similarity calculation unit 30 by calculating the image feature amount of the additional class and the additional class prototype.
[0053] FIG. 2 is a flowchart for explaining the processing procedure of additional learning by the image classification apparatus 100.
[0054] A sentence explaining the basic class is input to the sentence feature amount output unit 50. If there is no appropriate sentence explaining the basic class, the class name of the basic class may be used as the sentence. The sentence feature amount output unit 50 calculates the sentence feature amount of the basic class from the input sentence, and the sentence prototype generation unit 60 generates the sentence prototype of the basic class from the sentence feature amount of the basic class, and outputs the sentence prototype of the basic class to the additional class prototype generation unit 70 (S10).
[0055] The image feature amount output unit 10 takes as input the image of the basic class and outputs the image feature amount. The image prototype generation unit 20 takes as input the image feature amount and outputs the image prototype (weight vector) of the basic class to the additional class prototype generation unit 70 (S12).
[0056] The additional class prototype generation unit 70 holds the sentence prototype of the basic class and the image prototype (weight vector) of the basic class (S14).
[0057] Hereinafter, the processing procedure of additional learning for enabling the additional class to be classified after the additional class is given will be described. No major processing such as optimization is required for additional learning, and additional learning can be repeated.
[0058] Here, for simplicity of explanation, it is assumed that the additional classes are added one by one, but a plurality of them may be added at once.
[0059] As data of the additional class, a sentence describing the additional class j and K images are given. Here, K is an arbitrary integer of 1 or more. When there is no appropriate sentence describing the additional class, the class name of the additional class may be used as the sentence. The sentence is input to the sentence feature amount output unit 50, and the K images are input to the image feature amount output unit 10.
[0060] When the K images are input, the image feature amount output unit 10 calculates K image feature amounts FVn_Img(j,k) of the additional class j and gives them to the weight generation unit 80 (S16). Here, k = 0, 1, 2, ···, K - 1.
[0061] When the sentence describing the additional class is input, the sentence feature amount output unit 50 calculates a sentence prototype PVn_Com(j) of the additional class j and gives it to the additional class prototype generation unit 70 (S18).
[0062] The additional class prototype generation unit 70 selects sentence prototypes of M basic classes in the vicinity of the sentence prototype of the additional class j (for example, M = 3) (S20). Here, M is an arbitrary integer of 1 or more.
[0063] For the selected M basic classes i, the additional class prototype generation unit 70 calculates a movement vector MVb(i) from the sentence prototype PVb_Com(i) to the image prototype PVb_Img(i) as follows (S22). Here, i = 1, 2, ···, M. MVb(i)=PVb_Img(i)-PVb_Com(i)
[0064] The additional class prototype generation unit 70 calculates an average movement vector MVb_ave of the basic classes of the movement vectors of the M basic classes as follows (S24). MVb_ave=ΣMVb(i) / M
[0065] The additional class prototype generation unit 70 adds the average movement vector of the base class to the sentence prototype of the additional class j, calculates the additional class prototype PPVn_Img(j) of the additional class j as shown in the following formula, and provides it to the weight generation unit 80 (S26). PPVn_Img(j)=PVn_Com(j)+MVb_ave
[0066] FIG. 3 is a diagram for explaining an example of calculating the additional class prototype.
[0067] First, select the sentence prototypes (black circles) of the base classes B, C, and E that are near the sentence prototype of the additional class j. Next, calculate the movement vectors MVb(B), MVb(C), and MVb(E) from the sentence prototypes of the base classes B, C, and E to the image prototypes. Average MVb(B), MVb(C), and MVb(E) to calculate the average movement vector MVb_ave. Add the movement vector to the sentence prototype of the additional class j to calculate the additional class prototype of the additional class j.
[0068] FIG. 4 is a diagram for explaining an operation example of the weight generation unit 80.
[0069] The weight generation unit 80 averages the K image feature amounts of the additional class j input from the image feature amount output unit 10 and the additional class prototype of the additional class j input from the additional class prototype generation unit 70, calculates the image prototype PVn_Img(j) of the additional class j as shown in the following formula, and outputs the image prototype of the additional class j to the similarity calculation unit 30 (S28). PVn_Img(j)=(ΣFVn_Img(j,k)+PPVn_Img(j)) / (K + 1)
[0070] The similarity calculation unit 30 adds the image prototype of the additional class j as the weight vector of the additional class j to the fully connected layer (S30).
[0071] Thereby, the image classification device 100 can classify the additional class j in addition to the base class.
[0072] As another example, for instance, if it is known that the accuracy of the sentence prototype of the additional class j output by the sentence feature amount output unit 50 is high, then as shown in the following formula, the weight (α in the following formula) of the additional class prototype can be increased to calculate the weighted average image prototype of the additional class j. For example, set α to be greater than 1, such as setting α to 1.2. PVn_Img(j)=ΣFVn_Img(j,k) / K+α×PPVn_Img(j)
[0073] Also, here, the average of the movement vectors of the neighboring classes is used, but for example, statistical quantities other than the average, such as the median, maximum value, and minimum value of the movement vectors of the neighboring classes, may also be used.
[0074] Next, the classification (inference) processing procedure by the image classification device 100 after the additional class can be classified will be described.
[0075] FIG. 5 is a flowchart for explaining the classification processing procedure by the image classification device 100.
[0076] The input image is input to the image feature amount output unit 10. The image feature amount output unit 10 calculates the image feature amount of the input image and gives it to the similarity calculation unit 30 (S40).
[0077] The similarity calculation unit 30 calculates the similarity between the input image feature amount and the weight vectors of all classes, and gives the similarities of all classes to the classification unit 40 (S42).
[0078] The classification unit 40 selects the class having the maximum similarity from among the similarities of all classes (S44). Thereby, the class of the input image is determined.
[0079] As described above, the image classification apparatus 100 according to the embodiment calculates the additional class prototype by using the additional class sentence prototype, the average movement vector calculated from the basic class sentence prototype and the basic class image prototype. Further, the image prototype of the additional class is calculated as the weight vector of the similarity calculation unit 30 by averaging the additional class prototype and the K image feature amounts of the additional class. Thereby, even if the image data is a small number of additional classes, by using the high-precision basic class image prototype calculated from a large number of images, a high-precision additional class image prototype can be obtained. Further, according to the accuracy of the additional class sentence prototype output by the sentence prototype generation unit 60, by weighted-averaging the additional class image prototypes, the image classification apparatus 100 can classify the additional classes with high accuracy in addition to the basic classes.
[0080] In the above embodiment, the additional class prototype generation unit 70 selects the sentence prototype in the vicinity of the sentence prototype of the additional class j as the basic class that has been learned with a sufficient data amount. However, an additional class input before the additional class j may be selected as the sentence prototype in the vicinity.
[0081] (Modification Example 1) Here, a simpler configuration of the image classification apparatus 100 of the embodiment will be described. Specifically, a modification example in the case where there is only a class name (label) in the additional class and there is no image data will be described.
[0082] FIG. 6 is a diagram for explaining the image classification apparatus 100 of Modification Example 1. The difference from the embodiment is that there is no weight generation unit 80 and there is no processing flow related to the weight generation unit 80. Thereby, even when there is no image data of the additional class, the additional class can be learned by using the class name of the additional class as the sentence data.
[0083] When the feature space of the text feature quantity output unit 50 and the feature space of the image feature quantity output unit 10 are learned to be similar feature spaces, even when there is no image data for the additional class, the additional class can be learned by using the text prototype generated from the class name of the additional class.
[0084] The similar feature space means, for example, that the feature spaces of the text feature quantity output unit 50 and the image feature quantity output unit 10 have the same number of dimensions. When the text feature quantity output by the text feature quantity output unit 50 and the image feature quantity output by the image feature quantity output unit 10 in a certain class are mapped to the same feature space, the distance between the text feature quantity and the image feature quantity becomes close.
[0085] Even for an additional class that has only a class name and no image, by using the additional class prototype, the image classification device 100 can classify the additional class j in addition to the basic class.
[0086] In Modification 1, there is no weight generation unit 80. The additional class prototype generation unit 70 outputs the additional class prototype of the additional class j to the similarity calculation unit 30 as the weight vector of the additional class j. As described above, the text feature quantity and the image feature quantity have the same number of dimensions, and when the text feature quantity and the image feature quantity in a certain class are mapped to the same feature space, the distance becomes close. Therefore, the relationship between the text prototype and the image prototype of the additional class j generated from the text feature quantity and the image feature quantity respectively is the same. Thus, the additional class prototype generation unit 70 can generate an additional class prototype using the text prototype of the additional class j and the image prototype of the basic class.
[0087] FIG. 7 is a flowchart for explaining the additional learning processing procedure by the image classification device 100 of Modification 1. The steps S16 and S28 in FIG. 2 are omitted, and the difference is that the step S30 is replaced by the step S32. Since the other processing procedures are the same as those in FIG. 2, the description of the common processing procedures is omitted, and only the differences are described.
[0088] In step S32, the similarity calculation unit 30 adds the additional class prototype of the additional class j to the fully connected layer as the weight vector of the additional class j.
[0089] (Modification Example 2) Here, a more simplified configuration of Modification Example 1 will be described. FIG. 8 is a configuration diagram of the image classification apparatus 100 according to Modification Example 2. The difference from Modification Example 1 is that there is no additional class prototype generation unit 70. As a result, the sentence prototype output from the sentence prototype generation unit 60 (when there is only one sentence related to the class (for example, only the class name), the sentence feature amount output from the sentence feature amount output unit 50) is used as the weight vector of the similarity calculation unit 30.
[0090] Modification Example 2 can be applied when the feature space of the sentence feature amount output unit 50 and the feature space of the image feature amount output unit 10 are learned to be close feature spaces, and the sentence prototype output from the sentence prototype generation unit 60 can be directly used as the weight vector of the similarity calculation unit 30. Therefore, the configuration can be simplified and the processing load of the image classification apparatus 100 can be reduced.
[0091] Also, even when there is no image data of the additional class, additional learning can be performed only with the class name of the additional class. Furthermore, since it is not necessary to consider parts such as the background that have no relation to the class name in the image, the accuracy of the weight vector can be improved.
[0092] FIG. 9 is a flowchart for explaining the additional learning processing procedure by the image classification apparatus 100 according to Modification Example 2. Steps S10, S12, S14, S16, S20, S22, S24, S26, and S28 in the additional learning processing procedure of the embodiment in FIG. 2 are omitted, step S18 is replaced with step S19, and step S30 is replaced with step S34.
[0093] When the text of the additional class is input, the text feature quantity output unit 50 outputs the text feature quantity to the text prototype generation unit 60. The text prototype generation unit 60 calculates the text prototype PVn_Com(j) of the additional class j from the text feature quantity and provides it to the similarity calculation unit 30 (S19).
[0094] The similarity calculation unit 30 adds the text prototype of the additional class j as the weight vector of the additional class j to the fully connected layer (S34).
[0095] (Modification Example 3) Here, a simple configuration of the embodiment will be described. FIG. 10 is a diagram for explaining the image classification apparatus 100 of Modification Example 3. The difference from the embodiment of FIG. 1 is that there is no additional class prototype generation unit 70.
[0096] The weight generation unit 80 averages the text prototype output from the text prototype generation unit 60 (when there is only one text related to the class (for example, only the class name), the text feature quantity output from the text feature quantity output unit 50) and the image feature quantity output from the image feature quantity output unit 10 to generate the weight vector of the similarity calculation unit 30.
[0097] The feature space of the text feature quantity output unit 50 and the feature space of the image feature quantity output unit 10 are learned to be in a close feature space. When there is image data of the additional class, by using the text prototype of class j output using a text that is high-precision but conceptual information and the image feature quantity of the additional class j output using an image that is specific information, the weight generation unit 80 can generate an appropriate weight vector of the additional class for the similarity calculation unit 30.
[0098] FIG. 11 is a flowchart for explaining the additional learning processing procedure by the image classification apparatus 100 of Modification 3. Steps S10, S12, S14, S20, S22, S24, and S26 of the additional learning processing procedure of the embodiment in FIG. 2 are omitted, and the difference is that step S28 is replaced by step S29. Since the rest is the same as the processing procedure in FIG. 2, the description of the common processing procedure is omitted, and only the differences are described.
[0099] In step S29, the weight generation unit 80 averages the K image feature amounts of the additional class j input from the image feature amount output unit 10 and the text prototype (when there is only one text related to the class, the text feature amount output from the text feature amount output unit 50) input from the text prototype generation unit 60 to calculate the image prototype PVn_Img(j) of the additional class j, and outputs the image prototype of the additional class j to the similarity calculation unit 30.
[0100] FIG. 12 is a diagram for explaining an operation example of the weight generation unit 80 of Modification 3.
[0101] The weight generation unit 80 averages the K image feature amounts FVn_Img(j,k) of the additional class j input from the image feature amount output unit 10 and the text prototype PVn_Com(j) of the additional class j input from the text prototype generation unit 60 to calculate the image prototype PVn_Img(j) of the additional class j as follows. PVn_Img(j)=(ΣFVn_Img(j,k)+PVn_Com(j)) / (K + 1)
[0102] (Modification 4) Here, a modification example of the additional class prototype generation unit 70 will be described. It is assumed that a higher class than a given class is defined for the dataset used in the learning of the basic class.
[0103] FIG. 13 is a diagram for explaining the image classification apparatus 100 according to Modification 4. The difference from the embodiment is that the additional class prototype generation unit 70 is given upper class information. In the embodiment, the sentence prototypes of the basic classes in the vicinity of the sentence prototype of the additional class are used for calculating the additional class prototype. In Modification 4, the upper class is defined in advance, and the sentence prototypes of the basic classes belonging to the same upper class as the additional class are used for calculating the additional class prototype.
[0104] When comparing images of different classes with the same upper class, they have similar features. Therefore, when the deviation between the image prototype and the sentence prototype is large, by calculating the additional class prototype using the image prototype of another class belonging to the same upper class rather than the correlation by the sentence prototype, the accuracy can be improved.
[0105] Note that Modification 1 may be applied to Modification 4 to omit the weight generation unit 80.
[0106] FIG. 14 is a flowchart for explaining the additional learning processing procedure by the image classification apparatus 100 according to Modification 4. The difference is that step S20 in the additional learning processing procedure of the embodiment in FIG. 2 is replaced with step S21, and the rest is the same as the processing procedure in FIG. 2. Therefore, the common processing procedures will be omitted and only the differences will be explained.
[0107] In step S21, the additional class prototype generation unit 70 selects the sentence prototypes of M basic classes belonging to the upper class of the additional class j.
[0108] FIG. 15 is a diagram for explaining an example of the upper class. The upper class (Superclass) is a superordinate concept of the class (Classes). FIG. 15 is an example of the CIFAR100 dataset. For example, if the upper class is aquatic mammals and the additional class is a dolphin, the basic classes belonging to the same upper class are beavers, otters, seals, and whales. For example, aquatic mammals have similar features as images, such as having a tail fin.
[0109] FIG. 16 is a diagram for explaining the calculation of the prototype of the additional class by the additional class prototype generation unit 70.
[0110] Let the basic classes that are the same upper class as the additional class j be class A', class B', class B', class C', and class D'. Let the movement vectors from the sentence prototype to the image prototype be MVA', MVB', MVC', and MVD' in the order of the above basic classes. The additional class prototype generation unit 70 adds the average movement vector, which is the average of MVA', MVB', MVC', and MVD', to the sentence prototype of the additional class j to calculate the prototype of the additional class j.
[0111] For example, assume that the additional class j is a beaver, and that the dolphin, otter, seal, and whale included in the upper class of aquatic mammals are included in the basic classes. Then, the prototypes of the dolphin, otter, seal, and whale are A', B', C', and D', respectively.
[0112] As another example regarding the upper class, as shown in FIG. 17, the upper class may be defined using the classification levels of organisms generally defined, such as order, family, genus, and species.
[0113] When using classes with the same upper class, it can be expected that the closer the upper class is to the lower classification level, the closer the prototypes are located.
[0114] Also, here, the average of the movement vectors of classes with the same upper class is used, but for example, the median, maximum value, minimum value, etc. of the movement vectors may be used.
[0115] (Modification Example 5) FIG. 18 is a configuration diagram of the image classification apparatus 100 according to Modification Example 5. The configuration of Modification Example 5 is different from that of the embodiment in that the additional class prototype generation unit 70 is replaced by the basic class sentence prototype selection unit 72, and the operation of the weight generation unit 80 is different.
[0116] The basic class sentence prototype selection unit 72 selects a basic class sentence prototype near the additional class sentence prototype, and provides the basic class sentence prototype and the additional class sentence prototype to the weight generation unit 80.
[0117] The weight generation unit 80 averages the K image feature amounts of the additional class input from the image feature amount output unit 10, the L basic class sentence prototypes input from the basic class sentence prototype selection unit 72, and one additional class sentence prototype input from the basic class sentence prototype selection unit 72 to calculate an image prototype of the additional class, which is used as the weight vector of the additional class for the similarity calculation unit 30. Here, L is an arbitrary integer of 1 or more.
[0118] FIG. 19 is a flowchart for explaining the additional learning processing procedure by the image classification apparatus 100 of Modification 5. In Modification 5, steps S10, S12, and S14 of the additional learning processing procedure of the embodiment in FIG. 2 are replaced with steps S11 and S15 in FIG. 19, and steps S20, S22, S24, S26, and S28 in FIG. 2 are replaced with steps S23 and S27 in FIG. 19, which is different. Since the rest is the same as the processing procedure in FIG. 2, the description of the common processing procedure is omitted, and only the differences are described.
[0119] In step S11, the sentence prototype generation unit 60 outputs the basic class sentence prototype to the basic class sentence prototype selection unit 72.
[0120] In step S15, the additional class prototype generation unit 70 holds the basic class sentence prototype.
[0121] In step S23, the basic class sentence prototype selection unit 72 selects the sentence prototype of additional class j and L basic class sentence prototypes near the sentence prototype of additional class j.
[0122] In step S27, the weight generation unit 80 averages the K image feature amounts of the additional class j input from the image feature amount output unit 10, the L text prototypes of the basic classes input from the basic class text prototype selection unit 72, and the text prototype of one additional class to calculate the image prototype PVn_Img(j) of the additional class j, and outputs it to the similarity calculation unit 30 as the weight vector of the additional class j. FIG. 20 shows an operation example of the weight generation unit 80 of Modification 5. As an example, reference numerals 200a, 200b, 200c, 200d, and 200e indicate five image feature amounts of the additional class j, reference numerals 210a, 210b, and 210c indicate text prototypes of three basic classes, and reference numeral 220 indicates the text prototype of one additional class. Reference numeral 230 indicates the image prototype of the additional class j calculated by averaging the five image feature amounts of the additional class j, the three text prototypes of the basic classes, and the text prototype of one additional class.
[0123] It is known that an image prototype calculated using the image feature amounts of a small number of images of a certain class has insufficient ability (generalization performance) to adapt to unknown image data. Therefore, the generalization performance can be improved by using the text prototype of the additional class j and the text prototypes of neighboring basic classes.
[0124] (Modification 6) FIG. 21 is a configuration diagram of the image classification apparatus 100 of Modification 6. The configuration of Modification 6 is different in that the basic class text prototype selection unit 72 of Modification 5 is replaced by a basic class image prototype selection unit 74, and the operation of the weight generation unit 80 is different.
[0125] The basic class image prototype selection unit 74 selects L image prototypes of basic classes in the vicinity of the text prototype of the additional class, and gives the selected image prototypes of the basic classes and the text prototype of the additional class to the weight generation unit 80.
[0126] The weight generation unit 80 averages the K image feature amounts of the additional class input from the image feature amount output unit 10, the L image prototypes of the basic class input from the basic class image prototype selection unit 74, and the text prototype of one additional class input from the basic class image prototype selection unit 74 to calculate the image prototype of the additional class, which serves as the weight vector of the additional class for the similarity calculation unit 30.
[0127] In Modification 6, step S23 in FIG. 19 is replaced as follows. The basic class image prototype selection unit 74 selects the text prototype of the additional class j and the L image prototypes of the basic class in the vicinity of the text prototype of the additional class j.
[0128] In Modification 6, step S27 in FIG. 19 is replaced as follows. The weight generation unit 80 averages the K image feature amounts of the additional class j input from the image feature amount output unit 10, the L image prototypes of the basic class input from the basic class image prototype selection unit 74, and the text prototype of one additional class to calculate the image prototype PVn_Img(j) of the additional class j.
[0129] It is known that the image prototype calculated using a small number of image feature amounts has insufficient generalization performance. Therefore, the generalization performance can be improved by using the image prototypes of the basic classes in the vicinity of the additional class j.
[0130] (Modification 7) FIG. 22 is a configuration diagram of the image classification apparatus 100 according to Modification 7. The configuration of Modification 7 is different in that the basic class text prototype selection unit 72 of Modification 5 and the basic class image prototype selection unit 74 of Modification 6 are replaced by a basic class prototype selection unit 76, and the operation of the weight generation unit 80 is different.
[0131] The basic class prototype selection unit 76 selects the text prototype of the basic class and the image prototype of the basic class near the text prototype of the additional class, and provides the text prototype of the basic class, the image prototype of the basic class, and the text prototype of the additional class to the weight generation unit 80.
[0132] The weight generation unit 80 averages the K image feature amounts of the additional class input from the image feature amount output unit 10, the L text prototypes of the basic class selected by the basic class prototype selection unit 76, the M image prototypes of the basic class, and the 1 text prototype of the additional class input from the basic class prototype selection unit 76 to calculate the image prototype of the additional class, which is used as the weight vector of the additional class for the similarity calculation unit 30.
[0133] In Modification 7, step S23 in FIG. 19 is replaced as follows. The basic class prototype selection unit 76 selects the text prototype of the additional class j, the L text prototypes of the basic class near the text prototype of the additional class j, and the M image prototypes of the basic class near the text prototype of the additional class j.
[0134] In Modification 7, step S27 in FIG. 19 is replaced as follows. The weight generation unit 80 averages the K image feature amounts of the additional class input from the image feature amount output unit 10, the 1 text prototype of the additional class input from the basic class prototype selection unit 76, the L text prototypes of the basic class selected by the basic class prototype selection unit 76, and the M image prototypes of the basic class selected by the basic class prototype selection unit 76 to calculate the image prototype PVn_Img(j) of the additional class j.
[0135] It is known that the image prototype calculated using the image feature amounts of a small number of images has insufficient generalization performance. Therefore, the generalization performance can be improved by using the image prototypes of the basic class near the text prototype of the additional class j and the text prototypes of the basic class.
[0136] Of course, the various processes of the image classification apparatus 100 described above can be realized not only as an apparatus using hardware such as a CPU and a memory, but also by firmware stored in a ROM (Read Only Memory), a flash memory, etc., or software such as a computer. It is also possible to record and provide the firmware program and the software program on a computer-readable recording medium, to transmit and receive them to and from a server through a wired or wireless network, or to transmit and receive them as data broadcasting of terrestrial or satellite digital broadcasting.
[0137] As described above, the present invention has been described based on the embodiments. It is understood by those skilled in the art that the embodiments are examples, and various modifications are possible in the combination of each component and each processing process, and such modifications are also within the scope of the present invention.
Explanation of Reference Numerals
[0138] 10 Image feature amount output unit, 20 Image prototype generation unit, 30 Similarity calculation unit, 40 Classification unit, 50 Text feature amount output unit, 60 Text prototype generation unit, 70 Additional class prototype generation unit, 72 Basic class text prototype selection unit, 74 Basic class image prototype selection unit, 76 Basic class prototype selection unit, 80 Weight generation unit, 100 Image classification apparatus.
Claims
1. An image feature output unit that has been pre-trained from text and images, receives an image as input, and outputs an image feature; an image prototype generation unit that calculates the image feature amount for each class and outputs an image prototype for each class; A text feature output unit that has been pre-trained from text and images, receives a text describing a class as input, and outputs text features; a similarity calculation unit that holds the image prototype of a base class as a weight of the base class and the text feature of an additional class as a weight of the additional class, receives the image feature output from the image feature output unit as an input, and calculates a similarity; An image classification device comprising a classification unit that receives the similarity as an input and determines a classification of the image.
2. The image classification device according to claim 1 , further comprising a weight generation unit that generates weights for the additional class for the similarity calculation unit using the image features of the additional class output from the image feature output unit and the sentence features of the additional class output from the sentence feature output unit.
3. An image feature output step in which the system has been pre-trained from text and images, an image is input, and an image feature is output; an image prototype generating step of calculating the image feature quantity for each class and outputting an image prototype for each class; A text feature output step in which the system has been pre-trained from text and images, a text describing a class is input, and text feature values are output; a similarity calculation step of holding the image prototype of a base class as a weight of the base class and the text feature of an additional class as a weight of the additional class, inputting the image feature output from the image feature output step, and calculating a similarity; a classification step for determining a classification of the image using the similarity as an input.
4. An image feature output step in which the system has been pre-trained from text and images, an image is input, and an image feature is output; an image prototype generating step of calculating the image feature quantity for each class and outputting an image prototype for each class; A text feature output step in which the system has been pre-trained from text and images, a text describing a class is input, and text feature values are output; a similarity calculation step of holding the image prototype of a base class as a weight of the base class and the text feature of an additional class as a weight of the additional class, inputting the image feature output from the image feature output step, and calculating a similarity; a classification step of inputting the similarity and determining a classification of the image,
Citation Information
Patent Citations
Image classification apparatus, image classification method, and image classification program
JP2024129938A