Image classification device, image classification method, and image classification program

By employing feature extraction and similarity calculation methods, and utilizing the average feature matrix and cosine similarity, the problem of insufficient image classification accuracy is solved, achieving high-precision classification in both append-only learning and learning with a small number of images.

CN120917495APending Publication Date: 2025-11-07JVC KENWOOD CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480018781.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-14
Filing Date
2024-02-05
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of image classification is not high enough when learning from additional or a small number of images, making it difficult to improve effectively.

Method used

The first and second feature vectors of the input image are extracted by the feature extraction unit, and the average feature matrix is ​​calculated by the average feature calculation unit. The cosine similarity and weight matrix are combined by the feature similarity calculation unit, and the classification accuracy is improved by permuting the weight matrix.

Benefits of technology

It improves the accuracy of image classification, especially in cases of supplementary learning or learning from a small number of images, enhancing the adaptability and accuracy of the classifier.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120917495A_ABST
    Figure CN120917495A_ABST
Patent Text Reader

Abstract

A feature extraction unit (510) outputs a first feature vector and a second feature vector of an input image. An average first feature calculation unit (520a) and an average second feature calculation unit (520b) average the first feature vector and the second feature vector of some types to calculate an average first feature vector and an average second feature vector. And summarizing the average first feature vectors and the average second feature vectors of all categories to obtain an average first feature matrix and an average second feature matrix. A first feature similarity calculation unit and a second feature similarity calculation unit (532a, 532b) calculate a first similarity and a second similarity on the basis of a first feature vector and a second feature vector of an input image, and a first weight matrix and a second weight matrix. An average first feature calculation unit and an average second feature calculation unit (520a, 520b) replace the first weight matrix and the second weight matrix of the first feature similarity calculation unit and the second feature similarity calculation unit (532a, 532b) with an average first feature matrix and an average second feature matrix.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to an image classification technique. BACKGROUND

[0002] A person is able to learn new knowledge through long-term experience and is able to maintain without forgetting past knowledge. On the other hand, knowledge using a deep neural network (DNN) such as a convolutional neural network (CNN) is dependent on a data set used in learning, and in order to adapt to a change in data distribution, relearning of parameters of the DNN is required for the entire data set. In the DNN, as learning of a new task is performed, the estimation accuracy of a past task decreases. In this way, in the DNN, if continuous learning is performed, a catastrophic forgetting of a learning result of a past task in learning of a new task cannot be avoided.

[0003] As a method of avoiding catastrophic forgetting, incremental learning is proposed. Incremental learning is a learning method in which, when a new task or new data is generated, learning is performed by improving a model that has completed learning at present, rather than starting learning from the beginning.

[0004] In addition, a person is able to learn new knowledge from a small number of images. On the other hand, artificial intelligence using deep learning using a convolutional neural network or the like is dependent on big data (a large number of images) used in learning. It is known that, when learning of artificial intelligence using deep learning is performed using a small number of images, overfitting in which generalization performance is poor but local performance is good is caused.

[0005] As a method of avoiding overfitting, few shot learning is proposed. Few shot learning is a learning method in which basic knowledge is learned using big data in a basic task, and new knowledge is learned from a small number of images of a new task using the basic knowledge.

[0006] As a method of solving the problems of both incremental learning and few shot learning, there is few shot class incremental learning (non-patent literature 1). In addition, as one method of few shot learning, there is a technique of normalizing a feature vector and a weight vector and using cosine similarity (non-patent literature 2).

[0007] PRIOR ART DOCUMENTS

[0008] NON-PATENT LITERATURE

[0009] Non-Patent Literature 1: Tao, Xiaoyu, et al. "Few-shot class-incremental learning." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020;

[0010] Non-Patent Literature 2: Chen, Wei-Yu, et al. "A closer look at few-shot classification." arXiv preprint arXiv:1904.04232 (2019). SUMMARY

[0011] In the related art, there is a problem that the classification accuracy of an image is not high enough with respect to additional learning or learning with a small number of images.

[0012] The present embodiment is made in view of such a situation, and aims to provide an image classification technique capable of improving the classification accuracy of an image with respect to additional learning or learning with a small number of images.

[0013] To solve the above problem, an image classification device according to an embodiment includes: a feature extraction unit that outputs a first feature vector of an input image and outputs a second feature vector that is a different feature vector from the first feature vector; an average first feature calculation unit that calculates an average first feature vector by averaging the first feature vector of a certain class and aggregates the average first feature vectors of all classes to obtain an average first feature matrix; an average second feature calculation unit that calculates an average second feature vector by averaging the second feature vector of a certain class and aggregates the average second feature vectors of all classes to obtain an average second feature matrix; a first feature similarity calculation unit that calculates a first similarity from the first feature vector of the input image and a first weight matrix; and a second feature similarity calculation unit that calculates a second similarity from the second feature vector of the input image and a second weight matrix. The average first feature calculation unit replaces the first weight matrix of the first feature similarity calculation unit with the average first feature matrix, and the average second feature calculation unit replaces the second weight matrix of the second feature similarity calculation unit with the average second feature matrix.

[0014] In the above description, an example of "first" is "deep layer" or "first deep layer" in the embodiment, and an example of "second" is "shallow layer" or "second deep layer" in the embodiment.

[0015] Another embodiment of this method is an image classification method. This method includes: a feature extraction step, outputting a first feature vector of the input image and a second feature vector, wherein the second feature vector is a feature vector different from the first feature vector; an average first feature calculation step, averaging the first feature vectors of a certain category to calculate an average first feature vector, and summing the average first feature vectors of all categories to obtain an average first feature matrix; an average second feature calculation step, averaging the second feature vectors of a certain category to calculate an average second feature vector, and summing the average second feature vectors of all categories to obtain an average second feature matrix; a first feature similarity calculation step, calculating a first similarity based on the first feature vector and a first weight matrix of the input image; and a second feature similarity calculation step, calculating a second similarity based on the second feature vector and the second weight matrix of the input image. In the average first feature calculation step, the first weight matrix of the first feature similarity calculation step is replaced with the average first feature matrix; in the average second feature calculation step, the second weight matrix of the second feature similarity calculation step is replaced with the average second feature matrix.

[0016] In the above description, one example of "first" is "deep layer" or "first deep layer" in the implementation method, and one example of "second" is "shallow layer" or "second deep layer" in the implementation method.

[0017] Furthermore, any combination of the above-mentioned constituent elements, or any variation of the embodiment in terms of methods, apparatus, systems, recording media, computer programs, etc., is also valid as an embodiment.

[0018] According to this embodiment, an image classification technique can be provided that can improve the classification accuracy of images by learning from additional learning or learning from a small number of images. Attached Figure Description

[0019] Figure 1 This is a structural diagram of the image classification learning device involved in the implementation method;

[0020] Figure 2 This is an explanation based on Figure 1 A flowchart of the overall learning process of the image classification learning device;

[0021] Figure 3 Is with Figure 1 A structural diagram related to the learning of basic categories in an image classification learning device;

[0022] Figure 4 This is a learning instruction for basic categories using a basic dataset. Figure 1a flowchart illustrating detailed actions of the image classification learning device;

[0023] Figure 5 is a diagram illustrating deep feature vectors and shallow feature vectors;

[0024] Figure 6 is a flowchart illustrating a method of calculating average deep feature matrices and average shallow feature matrices;

[0025] Figure 7A is an example of a structure related to learning of an additional category of the image classification learning device;

[0026] Figure 7B is another example of a structure related to learning of an additional category of the image classification learning device;

[0027] Figure 7C is still another example of a structure related to learning of an additional category of the image classification learning device;

[0028] Figure 8A is a flowchart illustrating detailed actions of the image classification learning device 500 for learning of an additional category using an additional data set;

[0029] Figure 8B is a flowchart illustrating detailed actions of the image classification learning device for learning of an additional category using an additional data set;

[0030] Figure 8C is a flowchart illustrating detailed actions of the image classification learning device for learning of an additional category using an additional data set;

[0031] Figure 9 is a structural diagram of an image classification device;

[0032] Figure 10 is a flowchart illustrating detailed actions of the image classification device;

[0033] Figure 11 is a structural diagram of an image additional classification device;

[0034] Figure 12 is a flowchart illustrating detailed actions of the image additional classification device. DETAILED DESCRIPTION

[0035] Figure 1 is a structural diagram of the image classification learning device 500 to which the embodiment relates. The image classification learning device 500, after learning a basic category composed of a plurality of training data, continues learning of an additional category composed of a small number of training data, which is a few-shot class-incremental learning.

[0036] The image classification learning apparatus 500 includes a feature extraction section 510, an average deep feature calculation section 520a, an average shallow feature calculation section 520b, a classification section 530, a deep similarity scaling section 540a, a shallow similarity scaling section 540b, a learning section 550, a comprehensive similarity calculation section 560, and a classification determination section 570. The classification section 530 includes a deep feature similarity calculation section 532a and a shallow feature similarity calculation section 532b. The learning section 550 includes a deep loss operation section 552a, a shallow loss operation section 552b, a loss weighted addition section 554, and an optimization section 556.

[0037] Figure 2 is a flowchart illustrating an overall flow of learning based on the image classification learning apparatus 500. Referring to Figure 1 and Figure 2 The structure and actions of the image classification learning apparatus 500 are described.

[0038] First, the basic training data set and the additional training data set are described.

[0039] The basic training data set contains a plurality of basic classes (e.g., 100 to 1000 classes or so) and is a supervised data set in which each class is composed of a plurality of images (e.g., 3000 images). The basic training data set is set to be a data amount sufficient to learn a general classification task alone. Here, the number of basic classes is set to 60.

[0040] On the other hand, the additional training data set contains a small number of additional classes (e.g., 1 to 10 classes or so), and each additional class is a supervised data set composed of a small number of images (e.g., 1 to 10 or so). In addition, here, it is set to be a small number of images, but if it is a small number of classes, it can also be a large number of images. Here, the number of additional classes is set to 5.

[0041] Using the basic training data set, the feature extraction section 510 and the classification section 530 are learned based on the cosine similarity for the basic classes (S501). The learning session using the basic training data set is set to session 0. It is also referred to as the initial session.

[0042] The basic class weight vector of the feature extraction section 510 and the classification section 530 learned is not updated at the time of additional learning.

[0043] The images of the learned basic classes are classified (S502). This step does not necessarily need to be performed.

[0044] Next, the additional learning session s is repeated L times (s = 1, 2,..., L).

[0045] The weight vector of the additional category of the additional session s is learned by the classification section 530 based on the cosine distance using the additional training data set s (S503).

[0046] The basic category and the additional category that have been learned are classified (S504). This step does not necessarily need to be performed.

[0047] s is incremented by 1, and the step S503 is returned to, and the steps S503 to S504 are repeated until s = L, and if s exceeds L, the process is ended.

[0048] Here, L = 8 is assumed. In this case, 65 categories are learned after the end of the additional learning session 1, 70 categories are learned after the end of the additional learning session 2, and 100 categories are learned after the end of the additional learning session 8.

[0049] Figure 3 is a structural diagram related to learning of the basic category of the image classification learning apparatus 500. Figure 4 is a flowchart illustrating the detailed action of the image classification learning apparatus 500 for learning of the basic category using the basic data set. Referring to Figure 3 and Figure 4 The learning action of the basic category of the image classification learning apparatus 500 will be described in detail.

[0050] The learning is repeatedly performed N times (b = 1, 2,..., N) in units of batch size. For example, the batch size is set to 128. The cycle is repeatedly performed M times (e = 1, 2,..., M). The number of cycles is 400.

[0051] When an image is input to the feature extraction section 510, a deep feature vector and a shallow feature vector are extracted (S510).

[0052] First, the deep feature vector and the shallow feature vector will be described.

[0053] Figure 5 is a diagram illustrating the deep feature vector and the shallow feature vector.

[0054] The feature extraction section 510 includes CONV1 to CONV5, which are convolution layers of ResNet-18, and GAP1 (Global Average Pooling) and GAP2. GAP converts a feature map output from a convolution layer into a feature vector. In GAP1, a 7 x 7, 512-channel feature map is input from CONV5, and a 512-dimensional deep layer feature vector is output. In GAP2, a 14 x 14, 256-channel feature map is input from CONV4, and a 256-dimensional shallow layer feature vector is output. In addition, a 28 x 28, 128-channel feature map is output from CONV3, a 56 x 56, 64-channel feature map is output from CONV2, and a 112 x 112, 64-channel feature map is output from CONV1, respectively.

[0055] The deep layer feature vector is low resolution of 7 x 7 in terms of a feature map, and contains summary information of the entire image because a wide range of the entire image is convolved. On the other hand, the shallow layer feature vector is high resolution of 14 x 14 in terms of a feature map, and contains detailed information of a local image because a narrow range of the image is convolved. On the other hand, the deep layer feature vector contains a higher-dimensional feature vector than the shallow layer feature vector.

[0056] The feature extraction section 510 can also be a deep learning network having a plurality of weight parameters other than ResNet-18 such as VGG16 or ResNet-34, and the dimension of the feature vector can also be other than 512 dimensions and 256 dimensions. In addition, the feature map input to GAP2 can also be a feature map other than CONV4 such as CONV3 or CONV2. In addition, the feature extraction section 510 is configured to output two feature vectors, but can also output one or more than three feature vectors. Here, as the shallow layer feature vector, a feature map output of a layer of CONV4 is used, but for example, which layer of feature map to use can be determined in an initial session. For example, in the initial session, the accuracy is measured by learning all of CONV1 to CONV4 as the shallow layer feature vector, and the output of the layer that results in the best classification result is selected as the shallow layer feature vector.

[0057] The deep layer feature vector output from GAP1 of the feature extraction section 510 is input to the deep layer feature similarity calculation section 532a.

[0058] The shallow layer feature vector output from GAP2 of the feature extraction section 510 is input to the shallow layer feature similarity calculation section 532b.

[0059] Since the deep layer feature similarity calculation section 532a and the shallow layer feature similarity calculation section 532b have the same structure, they will be collectively described as the feature similarity calculation section 532.

[0060] The feature similarity calculation section 532 has a weight matrix of a linear layer (a fully connected layer) for deriving cosine similarity. The weight matrix has a weight of (D x NC) dimensions. Here, D is a weight vector of the same dimensions as the feature vector input to the linear layer. In the case of the deep feature similarity calculation section, D = 512, and in the case of the shallow feature similarity calculation section, D = 256. NC is the number of classes. Here, NC is the sum of 100 of the basic classes and the additional classes. NC can also be more than the sum of the basic classes and the additional classes.

[0061] The input feature vector is normalized, and the normalized feature vector is input to the linear layer. At this time, the weight vector of the linear layer is also normalized. As a result, the NC-dimensional cosine similarity of the weight vector of each class of the feature vector and the classification is derived. By normalizing the feature vector to calculate the cosine similarity, the within-class variance can be suppressed to improve the classification accuracy.

[0062] The deep feature similarity calculation section 532a calculates the deep cosine similarity from the input deep feature vector and the deep weight vector of each class and outputs it to the deep similarity scaling section 540a (S511a).

[0063] The deep similarity scaling section 540a scales the input deep cosine similarity by a to output the deep cosine similarity using the deep learning parameter (S512a).

[0064] The shallow feature similarity calculation section 532b calculates the shallow cosine similarity from the input shallow feature vector and the shallow weight vector of each class and outputs it to the shallow similarity scaling section 540b (S511b).

[0065] The shallow similarity scaling section 540b scales the input shallow cosine similarity by a to output the shallow cosine similarity using the shallow learning parameter (S512b).

[0066] Here, the deep learning parameter and the shallow learning parameter are scaled using the same value a, but can also be scaled with different values, i.e., a1 and a2.

[0067] The deep loss operation section 552a calculates the loss of the deep cosine similarity and the correct answer label (correct answer class) of the input image, i.e., the deep cross-entropy loss (S513a).

[0068] The shallow loss operation section 552b calculates the loss of the shallow cosine similarity and the correct answer label (correct answer class) of the input image, i.e., the shallow cross-entropy loss (S513b).

[0069] The loss-weighted addition unit 554 adds the deep-layer cross-entropy loss Ld and the shallow-layer cross-entropy loss Ls with weighting to calculate the total cross-entropy loss L (S514). Here, λ is set to a prescribed value of 0 to 1, for example, 0.2. Here, 0.2 is used as λ, but for example, which value of 0 to 1 is used can be decided in the initial session. For example, in the initial session, the accuracy is measured while learning all values of 0.05 degrees from 0 to 1 as λ, and the value that results in the best classification result can be selected as λ. If such a structure, the initial session can be performed by offline processing, and the additional session can be performed by online processing.

[0070] L = (1 - λ) x Ld + λ x Ls

[0071] The optimization unit 556 optimizes the weight parameters of the convolution layers of the feature extraction unit 510 and the weight matrix of the feature similarity calculation unit 532 by backpropagation using an optimization method such as stochastic gradient descent (SGD) or Adam to minimize the total cross-entropy loss (S515). In addition, the feature similarity calculation unit 532 is essentially the classification unit 530.

[0072] When the learning (cycle) ends, the average deep-layer feature calculation unit 520a calculates an average deep-layer feature matrix, and the average shallow-layer feature calculation unit 520b calculates an average shallow-layer feature matrix (S516).

[0073] The average deep-layer feature calculation unit 520a replaces the weight matrix of the deep-layer feature similarity calculation unit 532a with the average deep-layer feature matrix (S517a).

[0074] The average shallow-layer feature calculation unit 520b replaces the weight matrix of the shallow-layer feature similarity calculation unit 532b with the average shallow-layer feature matrix (S517b).

[0075] Figure 6 is a flowchart that describes the method of calculating the average deep-layer feature matrix and the average shallow-layer feature matrix. The method of calculating the average deep-layer feature matrix and the average shallow-layer feature matrix based on the average deep-layer feature calculation unit 520a and the average shallow-layer feature calculation unit 520b is described.

[0076] Here, the number of classes of the basic classes is set to K. As c = 1, 2,..., K, the average deep-layer feature vector and the average shallow-layer feature vector are calculated for each class. The entire image data of a certain class c included in the basic training data set is input to the feature extraction unit 510, the deep-layer feature vector of the entire image and the shallow-layer feature vector of the entire image are calculated, and the calculated deep-layer feature vector of the entire image and the shallow-layer feature vector of the entire image are obtained (S520).

[0077] The average deep feature vector is obtained by averaging all the deep feature vectors for a certain category c (S521a).

[0078] The average shallow feature vector is obtained by averaging all the shallow feature vectors for a certain category c (S521b).

[0079] The average deep feature vectors for all categories are summarized to obtain an average deep feature matrix (S522a).

[0080] The average shallow feature vectors for all categories are summarized to obtain an average shallow feature matrix (S522b).

[0081] Here, the average deep feature vectors for all categories are summarized as an average deep feature matrix of (D x NC) dimensions, and the weight matrix of the deep feature similarity calculation section 532a is replaced with the average deep feature matrix. In addition, the average shallow feature vectors for all categories are summarized as an average shallow feature matrix, and the weight matrix of the shallow feature similarity calculation section 532b is replaced with the average shallow feature matrix.

[0082] Not limited to this, the weight matrix of the deep feature similarity calculation section can be replaced with the average deep feature vector for a certain category. Similarly, the weight matrix of the shallow feature similarity calculation section can be replaced with the average shallow feature vector for a certain category.

[0083] Thus, by using the average deep feature matrix and the average shallow feature matrix, which are obtained using the feature extraction section 510 learned considering the entire image and the partial region of the image, as the weight matrix of the deep feature similarity calculation section 532a and the weight matrix of the shallow feature similarity calculation section 532b, respectively, it is possible to obtain a classifier that is independent of the batch size and the like of the learning process. Furthermore, the calculation of the average deep feature matrix and the average shallow feature matrix is independent of the amount of data, and can be used for both small data and large data.

[0084] Figure 7A is an example of a structure related to learning of an additional category of the image classification learning apparatus 500. Figure 8A is an example of a flowchart illustrating the detailed operation of the image classification learning apparatus 500 for learning of an additional category using an additional data set. Referring to Figure 7A and Figure 8A The learning operation of the image classification learning apparatus 500 for an additional category will be described in detail.

[0085] The feature extraction section 510 has the same configuration and the same parameters as the feature extraction section 510 obtained when learning the basic categories.

[0086] Here, the number of classes of the additional class is set to L, and the additional class c (c = 1, 2,..., L) has N (i = 1, 2,..., N) pieces of image data. When the N pieces of image data of the additional class c are input to the feature extraction section 510, the deep feature vector and the shallow feature vector are extracted for each piece of image data, and output to the average deep feature calculation section 520a and the average shallow feature calculation section 520b, respectively (S530).

[0087] The average deep feature calculation section 520a averages the deep feature vectors output from the feature extraction section 510 to calculate an average deep feature vector (S530-1a). The average shallow feature calculation section 520b averages the shallow feature vectors output from the feature extraction section 510 to calculate an average shallow feature vector (S530-1b).

[0088] The average deep feature calculation section 520a aggregates the average deep feature vectors of the additional classes to obtain an average deep feature matrix, and the average shallow feature calculation section 520b aggregates the average shallow feature vectors of the additional classes to obtain an average shallow feature matrix (S531).

[0089] The average deep feature matrix is substituted for the weight matrix of the additional classes of the deep feature similarity calculation section 532a (S532a).

[0090] The average shallow feature matrix is substituted for the weight matrix of the additional classes of the shallow feature similarity calculation section 532b (S532b).

[0091] As described above, without learning the image data of the additional classes, by calculating the average deep feature vector and the average shallow feature vector of the image data of the additional classes and substituting them for the weight matrix, it is possible to generate a classification section corresponding to the additional learning. Of course, as the image data of the additional classes from which the average deep feature vector is calculated, it is not necessary to use all of the image data of the basic classes and the additional classes, and it is possible to use only a part of the image data.

[0092] It is also possible to substitute the weight matrix of the deep feature similarity calculation section 532a and the weight matrix of the shallow feature similarity calculation section 532b generated in this way for the weight matrix of the deep feature similarity calculation section 532a and the weight matrix of the shallow feature similarity calculation section 532b of the image classification device 580 of Figure 9 or the image additional classification device 590. Figure 11

[0093] ​Here, the feature extraction section 510 uses a feature extraction section 510 having the same structure and the same parameters as the feature extraction section 510 obtained at the time of learning the basic classes, in order to learn the additional classes. In the case where classification is not required for the basic classes, it is not necessary to be the same structure and the same parameters as the feature extraction section 510 obtained at the time of learning the basic classes, and it can be arbitrary parameters as long as it is a learned feature extraction section 510 and has a plurality of levels.

[0094] Figure 7B is another example of a structure related to learning of additional classes of the image classification learning apparatus 500. Figure 8B is a flowchart for explaining in detail the action of the image classification learning apparatus 500 for learning of additional classes using an additional data set. Referring to Figure 7B and Figure 8B Another example of the learning action of the image classification learning apparatus 500 for additional classes is explained in detail.

[0095] The feature extraction section 510 has the same structure and the same parameters as the feature extraction section 510 obtained at the time of learning the basic classes.

[0096] The image data of the N (i = 1, 2,..., N) additional classes is input to the feature extraction section 510, and a deep feature vector and a shallow feature vector are extracted for each image data by the feature extraction section 510 (S530).

[0097] The average deep feature vectors of the additional classes are summarized to obtain an average deep feature matrix, and the average shallow feature vectors of the additional classes are summarized to obtain an average shallow feature matrix (S531).

[0098] The weight matrix of the additional classes of the deep feature similarity calculation section 532a is replaced with the average deep feature matrix obtained in S531 (S532a).

[0099] The weight matrix of the additional classes of the shallow feature similarity calculation section 532b is replaced with the average shallow feature matrix obtained in S531 (S532b).

[0100] The image data of the N (i = 1, 2,..., N) additional classes is input to the feature extraction section 510, and a deep feature vector and a shallow feature vector are extracted for each image data by the feature extraction section 510 (S530).

[0101] The deep feature similarity calculation section 532a calculates a deep cosine similarity from the input deep feature vector and the average deep feature vector of each class, and outputs to the deep similarity scaling section 540a (S533a).

[0102] The shallow feature similarity calculation section 532b calculates a shallow cosine similarity from the input shallow feature vector and the average shallow feature vector of each category, and outputs the shallow cosine similarity to the shallow similarity scaling section 540b (S533b).

[0103] The deep similarity scaling section 540a scales the input deep cosine similarity by a factor of a using the deep learning parameter, and outputs the deep cosine similarity (S534a).

[0104] The shallow similarity scaling section 540b scales the input shallow cosine similarity by a factor of a using the shallow learning parameter, and outputs the shallow cosine similarity (S534b).

[0105] The deep loss operation section 552a calculates a loss of the deep cosine similarity with respect to the correct answer label (correct answer category) of the input image, that is, a deep cross-entropy loss (S535a).

[0106] The shallow loss operation section 552b calculates a loss of the shallow cosine similarity with respect to the correct answer label (correct answer category) of the input image, that is, a shallow cross-entropy loss (S535b).

[0107] The loss weighted addition section 554 calculates a total cross-entropy loss L by weighted addition of the deep cross-entropy loss Ld and the shallow cross-entropy loss Ls (S536). Here, λ is set to a prescribed value of 0 to 1, for example, 0.2.

[0108] L = (1 - λ) x Ld + λ x Ls

[0109] The optimization section 556 optimizes the weight matrix of the feature similarity calculation section 532 by backpropagation using an optimization method such as stochastic gradient descent (SGD) or Adam, so as to minimize the total cross-entropy loss (S537).

[0110] Here, the learning rate for learning the additional category is set to be smaller than the learning rate for learning the basic category. In addition, the number of cycles for learning the additional category is set to be smaller than the number of cycles for learning the basic category.

[0111] As described above, by performing fine-tuning (adjustment learning) after calculating the average feature vector of the image data of the additional category and replacing it as the weight matrix of the feature similarity calculation section 532, it is possible to generate the classification section 530 corresponding to the additional learning. This is the same as learning the initial value of the weight matrix of the feature similarity calculation section 532 as the average feature vector. As a result, it is possible to obtain a more appropriate weight vector than the average feature vector.

[0112] Here, for both the weight matrix of the additional category of the deep feature similarity calculation section 532a and the weight matrix of the additional category of the shallow feature similarity calculation section 532b, fine tuning is performed after substitution into the average feature vector, but fine tuning can be performed on only either one.

[0113] In addition, here, for both the weight matrix of the additional category of the deep feature similarity calculation section 532a and the weight matrix of the additional category of the shallow feature similarity calculation section 532b, substitution into the average feature vector is performed, but substitution into the average feature vector can be performed on only either one. In the case where substitution into the average feature vector is not performed, the weight matrix of the additional category is initialized, for example, with a random value.

[0114] As described above, by changing the learning characteristics of the deep feature similarity calculation section 532a and the shallow feature similarity calculation section 532b, it is possible to change the learning tendencies of the deep feature similarity calculation section 532a and the shallow feature similarity calculation section 532b, the likelihood of an increase in accuracy when they are combined increases, and it is possible to improve the accuracy.

[0115] The weight matrix of the deep feature similarity calculation section 532a and the weight matrix of the shallow feature similarity calculation section 532b generated in this way can be substituted into the weight matrix of the deep feature similarity calculation section 532a and the weight matrix of the shallow feature similarity calculation section 532b of the image classification device 580 or the image additional classification device 590. Figure 9 Figure 11 The weight matrix of the deep feature similarity calculation section 532a and the weight matrix of the shallow feature similarity calculation section 532b of the image classification device 580 or the image additional classification device 590.

[0116] Here, the feature extraction section 510 used the feature extraction section 510 having the same parameters in the same structure as the feature extraction section 510 obtained at the time of learning the basic category in order to learn the basic category. In the case where classification is not required for the basic category, it is not necessary to be the same parameters in the same structure as the feature extraction section 510 obtained at the time of learning the basic category, but as long as it is a feature extraction section 510 that has been learned, as long as it has a plurality of levels, it can be arbitrary parameters.

[0117] As described above, the average feature vector is held as the weight vector for a plurality of resolutions such as the deep layer (low resolution) and the shallow layer (high resolution), and thus, for an image for which the features of the image cannot be expressed in the average feature vector of the low resolution, it is also possible to express the features of the image by using the average feature vector of the high resolution.

[0118] Figure 7C is another example of the structure related to learning of an additional category of the image classification learning device 500. Figure 8C is a flowchart illustrating the detailed operation of the image classification learning device 500 for another example of learning of an additional category using an additional data set. Referring to​Figure 7C and Figure 8C Another example of the learning operation of the additional category of the image classification learning device 500 will be described.

[0119] Here, instead of using average feature vectors of a plurality of resolutions such as deep (low resolution) and shallow (high resolution), average feature vectors of a plurality of groups of the same resolution are used. In this case, the images of a certain category are divided into a plurality of groups, and average feature vectors are calculated for each group to obtain a plurality of average feature vectors. In the case where the images of a certain category are divided into a plurality of groups, the division can be performed randomly, but by classifying based on a prescribed characteristic using principal component analysis or the like, average feature vectors corresponding to the prescribed characteristic can be calculated.

[0120] The feature extraction section 510 has the same structure and the same parameters as the feature extraction section 510 obtained when learning the basic categories.

[0121] Here, a case where the N (i = 1, 2,..., N) additional categories are divided into two groups will be described. The image data of the N (i = 1, 2,..., N) additional categories divided into two groups are input to the feature extraction section 510, and the first deep feature vector and the second deep feature vector are extracted for each image data by the feature extraction section 510 (S540).

[0122] The first average deep feature vectors of the additional categories of the first group are summarized to obtain a first average deep feature matrix, and the second average deep feature vectors of the additional categories of the second group are summarized to obtain a second average deep feature matrix (S541).

[0123] The weight matrix of the additional categories of the first deep feature similarity calculation section 532a is replaced with the first average deep feature matrix (S542a).

[0124] The weight matrix of the additional categories of the second deep feature similarity calculation section 533a is replaced with the second average deep feature matrix (S542b).

[0125] The image data of the N (i = 1, 2,..., N) additional categories divided into two groups is input to the feature extraction section 510 for a cycle number of M (e = 1, 2,..., M) times. The cycle number is 30.

[0126] The first deep feature similarity calculation section 532a calculates the first deep cosine similarity from the input first deep feature vectors and the first average deep feature vectors of the categories, and outputs to the first deep similarity scaling section 540a (S543a).

[0127] The second deep feature similarity calculation section 533a calculates a second deep cosine similarity from the input second deep feature vector and the second average deep feature vector of each category, and outputs it to the second deep similarity scaling section 541a (S543b).

[0128] The first deep similarity scaling section 540a scales the input first deep cosine similarity by the first deep learning parameter by a, and outputs the first deep cosine similarity (S544a).

[0129] The second deep similarity scaling section 541a scales the input second deep cosine similarity by the second deep learning parameter by a, and outputs the second deep cosine similarity (S544b).

[0130] The first deep loss operation section 552a calculates a loss of the first deep cosine similarity and the correct answer label (correct answer category) of the input image, that is, a first deep cross entropy loss (S545a).

[0131] The second deep loss operation section 553a calculates a loss of the second deep cosine similarity and the correct answer label (correct answer category) of the input image, that is, a second deep cross entropy loss (S545b).

[0132] The loss weighted addition section 554 calculates a total cross entropy loss L by weighted addition of the first deep cross entropy loss Ld1 and the second deep cross entropy loss Ld2 (S546). Here, λ is set to a predetermined value of 0 to 1.

[0133] L = (1 - λ) x Ld1 + λ x Ld2

[0134] The optimization section 556 optimizes the weight matrix of the feature similarity calculation section 532 by backpropagation using an optimization method such as stochastic gradient descent (SGD) or Adam, so as to minimize the total cross entropy loss.

[0135] As described above, by having the average feature vector of the plurality of groups as the weight vector, even for an image whose features cannot be expressed with one average feature vector, it is possible to express the features of the image by using the average feature vectors of the plurality of groups. Furthermore, here, it is assumed that the feature extraction section 510 extracts a deep feature vector, but it can also be assumed that the feature extraction section 510 extracts a shallow feature vector.

[0136] Figure 9 is a block diagram of an image classification device 580. Figure 9 The image classification device 580 of is constituted by the constituent elements necessary for classification by the image classification learning device 500. Figure 10 is a flowchart illustrating the detailed operation of the image classification device 580. Refer to Figure 9 and Figure 10The classification operation of the image classification device 580 will be described in detail.

[0137] The feature extraction unit 510 has the same structure and the same parameters as those of the feature extraction unit obtained at the time of learning the basic categories.

[0138] The weight matrix of the basic categories in the weight matrix of the deep feature similarity calculation unit 532a is replaced with the average deep feature matrix calculated by the calculation method shown in Figure 6 The weight matrix of the basic categories in the weight matrix of the deep feature similarity calculation unit 532a is replaced with the average deep feature matrix calculated by the calculation method shown in Figure 7A and Figure 8A The weight matrix of the basic categories in the weight matrix of the deep feature similarity calculation unit 532a is replaced with the average deep feature matrix calculated by the calculation method shown in Figure 6 The weight matrix of the basic categories in the weight matrix of the deep feature similarity calculation unit 532a is replaced with the average deep feature matrix calculated by the calculation method shown in Figure 7A and Figure 8A The weight matrix of the basic categories in the weight matrix of the deep feature similarity calculation unit 532a is replaced with the average deep feature matrix calculated by the calculation method shown in

[0139] When the input image is input to the feature extraction unit 510, the deep feature vector and the shallow feature vector are extracted (S550).

[0140] The deep feature similarity calculation unit 532a calculates the deep cosine similarity for each category from the input deep feature vector and the deep weight vector of each category possessed by the average deep feature matrix, and outputs to the deep similarity scaling unit 540a (S551a).

[0141] The deep feature similarity calculation unit 532a calculates the deep cosine similarity for each category from the input deep feature vector and the deep weight vector of each category possessed by the average deep feature matrix, and outputs to the deep similarity scaling unit 540a (S551a).

[0142] The deep similarity scaling unit 540a scales the input deep cosine similarity by the deep association parameter to β times, and outputs the deep cosine similarity of each category (S552a).

[0143] The deep similarity scaling unit 540a scales the input deep cosine similarity by the deep association parameter to β times, and outputs the deep cosine similarity of each category (S552a).

[0144] The deep similarity scaling unit 540a scales the input deep cosine similarity by the deep association parameter to β times, and outputs the deep cosine similarity of each category (S552a).

[0145] The comprehensive similarity calculation section 560 weights the total cosine similarity of the additional category (S554). The weighting parameter is set to w. If it is desired to make the accuracy of the additional category relatively higher than that of the basic category, w is set to >1.0, and if it is desired to make the accuracy of the basic category relatively higher than that of the additional category, w is set to <1.0. In the case where the accuracy of the basic category is made equal to that of the additional category, w is set to =1.0.

[0146] Here, the deep learning parameter at the time of learning and the shallow learning parameter are set to the same parameter a, and the deep association parameter β and the shallow association parameter γ at the time of classification are set to different parameters. Generally, the processing load is larger at the time of learning than at the time of classification. Therefore, the adjustment is not made at the time of learning the deep and shallow layers, but is made at the time of classification. Of course, a = β = γ can also be set, or a ≠ β = γ, a = γ ≠ β, a ≠ β ≠ γ can also be set. If there is no problem of processing efficiency, the deep learning parameter and the shallow learning parameter can also be set to different values to adjust at the time of learning.

[0147] In addition, in Figure 1 , 3 and 9, it is possible to omit the shallow feature similarity calculation section 532b, the shallow similarity scaling section 540b, the shallow loss operation section 552b, and the loss weighting addition section 554. In this case, the deep learning parameter and the deep association parameter are set to different values a ≠ β. By setting a in such a manner that β = 1, the scaling processing at the time of classification can be reduced. For example, a = 20, β = 1 is set such that the learning parameter a is larger than the association parameter β. A is set to be larger in order to increase the resolution of the cosine similarity at the time of learning. At the time of classification, since the deep cosine similarity that has been learned is used, scaling is not necessary, and scaling can be performed weakly.

[0148] The classification decision section 570 selects the category having the largest total cosine similarity from the total cosine similarity of each category (S555).

[0149] Figure 11 is a block diagram of the image additional classification device 590. Figure 11 The image additional classification device 590 of Figure 9 is a structure in which the Figure 7A structure of the image classification device 580 of Figure 12 is a flowchart illustrating the detailed operation of the image additional classification device 590. Refer to Figure 11 and Figure 12 the classification operation of the additional category of the image additional classification device 590 is described in detail.

[0150] The feature extraction section 510 has the same structure and the same parameters as the feature extraction section 510 obtained at the time of learning the basic categories.

[0151] The additional learning session s is repeated L times (s = 1, 2,..., L).

[0152] Image data of N (i = 1, 2,..., N) additional categories is input to the feature extraction section 510, and a deep feature vector and a shallow feature vector are extracted for each of the image data by the feature extraction section 510 (S560).

[0153] The average deep feature vectors of the additional categories are aggregated to obtain an average deep feature matrix, and the average shallow feature vectors of the additional categories are aggregated to obtain an average shallow feature matrix (S561).

[0154] The average deep feature matrix is substituted for the weight matrix of the additional categories of the deep feature similarity calculation section 532a (S562a).

[0155] The average shallow feature matrix is substituted for the weight matrix of the additional categories of the shallow feature similarity calculation section 532b (S562b).

[0156] When an input image is input to the feature extraction section 510, a deep feature vector and a shallow feature vector are extracted (S563).

[0157] The deep feature similarity calculation section 532a calculates a deep cosine similarity for each category from the input deep feature vector and the average deep feature vector of each category, and outputs to the deep similarity scaling section 540a (S564a).

[0158] The shallow feature similarity calculation section 532b calculates a shallow cosine similarity for each category from the input shallow feature vector and the average shallow feature vector of each category, and outputs to the shallow similarity scaling section 540b (S564b).

[0159] The deep similarity scaling section 540a scales the input deep cosine similarity by the deep association parameter to β times, and outputs the deep cosine similarity of each category (S565a).

[0160] The shallow similarity scaling section 540b scales the input shallow cosine similarity by the shallow association parameter to γ times, and outputs the shallow cosine similarity of each category (S565b).

[0161] The merged similarity calculation section 560 adds the deep cosine similarity and the shallow cosine similarity to calculate the total cosine similarity of each class (S566). Here, the deep cosine similarity and the shallow cosine similarity are simply added, but the calculation method is not limited thereto. For example, weighted addition or multiplication, or the like can be performed.

[0162] The comprehensive similarity calculation section 560 weights the total cosine similarity of the additional class (S567). The weighting parameter is set to w.

[0163] The classification decision section 570 selects a class having the largest total cosine similarity from the total cosine similarity of each class (S568). Here, the class having the largest total cosine similarity is selected in accordance with the total cosine similarity, but the selection method is not limited thereto. For example, a plurality of upper classes can be selected.

[0164] Thus, the additional class can be classified without requiring a very heavy loss calculation, optimization processing. In Figure 11 In the above, an example in which the average deep feature matrix and the average shallow feature matrix are calculated for the additional class assuming that the content of the basic class does not change is shown, but in a case where the content of the basic class is changed, the average deep feature matrix and the average shallow feature matrix can be calculated for the basic class.

[0165] Further, in each embodiment, the deep similarity scaling section 540a and the shallow similarity scaling section 540b and the first deep similarity scaling section 540a and the second deep similarity scaling section 541a are not essential. Either of the deep similarity scaling section 540a and the shallow similarity scaling section 540b can be omitted, or both of them can be omitted. Either of the first deep similarity scaling section 540a and the second deep similarity scaling section 541a can be omitted, or both of them can be omitted.

[0166] That is, in a case where the deep similarity scaling section 540a is not provided, the deep cosine similarity calculated by the deep feature similarity calculation section 532a is output to the deep loss operation section 552a or the merged similarity calculation section 560. In a case where the first deep similarity scaling section 540a is not provided, the first deep cosine similarity calculated by the first deep feature similarity calculation section 532a is output to the first deep loss operation section 552a. In a case where the shallow similarity scaling section 540b is not provided, the shallow cosine similarity calculated by the shallow feature similarity calculation section 532b is output to the shallow loss operation section 552b or the merged similarity calculation section 560. In a case where the second deep similarity scaling section 541a is not provided, the second deep cosine similarity calculated by the second deep feature similarity calculation section 533a is output to the second deep loss operation section 553a.

[0167] The various processes of the image classification learning apparatus 500, the image classification apparatus 580, and the image additional classification apparatus 590 described above can of course be implemented as an apparatus using hardware such as a CPU, a memory, and the like, and can also be implemented by firmware stored in a ROM (Read Only Memory), a flash memory, and the like, and software such as a computer. The firmware program and the software program can be provided recorded in a recording medium readable by a computer or the like, can be transmitted and received by a server through a wired or wireless network, and can also be transmitted and received as data broadcast of a terrestrial wave or satellite digital broadcast.

[0168] The present application has been described above based on the embodiments. It will be understood by those skilled in the art that the embodiments are illustrative, and that various modifications of the constituent elements and combinations of the processes can be made, and that such modifications are also within the scope of the present application.

[0169] Industrial applicability

[0170] The present application can be utilized in image classification technology.

[0171] Explanation of symbols

[0172] 500 image classification learning apparatus, 510 feature extraction section, 520a average deep feature calculation section, 520b average shallow feature calculation section, 530 classification section, 532a deep feature similarity calculation section, 532b shallow feature similarity calculation section, 540a deep similarity scaling section, 540b shallow similarity scaling section, 550 learning section, 552a deep loss operation section, 552b shallow loss operation section, 554 loss weighted addition section, 556 optimization section, 560 comprehensive similarity calculation section, 570 classification decision section, 580 image classification apparatus, 590 image additional classification apparatus.

Claims

1. An image classification apparatus characterized by comprising: comprising: a feature extraction section that outputs a first feature vector of an input image and outputs a second feature vector that is a different feature vector from the first feature vector; an average first feature calculation section that calculates an average first feature vector by averaging the first feature vectors of a certain class and that aggregates the average first feature vectors of all classes to obtain an average first feature matrix; an average second feature calculation section that calculates an average second feature vector by averaging the second feature vectors of a certain class and that aggregates the average second feature vectors of all classes to obtain an average second feature matrix; a first feature similarity calculation section that calculates a first similarity from the first feature vector of the input image and a first weight matrix; a second feature similarity calculation section that calculates a second similarity from the second feature vector of the input image and a second weight matrix, the average first feature calculation section substitutes the first weight matrix of the first feature similarity calculation section with the average first feature matrix, the average second feature calculation section substitutes the second weight matrix of the second feature similarity calculation section with the average second feature matrix.

2. The image classification apparatus according to claim 1, wherein the first feature vector and the second feature vector have different resolutions.

3. The image classification apparatus according to claim 1 or 2, further comprising: a comprehensive similarity calculation section that adds the first similarity and the second similarity to calculate a total similarity; and a classification decision section that decides a class of the input image based on the total similarity.

4. The image classification apparatus according to claim 1 or 2, further comprising: a first loss operation section that calculates a first loss from the first similarity and a correct answer label of the input image; a second loss operation section that calculates a second loss from the second similarity and the correct answer label of the input image; a loss weighted addition section that adds the first loss and the second loss to calculate a total loss; and an optimization section that optimizes the first weight matrix of the first feature similarity calculation section and the second weight matrix of the second feature similarity calculation section to minimize the total loss. comprising: a feature extraction step that outputs a first feature vector of an input image and outputs a second feature vector that is a different feature vector from the first feature vector; an average first feature calculation step that calculates an average first feature vector by averaging the first feature vectors of a certain class and that aggregates the average first feature vectors of all classes to obtain an average first feature matrix; 5. An image classification method characterized by, an average second feature calculation step that calculates an average second feature vector by averaging the second feature vectors of a certain class and that aggregates the average second feature vectors of all classes to obtain an average second feature matrix; a first feature similarity calculation step that calculates a first similarity from the first feature vector of the input image and a first weight matrix; a second feature similarity calculation step that calculates a second similarity from the second feature vector of the input image and a second weight matrix, the average first feature calculation step substitutes the first weight matrix of the first feature similarity calculation step with the average first feature matrix, the average second feature calculation step substitutes the second weight matrix of the second feature similarity calculation step with the average second feature matrix. a second feature similarity calculation step of calculating a second similarity from the second feature vector of the input image and a second weight matrix, in the average first feature calculation step, the first weight matrix of the first feature similarity calculation step is replaced by the average first feature matrix, in the average second feature calculation step, the second weight matrix of the second feature similarity calculation step is replaced by the average second feature matrix.

6. An image classification program characterized by comprising: comprising: a feature extraction step of outputting a first feature vector of an input image and outputting a second feature vector which is a different feature vector from the first feature vector; an average first feature calculation step of calculating an average first feature vector by averaging the first feature vectors of a certain class and summarizing the average first feature vectors of all classes to obtain an average first feature matrix; an average second feature calculation step of calculating an average second feature vector by averaging the second feature vectors of a certain class and summarizing the average second feature vectors of all classes to obtain an average second feature matrix; a first feature similarity calculation step of calculating a first similarity from the first feature vector of the input image and a first weight matrix; a second feature similarity calculation step of calculating a second similarity from the second feature vector of the input image and a second weight matrix, in the average first feature calculation step, the first weight matrix of the first feature similarity calculation step is replaced by the average first feature matrix, in the average second feature calculation step, the second weight matrix of the second feature similarity calculation step is replaced by the average second feature matrix.