Apparatus and method for performing classification using a classification model, and computer-readable storage medium

The classification apparatus and method enhance video image classification accuracy by using a pre-trained model with feature extraction, contribution calculation, and fusion, addressing low quality and occlusion issues, achieving up to 10% improvement in face recognition tasks.

JP7707754B2Active Publication Date: 2025-07-15FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2021137216
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-30
Filing Date
2021-08-25
Publication Date
2025-07-15
Estimated Expiration
2041-08-25

AI Technical Summary

Technical Problem

Existing object classification methods based on video images face challenges due to low quality and occlusions, leading to decreased performance.

Method used

A classification apparatus and method using a pre-trained classification model with feature extraction, contribution calculation, and feature fusion units to enhance classification accuracy by considering the contribution of each image in a group, employing a contribution loss function during training to account for global information.

Benefits of technology

Improves classification accuracy by up to 10% compared to conventional methods, particularly in video-based human face recognition tasks, by accurately determining the influence of each image on the classification result and fusing features based on contributions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007707754000010
    Figure 0007707754000010
  • Figure 0007707754000011
    Figure 0007707754000011
  • Figure 0007707754000012
    Figure 0007707754000012
Patent Text Reader

Abstract

To provide a device and a method for performing classification using a pre-trained classification model and a computer-readable storage medium.SOLUTION: A device includes: a feature extraction unit that extracts, using a feature extraction layer of a pre-trained classification model, features of each image among a plurality of images included in a target image group waiting for being classified; a contribution calculation unit that calculates, using a contribution calculation layer of the pre-trained classification model, contribution of each image among the plurality of images to a classification result of the target image group; a feature fusion unit that fuses, on the basis of contribution of the plurality of images calculated by the contribution calculation unit, features of the plurality of images extracted by the feature extraction unit and acquires the features after the fusion as the features of the target image group; and a classification unit that classifies the target image group on the basis of the features of the target image group.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing, and more particularly, to an apparatus and method for performing classification using a classification model and a computer-readable storage medium.

Background Art

[0002] For example, object classification based on a group of video images (e.g., face recognition) has been widely applied in fields such as video surveillance and security authentication, and thus has been attracting increasing attention in the academic and industrial communities. Different from object classification based on static images, since the video image quality is relatively low, for example, the change in the posture of the object is large, and blocking (occlusion) is likely to occur, the classification performance may decrease.

Summary of the Invention

Problems to be Solved by the Invention

[0003] In view of the above problems, an object of the present invention is to provide an apparatus and method for training a classification model and an apparatus and method for performing classification using the classification model, which can solve one or more drawbacks existing in the prior art.

Means for Solving the Problems

[0004] According to one aspect of the present invention, there is provided an apparatus for performing classification using a pre-trained classification model, the apparatus including: a feature extraction unit that extracts features of each image among a plurality of images included in a target image group waiting for classification, using a feature extraction layer of the pre-trained classification model; a contribution calculation unit that calculates a contribution of each image among the plurality of images to a classification result of the target image group, using a contribution calculation layer of the pre-trained classification model; a feature fusion unit that performs fusion on the features of the plurality of images extracted by the feature extraction unit based on the contributions of the plurality of images calculated by the contribution calculation unit, and obtains the fused features as the features of the target image group; and It includes a classification unit that classifies the target image group based on the characteristics of the target image group.

[0005] According to another aspect of the present invention, a method for classification using a pre-trained classification model is provided. The method includes: A feature extraction step of extracting the features of each image among a plurality of images included in a target image group waiting for classification using a feature extraction layer of the pre-trained classification model; A contribution calculation step of calculating the contribution of each image among the plurality of images to the classification result of the target image group using a contribution calculation layer of the pre-trained classification model; A feature fusion step of fusing the features of the plurality of images extracted in the feature extraction step based on the contributions of the plurality of images calculated in the contribution calculation step, and obtaining the fused features as the features of the target image group; and A classification step of classifying the target image group based on the features of the target image group.

[0006] According to another aspect of the present invention, a computer program code and a computer program product for realizing the above-mentioned method according to the present invention, and a computer-readable storage medium storing the computer program code for realizing the above-mentioned method according to the present invention are provided.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 4A

Figure 4B

Figure 5

Figure 6

Figure 7

Best Mode for Carrying Out the Invention

[0008] Hereinafter, with reference to the accompanying drawings, preferred embodiments for carrying out the present invention will be described in detail. Note that such embodiments are merely illustrative and do not limit the present invention.

[0009] First, with reference to FIGS. 1 and 2, a realization example of an apparatus that performs classification using a pre-trained classification model in an embodiment of the present invention will be described. FIG. 1 is a block diagram of an example of the functional configuration of an apparatus 100 that performs classification using a pre-trained classification model in an embodiment of the present invention. FIG. 2 is a block diagram of an example of the architecture of one specific realization method of an apparatus that performs classification using a pre-trained classification model in an embodiment of the present invention.

[0010] As shown in FIGS. 1 and 2, in an embodiment of the present invention, an apparatus 100 for performing classification using a pre-trained classification model may include a feature extraction unit 102, a contribution calculation unit 104, a feature fusion unit 106, and a classification unit 108.

[0011] The feature extraction unit 102 may be configured to extract features of each of a plurality of images included in a target image group waiting for classification, using a feature extraction layer of a pre-trained classification model. For example, one target image group may correspond to one video segment. In such a case, one target image group may include all or some of the frames of the corresponding video segment. Also, for example, each image included in the same target image group may all be related to the same object. However, the same target image group may include a plurality of images related to two or more objects.

[0012] Also, for example, as shown in FIG. 2, the target image group may consist of human face images. For example, the target image group may be all or some of the frames of a video segment including a human face. However, the target image group is not limited thereto, and the target image group may include other images, but detailed description thereof is omitted here.

[0013] The pre-trained classification model may be any suitable pre-trained classification model, for example, a pre-trained deep learning network model such as a pre-trained convolutional neural network model.

[0014] FIG. 2 is a diagram showing an architecture example of one specific implementation manner of the apparatus 100 in an embodiment of the present invention when a pre-trained convolutional neural network model is adopted as a pre-trained classification model. As shown in FIG. 2, the feature extraction layer of the pre-trained classification model may include one or more convolutional layers C of the convolutional neural network model and one fully connected layer FC1. Note that the feature extraction layer of the pre-trained classification model is not limited to the example shown in FIG. 2, and those skilled in the art may set the corresponding feature extraction layer according to actual needs, but the detailed description thereof is omitted here.

[0015] The contribution calculation unit 104 may be configured to calculate the contribution of each of the above-mentioned multiple images to the classification result of the target image group by using the contribution calculation layer of the pre-trained classification model. For example, the contribution may represent the degree of influence of each image on the classification result of the target image group, for example, the degree of positive influence. For example, for a certain image, the greater the degree of positive influence of the image on the classification result of the target image group, or the greater the possibility that the target image group is accurately classified by the image, the greater the contribution of the image.

[0016] As shown in FIG. 2, when a pre-trained convolutional neural network model is adopted as a pre-trained classification model, the contribution calculation layer may include one or more convolutional layers C of the convolutional neural network model and one fully connected layer FC2. Note that the contribution calculation layer of the pre-trained classification model is not limited to the example shown in FIG. 2. For example, the contribution calculation layer may include only one fully connected layer FC2. In addition, those skilled in the art may set the corresponding feature extraction layer according to actual needs, but the detailed description thereof is omitted here.

[0017] In addition, although FIG. 2 shows that the contribution calculation unit 104 calculates the contribution of an image based on the features of the image at a certain stage by feature extraction, in actual applications, the contribution calculation unit 104 can also directly calculate the contribution of an image based on the images included in the target image group.

[0018] Also, as can be understood by those skilled in the art, the structures and parameters of the different convolutional layers and fully connected layers shown in FIG. 2 may be different.

[0019] The feature fusion unit 106 may be configured to perform fusion on the features of a plurality of images extracted by the feature extraction unit 102 based on the contributions of the plurality of images calculated by the contribution calculation unit 104, and obtain the fused features as the features of the target image group.

[0020] The classification unit 108 may be configured to classify the target image group based on the features of the target image group. For example, the classification unit 108 can recognize the target image group based on the features of the target image group.

[0021] According to an embodiment of the present invention, the feature fusion unit 106 may further perform weighted averaging on the features of a plurality of images extracted by the feature extraction unit 102 based on the contributions of the plurality of images included in the target image group calculated by the contribution calculation unit 104, and use the obtained result as the features of the target image group. For example, when the target image group corresponds to a video segment, the features of the target image group can be referred to as "video layer features".

[0022] For example, the feature fusion unit 106 can obtain the features F of the target image group based on the following formula (1). V can be obtained.

Equation

[0023] In formula (1), f1, f2 and f m respectively represent the features extracted by the feature extraction unit 102 of the first image I1, the second image I2 and the m-th image I in the target image group. m of the target image group, and w irepresents the contribution calculated by the contribution calculation unit 104 for the i-th image in the corresponding target image group.

[0024] For example, according to an embodiment of the present invention, the feature fusion unit 106 further performs fusion on the features of one or more of the above-mentioned images based on the contributions of one or more images among the plurality of images included in the target sample, the contributions of which are equal to or greater than a predetermined threshold, and can be configured such that the fused features are the features of the target image group. For example, the feature fusion unit 106 can perform a weighted average on the features of one or more of the above-mentioned images based on the contributions of one or more images among the plurality of images included in the target sample, the contributions of which are equal to or greater than a predetermined threshold, and obtain the fused features as the features of the target image group.

[0025] In addition, the exemplary method of obtaining the features of the target image group by the feature fusion unit 106 performing fusion on all or some of the sample image features included in the target image group has been described above. However, the method of obtaining the features of the target image group is not limited to the above exemplary method, and those skilled in the art may adopt an appropriate method according to actual needs to obtain the features of the target image group. For example, the features of the image with the largest contribution in the target image group can also be used as the features of the target image group.

[0026] As described above, in the embodiment of the present invention, the apparatus 100 for performing classification using a pre-trained classification model calculates the contribution of each image in the target image group, and performs fusion on the features of each image in the target image group based on the calculated contribution, so that classification can be performed on the target image group based on the fused features. Compared with the prior art that simply performs classification on the target image group based on the average value of the features of each image included in the target image group, the apparatus 100 according to the embodiment of the present invention considers the contribution of the corresponding image in the target image group to the classification result, and performs classification on the target image group based on the features of one or more images in the target image group, thereby improving the accuracy of classification.

[0027] According to the experimental-based analysis, the contribution of an image is related to the quality of the image. The higher the quality of the image, the greater the corresponding contribution. It should be noted that the contribution of an image is not equivalent to the quality of the image. For example, as described above, the contribution can represent the degree of influence of each image on the classification result of the target image group, for example, the degree of positive influence.

[0028] According to the embodiment of the present invention, for each image among the plurality of images included in the target image group, the contribution of the image to the classification result of the target image group can be represented by a scalar. For example, the contribution of each image can be represented by a numerical value greater than 0. For example, the contribution of each image can be represented by a numerical value within a predetermined range (for example, 0 to 20). This predetermined range can be determined based on experience or experiments.

[0029] Alternatively, according to an embodiment of the present invention, for each image among a plurality of images included in a target image group, the contribution of the image to the classification result of the target image group includes the contribution of the features of each dimension of the image to the classification result of the target image group. For example, when the features of a certain image are N-dimensional (for example, 512-dimensional), the contribution of the image can be represented by a single N-dimensional contribution vector. Here, each element in the contribution vector represents the contribution of each dimension of the features of the corresponding image to the classification result. By calculating the contribution for each dimension of the features of the image, for example, the classification accuracy can be further improved.

[0030] According to an embodiment of the present invention, a pre-trained classification model may be obtained by training an initial classification model in the following manner using a training sample set including at least one sample image group, that is, extracting the features of each sample image in the at least one sample image group using the feature extraction layer of the initial classification model; for each sample image group, calculating the contribution of each sample image included in the sample image group to the classification result of the sample image group using the contribution calculation layer of the initial classification model; for each sample image group, performing fusion on the features of each sample image of the sample image group based on the contribution of each sample image of the sample image group to obtain the fused features as the features of the sample image group; and using the features of each sample image group, training the initial classification model based on the loss function of the initial classification model until a predetermined convergence condition is satisfied to obtain a pre-trained classification model.

[0031] For example, the predetermined convergence condition may be any one of the following, that is, the training has reached a predetermined number of times; minimization of the loss function; and the loss function is below a predetermined threshold.

[0032] As an example, an initial classification model can be generated based on any suitable untrained classification model. Alternatively, for example, an initial classification model can also be generated based on any suitable conventional trained classification model (e.g., VGGnet model, Resnet model, etc.). For example, one branch can be added to the conventional trained classification model as a contribution calculation layer. By generating the initial classification model based on the conventional trained classification model, the training process can be simplified. As an example, in the training process of the initial classification model, the parameters of the feature extraction layer of the initial classification model may be fixed. For example, this can simplify the training process. However, in the training process of the initial classification model, the parameters of the feature extraction layer of the initial classification model do not have to be fixed.

[0033] According to an embodiment of the present invention, the loss function may include a classification loss function representing the classification loss of the initial classification model. For example, a loss function similar to Softmax may be adopted as the classification loss function. For example, the classification loss function L id can be represented by the following formula (2).

Equation

[0034] In formula (2), N represents the number of groups of sample images in one mini-batch, θ represents the angle between the features of the group of sample images and their corresponding weights, and s and m are a scaling factor and an edge factor, respectively. The definitions of the parameters in formula (2) are almost the same as the definitions of the corresponding parameters in Reference 1 (ArcFace: Additive Angular Margin Loss for Deep Face Recognition), except for the definition of θ. Note that in Reference 1, θ represents the angle between the features of the sample image and its corresponding weight. However, as described above, in formula (2), θ represents the angle between the features of the group of sample images (e.g., video layer features) and their corresponding weights.

[0035] As described above, by training the initial classification model using the classification loss function, the true value of the contribution or quality of the training data set (i.e., the sample image group) is not required. This can significantly reduce the cost required for preparing the training data set.

[0036] Alternatively, according to an embodiment of the present invention, the loss function may include a classification loss function and a contribution loss function. Here, the contribution loss function can be used to represent the distance between the features of each sample image group and the feature center of the class to which the corresponding sample image group is classified. For example, the loss function L can be represented by the following formula (3).

Number

[0037] In formula (3), λ≥0, which represents a trade-off factor. The larger λ is, the greater the proportion occupied by the contribution loss function L in the training process. c For example, the contribution loss function L c can be represented by the following formula (4).

Number

[0038] In formula (4) (Outer 1) TIFF0007707754000005.tif17170 represents the features of the i-th sample image group, (Outer 2) TIFF0007707754000006.tif15170 represents the center of the features of the class y to which the i-th sample image group in the training sample set or the training sample subset is classified. i During the training process, (Outer 3) TIFF0007707754000007.tif17170 can be updated in real time. For example, (Outer 4) TIFF0007707754000008.tif16170 represents the center of the features for class y in the training sample set i When representing the center of the features of class y, among the groups of sample images used in the training process in the training sample set, for one or more groups of sample images classified into class y i it can be obtained by calculating the average of the features (e.g., video layer features) of these groups of sample images. Also, for example, (Outer 5) TIFF0007707754000009.tif14170 represents the center of the features for class y in the training sample subset i When representing the center of the features of class y, among the groups of sample images used in the training process in the training sample subset, for one or more groups of sample images classified into class y i it can be obtained by calculating the average of the features (e.g., video layer features) of these groups of sample images.

[0039] In the training process of the conventional classification model, in view of problems such as the processing capacity of the training device, usually, a training method following mini - batch is adopted, so global information is ignored. As described above, by introducing the contribution loss function in the training process, considering the global information obtained from the training sample set or the training sample subset, training the classification model can, for example, improve the accuracy of the obtained trained classification model.

[0040] To better explain the advantageous effects brought by the introduction of the contribution loss function, hereinafter, based on FIGS. 3A to 3C, together with an example of video - based human face recognition, this advantageous effect will be described.

[0041] FIG. 3A is a diagram showing a training sample subset of a predetermined class T (i.e., a specific human), as well as the actual feature distribution and contribution distribution of each sample image in the training sample subset. FIGS. 3B and 3C are diagrams showing the feature distribution and contribution distribution of a plurality of sample images (i.e., a plurality of sample images included in one mini-batch) used in one training process in the above-mentioned training sample subset, respectively, in the case where no contribution loss function is introduced and in the case where a contribution loss function is introduced.

[0042] In FIGS. 3A to 3C, "●", "▲", and "★" represent sample images. Here, the actual contribution of the sample image represented by "●" is relatively low, and the actual contribution of the sample image represented by "★" is relatively high. Also, "■" represents the actual feature distribution center of class T, and "◆" represents the feature distribution center of class T obtained by calculation in one training process. Also, since the sample image represented by "★" is not used in the training process corresponding to FIGS. 3B and 3C, "★" is not shown in FIGS. 3B and 3C. As can be seen from FIGS. 3A to 3C, compared with the case where no contribution loss function is introduced, when a contribution loss function is introduced, the feature distribution center of class T obtained by calculation in one training process is closer to the actual feature distribution center of class T, and the contribution of each sample image obtained by calculation in one training process is closer to its actual contribution. Therefore, by introducing a contribution loss function, the contribution of each sample image can be calculated more accurately, and for example, the classification accuracy of the obtained pre-trained classification model can be improved.

[0043] Hereinafter, based on FIGS. 4A and 4B, together with a specific example of human face recognition based on a video, the advantageous effects in terms of the classification accuracy of the apparatus 100 for performing classification using a pre-trained classification model in the embodiments of the present invention will be described. In FIGS. 4A and 4B, the pre-trained classification model adopted by the apparatus 100 according to the embodiments of the present invention is a classification model based on ResNet50, and the pre-trained classification model is represented by "CAN".

[0044] FIG. 4A shows a comparison between the classification accuracy of the device 100 according to an embodiment of the present invention and the classification accuracy of a device based on ArcFace when using the NIST IJB-C dataset. As can be seen from FIG. 4A, when the FAR (False Accept Rate) = 0.001%, the TAR (True Accept Rate) of the device 100 according to the embodiment of the present invention is improved by about 7% compared to the device based on ArcFace.

[0045] FIG. 4B shows a comparison between the classification accuracies of the device 100 according to an embodiment of the present invention, a device based on VGG Face, and a device based on TBE-CNN in the case of the COX face dataset. In FIG. 4B, V2S_1, V2S_2, and V2S_3 respectively show the face recognition results when video capture is performed using different imaging devices. As can be seen from FIG. 4B, in the case of V2S_1, the recognition rate of the device 100 according to the embodiment of the present invention is improved by about 10% and about 5% respectively compared to the device based on VGG Face and the device based on TBE-CNN.

[0046] As described above, an apparatus for performing classification using a pre-trained classification model in an embodiment of the invention has been described. Hereinafter, corresponding to the above-described apparatus embodiments, the present invention provides an embodiment of a method for performing classification using a pre-trained classification model.

[0047] FIG. 5 is a flowchart of an exemplary flow of a method 500 for performing classification using a pre-trained classification model in an embodiment of the present invention. As shown in FIG. 5, the method 500 for performing classification using a pre-trained classification model in an embodiment of the present invention starts at a start step S502 and ends at an end step S512. The method 500 according to an embodiment of the present invention may include a feature extraction step S504, a contribution calculation step S506, a feature fusion step S508, and a classification step S510.

[0048] In the feature extraction step S504, the feature extraction layer of the pre-trained classification model can extract the features of each of the plurality of images included in the target image group waiting for classification. For example, one target image group may correspond to one video segment. In such a case, one target image group may include all or some of the frames of the corresponding video segment. For example, since the feature extraction step S504 can be implemented by the above-described feature extraction unit 102, the specific details thereof are omitted here.

[0049] In the contribution calculation step S506, the contribution calculation layer of the pre-trained classification model can calculate the contribution of each of the above-mentioned plurality of images to the classification result of the target image group. For example, the contribution can represent the degree of influence of each image on the classification result of the target image group, for example, the degree of positive influence. For example, for a certain image, the greater the degree of positive influence of the image on the classification result of the target image group, the greater the contribution of the image. For example, since the contribution calculation step S506 can be implemented by the above-described contribution calculation unit 104, the specific details thereof are omitted here.

[0050] In the feature fusion step S508, based on the contributions of the plurality of images included in the target image group calculated in the contribution calculation step S506, fusion is performed on the features of the plurality of images included in the target image group extracted in the feature extraction step S504, so that the fused features can be obtained as the features of the target image group. For example, since the feature fusion step S508 can be implemented by the above-described feature fusion calculation unit 106, the specific details thereof are omitted here.

[0051] In the classification step S510, classification can be performed on the target image group based on the characteristics of the target image group. For example, in the classification step S510, recognition can be performed on the target image group based on the characteristics of the target image group. Also, for example, since the classification step S510 can be implemented by the classification unit 108 described above, detailed description thereof is omitted here.

[0052] According to an embodiment of the present invention, in the feature fusion step S508, based on the contributions of the plurality of images included in the target image pixels calculated in the contribution calculation step S506, weighted averaging is performed on the features of the plurality of images extracted in the feature extraction step S504, and the obtained result can be used as the feature of the target image group. For example, in the feature fusion step S508, the feature F of the target image group can be obtained by the above formula (1). V can be obtained.

[0053] Alternatively, according to an embodiment of the present invention, in the feature fusion step S508, based on the contributions of one or more images among the plurality of images included in the target sample whose contributions are equal to or greater than a predetermined threshold, by performing fusion on the features of the above one or more images, the fused feature can be obtained as the feature of the target image group. For example, in the feature fusion step S508, based on the contributions of one or more images among the plurality of images included in the target sample whose contributions are equal to or greater than a predetermined threshold, by performing weighted averaging on the features of the above one or more images, the fused feature can be obtained as the feature of the target image group.

[0054] As described above, similar to the apparatus 100 for performing classification using a pre-trained classification model in an embodiment of the present invention, a method 500 for performing classification using a pre-trained classification model in an embodiment of the present invention calculates the contribution of each image in a target image group, and performs fusion on the features of each image in the target image group based on the calculated contribution, so that classification can be performed on the target image group based on the fused features. Compared with the prior art that performs classification on a target image group based on, for example, the average value of the features of each image included in the target image group or the features of the image with the best quality in the target image group, the method 500 according to the embodiment of the present invention can improve the classification accuracy by considering the contribution of each image to the classification result and performing classification on the target image group based on the features of one or more images in the target image group.

[0055] According to an embodiment of the present invention, for each image among a plurality of images included in a target image group, the contribution of the image to the classification result of the target image group can be represented by a scalar. For example, the contribution of each image can be represented by a numerical value greater than 0.

[0056] According to an embodiment of the present invention, for each image among a plurality of images included in a target image group, the contribution of the image to the classification result of the target image group includes the contribution of the features of each dimension of the image to the classification result of the target image group. For example, when the features of a certain image are N-dimensional (for example, 512-dimensional), the contribution of the image can be represented by an N-dimensional contribution vector. Here, each element in the contribution vector represents the contribution of each dimension of the features of the corresponding image to the classification result. By calculating the contribution for each dimension of the features of the image, for example, the classification accuracy can be further improved.

[0057] According to an embodiment of the present invention, a pre-trained classification model may be obtained by training an initial classification model in the following manner using a training sample set including at least one sample image group, that is, extracting features of each sample image in the at least one sample image group using a feature extraction layer of the initial classification model; for each sample image group, calculating the contribution of each sample image included in the sample image group to the classification result of the sample image group using a contribution calculation layer of the initial classification model; for each sample image group, performing fusion on the features of each sample image in the sample image group based on the contributions of each sample image in the sample image group to obtain the fused features as the features of the sample image group; and using the features of each sample image group to train the initial classification model based on a loss function for the initial classification model until a predetermined convergence condition is satisfied, thereby obtaining a pre-trained classification model.

[0058] For example, the predetermined convergence condition may be any one of the following, that is, the training has reached a predetermined number of times; minimization of the loss function; and the loss function is less than or equal to a predetermined threshold.

[0059] According to an embodiment of the present invention, the loss function may include a classification loss function and a contribution loss function. Here, the contribution loss function can be used to represent the distance between the features of each sample image group and the feature center of the class to which the corresponding sample image group is classified. For example, the loss function L can be represented by the above formula (3).

[0060] In the training process of a conventional classification model, in view of problems such as the processing capacity of the training device, usually, a method of training according to mini-batch is adopted, so global information is ignored. As described above, by introducing a contribution loss function in the training process, training the classification model in consideration of the global information obtained from the training sample set or a subset of the training samples can improve, for example, the classification accuracy of the obtained trained classification model.

[0061] The above described the embodiments of the apparatus 100 and method 500 for classification using the pre-trained classification model in the embodiments of the present invention. Further, according to the present invention, an apparatus for training the initial training may be provided. FIG. 6 is a block diagram of a functional configuration example of an apparatus 600 for training the initial classification model in an embodiment of the present invention.

[0062] As shown in FIG. 6, the apparatus 600 for training the initial classification model in the embodiments of the present invention may include a second feature extraction unit 602, a second contribution calculation unit 604, a second feature fusion unit 606, and a training unit 608.

[0063] The second feature extraction unit 602 may be configured to extract features of each sample image in at least one sample image group included in the training sample set using the feature extraction layer of the initial classification model.

[0064] The second contribution calculation unit 604 may be configured to calculate, for each sample image group, the contribution of each sample image included in the sample image group to the classification result of the sample image group using the contribution calculation layer of the initial classification model.

[0065] The second feature fusion unit 606 may be configured to, for each sample image group, perform fusion on the features of each sample image in the sample image group extracted by the second feature extraction unit 602 based on the contribution of each sample image in the sample image group calculated by the second contribution calculation unit 604, so as to obtain the fused features as the features of the sample image group.

[0066] The training unit 608 may be configured to use the features of each sample image group to train the initial classification model based on the loss function for the initial classification model until a predetermined convergence condition is satisfied, so as to obtain a pre-trained classification model.

[0067] Regarding the details of the training of the apparatus 600 according to the embodiment of the present invention for the initial classification model, it is the same as the details of the training for the initial classification model in the apparatus 100 and the method 500 that perform classification using the pre-trained classification model in the above-described embodiments. Therefore, the detailed description thereof is omitted here.

[0068] The apparatus 600 for training the initial classification model in the embodiment of the present invention has strong versatility and can be easily applied to any appropriate initial classification model. Further, the apparatus 600 for training the initial classification model in the embodiment of the present invention can improve the classification accuracy of the obtained pre-trained classification model by training the initial classification model based on one or more images in the corresponding sample image group in consideration of the contribution of each sample image.

[0069] Note that above, the apparatus and method for performing classification using the pre-trained classification model in the embodiment of the present invention, and the functional arrangement and operations of the apparatus for training the initial classification model have been described. However, these are merely examples, and those skilled in the art can make changes to the above-described embodiments based on the principles of the present invention, for example, increasing or decreasing and combining functional modules and operations in each embodiment. Also, all such changes belong to the scope of the present invention.

[0070] Also, since the method embodiments here correspond to the apparatus embodiments above, for the content not described in detail in the method embodiments, reference can be made to the description of the corresponding parts in the apparatus embodiments. Therefore, the detailed description thereof is omitted here.

[0071] Further, the present invention further provides a storage medium and a program product. It should be understood that the machine-executable instructions in the storage medium and the program product according to the embodiment of the present invention can be further configured to implement the method for performing classification using the pre-trained classification model described above. Therefore, for the content not described in detail here, reference can be made to the description of the corresponding parts above. Therefore, the detailed description thereof is omitted here.

[0072] As is apparent, each operation process of the method according to the present invention can be realized by a computer-executable program stored in various machine-readable storage media.

[0073] Further, the object of the present invention may also be realized in the following manner, that is, a storage medium storing the above-described executable program code is directly or indirectly provided to a system or device, and a computer or a central processing unit (CPU) in the system or device reads and executes the above-described program code. At this time, if the system or device has a function capable of executing the program, the implementation manner of the present invention is not limited to the program, and the program may be in any form, for example, an object-oriented program, a program executed by an interpreter, a script program provided to an operating system, or the like.

[0074] These machine-readable storage media may include, but are not limited to, various memories and storage units, semiconductor devices, magnetic disks such as optical, magnetic, magneto-optical disks, and other media suitable for storing information.

[0075] Further, the above-described series of processes and devices can be realized by software and / or firmware. When realized by software and / or firmware, a program constituting the software is installed from a storage medium or a network into a computer having a dedicated hardware configuration, for example, a general-purpose personal computer 700 shown in FIG. 7, and the computer can execute various functions when various programs are installed.

[0076] FIG. 7 is a configuration diagram of a hardware configuration (general-purpose machine) 700 capable of realizing an information processing method and apparatus in an embodiment of the present invention.

[0077] The general-purpose machine 700 may be, for example, a computer system. Note that the general-purpose machine 700 is merely an example and does not limit the scope of application or functions of the method and apparatus according to the present invention. Also, the general-purpose machine 700 does not depend on any module, assembly, etc. in the above-described method and apparatus or their combination.

[0078] In FIG. 7, the central processing unit (CPU) 701 performs various processes based on a program stored in the ROM 702 or a program loaded from the storage unit 708 into the RAM 703. In the RAM 703, data and the like necessary when the CPU 701 performs various processes can also be stored according to needs.

[0079] The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The input / output interface 705 is also connected to the bus 704.

[0080] In addition, the following components are further connected to the input / output interface 705, that is, an input unit 706 including a keyboard and the like, a display unit including a liquid crystal display (LCD) and the like, an output unit 707 including a speaker and the like, a storage unit 708 including a hard disk and the like, and a communication unit 709 including a network interface card, for example, a LAN card, a modem, and the like. The communication unit 709 performs communication processing via a network such as the Internet or a LAN.

[0081] The drive 710 may be connected to the input / output interface 705 according to needs. A removable medium 711, for example, a semiconductor memory, etc., can be set in the drive 710 as needed, and the computer program read from it can be installed in the storage unit 708.

[0082] Furthermore, the present invention further provides a program product including machine-readable instruction codes. When such instruction codes are read and executed by a machine, the method in the above-described embodiment of the present invention can be executed. Correspondingly, various storage media that carry such a program product, such as magnetic disks (including floppy disks (registered trademark)), optical disks (including CD-ROMs and DVDs), magneto-optical disks (including MDs (registered trademark)), and semiconductor memories, are also included in the present invention.

[0083] The above-described storage media may include, for example, magnetic disks, optical disks, magneto-optical disks, semiconductor memories, etc., but are not limited thereto.

[0084] In addition, each operation (process) in the above-described method can also be realized in the form of a computer-executable program stored in various machine-readable storage media.

[0085] In addition, regarding the above embodiments, etc., the following is further disclosed as an additional note.

[0086] (Additional Note 1) An apparatus for performing classification using a pre-trained classification model, a feature extraction unit that extracts features of each image among a plurality of images included in a target image group waiting for classification using a feature extraction layer of the pre-trained classification model; a contribution calculation unit that calculates the contribution of each image among the plurality of images to the classification result of the target image group using a contribution calculation layer of the pre-trained classification model; a feature fusion unit that performs fusion on the features of the plurality of images extracted by the feature extraction unit based on the contributions of the plurality of images calculated by the contribution calculation unit, and obtains the fused features as the features of the target image group; and a classification unit that classifies the target image group based on the features of the target image group.

[0087] (Appendix 2) The apparatus according to Appendix 1, wherein the feature fusion unit further performs weighted averaging on the features of the plurality of images extracted by the feature extraction unit based on the contributions of the plurality of images calculated by the contribution calculation unit, and uses the obtained result as the feature of the target image group.

[0088] (Appendix 3) The apparatus according to Appendix 1, wherein the feature fusion unit further performs fusion on the features of one or more images among the plurality of images based on the contributions of the one or more images whose contributions are equal to or greater than a predetermined threshold, and obtains the fused features as the features of the target image group.

[0089] (Appendix 4) The apparatus according to any one of Appendices 1 to 3, wherein for each image among the plurality of images, the contribution of the image to the classification result of the target image group is represented by a scalar.

[0090] (Appendix 5) The apparatus according to any one of Appendices 1 to 3, wherein for each image among the plurality of images, the contribution of the image to the classification result of the target image group includes the contribution of the features of each dimension of the image to the classification result of the target image group.

[0091] (Appendix 6) The apparatus according to any one of Appendices 1 to 3, wherein the pre-trained classification model is obtained by training an initial classification model in the following manner using a training sample set including at least one sample image group, that is, extracting the features of each sample image in the at least one sample image group using the feature extraction layer of the initial classification model; For each sample image group, using the contribution calculation layer of the initial classification model, calculate the contribution of each sample image included in the sample image group to the classification result of the sample image group; For each sample image group, based on the contributions of the sample images in the sample image group, perform fusion on the features of the sample images in the sample image group, and obtain the fused features as the features of the sample image group; and An apparatus for training the initial classification model based on a loss function for the initial classification model using the features of each sample image group so that a predetermined convergence condition is satisfied to obtain the pre-trained classification model.

[0092] (Appendix 7) The apparatus according to Appendix 6, wherein the loss function includes a classification loss function and a contribution loss function, the classification loss function represents the classification loss of the initial classification model, and the contribution loss function represents the distance between the features of each sample image group and the feature center of the class to which the corresponding sample image group is classified. An apparatus.

[0093] (Appendix 8) The apparatus according to Appendix 6, wherein in the training process of the initial classification model, the parameters of the feature extraction layer of the initial classification model are fixed. An apparatus.

[0094] (Appendix 9) A method for performing classification using a pre-trained classification model, a feature extraction step of extracting the features of each image among a plurality of images included in a target image group waiting for classification using the feature extraction layer of the pre-trained classification model; a contribution calculation step of calculating the contribution of each image among the plurality of images to the classification result of the target image group using the contribution calculation layer of the pre-trained classification model; A feature fusion step of performing fusion on the features of the plurality of images extracted in the feature extraction step based on the contributions of the plurality of images calculated in the contribution calculation step, and obtaining the fused features as the features of the target image group; and A method including a classification step of classifying the target image group based on the features of the target image group.

[0095] (Appendix 10) The method according to Appendix 9, wherein In the feature fusion step, weighted averaging is performed on the features of the plurality of images extracted in the feature extraction step based on the contributions of the plurality of images calculated in the contribution calculation step, and the obtained result is used as the features of the target image group.

[0096] (Appendix 11) The method according to Appendix 9, wherein In the feature fusion step, fusion is performed on the features of one or more images among the plurality of images based on the contributions of the one or more images whose contributions are equal to or greater than a predetermined threshold, and the fused features are obtained as the features of the target image group.

[0097] (Appendix 12) The method according to any one of Appendices 9 to 11, wherein For each image among the plurality of images, the contribution of the image to the classification result of the target image group is represented by a scalar.

[0098] (Appendix 13) The method according to any one of Appendices 9 to 11, wherein For each image among the plurality of images, the contribution of the image to the classification result of the target image group includes the contribution of the features of each dimension of the image to the classification result of the target image group.

[0099] (Appendix 14) The method according to any one of Supplementary Notes 9 to 11, wherein the pre-training classification model is obtained by training an initial classification model in the following manner using a training sample set including at least one sample image group, that is, extracting the features of each sample image in the at least one sample image group using the feature extraction layer of the initial classification model; for each sample image group, calculating the contribution of each sample image included in the sample image group to the classification result of the sample image group using the contribution calculation layer of the initial classification model; for each sample image group, performing fusion on the features of each sample image of the sample image group based on the contribution of each sample image of the sample image group, and obtaining the fused features as the features of the sample image group; and using the features of each sample image group, training the initial classification model based on the loss function for the initial classification model so that a predetermined convergence condition is satisfied, and obtaining the pre-training classification model.

[0100] (Supplementary Note 15) The method according to Supplementary Note 14, wherein the loss function includes a classification loss function and a contribution loss function, the classification loss function represents the classification loss of the initial classification model, and the contribution loss function represents the distance between the features of each sample image group and the feature center of the class to which the corresponding sample image group is classified.

[0101] (Supplementary Note 16) The method according to Supplementary Note 14, wherein, in the training process of the initial classification model, the parameters of the feature extraction layer of the initial classification model are fixed.

[0102] (Supplementary Note 17) A computer-readable storage medium storing program instructions, When the program instruction is executed by a computer, it can execute the method described in any one of Appendices 9 to 16.

[0103] As described above, the preferred embodiments of the present invention have been explained. However, the present invention is not limited to these embodiments, and any changes to the present invention belong to the technical scope of the present invention as long as they do not depart from the gist of the present invention.

Claims

1. An apparatus for performing classification using a pre-trained classification model, comprising: a feature extraction unit that extracts features of each image among a plurality of images included in a target image group awaiting classification, using a feature extraction layer of the pre-trained classification model; a contribution calculation unit that calculates the contribution of each image among the plurality of images to the classification result of the target image group, using a contribution calculation layer of the pre-trained classification model; a feature fusion unit that performs fusion on the features of the plurality of images extracted by the feature extraction unit based on the contributions of the plurality of images calculated by the contribution calculation unit, and obtains the fused features as the features of the target image group; and a classification unit that performs classification on the target image group based on the features of the target image group, wherein the pre-trained classification model is obtained by training an initial classification model in the following manner using a training sample set including at least one sample image group, that is, extracting features of each sample image in the at least one sample image group using a feature extraction layer of the initial classification model; for each sample image group, calculating the contribution of each sample image included in the sample image group to the classification result of the sample image group, using a contribution calculation layer of the initial classification model; for each sample image group, performing fusion on the features of each sample image in the sample image group based on the contributions of each sample image in the sample image group, and obtaining the fused features as the features of the sample image group; and using the features of each sample image group, training the initial classification model based on a loss function for the initial classification model so that a predetermined convergence condition is satisfied, and obtaining the pre-trained classification model, wherein the loss function includes a classification loss function and a contribution loss function, the classification loss function represents the classification loss of the initial classification model, and the contribution loss function represents the distance between the features of each sample image group and the feature center of the class to which the corresponding sample image group is classified.

2. The apparatus according to claim 1, The feature fusion unit further performs a weighted average on the features of the plurality of images extracted by the feature extraction unit based on the contributions of the plurality of images calculated by the contribution calculation unit, and uses the obtained result as the feature of the target image group.

3. The apparatus according to claim 1, wherein the feature fusion unit further performs fusion on the features of one or more images among the plurality of images based on the contributions of the one or more images whose contributions are equal to or greater than a predetermined threshold value, and obtains the fused features as the features of the target image group.

4. The apparatus according to any one of claims 1 to 3, wherein, for each image among the plurality of images, the contribution of the image to the classification result of the target image group is represented by a scalar.

5. The apparatus according to any one of claims 1 to 3, wherein, for each image among the plurality of images, the contribution of the image to the classification result of the target image group includes the contributions of the features of each dimension of the image to the classification result of the target image group.

6. The apparatus according to claim 1, wherein, in the training process of the initial classification model, the parameters of the feature extraction layer of the initial classification model are fixed.

7. A method for performing classification using a pre-trained classification model, comprising: a feature extraction step of extracting features of each image among a plurality of images included in a target image group to be classified using a feature extraction layer of the pre-trained classification model; a contribution calculation step of calculating the contribution of each image among the plurality of images to the classification result of the target image group using a contribution calculation layer of the pre-trained classification model; a feature fusion step of performing fusion on the features of the plurality of images extracted in the feature extraction step based on the contributions of the plurality of images calculated in the contribution calculation step, and obtaining the fused features as the features of the target image group; and a classification step of performing classification on the target image group based on the features of the target image group, wherein the pre-trained classification model is obtained by training an initial classification model in the following manner using a training sample set including at least one sample image group, that is, Using the feature extraction layer of the initial classification model, extract the features of each sample image in the at least one sample image group; For each sample image group, using the contribution calculation layer of the initial classification model, calculate the contribution of each sample image included in the sample image group to the classification result of the sample image group; For each sample image group, based on the contributions of the sample images in the sample image group, perform fusion on the features of the sample images in the sample image group, and obtain the fused features as the features of the sample image group; and Using the features of each sample image group, based on the loss function for the initial classification model, train the initial classification model so that a predetermined convergence condition is satisfied, and obtain the pre-trained classification model. The loss function includes a classification loss function and a contribution loss function. The classification loss function represents the classification loss of the initial classification model. The contribution loss function represents the distance between the features of each sample image group and the feature center of the class to which the corresponding sample image group is classified. **Claim 8** A program for causing a computer to execute the method according to claim 7.