Image processing apparatus, image processing method, and image processing program

The image processing technology addresses catastrophic forgetting in CNNs by using two embedded learning units to combine embedding vectors and optimize neural networks, achieving effective reduction of forgetting and enabling continuous learning.

JP7683426B2Active Publication Date: 2025-05-27JVC KENWOOD CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021140818
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-08-31
Publication Date
2025-05-27
Estimated Expiration
2041-08-31

AI Technical Summary

Technical Problem

Existing Convolutional Neural Networks (CNNs) suffer from catastrophic forgetting, where the learning results of old tasks are forgotten during the learning of new tasks, leading to insufficient reduction of catastrophic forgetting.

Method used

An image processing apparatus and method that utilize two embedded learning units to continuously learn and optimize neural networks, combining embedding vectors from different neural networks to reduce catastrophic forgetting through regularization and metric losses.

Benefits of technology

The proposed solution effectively reduces catastrophic forgetting in CNNs, allowing for continuous learning while maintaining the accuracy of old tasks and adapting to new tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683426000001
    Figure 0007683426000001
  • Figure 0007683426000002
    Figure 0007683426000002
  • Figure 0007683426000003
    Figure 0007683426000003
Patent Text Reader

Abstract

To provide an image processing technology based on machine learning, with which it is possible to reduce fatal oblivion.SOLUTION: A schema embedding learning unit 210 continuously trains a first neural network. A deep embedding learning unit 220 continuously trains a second neural network. A synthesis unit 230 synthesizes, to input data, a first embedding vector derived from the first neural network of the schema embedding learning unit 210 and a second embedding vector derived from the second neural network of the deep embedding learning unit 220, and outputs a synthesized embedding vector. A classification unit 240 classifies the input data on the basis of the synthesized embedding vector. The schema embedding learning unit 210 and the deep embedding learning unit 220 convert the embedment outputted from the neural networks into mutually different embedding matrices and calculate regularization losses on the basis of the mutually different embedding matrices.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing technology based on machine learning.

Background Art

[0002] Humans can learn new knowledge through long-term experience and can maintain old knowledge without forgetting it. On the other hand, the knowledge of a Convolutional Neural Network (CNN) depends on the dataset used for learning, and in order to adapt to changes in the data distribution, it is necessary to relearn the CNN parameters for the entire dataset. In a CNN, as learning progresses for a new task, the estimation accuracy for an old task decreases. Thus, in a CNN, when continuous learning is performed, catastrophic forgetting, where the learning results of old tasks are forgotten during the learning of new tasks, cannot be avoided.

[0003] As a method for avoiding catastrophic forgetting, incremental learning or continual learning has been proposed. Incremental learning or continual learning is a learning method in which, when a new task or new data occurs, instead of learning the model from scratch, the currently learned model is improved and learned. One method of incremental learning is regularization-based incremental learning, which uses a regularization loss for learning (Patent Document 1).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Non-Patent Documents

[0005]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] In the technology described in Patent Document 1, there was a problem that sufficient catastrophic forgetting could not be reduced.

[0007] The present invention has been made in view of such a situation, and an object thereof is to provide an image processing technology based on machine learning capable of reducing catastrophic forgetting.

Means for Solving the Problems

[0008] To solve the above problems, an image processing apparatus according to an aspect of the present invention includes a first embedded learning unit that continuously learns a first neural network, a second embedded learning unit that continuously learns a second neural network, a first embedded vector derived from the first neural network of the first embedded learning unit, and a second embedded vector derived from the second neural network of the second embedded learning unit for input data, and a combining unit that combines the first embedded vector and the second embedded vector to output a combined embedded vector, and a classification unit that classifies the input data based on the combined embedded vector. The first and second embedded learning units include an embedding conversion unit that converts an embedding output from a neural network into an embedding matrix, a regularization loss calculation unit that calculates a regularization loss based on the embedding matrix, a metric loss calculation unit that calculates a metric loss based on the embedding matrix, and an optimization unit that optimizes the neural network based on the metric loss and the regularization loss. The embedding conversion unit of the first embedded learning unit and the embedding conversion unit of the second embedded learning unit convert the embedding into different embedding matrices, and the regularization loss calculation unit of the first embedded learning unit and the regularization loss calculation unit of the second embedded learning unit calculate the regularization loss based on different embedding matrices.

[0009] Another aspect of the present invention is an image processing method. This method includes a first embedding learning step of continuously training a first neural network, a second embedding learning step of continuously training a second neural network, and for input data, synthesizing a first embedding vector derived from the first neural network of the first embedding learning unit and a second embedding vector derived from the second neural network of the second embedding learning unit to output a synthesized embedding vector, and a classification step of classifying the input data based on the synthesized embedding vector. The first and second embedding learning steps include an embedding conversion step of converting an embedding output from a neural network into an embedding matrix, a regularization loss calculation step of calculating a regularization loss based on the embedding matrix, a metric loss calculation step of calculating a metric loss based on the embedding matrix, and an optimization step of optimizing the neural network based on the metric loss and the regularization loss. The embedding conversion step of the first embedding learning step and the embedding conversion step of the second embedding learning step convert the embedding into different embedding matrices, and the regularization loss calculation step of the first embedding learning step and the regularization loss calculation step of the second embedding learning step calculate the regularization loss based on different embedding matrices.

[0010] In addition, any combination of the above components, as well as those obtained by converting the expression of the present invention among methods, apparatuses, systems, recording media, computer programs, etc., are also effective as aspects of the present invention.

Effects of the Invention

[0011] According to the present invention, it is possible to provide an image processing technology based on machine learning that can reduce catastrophic forgetting.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8A

Figure 8B

Figure 9A

Figure 9B

Figure 10

Embodiments for Carrying Out the Invention

[0013] (First Embodiment) FIG. 1 is a configuration diagram of a machine learning device 100 according to the first embodiment. The machine learning device 100 includes a neural network processing unit 10 for learning target, an embedding conversion unit 20, a representative vector storage unit 30, a learned neural network processing unit 40, an embedding conversion unit 50, a loss calculation unit 60, a metric loss calculation unit 70, a regularization loss calculation unit 80, and an optimization unit 90.

[0014] In this embodiment, machine learning that combines continuous learning and metric learning is performed. Here, although an image is used as an example of input data, the input data is not limited to images. There is metric learning as a method for learning an embedding space (feature space) that takes into account the relationship between images (see, for example, Non-Patent Document 1). Metric learning is used in various fields such as information retrieval, data classification, and image recognition. Continuous learning that uses a regularization loss can be combined with metric learning that uses a metric loss.

[0015] FIG. 2 is a flowchart for explaining the overall flow of learning by the machine learning device 100. The configuration and operation of the machine learning will be described with reference to FIGS. 1 and 2.

[0016] First, the neural network processing unit 10 to be learned learns basic classes using a basic training dataset and updates the neural network to be learned (S10). Let the learning session i be 0. This is also called the initial session.

[0017] Subsequently, the embedding conversion unit 20 derives representative vectors of the basic classes using the basic training dataset based on the neural network to be learned and stores them in the representative vector storage unit 30 (S20).

[0018] Subsequently, the neural network processing unit 10 to be learned stores the neural network to be learned as a learned neural network and provides it to the learned neural network processing unit 40 (S30).

[0019] Next, the learning session i of the continuous learning is repeated N times (i = 1, 2,..., N) (S40). The learned neural network stored in the learning session (i - 1) is used as the neural network to be learned in the learning session i.

[0020] The neural network processing unit 10 to be learned continues to learn additional classes using the additional training dataset and updates the neural network to be learned (S50).

[0021] Subsequently, the embedding conversion unit 20 derives representative vectors of the additional classes using the additional training dataset based on the neural network to be learned and stores them in the representative vector storage unit 30 (S60).

[0022] Subsequently, the neural network processing unit 10 to be learned stores the neural network to be learned as a learned neural network and provides it to the learned neural network processing unit 40 (S70).

[0023] Increment i by 1 (S80), return to step S40, repeat steps S50 to S70 until i = N, and end if i exceeds N.

[0024] The configuration and operation of the continuous learning will be described in more detail.

[0025] The basic training dataset is a supervised dataset that includes a large number of basic classes (for example, about 100 to 1000 classes), and each class is composed of a large number of images (for example, 3000 images). It is assumed that the basic training dataset has a sufficient amount of data to independently learn a general classification task.

[0026] In contrast, the additional training dataset is a supervised dataset that includes a small number of additional classes (for example, about 2 to 10 classes), and each additional class is composed of a small number of images (for example, about 1 to 5 images). Input the training data consisting of three images: an anchor image belonging to a certain class, a positive image belonging to the same class as the anchor image, and a negative image belonging to a different class from the anchor image, into the neural network to be learned. Here, the reason for setting the number of minority classes to 2 is that even if the class to be learned is 1, a class that is not the target of learning as a negative image is required. Also, here, although it is assumed to be a small number of images, it can be a large number of images as long as it is a small number of classes.

[0027] Figure 3 is a diagram for explaining the structure of the neural network model to be learned in learning session i. The neural network to be learned is a deep neural network including a convolutional layer and a pooling layer. It has a configuration including CONV-1 to CONV-5 which are the convolutional layers of ResNet-18 shown in Figure 3, and outputs an embedding map of 7×7×512 dimensions (hereinafter, also simply referred to as "embedding").

[0028] The learned neural network has the same configuration as the neural network to be learned. The learned neural network is a neural network that has already completed learning in learning session (i - 1).

[0029] The embedding conversion unit 20 performs smoothing by global average pooling for each 7×7 embedding map output by the neural network to be learned, and outputs a 512-dimensional embedding vector. The embedding conversion unit 50 performs smoothing by global average pooling for each 7×7 embedding map output by the learned neural network, and outputs a 512-dimensional embedding vector. Since the embedding conversion unit 20 and the embedding conversion unit 50 have the same configuration and perform the same operation, they can also be made into one configuration.

[0030] The loss calculation unit 60 calculates the overall loss L by adding the metric loss Lml calculated by the metric loss calculation unit 70 and the regularization loss Lr calculated by the regularization loss calculation unit 80 as follows. L = Σ(Lml + Lr) Here, Σ indicates taking the sum over the input images.

[0031] The metric loss calculation unit 70 calculates the metric loss. Here, the triplet loss is used as the metric loss. The triplet loss Lml is calculated by the following formula based on the embedding vectors of the anchor image, the positive image, and the negative image. Lml = dp - dn + α Here, dp is the Euclidean distance between the embedding vectors of the anchor image and the positive image. dn is the Euclidean distance between the embedding vectors of the anchor image and the negative image. α is the offset.

[0032] The regularization loss calculation unit 80 calculates the regularization loss. The regularization loss Lr is the sum of the embedding loss Lre and the embedding vector loss Lrv. As shown in the following formula, the embedding loss Lre is the difference in the embeddings before and after the learning session when the image is input to the neural network, and the embedding vector loss Lrv is the difference in the embedding vectors before and after the learning session when the image is input to the neural network. Lr = Lre + Lrv Lre = ||E(i) - E(i−1)|| Lrv = ||V(i) - V(i−1)|| Here, E(i) is the embedding output by the neural network to be learned, and E(i−1) is the embedding output by the learned neural network. V(i) is the embedding vector output by the neural network to be learned, and V(i−1) is the embedding vector output by the learned neural network. ||·|| is a symbol indicating the calculation of the Frobenius norm.

[0033] The regularization loss is not optimized only for the additional training data to be learned in the learning session i, but is used to maintain the previous classification results for the training data learned in the previous learning session (i - 1). The embedding loss Lre is a loss considering the detailed characteristics of the image, and the embedding vector loss Lrv is a loss considering the general characteristics of the image.

[0034] The optimization unit 90 optimizes the weights of the neural network to be learned so as to minimize the overall loss L. When the optimization is completed, the overall loss L is reset.

[0035] The representative vector of the class is obtained as follows.

[0036] The neural network processing unit 10 to be learned sequentially inputs the images in the training data set into the neural network to be learned to derive an embedding, and the embedding conversion unit 20 converts the embedding into an embedding vector. The embedding conversion unit 20 calculates a representative vector for each class based on the embedding or the embedding vector, and stores it in the representative vector storage unit 30.

[0037] The embedding conversion unit 20 calculates the average of the embedding vectors of all the images belonging to the class as the representative vector of the class. The embedding conversion unit 20 associates the class with the calculated representative vector and stores it in the representative vector storage unit 30. The representative vector of the class may be the centroid of the embedding vectors of all the images belonging to the class, the average of the embeddings of all the images belonging to the class, or the centroid of the embeddings of all the images belonging to the class. Here, in order to calculate the representative vector of the class, the embedding vectors of all the images belonging to the class are used. However, for example, N may be randomly selected from the embedding vectors of all the images belonging to the class and used. If the number of embedding vectors of all the images is NA, then N ≤ NA.

[0038] Figure 4 is a flowchart for explaining the operation of continuous learning.

[0039] The neural network processing unit 10 to be learned inputs training data consisting of a set of three images, namely an anchor image, a positive image, and a negative image, into the neural network to be learned (S100). A metric loss is calculated using the three images of the anchor image, the positive image, and the negative image. The anchor image is the image to be learned. The positive image and the negative image are not the objects of learning. That is, the regularization loss operation and the calculation of the representative vector are performed using the anchor image. The positive image and the negative image are not used for the regularization loss operation and the calculation of the representative vector. The training data is randomly selected from the training data set.

[0040] The neural network processing unit 10 to be learned supplies the embedding output by the neural network to be learned to the embedding conversion unit 20 and the regularization loss calculation unit 80 (S110).

[0041] The embedding conversion unit 20 smooths the embedding output by the neural network to be learned to calculate an embedding vector, and supplies the calculated embedding vector to the metric loss calculation unit 70 and the regularization loss calculation unit 80 (S120).

[0042] The metric loss calculation unit 70 calculates a metric loss based on the embedding vector supplied from the embedding conversion unit 20 (S130).

[0043] The learned neural network processing unit 40 inputs the training data into the learned neural network (S140).

[0044] The learned neural network processing unit 40 supplies the embedding output by the learned neural network to the embedding conversion unit 50 and the regularization loss calculation unit 80 (S150).

[0045] The embedding conversion unit 50 smooths the embedding output by the learned neural network to calculate an embedding vector, and supplies the calculated embedding vector to the regularization loss calculation unit 80 (S160).

[0046] The regularization loss calculation unit 80 calculates an embedding loss based on the embeddings supplied from the neural network processing unit 10 to be learned and the learned neural network processing unit 40, calculates an embedding vector loss based on the embedding vectors supplied from the embedding conversion unit 20 and the embedding conversion unit 50, and adds the embedding loss and the embedding vector loss to calculate a regularization loss (S170).

[0047] The loss calculation unit 60 adds the metric loss and the regularization loss to calculate an overall loss (S180).

[0048] The optimization unit 90 determines whether to optimize (S190). If not, it returns to step S100. If so, it proceeds to step S200. Optimization is performed on an epoch basis. An epoch is a unit in which all images in the training dataset are used as anchor images. If the number of images is large, optimization may be performed in batch units. If not, the overall loss is added up. If so, learning is performed to minimize the sum of the overall losses, and the overall loss is reset when optimization is complete.

[0049] The optimization unit 90 optimizes the neural network to be learned based on the sum of the overall losses (S200).

[0050] The optimization unit 90 determines whether to complete learning (S210). If learning continues, it returns to step S100. If learning is complete, it proceeds to step S220. Learning is completed when learning for a predetermined number of epochs is finished. The predetermined number of epochs is set before training.

[0051] When the learning is completed, the neural network processing unit 10 to be learned sets the neural network to be learned as a learned neural network in the learned neural network processing unit 40 (S220).

[0052] As described above, according to the machine learning device 100 of the present embodiment, as the regularization loss, by using the embedding vector loss considering the outline characteristics of the image while considering the embedding loss considering the detailed characteristics of the image, continuous learning can be performed considering both the detailed characteristics and the outline characteristics of the image. In addition, by using the detailed characteristics and the outline characteristics of the image together, the outline characteristics are emphasized, and the entanglement of the embeddings in the embedding space can be suppressed.

[0053] Further, according to the machine learning device 100 of the present embodiment, by saving the representative vectors of the previous learning session and using them in the next learning session, additional classes can be continuously learned while retaining the classes learned in the past sessions.

[0054] Hereinafter, some modifications of the present embodiment will be described.

[0055] (Modification 1) When calculating the regularization loss Lr, only the embedding loss Lre is used without using the embedding vector loss Lrv.

[0056] By using the embedding loss considering the detailed characteristics of the image as the regularization loss, continuous learning can be performed considering the detailed characteristics of the image. That is, continuous learning can be performed considering even the individual components of 7×7 of the embedding.

[0057] (Modification 2) When calculating the regularization loss Lr, the embedding loss Lre and the embedding vector loss Lrv are weighted and added. Lr = β * Lre + Lrv Here, β is a weight parameter from 0.0 to 1.0.

[0058] The degree to which the situation where classes are mixed in the embedding space of the neural network can be removed can be controlled by the weight parameter β.

[0059] (Modification Example 3) Increase the initial value of the weight parameter β in Modification Example 2 and decrease β as the number of learning sessions increases. β may be set to 0 in a predetermined learning session.

[0060] In the first stage of learning, detailed characteristics are considered. However, generally, as the learning session progresses to a certain extent, the entanglement of the embedding space increases. Therefore, by increasing the ratio of the outline characteristics, the entanglement of the embedding space can be suppressed.

[0061] (Modification Example 4) At first, the representative vector of the class may be used as the embedding output by the neural network, and may be switched to the smoothed embedding vector at a predetermined learning session.

[0062] In the first half of learning, the representative vector considering detailed characteristics is left using the embedding. However, generally, in the second half of learning, since the entanglement of the embedding space increases, the representative vector considering the outline characteristics can be switched using the embedding vector to suppress the entanglement of the embedding space. Also, by this switching, the memory capacity can be reduced.

[0063] (Modification Example 5) In the above embodiment, the embedding loss considering the detailed characteristics of the image is calculated using the embedding output by the neural network, and the embedding vector loss considering the outline characteristics of the image is calculated using the smoothed embedding vector. However, this may be generalized to calculate two types of losses using the following two embedding matrices. In this modification example, the configurations and operations of the embedding conversion units 20 and 50, the metric loss calculation unit 70, and the regularization loss calculation unit 80 in the embodiment are different as follows.

[0064] The embedding conversion units 20 and 50 apply two different conversions to a 7×7 embedding map and convert it into two embedding matrices.

[0065] Figs. 5(a) to 5(f) are diagrams for explaining examples of conversions applied to a 7×7 embedding map. The values within each black frame in Figs. 5(a) to 5(f) are smoothed (averaged), and the smoothed portions become the columns of the embedding matrix.

[0066] Fig. 5(a) shows global average pooling that smooths all the values within the embedding map. This corresponds to smoothing a 7×7 embedding to obtain an embedding vector as in the above-described embodiment.

[0067] Fig. 5(b) shows horizontal average pooling that smooths the horizontal components within the embedding map. Here, all the horizontal components are smoothed and utilized, but only the even rows or only the odd rows may be utilized. Fig. 5(c) shows vertical average pooling that smooths the vertical components within the embedding map. Here, all the vertical components are smoothed and utilized, but only the even columns or only the odd columns may be utilized.

[0068] Figs. 5(d) and 5(e) show vertical average pooling that smooths a predetermined rectangular region within the embedding map. It is not necessary to utilize all the values within the embedding map as shown in Fig. 5(d), and the values within the embedding map may be utilized repeatedly as shown in Fig. 5(e). In Fig. 5(d), the 3×3 rectangular regions in the upper left, lower left, upper right, and lower right are smoothed, and the central row and column are not utilized. In Fig. 5(e), overlapping is allowed in the central row and column, and the 4×4 rectangular regions in the upper left, lower left, upper right, and lower right are smoothed. Each rectangular region is not limited to these sizes.

[0069] FIG. 5(f) shows concentric circle average pooling for dividing and smoothing the embedded map based on the distance from the center of the embedded map. In FIG. 5(f), smoothing is performed by overlapping rectangular regions of 3×3, 5×5, and 7×7 according to the distance from the center. Here, smoothing is performed by overlapping rectangular regions of 3×3, 5×5, and 7×7, but smoothing may also be performed without overlapping. Also, only the 5×5 rectangular region in FIG. 5(f) may be used. In this case, it corresponds to the case where the region of global average pooling in FIG. 5(a) is reduced.

[0070] The number and shape of the regions to be smoothed are not limited to those illustrated in FIGS. 5(a) to 5(f).

[0071] For example, as a first transformation, the embedding conversion units 20 and 50 convert the embedding output by the neural network by global average pooling in FIG. 5(a) into a first embedding matrix. The first embedding matrix is a 512×1 matrix. As a second transformation, the embedding conversion units 20 and 50 convert the embedding output by the neural network by 4-region equal average pooling for smoothing the inside of the black frame shown in FIG. 5(d) into a second embedding matrix. The second embedding matrix is a 512×4 matrix.

[0072] Here, global average pooling and 4-region equal average pooling are combined, but not limited to this combination. For example, the transformations in FIGS. 5(a) to 5(f) may be freely combined. Also, as a combination target, the embedded map itself can be included as an equivalent transformation. In the first embodiment, as a first transformation, an equivalent transformation is used to calculate the embedding loss Lre with the first embedding matrix as the embedding itself, and as a second transformation, global average pooling in FIG. 5(a) is used to calculate the embedding vector loss Lrv with the second embedding matrix as the embedding vector, and the sum of the embedding loss Lre and the embedding vector loss Lrv is calculated as the regularization loss Lr.

[0073] In calculating the metric loss, the metric loss calculation unit 70 uses, instead of the embedding vector in the embodiment, the embedding matrix with the smaller number of columns among the first embedding matrix or the second embedding matrix obtained by converting the embedding output by the neural network to be learned.

[0074] The regularization loss calculation unit 80 sets the regularization loss Lr to the sum of the first embedding matrix loss Lr1 and the second embedding matrix loss Lr2. Lr = Lr1 + Lr2 Lr1 = ||M1(i) - M1(i-1)|| Lr2 = ||M2(i) - M2(i-1)|| Here, M1(i) is the first embedding matrix obtained by converting the embedding output by the neural network to be learned, and M1(i-1) is the first embedding matrix obtained by converting the embedding output by the learned neural network. M2(i) is the second embedding matrix obtained by converting the embedding output by the neural network to be learned, and M2(i-1) is the second embedding matrix obtained by converting the embedding output by the learned neural network.

[0075] Instead of the embedding vector, the embedding conversion unit 20 obtains the representative vector of the class using the embedding matrix used by the metric loss calculation unit 70 and stores it in the representative vector storage unit 30. However, the metric loss calculation unit 70 may use the embedding vector as in the embodiment. In that case, the representative vector obtained from the embedding vector is stored in the representative vector storage unit 30.

[0076] By making the combination of the regularization losses arbitrary, without being limited to the combination of the detailed characteristics and the outline characteristics, by calculating and optimizing the loss by combining two characteristics with different features, it is possible to suppress the entanglement of the embeddings in the embedding space according to the image characteristics to be classified. Furthermore, the loss may be calculated and optimized by combining two or more different characteristics. For example, for the conversions shown in FIGS. 5(a) to 5(f), three or more conversions may be arbitrarily combined, such as by changing the number of divisions.

[0077] (Second Embodiment) FIG. 6 is a configuration diagram of an image processing apparatus 200 according to the second embodiment. The image processing apparatus 200 includes a general embedding learning unit 210, a detailed embedding learning unit 220, a synthesis unit 230, and a classification unit 240.

[0078] Hereinafter, the configuration and operation of the image processing apparatus 200 according to the second embodiment will be described. Regarding the configuration and operation common to the machine learning apparatus 100 of the first embodiment, the description will be omitted as appropriate, and the configuration and operation different from those of the machine learning apparatus 100 of the first embodiment will be described.

[0079] Similar to the first embodiment, in the second embodiment, machine learning that combines continuous learning and metric learning is also performed.

[0080] FIG. 7 is a flowchart for explaining the overall flow of learning by the image processing apparatus 200. The configuration and operation of the machine learning will be described with reference to FIGS. 6 and 7.

[0081] First, a basic class is learned using a basic training dataset, and the neural network is updated (S300). Let the learning session i be 0.

[0082] Subsequently, based on the neural network, a representative vector of the basic class is derived and saved using the basic training dataset (S310). The method for deriving the representative vector will be described later.

[0083] Up to this point is regarded as learning session 0, which is also called the basic learning session.

[0084] Next, the learning session i of continuous learning is repeated N times (i = 1, 2,..., N) (S320). The learned neural network in the learning session (i - 1) is used as the neural network to be learned in the learning session i.

[0085] In learning session i (i ≥ 1), the following processes are executed in the summary embedding learning unit 210 and the detailed embedding learning unit 220.

[0086] Using the additional training dataset, an embedding or embedding vector is derived from the neural network before continuous learning and saved (S330). The summary embedding learning unit 210 derives and saves the embedding vector, and the detailed embedding learning unit 220 derives and saves the embedding. Here, the neural network before continuous learning is the learned neural network in learning session (i - 1).

[0087] Subsequently, the summary embedding learning unit 210 and the detailed embedding learning unit 220 continuously learn additional classes using the additional training dataset and update the neural network (S340).

[0088] Subsequently, the summary embedding learning unit 210 and the detailed embedding learning unit 220 derive and save the representative vectors of the additional classes using the additional training dataset based on the neural network (S350).

[0089] Increment i by 1 (S360), return to step S320, repeat steps S330 - S350 until i = N, and end if i exceeds N.

[0090] The configuration and operation of continuous learning will be described in more detail. Assuming that learning session 0 is completed, learning session i (i ≥ 1) of continuous learning will be described.

[0091] When a training dataset is input, the summary embedding learning unit 210 learns the summary embedding of the classes included in the training dataset and outputs the representative vector of each class. Also, when an image to be classified is input, the summary embedding learning unit 210 outputs an embedding vector.

[0092] When the training dataset is input, the detailed embedding learning unit 220 learns the detailed embedding of the classes included in the training dataset and outputs the representative vector of each class. Also, when an image to be classified is input, the detailed embedding learning unit 220 outputs an embedding vector.

[0093] The learning rate of the detailed embedding learning unit 220 is set to be larger than the learning rate of the general embedding learning unit 210. For example, if the learning rate of the general embedding learning unit 210 is set to 1e -6 and the learning rate of the detailed embedding learning unit 220 is set to 1e -4 .

[0094] The combining unit 230 operates only when an image to be classified is input. The combining unit 230 concatenates the embedding vector output by the general embedding learning unit 210 and the embedding vector output by the detailed embedding learning unit 220 to derive a combined embedding vector. Also, for each class, the combining unit 230 concatenates the representative vector output by the general embedding learning unit 210 and the representative vector output by the detailed embedding learning unit 220 to derive a combined representative vector.

[0095] The classification unit 240 operates only when an image to be classified is input. The classification unit 240 performs class classification using the combined representative vector and the combined embedding vector. In class classification, the NCM (Nearest Class Mean) method is used. That is, the input image is classified into the class having the combined representative vector with the shortest distance to the combined embedding vector. The distance is the Mahalanobis distance as follows. d M (x,x’)=(x - x’) T M(x - x’) Here, x is the combined embedding vector, x' is the combined representative vector, and M is the scatter covariance matrix. The scatter covariance matrix may be a positive semi-definite matrix.

[0096] FIG. 8A shows the configuration of the general embedding learning unit 210, and FIG. 8B is a diagram showing the configuration of the detailed embedding learning unit 220.

[0097] The memory units 330a and 330b are connected to each part of the general embedding learning unit 210 and the detailed embedding learning unit 220. Before the learning session 1 starts, it is assumed that the representative vectors of the basic classes derived in the learning session 0 are stored in the memory units 330a and 330b.

[0098] Before the learning of the neural networks 310a and 310b starts in each learning session, using the additional training data set, the embeddings and embedding vectors are derived from the neural networks 310a and 310b before continuous learning and stored in the memory units 330a and 330b.

[0099] The neural networks 310a and 310b become the neural networks to be learned in the learning session i. The structure of the neural network model to be learned is the same as that shown in FIG. 3 in Embodiment 1. After the learning session 0 is completed, the learned neural network of the learning session 0 is replicated to the neural network of the learning session 1.

[0100] The embedding conversion units 320a and 320b perform smoothing by global average pooling for each 7×7 embedding map output by the neural networks 310a and 310b to be learned, and output 512-dimensional embedding vectors. The embedding conversion units 320a and 320b perform smoothing by global average pooling for each 7×7 embedding map output by the neural networks 310a and 310b before continuous learning, and output 512-dimensional embedding vectors.

[0101] The loss calculation units 340a and 340b add the metric loss Lml calculated by the metric loss calculation units 350a and 350b and the regularization loss Lr calculated by the regularization loss calculation units 360a and 360b to calculate the overall loss L as follows. L = Σ(Lml + Lr) Here, Σ indicates taking the sum with respect to the input image.

[0102] The metric loss calculation units 350a and 350b calculate the metric loss. Here, the triplet loss is used as the metric loss. The triplet loss Lml is calculated by the following equation based on the embedding vector of the anchor image, the embedding vector of the positive image, and the embedding vector of the negative image. Lml = dp - dn + α Here, dp is the Euclidean distance between the embedding vector of the anchor image and the embedding vector between the positive images. dn is the Euclidean distance between the embedding vector of the anchor image and the embedding vector between the negative images. α is an offset.

[0103] The regularization loss calculation units 360a and 360b calculate the regularization loss. The regularization loss Lr is the embedding loss Lre or the embedding vector loss Lrv. In the regularization loss calculation unit 360a of the overview embedding learning unit 210, the regularization loss Lr is the embedding vector loss Lrv. In the regularization loss calculation unit 360b of the detailed embedding learning unit 220, the regularization loss Lr is the embedding loss Lre.

[0104] As shown in the following equation, the embedding loss Lre is the difference before and after the learning session of the embedding output when the image is input to the neural network, and the embedding vector loss Lrv is the difference before and after the learning session of the embedding vector output when the image is input to the neural network. Lre = ||E(i) - E(i−1)|| Lrv = ||V(i) - V(i−1)|| Here, E(i) is the embedding output by the neural network to be learned, and E(i−1) is the embedding output by the neural network before continuous learning. V(i) is the embedding vector output by the neural network to be learned, and V(i−1) is the embedding vector output by the neural network before continuous learning. E(i−1) and V(i−1) are obtained from the storage units 330a and 330b.

[0105] The optimization units 380a and 380b optimize the weights of the neural network to be learned so as to minimize the overall loss L. When the optimization is completed, the overall loss L is reset.

[0106] The representative vector of the class is obtained as follows.

[0107] Images in the training dataset are sequentially input to the neural networks 310a and 310b to be learned to derive embeddings, and the embedding conversion units 320a and 320b convert the embeddings into embedding vectors. The representative vector derivation units 370a and 370b calculate representative vectors for each class based on the embeddings or embedding vectors and supply them to the synthesis unit 230.

[0108] The representative vector derivation units 370a and 370b calculate the average of the embedding vectors of all images belonging to the class as the representative vector of the class. The representative vector derivation units 370a and 370b associate the class with the calculated representative vector and store it in the storage units 330a and 330b. The representative vector of the class may be the centroid of the embedding vectors of all images belonging to the class, the average of the embeddings of all images belonging to the class, or the centroid of the embeddings of all images belonging to the class.

[0109] FIG. 9A is a flowchart for explaining the operation of continuous learning by the overview embedding learning unit 210.

[0110] Training data consisting of a set of three images, an anchor image, a positive image, and a negative image, is input to the neural network 310a to be learned (S400).

[0111] The neural network 310a to be learned outputs an embedding (S410).

[0112] The embedding conversion unit 320a smoothes the embedding output by the neural network 310a to be learned to calculate an embedding vector, and supplies the calculated embedding vector to the metric loss calculation unit 350a and the regularization loss calculation unit 360a (S420). The embedding conversion unit 320a supplies the embedding vectors of the anchor image, positive image, and negative image to the metric loss calculation unit 350a, and supplies the embedding vector of the anchor image to the regularization loss calculation unit 360a.

[0113] The metric loss calculation unit 350a calculates a metric loss based on the embedding vector supplied from the embedding conversion unit 320a (S430).

[0114] The embedding conversion unit 320a obtains an embedding vector obtained by smoothing the embedding output by the neural network 310a before continuous learning saved in S330, and supplies the obtained embedding vector to the regularization loss calculation unit 360a (S440).

[0115] The regularization loss calculation unit 360a calculates an embedding vector loss based on the embedding vector obtained by smoothing the embedding output by the neural network 310a to be learned supplied from the embedding conversion unit 320a and the embedding vector obtained by smoothing the embedding output by the neural network 310a before continuous learning, and outputs the embedding vector loss as a regularization loss (S450).

[0116] The loss calculation unit 340a calculates an overall loss by adding the metric loss and the regularization loss (S460).

[0117] The optimization unit 380a determines whether to optimize (S470). If not, it returns to step S400. If so, it proceeds to step S480. Optimization is performed on an epoch basis, but when the number of images is large, it may be optimized in batch units. If not optimized, the overall loss is accumulated. If optimized, learning is performed to minimize the sum of the overall losses, and the overall loss is reset after optimization is completed.

[0118] The optimization unit 380a optimizes the neural network to be learned based on the sum of the overall losses (S480).

[0119] The optimization unit 380a determines whether to complete learning (S490). If learning continues, it returns to step S400. If learning is completed, it ends. Learning is completed when learning for a predetermined number of epochs is finished.

[0120] FIG. 9B is a flowchart for explaining the operation of continuous learning by the detailed embedding learning unit 220.

[0121] Training data consisting of a set of three images, an anchor image, a positive image, and a negative image, is input to the neural network 310b to be learned (S500).

[0122] The neural network 310b to be learned outputs an embedding (S510). The embedding conversion unit 320b is supplied with the embeddings of the anchor image, positive image, and negative image. The regularization loss calculation unit 360b is supplied with the embedding of the anchor image.

[0123] The embedding conversion unit 320b smooths the embedding output by the neural network 310b to be learned to calculate an embedding vector, and supplies the calculated embedding vector to the metric loss calculation unit 350b (S520).

[0124] The metric loss calculation unit 350b calculates a metric loss based on the embedding vector supplied from the embedding conversion unit 320b (S530).

[0125] The embedding conversion unit 320b supplies the embedding output by the neural network 310b before continuous learning saved in S330 to the regularization loss calculation unit 360b (S540).

[0126] The regularization loss calculation unit 360b calculates an embedding loss based on the embeddings output from the neural network 310b to be learned and the neural network 310b before continuous learning, and outputs the embedding loss as a regularization loss (S550).

[0127] The loss calculation unit 340b calculates the overall loss by adding the metric loss and the regularization loss (S560).

[0128] The optimization unit 380b determines whether to optimize (S570). If not, it returns to step S500. If so, it proceeds to step S580. Optimization is performed on an epoch basis, but if the number of images is large, it may be optimized in batches. If not, the overall loss is accumulated. If so, learning is performed to minimize the sum of the overall losses, and the overall loss is reset after optimization is completed.

[0129] The optimization unit 380b optimizes the neural network to be learned based on the sum of the overall losses (S580).

[0130] The optimization unit 380b determines whether to complete learning (S590). If learning continues, it returns to step S500. If learning is completed, it ends. Learning is completed when learning for a predetermined number of epochs is finished.

[0131] FIG. 10 is a flowchart for explaining the operation of class classification by the classification unit 240 of the image processing apparatus 200.

[0132] An image is input to the outline embedding learning unit 210 and the detailed embedding learning unit 220 (S600).

[0133] The neural network 310a of the summary embedding learning unit 210 outputs an embedding (S610). The embedding conversion unit 320a of the summary embedding learning unit 210 smooths the embedding output by the neural network 310a and outputs an embedding vector (S620).

[0134] The neural network 310b of the detailed embedding learning unit 220 outputs an embedding (S630). The embedding conversion unit 320b of the detailed embedding learning unit 220 smooths the embedding output by the neural network 310b and outputs an embedding vector (S640).

[0135] The combining unit 230 combines the embedding vector output by the summary embedding learning unit 210 and the embedding vector output by the detailed embedding learning unit 220 to derive a combined embedding vector (S650).

[0136] The combining unit 230 combines the representative vector output by the summary embedding learning unit 210 and the representative vector output by the detailed embedding learning unit 220 to derive a combined representative vector (S660).

[0137] The classification unit 240 predicts a class based on the combined embedding vector and the combined representative vector and classifies the image (S670).

[0138] As described above, according to the image processing apparatus 200 of the present embodiment, by using together the embedding output by the neural network that learns based on the embedding loss considering the detailed characteristics of the image and the embedding output by the neural network that learns based on the embedding vector loss considering the conceptual characteristics of the image, it is possible to perform continuous learning while maintaining the conceptual characteristics of the image while retaining the detailed characteristics of the image.

[0139] In addition, according to the image processing apparatus 200 of the present embodiment, by saving the representative vectors of previous learning sessions and using them in the next learning session, it is possible to continuously learn additional classes while retaining the classes learned in past sessions.

[0140] Hereinafter, several modification examples of the present embodiment will be described.

[0141] (Modification Example 1) A modification example of the synthesis unit 230 will be described.

[0142] The synthesis unit 230 derives a synthesized embedding vector V by weighted averaging of the embedding vector Vo output from the outline embedding learning unit 210 and the embedding vector Vd output from the detailed embedding learning unit 220.

[0143] For each class, the synthesis unit 230 derives a synthesized representative vector U by weighted averaging of the representative vector Uo output from the outline embedding learning unit 210 and the representative vector Ud output from the detailed embedding learning unit 220.

[0144] V = γVo+(1 - γ)Vd U = γUo+(1 - γ)Ud Here, γ is a weight parameter ranging from 0.0 to 1.0.

[0145] The importance of the detailed characteristics and the conceptual characteristics in the embedding space can be controlled by the weight parameter γ.

[0146] (Modification Example 2) Set the initial value of the weight parameter γ in Modification Example 1 to be small, and increase γ as the number of learning sessions increases. γ may be set to 1 in a predetermined learning session.

[0147] In the initial stage of learning, importance is attached to the detailed characteristics for learning. However, generally, as the learning session progresses to a certain extent, the entanglement of the embedding space increases. Therefore, by increasing the ratio of the conceptual characteristics, the entanglement of the embedding space can be suppressed.

[0148] (Modification Example 3) In the above-described embodiment, an embedding loss considering the detailed characteristics of an image is calculated using the embedding output by the neural network, and an embedding vector loss considering the outline characteristics of the image is calculated using the smoothed embedding vector. However, this can be generalized to calculate two types of losses using the following two embedding matrices. In this modification example, the configurations and operations of the embedding conversion units 320a and 320b, the metric loss calculation units 350a and 350b, and the regularization loss calculation units 360a and 360b in the embodiment are different as follows.

[0149] The embedding conversion unit 320a of the outline embedding learning unit 210 and the embedding conversion unit 320b of the detailed embedding learning unit 220 apply two different conversions to the 7×7 embedding map and convert it into two embedding matrices.

[0150] Examples of the conversions applied to the 7×7 embedding map are as described in FIGS. 5(a) to 5(f), similar to Modification Example 5 of the first embodiment.

[0151] For example, as the first conversion in the outline embedding learning unit 210, the embedding conversion unit 320a converts the embedding output by the neural network into the first embedding matrix by the global average pooling in FIG. 5(a). The first embedding matrix is a 512×1 matrix. As the second conversion in the detailed embedding learning unit 220, the embedding conversion unit 320b converts the embedding output by the neural network into the second embedding matrix by the 4-region equal average pooling that smooths the inside of the black frame shown in FIG. 5(d). The second embedding matrix is a 512×4 matrix.

[0152] Here, the global average pooling and the 4-region equal average pooling are combined, but it is not limited to this combination. For example, regarding the conversions shown in FIGS. 5(a) to 5(f), the number of divisions may be freely changed and combined as long as the dimension of the first embedding matrix is less than that of the second embedding matrix. Also, as a combination target, the embedding map itself can be included as an equivalent conversion. In the second embodiment, in the outline embedding learning unit 210, the global average pooling in FIG. 5(a) is used as the first conversion, and the first embedding matrix is used as an embedding vector to calculate the embedding vector loss Lrv. In the detailed embedding learning unit 220, the equivalent conversion is used as the second conversion, and the second embedding matrix is used as the embedding itself to calculate the embedding loss Lre.

[0153] The metric loss calculation unit 350a of the outline embedding learning unit 210 uses the first embedding matrix obtained by converting the embedding output by the neural network to be learned instead of the embedding vector in the embodiment in calculating the metric loss.

[0154] The metric loss calculation unit 350b of the detailed embedding learning unit 220 uses the second embedding matrix obtained by converting the embedding output by the neural network to be learned instead of the embedding vector in the embodiment in calculating the metric loss.

[0155] The regularization loss calculation unit 360a of the outline embedding learning unit 210 sets the regularization loss Lr as the first embedding matrix loss Lr1.

[0156] The regularization loss calculation unit 360b of the detailed embedding learning unit 220 sets the regularization loss Lr as the second embedding matrix loss Lr2.

[0157] Lr1 = ||M1(i) - M1(i−1)|| Lr2 = ||M2(i) - M2(i−1)|| Here, M1(i) is the first embedding matrix obtained by converting the embedding output by the neural network to be learned, and M1(i−1) is the first embedding matrix obtained by converting the embedding output by the learned neural network before continuous learning. M2(i) is the second embedding matrix obtained by converting the embedding output by the neural network to be learned, and M2(i−1) is the second embedding matrix obtained by converting the embedding output by the neural network before continuous learning.

[0158] The representative vector derivation units 370a and 370b obtain the representative vectors of the classes using the embedding matrices used by the metric loss calculation units 350a and 350b instead of the embedding vectors, and store them in the storage units 330a and 330b. However, the metric loss calculation units 350a and 350b may use the embedding vectors as in the embodiment. In that case, the representative vectors obtained from the embedding vectors are stored in the storage units 330a and 330b.

[0159] The synthesis unit 230 synthesizes the embedding vectors after converting the embedding matrices output by the overview embedding learning unit 210 and the detailed embedding learning unit 220 into embedding vectors. The synthesis unit 230 converts the first embedding matrix output by the overview embedding learning unit 210 into a first embedding vector, converts the second embedding matrix output by the detailed embedding learning unit 220 into a second embedding vector, and synthesizes the first embedding vector and the second embedding vector to derive a synthesized embedding vector.

[0160] By making the combination of the regularization losses arbitrary, without being limited to the combination of the detailed characteristics and the overview characteristics, by calculating and optimizing the loss by combining two characteristics having different features, it is possible to suppress the entanglement of the embeddings in the embedding space according to the image characteristics to be classified. Further, the loss may be calculated and optimized by combining two or more different characteristics. For example, for the conversions shown in FIGS. 5(a) to 5(f), three or more conversions may be arbitrarily combined by changing the number of divisions.

[0161] Of course, the various processes of the machine learning device 100 and the image processing device 200 described above can be realized as a device using hardware such as a CPU and a memory, and can also be realized by firmware stored in a ROM (Read Only Memory), a flash memory, or the like, or software such as a computer. It is also possible to record the firmware program and the software program on a computer-readable recording medium and provide them, or to transmit and receive them to and from a server through a wired or wireless network, or to transmit and receive them as data broadcasting of terrestrial or satellite digital broadcasting.

[0162] As described above, the present invention has been described based on the embodiments. It is understood by those skilled in the art that the embodiments are examples, and various modifications are possible for the combinations of their respective components and processing processes, and such modifications are also within the scope of the present invention.

Explanation of Reference Numerals

[0163] 10 Neural network processing unit for learning target, 20 Embedding conversion unit, 30 Representative vector storage unit, 40 Trained neural network processing unit, 50 Embedding conversion unit, 60 Loss calculation unit, 70 Metric loss calculation unit, 80 Regularization loss calculation unit, 90 Optimization unit, 100 Machine learning device, 200 Image processing device, 310a, 310b Neural network, 320a, 320b Embedding conversion unit, 330a, 330b Storage unit, 340a, 340b Loss calculation unit, 350a, 350b Metric loss calculation unit, 360a, 360b Regularization loss calculation unit, 370a, 370b Representative vector derivation unit, 380a, 380b Optimization unit.

Claims

1. a first embedded learning unit that continuously trains a first neural network; a second embedded learning unit that continuously trains a second neural network; a combining unit that combines a first embedded vector derived from the first neural network of the first embedded learning unit and a second embedded vector derived from the second neural network of the second embedded learning unit with respect to input data to output a combined embedded vector; a classification unit that classifies the input data based on the combined embedded vector; and the first and second embedded learning units each include an embedding conversion unit that converts an embedding output from a neural network into an embedding matrix; a regularization loss calculation unit that calculates a regularization loss, which is a difference between before and after learning of the embedding matrix, based on the embedding matrix; a metric loss calculation unit that calculates a metric loss based on the embedding matrix; and an optimization unit that optimizes the neural network based on the metric loss and the regularization loss, wherein the embedding conversion unit of the first embedded learning unit and the embedding conversion unit of the second embedded learning unit convert the embedding into different embedding matrices from each other, the regularization loss calculation unit of the first embedded learning unit and the regularization loss calculation unit of the second embedded learning unit calculate the regularization loss based on different embedding matrices from each other, and the embedding conversion unit of the first embedded learning unit and the embedding conversion unit of the second embedded learning unit generate the embedding matrices having different dimensions from each other by smoothing different regions of the embedding output from the neural network, an image processing apparatus characterized by this.

2. A first embedded learning unit that continuously trains a first neural network; a second embedded learning unit that continuously trains a second neural network; a combining unit that combines a first embedded vector derived from the first neural network of the first embedded learning unit and a second embedded vector derived from the second neural network of the second embedded learning unit with respect to input data to output a combined embedded vector; a classification unit that classifies the input data based on the combined embedded vector; and the first and second embedded learning units each include An embedding conversion unit that converts an embedding output from a neural network into an embedding matrix, A regularization loss calculation unit that calculates a regularization loss, which is the difference before and after learning of the embedding matrix, based on the embedding matrix, A metric loss calculation unit that calculates a metric loss based on the embedding matrix, An optimization unit that optimizes the neural network based on the metric loss and the regularization loss, and the embedding conversion unit of the first embedding learning unit and the embedding conversion unit of the second embedding learning unit convert the embedding into different embedding matrices, the regularization loss calculation unit of the first embedding learning unit and the regularization loss calculation unit of the second embedding learning unit calculate the regularization loss based on different embedding matrices, the embedding conversion unit of the first embedding learning unit and the embedding conversion unit of the second embedding learning unit generate embedding matrices with different dimensions by smoothing different regions of the same-sized embeddings output by the neural network. An image processing apparatus characterized by this.

3. The embedding matrix generated by the embedding conversion unit of the first embedding learning unit is an equivalent conversion of the embedding output by the neural network, and the embedding matrix generated by the embedding conversion unit of the second embedding learning unit is an embedding vector obtained by smoothing the embedding output by the neural network. The image processing apparatus according to claim 1 or 2, characterized by this.

4. A first embedding learning step of continuously learning a first neural network, A second embedding learning step of continuously learning a second neural network, A synthesis step of synthesizing a first embedding vector derived from the first neural network of the first embedding learning unit and a second embedding vector derived from the second neural network of the second embedding learning unit with respect to input data, and outputting a synthesized embedding vector, A classification step of classifying the input data based on the synthesized embedding vector, and the first and second embedding learning steps include an embedding conversion step of converting an embedding output from a neural network into an embedding matrix, A regularization loss calculation step of calculating a regularization loss, which is the difference before and after learning of the embedding matrix, based on the embedding matrix; A metric loss calculation step of calculating a metric loss based on the embedding matrix; An optimization step of optimizing the neural network based on the metric loss and the regularization loss, wherein the embedding conversion steps of the first embedding learning step and the embedding conversion steps of the second embedding learning step convert the embedding into different embedding matrices from each other; wherein the regularization loss calculation steps of the first embedding learning step and the regularization loss calculation steps of the second embedding learning step calculate the regularization loss based on different embedding matrices from each other; An image processing method in which a computer executes each step, characterized in that the embedding conversion steps of the first embedding learning step and the embedding conversion steps of the second embedding learning step generate embedding matrices with different dimensions from each other by smoothing different regions of the embedding output by the neural network.

5. A first embedding learning step of continuously learning a first neural network; A second embedding learning step of continuously learning a second neural network; A synthesis step of synthesizing a first embedding vector derived from the first neural network of the first embedding learning unit and a second embedding vector derived from the second neural network of the second embedding learning unit with respect to input data and outputting a synthesized embedding vector; A classification step of classifying the input data based on the synthesized embedding vector, and causing a computer to execute; The first and second embedding learning steps include: An embedding conversion step of converting an embedding output from a neural network into an embedding matrix; A regularization loss calculation step of calculating a regularization loss, which is the difference before and after learning of the embedding matrix, based on the embedding matrix; A metric loss calculation step of calculating a metric loss based on the embedding matrix; An optimization step of optimizing the neural network based on the metric loss and the regularization loss, and causing a computer to execute. The embedding conversion step of the first embedding learning step and the embedding conversion step of the second embedding learning step convert the embedding into the embedding matrices different from each other, The regularization loss calculation step of the first embedding learning step and the regularization loss calculation step of the second embedding learning step calculate the regularization loss based on the embedding matrices different from each other, The embedding conversion step of the first embedding learning step and the embedding conversion step of the second embedding learning step generate the embedding matrices different in dimension from each other by smoothing different regions of the embedding output by the neural network, and an image processing program characterized by this.

Citation Information

Patent Citations

  • Integral decision method for plural feature quantity

    JP1993101028A

  • Image processing apparatus, and image processing method

    JP2019032773A

  • Information processing apparatus, and information processing method

    JP2021051589A

  • Information processing device and method, and device for performing classification by using model

    JP2021077352A

  • Neural network learning device, neural network learning method and storage medium storing program

    WO2017145852A1