Image processing device, method for producing a learning model, and inference method

The learning model enhances explainability in image retrieval by generating prototype vectors that align with pixel vectors in the feature map, addressing the lack of transparency in existing models and improving similarity searches.

JP7840015B2Active Publication Date: 2026-04-03GLORY LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing machine learning models, particularly those using ProtoPNet, lack explainability in image retrieval tasks, especially when searching for similar images that may include classes other than known classes, limiting their application to classification problems.

Method used

A learning model is developed that generates multiple prototype vectors representing image features, calculates integrated similarity vectors, and optimizes the model to improve explainability by adjusting prototype vectors to match pixel vectors in the feature map, using evaluation functions to enhance transparency and similarity searches.

Benefits of technology

The model improves the explainability of image retrieval processes by learning prototype vectors that represent specific image features, enhancing the ability to understand the reasoning behind similarity searches, thus improving transparency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007840015000021
    Figure 0007840015000021
  • Figure 0007840015000022
    Figure 0007840015000022
  • Figure 0007840015000023
    Figure 0007840015000023
Patent Text Reader

Abstract

To provide a technology that allows improvement in explainability of decision-making basis even during image retrieval processing for retrieving an image similar to an inference target image that belongs to an unclassified class.SOLUTION: A learning model generates a plurality of prototype vectors (a parameter column indicating a candidate for a concept of an image feature) and generates an integrated similarity vector that indicates similarity between an input image and each prototype for a plurality of prototypes in accordance with similarity between one prototype vector and each pixel vector in a feature map acquired from a CNN. An image processing apparatus 30 obtains prototype belongingness (distributed prototype belongingness) for each image by distributing prototype belongingness of a belonging prototype of each class to each of two or more images that belong to one class. Then, the learning model is subjected to machine learning in accordance with the distributed prototype belongingness of each prototype vector for each image so that each prototype vector is brought closer to any pixel vector in the feature map corresponding to each image.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device (particularly an image processing device that improves explainability in machine learning) and related technologies. [Background technology]

[0002] In recent years, inference processing technologies using machine learning, such as deep learning, have been rapidly evolving (see Patent Document 1, etc.).

[0003] However, due to the extreme complexity of machine learning models, there is a problem in that the basis on which the inference results obtained by these models are derived is not always clear (and therefore not easily explained).

[0004] In particular, in situations where the inference results have a significant impact, it is necessary to improve the explainability of the reasoning behind the decision.

[0005] For example, one machine learning technique that applies ProtoPNet to classification problems can improve the explainability of the reasoning behind decisions, particularly "transparency" (whether the inference result can be explained in a concept understandable to humans). In this technique, a "prototype" (a learning parameter representing some image feature) is used within the learning model. More specifically, machine learning is performed so that the "prototype" approaches the image feature (pixel vector in the feature map) at each planar position within the feature map. This feature map is output from the convolutional neural network within the learning model when a training input image is input to the learning model. The image features at each planar position within the feature map represent the image features of a subregion (image patch) within the input image.

[0006] By performing inference using a learning model obtained through such machine learning, it becomes possible to demonstrate, as the basis for judgment, that the image being inferred is similar to a specific subregion (image patch) at a specific location in a particular image. In other words, transparency can be increased. [Prior art documents] [Patent Documents]

[0007] [Patent Document 1] Japanese Patent Publication No. 2018-200531 [Overview of the project] [Problems that the invention aims to solve]

[0008] However, the aforementioned technique applies ProtoPNet to a classification problem. In other words, this technique can only be used to solve a classification problem that determines which of several (known) classes an image to be inferred belongs to.

[0009] In particular, the technology in question (the conventional technology using ProtoPNet) targets a classification problem, and it is assumed that each prototype is uniquely associated with only one specific class. Machine learning is then performed to optimize the loss function (evaluation function) based on this assumption.

[0010] Therefore, this technology cannot be used for purposes other than classification in its current form. For example, it cannot be used as is in image search processing, such as searching for similar images among multiple images that are similar to images of a class other than a known class (images of an unclassified class) that are used as the image to be inferred.

[0011] Therefore, the object of this invention is to provide a technology that can improve the explainability of the basis for judgment even in image retrieval processing that searches for similar images from among multiple images with respect to an image to be inferred that may include images other than those of known classes (images of unclassified classes). [Means for solving the problem]

[0012] To solve the above problems, the image processing apparatus according to the present invention comprises a control unit that performs machine learning on a learning model configured with a convolutional neural network, wherein the learning model generates a feature map obtained from a predetermined layer in the convolutional neural network in response to an input image, the feature map showing the feature quantities for each subregion in the input image for multiple channels. The process to do, Multiple prototype vectors are generated, which are parameter sequences learned as prototypes representing candidate image feature concepts composed of the aforementioned multiple channels. The process to do, Based on the similarity between each pixel vector, which is a vector representing the image features across multiple channels at each planar position of each pixel in the feature map, and a single prototype vector, an integrated similarity vector is generated that shows the similarity between the input image and each prototype for multiple prototypes. The process of doing so, and machine learning to get the computer to do it. The control unit, in the learning stage of the learning model based on multiple images for learning, determines a belonging prototype, which is a prototype belonging to one class, and a prototype belonging degree, which indicates the degree to which the belonging prototype belongs to one class, for each of the multiple classes labeled on the multiple images for learning. It also determines a distributed prototype belonging degree, which is the prototype belonging degree for each image, by distributing the prototype belonging degree of the belonging prototype in each class to each of two or more images within the same class based on a predetermined criterion. When performing a learning process based on multiple integrated similarity vectors corresponding to the multiple images, it ensures that each prototype vector approaches one of the pixel vectors in the feature map corresponding to each image, according to the distributed prototype belonging degree of each image. to, The aforementioned learning model is subjected to machine learning.

[0013] The control unit performs a prototype selection process to select a prototype belonging to the first class from among the plurality of learning images based on a comparison between each of the plurality of comparison images belonging to classes other than the first class. Based on the number of prototypes selected for each belonging prototype in the prototype selection process, the control unit determines the degree of prototype belonging to each belonging prototype for the first class. The prototype selection process includes a unit selection process to obtain a difference vector by subtracting the integrated similarity vector obtained by inputting one of the plurality of comparison images into the learning model from the integrated similarity vector obtained by inputting the predetermined image into the learning model, and selecting the prototype corresponding to the largest component among the plurality of components in the difference vector as the belonging prototype belonging to the class of the predetermined image. The control unit may also include a selection count calculation process to select at least one belonging prototype belonging to the first class and count the number of selected prototypes for each belonging prototype by performing the unit selection process for the plurality of comparison images while changing the one comparison image to another comparison image.

[0014] The control unit may reduce the degree to which a single prototype belongs in each of the two or more classes if the prototype belongs to two or more classes.

[0015] The control unit, in determining the assigned prototype affiliation degree for each of the N images belonging to the first class by distributing the prototype affiliation degree of one affiliated prototype belonging to the first class to N images belonging to the first class, determines a first distance, which is the distance from a plurality of pixel vectors in the feature map corresponding to one of the N images to the pixel vector that is most similar to the prototype vector of the first affiliated prototype, and determines a second distance, which is the distance from a plurality of pixel vectors in the feature map corresponding to another of the N images to the pixel vector that is most similar to the prototype vector of the first affiliated prototype. If the first distance is greater than the second distance, the assigned prototype affiliation degree for the first image may be determined to be a smaller value than the assigned prototype affiliation degree for the other images.

[0016] The control unit may, after the machine learning of the learning model is completed, modify the learning model by replacing each prototype vector with the most similar pixel vector, which is the pixel vector that is most similar to each prototype vector among the multiple pixel vectors in the multiple feature maps relating to the multiple images.

[0017] The evaluation function used for machine learning of the learning model has a first evaluation term which is an evaluation term related to clarity. The control unit may optimize the first evaluation term so as to maximize the value obtained by inputting the first image to the learning model and the second image to the learning model, with respect to a first image and a second image relating to any combination of the plurality of images for learning. The control unit sorts the absolute values ​​of the plurality of components in the difference vector between the first and second vectors in descending order. The size Dn of the partial difference vector reconstructed using only the top n components after sorting the plurality of components of the difference vector in descending order is determined for each of a plurality of values ​​n (n=1,...,,Nd; where the value Nd is a predetermined integer less than or equal to the number of dimensions Nc of the integrated similarity vector). The control unit may then machine learn the learning model by optimizing the first evaluation term so as to maximize the value obtained by dividing the sum of the plurality of sizes Dn corresponding to each of the plurality of values ​​n by the inter-vector distance between the two vectors.

[0018] The evaluation function further comprises a second evaluation term, which is an evaluation term for distance learning based on the plurality of integrated similarity vectors corresponding to the plurality of images, wherein the first evaluation term is expressed as a sum obtained by adding up pair-specific first evaluation terms, which are obtained for each pair of images, over multiple pairs of images, and the second evaluation term is expressed as a sum obtained by adding up pair-specific second evaluation terms, which are obtained for each pair of images, which are obtained for distance learning based on the plurality of integrated similarity vectors, over multiple pairs of images, and the control unit may adjust the magnitude of the pair-specific first evaluation term such that the absolute value of the partial derivative of the pair-specific first evaluation term with respect to the inter-vector distance for each pair of images does not exceed the absolute value of the partial derivative of the pair-specific second evaluation term with respect to the inter-vector distance for each pair of images.

[0019] After the machine learning of the learning model is completed, the control unit may search for an image similar to the input image to be searched from the plurality of learning images based on the integrated similarity vector output from the learning model in response to inputting the input image to be searched into the learning model and the plurality of integrated similarity vectors output from the learning model in response to inputting the plurality of learning images into the learning model.

[0020] After the machine learning of the learning model is completed and each prototype vector is replaced with the most similar pixel vector, the control unit may search for an image similar to the input image to be searched from the plurality of learning images based on the integrated similarity vector output from the learning model in response to inputting the input image to be searched into the learning model and the plurality of integrated similarity vectors output from the learning model in response to inputting the plurality of learning images into the learning model.

[0021] In order to solve the above problems, a method for producing a learning model according to the present invention machine-learns and produces the following learning model. The learning model generates a feature map obtained from a predetermined layer in a convolutional neural network in the learning model in response to input of an input image, the feature map indicating feature amounts for each partial region in the input image for a plurality of channels. The process to do, Generate a plurality of prototype vectors, which are parameter sequences learned as prototypes indicating candidates for specific image feature concepts composed of the plurality of channels. The process to do, Based on the similarity between each pixel vector, which is a vector representing the image features across the plurality of channels at each planar position of each pixel in the feature map, and one prototype vector, generate an integrated similarity vector indicating the similarity between the input image and each prototype for the plurality of prototypes. The process of doing so, and machine learning to get the computer to do it. It is a model. The method for producing the learning model is as follows: a) ComputersThe step of determining, for each of the multiple classes labeled with the multiple training images, a "belonging prototype," which is a prototype belonging to a particular class, and a "prototype belonging degree," which indicates the degree to which the belonging prototype belongs to that particular class, based on multiple training images. a) b) Computers The step of determining the distributed prototype affiliation degree, which is the prototype affiliation degree for each image, by distributing the prototype affiliation degree of each class's belonging prototypes to each of two or more images within the same class, based on predetermined criteria. b) c) Computers When performing a learning process based on multiple integrated similarity vectors corresponding to the multiple images, each prototype vector is made to approach one of the pixel vectors in the feature map corresponding to each image, according to the degree of belonging of each image's distributed prototype. to, Steps to machine learn the aforementioned learning model c) It is equipped with the following.

[0022] The method for producing the aforementioned learning model is d) Computers After the machine learning of the aforementioned learning model is completed, the learning model is modified by replacing each prototype vector with the most similar pixel vector, which is the pixel vector that is most similar to each prototype vector among the multiple pixel vectors in the multiple feature maps relating to the multiple images. d) It may also be provided as an additional feature.

[0023] Step c) above is performed with respect to the first image and the second image relating to any combination of the plurality of images for training, c-1) Computers The steps include: obtaining a first vector, which is an integrated similarity vector obtained by inputting the first image into the learning model, and a second vector, which is an integrated similarity vector obtained by inputting the second image into the learning model; and c-2) Computers (c-3) The steps of sorting the absolute values ​​of multiple components in the difference vector between the first vector and the second vector in descending order. ComputersThe steps include determining the magnitude Dn of the partial difference vector, which is reconstructed using only the top n components after sorting the multiple components of the difference vector in descending order, for each of the multiple values ​​n (n=1,...,,Nd; where the value Nd is a predetermined integer less than or equal to the number of dimensions Nc of the unified similarity vector), and c-4) Computers The sum of the multiple magnitudes Dn corresponding to each of the multiple values ​​n, divided by the distance between the two vectors, is normalized to maximize the value obtained by this process. to, The system may also include the step of machine learning the aforementioned learning model.

[0024] To solve the above problems, the inference method according to the present invention uses a learning model produced by any of the above learning model production methods, Computer Perform inference processing on the new input image. [Effects of the Invention]

[0025] According to the present invention, each prototype vector is learned to approach one of the pixel vectors in the feature map corresponding to each image, according to the degree of belonging of the distributed prototype to each image. Therefore, each prototype is learned to represent features that are close to the image features (pixel vectors) of a specific region of a specific image. Consequently, it is possible to improve the explainability (especially transparency (the ability to explain with concepts that can be understood by humans)) of the learning results in the learning model. [Brief explanation of the drawing]

[0026] [Figure 1] This is a schematic diagram showing an image processing system. [Figure 2] This diagram shows the hierarchical structure of the learning model. [Figure 3] This diagram shows the data structure and other aspects of the learning model. [Figure 4] This is a conceptual diagram showing an example of the configuration of a feature extraction layer. [Figure 5] This is a flowchart showing the processing steps of an image processing device (controller, etc.). [Figure 6]This is a conceptual diagram illustrating the overview of the learning process. [Figure 7] This figure shows the feature space, etc., before learning progress. [Figure 8] This figure shows the feature space and other aspects after learning has progressed. [Figure 9] This flowchart shows the details of the learning process. [Figure 10] This flowchart shows the details of the learning process. [Figure 11] This is a flowchart that shows some of the processes in Figure 9 in detail. [Figure 12] This diagram conceptually illustrates the learning process related to the evaluation term Lclst. [Figure 13] This is a conceptual diagram illustrating the prototype selection process. [Figure 14] This figure shows an example of the process of averaging prototype affiliation by class. [Figure 15] This is a conceptual diagram illustrating debiasing (bias suppression) processing. [Figure 16] This figure shows an example of debiasing. [Figure 17] This figure shows another example of debiasing. [Figure 18] This is a conceptual diagram illustrating the distribution process. [Figure 19] This figure shows an example of the distribution process. [Figure 20] This is a conceptual diagram illustrating the process of replacing prototype vectors. [Figure 21] This is a diagram illustrating the inference process. [Figure 22] This figure shows an example of the inference processing result. [Figure 23] This figure shows an example of how explanatory information is displayed. [Figure 24] This figure shows another example of how explanatory information can be displayed. [Figure 25] This figure shows part of the process by which the criteria for determining that two images are not similar are generated. [Figure 26] This figure shows the rearrangement of difference vectors. [Figure 27]This figure shows an example of how explanatory information is displayed regarding inference results that are not similar. [Figure 28] This figure shows an example of how explanatory information is displayed regarding inference results that are not similar. [Figure 29] This figure shows the level of explanation provided by a predetermined number of prototypes (before improvement). [Figure 30] This is a conceptual diagram illustrating how clarity is improved. [Figure 31] This is a diagram illustrating Lint, an evaluation item used to improve clarity. [Figure 32] This figure shows the learning process related to the evaluation term Lint. [Figure 33] Figure 29 shows an example of improved clarity for the image pair. [Figure 34] This figure shows the repulsive and attractive forces acting on a negative pair. [Modes for carrying out the invention]

[0027] Embodiments of the present invention will be described below with reference to the drawings.

[0028] <1. First Embodiment> <1-1. System Overview, etc.> Figure 1 is a schematic diagram of the image processing system 1. As shown in Figure 1, the image processing system 1 comprises multiple (many) imaging devices 20 for capturing images and an image processing device 30 for processing the captured images. The image processing device 30 is a device that performs various processes for identifying or classifying the subject (in this case, the subject person) in the captured image.

[0029] Images captured by each imaging device 20 are input to the image processing device 30 via a communication network (LAN and / or the Internet, etc.). The image processing device 30 then performs image retrieval processing, etc., to search for similar images to a given inference target image (a certain captured image, etc.) from among multiple images (known images (training captured images, etc.)).

[0030] More specifically, as shown in the flowchart of Figure 5, first, the image processing device 30 trains (machine learns) a learning model 400, which will be described later, based on multiple training images taken of multiple objects (multiple types of objects, such as birds). Through such machine learning, a trained learning model 400 (also referred to as 420) is generated (step S11). Figure 5 is a flowchart showing the processing of the image processing device 30 (controller 31, etc.).

[0031] Subsequently, the image processing device 30 performs inference processing using the trained training model 420 (step S12). Specifically, the image processing device 30 uses the trained training model 420 to perform image search processing, such as searching (extracting) an image from a plurality of training images that is most similar to a certain image to be inferred (an image containing an object most similar to an object in a certain image to be inferred). Such processing is also referred to as the process of identifying an object (animal, person, etc.) in a certain image.

[0032] Furthermore, the image processing device 30 performs an explanatory information generation process (step S13) for the inference result.

[0033] Here, while we primarily use captured images as examples for the images to be inferred and the multiple images for training, we are not limited to these. For example, the multiple images for training and the images to be inferred may be images other than captured images (e.g., computer graphics images or handwritten images). Furthermore, the captured images may be images captured by the image capture device 20 of the image processing system 1, or images captured by an image capture device other than the image capture device 20 of the image processing system 1.

[0034] <1-2. Image processing device 30> Refer to Figure 1 again. As shown in Figure 1, the image processing device 30 comprises a controller 31 (also called a control unit), a storage unit 32, a communication unit 34, and an operation unit 35.

[0035] The controller 31 is a control device built into the image processing device 30 that controls the operation of the image processing device 30.

[0036] The controller 31 is configured as a computer system equipped with one or more hardware processors (for example, a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit)). The controller 31 performs various processes by executing a predetermined software program (hereinafter also simply referred to as a program) stored in a storage unit (ROM and / or a non-volatile storage unit such as a hard disk) 32 using the CPU, etc. The program (more specifically, a group of program modules) may be recorded on a portable recording medium such as a USB memory stick and read from the recording medium to be installed in the image processing device 30. Alternatively, the program may be downloaded via a communication network or the like and installed in the image processing device 30.

[0037] Specifically, the controller 31 performs a learning process to train the learning model 400, and an inference process (such as image search processing) using the trained learning model 400 (420). The controller 31 also performs explanatory processing to show the basis for the inference process.

[0038] The storage unit 32 consists of a storage device such as a hard disk drive (HDD) and / or a solid-state drive (SSD). The storage unit 32 stores the learning model 400 (including learning parameters and programs related to the learning model) (and consequently the trained model 420), etc.

[0039] The communication unit 34 is capable of performing network communication via a network. Various protocols such as TCP / IP (Transmission Control Protocol / Internet Protocol) are used in this network communication. By using this network communication, the image processing device 30 can exchange various types of data (such as captured image data and ground truth data) with a desired partner (for example, the imaging device 20 or an information storage device not shown).

[0040] The operation unit 35 includes an operation input unit 35a that receives operation input to the image processing device 30, and a display unit 35b that outputs various information. A mouse and keyboard can be used as the operation input unit 35a, and a display (such as a liquid crystal display) can be used as the display unit 35b. A touch panel that functions as both part of the operation input unit 35a and part of the display unit 35b may also be provided.

[0041] Furthermore, the image processing device 30 is also called a learning model generator because it has the function of generating a learning model 400 by machine learning using training data (image data of multiple images for training). In addition, the image processing device 30 is also a device that performs inference regarding object identification and / or classification using the learned learning model 400, so it is also called an inference device.

[0042] Furthermore, although various processes (functions) are implemented by a single image processing device 30 here, this is not limited to that. For example, various processes may be implemented by multiple devices. For instance, the training process of the learning model 400 and the inference process using the trained model 400 (420) may be performed by separate devices.

[0043] <1-3. Learning Model 400> As described above, the image processing device 30 is equipped with a learning model 400. Here, the learning model 400 is a neural network model consisting of multiple layers, specifically a convolutional neural network (CNN) model. The learning model 400 is then trained by deep metric learning. Specifically, the parameters (learning parameters) of various image filters (image filters of the convolutional layer) for feature extraction in multiple layers (especially multiple hidden layers) of the convolutional neural network model are adjusted.

[0044] As mentioned above, the trained model 400 after being trained by machine learning is also called a pre-trained model. The pre-trained trained model 400 (pre-trained model 420) is generated by adjusting the training parameters of the trained model 400 (learner) using a predetermined machine learning method.

[0045] In this application, generating a pre-trained model 400 (420) means manufacturing (producing) a pre-trained model 400, and "method for generating a pre-trained model" means "method for producing a pre-trained model."

[0046] Figures 2 and 3 illustrate the configuration of the learning model 400. Figure 2 shows the hierarchical structure of the learning model 400, and Figure 3 shows the data structure and other elements within the learning model 400.

[0047] As shown in Figure 2, the learning model 400 has a hierarchical structure in which multiple layers (hierarchies) are connected hierarchically. Specifically, the learning model 400 comprises an input layer 310, a feature extraction layer 320, a similarity map generation layer 330, an integrated similarity vector generation layer 370, and an output layer 380.

[0048] <Input layer 310> The input layer 310 is the layer that receives the input image 210. The input image 210 is a photograph of an object (for example, an image of a bird). For example, a color image (3 channels) having a pixel array of width (horizontal) W0 pixels and height (vertical) H0 pixels (a rectangular pixel array) is input as the input image 210. In other words, the input image 210 is generated as W0 × H0 × C0 voxel data (where C0 = 3).

[0049] <Feature extraction layer 320> The learning model 400 includes a feature extraction layer 320 following the input layer 310. The feature extraction layer 320 is composed of a convolutional neural network (CNN) 220 (Figure 3). A feature map 230 is generated by processing the input image 210 with the feature extraction layer 320.

[0050] The feature extraction layer 320 includes multiple convolutional layers and multiple pooling layers (such as average pooling and / or maximum pooling). In this convolutional neural network, multiple hidden layers are provided. For example, a part (feature extraction portion) of various convolutional neural network configurations (such as VGG or ResNet) can be used as the feature extraction layer 320.

[0051] For example, in VGG16, the feature extraction layers (13 convolutional layers and 5 pooling layers) (see Figure 4) that are provided up to the final pooling layer following the final convolutional layer CV13 (the pooling layer immediately preceding the fully connected layer (3 layers)) are provided as feature extraction layer 320. In other words, the 18 layers starting from the input layer 310 are provided as feature extraction layer 320 in the convolutional neural network. In Figure 4, a part of the configuration of VGG16 (which has 13 convolutional layers, 5 pooling layers, and 3 fully connected layers) (the feature extraction portion up to the final pooling layer) is shown as an example of feature extraction layer 320. Note that in Figure 4, activation functions and other elements are omitted as appropriate.

[0052] Alternatively, all (or part) of the feature extraction layers provided in other convolutional neural networks such as ResNet (Residual Network) may be provided as the feature extraction layer 320 in the convolutional neural network. ResNet is a convolutional neural network that includes summing residuals between layers. The feature extraction layer in ResNet consists of multiple residual blocks composed of combinations of convolutional layers, activation functions, and skip connections (shortcut connections). In a typical convolutional neural network, a fully connected layer is provided after the feature extraction layer as a layer that performs classification processing based on the features extracted in the feature extraction layer (also called a classification layer). All (or part) of the feature extraction layers provided immediately before such a fully connected layer may be provided as the feature extraction layer 320 in the convolutional neural network.

[0053] Feature map 230 is a feature map output from a predetermined layer (in this case, the final pooling layer) in the convolutional neural network of the learning model 400. Feature map 230 is generated as a feature map having multiple channels. Feature map 230 is generated as a 3D array data (W1×H1×C1 voxel data) having C1 channels, each consisting of a 2D array data of pixel arrays (rectangular pixel arrays) with width W1 pixels and height H1 pixels. The size (W1×H1) of each channel in feature map 230 is, for example, 14×14. The number of channels C1 in feature map 230 is, for example, 512. However, it is not limited to this, and the size of each channel and the number of channels may be other values. For example, the number of channels C1 may be 256 or 1024.

[0054] In this configuration, the feature extraction layer 320 is composed of a repeating arrangement of one or more convolutional layers and one pooling layer. In each convolutional layer, features in the image are extracted by a filter that performs convolution. In addition, each pooling layer performs a pooling process (average pooling or maximum pooling, etc.) to extract the average pixel value or maximum pixel value for each small pixel range (for example, a 2x2 pixel range), thereby reducing the pixel size (for example, by half in both the vertical and horizontal directions) (the amount of information is condensed).

[0055] Then, by applying this processing (convolution and pooling) by the feature extraction layer 320 to the input image 210, a feature map 230 is generated. In this way, the feature map 230 is generated by an intermediate layer in the convolutional neural network that includes multiple convolutional layers and multiple pooling layers, which are placed after the input layer 310. According to this, various features of the image in the input image 210 are extracted for each channel in the feature map 230. In addition, the features of the image in the input image 210 are extracted while retaining their approximate position within the 2D image of each channel in the feature map 230.

[0056] Thus, the feature extraction layer 320 is a layer that generates a feature map 230 obtained from a predetermined layer within the convolutional neural network (CNN) 200 in response to the input image 210. The feature map 230 is voxel data that shows the feature quantities of each subregion in the input image 210 for multiple (C1) channels CH.

[0057] <Similarity Map Generation Layer 330> The similarity map generation layer 330 is a processing layer that generates a similarity map 270 based on the feature map 230 and multiple prototype vectors 250 (see Figure 3). Each prototype vector 250 is also expressed as prototype vector p (or pk).

[0058] Each prototype vector p (the kth prototype vector pk) (see Figure 3) is a sequence of parameters that is learned as a prototype PT (the kth prototype PTk) representing a candidate for a specific image feature concept composed of multiple channels CH. Each prototype vector p is a vector composed of multiple parameters to be learned, and has the same number of dimensions as the number of channels (dimensions in the depth direction) C1 of the feature map 230. Roughly speaking, each prototype vector p is a vector that is learned to represent a specific image feature of a specific image, and is learned to approach any pixel vector q (described below) in the feature map 230 of any image.

[0059] In the learning model 400, multiple prototype vectors p are generated (Nc (for example, 512)). In other words, multiple (Nc) prototype vectors pk (k=1,...,Nc) are generated.

[0060] On the other hand, each pixel vector q (qwh) in the feature map 230 is a vector that represents image features across multiple channels CH at each planar position (w,h) of each pixel in the feature map 230. In Figure 3, the pixel vector q at a certain position (w,h) in the feature map 230 (more specifically, the columnar space extending in the depth direction corresponding to the pixel vector q) is shown with hatching. The number of dimensions (number of dimensions in the depth direction) of each pixel vector q is the same as the number of channels (CH1) in the feature map 230. Each pixel vector q in the feature map 230 represents a specific image feature of a specific region in the original image of the feature map 230. In other words, each pixel vector q is a vector that shows the features of a subregion in a specific image (a subregion image feature representation vector).

[0061] The similarity map generation layer 330 generates a planar similarity map 260 (planar map (2D map)) that shows the similarity Sim(qwh,pk) between each such pixel vector qwh and a prototype vector pk for each planar position. The planar similarity map 260 corresponding to the kth prototype vector pk is also called the kth planar similarity map. Furthermore, the similarity map generation layer 330 generates a similarity map (3D map) 270 constructed with respect to multiple prototype PTk (multiple prototype vectors pk) of the planar similarity map 260. Here, the similarity Sim(q,pk) is a function that calculates the similarity between the prototype vector pk and each of the multiple pixel vectors q (specifically qwh) in the feature map 230. This function is, for example, cosine similarity. However, it is not limited to this, and other functions (various distance functions, etc.) may be used as the function to calculate the similarity Sim.

[0062] As shown in Figure 3, the similarity between a certain prototype vector (the kth prototype vector pk) and each pixel vector 240 (qwh) is placed at each position (w,h) within a single planar similarity map 260 (planar map (2D map)). A similarity map 270 (3D map) is generated by creating multiple (Nc) such planar similarity maps 260 for each prototype vector pk (k=1,...,Nc). In other words, the similarity map 270 is a 3D map in which multiple planar maps 260 are stacked in the depth direction. It should be noted that the similarity map 270 can also be described as a kind of (broadly defined) "feature map" (although it is a different map from the feature map 230).

[0063] <Integrated Similarity Vector Generation Layer 370> The integrated similarity vector generation layer 370 is a processing layer that generates an integrated similarity vector 280 based on the similarity map 270.

[0064] The integrated similarity vector 280 is an Nc-dimensional vector. The integrated similarity vector 280 is also denoted as the integrated similarity vector s. The k-th component Sk of the integrated similarity vector s is calculated by applying GMP processing to the planar similarity map 260 corresponding to the k-th prototype vector. That is, the k-th component Sk is the maximum value among several values ​​in the planar similarity map 260 corresponding to the k-th prototype vector. This k-th component Sk of the integrated similarity vector 280 represents the similarity between the feature map 230 (more specifically, a certain pixel vector q in the feature map 230) and the k-th prototype vector, and is expressed by equation (1). More specifically, the k-th component Sk is the maximum value among the similarities between the k-th prototype vector pk and any pixel vector q in the feature map 230.

[0065]

number

[0066] Global Max Pooling (GMP) is a type of Max Pooling process.

[0067] Max pooling is a process that extracts the largest value (maximum pixel value) from among multiple pixels corresponding to the kernel (filter) size as the feature value (output value). In max pooling, the maximum value is often extracted from multiple pixels (for example, 4 pixels) corresponding to a filter size smaller than the channel size (for example, 2x2 size).

[0068] Global Max Pooling (GMP) processing is a maximum pooling process that applies to the "entire channel" (in this case, the entire plane similarity map 260). GMP processing (global maximum pooling) is a maximum pooling process that extracts the maximum value from among multiple pixels (all pixels in the channel) (for example, 196 pixels) corresponding to the channel size (the size of the plane similarity map 260) and the same filter size (for example, W1 × H1 = 14 × 14).

[0069] By applying this GMP (Global Max Pooling) process to each of the multiple planar similarity maps 260, the maximum pixel value for each channel (for each prototype) of the feature map being processed (in this case, the similarity map 270) is extracted (for each channel). When GMP processing is applied to a similarity map 270 having Nc channels (for example, 512 prototypes), Nc (for example, 512 maximum values) for each channel (for each prototype) are output. In other words, the unified similarity vector 280 is generated as a vector with Nc dimensions (for example, 512 dimensions). This unified similarity vector 280 is a vector that aggregates the similarity Sk between the input image and each prototype (integrated for multiple prototypes PT). The unified similarity vector 280 is a vector that shows the similarity (in other words, the features of the image) between the input image and each prototype, and can also be described as a kind of "feature (quantity) vector".

[0070] In this way, a unified similarity vector 280 is generated based on the similarity between each pixel vector, which is a vector representing image features across multiple channels at each planar position of each pixel in the feature map 230, and one prototype vector. The unified similarity vector 280 is a vector that shows the similarity between the input image 210 and each prototype for multiple prototypes.

[0071] Furthermore, each component Sk of the integrated similarity vector 280 for a given input image 210 can also be expressed as an index value representing the similarity (or distance) between the most similar pixel vector q (also denoted as qnk) within the input image 210 and the kth prototype vector pk (see also Figure 25). The most similar pixel vector qnk is the pixel vector q that is most similar to the kth prototype vector pk among the multiple pixel vectors q in the feature map 230 output from the CNN 220 (feature extraction layer 320) for the input image 210. Each component Sk indicates "the degree to which the image feature represented by the kth prototype vector pk exists in the input image 210." In other words, the similarity Sk can also be expressed as the degree of existence of the prototype PTk (or its concept) in the input image. Moreover, such an integrated similarity vector 280 may be generated directly based on multiple prototype vectors p and each pixel vector q in the feature map 230, without generating a similarity map 270.

[0072] <Output layer 380> The output layer 380 is a processing layer that outputs the unified similarity vector 280 as is. In other words, the mapping (unified similarity vector 280) applied by the learning model 400 to the input image 210 is output from the output layer 380.

[0073] <1-4. Training process for learning model 400> In step S11 (Figure 5), the learning process (machine learning process) of the learning model 400 is executed. This learning process (the learning stage of the learning model 400) is performed based on multiple images used for training.

[0074] First, the image processing device 30 generates multiple images for training by performing size adjustment processing (resizing processing) on ​​each of the multiple captured images acquired from the camera 20, etc., and prepares these multiple images as an input image group for the learning model 400. It is assumed that each of the multiple images for training is pre-labeled (ground truth data) with its class (for example, "bird species"). For example, if the subjects of the multiple images include multiple types of birds, the type of bird that is the subject of each image ("pelican," "green jay," etc.) is pre-assigned as the class of each image. These pre-labeled multiple images (data) are used as training data (training data with ground truth labels).

[0075] In this embodiment, the image processing device 30 basically performs metric learning (also called distance learning) as a machine learning process. More specifically, deep metric learning using a deep neural network (particularly a convolutional neural network) is used. In this metric learning, a learning model 400 is used that outputs a feature vector in the feature space (feature quantity space) for the input image. Such a learning model 400 can also be described as a model that shows the transformation (mapping) from the input image (input) to the feature vector (output).

[0076] Multiple images for training (a group of input images) are sequentially input to the training model 400, and multiple outputs from the training model 400, i.e., multiple feature vectors (a group of feature vectors) in the feature space, are sequentially output. Ideally, in the feature space, multiple feature vectors corresponding to multiple input images of the same class (for example, the same type of bird) are placed close to each other, and multiple feature vectors corresponding to multiple input images of different classes (different types of birds) are placed far apart from each other. However, the distribution of the group of feature vectors based on the output from the training model 400 before training (see Figure 7) deviates from this ideal distribution state (see Figure 8). In Figures 7 and 8, each dot shape (small square or small circle, etc.) placed within the rectangle representing the feature space on the far right represents each feature vector placed within that feature space. Feature vectors of the same class (multiple feature vectors corresponding to multiple images belonging to the same class) are shown with the same shape (white circles, etc.). Conversely, feature vectors from different classes (multiple feature vectors corresponding to multiple images belonging to different classes) are represented by different shapes (e.g., shapes with different hatching patterns).

[0077] Next, in metric learning, the learning model 400 is trained to optimize (minimize) evaluation functions such as Triplet Loss. This trains the learning model 400 (mapping relationship) so that the similarity of input images in the input space corresponds to the distance (distance between feature vectors) in the feature space. In other words, the distribution of feature vectors in the feature space gradually changes as learning progresses. If very good machine learning is performed, the distribution of feature vectors in the feature space will gradually approach the ideal distribution state described above (see Figure 8).

[0078] Figure 6 is a conceptual diagram showing an overview of the learning process in this embodiment. As shown in Figure 6, in this embodiment, two types of metric learning (distance learning) are performed (see the upper part of Figure 6). One is metric learning that treats the integrated similarity vector 280 as a feature vector. The other is metric learning that treats the sub-feature vector 290 as a feature vector. In this embodiment, metric learning with respect to the integrated similarity vector 280 is performed as the primary metric learning, and metric learning with respect to the sub-feature vector 290 is performed as a secondary (supplementary) metric learning. In addition, in this embodiment, a learning process is also performed to bring each prototype vector pk(250) closer to the image features (any pixel vector q(240)) of a specific subregion of a specific image (see the lower part of Figure 6).

[0079] Now, in order to realize the learning process shown in Figure 6, in this embodiment, the learning model 400 is trained to optimize (minimize) the (overall) evaluation function (loss function) L, which has three types of evaluation terms (evaluation functions) Ltask, Lclst, and Laux (described later). The evaluation function L is expressed as a linear combination of the three types of evaluation terms Ltask, Lclst, and Laux, for example, as shown in equation (2) below. The values ​​λc and λa are hyperparameters for balancing the evaluation terms.

[0080]

number

[0081] The following sections will explain each evaluation item—Ltask, Lclst, and Laux—in turn.

[0082] <Evaluation Item Ltask> The evaluation term Ltask is an evaluation term for distance learning (metric learning) based on multiple integrated similarity vectors 280 corresponding to multiple images. The evaluation term Ltask can be expressed, for example, by the following equation (3). Note that the symbol represented by parentheses with a "+" in the lower right corner means that the larger of the value v inside the parentheses and zero will be output. In other words, this symbol represents max(v,0).

[0083]

number

[0084] Here, distance dap is the distance between the feature vector (in this case, the integrated similarity vector 280) corresponding to a certain image (anchor image) and the feature vector (in this case, the integrated similarity vector 280) corresponding to another image (positive image) belonging to the same class. On the other hand, distance dan is the distance between the feature vector (in this case, the integrated similarity vector 280) corresponding to the same image (anchor image) and the feature vector (in this case, the integrated similarity vector 280) corresponding to an image (negative image) belonging to a different class. The combination of an anchor image and a positive image is also called a positive pair, and the combination of an anchor image and a negative image is also called a negative pair. Distance dap is the distance between the integrated similarity vectors 280 of a positive pair, and distance dan is the distance between the integrated similarity vectors 280 of a negative pair.

[0085] Equation (3) shows an evaluation function that reduces the distance dap between the element of interest (anchor) and the same-classified element (positive) to below a certain level, while increasing the distance dan between the element of interest and the different-classified element (negative) to above a certain level. The value m is a hyperparameter that indicates the margin. The intention is to increase the distance between feature vectors of negative pairs to a value of (β+m) or more, and to bring the distance between feature vectors of positive pairs to within a value of (β-m). The value β is a learning parameter for each class (each anchor), and is an adjustment parameter that adjusts the degree of adjustment of positions in the feature space between classes (between anchors).

[0086] In equation (3), the Ltask is calculated by averaging the sum Σ obtained by adding the values ​​in curly braces over multiple training images, and then dividing the sum by the total number of such images N (number of anchors).

[0087] By performing the learning process in a manner that minimizes such an evaluation function Ltask, distance learning based on the integrated similarity vector 280 is realized. Specifically, in the feature space, multiple integrated similarity vectors 280 corresponding to multiple input images of the same class (for example, the same type of bird) are placed close to each other. On the other hand, multiple integrated similarity vectors 280 corresponding to multiple input images of different classes (different types of birds) are placed far apart from each other.

[0088] <Evaluation item Laux> The evaluation term Laux is an evaluation term for distance learning based on multiple sub-feature vectors 290 corresponding to multiple images.

[0089] The evaluation term Laux can be expressed, for example, by equation (4) below. Equation (4) is the evaluation function (evaluation term) that implements metric learning on the sub-feature vector 290 as described above.

[0090]

number

[0091] Each value in equation (4) is equivalent to the corresponding values ​​in equation (3). The difference is that the distance related to the sub-feature vector 290 is considered instead of the distance related to the integrated similarity vector 280.

[0092] Here, the distance d'ap is the distance between two positive pair of feature vectors (in this case, sub-feature vectors 290) relating to a given image (anchor image). On the other hand, the distance d'an is the distance between two negative pair of feature vectors (in this case, sub-feature vectors 290) relating to the same image (anchor image). The value m' is a hyperparameter indicating the margin. The intention is to bring feature vectors of the same class closer together (within the value (β'-m')) and to keep feature vectors of different classes further apart (beyond the value (β'+m')). The value β' is a learning parameter for each class (each anchor), and is an adjustment parameter that adjusts the degree of position adjustment in the feature space between classes (between anchors).

[0093] The evaluation term Laux is an auxiliary evaluation term. In this embodiment, the evaluation term Laux is considered, but it is not necessarily required (the evaluation term Laux may be omitted). However, by considering the evaluation term Laux, it is possible to construct the CNN (feature extraction layer 320) more appropriately, and consequently, to improve inference accuracy.

[0094] Furthermore, the evaluation function (evaluation term) for implementing metric learning with respect to the sub-feature vector 290 is not limited to the evaluation function in equation (4), but may be other evaluation functions (for example, various triplet losses (loss functions that make the distance between negative pairs greater than the distance between positive pairs)). The same applies to the evaluation function (equation (3)) for implementing metric learning with respect to the integrated similarity vector 280.

[0095] <Evaluation Term Lclst ​​and Learning Process> The evaluation term Lclst ​​is an evaluation term that brings each prototype vector p closer to an image feature (any pixel vector q) in a subregion of any image. More specifically, the evaluation term Lclst ​​is an evaluation term that brings the prototype vector pk of the prototype PTk to which each image i belongs closer to any pixel vector q in the feature map 230 corresponding to each image i (see Figure 6, lower panel). As will be described later, each prototype vector pk is brought closer to any pixel vector q in the feature map 230 corresponding to each image i, depending on the distribution prototype belonging degree Bi(k) of each prototype PTk to which each image i belongs. Therefore, each prototype vector pk may be brought closer to two or more pixel vectors q that exist in different images.

[0096] In the learning process of this embodiment, distance learning (metric learning) is basically performed on the integrated similarity vector 280. Specifically, by considering the evaluation term Ltask, the learning model 400 is trained so that the distribution of the integrated similarity vector 280 in the feature space approaches an ideal distribution state.

[0097] In the learning process of this embodiment, the learning model 400 is further trained to bring each prototype vector pk closer to the image features (any pixel vector q) of any subregion of any image by considering the evaluation term Lclst.

[0098] The following will primarily explain this type of learning process (learning process based on the evaluation term Lclst).

[0099] Figures 9 and 10 are flowcharts detailing the learning process (step S11 (Figure 5)) of this embodiment. In the flowchart of Figure 9, the learning process related to the evaluation term Lclst ​​is mainly shown, while the learning processes related to other evaluation terms Ltask and Laux are shown in a simplified manner. Figure 11 is a flowchart that shows in detail the process of a part of Figure 9 (step S21). Furthermore, Figure 12 is a diagram that conceptually shows the learning process related to the evaluation term Lclst.

[0100] The learning process related to the evaluation term Lclst ​​can be broadly divided into three (or four) processes (steps S41, S42, (S43, S44)), as shown in Figures 9 and 12.

[0101] In step S41, the controller 31 determines, for each of the multiple classes to which the multiple images for training are labeled, the class-belonging prototype PT(PTk) and the degree of belonging Bk of the said class-belonging prototype PT(PTk) to that class. Specifically, for each of the multiple classes (classes of interest), the belonging prototype PTk, which is a prototype belonging to one class (class of interest), and the prototype belonging degree Bk, which indicates the degree to which the said belonging prototype PTk belongs to that class, are determined. The prototype belonging degree Bk can also be expressed as the degree to which each prototype PTk represents the image features of one class. Step S41 includes steps S21 and S22 (Figure 9).

[0102] In step S42, the controller 31 calculates the distributed prototype belonging degree Bik, which is the prototype belonging degree for each image. The distributed prototype belonging degree Bik is the degree of belonging obtained by distributing the prototype belonging degree Bk of the belonging prototype PTk in each class to each of two or more images within the same class (for example, the same type of bird) based on predetermined criteria. The distributed prototype belonging degree Bik is also referred to as the distributed prototype belonging degree Tik. The distributed prototype belonging degree Bik (Tik) is also called the image-specific prototype belonging degree. Step S42 includes step S23.

[0103] In steps S43 and S44, the controller 31 performs a learning process so that each prototype vector p approaches one of the pixel vectors in the feature map corresponding to each image, according to the distribution prototype belongingness Bik of each image. In other words, the learning process is performed so that each prototype vector p approaches the closest pixel vector q among the multiple pixel vectors q in the feature map 230 of each image.

[0104] When the controller 31 performs a learning process (such as metric learning) based on multiple integrated similarity vectors 280 corresponding to multiple images, it also performs the learning processes of steps S43 and S44. In other words, when a learning process (including distance learning) based on the integrated similarity vectors 280 is performed, the learning model 400 is also trained so that each prototype vector p approaches one of the pixel vectors in the feature map corresponding to each image. Through this machine learning, each parameter in the learning model 400 (each parameter related to the prototype vector p and the convolutional neural network 220, etc.) is learned. Steps S43 and S44 include steps S24, S25, and S26.

[0105] In detail, in step S43, the evaluation function L (including evaluation terms Ltask, Lclst, etc.) is determined (steps S24, S25), and in step S44, a learning process based on the evaluation function L is executed (step S26).

[0106] The following will explain the process in detail, starting from step S41. First, the process in step S21 (see also Figure 11) within step S41 is executed. In step S21, the belonging prototype PTk and its degree of belonging Bk belonging to each class are (provisionally) determined for each of the multiple classes.

[0107] Specifically, in steps S211 to S215 (Figure 11), the controller 31 executes a prototype selection process to select a prototype to which one class (the class of interest) belongs. Figure 13 is a conceptual diagram showing the prototype selection process. The prototype selection process will be explained with reference to Figure 13.

[0108] In this prototype selection process, first, a predetermined image IMGi belonging to one class (the class of interest) from among multiple training images is compared with multiple comparison images IMGj belonging to classes other than the predetermined class (the negative class). Then, based on the comparison results, the prototype belonging to the predetermined class (and the prototype belonging to the predetermined image PT) is selected (steps S211 to S213).

[0109] In Figure 13, the i-th image IMGi on the left is focused on as a designated image (focus image) belonging to one class (focus class) from among multiple training images. Furthermore, multiple comparison images IMGj (see the right side of Figure 13) belonging to classes other than the designated one (negative classes) exist as comparison targets for this designated image. For example, one sample (designated image) within a certain class (e.g., a bird of the species "Green Jay") is the i-th image IMGi (focus image). Also, if the number of samples (training images) in a training unit minibatch is 100, and 5 samples are prepared for each of the 20 classes, then the number of samples belonging to classes other than the designated one (negative classes) is 95. In this case, the single i-th image IMGi (focus image in the left column of Figure 13) is compared with each of the 95 images IMGj (see the right column of Figure 13) in the negative class.

[0110] In detail, the controller 31 first performs a unit selection process (steps S211, S212 (Figure 11)) on one of the multiple comparison images (for example, the top image in the right column of Figure 13).

[0111] In step S211, the difference vector Δs (=si-sj) is calculated. Vector si is the unified similarity vector si(280) obtained by inputting the image of interest (i-th image) IMGi into the learning model 400. Vector sj is the unified similarity vector sj(280) obtained by inputting one of the multiple comparison images (j-th image) IMGj into the learning model 400. The difference vector Δs is the vector obtained by subtracting vector sj from vector si.

[0112] Next, in step S212, the prototype corresponding to the largest component (the largest positive component) among the multiple (Nc) components ΔSk in the difference vector Δs is selected as the belonging prototype to the predetermined image class (the class of interest). For example, in the difference vector Δs in the lower left of Figure 13, the prototype PT1 corresponding to the largest component ΔS1 among the multiple components ΔSk is selected as the belonging prototype PT.

[0113] In the unified similarity vector si, the value of the component corresponding to the prototype PT that represents the features of the target image (i-th image) IMGi of the class of interest appears as a larger value than the values ​​of the other components. On the other hand, in the unified similarity vector sj, the value of the component corresponding to the prototype PT that represents the features of the comparison image (j-th image) IMGj of classes other than the class of interest (negative classes) appears as a larger value than the values ​​of the other components. Therefore, the prototype PT that prominently represents the features of the class of interest appears as a large component with respect to the target image IMGi, and conversely appears as a small component with respect to the comparison image IMGj. Thus, the prototype PT corresponding to the component with a relatively large value (for example, the largest value) among the multiple components ΔSk of the difference vector Δs is the prototype PT that prominently represents the features of the class of interest. In step S212, taking these characteristics into consideration, the prototype PT corresponding to the largest component among the multiple components ΔSk of the difference vector Δs is selected as the prototype PT to which the class of interest belongs.

[0114] Thus, the unit selection process (S211, S212) is a process that selects a prototype belonging to a class based on a comparison process between one image (the image of interest) within that class and one comparison target image (an image from another class) (a process that selects on a unit basis of the image of interest).

[0115] This unit selection process is also performed for the remaining (for example, 94) comparison images out of the multiple comparison images (for example, 95) (S213). In other words, the unit selection process is performed for multiple comparison images while changing one comparison image to another. Note that in the second difference vector Δs near the lower center of Figure 13, prototype PT3 corresponding to the largest component ΔS3 among the multiple components ΔSk is selected as the associated prototype PT. For the remaining (93 (=95-2)) difference vectors Δs, the prototype PT corresponding to the largest component is also selected as the associated prototype PT.

[0116] In the bottom row of Figure 13, prototype PTs selected through multiple (for example, 95) unit selection processes are listed. For example, through these multiple unit selection processes, prototypes PT1, PT3, and PT4 are selected as affiliated prototype PTs (affiliated prototype PTs of the class of interest) based on a predetermined image (first image IMG1). The number of times prototypes PT1, PT3, and PT4 are selected (selection counts) are "40", "30", and "25", respectively (see the bottom row of Figure 13 and the upper left of Figure 14). Note that the prototype PTs (image representations (blue head, blue tail of a bird, etc.)) in the bottom row of Figure 13, etc., are shown in their ideal state after training is complete (image features corresponding to specific subregions of a specific sample image). The content of each prototype PT gradually changes during training.

[0117] In this way, by comparing a predetermined image belonging to one class with multiple comparison images from other classes, at least one affiliated prototype PT (affiliated prototype PT based on the predetermined image) belonging to that class is selected based on the predetermined image belonging to that class. A selection count calculation process is also performed to count the number of selected prototypes for each affiliated prototype.

[0118] Next, in step S214, the controller 31 performs the processing in steps S211 to S213 for each of the reference images within the class of interest, treating each reference image as the image of interest. Specifically, the controller 31 performs prototype selection processing, etc., for each of the (N-1) images (for example, 4 images) other than the predetermined image belonging to a class (class of interest), based on comparison with each of the multiple comparison target images (for example, 95 images).

[0119] For example, the unit selection process for the second image of the focus class (the new focus image) is repeatedly performed on 95 negative images, resulting in the selection of prototypes PT1, PT2, and PT4 as affiliated prototypes based on the second image. The number of times each prototype PT1, PT2, and PT4 are selected (selection counts) are "30", "30", and "35", respectively (see upper right of Figure 14).

[0120] Similarly, by repeatedly performing the unit selection process for the third image of a certain class on 95 negative images, for example, prototypes PT1, PT2, PT3, and PT4 are selected as affiliated prototypes PT based on the third image, and the number of selections (selection counts) for each prototype PT1, PT2, PT3, and PT4 are "35", "30", "25", and "5", respectively.

[0121] Next, in step S215, the controller 31 aggregates the number of selected prototype PTs for each of the N images (5 in this case) belonging to the class of interest. Specifically, the controller 31 determines the number of selected prototypes for each belonging prototype in the class of interest as the average value obtained by dividing the total number of selected prototypes selected in the prototype selection process by the number of images (N in this case, 5) belonging to the class of interest.

[0122] For example, consider a case where, for five reference images of a class, the number of prototypes PT1 belonging to that class are selected as "40", "30", "35", "30", and "40". In this case, the average value obtained by dividing the sum of these five values, "175", by the number of reference images, "5", is "35", which is determined as the number of selected prototypes PT1 belonging to that class (average number of selections) (see PT1 in the lower part of Figure 14).

[0123] Then, in step S216, the controller 31 determines the prototype affiliation degree Bk(Ya) of each affiliated prototype PTk for that class based on the number of times each affiliated prototype has been selected (number of selections) in the prototype selection process.

[0124] Specifically, the controller 31 calculates the number of selected prototypes PT belonging to each class (the class of interest) by the number of comparison images (for example, 95), and uses this value to determine the prototype belonging degree Bk of each prototype belonging to that class (Yc). (Yc) The following is calculated (provisionally). For example, let's assume that the number of selections (average number of selections) for each of the prototypes PT1, PT2, PT3, and PT4 belonging to the class CS1 is "35", "15", "15", and "30", respectively (Figure 14, bottom). In this case, the degrees of belonging B1, B2, B3, and B4 for each of the prototypes PT1, PT2, PT3, and PT4 belonging to class CS1 are calculated as "35 / 95", "15 / 95", "15 / 95", and "30 / 95".

[0125] In this way, the belonging prototype PTk of a particular class and the degree of belonging of that belonging prototype PTk to that particular class, Bk, are (provisionally) calculated.

[0126] Furthermore, in step S217, the processes in steps S211 to S216 are repeatedly executed while changing the class of interest to a different class. In other words, the controller 31 also executes a prototype selection process to select a prototype belonging to a class other than the one mentioned above. Through this, the controller 31 determines the prototypes belonging to each of the multiple classes and the degree of belonging of those prototypes.

[0127] For example, prototypes PT4 and PT5 are determined to belong to another class CS2, and their degrees of belonging, B4 and B5, are calculated as "30 / 95" and "65 / 95" respectively (see upper right of Figure 16). For other classes such as CS3, their respective prototypes and degrees of belonging are also calculated.

[0128] Next, we will explain step S22 of step S41 (Figure 9).

[0129] In this case, among the prototype PTs belonging to each class determined in step S21 above, there may be prototype PTs that are selected across multiple classes (Y1, Y2, ...). In other words, among multiple prototype PTs, a particular prototype PT may be selected as a biased group of prototype PTs.

[0130] For example, prototype PT4 may be selected as the affiliated prototype PT for 10 classes, CS1 to CS10 (see Figure 16, upper panel). Such a prototype PT (PT4) may have features common to many images, such as background image features. In other words, the prototype PT4, etc., does not necessarily prominently represent the image features of a particular class.

[0131] Therefore, in this embodiment, in order to suppress the influence of such prototype PTs, the controller 31 suppresses the degree of belonging of a prototype PT selected as a belonging prototype vector across multiple classes (step S22). In other words, if a particular prototype PT is selected in a biased manner as a belonging prototype PT, a process to suppress the bias of prototype PTs (De-Bias processing) is performed (see Figure 15). Specifically, if a single prototype PTk belongs to two or more classes, the controller 31 suppresses the degree of belonging of that single prototype (prototype belonging degree Bk) in each of those two or more classes (each class Yc). (Yc) Reduce (normalize) the degree of belonging to a prototype PT that belongs to multiple classes. For example, change the degree of belonging to a prototype PT that belongs to only one class to a smaller value than the degree of belonging to a prototype PT that belongs to only one class.

[0132] Specifically, according to equation (5), etc., the degree of prototype affiliation Bk of each affiliated prototype PTk in each class Yc is calculated. (Yc) This will be corrected.

[0133]

number

[0134] Equation (5) is the value of its left-hand side (new degree of belonging Bk (Yc) This means replacing ) with the value on the right side. The fractional value on the right side of equation (5) is the numerator (the original degree of belonging Bk of each belonging prototype PTk in a certain class Yc). (Yc) This is the value obtained by dividing ) by the denominator. The denominator is the original degree of affiliation Bk of each affiliation prototype PTk. (Yc) It is the larger of the sum of all classes (total value) and a predetermined value ε.

[0135] For example, as shown in Figure 16, suppose that prototype PT4 is selected as a member prototype PT across a large number of classes (e.g., 10), and the sum of the above values ​​(the sum of its membership degree B4) across those 10 classes is "300 / 95". In this case, when the predetermined value ε=1, the original membership degree B4 "30 / 95" of a member prototype PT4 in a certain class CS1 is reduced (normalized) to "1 / 10" (=(30 / 95) / 300 / 95) by equation (5). Similarly, in other classes such as CS2, the membership degree B4 of a member prototype PT4 in each class is calculated by dividing the original membership degree in each class by the larger of the above sum and the predetermined value ε (e.g., "1"). Note that if the original membership degree is divided by "1", the modified membership degree will be the same value as the original membership degree.

[0136] Subsequently, within each class, the sum of the affiliation degrees Bk of multiple affiliated prototype PTs (of different types) is adjusted to equal "1". For example, as shown in the middle of Figure 16, consider the case where the affiliation degrees B1, B2, B3, and B4 of each affiliated prototype PT1, PT2, PT3, and PT4 in class CS1 before adjustment (and after the bias correction described above) are "35 / 95", "15 / 95", "15 / 95", and "1 / 10", respectively. In this case, the affiliation degrees B1, B2, B3, and B4 are adjusted to "70 / 149", "30 / 149", "30 / 149", and "19 / 149", respectively (see the bottom of Figure 16). In other classes (CS2, etc.), the affiliation degrees of their respective affiliated prototype PTs are also adjusted to equal "1".

[0137] Thus, in step S22, the main process involves adjusting the degree of prototype affiliation across classes. In step S22, the degree of affiliation of a prototype PT that belongs to multiple classes is reduced compared to the degree of affiliation of a prototype PT that belongs to only a single class.

[0138] Note that while the example here mainly uses the case where ε=1, it is not limited to this, and the value ε may be smaller than 1 (for example, 0.1). In that case, for example, the degree of belonging of prototype PTs belonging to a single class (or a small number of classes) is temporarily changed to "1" (or "1 / 2", etc.), and the degree of belonging of prototype PTs belonging to multiple classes is changed to a relatively small value. Then, the total value of the degrees of belonging of prototype PTs belonging to the same class is normalized to "1".

[0139] In this case, each Bk is modified as shown in Figure 17, for example.

[0140] The top row of Figure 17 shows the same state as the top row of Figure 16.

[0141] In the middle section of Figure 17, the original degree of belonging B4 "30 / 95" for the affiliated prototype PT4 is similarly reduced (normalized) to "1 / 10" (=(30 / 95) / 300 / 95). On the other hand, due to the small value of ε in equation (5), the degree of belonging B1 for the affiliated prototype PT1 is changed to "1" (=(35 / 95) / (35 / 95)). The degrees of belonging B2 and B3 for the other affiliated prototypes PT2 and PT3 are also corrected to "1". With such corrections, the degrees of belonging ("1") for prototypes PT1, PT2, and PT3 belonging to a single class are (more reliably) changed to a value that is (relatively) larger than the degree of belonging for prototype PT4, which belongs to two or more classes. Subsequently, within each class, the sum of the degrees of belonging Bk of multiple affiliated prototypes PT (of different types) is adjusted to "1". For example, as shown in the bottom row of Figure 17, in class CS1, the membership degrees B1, B2, and B3 are all changed to "10 / 31", and the membership degree B4 is changed to "1 / 31".

[0142] In step S22, the above processing is performed.

[0143] In the next step S42 (step S23) (Figure 9), the controller 31 determines the distributed prototype belonging degree Bik (Tik), which is the prototype belonging degree for each image (see Figures 18 and 19). As described above, the distributed prototype belonging degree Bik is the belonging degree obtained by distributing the prototype belonging degree Bk of the belonging prototype PTk in each class to each of two or more images within the same class (for example, the same type of bird) based on predetermined criteria.

[0144] Specifically, the controller 31 distributes the prototype affiliation degree Bk of one affiliated prototype belonging to one class to N images belonging to that class, and determines the distributed prototype affiliation degree Bik(Tik) for each of the N images (IMGi). In this process, the original prototype affiliation degree Bk is distributed to each image IMGi such that the smaller the distance (higher the similarity) between the prototype vector p of the affiliated prototype PT and the most similar pixel vector q in each image IMGi, the larger the distributed prototype affiliation degree Bik for each image.

[0145] For example, if the first distance D1 (described below) is greater than the second distance D2 (described below), the controller 31 determines the distribution prototype belonging degree (e.g., T1k) for one image (e.g., IMG1) as a smaller value than the distribution prototype belonging degree (e.g., T2k) for the other image (e.g., IMG2). Here, the first distance D1 is the distance (e.g., C1k) between the pixel vector (most similar pixel vector) that is most similar to the prototype vector pk of the assigned prototype PTk among multiple pixel vectors q in the feature map corresponding to "one image" (e.g., IMG1) among the N images, and the prototype vector pk. The second distance D2 is the distance (e.g., C2k) between the pixel vector (most similar to the prototype vector pk of the assigned prototype) among multiple pixel vectors in the feature map corresponding to "the other image" (e.g., IMG2) among the N images, and the prototype vector pk. Each distance D1 and D2 is the distance between the respective most similar pixel vector q and the prototype vector pk (see equation (6) described later). If k=2 and C12>C22, the distribution prototype affiliation degree T12 is determined to be smaller than the distribution prototype affiliation degree T22.

[0146] In detail, this distribution process can be viewed as a discrete optimal transport problem. Figuratively speaking, this distribution process is a problem (a discrete optimal transport problem) of allocating the required quantities for multiple destinations to multiple delivery stores in order to minimize the total transport cost (the sum of the transport costs corresponding to the transport distance between each delivery store and each destination and the transport quantity from each delivery store to each destination). In this embodiment, each image i can be considered as each delivery store, each destination as each prototype PTk, and the required quantity at each destination as the degree of belonging Bk of each prototype PTk. The degrees of belonging Bk of multiple prototype PTs are distributed (assigned) to multiple images in order to minimize an evaluation value corresponding to the transport cost. This evaluation value is an evaluation value (see equation (7)) corresponding to the distance Cik (see equation (6)) and the magnitude (allocation amount) of the distributed prototype degree of belonging Tik. Distance Cik is the distance between each prototype vector pk and the most similar pixel vector q (in the feature map 230 corresponding to each image i), and the distributed prototype belonging degree Tik is the belonging degree distributed to each image i for each prototype PTk. Such discrete optimal transport problems can be solved by methods such as the Sinkhorn-Knopp algorithm.

[0147]

number

[0148] In equation (6), Cik is the jth pixel vector qj of the i-th image, where q is each of the multiple pixel vectors q in the feature map 230 of the i-th image. (i) Cik is the minimum distance between the k-th prototype vector pk and the pixel vector q (most similar pixel vector) that is most similar to the k-th prototype vector pk in the feature map 230 of the i-th image. In short, Cik is the minimum distance between any pixel vector q (qj) in the i-th image. (i) This is the minimum distance between the vector and the prototype vector pk. Note that Cik in equation (6) is equivalent to Sk in equation (1).

[0149]

Number

[0150] Equation (7) is an equation showing the evaluation value in the above distribution process. Tik(Bik) represents the assigned prototype membership degree assigned (allocated) to the i-th image IMGi among the membership degrees Bk of the k-th prototype PTk in the target class. The value of Equation (7) corresponds to the "total transportation cost" in the above discrete optimal transportation problem.

[0151]

Number

[0152] The right side of Equation (8) (the upper equation) represents the minimized evaluation value among a plurality of evaluation values (see Equation (7)) obtained by varying T(Tij) in the distribution process for a certain class Yc.

[0153] Equation (8) also shows two conditions regarding Tik. One condition is that the sum of the assigned prototype membership degrees Tik of the k-th prototype PTk for the i-th image, for a plurality (Ns) of images within the class (for each prototype PTk), is equal to Ns times the membership degree Bk of the k-th prototype PTk to the class Yc. (Yc) The other condition is that the sum of the assigned prototype membership degrees Tik of the k-th prototype PTk for the i-th image, for a plurality of prototypes PTk (for each image), is equal to "1". Here, Ns is the number of images (assigned images) within the same class.

[0154] That is, the value Lclst of Equation (8) is the value obtained by minimizing (optimizing) the evaluation value of Equation (7) while following the two conditions in Equation (8). In step S23 and the like, the solution and the like of the above distribution problem (discrete optimal transportation problem) are used. Specifically, the optimal solution (including the approximate optimal solution) shown in Equation (8), and Cik, Tik (distribution results, etc.) constituting the optimal solution are used.

[0155] Figure 19 shows the distribution result.

[0156] Here, we assume that in a single class, there are three affiliated prototypes PT1, PT2, and PT3, and that the degree of belonging Bk (B1, B2, B3) of each prototype PTk to that class is "5 / 12", "3 / 12", and "4 / 12". Furthermore, the distance C12 (the minimum distance between the pixel vector q and the prototype vector p2 in image IMG1) is very large, and the distance C22 (the minimum distance between the pixel vector q and the prototype vector p2 in image IMG2) is very small. Also, the distance C13 is very small, and the distance C23 is very large.

[0157] Thus, when distance C12 is greater than distance C22, the controller 31 determines the assigned prototype belonging degree T12(B12) for image IMG1 as a smaller value than the assigned prototype belonging degree T22(B22) for image IMG2. For example, the assigned prototype belonging degree T12 is "0 (zero)" and the assigned prototype belonging degree T22 is "1 / 2". In short, among two or more images in the same class, images with a relatively small similarity to the prototype vector pk are assigned a relatively small belonging degree.

[0158] Furthermore, if distance C13 is less than distance C23, the controller 31 determines the assigned prototype belonging degree T13 for image IMG1 as a larger value than the assigned prototype belonging degree T23 for image IMG2. For example, the assigned prototype belonging degree T12 is "2 / 3" and the assigned prototype belonging degree T22 is "0 (zero)". In short, among two or more images in the same class, images with a relatively high similarity to the prototype vector pk are assigned a relatively large belonging degree.

[0159] Furthermore, the allocation prototype affiliation degree T11 for image IMG1 and the allocation prototype affiliation degree T21 for image IMG2 are determined based on factors such as the relative magnitudes of distances C11 and C12.

[0160] Each Tik is determined to satisfy the two conditions of equation (8). As a result, for example, the distribution prototype affiliations T11, T12, and T13 are determined as "1 / 3", "0 (zero)", and "2 / 3", and the distribution prototype affiliations T21, T22, and T23 are determined as "1 / 2", "1 / 2", and "0 (zero)" (see the leftmost part of Figure 19).

[0161] In this way, the distribution process for a single class is carried out.

[0162] The controller 31 applies this distribution process to other classes and repeats the same distribution process to obtain evaluation values ​​(optimized evaluation values ​​in equation (8)) for each of the multiple classes.

[0163] Then, the controller 31 calculates multiple evaluation values ​​Lclst ​​for multiple classes. (Yc) The evaluation term Lclst ​​of the evaluation function L is calculated by further adding the (optimized evaluation value) (see equation (9)) (step S24).

[0164]

number

[0165] Equation (9) is an expression that shows the evaluation function (evaluation term) Lclst. The evaluation term Lclst ​​in equation (9) is the class-specific evaluation term Lclst ​​defined in equation (8). (Yc) This is the sum of values ​​obtained by adding up values ​​across multiple classes.

[0166] Furthermore, in step S25 (Figure 9), the controller 31 calculates the evaluation function L by adding the evaluation terms Ltask and Laux in addition to the evaluation term Lclst ​​obtained by equation (9), and adding them according to equation (2).

[0167] In step S26, the controller 31 performs a learning process (machine learning) to minimize (optimize) this evaluation function L. More specifically, this learning process is performed by repeatedly executing steps S21 to S25.

[0168] In this case, the controller 31, in particular, evaluates the Lclst ​​(and Lclst (Yc) The learning process is performed to minimize the following: In other words, the learning process is performed so that each prototype vector pk approaches one of the pixel vectors q in the feature map corresponding to each image i, according to the distribution prototype belongingness Tik (Bik) of each image i. In other words, the learning process is performed so that each prototype vector pk approaches the closest pixel vector q among the multiple pixel vectors q in the feature map 230 of each image. As a result, the learning model 400 (each prototype vector p and CNN220 etc. (especially prototype vector p)) is trained so that each prototype vector pk approaches one of the image features of the multiple images used for training.

[0169] Through this process, the learning model 400 is trained (machine learning) and a trained model 420 is generated.

[0170] Through the learning process described above, the learning model 400 is trained to optimize (minimize) the evaluation function L, which has evaluation terms Ltask, Lclst, and Laux. More specifically, the learning model 400 is trained to optimize (minimize) each of the evaluation terms Ltask, Lclst, and Laux.

[0171] Distance learning on the integrated similarity vector 280 proceeds by minimizing the evaluation term Ltask. As a result, in the feature space of the integrated similarity vector 280, multiple feature vectors corresponding to multiple input images of the same class (for example, the same type of bird) are placed close to each other. On the other hand, multiple feature vectors corresponding to multiple input images of different classes (different types of birds) are placed far apart from each other.

[0172] Furthermore, distance learning on the sub-feature vector 290 proceeds through the action of minimizing the evaluation term Laux. According to this distance learning, in the feature space for the sub-feature vector 290, multiple feature vectors corresponding to multiple input images of the same class are placed close to each other, while multiple feature vectors corresponding to multiple input images of different classes are placed far apart from each other.

[0173] The sub-feature vector 290 is a vector obtained by aggregating the outputs (feature maps 230) from the CNN220 channel by channel. In other words, the sub-feature vector 290 is an output vector from a location within the learning model 400 that is close to the output location from the CNN220 (compared to the integrated similarity vector 280). In this embodiment, distance learning is performed using the sub-feature vector 290 having such characteristics. Therefore, it is possible to construct a CNN220 with appropriate feature extraction capabilities more accurately compared to the case where the evaluation term Laux is not considered (when only the evaluation term Ltask is considered).

[0174] Furthermore, by minimizing the evaluation term Lclst, each prototype vector pk is learned to approach the most similar pixel vector q, etc. Therefore, each prototype vector pk is learned to reflect the image features of a specific subregion of a particular image. In other words, each trained prototype vector pk is learned to represent the concept (the concept of the image feature) of each prototype PTk.

[0175] In particular, each prototype vector pk is learned to approach the image-specific most similar pixel vector q within each image, depending on the distribution prototype belonging degree Tik of each image. Furthermore, each prototype PT can belong to two or more classes. Therefore, each prototype vector p can be learned to reflect similar features between different images of different classes. Also, since it is not necessary to prepare a predetermined number of dedicated prototype vectors for each class, it is possible to construct prototype vectors p efficiently.

[0176] Furthermore, each prototype vector p is learned to approach different images within the same class according to the image-specific prototype belonging degree (which differs for each image). More specifically, each prototype vector p is learned to reflect image features according to the image-specific prototype belonging degree (which differs for each image) even within the same class. Therefore, it is possible to construct prototype vectors p efficiently.

[0177] Furthermore, compared to conventional techniques that use ProtoPNet for class classification, it is not necessary to prepare each prototype vector p as a prototype vector p dedicated to each class. In other words, the relationship between the prototype vector p and the class does not need to be fixed. Therefore, as mentioned above, it is possible to realize a learning process (such as distance learning) that approaches a certain image feature without fixing the relationship between the prototype vector p and the class. Consequently, in processes such as similar image search for unclassified images, it becomes possible to extract image features based on the prototype vector p and explain the reasoning behind the inference. Therefore, the explainability of the reasoning (especially "transparency": the property of being able to explain the inference result with a concept that can be understood by humans) can be improved.

[0178] <Replacement process for prototype vectors> Once the machine learning of the learning model 400 is completed as described above, the process proceeds to step S28 (Figure 10). In step S28, the controller 31 replaces each prototype vector pk in the learning model 400 (trained model 420) with the most similar pixel vector q (also denoted as qmk) (see Figure 20). Specifically, all (Nc) prototype vectors pk (k=1,...,Nc) in the trained model 420 after training are replaced with the most similar pixel vector qmk for each of them. The most similar pixel vector qmk is the pixel vector that is most similar to each prototype vector pk among multiple pixel vectors q in multiple feature maps for multiple images used for training. The most similar pixel vector qmk is a pixel vector corresponding to a specific region within a specific image (a vector indicating the image features of that specific region).

[0179] Specifically, the controller 31 first inputs an image (the i-th image) into the trained 420 to obtain a feature map 230. Then, the controller 31 finds the pixel vector q that is most similar to the prototype vector pk (for example, p1) of interest from among the multiple pixel vectors q in the feature map 230. The pixel vector q that is most similar to the feature map 230 of the i-th image is also called the image-specific most similar pixel vector q. The similarity between the two vectors q and pk can be calculated using cosine similarity, etc. (see equation (1) or equation (6)).

[0180] The controller 31 repeats the same operation for multiple images. This extracts multiple feature maps 230 corresponding to multiple (e.g., 100) images for training, and for each of these feature maps 230, the image-specific most similar pixel vector q for the prototype vector of interest pk (e.g., p1) is determined.

[0181] The controller 31 then identifies the image-specific most similar pixel vector q that is most similar to the prototype vector pk of interest from among multiple (for example, 100) image-specific most similar pixel vectors q for multiple images, and designates it as the most similar pixel vector q(qmk). The controller 31 also identifies the image containing this most similar pixel vector qmk (for example, the first image) as the most similar image (the image that best possesses the features of the prototype vector pk).

[0182] In this way, the controller 31 finds the most similar pixel vector qmk for the prototype vector pk of interest.

[0183] Then, the controller 31 replaces the prototype vector pk of interest in the trained model 420 with the most similar pixel vector qmk for that prototype vector pk of interest (see Figure 20).

[0184] Furthermore, the most similar pixel vector qmk is obtained for each of the other prototype vectors pk in the same manner, and each prototype vector pk is replaced by its respective most similar pixel vector qmk.

[0185] This replacement modifies the trained model 420, completing the modified trained model 420 (Step S29 (Figure 10)).

[0186] In this way, the process in step S11 (Figure 5) (the learning phase of the learning model 400) is executed.

[0187] <1-5. Inference processing using learning model 400> Next, inference processing is performed using the trained model 400 (trained model 420) generated in step S11 (step S12 (Figure 5)).

[0188] For example, the inference process involves searching for images similar to a new image (the image to be inferred) 215 from among multiple images 213. More specifically, the inference process involves searching for images from among multiple images 213 (in this case, multiple images for training) that have a degree of similarity to the source image (also called the query image) 215 that is above a predetermined level (in other words, the distance between feature vectors (integrated similarity vector 280) in the feature space is below a predetermined distance). Alternatively, the inference process may involve searching for images similar to the query image in order of similarity.

[0189] The image to be inferred (query image) may be an image belonging to a class other than the class used for labeling the training data (image data of multiple images used for training, etc.) (known class). In other words, the image to be inferred may be an image of a known class or an image of an unclassified class. The inference process according to this embodiment (inference process using the above-described learning model 400) is particularly significant in that it can search for images similar to images of an unclassified class (not only can it successfully search for images similar to images of a known class), but it can also successfully search for images similar to images of an unclassified class.

[0190] The following explanation of this inference process will be given with reference to Figures 21 and 22. Figure 21 is a diagram illustrating the inference process using the integrated similarity vector 280(283) as the feature vector in the feature space. Figure 22 is a diagram showing an example of the inference process result.

[0191] First, the image processing device 30 inputs multiple training images (gallery images 213) into the trained model 420 and obtains the output from the trained model 420. Specifically, as shown in Figure 21 (especially the right side), each integrated similarity vector 280 (283) is obtained as the output (feature vector) for each input image 210 (213). The multiple integrated similarity vectors 283 are multiple integrated similarity vectors 280 output from the trained model 420 in response to the input of multiple training images into the trained model 420. Furthermore, each integrated similarity vector 280 (283) is generated, for example, as a 512-dimensional vector. Such integrated similarity vectors 283 (feature vectors) are obtained for each of the multiple input images 213 as vectors representing the features of each input image 213.

[0192] Similarly, the image processing device 30 inputs the target input image (query image) 215 to the learning model 420 and obtains the integrated similarity vector 280 (285) output from the learning model 420 as a feature vector (see left side of Figure 21). The integrated similarity vector 285 is the integrated similarity vector 280 output from the trained model 420 in response to inputting the query image 215 to the trained model 420. Note that the query image 215 is, for example, a different image from the multiple input images 213 (gallery images) (such as an image newly assigned for the search). However, it is not limited to this, and the query image 215 may be any of the multiple input images 213 (gallery images).

[0193] The image processing device 30 then searches for an image similar to the query image 215 from among multiple training images based on the unified similarity vector 285 and multiple unified similarity vectors 283.

[0194] Specifically, the image processing device 30 calculates the degree of similarity (for example, Euclidean distance, or the inner product (cosine similarity) between the vectors) between the feature vector 285 of the query image 215 and each of the multiple feature vectors 283 relating to the multiple input images 213. Furthermore, the multiple feature vectors 283 are sorted in descending order of similarity (descending order of similarity). More specifically, the multiple feature vectors 283 are sorted in ascending order of Euclidean distance (or descending order of cosine similarity).

[0195] Next, the image processing device 30 identifies one or more feature vectors 283 whose distance from the feature vector 285 in the feature space is less than or equal to a predetermined distance (i.e., the degree of similarity is greater than or equal to a predetermined degree) as feature vectors 285 of images that are (particularly) similar to the query image 215. In other words, the image processing device 30 recognizes the subjects in one or more input images 213 corresponding to the identified one or more feature vectors 285 as subjects similar to the subjects in the query image 215.

[0196] Furthermore, the image processing device 30 identifies the feature vector 283 with the smallest distance from the feature vector 285 in the feature space as the feature vector 285 of the image most similar to the query image 215. In other words, the image processing device 30 recognizes the subject in one input image 213 corresponding to the identified feature vector 285 as the subject most similar to the subject in the query image 215.

[0197] Figure 22 shows how multiple feature vectors 283 (shown as hatched white circles in Figure 22) corresponding to multiple input images 213 are distributed in the feature space. In Figure 22, three feature vectors 283 (V301, V302, V303) exist within a predetermined distance range from the feature vector 285 (see white star) of the query image 215.

[0198] In this case, for example, three images 213 corresponding to the three feature vectors 283 (V301, V302, V303) are extracted as similar images. Multiple feature vectors 283, including the three feature vectors 283, are sorted in descending order of similarity to feature vector 285 (ascending order of distance). Here, the three images 213 corresponding to the top three feature vectors 283 are recognized as images of subjects particularly similar to the subject of the query image 215.

[0199] Additionally, one image 213, which corresponds to the feature vector 283 (V301) that is closest to feature vector 285, is extracted as the similar image that is most similar to the query image 215.

[0200] However, this is not the only option; multiple input images 213 may simply be sorted in ascending order of distance (descending order of similarity) to the query image 215 (with respect to the feature vector 285). Even in this case, the image processing device 30 is essentially performing a process to find subjects similar to the subject in the query image in order of similarity (similar image search process). This process can also be described as an inference process for recognizing the subject in the query image.

[0201] Furthermore, in this embodiment, images similar to the image to be inferred are searched for from a plurality of training images, but this is not limited to this. For example, images similar to the image to be inferred may be searched from an image that includes images other than the plurality of training images.

[0202] <1-6. Explanation process for inference results 1> Next, we will explain the explanation process for the inference results (step S13 (Figure 5)).

[0203] For example, consider the case where it is inferred that "one image 213a (also referred to as 214) corresponding to feature vector 283 (V301) which is closest to feature vector 285 is the most similar image to query image 215" (see Figure 22). In other words, consider the case where the distance D between the unified similarity vector 285 corresponding to query image 215 and the unified similarity vector 284 corresponding to image 214 is determined to be the smallest (maximum similarity between the two vectors) among the distances D related to multiple combinations.

[0204] In this case, the image processing device 30 (controller 31) generates explanatory information that explains the basis for its inference (the basis on which the image processing device 30 inferred that image 213a is similar to query image 215). This explanatory information is then displayed on the display screen.

[0205] Figure 23 shows an example of how such explanatory information is displayed (an example of a display screen). For example, the entirety of Figure 23 is displayed on the display unit 35b. In this example, the image found to be most similar to the query image 215 (see top left) is displayed on the top right. In addition, "Judgment Basis 1" shows a partial image corresponding to the prototype PTmax (described below), and "Judgment Basis 2" shows a partial region image similar to the prototype PTmax within the query image 215 (and the position of that partial region image within the query image 215).

[0206] To achieve this display, the controller 31 first sorts the multiple components Sk of the feature vector (integrated similarity vector) 285 corresponding to the query image 215 in descending order. The prototype PT (also written as PTmax) corresponding to the largest component Smax among these multiple components Sk is the primary basis for similarity judgment. In other words, the primary basis for image similarity judgment is that the query image contains image features similar to a specific image feature related to prototype PTmax.

[0207] In particular, the replacement process described above (step S28) replaces the prototype vector pk of each prototype PTk with the nearest similar pixel vector qmk (see the dashed rectangle in Figure 23). In other words, each prototype vector pk is overwritten with the nearest similar pixel vector qmk. Therefore, each component Sk of the unified similarity vector 280 in the feature space represents its similarity to the nearest similar pixel vector qmk. Consequently, the nearest similar pixel vector q(qmax), which has been overwritten with the prototype vector p of prototype PT(PTmax), is the primary basis for similarity judgment.

[0208] Therefore, the controller 31 presents the user with the image features corresponding to the prototype PT (PTmax) having the largest component Smax (i.e., the image features corresponding to the overwritten most similar pixel vector qmax) as the basis for its decision.

[0209] For example, if the most similar pixel vector qmax, which is overwritten for the prototype vector p(pmax) of prototype PTmax, corresponds to a specific subregion R1 within image IMG1 (see the top row of Figure 20 and the bottom of Figure 23), the controller 31 presents the image of that specific subregion R1 to the user as an image showing the basis for the similarity judgment. Specifically, the controller 31 displays the subregion image corresponding to prototype PTmax (more specifically, the subregion image corresponding to the most similar pixel vector qmax) as "Judgment Basis 1".

[0210] Furthermore, the controller 31 identifies a region (specific similarity region Rq) similar to the specific subregion R1 within the query image (inference target image) 215, and presents the image of the specific similarity region Rq to the user. Specifically, as "judgment basis 2," the controller 31 overlays a rectangle (enclosing rectangle) surrounding the specific similarity region Rq within the query image 215 onto the query 215, thereby presenting the user with both the image features of the specific similarity region Rq and its position within the query image 215. Also as "judgment basis 2," the controller 31 displays an enlarged image of the specific similarity region Rq. This enlarged image is displayed near the display area of ​​the query image 215 that includes the enclosing rectangle (in this case, on the left side).

[0211] More specifically, the controller 31 first searches for the image feature most similar to the most similar pixel vector qmax within the query image 215. Specifically, from among the multiple pixel vectors q in the feature map 230 obtained by inputting the query image 215 into the trained model 420, the pixel vector q most similar to the most similar pixel vector qmax (= prototype vector pmax of prototype PTmax) is extracted. Then, the controller 31 displays the subregion image (specific similar region) corresponding to the extracted pixel vector q and its position within that image as "judgment basis 2".

[0212] Based on the presentation of such basis for judgment (presented by the device), the following understanding can be obtained, particularly based on "basis for judgment 1". Specifically, the user can understand that the image processing device 30 made a "similar" judgment based on the image features corresponding to the prototype PTmax (such as the image features of the "blue head" in the specific subregion R1).

[0213] Furthermore, based on "Judgment Basis 2," the following understanding can be obtained. Specifically, the user can identify the subregion (specific similarity region Rq) that the device inferred to be similar to the specific subregion R1 within the query image 215. By comparing the image features of the specific similarity region Rq with the image features of the specific subregion R1, the user can confirm that the image features of the specific subregion R1 ("blue head," etc.) are present in the image features of the specific similarity region Rq, and understand that the inference result is correct.

[0214] According to this explanatory information, the similarity of the query image 215 (the image to be inferred) can be explained very appropriately using the similarity with each replaced prototype vector pmax. Here, each replaced prototype vector pmax represents the image features of the most similar pixel vector qmax (i.e., the image features of a specific subregion (R1, etc.) of a specific image (IMG1, etc.) among the multiple images used for training). Therefore, the integrated similarity vector 280 output from the trained model 420 in the inference process represents the similarity with the image features of the replaced most similar pixel vector qmax (not the prototype vector pmax before replacement). Consequently, the image processing device 30 can explain to the user (human) that it has determined similarity based on whether or not it is similar to the image features ("blue head, etc.") of a specific subregion (R1, etc.) of the specific image (IMG1, etc.). In other words, it is possible to accurately obtain high "transparency" (the property of being able to explain the inference results in a concept that can be understood by humans).

[0215] Figure 23 shows a case where the query image 215 (the image to be inferred) belongs to a previously classified class, but it is not limited to this case; the query image 215 does not have to belong to a previously classified class.

[0216] Figure 24 shows a different display example from Figure 23. In Figure 24, the case where query image 215 does not belong to any previously classified class is shown.

[0217] Figure 24 shows another example of how explanatory information can be displayed (an example of a display screen).

[0218] Query image 215 is an image of a bird belonging to the unclassified class (a specific species of bird with reddish legs). The image most similar to such query image 215 is searched from a set of training images. In this case, the training classes include a first class Y1 (a certain species of bird with orange legs) and a second class Y2 (another species of bird with red legs).

[0219] As "Judgment Basis 1," a partial image corresponding to prototype PTmax is shown. This prototype PTmax is a prototype PT that belongs to multiple (in this case, two) classes Y1 and Y2.

[0220] Furthermore, as "basis for judgment 2," the query image 215 (an image of a bird with "reddish feet") shows a subregion image that is most similar to the prototype PTmax (a subregion image showing "reddish feet").

[0221] As described above, when distance learning of the integrated similarity vector 280 is performed based on the evaluation term Ltask, the evaluation term Lclst ​​is also taken into consideration. According to this, the prototype PT (and consequently PTmax) is learned to approach a certain pixel vector q in each image within each class, depending on the degree to which the prototype PT belongs to each class. For example, the prototype vector pmax of the prototype PTmax is learned to approach both the pixel vector q1 (corresponding to the "orange foot" in the image belonging to class Y1) and the pixel vector q2 (corresponding to the "red foot" in the image belonging to class Y2) (see the large dashed rectangle at the bottom of Figure 24). As a result, not only the bird images with "orange feet" but also the bird images with "red feet" are placed close together in the feature space of the integrated similarity vector 280. However, in the replacement process of step S28, the prototype vector pmax is assumed to have been replaced by pixel vector q1 of the pixel vectors q1 and q2.

[0222] The image processing device 30 presents explanatory information to the user as shown in Figure 24.

[0223] According to this explanatory information, the similarity of the query image 215 (the image to be inferred) can be explained very appropriately using the similarity with each replaced prototype vector pmax. Here, each replaced prototype vector pmax represents the image feature of the most similar pixel vector qmax (simply put, the "orange leg"). Therefore, the image processing device 30 can explain that it has determined similarity based on whether or not it is similar to the specific image feature ("orange leg"). In other words, it is possible to accurately obtain high transparency (the property of being able to explain the inference result in a concept that is understandable to humans).

[0224] Furthermore, a user who has access to such explanatory information can understand that, based on "Judgment Basis 1," images similar to query image 215 were searched based on prototype PTmax. In other words, the user can understand that the device determined the similarity of the images based on whether or not they are similar to a specific image feature ("orange foot") of the replaced prototype PTmax (which in this case is equal to pixel vector q1).

[0225] However, this is not the only way to interpret the situation; users can also make the following interpretations based on the information included in "Basis for Judgment 2."

[0226] In "Judgment Basis 2," the image features of similar areas within the query image (specifically, "reddish feet") are shown. Although the user cannot know that prototype PTmax is trained to reflect pixel vector q2 ("red feet"), the user can know the image features of prototype PTmax ("orange feet") and the image features in query image 215 ("reddish feet"). Based on this information, the user can infer (comprehend) that prototype PTmax is actually trained to represent the image feature "bright reddish feet" (an image feature that includes both "orange" and "red"). Therefore, the user can also interpret that the image processing device 30 is judging similarity based on the image feature "bright reddish feet." In particular, as mentioned above, relying on the effect of the evaluation value Lclst, the group of integrated similarity vectors 280 corresponding to groups of images having similar features (features corresponding to pixel vectors q1 and q2) are located close together in the feature space. Considering this, such an interpretation has a certain degree of rationality.

[0227] <1-7. Explanation process for inference results 2> In the above, the image processing device 30 explains the reason why query image 215 is similar to a specific image among the multiple images being searched. However, the image processing device 30 may also explain the reason why query image 215 is not similar to a specific image among the multiple images being searched. Such embodiments will be described below.

[0228] The following describes how the image processing device 30 explains the basis for its determination that the first image G1 and the second image G2 are not similar (see Figure 25). For example, the first image is the query image 215, and the second image is an image that has been determined (inferred) to be not similar to the query image 215 (such as one of several training images).

[0229] When determining similarity, the controller 31 obtains a feature vector (integrated similarity vector) 280 (also denoted as s1) corresponding to the first image G1 and a feature vector (integrated similarity vector) 280 (also denoted as s2) corresponding to the second image G2 (see Figure 25).

[0230] As described above, each component Sk of the integrated similarity vector 280 for a given input image 210 represents the distance between the most similar pixel vector qnk within the input image 210 and the kth prototype vector pk (see Figure 25). Each component Sk indicates the degree to which the image feature represented by the kth prototype vector pk exists in the input image. The most similar pixel vector qnk is the pixel vector q in the feature map 230 output from the CNN 220 (feature extraction layer 320) for the input image 210 that is most similar to the kth prototype vector pk.

[0231] For example, if the distance (in this case, the Euclidean distance) D (equation (10) below) between the two vectors s1 and s2 relating to both images G1 and G2 is less than a predetermined value, the image processing device 30 determines that the first image G1 and the second image G2 are not similar. In equation (10), the numbers enclosed in parentheses to the right of each component Sk of the unified similarity vector s(280) are added to identify whether the component relates to the first image or the second image.

[0232]

number

[0233] Next, the controller 31 compares both vectors s1 and s2 for each component Sk. More specifically, the controller 31 finds the difference vector Δs (=Δs12=s1-s2) of both vectors 283 (see Figure 25, middle right), and sorts the multiple (Nc) components ΔSk (k=1,...,Nc) (specifically their absolute values) of the difference vector Δs in descending order (see Figure 26).

[0234] A small k-th component ΔSk ​​(absolute value) of the difference vector Δs means that the k-th component Sk of one vector (e.g., vector s1) and the k-th component Sk of the other vector (e.g., vector s2) are close in value. Conversely, a large k-th component ΔSk ​​(absolute value) of the difference vector Δs means, for example, that the k-th component Sk of one vector s1 is large and the k-th component Sk of the other vector s2 is small. In other words, the prototype PTk (image feature concept) corresponding to the k-th component Sk is largely present in one image (e.g., image 1 G1), while the prototype PTk is not (very) present in the other image (image 2 G2). That is, the image feature of the prototype vector p of the k-th component Sk is (sufficiently) present in one image (e.g., image 1), while the image feature of the prototype vector p of the k-th component Sk is (almost) absent in the other image (image 2). Therefore, the larger the kth component ΔSk ​​of the difference vector Δs, the better that its kth prototype PTk is able to explain the differences between the two images G1 and G2.

[0235] Here, after the multiple (Nc) components ΔSk (absolute values) of the difference vector Δs are sorted in descending order, the image processing device 30 determines that the prototype PTk corresponding to the largest ΔSk ​​(ΔS2 in Figure 26) is the prototype PT (top-ranking prototype PT) that best explains the difference between the two images G1 and G2 (first-priority). The image processing device 30 also determines that the prototype PTk corresponding to the second largest ΔSk ​​(ΔS3 in Figure 26) is the prototype PT (top-ranking prototype PT) that best explains the difference between the two images G1 and G2 (second-priority). Similarly, each prototype PTk is ranked according to the rank of its corresponding ΔSk (absolute value).

[0236] Figure 27 shows an example of how explanatory information explaining the differences is displayed (an example of a display screen). For example, the entirety of Figure 27 is displayed on the display unit 35b.

[0237] In FIG. 27, both images G1 and G2 are displayed on the left side, and each component ΔSk (absolute value) of the difference vector Δs after rearrangement is displayed in a graph format. Note that the values of each component ΔSk etc. may also be displayed together.

[0238] Also, image features etc. of the prototype vectors pk corresponding to the top several (here, 3) of the plurality of components ΔSk are displayed (refer to the rightmost column in FIG. 27). Specifically, together with the image gk including the most similar pixel vector qmk replacing the prototype vector pk (the image having the features of the prototype vector pk most), the partial image corresponding to the most similar pixel vector qmk is shown surrounded by a rectangle.

[0239] For example, for the topmost component ΔS2, the image g2 and the partial image (the region surrounded by a rectangle) within the image g2 are shown. Similarly, for the second-ranked component ΔS3, the image g3 and the partial image within the image g3 are shown, and for the third-ranked component ΔS7, the image g7 and the partial image within the image g7 are shown.

[0240] Also, FIG. 28 is a display example showing further detailed information. It is preferable that not only the display screen of FIG. 27 but also the display screen of FIG. 28 be displayed.

[0241] The total of nine heatmaps in three vertical and three horizontal directions in FIG. 28 show the ignition positions in each most similar image (refer to the leftmost column), the ignition position in the first image G1 (refer to the central column), and the ignition position in the second image G2 (refer to the rightmost column) with respect to these top three prototype vectors p2 (refer to the uppermost row), p3 (refer to the middle row), and p7 (refer to the lowermost row). Each heatmap (three heatmaps arranged horizontally) corresponding to the prototype vector pk of each row corresponds to the k-th plane similarity map 260 (FIG. 3) for the images of each column. Note that the ignition position (the position (w, h) of the pixel vector q with a high degree of similarity to each prototype vector pk) within each image is shown in a relatively bright color. However, the scale is different for each heatmap (specifically, it is scale-converted so that the maximum similarity within each heatmap becomes 100%). Therefore, caution is required when comparing the brightness between the nine heatmaps (or it is preferable not to compare the brightness between the nine heatmaps).

[0242] In the uppermost row and leftmost column of FIG. 28, it is shown that the ignition position in the most similar image g2 with respect to the prototype vector p2 exists near the bird's feet. That is, it can be understood that the prototype vector p2 represents the image features near the bird's feet in the most similar image g2. Also, in the central column of the uppermost row (in the horizontal direction), it is shown that the ignition position in the first image G1 of the prototype vector p2 exists near the bird's feet. In the rightmost column of the uppermost row, it is shown that the ignition position in the second image G2 of the prototype vector p2 exists near the bird's feet. However, the similarity at the ignition position in the first image G1 is high (for example, "0.489"), and the similarity at the ignition position in the second image G2 is low (for example, "0.201"). That is, the image features of the prototype vector p2 appear relatively large in the first image G1 (the upper side of the leftmost column in FIG. 27), while they do not appear much in the second image G2 (the lower side of the leftmost column in FIG. 27).

[0243] In the leftmost column of the middle section of Figure 28, it is shown that the firing position of the prototype vector p3 in the most similar image g3 is located in the background (above the bird's head). In other words, it can be seen that the prototype vector p3 represents some of the image features of the background in the most similar image g3. Also, in the middle column (left-right direction), the firing positions in the first image G1 and the second image G2 are shown as bright regions. However, the similarity to the firing position in the first image G1 is high (e.g., "0.406"), while the similarity to the firing position in the second image G2 is low (e.g., "0.168"). In other words, the image features of the prototype vector p3 are relatively prominent in the first image G1, but not so prominent in the second image G2.

[0244] In the bottom left column of Figure 28, it is shown that the firing position of the prototype vector p7 in the most similar image g7 is located near the tip of the bird's beak. In other words, it can be seen that the prototype vector p7 represents the image features (such as the elongated shape) near the tip of the bird's beak in the most similar image g7. In the bottom (left-right) center column, the firing positions in the first image G1 and the second image G2 are shown as bright regions. However, the similarity to the firing position in the first image G1 is low (e.g., "0.147"), while the similarity to the firing position in the second image G2 is high (e.g., "0.361"). In other words, the image features of the prototype vector p7 are relatively prominent in the second image G2, while they are not very prominent in the first image G1.

[0245] This type of explanatory information is presented to the user by the image processing device 30.

[0246] In particular, the image processing device 30 presents the user with the top few (in this case, three) prototype vectors pk (concepts) as the basis for its judgment that the two images G1 and G2 are not similar to each other. Upon receiving such a presentation, the user can understand that the two images were judged to be dissimilar based on the concepts expressed by these top few prototype vectors pk, for example, the concept of the highest-level prototype vector p2 (image features near the bird's foot in the most similar image g2).

[0247] Here, we assume that the image being compared to query image 215 (the image to be inferred), which belongs to the unclassified class, is one of the multiple training images. However, this is not limited to this, and the image being compared may be an image other than the multiple training images. Furthermore, the image being compared may also be an image of the unclassified class. Thus, when comparing two images, each of the two images may be one of the multiple training images, or an image other than the multiple training images. Furthermore, the two images may both be images of the unclassified class.

[0248] <1-8. Effects of the Embodiment> In the above embodiment, during the learning phase of the learning model 400, distance learning with respect to the integrated similarity vector 280 proceeds by minimizing the evaluation term Ltask. As a result, in the feature space of the integrated similarity vector 280, multiple feature vectors corresponding to multiple input images of the same class (for example, the same type of bird) are placed close to each other. On the other hand, multiple feature vectors corresponding to multiple input images of different classes (different types of birds) are placed far apart from each other.

[0249] Furthermore, by minimizing the evaluation term Lclst, each prototype vector pk is learned to approach the most similar pixel vector q, etc. Therefore, each prototype vector pk is learned to reflect the image features of a specific subregion of a particular image. In other words, each trained prototype vector pk is learned to represent the concept (the concept of the image feature) of each prototype PTk.

[0250] In particular, each prototype vector pk is learned to approach one of the pixel vectors in the feature map corresponding to each image (the image-specific most similar pixel vector q within each image), depending on the distribution prototype belonging degree Tik of each image. Therefore, each prototype is learned to represent features (pixel vectors) that are similar to the image features (pixel vectors) of a specific region of a particular image. Consequently, it is possible to improve the explainability (especially transparency (the ability to explain with concepts that are understandable to humans)) of the learning results in the learning model.

[0251] Furthermore, each prototype PT can belong to two or more classes. Therefore, each prototype vector p can be learned to reflect similar image features between different images of different classes. Also, since it is not necessary to prepare a predetermined number of dedicated prototype vectors for each class, it is possible to construct prototype vectors p efficiently.

[0252] For example, consider a case where a prototype vector pk belongs to both a first class and a second class. In this case, the prototype vector pk is learned to reflect image features similar to both the pixel vector q1 corresponding to the first image feature of an image in the first class and the pixel vector q2 corresponding to the second image feature of an image in the second class. Furthermore, through a synergistic effect with distance learning on the integrated similarity vector 280, the integrated similarity vector 280 corresponding to an image with image features similar to pixel vector q1 and the integrated similarity vector 280 corresponding to an image with image features similar to pixel vector q2 can be placed in close proximity in the feature space.

[0253] Furthermore, each prototype vector p is learned to approach different images within the same class according to the image-specific prototype belonging degree (which differs for each image). More specifically, each prototype vector p is learned to reflect image features according to the image-specific prototype belonging degree (which differs for each image) even within the same class. Therefore, it is possible to construct prototype vectors p efficiently.

[0254] Furthermore, this type of learning process makes it possible to construct (generate) a learning model that can be used in image search processing, where images other than those of known classes (images of unclassified classes) are used as the inference target images, and similar images similar to those inference target images are searched for from among multiple images.

[0255] Furthermore, in step S22 (Figure 9) above, if a single prototype belongs to two or more classes, the degree to which that single prototype belongs to each of those two or more classes is reduced. This makes it possible to reduce the importance of a prototype that belongs to multiple (especially many) classes, considering that it is highly likely to be background information or similar.

[0256] Furthermore, in step S28 (Figure 10) described above, the trained model 420 is modified by replacing each prototype vector with a pixel vector corresponding to a specific region within a specific image. In the modified trained model 420, the similarity between the replaced prototype vector pk (i.e., its most similar pixel vector qmk) and the image features of any subregion within each input image is directly represented in the integrated similarity vector 280. Specifically, a similarity map 270 and an integrated similarity vector 280 are formed, showing the similarity between each pixel vector q of the input image and the replaced prototype vector pk.

[0257] Therefore, the features of the input image can be explained more appropriately using the similarity between the prototype vector pk and the most similar pixel vector qmk (image features of a specific region of a specific image among multiple training images). For example, as shown in Figure 23, the image processing device 30 can explain its similarity to the image features of the most similar pixel vector qmax (image features of a specific subregion R1 of IMG1: "blue head," etc. (see Figure 23)) as the basis for its similarity judgment. In other words, it is possible to accurately obtain very high transparency (the property of being able to explain the inference results with concepts that are understandable to humans).

[0258] <2. Second Embodiment> The second embodiment is a modification of the first embodiment. The following description will focus on the differences from the first embodiment.

[0259] In the first embodiment described above, the basis for determining that the query image 215 is "not similar" to a given image is also explained. In other words, the basis for determining that a given pair of images is not similar is explained.

[0260] FIG. 29 is a graph showing an evaluation result of evaluating to what extent the difference between a certain image pair (images G1, G2) can be explained by using a predetermined number (for example, 20) of the top prototype vectors p out of Nc prototype vectors p. However, in FIG. 29 and the like, image pairs different from the image pairs in FIGS. 25 to 26 are exemplified.

[0261] Here, the distance Dn between both images expressed by the components of the top n prototypes PT out of Nc prototype vectors p is expressed by the following equation (11). It can also be expressed that after the plurality of components ΔSk (absolute values) of the difference vector Δs (= s1 - s2) are sorted in descending order, the value Dn is the magnitude of the partial difference vector reconstructed only by the top n components ΔSk.

[0262]

Equation

[0263] And, the ratio Dr occupied by the distance Dn (Equation (10)) expressed by the components of the top n prototypes PT with respect to the total distance D (Equation (10)) of the image pair is represented by the following equation (12).

[0264]

Equation

[0265] This value (distance ratio) Dr is an evaluation value for evaluating to what extent the difference between a certain image pair (images G1, G2) can be explained by using the top n prototype vectors p (concepts). According to this value Dr, it is possible to evaluate to what extent the distance Dn has been reached by the top n prototype vectors p with respect to the distance D (100%) between the feature vectors of two images i and j.

[0266] The graph in Figure 29 shows the value Dr for two images G1 and G2 (the image pair in Figure 29) of different classes. This value Dr is calculated based on two integrated similarity vectors 280(s1,s2) obtained using the trained model 420 obtained by the training process of the first embodiment described above. In the graph, the horizontal axis represents the value n (number of prototypes considered), and the vertical axis represents the value Dr.

[0267] The graph in Figure 29 shows the values ​​of Dr for each of the top 1 to top 20 prototype vectors p before improvement. Specifically, it shows the value of Dr when using the top 1 prototype vector p, the value of Dr when using the top 2 prototype vectors p, (...omitted...), and the value of Dr when using the top 20 prototype vectors p.

[0268] In the graph in Figure 29, the top-ranked prototype achieves approximately 10% of the Dr values, while the top two prototypes achieve approximately 13%. Furthermore, the Dr values ​​corresponding to the top 20 prototypes are less than 40%. This means that even using the top 20 concepts, 40% of the overall differences (total distance D) cannot be explained. In other words, the generated prototypes may not necessarily be appropriate prototypes. Therefore, there is room for improvement in the "clarity" aspect of the explainability (interpretability) of the basis for judgment (the property of being able to explain the basis for judgment with a small number of concepts).

[0269] Therefore, this second embodiment provides a technology that can improve the "clarity" of the basis for judgment.

[0270] In the second embodiment, the evaluation function L of equation (13) is used instead of the evaluation function L of equation (2).

[0271]

number

[0272] The evaluation function L in equation (13) is an evaluation function to which the new evaluation term Lint has also been added.

[0273] The following explains this evaluation item, Lint.

[0274] Figure 30 shows the graph in Figure 29 with the horizontal axis extended to include all (Nc) prototype PTs. The upper graph in Figure 30 shows a state with poor "clarity," while the lower graph in Figure 30 shows a state with improved "clarity" compared to the upper graph. As shown in Figure 30, clarity can be improved by raising the graph upwards, or more specifically, by minimizing the area of ​​the shaded portion of the graph.

[0275] Therefore, the value Lia in equation (14) is defined, and learning proceeds to minimize this value Lia.

[0276]

number

[0277] The value Lia in equation (14) (when Nd = Nc) corresponds to the area of ​​the shaded region (hatched area) in the graph of Figure 30. As shown in Figure 31, here the area of ​​the shaded region is approximated by the area of ​​a collection of band-shaped regions. The area of ​​each band-shaped region is the value obtained by multiplying the vertical length (1 - Dn / D) by the width (for example, "1"). The value Lia in equation (14) is obtained by adding the length (1 - Dn / D) while varying the value n from 1 to Nd, and then dividing the result by the value Nd. Here, the value Nd is the value Nc (total number of prototype PTs). This value Lia is determined for each image pair.

[0278] Then, the evaluation term Lint is obtained by calculating this value Lia for all image pairs related to the multiple images used for training, adding them together, and dividing the result by the number of pairs Np (see equation (15)).

[0279]

number

[0280] The learning model 400 (prototype vector p, etc.) is optimized by learning to minimize the evaluation function that includes such an evaluation term Lint. In other words, the learning model 400 is optimized to minimize the evaluation value Lint in equation (15) (and the evaluation value Lia in equation (14)).

[0281] Here, minimizing (optimizing) the evaluation value Lint (and the evaluation value Lia in equation (14)) in equation (15) is equivalent to maximizing the normalized value (see equation (14)) obtained by dividing the sum of multiple magnitudes Dn (see equation (11)) corresponding to multiple values ​​n by the inter-vector distance D between the two vectors s1 and s2. The multiple magnitudes Dn are values ​​obtained by calculating the distance Dn (see equation (11)) for each of the multiple values ​​n (n=1,...,Nd; where the value Nd is a predetermined integer less than or equal to the number of dimensions Nc of the unified similarity vector). Furthermore, the distance Dn is the magnitude of the partial difference vector reconstructed using only the top n components ΔSk (absolute value) of the difference vector Δs (=s1-s2) after sorting the multiple components ΔSk in descending order.

[0282] Next, we will explain this process in more detail, referring to the flowchart in Figure 32. Figure 32 is a diagram showing the learning process related to the evaluation term Lint.

[0283] The process shown in Figure 32 (step S60) is executed during (or before) the process in step S25 (Figure 9). In step S25, the evaluation function Lint calculated in step S60 is added to the evaluation function L, and machine learning is performed in step S26, etc.

[0284] Therefore, in step S61, the controller 31 first focuses on one of the image pairs among all the image pairs related to the multiple images used for training.

[0285] Then, the controller 31 calculates the unified similarity vector 280 for each image in the image pair (also referred to as the pair of images of interest) (step S62). Specifically, the controller 31 calculates both the unified similarity vector 280 (first vector s1) obtained by inputting the first image G1 into the learning model 400, and the unified similarity vector 280 (second vector s2) obtained by inputting the second image G2 into the learning model 400.

[0286] Next, the controller 31 sorts the absolute values ​​of the multiple components ΔSk in the difference vector Δs between the first vector s1 and the second vector s2 in descending order (step S63). Note that the absolute value of each component ΔSk ​​can also be expressed as the absolute value of the difference obtained by differentiating both vectors s1 and s2 for each component (prototype component) (the magnitude of the difference for each prototype component).

[0287] Furthermore, the controller 31 determines the magnitude Dn of the partial difference vector, which is reconstructed using only the top n components of the difference vector Δs after sorting in descending order, for each of several values ​​n (n=1,...,Nd) (step S64). Here, the value Nd is a predetermined integer less than or equal to the number of dimensions Nc of the integrated similarity vector. In this case, the value Nd is set to a value equal to Nc.

[0288] Then, the controller 31 calculates the value Lia according to equation (14) (step S65).

[0289] Furthermore, in step S65, the controller 31 repeatedly executes steps S61 to S64 while changing the pair of images of interest, thereby obtaining the value Lia for all image pairs. In other words, the controller 31 performs the processing of steps S61 to S64 with respect to the first and second images relating to any combination of the multiple images used for training. Then, the controller 31 calculates the value Lint according to equation (15) (step S65).

[0290] Subsequently, in step S25, the evaluation function L, which also includes the evaluation term Lint, is calculated according to equation (13).

[0291] Then, steps S21 to S25 and steps S61 to S65 are repeatedly executed, and the learning process (machine learning) of the learning model 400 is carried out to minimize (optimize) the evaluation function L (step S26). Specifically, the learning process is carried out so that the evaluation term Lia in equation (14) and the evaluation term Lint in equation (15) are minimized. Note that minimizing the evaluation term Lia in equation (14) is equivalent to maximizing the normalized value (Dn / D) obtained by dividing the sum of multiple magnitudes Dn corresponding to multiple values ​​n by the vector distance D between the two vectors.

[0292] Figure 33 shows an example of improvement regarding the image pair in Figure 29. Specifically, it shows how the clarity of the inference basis (the reason why they are not similar) is improved when inference processing is performed on the image to be inferred (in detail, the image pair in Figure 29) based on the trained model 420 obtained by the training process according to the second embodiment.

[0293] The upper part of Figure 33 shows the values ​​of Dr for each of the top 1 to top 20 prototype vectors p before improvement. As mentioned above, the upper graph shows that even using the top 20 prototype vectors p (concepts) only explains slightly less than 40% of the overall differences (total distance D).

[0294] On the other hand, the lower part of Figure 33 shows the values ​​of Dr for each of the top 1 to top 20 prototype vectors p after improvement (when inference is performed using the trained model 420 which has been trained with the evaluation term Lint as described above). In the lower graph, by using the top 20 prototype vectors p (concepts), approximately 60% of the overall differences (total distance D) can be explained. The degree of explanation (value Dr) has improved by about 20% compared to before the improvement.

[0295] Furthermore, after the improvements, the top-performing prototype achieved approximately 22% of the Dr value, and the top two prototypes achieved approximately 28%. Additionally, the top seven prototypes achieved approximately 40% of the Dr value. Thus, the rationale for the decisions can be explained with fewer concepts, demonstrating improved "clarity."

[0296] Thus, according to the second embodiment, it is possible to improve clarity (the ability to explain the basis for a decision with fewer concepts) within explainability (interpretability).

[0297] Furthermore, although not shown in the diagram, the improved top-level prototype vector has changed to a different prototype vector than the pre-improvement top-level prototype vector. In other words, the top-level prototype vector has changed to a prototype vector that better explains the difference between the two images G1 and G2.

[0298] In the above embodiment, the value Nd in equation (14) is set to the value Nc, but the invention is not limited to this, and the value Nd may be smaller than the value Nc. When the value Nd is smaller than the value Nc, the value Lia in equation (14) corresponds to the area to the left of n=Nd in the shaded region of the graph in Figure 31. Even in this case, it is possible to obtain a certain degree of effect in minimizing the area of ​​the shaded region. However, in this case, the area of ​​the shaded region to the right of the value Nd is not necessarily minimized. Therefore, it is preferable that the value Nd is the value Nc.

[0299] <Modified form of the second embodiment> Incidentally, in the learning process that minimizes the evaluation term Ltask (see equation (3)) related to distance learning, a force (repulsion) acts that expands the distance D between negative pairs in the feature space (the distance between the unified similarity vectors 280 related to the negative pair) (see the double arrow in Figure 34). The value obtained by partially differentiating the evaluation term Ldist (described later) with respect to distance D (δLdist / δD) can also be expressed as the repulsive force acting between negative pairs (anchor-negative) due to the evaluation term Ltask. Here, the value Ldist is the term relating to the anchor-negative distance dan in equation (3) (see equation (20) (described later)). The absolute value of the value (δLdist / δD) is "1".

[0300] On the other hand, it can be derived from equation (14) etc. that the value obtained by partially differentiating the clarity evaluation term Lint (equation (15)) with respect to distance D (δLint / δD) is always positive. Optimization using the evaluation term Lint has the effect of reducing the distance D between the 280 integrated similarity vectors in the feature space. This value (δLint / δD) can also be expressed as the attractive force due to the evaluation term Lint (see the inward arrow in Figure 34). Furthermore, it can be derived from equation (14) etc. that the value obtained by partially differentiating the evaluation term Lint in equation (15) with respect to distance D (δLint / δD) is inversely proportional to distance D (see equation (16)). In other words, as distance D decreases, the attractive force due to the evaluation term Lint increases.

[0301]

number

[0302] To improve clarity, the evaluation term Lint needs to be reduced. Reducing the evaluation term Lint reduces the distance D, and the attractive force due to the evaluation term Lint increases.

[0303] Therefore, if we adopt the evaluation term Lia in equation (14) (and the evaluation term Lint in equation (15)) as is, the attractive force of the evaluation term Lint may outweigh the repulsive force of the evaluation term Ltask, potentially causing negative pairs to move closer to each other (moving in a way that reduces the distance D between them). In other words, even though we want to increase the distance D between negative pairs in distance learning, if the attractive force of the evaluation term Lint is too strong, the distance D between negative pairs may actually decrease. That is, the accuracy of distance learning may deteriorate.

[0304] Therefore, in this modified example, the evaluation term Lint is changed as shown in equation (17). Equation (17) is obtained by replacing the value Lia in equation (15) with a new value Lia (hereinafter also referred to as value Lib).

[0305]

number

[0306] However, the value Lib is expressed by the following equation (18). The value Lib is the value obtained by multiplying the value Lia from equation (14) by the coefficient w (where w is less than or equal to 1). The value Lib can also be expressed as the value after the original value Lia has been modified (reduced, etc.) using the coefficient w (the modified value Lia).

[0307]

number

[0308] Furthermore, the coefficient w is expressed by the following equation (19).

[0309]

number

[0310] The coefficient w is determined for each pair (similar to the value Lia before modification). The coefficient w in equation (19) corresponds to the attractive force due to evaluation term Lint divided by the repulsive force due to evaluation term Ltask. The attractive force due to evaluation term Lint is expressed as the partial derivative of the evaluation term Lia before modification with respect to distance D (δLia / δD), and the repulsive force due to evaluation term Ltask is expressed as the partial derivative of the evaluation term Ltask (specifically, the evaluation term Ldist for negative pairs) with respect to distance D (δLdist / δD).

[0311] In other words, the coefficient w is the value obtained by dividing the repulsive force (δLdist / δD) by the excessively large attractive force (δLia / δD). If the attractive force (δLia / δD) is greater than the repulsive force (δLdist / δD), the value of Lia is adjusted by the value w (a value less than 1) to reduce it, and the corrected value Lia(Lib) is calculated. In other words, the value of Lia is adjusted so that the corrected attractive force (δLib / δD) does not exceed the repulsive force (δLdist / δD).

[0312] The evaluation term Ldist is expressed by equation (20).

[0313]

number

[0314] Here, the evaluation term Ltask (Equation (3)) is expressed as the sum (more specifically, the average of the sums) obtained by adding up pairwise evaluation terms Lta (described below), which are obtained for each image pair based on evaluation terms related to distance learning based on multiple integrated similarity vectors 280, for multiple image pairs. The pairwise evaluation term Lta is the value inside the summation symbol Σ in Equation (3) (the value to be added).

[0315] Furthermore, the evaluation term Lint (Equation (15)) before the correction is expressed as the sum of the pairwise evaluation term Lia added up over multiple pairs (Np pairs) of image pairs (more specifically, the average of these sums). The pairwise evaluation term Lia (Equation (14)) is the value obtained for each image pair as an evaluation term relating to the sum of multiple sizes Dn (more specifically, the normalized value obtained by dividing this sum by the inter-vector distance D between the two vectors).

[0316] The coefficient w is a value used to adjust the pairwise evaluation term Lia so that it does not become too large. Specifically, the magnitude of the pairwise evaluation term Lia is adjusted by the coefficient w so that the absolute value of the partial derivative of the pairwise evaluation term Lia with respect to the inter-vector distance D with respect to each image pair (corresponding image pair) (attraction based on the pairwise evaluation term Lia) does not exceed the absolute value of the partial derivative of the pairwise evaluation term Lta with respect to the inter-vector distance D with respect to each image pair (especially negative pairs) (repulsion based on the pairwise evaluation term Lta). In other words, the pairwise evaluation term Lia is changed to a new pairwise evaluation term Lia (i.e., pairwise evaluation term Lib).

[0317] The repulsive force (δLdist / δD) corresponds to "the absolute value of the partial derivative of the pairwise evaluation term Lta with respect to the inter-vector distance D with respect to each image pair (negative pair) (repulsive force based on the pairwise evaluation term Lta)."

[0318] Then, the adjusted value Lia (value Lib) is calculated for all image pairs related to the multiple images used for training, and the sum of these values ​​is divided by the number of pairs Np to obtain the evaluation term Lint (see equation (17)).

[0319] The learning model 400 (including the prototype vector p) is optimized by training it to minimize the evaluation function which includes such an evaluation term, Lint.

[0320] According to this, it is possible to suppress the decrease (degradation) in the accuracy of the learning model 400, and consequently inference accuracy.

[0321] <3. Third Embodiment: No Replacement> The third embodiment is a modification of the first and second embodiments. The following will focus on explaining the differences from the first embodiment and the like.

[0322] In each of the above embodiments, the search process (inference process) is performed after the machine learning of the trained model 400 is completed and each prototype vector pk has been replaced with the most similar pixel vector qmk. That is, each prototype vector pk in the trained model 420 has been replaced with the most similar pixel vector qmk (see step S28 (Figure 10) and Figure 20).

[0323] As described above, according to the explanatory information in the first embodiment (see Figure 23, etc.), the similarity between the images being compared can be explained very appropriately using the similarity with respect to the replaced prototype vector pmax (i.e., the most similar pixel vector qmax). For example, the image processing device 30 can explain that it has determined similarity based on whether or not it is similar to the image features of the most similar pixel vector qmax (image features of a specific subregion R1 of IMG1: "blue head," etc. (see Figure 23)). In other words, it is possible to accurately obtain very high transparency (the property of being able to explain the inference results in a concept that is understandable to humans).

[0324] Furthermore, the top few prototype vectors pk can be used as information to explain the differences between two images. In both the first and second embodiments, all prototype vectors pk, including the top few, are replaced by their corresponding most similar pixel vectors q. Therefore, in the inference process using the trained model 420 after the replacement, inference is performed based on the similarity with each replaced most similar pixel vector q. Thus, the differences between two images can be explained based on the similarity (difference) between the image features of the most similar pixel vector q. In other words, it is possible to accurately obtain very high transparency (the property of being able to explain the inference results with concepts that are understandable to humans).

[0325] On the other hand, in this third embodiment, the process of replacing the prototype vector with the most similar pixel vector q is not performed. After the machine learning of the learning model 400 is completed, the search process (inference process) is executed without each prototype vector pk being replaced with the most similar pixel vector.

[0326] In this case, the prototype vector pmax in Judgment Basis 1 (see Figure 23, etc.) does not perfectly correspond to the image of a specific subregion of a particular image. Although the prototype vector pmax is trained to approximate a certain image feature, it is not replaced by the most similar pixel vector qmax. In other words, the inference process itself is performed using the prototype vector p before it is replaced by the most similar pixel vector q, not the prototype vector p that has been replaced by the most similar pixel vector q. That is, complete agreement between the prototype vector pmax and the most similar pixel vector qmax cannot be guaranteed. Therefore, it is difficult to obtain very high transparency.

[0327] However, even in this embodiment (the embodiment of the third embodiment), it is possible to obtain a certain degree of transparency.

[0328] For example, in the process of searching for similar images to the query image 215, the image processing device 30 can indicate that "the basis for the similarity judgment is that the image is similar to the features of the prototype vector pmax," and "the image features of the top n pixel vectors (for example, the top two q1 and q2) that are similar to the prototype vector pmax."

[0329] For example, the image processing device 30 can display the image features corresponding to pixel vector q1 ("orange foot") and the image features corresponding to pixel vector q2 ("red foot") in the "Judgment Basis 1" column on the display screen in Figure 24, as "the image features of the top n pixel vectors similar to the prototype vector pmax". In this case, the user can understand that the basis for the similarity judgment by the image processing device 30 is that "it is similar to the prototype vector that reflects these two types of image features (simply put, it is similar to the common image feature of pixel vectors q1 and q2 ("a bright reddish foot"))".

[0330] Therefore, it is possible to obtain a certain degree of "transparency" (the property of being able to explain the inference results in a concept that is understandable to humans).

[0331] Alternatively, the image processing device 30 can present explanatory information similar to that in Figures 27 and 28 when explaining why two images are not similar. That is, it can show that "the top few (e.g., three) prototype vectors p are the basis for the dissimilarity," and "for each of the top few prototype vectors p, the image features of the top n (e.g., only the topmost) similar pixel vectors." However, in the third embodiment, the prototype vectors p are not replaced by the most similar pixel vectors q, and the similarity between each image and the prototype vectors p themselves is not fully reflected in the unified similarity vector 280. Therefore, it is not possible to obtain the same high level of transparency as in the first embodiment, etc. Specifically, it is not possible to explain that the image features of a particular pixel vector q that corresponds one-to-one with a certain prototype vector p are the basis for the dissimilarity. However, even in this case, it is possible to obtain a certain degree of "transparency" (the property of being able to explain the inference results with a concept that is understandable to humans) by showing "for each of the top few prototype vectors p, the image features of the top n (e.g., only the topmost) similar pixel vectors."

[0332] <4. Variations, etc.> The embodiments of this invention have been described above, but this invention is not limited to those described above.

[0333] For example, in each of the embodiments described above, the image most similar to the estimated target image is searched from among multiple images of multiple types of objects (e.g., multiple types of birds), but this is not limited to this. For example, the image most similar to the estimated target image may be searched from among multiple images of multiple objects (e.g., multiple people). More specifically, an image containing a person (which may be the same person) that most closely resembles a certain person (such as a lost child or a criminal (suspect)) in the estimated target image may be searched from among the multiple images as the image most similar to the estimated target image. In other words, the same class may consist of "objects of the same type" or "identical objects".

[0334] Furthermore, in each of the above embodiments, an image retrieval process is exemplified in which, with respect to an inference target image (such as a certain photographed image) that may include images other than those of known classes (images of unclassified classes), similar images are searched for from a plurality of images (known images (such as photographed images for training)). However, the present invention is not limited thereto. For example, the concept of the present invention may be applied to a classification problem (a classification process that classifies a certain estimated target image into one of the known classes) that infers which of the known classes an inference target image belonging to one of the known classes belongs to.

[0335] Specifically, in the same manner as in each of the embodiments described above, the controller 31 obtains a unified similarity vector 285 for the image to be inferred and a unified similarity vector 283 for multiple training images. Then, based on the positional relationship of these vectors in the feature space, methods such as the k-nearest neighbor method may be used. More specifically, a classification process is performed to estimate the class to which the image to be inferred belongs, based on the top few (k) images extracted in descending order of similarity (in order of proximity to the image to be inferred). In detail, it is estimated that the class to which the image to be inferred belongs is the class that is most frequently assigned to the top k images (training images) (for example, k=1, 3, 5, etc.).

[0336] In this modified example, it is possible to efficiently represent the features of multiple images with fewer concepts (prototypes) compared to conventional techniques that perform classification using ProtoPNet. Conventional techniques that perform classification using ProtoPNet require a predetermined number of dedicated prototypes for each class, thus requiring a very large number of prototypes (= predetermined number × number of classes). For example, 2000 prototypes (= 10 × 200 classes) are required. In contrast, in this embodiment, it is not necessary to prepare dedicated prototypes for each class; it is sufficient to prepare a prototype common to multiple classes, thus reducing the number of prototypes. For example, it is possible to achieve the same (or better) inference accuracy with around 512 prototypes.

[0337] Furthermore, even in image search processing (processing other than class classification) as in the embodiments described above, it is not necessary to prepare a predetermined number of dedicated prototypes for each class. Therefore, it is possible to construct prototype vectors p efficiently (with a relatively small number of prototype vectors p).

[0338] Thus, the present invention may be applied to processes other than image retrieval processing for images similar to the image to be inferred (particularly inference processing using distance learning). For example, the learning process based on multiple integrated similarity vectors 280 as described above may also be classification learning using distance information (classification learning using KNN nearest neighbor method, etc.). Furthermore, the above concept may be applied to biometric authentication or anomaly detection processing using distance information. [Explanation of symbols]

[0339] 30 Image Processing Device 210 Each input image 220 Convolutional Neural Networks (CNNs) 230 Feature Map 240 pixel vector 250,p,pk prototype vector 260 Planar Similarity Maps 270 Similarity Maps 280,283,285,s Integrated Similarity Vector 290 sub-feature vectors 400,420 Learning Models q pixel vector The k-th component of the Sk unified similarity vector

Claims

1. An image processing device, A control unit that performs machine learning on a learning model comprising a convolutional neural network. Equipped with, The aforementioned learning model, A process to generate a feature map obtained from a predetermined layer in the convolutional neural network in response to an input image, which shows the feature quantities for each subregion in the input image for multiple channels. A process to generate multiple prototype vectors, which are sequence parameters learned as prototypes representing candidate image feature concepts composed of the aforementioned multiple channels, A process to generate an integrated similarity vector that shows the similarity between the input image and each prototype for multiple prototypes, based on the similarity between each pixel vector, which is a vector representing the image features across multiple channels at each planar position of each pixel in the feature map, and a single prototype vector, It is a machine learning model for making a computer perform the task, The control unit, in the learning phase of the learning model based on multiple images for learning, The belonging prototype, which is a prototype belonging to a class, and the prototype belonging degree, which indicates the degree to which the belonging prototype belongs to the class, are determined for each of the multiple classes labeled on the multiple images used for training, The prototype affiliation degree of each class is distributed to each of two or more images within the same class based on predetermined criteria, and the distributed prototype affiliation degree, which is the prototype affiliation degree for each image, is calculated. An image processing apparatus characterized in that, when performing a learning process based on multiple integrated similarity vectors corresponding to the multiple images, the learning model is machine-trained so that each prototype vector approaches one of the pixel vectors in the feature map corresponding to each image, according to the degree of belonging of each image's distributed prototype.

2. The control unit, A prototype selection process is performed to select a prototype belonging to the first class from among the multiple images used for training, based on a comparison between a predetermined image belonging to the first class and multiple comparison images belonging to classes other than the first class. Based on the number of prototypes selected for each affiliated prototype in the aforementioned prototype selection process, the degree of prototype affiliation of each affiliated prototype for the aforementioned class is determined. The aforementioned prototype selection process is: The process includes a unit selection process in which a difference vector is obtained by subtracting the integrated similarity vector obtained by inputting one of the multiple comparison images into the learning model from the integrated similarity vector obtained by inputting the predetermined image into the learning model, and the prototype corresponding to the largest component among the multiple components in the difference vector is selected as the belonging prototype belonging to the class of the predetermined image. The image processing apparatus according to claim 1, characterized in that it includes a selection count calculation process which selects at least one belonging prototype belonging to the one class and counts the number of selected belonging prototypes by performing the unit selection process on the plurality of comparison images while changing the one comparison image to another comparison image.

3. The image processing apparatus according to claim 2, characterized in that, when a prototype belongs to two or more classes, the control unit reduces the degree to which the prototype belongs in each of the two or more classes.

4. The control unit, in distributing the prototype belonging degree of one belonging prototype belonging to one class to N images belonging to one class and determining the distributed prototype belonging degree for each of the N images, A first distance is calculated, which is the distance between one of the N images and the pixel vector in the feature map corresponding to one of the N images that is most similar to the prototype vector of the prototype to which the one belongs. A second distance is calculated, which is the distance between the pixel vector that is most similar to the prototype vector of the first affiliated prototype among multiple pixel vectors in the feature map corresponding to the other N images. The image processing apparatus according to claim 2 or 3, characterized in that, when the first distance is greater than the second distance, the degree of belonging of the distribution prototype to one image is determined to be a smaller value than the degree of belonging of the distribution prototype to the other image.

5. The image processing apparatus according to claim 1, characterized in that, after the machine learning of the learning model is completed, the control unit modifies the learning model by replacing each prototype vector with the most similar pixel vector, which is the pixel vector that is most similar to each prototype vector among a plurality of pixel vectors in a plurality of feature maps relating to the plurality of images.

6. The evaluation function used in the machine learning of the aforementioned learning model has a first evaluation term which is an evaluation term related to clarity, The control unit, with respect to the first image and the second image relating to any combination of the plurality of learning images, Obtain both a first vector, which is an integrated similarity vector obtained by inputting the first image into the learning model, and a second vector, which is an integrated similarity vector obtained by inputting the second image into the learning model. The absolute values ​​of the multiple components in the difference vector between the first vector and the second vector are sorted in descending order. The magnitude Dn of the partial difference vector, which is reconstructed using only the top n components after sorting the plurality of components of the difference vector in descending order, is determined for each of the plurality of values ​​n (n = 1, ..., Nd; where the value Nd is a predetermined integer less than or equal to the number of dimensions Nc of the unified similarity vector). The image processing apparatus according to claim 1, characterized in that the first evaluation term is optimized to maximize the value obtained by normalizing the sum of the multiple magnitudes Dn corresponding to each of the multiple values ​​n by the inter-vector distance between the two vectors, thereby performing machine learning on the learning model.

7. The aforementioned evaluation function is, A second evaluation term is an evaluation term for distance learning based on the multiple integrated similarity vectors corresponding to the multiple images, It further possesses, The first evaluation term is expressed as the sum of the pair-specific first evaluation terms obtained for each image pair, which are the evaluation terms relating to the sum of the multiple sizes, and these are added together for multiple pairs of images. The second evaluation term is expressed as a sum obtained by adding up pairwise second evaluation terms, which are obtained for each image pair based on the evaluation terms related to distance learning based on the multiple integrated similarity vectors, for multiple sets of image pairs. The image processing apparatus according to claim 6, characterized in that the control unit adjusts the magnitude of the first evaluation term for each pair such that the absolute value of the partial derivative of the first evaluation term for each pair with respect to the inter-vector distance with respect to each image pair does not exceed the absolute value of the partial derivative of the second evaluation term for each pair with respect to the inter-vector distance with respect to each image pair.

8. The image processing apparatus according to claim 1, wherein the control unit searches for an image similar to the input image to be searched from among the plurality of images for training, based on an integrated similarity vector output from the learning model in response to inputting the input image to be searched into the learning model after the machine learning of the learning model has been completed, and a plurality of integrated similarity vectors output from the learning model in response to inputting the plurality of images for training into the learning model.

9. The image processing apparatus according to claim 5, characterized in that, after the machine learning of the learning model is completed and each prototype vector has been replaced with the most similar pixel vector, the control unit searches for an image similar to the input image to be searched from among the plurality of learning images based on an integrated similarity vector output from the learning model in response to inputting the input image to be searched to the learning model, and a plurality of integrated similarity vectors output from the learning model in response to inputting the plurality of learning images to the learning model.

10. A method for producing learning models, The aforementioned learning model, A process to generate a feature map obtained from a predetermined layer in the convolutional neural network within the learning model in response to an input image, which shows the feature quantities for each subregion within the input image for multiple channels. A process to generate multiple prototype vectors, which are sequence parameters learned as prototypes representing candidate image feature concepts composed of the aforementioned multiple channels, A process to generate an integrated similarity vector that shows the similarity between the input image and each prototype for multiple prototypes, based on the similarity between each pixel vector, which is a vector representing the image features across multiple channels at each planar position of each pixel in the feature map, and a single prototype vector, It is a machine learning model for making a computer perform the task, The method for producing the aforementioned learning model is: a) A computer determines, based on multiple images for training, a belonging prototype which is a prototype belonging to a class, and a prototype belonging degree which indicates the degree to which the belonging prototype belongs to the class, for each of the multiple classes labeled on the multiple images for training. b) The computer determines the distributed prototype affiliation, which is the prototype affiliation for each image, by distributing the prototype affiliation of the prototypes belonging to each class to each of two or more images within the same class, based on predetermined criteria. c) Steps of machine learning the learning model such that when the computer performs a learning process based on multiple integrated similarity vectors corresponding to the multiple images, each prototype vector approaches one of the pixel vectors in the feature map corresponding to each image, according to the degree of distribution prototype affiliation of each image, A method for producing a learning model, characterized by comprising the following features.

11. d) After the machine learning of the learning model is completed, the computer processes each of the prototypes Step d) modifies the learning model by replacing each prototype vector with the most similar pixel vector, which is the pixel vector that is most similar to each prototype vector among the multiple pixel vectors in the multiple feature maps relating to the multiple images. A method for producing a learning model according to claim 10, further comprising the above.

12. Step c) above relates to the first image and the second image relating to any combination of the plurality of images for learning, c-1) A step in which the computer obtains both a first vector, which is an integrated similarity vector obtained by inputting the first image into the learning model, and a second vector, which is an integrated similarity vector obtained by inputting the second image into the learning model, c-2) The computer sorts the absolute values ​​of multiple components in the difference vector between the first vector and the second vector in descending order. c-3) A step in which the computer determines the magnitude Dn of a partial difference vector reconstructed using only the top n components after sorting the plurality of components of the difference vector in descending order, for each of a plurality of values ​​n (n = 1, ..., Nd; where the value Nd is a predetermined integer less than or equal to the number of dimensions Nc of the unified similarity vector), c-4) A step in which the computer performs machine learning on the learning model such that the value obtained by normalizing the sum of the multiple magnitudes Dn corresponding to each of the multiple values ​​n by the distance between the two vectors is maximized, A method for producing a learning model according to claim 10, characterized by comprising:

13. An inference method characterized in that a computer performs inference processing on a new input image using a learning model produced by a learning model production method according to any one of claims 10 to 12.

Citation Information

Patent Citations

  • Teacher data generation device, teacher data generation method, teacher data generation program, and object detection system

    JP2018200531A

  • Learning device, learning method, and program

    JP2022041434A

  • Cross-transformer neural network system for few-shot similarity determination and classification

    US20210383226A1