Learning apparatus, inference apparatus, learning method, and inference method
Patent Information
- Application Number
- JP2022175029
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2025-11-14
AI Technical Summary
Existing authentication technologies using image similarity learning methods struggle with low accuracy improvements for pairs significantly above or below the decision threshold, and existing methods focusing on pairs near the threshold provide limited enhancements.
A technique that calculates feature vectors from images, determines positive and negative pair similarities, and uses a loss function with a slope that is large near the threshold and decreases away from it to improve authentication accuracy by focusing on pairs near the threshold.
Enhances authentication accuracy by intensively learning pairs near the threshold while also addressing pairs further away, reducing the likelihood of misjudgments and maintaining overall similarity.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a technique for authentication using an image. [Background technology]
[0002] Conventionally, there is a learning method for image-based authentication technology that uses pair similarity (Patent Document 1). In addition, there is a learning method that emphasizes pairs near a judgment threshold to improve accuracy at a certain level of misjudgment (Non-Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special table 2020-504891 [Non-patent literature]
[0004] [Non-Patent Document 1] J. Liu, H. Qin, Y. Wu, and D. Liang, “AnchorFace: Boosting TAR@FAR for Practical Face Recognition”, AAAI, vol. 36, no. 2, pp. 1711-1719, Jun. 2022. Summary of the Invention [Problem to be solved by the invention]
[0005] However, the technology described in Patent Document 1 has a problem in that it is used for learning regardless of the similarity of the pair. Also, the technology described in Non-Patent Document 1 has a problem in that the learning improvement effect is small for pairs that greatly exceed or are greatly below the threshold. The present invention provides a technology for improving authentication accuracy by focusing on pairs that show similarity near the threshold while also learning other cases. [Means for solving the problem]
[0006] One aspect of the present invention includes a first acquisition means for acquiring one or more pairs of an image and a label; a first calculation means for calculating a feature vector from the image using a feature extractor; a second calculation means for calculating a similarity between the feature vector calculated by the first calculation means from the image acquired by the first acquisition means and a feature vector of the same label as the feature vector calculated by the first calculation means as a positive pair similarity, and a similarity between the feature vector of a different label as a negative pair similarity; a determination means for determining a threshold value for the similarity; a third calculation means for calculating a loss value for a positive pair similarity smaller than the threshold; a fourth calculation means for calculating a loss value for a negative pair similarity larger than the threshold; and a learning means for learning parameters of the feature extractor that reduce the loss value calculated by the third calculation means and the loss value calculated by the fourth calculation means, wherein the third calculation means or the fourth calculation means uses a loss function including a function that generates a function whose absolute value of the gradient is larger in the vicinity of a predetermined threshold and whose absolute value of the gradient is smaller even when the function is away from the threshold. Effect of the Invention
[0007] According to the present invention, it is possible to improve authentication accuracy by focusing on pairs that show similarities near a threshold value while also learning other examples. [Brief description of the drawings]
[0008] [Figure 1] FIG. 4A is a block diagram showing an example of the hardware configuration of a computer device 100a, and FIG. 4B is a block diagram showing an example of the functional configuration of a learning device 100b. [Diagram 2] 10A is a flowchart of a learning loop process, and FIG. 10B is a flowchart showing details of the process in step S204a. [Diagram 3] A diagram to explain positive loss and negative loss. [Figure 4]1A is a block diagram showing an example of the functional configuration of the learning device 400a, FIG. 1B is a flowchart showing details of the processing of step S202b, FIG. 1C is a graph of the loss when r in (Equation 3) is changed depending on the race, and FIG. 1D is a graph of the loss gradient. [Diagram 5] 1A is a block diagram showing an example of the functional configuration of a learning device 500a, FIG. 1B is a flowchart showing details of the process of step S204a, and FIG. 1C is a diagram explaining the process of determining quality based on attributes. [Figure 6] FIG. 6A is a block diagram showing an example of the functional configuration of an inference device 600a, and FIG. 6B is a block diagram showing an example of the functional configuration of an inference device 600b. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.
[0010] [First embodiment] First, a hardware configuration example of a computer device 100a applicable to a learning device or an inference device will be described with reference to the block diagram of Fig. 1(a). Note that the hardware configuration of a computer device applied to a learning device and the hardware configuration of a computer device applied to an inference device may be the same or different. Also, a single computer device may execute the functions of both the learning device and the inference device.
[0011] A CPU (Central Processing Unit) 101a executes various processes using computer programs and data stored in a ROM 102a and a RAM 103a, thereby controlling the operation of the entire computer device 100a and executing or controlling various processes that will be described as processes performed by a learning device and an inference device.
[0012] The ROM (Read Only Memory) 102a stores setting data for the computer device 100a, computer programs and data related to the startup of the computer device 100a, computer programs and data related to the basic operation of the computer device 100a, and the like.
[0013] The RAM (Random Access Memory) 103a has an area for storing computer programs and data loaded from the ROM 102a or the external storage device 104a. The RAM 103a also has an area for storing computer programs and data received from the outside via the communication interface 107a. The RAM 103a also has a work area used when the CPU 101a executes various processes. In this way, the RAM 103a can provide various areas as needed.
[0014] The external storage device 104a is a large-capacity information storage device such as a hard disk drive. The external storage device 104a stores an OS (operating system), computer programs and data for causing the CPU 101a to execute or control various processes described as processes performed by the learning device and the inference device, etc. The computer programs and data stored in the external storage device 104a are loaded into the RAM 103a as appropriate under the control of the CPU 101a, and become targets for processing by the CPU 101a.
[0015] The external storage device 104a may include a memory card, an optical disk such as a flexible disk (FD) or a compact disk (CD) that is detachable from the computer device 100a, a magnetic or optical card, an IC card, a memory card, or the like.
[0016] The input device 109a is connected to the input device interface 105a. The input device 109a is a user interface such as a keyboard, a mouse, a touch panel screen, etc., and can be operated by a user to input various instructions to the CPU 101a.
[0017] The monitor 110a is connected to the output device interface 106a. The monitor 110a has a liquid crystal screen or a touch panel screen, and displays the processing results of the CPU 101a as images, characters, etc. A projection device such as a projector may be provided in addition to or instead of the monitor 110a.
[0018] 1(a), it is connected to a network line 111a such as a LAN or the Internet. A network camera (NW camera) 112a that captures moving images and still images periodically or irregularly is connected to the network line 111a.
[0019] Images captured by the network camera 112a (images of each frame in a moving image, or still images captured periodically or irregularly) are received by the external storage device 104a or RAM 103a via the network line 111a and the communication interface 107a.
[0020] The CPU 101a, the ROM 102a, the RAM 103a, the external storage device 104a, the input device interface 105a, the output device interface 106a, and the communication interface 107a are all connected to a system bus 108a.
[0021] <Learning device> An example of the functional configuration of the learning device 100b according to this embodiment will be described with reference to the block diagram of FIG. 1(b). In the following, the functional units shown in FIG. 1(b) may be described as the main processing units, but in reality, the functions of the functional units are realized by the CPU 101a executing a computer program corresponding to the functional units. One or more of the functional units shown in FIG. 1(b) may be implemented in hardware.
[0022] The learning device according to this embodiment learns a face recognition task of determining whether or not faces captured in two images belong to the same person. Specifically, feature vectors are calculated from each image, and learning is performed so that the similarity between the two feature vectors becomes high when the faces of the same person appear in each image. Note that in this embodiment, the similarity is determined by the cosine similarity (hereinafter, Cos similarity) of vectors. However, Euclidean distance or the like may be used as the similarity, and the definition of the similarity is not limited to a specific form.
[0023] The acquisition unit 101b acquires one or more pairs (learning data) of an image and a label (person label) of a person appearing in the image. Hereinafter, one or more pairs (pairs of an image and a person label) acquired by the acquisition unit 101b are called a mini-batch. Also, a pair of an image and a person label may be called a sample. The image includes a face image of a person. The person label is information indicating a person's ID, and the same label value is used when the image is the same person. The pair of the image and the person label is stored in the external storage device 104a or the like. Note that the person label may have a folder structure, and may be configured such that images of the same person are stored together in a folder, and the same folder indicates that the images are of the same person. The method of assigning the person label is not limited to the above.
[0024] The calculation unit 102b calculates a feature vector used for authentication from the image. A neural network is used as a feature extractor to calculate the feature vector. For example, Convolutional Neural Networks (CNN), which is a type of neural network, is used as the feature extractor. CNN extracts abstracted information from the input image by repeatedly performing a process consisting of a convolution process, an activation process, and a pooling process on the input image. At this time, a processing unit consisting of a convolution process, an activation process, and a pooling process is often called a hierarchy. As the activation process used at this time, there are several known methods, and for example, a method called Rectified Linear Unit (ReLU) may be used. In addition, there are several known methods for the pooling process, and for example, a method called maximum pooling (Max pooling) may be used. For example, ResNet, etc. introduced in non-patent literature (K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. In ECCV, 2016) may be used as the structure of the CNN. In addition, a neural network described in a non-patent document (Alexey Dosovitskiy, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.), known as VisionTransformer (ViT), may be used. Note that the configuration of the neural network is not limited to these. The calculation unit 102b holds information such as the structure and weights of the neural network in the RAM 103a or the like.
[0025] The memory bank unit 103b holds the feature vectors in association with the person labels. The memory bank unit 103b holds m feature vectors for each person label, where m is a predetermined constant. The memory bank unit 103b holds the person labels and the feature vectors in the RAM 103a or the like.
[0026] The calculation unit 104b acquires the feature vector acquired by the calculation unit 102b from the image acquired by the acquisition unit 101b. The calculation unit 104b then obtains a pair (positive pair) similarity between the acquired feature vector and a feature vector with the same person label as the person label of the image among the feature vectors held in the memory bank unit 103b as a "positive pair similarity". The calculation unit 104b also obtains a pair (negative pair) similarity between the acquired feature vector and a feature vector with a different person label from the person label of the image among the feature vectors held in the memory bank unit 103b as a "negative pair similarity".
[0027] The threshold determination unit 105b acquires a threshold used for loss calculation described later. For example, the threshold determination unit 105b may acquire a threshold stored in advance in the RAM 103a. Alternatively, the threshold determination unit 105b may acquire, for example, the similarity of the top N% points of the negative pairs obtained by the calculation unit 104b as the threshold. N is a predetermined constant. Also, for example, the threshold determination unit 105b may acquire, for example, a false acceptance rate (for example, 0.001%) during face recognition operation as the threshold. The method of acquiring the threshold is not limited to these.
[0028] Calculation unit 106b calculates a loss for a positive pair similarity smaller than the threshold obtained by threshold determination unit 105b. Calculation unit 107b calculates a loss for a negative pair similarity larger than the threshold obtained by threshold determination unit 105b.
[0029] The calculation unit 108b learns a representative vector having the same number of dimensions as the feature vector calculated by the calculation unit 102b for each person label. Then, the calculation unit 108b calculates a loss that approaches a representative vector of the same person label and moves away from a representative vector of a different person label among these representative vectors. The calculation unit 108b calculates the loss using a loss function such as ArcFace shown in non-patent document (J. Deng, J. Guo, N. Xue, and S. Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In CVPR, 2019).
[0030] The learning unit 109b updates (learns) the neural network used by the calculation unit 102b so as to minimize the loss. The losses used are those from the calculation units 106b, 107b, and 108b. The neural network can be updated (learned) using a general backpropagation method or the like.
[0031] The update unit 110b uses the feature vector obtained from the image acquired by the acquisition unit 101b to update the feature vector stored in the memory bank unit 103b. The memory bank unit 103b stores a maximum of m feature vectors for each person label. If the number of feature vectors stored in the memory bank unit 103b corresponding to the person label of the image is less than m, a new feature vector is stored in the memory bank unit 103b. If m feature vectors are stored in the memory bank unit 103b corresponding to the person label of the image, the oldest registered feature vector is deleted and a new feature vector is stored.
[0032] <Learning loop processing> The learning loop process according to this embodiment will be described with reference to the flowchart in FIG. 2(a). Step S201a is the start of an epoch loop. One cycle of using all of the learning data stored in the external storage device 104a for the learning process is called one epoch. The number of repeated epochs (repeated epoch number) is determined in advance. To count the number of epoch repetitions, the CPU 101a initializes a variable i to 1. If the value of the variable i is equal to or less than the repeated epoch number, the process proceeds to step S202a, and if the value of the variable i is greater than the repeated epoch number, the process exits the loop and ends.
[0033] Step S202a is the start of the mini-batch loop. The number of samples forming a mini-batch is determined in advance, and the learning data is divided into mini-batches by that number of samples. It is assumed that numbers are assigned to each mini-batch in sequence, starting from 1. To refer to this using variable j, CPU 101a first initializes the value of variable j to 1. If the value of variable j is equal to or less than the total number of mini-batches, the process proceeds to step S203a. If the value of variable j is greater than the total number of mini-batches, the process exits the loop and proceeds to step S206a.
[0034] In step S203a, the j-th mini-batch of the learning data is acquired. Specifically, the acquisition unit 101b acquires a set of a learning image and a person label. The acquisition unit 101b may perform image processing, known as Data Augmentation, on the acquired image. For example, the acquisition unit 101b may perform processing such as changing the color of the image or adding noise to pixel values.
[0035] In step S204a, learning is performed using the mini-batches obtained in step S203a. Details of the process in step S204a will be described later with reference to the flowchart in FIG.
[0036] Step S205a is the end of the mini-batch loop, where the CPU 101a adds 1 to the value of the variable j, and the process proceeds to step S202a. Step S206a is the end of the epoch loop, where the CPU 101a adds 1 to the value of the variable i, and the process returns to step S201a.
[0037] <Learning step processing> The details of the learning step process performed in step S204a above will be described with reference to the flowchart in Fig. 2(b). In step S201b, the calculation unit 102b calculates a feature vector from the image included in the mini-batch acquired by the acquisition unit 101b.
[0038] In step S202b, the calculation unit 104b calculates a pair (positive pair) similarity (positive pair similarity) between a feature vector with the same person label as the person label of the image included in the mini-batch acquired by the acquisition unit 101b among the feature vectors held in the memory bank unit 103b and the acquired feature vector. The calculation unit 104b also calculates a pair (negative pair) similarity (negative pair similarity) between a feature vector with a person label different from the person label of the image among the feature vectors held in the memory bank unit 103b and the acquired feature vector.
[0039] More specifically, the calculation unit 104b obtains, for each sample of the mini-batch, a feature vector of a person label that is the same as the person label included in the sample from the memory bank unit 103b, and calculates the similarity between the feature vectors as a positive similarity. The calculation unit 104b also obtains a feature vector of a person label that is different from the person label included in the sample, and calculates the similarity between the feature vectors as a negative pair similarity.
[0040] In step S203b, the threshold value determining unit 105b obtains a threshold value used for loss calculation, which will be described later. In this embodiment, the threshold value determining unit 105b obtains a predetermined threshold value th.
[0041] In step S204b, the calculation unit 106b calculates the loss of the positive pair. Specifically, the calculation unit 106b calculates the loss of the positive pair using the following (Equation 1) and (Equation 2).
[0042]
number
[0043]
number
[0044] Here, (Equation 1) is a formula for calculating the loss of one positive pair (positive loss), and (Equation 2) is a formula for summarizing the losses of all positive pairs. p indicates the loss of one positive pair. sim p indicates the similarity of the positive pair. th indicates the threshold obtained in step S203b. r is a preset constant, and the smaller r is, the greater the loss for pairs that are farther away from the threshold. N is the exponent of the radical root, and a fixed value such as N=2 may be used. If the similarity of the positive pair is smaller than the threshold, the formula for the radical root in the upper part of (Formula 1) is applied, and otherwise the loss of the positive pair is zero.
[0045] Also, L p indicates the loss of all positive pairs. N p indicates the number of positive pairs that show a similarity below the threshold. In other words, it is the number of positive pairs to which the upper part of (Equation 1) is applied.
[0046] In step S205b, the calculation unit 107b calculates the loss of the negative pair. Specifically, the calculation unit 107b calculates the loss of the negative pair by using the following (Equation 3) and (Equation 4).
[0047]
number
[0048]
number
[0049] Here, (Equation 3) is a formula for calculating the loss of one negative pair (negative loss), and (Equation 4) is a formula for combining the losses of all negative pairs. n indicates the loss of one negative pair. sim n indicates the similarity of the negative pair. th indicates the threshold value obtained in step S203b. r is a preset constant, and the smaller r is, the greater the loss for pairs that are farther away from the threshold. N is the exponent of the radical root, and a fixed value such as N=2 may be used. When the similarity of the negative pair is greater than the threshold, the formula for the radical root in the upper part of (Equation 3) is applied, and otherwise the loss of the negative pair is zero. L n indicates the loss of the entire negative pair. N n is the number of negative pairs that show a similarity greater than the threshold. In other words, it is the number of negative pairs to which the upper part of (Equation 3) is applied.
[0050] In step S206b, the calculation unit 108b learns a representative vector having the same number of dimensions as the feature vector calculated by the calculation unit 102b for each person label, and obtains the loss of the above-mentioned ArcFace, etc.
[0051] In step S207b, the learning unit 109b updates (learns) the neural network used by the calculation unit 102b by the backpropagation method so as to minimize the loss L shown in the following equation 5.
[0052]
number
[0053] Here, L clsis the loss (representative vector loss) calculated in step S206b. λ1 and λ2 are weighting parameters and are predefined constants.
[0054] In the backpropagation method, the update amount of each weight of the neural network is calculated from the gradient value, which is the first derivative of the loss function. p =th and sim n Since =th is a non-differentiable point, if the similarity takes this value, the learning is continued with the gradient value set to 0 (the same applies to the non-differentiable points of the equations of other loss functions). In step S208b, the update unit 110b updates the feature vector held in the memory bank unit 103b by using the feature vector of the mini-batch.
[0055] <Detailed explanation of loss function> The positive loss and negative loss described in (Equation 1) and (Equation 3) will be explained in detail with reference to FIG. 3. FIG. 3(a) is a histogram showing the number of pairs for each similarity. 301a shows a negative pair, 302a shows a positive pair, and 303a shows a threshold. Since losses are generated for positive pairs on the left side of the threshold and negative pairs on the right side of the threshold, the graph of loss and similarity is as shown in FIG. 3(b). 301b is the loss of the positive pair, and 302b is the loss of the negative pair. The derivatives of those loss functions are as shown in FIG. 3(c). 301c is the derivative of the positive loss, and 302c is the derivative of the negative loss. From FIG. 3(c), it can be seen that the absolute value of the gradient is large around the threshold 303a. This shows that learning of positive pairs and negative pairs is strongly performed around the threshold. In addition, as the threshold is moved away from the threshold, the absolute value of the gradient rapidly decreases, but the gradient does not disappear to zero, and learning of pairs far from the threshold is also performed. Therefore, while focusing on learning pairs close to the threshold, it is possible to learn to improve pairs far from the threshold.
[0056] In the above formula, radical roots are used. However, other functions may be used. For example, a logarithmic function may be used. For example, if a logarithmic function is used for the loss functions of (Formula 1) and (Formula 3), the following (Formula 6) and (Formula 7) are obtained, respectively.
[0057]
number
[0058]
number
[0059] Unlike (Equation 1) and (Equation 3), 1 is added to avoid zero in the logarithmic function. Alternatively, a negative inverse proportional function may be used as in (Equation 8) and (Equation 9) below.
[0060]
number
[0061]
number
[0062] Alternatively, an arctangent function may be used as shown in the following (Equation 10) and (Equation 11).
[0063]
number
[0064]
number
[0065] As mentioned above, the loss function using the arctangent has no non-differentiable points and is non-zero across the entire range of similarity. All of the above function examples have the property that "the absolute value of the gradient is (more) large near a certain threshold, and the absolute value of the gradient is (more) small even when it is farther away from the threshold." This is because the gradient (derivative of the loss) is a "power function with a negative exponent." A power function is when x a where a is called the exponent of the power function. A "power function with a negative exponent" also has the effect of rapidly suppressing the output when the input increases, since its derivative is also a "power function with a negative exponent." Therefore, it has the effect of creating a gradient that is large near the threshold, but rapidly decreases as the input moves away from the threshold. For example, if the loss function is the logarithmic function log(x), then the first derivative is x -1 And the second derivative is -x -2 The absolute value of the first derivative |x -1 Since | decreases, the slope of the logarithm also decreases as x increases. In addition, the absolute value of the second derivative |-x -2 | also decreases, so the slope of the logarithmic function also decreases "rapidly" as x increases. This tendency also holds true for the radical roots and inverse proportional functions in the above equations, because their first derivatives are "power functions with negative exponents." In addition, the arctangent tan -1 The first derivative of (x) is (1+x 2 ) -1 This tendency holds true because it is a "power function with a negative exponent." The first derivative of the loss function using the arctangent of (Equation 11) is known as the Cauchy distribution. In the first derivative, the "power function with a negative exponent" is not limited to radical roots, inverse proportional functions, logarithmic functions, arctangents, etc.
[0066] Furthermore, compared to the sigmoid function used in Non-Patent Document 1, the gradient of the sigmoid function is composed of an exponential function. Therefore, the gradient decreases more rapidly than with a power function, causing the gradient to disappear. Therefore, by using a "power function with a negative exponent", the characteristic that "the absolute value of the gradient remains small even when it is far from the threshold" can be obtained. FIG. 3(g) is a diagram comparing the gradient with the sigmoid function, taking as an example the negative loss using the logarithmic function of (Equation 7). The negative loss of the sigmoid function can be calculated by the following (Equation 12).
[0067]
number
[0068] Here, t is a parameter for controlling the gradient of the sigmoid function. FIG. 3(g) is a comparison diagram of the gradients of (Formula 7) (logarithmic function) and (Formula 12) (sigmoid function). The amount of update of the weights of the neural network according to the magnitude of the gradient can be adjusted by the learning rate, etc., so the magnitude of the gradient is not important. Therefore, in order to make it easier to compare, both gradients are normalized to the range of 0 to 1 by dividing them by the maximum value. FIG. 3(g) also illustrates the range of similarity 0 or more. 303a in FIG. 3(g) is a threshold value, which is 0.3 in this example. 304g, 305g, and 306g in FIG. 3(g) respectively illustrate the gradients of the sigmoid function (Formula 12) at t=0.03, t=0.1, and t=0.2. 302g is the gradient of the logarithmic function (Formula 7) at r=0.03. As shown in the figure, in 304g, where the parameter t that controls the gradient is small, the gradient occurs only near the threshold, and the gradient disappears to zero abruptly. On the other hand, if the parameter t is increased to forcibly avoid the disappearance of the gradient, the gradient gradually occurs in the range of large similarity. However, the property of the gradient decreasing abruptly is lost, and the property of focusing on learning near the threshold is lost. In contrast, in 302g of (Equation 7), although the gradient decreases abruptly as it moves away from the threshold, the gradient does not disappear to zero, but a small gradient occurs. Note that in Figure 3(g), negative loss is shown as an example, but the same is true for positive loss. In addition, the logarithmic function (Equation 7) and the sigmoid function were compared, but the gradient is also a "power function" in functions such as power root, inverse proportional, and arctangent, and the same tendency is shown.
[0069] In addition, the "function that generates a gradient with a large absolute value near a predetermined threshold and a small absolute value even when it is far from the threshold" may be combined with a sigmoid function. For example, a logarithmic function or the like may be used in a predetermined domain. For example, a configuration may be used in which a sudden loss of a sigmoid function is used near the threshold and a logarithmic function is used in the portion far from the threshold. Alternatively, the sum of a logarithmic function or the like and a sigmoid function may be used as a loss function. In this case, a logarithmic function or the like is used for a certain term of a polynomial. The form in which the "function that generates a gradient with a large absolute value near a predetermined threshold and a small absolute value even when it is far from the threshold" is included in the loss function is not limited to these.
[0070] In addition, in the above, the same function is used for both the positive loss and the negative loss, but different functions may be used for each. Alternatively, a function that does not have the property of "decreasing the absolute value of the gradient" may be used for one of them. In other words, the gradient may not decrease as it moves away from the threshold, and all similarity pairs may be used for learning. This allows for learning to focus on the area near the threshold for only one of the negative and positive pairs, while learning overall, not limited to the area near the threshold, for the other.
[0071] <Example of loss function using multiple thresholds> In the above example, the threshold determination unit 105b acquires only one threshold, and the same threshold is used for both the positive loss and the negative loss. However, different thresholds may be used for each. Specifically, as shown in the following (Equation 13) and (Equation 14). Here, an example using a logarithmic function is shown, but functions such as a power root may also be used.
[0072]
number
[0073]
number
[0074] Here, th p is the positive threshold, and th n is the negative threshold. A histogram of similarity when two thresholds are used is shown in FIG. 3(d). 304d is the positive threshold, and 303d is the negative threshold. By setting the positive threshold 304d higher than the negative threshold 303d, it is possible to raise the positive pairs with a margin for the negative threshold 303d. This reduces the possibility that the positive pairs will mistakenly fall below the threshold when the face recognition system is operated using the negative threshold 303d, and prevents unauthorized recognition.
[0075] In addition, a small loss may be applied to positive pairs larger than the positive threshold 304d. That is, a loss is also applied to positive pairs larger than the positive threshold 304d in FIG. 3(d). Specifically, the following (Equation 15) is used.
[0076]
number
[0077] Here, u is a parameter that adjusts the strength of the loss applied to pairs larger than the positive threshold. Increasing u applies more loss to positive pairs close to the threshold, while decreasing u applies more loss to positive pairs farther away from the threshold.
[0078] FIG. 3(e) shows the shape of the loss function. 304d is the positive threshold, and 303d is the negative threshold. 301e is the loss function for positive pairs below the same positive threshold as above. On the other hand, 305e is the loss function for positive pairs above the positive threshold. Loss function 305e generates a smaller loss value than loss function 301e. Also, 302e is a negative loss, which is a visualization of the above-mentioned (Equation 14).
[0079] FIG. 3(f) is a graph of the derivative of the loss. 302f is the derivative of the loss of the negative pair exceeding the negative threshold corresponding to the negative loss 302e. 301f is the derivative of the loss of the positive pair below the positive threshold corresponding to the loss function 301e. 305f is the derivative for the positive pair exceeding the positive threshold corresponding to the loss function 305e. The absolute value of the gradient decreases as the derivative 302f, the derivative 301f, and the derivative 305f move away from the threshold, so that pairs close to the threshold are learned with emphasis, while pairs far from the threshold are also learned gradually. In addition, comparing the derivative 301f and the derivative 305f, the update of the gradient by the positive pair exceeding the threshold is relatively small. Therefore, the overall similarity of the positive pairs can be maintained high while giving priority to the update of the positive pairs below the threshold.
[0080] In addition, the positive threshold or the negative threshold may be configured to have two or more. For example, when the negative threshold is configured to have two, the negative loss is expressed as the following (Equation 16).
[0081]
number
[0082] Here, th1 n is the first negative threshold, th2 n is the second negative threshold. When the first negative threshold is exceeded, the loss becomes zero. Between the first and second negative thresholds, a loss similar to the negative loss described above is applied. This prevents negative pairs with a similarity greater than a certain value from being used in learning. Therefore, it is possible to ignore high similarities that arise when what should have been a positive pair are processed as a negative pair due to errors in person labels, etc. Similarly, it is possible to ignore low similarities caused by noise in the positive loss by setting losses below a certain value to zero. When learning in this way, th2 nThe face recognition system is operated using the negative threshold value of . It should be noted that although (Equation 13) to (Equation 16) are examples using logarithmic functions, radical roots, etc. may also be used.
[0083] <Modifications of the learning device> In the above, the representative vector loss is also used as a loss, but it may not be used. Alternatively, it may be combined with other face recognition losses. For example, the triplet loss described in Non-Patent Document (Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clustering. In CVPR, 2015) may be used. In that case, the configuration may be changed to delete the calculation unit 108b and calculate the triplet loss. Alternatively, other losses may be combined, and the combination of losses is not limited to these.
[0084] In the above, pairs are formed from feature vectors obtained from mini-batches and feature vectors from memory banks, but pairs may be formed from only mini-batches. In other words, positive pairs are formed from feature vectors with the same person label in the mini-batch, and negative pairs are formed from feature vectors with different person labels. In this case, the configuration may be changed to exclude the memory bank unit 103b and the update unit 110b. Furthermore, in the above, the acquisition unit 101b acquires one or more pairs of images and person labels, but it may also acquire "image pairs" and "pair labels." The "pair labels" are either positive pairs or negative pairs. In this case, the calculation unit 104b does not need to create pairs, and the configuration may be changed to obtain the similarity of each pair.
[0085] In the above, old feature vectors are also used for learning unless they are deleted from the memory bank unit 103b. However, a "validity time" may be set for the feature vector, and feature vectors that have exceeded the validity time may not be used for learning. Specifically, the memory bank unit 103b holds a validity time (integer value) for each feature vector. When a feature vector is registered by the update unit 110b in step S208b, a validity time (a predetermined positive integer value) is set for the feature vector, and as time passes, 1 is subtracted from the validity time of the feature vector stored in the memory bank unit 103b. Then, when calculating the pair similarity in step S202b, pairs are generated using only feature vectors with a validity time of 1 or more among the feature vectors stored in the memory bank unit 103b. This makes it possible to prevent old feature vectors registered in the memory bank unit 103b from being used for learning.
[0086] <Inference device> Next, an example of the functional configuration of an inference device 600a that uses a neural network trained by the learning device 100b in Fig. 1(b) to determine whether faces shown in two images are the same person's face will be described with reference to the block diagram in Fig. 6(a). First, an example of an inference device that registers person images in advance and determines whether a match image is one of the registered person images will be described.
[0087] The acquisition unit 601a acquires a registered image and a person label. The person label is information indicating a person. The registered image and the person label may be acquired from the external storage device 104a. Alternatively, a captured image transmitted from the network camera 112a may be acquired as a registered image, and a "person label of a person in the captured image" input by a user operating the input device 109a may be acquired.
[0088] The acquiring unit 602a acquires a match image. The acquiring unit 602a may acquire an image stored in the external storage device 104a as the match image, or may acquire a captured image transmitted from the network camera 112a as the match image.
[0089] The calculation unit 603a calculates a feature vector of the registered image acquired by the acquisition unit 601a and a feature vector of the match image acquired by the acquisition unit 602a, using the "neural network used by the calculation unit 102b" trained by the learning device 110b. The database unit 604a holds the feature vector calculated by the calculation unit 603a for the registered image acquired by the acquisition unit 601a and the corresponding person label in association with each other.
[0090] The matching unit 605a matches feature vectors with each other. Specifically, the matching unit 605a calculates cosine similarities between the feature vector calculated by the calculation unit 603 for the match image acquired by the acquisition unit 602a and each feature vector registered in the database unit 604a, and outputs the calculated cosine similarities.
[0091] The matching unit 605a may output information indicating whether each of the calculated cos similarities exceeds a predetermined threshold. Of the feature vectors registered in the database unit 604a, a person corresponding to a feature vector whose cos similarity with the feature vector of the match image is equal to or greater than the threshold can be determined to be the same person as the person in the match image. On the other hand, of the feature vectors registered in the database unit 604a, a person corresponding to a feature vector whose cos similarity with the feature vector of the match image is less than the threshold can be determined to be not the same person as the person in the match image. Thus, the matching unit 605a may make this determination and output the result of the determination.
[0092] In addition, the matching unit 605a may output a person label stored in the database unit 604a in association with a feature vector registered in the database unit 604a whose cosine similarity with the feature vector of the matching image is equal to or greater than a threshold value.
[0093] In the above, the acquisition unit 601a acquires a person label, but it may acquire only a registered image. In this case, the database unit 604a holds only the feature vector. When configured in this way, the matching unit 605a does not output the person label of the matched feature vector, but outputs information indicating whether or not a registered image of a person identical to the person in the matching image is registered in the database unit 604a.
[0094] In addition, although the above describes an example in which a person is registered in advance, it may be configured to match two images, a registered image and a match image. In this case, the database unit 604a is not necessary. That is, the acquisition unit 601a and the acquisition unit 602a acquire a registered image and a match image, respectively, and the calculation unit 603a obtains a feature vector of the registered image and a feature vector of the match image. The matching unit 605a obtains a similarity between the feature vector of the registered image and the feature vector of the match image, and outputs information indicating whether the similarity exceeds a predetermined threshold value. Alternatively, the matching unit 605a may be configured to output the obtained similarity.
[0095] <Effects of this embodiment> In this way, according to the present embodiment, since learning of pairs showing similarity near a threshold is emphasized, it is possible to improve the accuracy of authentication when operating near the threshold. In particular, by using a loss function including "a function that generates a large absolute value of the gradient near a predetermined threshold and a small absolute value of the gradient even when it is far from the threshold", it is possible to focus on learning cases close to the threshold while also learning cases far from the threshold in a gradual manner.
[0096] In addition, by using different thresholds for positive loss and negative loss, the possibility that positive pairs will mistakenly fall below the threshold when the authentication system is operated using a negative threshold is reduced, thereby preventing unauthorized authentication.
[0097] In addition, by applying a small loss to positive pairs greater than the positive threshold, it is possible to maintain a high similarity for all positive pairs while prioritizing learning of positive pairs near and below the threshold.
[0098] [Second embodiment] In this embodiment, the difference from the first embodiment will be described, and unless otherwise noted below, it is the same as the first embodiment. In the first embodiment, the learning data was a set of an image and a person label. In this embodiment, an embodiment in which attributes are added to the image and person label and used as learning data will be described. In this embodiment, attributes are used to configure pairs of only combinations of predetermined attributes, so that learning of a set of predetermined attributes is focused. In addition, the loss is strengthened by increasing the gradient of the loss of a specific attribute, and similarly, learning of a specific attribute is focused.
[0099] <Pairing based on race and other personal information> In the following, a learning device that uses "race" as an attribute and configures pairs so that both sides of the pair are of the same race for learning is described. Since people of the same race are more similar and therefore more difficult to distinguish, this allows focused learning of difficult pairs. An example of the functional configuration of the learning device 400a according to this embodiment is described using the block diagram in FIG. 4(a).
[0100] The acquisition unit 401a acquires one or more pairs (learning data) of an image, a person label, and an attribute. Hereinafter, one or more "pairs of an image, a person label, and an attribute" acquired by the acquisition unit 401a are referred to as a mini-batch. Also, the "pairs of an image, a person label, and an attribute" may be referred to as a sample. The attributes are metadata of an image, and in this embodiment, the "race" of the face in the image is held. The race has values such as "East Asia", "Southeast Asia", "White", and "Black". The "pairs of an image, a person label, and an attribute" are stored in the external storage device 104a or the like.
[0101] The memory bank unit 403a holds feature vectors in association with person labels and attributes. The memory bank unit 403a holds m feature vectors (upper limit number) for each person label, where m is a predetermined constant. Note that m feature vectors may be held for each pair of person label and attribute. This allows feature vectors to be held for each person label without bias in attributes. In this case, when the memory bank update unit 110b additionally stores feature vectors, if the number of feature vectors for each pair of person label and attribute exceeds m, old feature vectors are deleted and stored. The memory bank unit 403a holds person labels, attributes, and feature vectors in the RAM 103a or the like.
[0102] The filter unit 410a identifies a feature vector stored in the memory bank unit 403a based on the attribute for each image acquired by the acquisition unit 401a. Specifically, the filter unit 410a acquires a feature vector having the same attribute as the attribute corresponding to each image acquired by the acquisition unit 401a from the memory bank unit 403a. For example, when the attribute of the image acquired by the acquisition unit 401a is "white person", the filter unit 410a acquires a feature vector corresponding to the attribute "white person" from the memory bank unit 403a.
[0103] The calculation unit 404a acquires, for each image acquired by the acquisition unit 401a, a feature vector acquired by the filter unit 410a for that image. Then, for the acquired feature vector, the calculation unit 404a calculates a similarity with a feature vector of the same person label as a positive pair similarity, and calculates a similarity with a feature vector of a different person label as a negative pair similarity.
[0104] In this embodiment, like the first embodiment, step S204a is performed according to the flowchart in Fig. 2(b), but step S202b is performed according to the flowchart in Fig. 4(b), which is different from the first embodiment. Details of the process in step S202b according to this embodiment will be described with reference to the flowchart in Fig. 4(b).
[0105] Step S401b is the start of a loop of learning samples. It is assumed that a number is assigned to each sample in a mini-batch, starting from 1. To refer to this using variable k, the CPU 101a first initializes the value of variable k to 1. If the value of variable k is equal to or less than the number of samples, the process proceeds to step S402b. If the value of variable k exceeds the number of samples, the process exits the loop and ends.
[0106] In step S402b, the acquisition unit 401a acquires the k-th sample from the mini-batch. In step S403b, the filter unit 410a identifies an attribute to be paired with the attribute of the sample acquired in step S402b. Specifically, the filter unit 410a identifies the attribute of the sample. That is, when the attribute of the sample is "white person", the filter unit 410a identifies it as "white person".
[0107] In step S404b, the filter unit 410a acquires a feature vector corresponding to the attribute acquired in step S403b from the memory bank unit 403a. For example, when "white" is specified in step S403b, the filter unit 410a acquires a feature vector corresponding to "white" from the memory bank unit 403a.
[0108] In step S405b, the calculation unit 404a calculates a positive pair similarity and a negative pair similarity using the feature vector of the sample and the feature vector acquired in step S404b. More specifically, the calculation unit 404a calculates a positive pair similarity from the similarity between feature vectors of the same person label as the person label of the sample. In addition, the calculation unit 404a calculates a negative pair similarity from the similarity between feature vectors of a person label different from the person label of the sample. Step S406b is the end of the sample loop, and the CPU 101a adds 1 to the variable k, and the process proceeds to step S401b.
[0109] In the above, only "race" was used as personal information, but other attributes such as "gender" and "age" may be used. It is also possible to use "place of birth" as an alternative to "race". It is also possible to use attributes that combine these. In other words, when using both race and gender, only pairs in which the two are the same are formed. In addition, categorical values divided into 10-year units may be used for age, etc. Personal information is not limited to these.
[0110] In the above, pairs with the same personal information attribute are formed. However, pairs with similar personal information may be formed. For example, if the racial attribute is "East Asia", "Southeast Asia", "White", "Black", etc., "East Asia" and "Southeast Asia" are similar races, so "East Asia" and "Southeast Asia" may be paired. In addition, in the case of age, etc., pairs may be formed with a predetermined range of ±5 years as similar ages. In addition, since there is a large change in the younger generation depending on age, if the age is less than 20 years old, the range for determining similarity may be changed to ±3 years. Specifically, the filter unit 410a may be configured to obtain a feature vector similar to the attribute of the sample from the memory bank unit 403a. At this time, in step S403b of FIG. 4(b), all similar attributes are obtained, and in step S404b, a feature vector that matches any of them may be obtained from the memory bank unit 403a. However, when using continuous values such as age, all ages may be found in step S403b, but a range conditional expression may be generated in step S403b, and feature vectors that satisfy the conditional expression may be retained in step S404b. In this way, the filter may be not only a value match, but also a complex conditional expression such as a range.
[0111] In the above, the filter unit 410a supplies the feature vector to the calculation unit 404a. However, the filter unit 410a may supply "identifier information for identifying a feature vector that meets the conditions" to the calculation unit 404a. The calculation unit 404a may then select the pair similarity based on the identifier information. That is, the calculation unit 404a may calculate the similarity for all pairs, and then use the identifier information to obtain the pair similarity to be used. The advantage of adopting such a configuration is that a similarity-based pair selection method can be used at the same time. That is, a negative pair showing a similarity higher than a predetermined similarity, or a positive pair showing a similarity lower than a predetermined similarity, can be said to be a difficult pair. If such difficult pairs are selected, focused learning of difficult pairs can be performed. When such a selection method is used at the same time, it is more efficient to calculate the similarity of all pairs and then select pairs with the same attribute. The method of transmitting the feature vector identified by the filter unit 410a is not limited to these.
[0112] <Loss Enhancement by Attributes> Next, we will explain a method of focusing on learning a specific attribute by strengthening the loss of that attribute. Even when learning is limited to the same attribute, each attribute value has strengths and weaknesses. For example, a weak attribute may result in a higher similarity for negative pairs. Therefore, by changing r in (Equation 3) described above depending on the attribute, the loss can be strengthened or weakened depending on the attribute. Figure 4(c) is a graph of the loss when r in (Equation 3) is changed depending on the race. In this example, we assume a case where Asians are weak and Caucasians are strong. The loss for Asians is 401c, and the loss for Caucasians is 403c. 402c is the loss for other races. Even with the same similarity, Asians who are weak have a larger loss. This is because r in (Equation 3) is small for Asians who are weak attributes, and large for Caucasians who are strong attributes. In this way, by changing the aforementioned r depending on the attribute, the strength of the loss can be changed depending on the attribute. Also, Figure 4(d) is a graph of the loss gradient, and it can be seen that Asians who are weak attributes have a larger gradient than other attributes and are strongly learned. It can be seen that by changing r depending on the attribute, the magnitude of the gradient can also be changed.
[0113] In this case, r is changed according to the attribute of the sample acquired by the acquisition unit 401a as described above. However, when both attributes of a pair are applicable, the corresponding r may be used. Alternatively, when either of the attributes of one of the pair is applicable, the corresponding r may be used. In addition, although the above description was given taking the case of negative loss as an example, it may be used in the case of positive loss in (Formula 1).
[0114] <Pair configuration based on image quality such as masks> Next, assuming "presence or absence of mask" as an attribute, an example of the functional configuration of a learning device 500a that performs learning by forming pairs such that only one side of a pair has a mask will be described with reference to the block diagram of FIG. 5(a).
[0115] When a face recognition system that judges whether or not the person in two images is the same is operated, the image on one side is a registration image and therefore the person is not wearing a mask. Therefore, in many cases, the only pairs evaluated during operation are "without mask and with mask" and "without mask and without mask," and "with mask and with mask" are not evaluated. Therefore, in the following, learning is performed by forming pairs in which only one side has a mask. In addition, in FIG. 5(a), the same reference numbers are used for the functional units that are the same as those shown in FIG. 4(a), and the description of these functional units is omitted.
[0116] The filter unit 510a includes a determination unit 511a. The determination unit 511a determines the quality of an image acquired by the acquisition unit 401a from the attribute of the image. For example, if the attribute of the image is "no mask", the determination unit 511a determines the image as "high quality", and if the attribute of the image is "masked", the determination unit 511a determines the image as "low quality". If the image acquired by the acquisition unit 401a is "low quality", the filter unit 510a acquires a feature vector determined to be "high quality" from the memory bank unit 403a. On the other hand, if the image acquired by the acquisition unit 401a is "high quality", the filter unit 510a acquires a feature vector of "low quality" or "high quality" from the memory bank unit 403a.
[0117] In this embodiment, like the first embodiment, processing is performed according to the flowchart in Fig. 2(a), but in step S204a, processing is performed according to the flowchart in Fig. 5(b), which is different from the first embodiment. Details of the processing in step S204a according to this embodiment will be described with reference to the flowchart in Fig. 5(b). In Fig. 5(b), the same step numbers are assigned to the same processing steps as those shown in Fig. 4(b), and descriptions of these processing steps will be omitted.
[0118] In step S503b, the determination unit 511a determines the quality from the attribute of the sample acquired in step S402b. Specifically, the determination unit 511a determines the quality as "low quality" when the attribute of the sample is "masked" and determines the quality as "high quality" when the attribute of the sample is "unmasked".
[0119] In step S504b, filter unit 510a determines a quality to be paired with the quality determined in step S503b. Specifically, if the sample in step S503b is "low quality", filter unit 510a determines it as "high quality", and if the sample in step S503b is "high quality", filter unit 510a determines it as both "low quality" and "high quality".
[0120] In step S505b, the filter unit 510a obtains the feature vector of the corresponding quality obtained in step S504b from the memory bank unit 403a. Specifically, the filter unit 510a determines the quality by the determination unit 511a based on the attribute of the feature vector of the memory bank unit 403a, and obtains the feature vector of the corresponding quality obtained in step S504b. Note that the quality determination result of the memory bank unit 403a may be cached and used. In the above, an example in which "mask" is used as an attribute has been described, but the quality may be determined by an attribute as shown in FIG. 5(c).
[0121] The occlusion attribute is an attribute related to the occlusion of the face of a person shown in the image. In addition to the presence or absence of a mask, possible attributes include the presence or absence of "sunglasses" or "hat." Regarding "makeup," normal makeup is acceptable, but "abnormal" makeup such as face paint, as seen when watching soccer games, may be rated as low quality. It is also possible that parts of the face may be occluded by other items or hands, in which case the quality may be determined by the proportion of the occluded area.
[0122] Appearance attributes are attributes related to how a person appears in an image. "Face size" is the length of the short side of the rectangular area of the face, and quality is determined by whether or not that length is greater than a specified size, such as 100 pixels. "Eye width" is the number of pixels between the two eyes, and quality is determined by whether or not it is greater than a specified number of pixels. "Eyes closed" is an attribute of whether the eyes are closed or not, and when the eyes are open, it is determined that the quality is high. "Expression" is the facial expression of a person, and an expression close to a straight face is determined to be high quality. On the other hand, abnormal cases such as an extremely wide open mouth, or a smile, may be determined to be low quality. "Facial orientation" is the orientation of the face, and a rotation angle of pitch, roll, and yaw within a predetermined threshold (±5 degrees) is determined to be high quality, and anything other than that is determined to be low quality.
[0123] The image quality attribute is an attribute related to image quality. "Brightness" is the brightness of the face, and whether it is bright or not is judged by threshold processing (processing to compare the magnitude with a threshold) of numerical values such as brightness. "Image capture device" is information about the device that took the image, and if it is a specified camera, it may be judged as high quality, and if it is not, it may be judged as low quality. Here, if it is taken with a single-lens reflex camera, it is high quality, and if it is not, it is low quality. "Noise" is blur, blur, etc., and if there is a specified amount or more of them, it is judged as low quality. Alternatively, other noise such as salt-and-pepper noise may be added to make the judgment. Attributes for judging quality are not limited to these. In addition, image quality attributes may include resolution.
[0124] It is not necessary to use all of these, and they may be selectively combined. A possible combination method is to combine them using AND conditions. For example, only two conditions, "mask" and "sunglasses", may be used, and the image may be judged as "high quality" only when both are judged as high quality, and may be judged as "low quality" when either is judged as low quality. However, if many conditions are all used as AND conditions, the number of images judged as high quality will be extremely small, and the number of pairs that can be used for learning will decrease. Therefore, the feature vector to be left for each attribute may be determined and left using OR conditions. That is, the feature vector to be left is obtained from the memory bank unit 403a depending on the presence or absence of "mask". Next, the feature vector to be left is obtained from the memory bank unit 403a depending on the presence or absence of "sunglasses". The feature vector that is judged to be left under either one of the conditions may be left to form a pair. Also, multiple attributes may be treated as AND conditions, and several sets of attributes may be created, and the feature vectors to be left for these may be left using OR conditions. For example, "mask" and "sunglasses" may be treated as one AND condition, and "eye width" and "brightness" may be treated as another AND condition. Then, the feature vector to be kept is obtained by ANDing these two conditions, and the result is integrated by ORing the feature vector to be kept. The combination of attributes is not limited to these.
[0125] In the above, pairs are formed excluding "low quality and low quality". However, pairs may also be formed excluding "high quality and high quality". Specifically, when the image acquired by the acquisition unit 401a is "high quality", only the feature vector of "low quality" may be acquired from the memory bank unit 403a. Since "high quality and high quality" pairs are easy to match, learning can be limited to the difficult "low quality and high quality" pairs. This can improve the accuracy of difficult pairs.
[0126] In the above, the acquiring unit 401a also holds attributes in the external storage device 104a, but it may be configured to obtain the attributes from an image using an attribute determiner or the like. For example, if the attribute is whether or not a mask is worn, a neural network that estimates the presence or absence of a mask for an image may be prepared, and the results estimated by the neural network may be used. Alternatively, if the attribute is brightness, it may be possible to calculate the luminance value of the image. The configuration of the attribute estimator is not limited to these.
[0127] In the above, the quality is estimated from the attribute, but the quality may be estimated directly from the image. That is, a neural network may be trained to output the quality when an image is input, and the output of the neural network may be obtained as the quality attribute. A quality score for an image may be calculated by a neural network that classifies a data set of high quality images and low quality images. This quality score may be held as an attribute. In this case, the determination unit 511a may hold a threshold value for the quality score, and may determine that the quality score is high quality when the quality score exceeds the threshold value, and low quality otherwise. The quality may be included in the attribute, and the method of obtaining the quality from an image is not limited to these.
[0128] <Inference device> Next, a functional configuration example of an inference device 600b that uses a neural network trained by the learning device according to this embodiment to determine whether faces shown in two images are the faces of the same person will be described with reference to FIG. 6(b). Basically, it can be configured in the same way as the inference device having the functional configuration example shown in FIG. 6(a). Here, an inference device that uses a quality judgment process similar to that of the judgment unit 511a used in "pair formation based on image quality such as masks" during inference as well, thereby preventing formation of unlearned pairs, will be described. In FIG. 6(b), the same functional units as those shown in FIG. 6(a) are assigned the same reference numbers, and descriptions of these functional units will be omitted.
[0129] The determination unit 601b obtains the attributes of the image obtained by the obtaining unit 601a. Specifically, the determination unit 601b may be configured to obtain the attributes of the image from the image using an attribute determiner or the like. For example, the attribute determiner estimates whether or not a mask is being worn in the image. Alternatively, the determination unit 601b may be configured to store the attributes of the image in the external storage device 104a in advance and obtain the attributes. The method of obtaining the attributes of the image by the determination unit 601b is not limited to the above.
[0130] The determining unit 602b determines the quality of the image based on the same criteria as the determining unit 511a used in the learning device. Specifically, the determining unit 602b may perform the same process as the determining unit 511a. For example, if the determining unit 602b is configured to determine the quality based on the presence or absence of a mask, the image is determined to be low quality if the mask is present, and high quality if the mask is absent.
[0131] Alternatively, the determination unit 602b may use stricter criteria than the determination unit 511a. For example, the criteria for the numerical range may be made stricter. That is, the determination unit 511a determines that the image is of high quality when the "eye width" is 50 pixels or more, but the determination unit 602b determines that the image is of high quality when the value is a numerical value greater than 50 pixels (for example, 70 pixels). Alternatively, the criteria may be made stricter by increasing the number of attributes to be used. That is, the determination unit 602b may use attributes that the determination unit 511a did not use, and may determine that the image is of "high quality" when all attributes are determined to be of high quality. This allows higher quality images to be held in the registration database unit 604a, improving the accuracy of matching.
[0132] When the determining unit 602b determines that the image is of low quality, the notifying unit 603b notifies the result of the determination. Specifically, the notifying unit 603b displays on the monitor 110a or the like that the registration has failed because the image is of low quality. In this case, the feature vector of the image acquired by the acquiring unit 601a is not registered in the database unit 604a.
[0133] In the above, the database unit 604a is a component, but this is not essential. When the database unit 604a is not provided, the matching unit 605a calculates the similarity between the feature vectors of the images obtained by the acquisition unit 601a and the acquisition unit 602a. However, when the determination unit 602b determines that the image is of low quality, the matching unit 605a may not calculate the similarity. Alternatively, the calculation unit 603a may not calculate the feature vector of the image.
[0134] <Effects of this embodiment> Thus, according to this embodiment, in the example shown in "Pair formation based on personal information such as race", pairs with similar personal information can focus on learning difficult pairs, improving accuracy. In the example shown in "Strengthening loss based on attributes", learning can be focused on weak attributes, reducing the variation in accuracy due to attributes. In the example shown in "Pair formation based on image quality such as masks", learning can be focused on pairs that appear during operation of the face recognition system, improving accuracy during operation. In addition, by checking the quality of registered images in the inference device using a standard that is equal to or stricter than the image quality used during learning, matching with pairs that have not been learned can be avoided, preventing erroneous recognition and non-recognition.
[0135] [Third embodiment] The learning device and the inference device of the above embodiment are targeted at face recognition tasks, but may be applied to other recognition tasks that apply distance learning. For example, it may be applied to other biometric authentication such as determining whether two images of eyes are of the same person based on images of the iris or other pupils. The types of recognition tasks that the learning device and the inference device handle are not limited to these.
[0136] In addition, the numerical values, processing timing, processing order, processing subject, data (information) acquisition method / destination / source / storage location, etc. used in each of the above embodiments are given as examples to provide a concrete explanation, and are not intended to be limited to these examples.
[0137] In addition, a part or all of the embodiments described above may be used in appropriate combination. In addition, a part or all of the embodiments described above may be used selectively.
[0138] (Other embodiments) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.
[0139] The invention of this specification includes the following learning device, inference device, learning method, inference method, and computer program.
[0140] (Item 1) A first acquisition means for acquiring one or more pairs of an image and a label; a first calculation means for calculating a feature vector from an image using a feature extractor; a second calculation means for calculating a similarity between the feature vector calculated by the first calculation means from the image acquired by the first acquisition means and a feature vector with the same label as the feature vector calculated by the first calculation means as a positive pair similarity, and for calculating a similarity between the feature vector with a different label as a negative pair similarity; a determination means for determining a similarity threshold; A third calculation means for calculating a loss value for a positive pair similarity smaller than the threshold value; A fourth calculation means for calculating a loss value for a negative pair similarity greater than the threshold value; a learning means for learning parameters of the feature extractor that make the loss value calculated by the third calculation means and the loss value calculated by the fourth calculation means smaller; Equipped with The third calculation means or the fourth calculation means Use a loss function that generates a function that has a larger absolute value of the gradient near a given threshold and a smaller absolute value of the gradient even when it is far from the threshold. A learning device characterized by:
[0141] (Item 2) The determining means obtains a first threshold value and a second threshold value; The third calculation means calculates a loss value for a positive pair similarity smaller than the first threshold, The fourth calculation means calculates a loss value for a negative pair similarity greater than the second threshold. 2. The learning device according to item 1,
[0142] (Item 3) The third calculation means calculates a loss value for a positive pair similarity greater than the first threshold value that is smaller than a loss value for a positive pair similarity less than the first threshold value. 3. The learning device according to item 2,
[0143] (Item 4) The learning device described in item 1, characterized in that the loss function uses a function in a predetermined domain that generates a gradient whose absolute value is larger near a predetermined threshold and a gradient whose absolute value is smaller even when the function is away from the threshold.
[0144] (Item 5) The learning device described in item 1, characterized in that the loss function uses a function in polynomial terms that generates a function with a larger absolute value of the gradient near a predetermined threshold and a smaller absolute value of the gradient even when the function is away from the threshold.
[0145] (Item 6) 2. The learning device according to item 1, wherein the loss function is a function whose derivative has a negative exponent.
[0146] (Item 7) The first acquisition means acquires one or more positive pair images or negative pair images, The learning device according to item 1, wherein the second calculation means calculates a similarity between feature vectors of the positive pair of images as a positive pair similarity, and calculates a similarity between feature vectors of the negative pair of images as a negative pair similarity.
[0147] (Item 8) moreover, a memory bank means for holding feature vectors in association with labels; an updating means for updating the feature vector of the memory bank means with the feature vector of the image acquired by the first acquiring means; Equipped with the second calculation means calculates a similarity between the feature vector calculated by the first calculation means from the image acquired by the first acquisition means and a feature vector with the same label as the feature vector calculated by the first calculation means from the image acquired by the first acquisition means, as a positive pair similarity, and calculates a similarity between the feature vector calculated by the first calculation means and a feature vector with a different label as a negative pair similarity.
[0148] (Item 9) the memory bank means holds an upper limit number of feature vectors to be held for each label; 9. The learning device according to item 8, wherein the update means deletes the oldest feature vector and holds a new feature vector when the number of feature vectors of the labels held by the memory bank means exceeds an upper limit number.
[0149] (Item 10) A second acquisition means for acquiring a registration image; A third acquisition means for acquiring a match image; A fifth calculation means for calculating a feature vector from an image using a feature extractor trained by the learning device according to any one of items 1 to 9; a matching means for matching the feature vector of the registered image calculated by the fifth calculation means with the feature vector of the match image calculated by the fifth calculation means based on a similarity between the feature vector of the registered image calculated by the fifth calculation means and the feature vector of the match image calculated by the fifth calculation means; An inference device comprising:
[0150] (Item 11) A first acquisition means for acquiring one or more pairs of an image, a label, and an attribute; a first calculation means for calculating a feature vector from an image using a feature extractor; a memory bank means for storing feature vectors by associating labels with attributes; filter means for identifying a feature vector of said memory bank means based on said attributes; a second calculation means for calculating a similarity between the feature vector specified by the filter means and a feature vector with the same label as the feature vector calculated by the first calculation means from the image acquired by the first acquisition means as a positive pair similarity, and calculating a similarity between the feature vector specified by the filter means and a feature vector with a different label as a negative pair similarity; a threshold value determining means for determining a threshold value of the similarity; A third calculation means for calculating a loss value for a positive pair similarity smaller than the threshold value; A fourth calculation means for calculating a loss value for a negative pair similarity greater than the threshold value; an updating means for updating a feature vector of the memory bank means with a feature vector of the image acquired by the first acquiring means; a learning means for learning parameters of the feature extractor that make the loss value calculated by the third calculation means and the loss value calculated by the fourth calculation means smaller; A learning device comprising:
[0151] (Item 12) 12. The learning device according to item 11, wherein the filter means acquires feature vectors of similar attributes from the memory bank means based on person information of attributes of the image acquired by the first acquisition means.
[0152] (Item 13) moreover, The filtering means includes a determining means for determining the quality of the image based on the attribute, The learning device described in item 11 is characterized in that, when the image is determined to be of high quality, one or both of the feature vectors determined to be of high quality or low quality are acquired from among the feature vectors held in the memory bank means, and when the image is determined to be of low quality, the feature vector determined to be of high quality from among the feature vectors held in the memory bank means is acquired.
[0153] (Item 14) 12. The learning device according to item 11, wherein the third calculation means or the fourth calculation means changes the magnitude of loss depending on an attribute.
[0154] (Item 15) 13. The learning device according to item 12, wherein the personal information is any one of race, place of birth, sex, and age.
[0155] (Item 16) Item 14. The learning device according to item 13, wherein the quality of the image is determined from one or more of image quality attributes, reflection attributes, and occlusion attributes.
[0156] (Item 17) The image quality attributes include resolution, brightness, imaging device, noise, The reflection attributes include face size, face direction, closed eyes, eye width, facial expression, 17. The learning device according to item 16, wherein the occlusion attribute is any one of the presence or absence of occlusion, a mask, sunglasses, a hat, and makeup.
[0157] (Item 18) the memory bank means holds an upper limit number of feature vectors to be held for each label; 12. The learning device according to item 11, wherein the update means deletes the oldest feature vector and holds a new feature vector when the number of feature vectors of the label held by the memory bank means exceeds an upper limit number.
[0158] (Item 19) A second acquisition means for acquiring a registration image; A third acquisition means for acquiring a match image; A fifth calculation means for calculating a feature vector from an image using a feature extractor trained by the learning device according to any one of items 11 to 18; a matching means for matching the feature vector of the registered image with the feature vector of the match image based on a similarity between the feature vector of the registered image and the feature vector of the match image; An inference device comprising:
[0159] (Item 20) An inference system including an inference device trained by a learning device, The learning device includes: A first acquisition means for acquiring one or more pairs of an image, a label, and an attribute; a first calculation means for calculating a feature vector from an image using a feature extractor; a memory bank means for storing feature vectors by associating labels with attributes; a filter means for determining a quality of an image from the attributes based on the attributes, and, if the image is determined to be of high quality, acquiring one or both of the feature vectors determined to be of high quality or low quality from among the feature vectors held in the memory bank means, and, if the image is determined to be of low quality, acquiring the feature vector determined to be of high quality from among the feature vectors held in the memory bank means; a second calculation means for calculating a similarity between the feature vector specified by the filter means and a feature vector with the same label as the feature vector calculated by the first calculation means from the image acquired by the first acquisition means as a positive pair similarity, and calculating a similarity between the feature vector specified by the filter means and a feature vector with a different label as a negative pair similarity; a threshold value determining means for determining a threshold value of the similarity; A third calculation means for calculating a loss value for a positive pair similarity smaller than the threshold value; A fourth calculation means for calculating a loss value for a negative pair similarity greater than the threshold value; an updating means for updating a feature vector of the memory bank means with a feature vector of the image acquired by the first acquiring means; a learning means for learning parameters of the feature extractor that make the loss value calculated by the fourth calculation means and the loss value calculated by the fifth calculation means smaller; Equipped with The inference device is the inference device according to item 19, further comprising: a matching means for matching the feature vector of the registered image with the feature vector of the match image based on a similarity between the feature vector of the registered image and the feature vector of the match image; An attribute determination means for obtaining attributes of the match image; a quality judgment means for judging the quality of an image using the same standard as that of the filtering means or a stricter standard; an output means for outputting notification information when the quality of the registered image is lower than a predetermined standard; An inference system comprising:
[0160] (Item 21) A learning method performed by a learning device, comprising: A first acquisition step in which a first acquisition means of the learning device acquires one or more pairs of an image and a label; a first calculation step in which a first calculation means of the learning device calculates a feature vector from an image using a feature extractor; a second calculation step in which a second calculation means of the learning device calculates a similarity between the feature vector calculated in the first calculation step from the image acquired in the first acquisition step and a feature vector with the same label as the feature vector calculated in the first calculation step as a positive pair similarity, and calculates a similarity between the feature vector with a different label as a negative pair similarity; a determination step in which a determination means of the learning device determines a similarity threshold; a third calculation step in which a third calculation means of the learning device calculates a loss value for a positive pair similarity smaller than the threshold; a fourth calculation step in which a fourth calculation means of the learning device calculates a loss value for a negative pair similarity greater than the threshold; a learning step in which a learning means of the learning device learns parameters of the feature extractor that make the loss value calculated in the third calculation step and the loss value calculated in the fourth calculation step smaller; Equipped with In the third calculation step or the fourth calculation step, Use a loss function that generates a function that has a larger absolute value of the gradient near a given threshold and a smaller absolute value of the gradient even when it is far from the threshold. A learning method comprising:
[0161] (Item 22) An inference method performed by an inference device, comprising: A second acquisition step in which a second acquisition means of the inference device acquires a registration image; a third acquisition step in which a third acquisition means of the inference device acquires a match image; A fifth calculation step in which a fifth calculation means of the inference device calculates a feature vector from an image using a feature extractor trained by the learning method described in item 21; a matching step in which a matching means of the inference device matches the feature vector of the registered image calculated in the fifth calculation step with the feature vector of the match image calculated in the fifth calculation step based on a similarity between the feature vector of the registered image calculated in the fifth calculation step and the feature vector of the match image calculated in the fifth calculation step; An inference method comprising:
[0162] (Item 23) A learning method performed by a learning device, comprising: a first acquisition step in which a first acquisition means of the learning device acquires one or more pairs of an image, a label, and an attribute; a first calculation step in which a first calculation means of the learning device calculates a feature vector from an image using a feature extractor; a filtering step in which a filter means of the learning device identifies a feature vector of a memory bank means for storing feature vectors by associating a label with an attribute based on the attribute; a second calculation step in which a second calculation means of the learning device calculates, from the image acquired in the first acquisition step, a similarity between the feature vector calculated in the first calculation step and a feature vector having the same label as the feature vector specified in the filtering step as a positive pair similarity, and calculates a similarity between the feature vector calculated in the first calculation step and a feature vector having a different label as a negative pair similarity; a threshold determination step in which a threshold determination means of the learning device determines a threshold value of a similarity; a third calculation step in which a third calculation means of the learning device calculates a loss value for a positive pair similarity smaller than the threshold; a fourth calculation step in which a fourth calculation means of the learning device calculates a loss value for a negative pair similarity greater than the threshold; an updating step in which an updating means of the learning device updates a feature vector of the memory bank means with a feature vector of the image acquired in the first acquisition step; a learning step in which a learning means of the learning device learns parameters of the feature extractor that make the loss value calculated in the third calculation step and the loss value calculated in the fourth calculation step smaller; A learning method comprising:
[0163] (Item 24) An inference method performed by an inference device, comprising: A second acquisition step in which a second acquisition means of the inference device acquires a registration image; a third acquisition step in which a third acquisition means of the inference device acquires a match image; A sixth calculation step in which a sixth calculation means of the inference device calculates a feature vector from an image using a feature extractor trained by the learning method described in Item 22; a matching step in which a matching means of the inference device matches the feature vector of the registered image with the feature vector of the match image based on a similarity between the feature vector of the registered image and the feature vector of the match image; An inference method comprising:
[0164] (Item 25) 24. A computer program for causing a computer to execute each step of the learning method according to item 21 or 23.
[0165] (Item 26) A computer program for causing a computer to execute each step of the inference method according to item 22 or 24.
[0166] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0167] 101b: Acquisition unit 102b: Calculation unit 103b: Memory bank unit 104b: Calculation unit 105b: Threshold determination unit 106b: Calculation unit 107b: Calculation unit 108b: Calculation unit 109b: Learning unit 110b: Update unit
Claims
1. a first acquisition means for acquiring one or more pairs of an image and a label; a first calculation means for calculating a feature vector from an image using a feature extractor; a second calculation means for calculating a similarity between the feature vector calculated by the first calculation means from the image acquired by the first acquisition means and a feature vector with the same label as the feature vector calculated by the first calculation means as a positive pair similarity, and for calculating a similarity between the feature vector and a feature vector with a different label as a negative pair similarity; a determination means for determining a similarity threshold; a third calculation means for calculating a loss value for a positive pair similarity smaller than the threshold; a fourth calculation means for calculating a loss value for a negative pair similarity greater than the threshold; a learning means for learning parameters of the feature extractor that make the loss value calculated by the third calculation means and the loss value calculated by the fourth calculation means smaller; Equipped with The third calculation means or the fourth calculation means Use a loss function that generates a function that generates a gradient whose absolute value is larger near a given threshold and a gradient whose absolute value is smaller even when it is far from the threshold. A learning device characterized by:
2. The determining means obtains a first threshold value and a second threshold value; the third calculation means calculates a loss value for a positive pair similarity that is smaller than the first threshold; The fourth calculation means calculates a loss value for a negative pair similarity greater than the second threshold.
2. The learning device according to claim 1 .
3. The third calculation means calculates a loss value for a positive pair similarity greater than the first threshold value that is smaller than a loss value for a positive pair similarity less than the first threshold value.
3. The learning device according to claim 2.
4. The learning device according to claim 1, wherein the loss function uses, in a predetermined domain, a function that generates a gradient whose absolute value is larger near a predetermined threshold and whose absolute value is smaller even when the gradient is farther from the threshold.
5. 2. The learning device according to claim 1, wherein the loss function uses a function in polynomial terms that generates a function whose absolute value of the gradient is larger near a predetermined threshold and whose absolute value of the gradient is smaller even when the function is far from the threshold.
6. 2. The learning device according to claim 1, wherein the loss function is a function whose derivative has a negative exponent.
7. the first acquisition means acquires one or more positive pair images or negative pair images, The learning device according to claim 1, characterized in that the second calculation means calculates the similarity of the feature vectors of the positive pair of images as a positive pair similarity, and calculates the similarity of the feature vectors of the negative pair of images as a negative pair similarity.
8. moreover, a memory bank means for storing feature vectors in association with labels; an updating means for updating the feature vector of the memory bank means with the feature vector of the image acquired by the first acquiring means; Equipped with 2. The learning device according to claim 1, wherein the second calculation means calculates a similarity between the feature vector calculated by the first calculation means from the image acquired by the first acquisition means and a feature vector having the same label as the feature vector stored in the memory bank means as a positive pair similarity, and calculates a similarity between the feature vector calculated by the first calculation means from the image acquired by the first acquisition means and a feature vector having a different label as a negative pair similarity.
9. the memory bank means holds an upper limit number of feature vectors to be held for each label; 9. The learning device according to claim 8, wherein the update means deletes the oldest feature vector and stores a new feature vector when the number of feature vectors of the label stored in the memory bank means exceeds an upper limit.
10. a second acquisition means for acquiring a registration image; a third acquisition means for acquiring a match image; a fifth calculation means for calculating a feature vector from an image using a feature extractor trained by the learning device according to claim 1; a matching means for matching the feature vector of the registered image calculated by the fifth calculation means with the feature vector of the match image calculated by the fifth calculation means based on the similarity between the feature vector of the registered image calculated by the fifth calculation means and the feature vector of the match image calculated by the fifth calculation means; An inference device comprising:
11. a first acquisition means for acquiring one or more pairs of an image, a label, and an attribute; a first calculation means for calculating a feature vector from an image using a feature extractor; a memory bank means for storing feature vectors by associating labels with attributes; a filter means for identifying a feature vector of the memory bank means based on the attribute; a second calculation means for calculating a similarity between the feature vector specified by the filter means and a feature vector with the same label as the feature vector calculated by the first calculation means from the image acquired by the first acquisition means as a positive pair similarity, and a similarity between the feature vector specified by the filter means and a feature vector with a different label as a negative pair similarity; a threshold value determining means for determining a threshold value of the similarity; a third calculation means for calculating a loss value for a positive pair similarity smaller than the threshold; a fourth calculation means for calculating a loss value for a negative pair similarity greater than the threshold; an updating means for updating the feature vector of the memory bank means with the feature vector of the image acquired by the first acquiring means; a learning means for learning parameters of the feature extractor that make the loss value calculated by the third calculation means and the loss value calculated by the fourth calculation means smaller; A learning device comprising:
12. 12. The learning device according to claim 11, wherein the filter means acquires feature vectors of similar attributes from the memory bank means based on person information of attributes of the image acquired by the first acquisition means.
13. moreover, the filtering means includes a determining means for determining the quality of an image based on the attribute; The learning device according to claim 11, characterized in that, when the image is determined to be of high quality, one or both of the feature vectors held in the memory bank means that are determined to be of high quality or low quality are acquired, and when the image is determined to be of low quality, the feature vectors held in the memory bank means that are determined to be of high quality are acquired.
14. 12. The learning device according to claim 11, wherein the third calculation means or the fourth calculation means changes the magnitude of the loss depending on an attribute.
15. 13. The learning device according to claim 12, wherein the personal information is any one of race, place of birth, sex, and age.
16. The learning device according to claim 13 , wherein the quality of the image is determined from at least one of an image quality attribute, a reflection attribute, and an occlusion attribute.
17. The image quality attributes include resolution, brightness, imaging device, noise, The reflection attributes include face size, face direction, eye closure, eye width, facial expression, The learning device according to claim 16 , wherein the occlusion attribute is any one of the presence or absence of occlusion, a mask, sunglasses, a hat, and makeup.
18. the memory bank means holds an upper limit number of feature vectors to be held for each label; 12. The learning device according to claim 11, wherein the update means deletes the oldest feature vector and holds a new feature vector when the number of feature vectors of the label held by the memory bank means exceeds an upper limit.
19. a second acquisition means for acquiring a registration image; a third acquisition means for acquiring a match image; a fifth calculation means for calculating a feature vector from an image using a feature extractor trained by the learning device according to claim 11; a matching means for matching the feature vector of the registered image with the feature vector of the match image based on the similarity between the feature vector of the registered image and the feature vector of the match image; An inference device comprising:
20. An inference system using an inference device trained by a learning device, The learning device a first acquisition means for acquiring one or more pairs of an image, a label, and an attribute; a first calculation means for calculating a feature vector from an image using a feature extractor; a memory bank means for storing feature vectors by associating labels with attributes; a filter means for determining the quality of an image from the attributes based on the attributes, and if the image is determined to be of high quality, acquiring one or both of the feature vectors determined to be of high quality or low quality from among the feature vectors held in the memory bank means, and if the image is determined to be of low quality, acquiring the feature vectors determined to be of high quality from among the feature vectors held in the memory bank means; a second calculation means for calculating a similarity between the feature vector specified by the filter means and a feature vector with the same label as the feature vector calculated by the first calculation means from the image acquired by the first acquisition means as a positive pair similarity, and a similarity between the feature vector specified by the filter means and a feature vector with a different label as a negative pair similarity; a threshold value determining means for determining a threshold value of the similarity; a third calculation means for calculating a loss value for a positive pair similarity smaller than the threshold; a fourth calculation means for calculating a loss value for a negative pair similarity greater than the threshold; an updating means for updating the feature vector of the memory bank means with the feature vector of the image acquired by the first acquiring means; a learning means for learning parameters of the feature extractor that make the loss value calculated by the third calculation means and the loss value calculated by the fourth calculation means smaller; Equipped with The inference device according to claim 19, further comprising: a matching means for matching the feature vector of the registered image with the feature vector of the match image based on the similarity between the feature vector of the registered image and the feature vector of the match image; attribute determination means for obtaining attributes of the match image; a quality determination means for determining the quality of an image using the same standard as that of the filtering means or a stricter standard; an output means for outputting notification information when the quality of the registered image is lower than a predetermined standard; An inference system comprising:
21. A learning method performed by a learning device, a first acquisition step in which a first acquisition means of the learning device acquires one or more pairs of an image and a label; a first calculation step in which a first calculation means of the learning device calculates a feature vector from an image using a feature extractor; a second calculation step in which second calculation means of the learning device calculates a similarity between the feature vector calculated in the first calculation step and a feature vector with the same label as the feature vector calculated in the first calculation step from the image acquired in the first acquisition step as a positive pair similarity, and calculates a similarity between the feature vector and a feature vector with a different label as a negative pair similarity; a determination step in which a determination means of the learning device determines a threshold value of similarity; a third calculation step in which a third calculation means of the learning device calculates a loss value for a positive pair similarity that is smaller than the threshold; a fourth calculation step in which a fourth calculation means of the learning device calculates a loss value for a negative pair similarity greater than the threshold; a learning step in which a learning means of the learning device learns parameters of the feature extractor that make the loss value calculated in the third calculation step and the loss value calculated in the fourth calculation step smaller; Equipped with In the third calculation step or the fourth calculation step, Use a loss function that generates a function that generates a gradient whose absolute value is larger near a given threshold and a gradient whose absolute value is smaller even when it is far from the threshold. A learning method characterized by:
22. An inference method performed by an inference device, a second acquisition step in which a second acquisition means of the inference device acquires a registration image; a third acquisition step in which a third acquisition means of the inference device acquires a match image; a fifth calculation step in which a fifth calculation means of the inference device calculates a feature vector from an image using a feature extractor trained by the learning method according to claim 21; a matching step in which a matching means of the inference device matches the feature vector of the registered image calculated in the fifth calculation step with the feature vector of the match image calculated in the fifth calculation step based on the similarity between the feature vector of the registered image calculated in the fifth calculation step and the feature vector of the match image calculated in the fifth calculation step; An inference method comprising:
23. A learning method performed by a learning device, a first acquisition step in which a first acquisition means of the learning device acquires one or more pairs of an image, a label, and an attribute; a first calculation step in which a first calculation means of the learning device calculates a feature vector from an image using a feature extractor; a filtering step in which a filter means of the learning device identifies a feature vector of a memory bank means that associates a label with an attribute and holds the feature vector based on the attribute; a second calculation step in which second calculation means of the learning device calculates, from the feature vectors identified in the filtering step, a similarity between the feature vector calculated in the first calculation step from the image acquired in the first acquisition step and a feature vector with the same label as the feature vector calculated in the first calculation step as a positive pair similarity, and calculates a similarity between the feature vector and a feature vector with a different label as a negative pair similarity; a threshold determination step in which a threshold determination means of the learning device determines a threshold value of the similarity; a third calculation step in which a third calculation means of the learning device calculates a loss value for a positive pair similarity that is smaller than the threshold; a fourth calculation step in which a fourth calculation means of the learning device calculates a loss value for a negative pair similarity greater than the threshold; an updating step in which an updating means of the learning device updates the feature vector of the memory bank means with the feature vector of the image acquired in the first acquisition step; a learning step in which a learning means of the learning device learns parameters of the feature extractor that make the loss value calculated in the third calculation step and the loss value calculated in the fourth calculation step smaller; A learning method comprising:
24. An inference method performed by an inference device, a second acquisition step in which a second acquisition means of the inference device acquires a registration image; a third acquisition step in which a third acquisition means of the inference device acquires a match image; a sixth calculation step in which a sixth calculation means of the inference device calculates a feature vector from an image using a feature extractor trained by the learning method according to claim 23; a matching step in which a matching means of the inference device matches the feature vector of the registered image with the feature vector of the match image based on the similarity between the feature vector of the registered image and the feature vector of the match image; An inference method comprising:
25. A computer program for causing a computer to execute each step of the learning method according to claim 21 or 23.
26. A computer program for causing a computer to execute each step of the inference method according to claim 22 or 24.