Information processing apparatus, information processing method, and program
Patent Information
- Application Number
- JP2024553943
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-31
- Filing Date
- 2022-10-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-10-31
AI Technical Summary
Existing machine learning models for detecting feature information in images are ineffective when the correct feature information used for training includes errors, as they do not account for deviations in the estimation process, leading to suboptimal detection performance.
An information processing device and method that acquires target images and correct feature information, detects estimated feature information using a detection model, and calculates a loss to update the model parameters, allowing it to tolerate deviations from correct feature information, thereby improving detection accuracy even when errors are present in the training data.
The solution enables highly effective detection by accounting for errors in the correct feature information, resulting in improved accuracy and robustness of the detection model, even when the training data contains errors.
Abstract
Description
Information processing device, information processing method, and recording medium
[0001] The present disclosure relates to an information processing device, an information processing method, and a recording medium.
[0002] A machine learning model has been proposed to detect a predetermined region in an image.
[0003] Patent Literature 1 describes detecting facial landmarks using a machine learning model for pose estimation, for example, detecting multiple eye landmarks near eyebrows to extract eye regions from a face image.
[0004] Patent Document 2 describes an iris extraction device that extracts the iris portion of an eye image, and also describes that the iris extraction device can be implemented using a deep neural network.
[0005] Patent Document 3 describes the learning of a model for detecting feature figures and feature points from an image. Specifically, it describes that learning is performed by comparing correct answer data of feature figures and feature points with detected feature figures and feature points.
[0006] JP 2022-39984 A JP 2022-44603 A International Publication No. 2022 / 059066
[0007] The present disclosure aims to improve upon the techniques described in the prior art documents mentioned above.
[0008] According to one aspect of the present disclosure, there is provided an information processing device comprising: an acquisition means for acquiring a target image and ground truth feature information; a detection means for detecting estimated feature information relating to a detection target from the target image using a detection model; and a loss calculation means for calculating a loss in order to update parameters of the detection model so as to generate a detection model that can tolerate deviations in the estimated feature information from the ground truth feature information.
[0009] According to one aspect of the present disclosure, there is provided an information processing method in which one or more computers acquire a target image and ground truth feature information, detect estimated feature information related to a detection target from the target image using a detection model, and calculate a loss for updating parameters of the detection model to generate a detection model that can tolerate deviations in the estimated feature information from the ground truth feature information.
[0010] According to one aspect of the present disclosure, there is provided a computer-readable recording medium having a program recorded thereon, the program causing a computer to function as: an acquisition means for acquiring a target image and ground truth feature information; a detection means for detecting estimated feature information relating to a detection target from the target image using a detection model; and a loss calculation means for calculating a loss in order to update parameters of the detection model so as to generate a detection model that can tolerate deviations in the estimated feature information from the ground truth feature information.
[0011] 1 is a diagram illustrating an overview of an information processing device according to a first embodiment. FIG. 2 is a diagram illustrating an overview of an information processing method according to the first embodiment. FIG. 3 is a diagram illustrating functions of a detection model according to the first embodiment. FIG. 4 is a diagram illustrating a target image and feature information. FIG. 5 is a block diagram illustrating a functional configuration of an information processing device according to the first embodiment. FIG. 6 is a flowchart illustrating a processing flow performed by an information processing device according to the first embodiment. FIG. 7 is a flowchart illustrating a processing flow performed by a loss calculation unit according to the first embodiment. FIG. 8 is a diagram illustrating a relationship between a distance d and a loss L according to the first embodiment. FIG. 9 is a diagram illustrating a computer for realizing an information processing device. FIG. 10 is a flowchart illustrating a processing flow performed by an information processing device according to a first modification. (a) to (c) are diagrams for explaining errors in correct feature information. FIG. 11 is a flowchart illustrating a processing flow performed by a loss calculation unit according to a second embodiment. FIG. 12 is a diagram illustrating a first example of a relationship between a distance d and a loss L according to a third embodiment. FIG. 13 is a diagram illustrating a second example of a relationship between a distance d and a loss L according to the third embodiment. FIG. 14 is a block diagram illustrating a functional configuration of an information processing device according to a fourth embodiment. FIG. 15 is a flowchart illustrating a processing flow performed by a feature extraction unit and a loss calculation unit according to a fourth embodiment when the distance d is less than a reference value ε. FIG. 10 is a flowchart illustrating the flow of processing executed by a loss calculation unit according to a fifth embodiment. FIG. 11 is a diagram illustrating a function g(x). FIG. 12 is a diagram illustrating a function g(d-ε). FIG. 13 is a block diagram illustrating the functional configuration of an information processing device according to a sixth embodiment. FIG. 14 is a diagram for explaining an authentication unit. FIG. 15 is a diagram illustrating the functional configuration of an authentication device according to a seventh embodiment. FIG. 16 is a flowchart illustrating the flow of processing performed by an authentication device according to the seventh embodiment.
[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In all drawings, like components are designated by like reference numerals, and descriptions thereof will be omitted where appropriate.
[0013] First Embodiment FIG. 1 is a diagram illustrating an overview of an information processing device 10 according to the first embodiment. The information processing device 10 according to this embodiment includes an acquisition unit 110, a detection unit 130, and a loss calculation unit 150. The acquisition unit 110 acquires a target image and ground truth feature information. The detection unit 130 detects estimated feature information related to a detection target from the target image using a detection model. The loss calculation unit 150 calculates a loss L for updating parameters of the detection model so as to generate a detection model that can tolerate deviations in the estimated feature information from the ground truth feature information.
[0014] According to this information processing device 10, highly effective detection can be achieved even when the correct feature information contains errors.
[0015] FIG. 2 is a diagram showing an overview of an information processing method according to this embodiment. The information processing method according to this embodiment is executed by one or more computers. The information processing method according to this embodiment includes steps S10 to S30. In step S10, a target image and ground truth feature information are acquired. In step S20, estimated feature information related to the detection target is detected from the target image using a detection model. In step S30, a loss L is calculated to update the parameters of the detection model so as to generate a detection model that can tolerate deviations in the estimated feature information from the ground truth feature information.
[0016] According to this information processing method, highly effective detection can be achieved even when the correct feature information contains errors.
[0017] This information processing method can be executed by the information processing device 10 according to this embodiment.
[0018] Hereinafter, a detailed example of the information processing device 10 according to this embodiment and the information processing method according to this embodiment will be described.
[0019] Machine learning of a detection model is performed using correct answer data. However, the correct answer data may contain errors. For example, a person may manually add shapes or points related to the area to be detected in an image and use them as correct answer data. This human work may cause errors to be included in the correct answer data. Furthermore, the degree of error that is acceptable in the detection performance of a detection model varies depending on the intended use of the detection model. In other words, attempting to accurately match the output result of a detection model to correct answer data using machine learning may not necessarily be effective in light of the intended use of the detection model.
[0020] The techniques described in Patent Documents 1 to 3 do not take into consideration the possibility that the correct answer may contain errors.
[0021] According to the information processing device 10 of this embodiment, the loss calculation unit 150 calculates the loss L for updating the parameters of the detection model so as to generate a detection model that can tolerate deviations in the estimated feature information from the correct feature information. Therefore, machine learning that takes into account the influence of errors that may be contained in the correct feature information achieves detection that is highly effective in light of the purpose of the detection model. Note that in machine learning, the loss L is a value based on the magnitude of deviation between the correct answer and the prediction result output by the detection model. In machine learning, the model parameters are updated to reduce the value of the loss L.
[0022] <Target Image and Feature Information> FIG. 3 is a diagram illustrating the function of the detection model 132 according to this embodiment. The information processing device 10 is a device that performs machine learning of the detection model 132. The detection model 132 is configured to include a neural network. The detection model 132 is a model for detecting feature information related to a detection target from a target image. The feature information detected by the detection model 132 is also particularly referred to as "estimated feature information." On the other hand, the feature information included as a correct answer in training data for machine learning is also particularly referred to as "correct feature information." The "estimated feature information" and "correct feature information" are collectively referred to simply as "feature information."
[0023] The input data of the detection model 132 includes a target image. The output data of the detection model 132 includes estimated feature information. The feature information is, for example, information indicating a region of the target image to be detected (hereinafter referred to as a "target region"). In other words, the feature information can be used to identify the target region in the target image.
[0024] FIG. 4 is a diagram illustrating a target image 90 and feature information. In the example of FIG. 4, the target image 90 is an image that can be used for, for example, iris authentication. In this case, the target image 90 is an image that includes an eye. The target image 90 may be an image that includes both eyes, or an image that includes only one of the right eye or the left eye. Furthermore, the target image 90 may be an image that includes not only the eye but also the area surrounding the eye. The eye included in the target image 90 is, for example, a human eye, but may also be the eye of a non-human organism.
[0025] In the example of FIG. 4 , the detection target is, for example, an iris. The feature information is information indicating one or more of a point, a circle, an ellipse, and a shape formed by a combination thereof. However, a circle (or ellipse) with a missing part or a hidden part in the image may be indicated as a circle (or ellipse) in the feature information. When the feature information indicates a shape, the shape indicated by the feature information is not limited to a point, a circle, an ellipse, etc., and may be any shape. In the example of FIG. 4 , the feature information indicates a circle C1, a circle C2, and points P1 to P8. Circle C1 corresponds to the outer edge of the pupil, and circle C2 corresponds to the outer edge of the iris. Points P1 to P8 correspond to multiple points located on the edge of the eyelid. The iris region in the target image 90 can be identified by these circles and points.
[0026] Circles C1 and C2 are each represented by their center coordinates and radii (i.e., three elements). However, the center coordinates of circle C1 and circle C2 may be the same. Points P1 to P8 are each represented by coordinates (i.e., two elements). The feature information may be a vector including these elements. This vector is called a feature vector. For example, in the example of FIG. 4, by combining two circles each represented by three elements and eight points each represented by two elements, the feature information may be represented by a feature vector having 22 elements.
[0027] Note that the target image 90 is not limited to an image that can be used for iris authentication. For example, the target image 90 may be an image that can be used for face authentication. In this case, the target image 90 is an image that includes a face, and the detection model 132 detects feature points that indicate the positions of the eyes, nose, mouth, etc. in the face as feature information.
[0028] Alternatively, the target image 90 may be an image including an object other than a living thing. In this case, the feature information is information indicating a specific region of the object.
[0029] <Processing Flow> Fig. 5 is a block diagram illustrating an example of the functional configuration of the information processing device 10 according to this embodiment. In the example of this figure, the information processing device 10 further includes an update unit 170. Fig. 6 is a flowchart illustrating an example of the flow of processing executed by the information processing device 10 according to this embodiment. Each functional configuration unit of the information processing device 10 will be described in detail below with reference to Figs. 5 and 6 .
[0030] The acquisition unit 110 acquires the target image 90 and correct feature information. The correct feature information is feature information that corresponds to the target image 90 as a correct answer. A combination of the target image 90 and the correct feature information corresponding to the target image 90 is called training data. The correct feature information is prepared, for example, by specifying in advance for the target image 90 one or more of points, circles, ellipses, and shapes formed by combinations of these to identify a target region. This specification is performed, for example, by a person performing an operation similar to drawing points or shapes on the target image 90 displayed on a display (i.e., by adding annotations). However, the correct feature information may be prepared by other methods.
[0031] Here, the correct feature information may contain errors. That is, the correct feature information does not need to be completely correct. The correct feature information may be used as a target in machine learning.
[0032] For example, a storage unit accessible from the acquisition unit 110 stores a plurality of pieces of training data prepared in advance. The target images 90 included in the plurality of pieces of training data are different from one another. The sizes of the detection targets in the plurality of target images 90 may be different from one another, but are preferably the same from one another. The acquisition unit 110 reads and acquires the training data from the storage unit. Note that the acquisition unit 110 may acquire the training data from another device instead of reading the training data from the storage unit.
[0033] In the example shown in FIG. 6 , in step S101, the acquisition unit 110 acquires and stores a training dataset including multiple pieces of training data. Furthermore, in step S102, the acquisition unit 110 acquires two or more pieces of training data from the training dataset as training data batches. In this way, the information processing device 10 divides the training dataset into multiple data batches and performs machine learning. That is, the information processing device 10 updates the parameters of the detection model 132 for each training data batch. The number of pieces of training data included in a training data batch can be determined in advance. However, the number of pieces of training data included in a training data batch does not necessarily need to be constant. By updating the parameters of the detection model 132 for each training data batch, the accuracy of each parameter update can be improved.
[0034] Next, in step S103, the detection unit 130 detects estimated feature information in each of the target images 90 included in the training data batch.
[0035] The detection unit 130 detects estimated feature information from each target image 90 using the detection model 132. Specifically, the detection unit 130 inputs one target image 90 included in the training data batch acquired by the acquisition unit 110 to the detection model 132. Then, the detection model 132 outputs estimated feature information corresponding to the input target image 90. The output estimated feature information is associated with the input target image 90, i.e., the target image 90 from which the estimated feature information was detected. The output estimated feature information may be associated with training data that includes the target image 90 from which the estimated feature information was detected. The detection unit 130 obtains estimated feature information for all target images 90 included in the training data batch by sequentially inputting multiple target images 90 to the detection model 132.
[0036] Next, in step S104, the loss calculation unit 150 uses the estimated feature information and the correct feature information to calculate the loss L. One loss L is calculated for each training data batch. The method by which the loss calculation unit 150 calculates the loss L will be described in detail later.
[0037] The update unit 170 uses the loss L to update the parameters of the detection model 132 used by the detection unit 130. For example, the update unit 170 updates the parameters of the detection model 132 by backpropagation using the loss L. In this way, it is possible to improve the detection accuracy of the detection target by the detection model 132.
[0038] 6 , the update unit 170 calculates a gradient for each parameter of the detection model 132 using the loss L. In step S106, the update unit 170 updates each parameter of the detection model 132 based on the calculated gradient and a preset learning rate. Various existing methods can be adopted as the method by which the update unit 170 calculates the gradient and the method by which the update unit 170 updates the parameters.
[0039] In step S107, it is determined whether the termination condition for the repetition is satisfied. If the termination condition is not satisfied (No in S107), the process returns to step S102. A series of processes from step S102 to step S106 is regarded as one iteration of a learning loop, and the learning of the detection model 132 progresses by repeating this series of processes many times. If the termination condition is satisfied (Yes in S107), in step S108, the parameters of the detection model 132 at that time are saved, and the learning ends. Note that the learned parameters of the detection model 132 may be saved at an intermediate stage of learning.
[0040] The termination condition may be that the number of iterations reaches a predetermined number, or that the rate of change of the loss L calculated in step S104 with respect to the number of iterations becomes equal to or less than a predetermined value. Alternatively, the termination condition may be that at least one of these conditions is satisfied.
[0041] <Method of Calculating Loss L> The method by which the loss calculation unit 150 calculates the loss L will be described in detail below.
[0042] 7 is a flowchart illustrating the flow of processing performed by the loss calculation unit 150 according to this embodiment. In step S201, the loss calculation unit 150 calculates the distance D for each piece of training data. The distance D is the distance between estimated feature information and supervised feature information corresponding to the estimated feature information. In other words, when estimated feature information is detected in a certain target image 90, the distance D is calculated using supervised feature information included in the same training data as the target image 90 and the detected estimated feature information.
[0043] For example, as described above, the feature information includes a feature vector. A feature vector included in the correct feature information is hereinafter referred to as a "correct feature vector k a ", and the feature vector included in the estimated feature information will be referred to as "estimated feature vector k p Then, the loss calculation unit 150 calculates the correct feature vector k a and estimated feature vector k p The distance D between the two points is calculated. The distance D can be obtained using the following equation (1), for example. Here, ||||n Is L n represents a norm, and n is, for example, 1. However, n is not particularly limited and may be a value that satisfies 1≦n<∞. D=||k p -k a || n ...(1)
[0044] After the distances D are calculated for all training data included in the training data batch to be processed, in step S202, the loss calculation unit 150 calculates the distance d based on the multiple distances D. Specifically, the loss calculation unit 150 calculates the sum or average of the multiple distances D as the distance d.
[0045] The loss calculation unit 150 uses the distance d obtained in this way to calculate the loss L as follows: That is, in step S203, the loss calculation unit 150 determines whether the distance d is less than a predetermined reference value ε.
[0046] If the distance d is less than a predetermined reference value ε (Yes in step S203), the loss calculation unit 150 calculates the loss L based on the first rule (step S204). On the other hand, if the distance d is equal to or greater than the reference value (No in step S203), the loss calculation unit 150 calculates the loss L based on a second rule different from the first rule (step S205). This makes it possible to more appropriately calculate the loss L. Furthermore, even if the correct feature information contains an error, the effect of the error on learning can be reduced.
[0047] Although the first rule and the second rule are not limited, it is preferable that the obtained loss L monotonically increases or is constant with respect to the distance d in the first rule, and it is preferable that the obtained loss L monotonically increases with respect to the distance d in the second rule.
[0048] Although the relationship between the first rule and the second rule is not limited, it is preferable that any loss L calculated by the first rule is smaller than any loss L calculated by the second rule. It is also preferable that any slope of the loss L obtained by the first rule with respect to the distance d is smaller than any slope of the loss L obtained by the second rule with respect to the distance d.
[0049] For example, the first rule is a first formula L=f to obtain the loss L by substituting the distance d. 1 (d), and the second rule is a second formula for substituting the distance d to obtain the loss L: L=f 2 (d), where f 1 (d) <f 2 It is preferable that (d) holds. Then, as described above, when the distance d is less than the reference value ε, the loss L is calculated based on the first rule. Then, for all d, the second formula L=f 2 In comparison with the case where the loss L is calculated using (d), the loss L calculated when the distance d is small can be made smaller. Such a loss L makes it less likely that a parameter update will occur that attempts to bring the estimated feature information and the ground truth feature information closer together than necessary.
[0050] FIG. 8 is a diagram illustrating an example of the relationship between the distance d and the loss L according to this embodiment. In the example of FIG. 8 , the loss calculation unit 150 calculates the L2 loss as the loss L when the distance d is less than the reference value ε, and calculates the L1 loss as the loss L when the distance d is equal to or greater than the reference value ε. That is, the first rule is to set the L2 loss as the loss L, and the second rule is to set the L1 loss as the loss L. The loss L may be a so-called Huber loss. By the loss calculation unit 150 calculating the L2 loss as the loss L when the distance d is less than the reference value ε and calculating the L1 loss as the loss L when the distance d is equal to or greater than the reference value ε, even if an error is included in the correct feature information, the influence of the error on learning can be reduced.
[0051] In the example of FIG. 8, the first rule is to calculate the loss L using a quadratic function of the distance d as in the following formula (2), and the second rule is to calculate the loss L using a linear function of the distance d as in the following formula (3). 2 , a 1 , and b 1 are each a predetermined real number. 2 and a 1 is not zero. L2 loss = a 2 ×d 2 ...(2) L1 loss = a 1 ×d+b 1 ...(3)
[0052] The reference value ε may be determined in advance, or may be specified by the information processing device 10 as will be described in the sixth embodiment.
[0053] When the reference value ε is determined in advance, the reference value ε can be determined based on the accuracy required for the detection model 132 in light of the intended use of the detection model 132. Alternatively, the reference value ε may be determined based on the degree of variability in feature information added (annotated) to the same target image 90 by multiple people, or the reference value ε may be determined based on the degree of variability in feature information added to different target images 90. In this case, it is preferable to determine the reference value ε by increasing the importance of variability for target images 90 that are significantly affected by errors. For example, if the target image 90 is an image that can be used for iris authentication, it is preferable to determine the reference value ε by increasing the importance of variability for target images 90 with small iris diameters.
[0054] As shown in Figure 8, it is preferable that the change in loss L with respect to the change in distance d is continuous at the reference value ε. In other words, it is preferable that a graph with one axis representing distance d and the other axis representing loss L is continuous at the reference value ε. This allows for more appropriate learning. Furthermore, it is preferable that the loss L be differentiable with respect to d when d = ε.
[0055] The formulas for calculating the L1 loss and the L2 loss may be predetermined. However, when the reference value ε is specified by the information processing device 10, the slope a of the linear function for calculating the L1 loss may be 1 , intercept b 1 , and the coefficient a of the quadratic function for calculating the L2 loss 2 The loss calculation unit 150 may specify at least one of the slope a and the distance d based on the reference value ε. In this case, the loss calculation unit 150 may, for example, determine the slope a so that the change in the loss L with respect to the change in the distance d is continuous at the reference value ε. 1 , intercept b 1 , and coefficient a 2 Identify at least one of the following:
[0056] After the loss L is calculated in step S204 or step S205, the loss calculation unit 150 outputs the loss L in step S206.
[0057] <Hardware Configuration> The hardware configuration of the information processing device 10 will be described below. Each functional component of the information processing device 10 (the acquisition unit 110, the detection unit 130, the loss calculation unit 150, and the update unit 170) may be realized by hardware that realizes the functional component (e.g., a hardwired electronic circuit, etc.), or may be realized by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it). Below, a case where each functional component of the information processing device 10 is realized by a combination of hardware and software will be further described.
[0058] FIG. 9 is a diagram illustrating a computer 1000 for implementing the information processing device 10. The computer 1000 is any computer. For example, the computer 1000 may be a system on chip (SoC), a personal computer (PC), a server machine, a tablet terminal, or a smartphone. The computer 1000 may be a dedicated computer designed to implement the information processing device 10, or may be a general-purpose computer. The information processing device 10 may be implemented by a single computer 1000 or by a combination of multiple computers 1000.
[0059] The computer 1000 includes a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. The bus 1020 is a data transmission path through which the processor 1040, the memory 1060, the storage device 1080, the input / output interface 1100, and the network interface 1120 transmit and receive data to and from each other. However, the method of interconnecting the processor 1040 and other components is not limited to bus connection. The processor 1040 may be any of various processors, such as a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The memory 1060 is a main storage device implemented using a random access memory (RAM) or the like. The storage device 1080 is an auxiliary storage device implemented using a hard disk, a solid state drive (SSD), a memory card, a read-only memory (ROM), or the like.
[0060] The input / output interface 1100 is an interface for connecting the computer 1000 to an input / output device. For example, an input device such as a keyboard and an output device such as a display are connected to the input / output interface 1100. The input / output interface 1100 may be connected to the input device or output device via a wireless connection or a wired connection.
[0061] The network interface 1120 is an interface for connecting the computer 1000 to a network. This communication network is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network). The network interface 1120 may be connected to the network via a wireless connection or a wired connection.
[0062] The storage device 1080 stores program modules that realize the various functional components of the information processing device 10. The processor 1040 reads these program modules into the memory 1060 and executes them to realize the functions corresponding to the respective program modules.
[0063] As described above, according to this embodiment, the loss calculation unit 150 calculates the loss L for updating the parameters of the detection model so as to generate a detection model that can tolerate deviations in the estimated feature information from the ground-truth feature information. Therefore, even if the ground-truth feature information contains errors, highly effective detection is achieved.
[0064] (Modification 1) Modification 1 is a modification of the first embodiment. The information processing device 10 according to this modification is the same as the information processing device 10 according to the first embodiment except for the points described below. The information processing method according to this modification is the same as the information processing method according to the first embodiment except for the points described below.
[0065] In the first embodiment, the acquiring unit 110 acquires a plurality of training data sets, each set consisting of a target image 90 and correct feature information, as training data batches. In the first embodiment, the loss calculating unit 150 calculates the distance D for each of the plurality of sets (training data), and calculates the sum or average of the plurality of distances D as the distance d.
[0066] In this modification, the distance d and the loss L are calculated for each training data. In this case, the parameters of the detection model 132 are updated for each training data.
[0067] 10 is a flowchart illustrating the flow of processing executed by the information processing device 10 according to this modification. Step S101 and steps S105 to S108 in FIG. 10 are the same as step S101 and steps S105 to S108, respectively, described in the first embodiment.
[0068] In step S302 following step S101, the acquisition unit 110 according to this modification acquires one piece of training data from the training data set. Then, in step S303, the detection unit 130 detects estimated feature information for each of the target images 90 included in the training data. The method by which the detection unit 130 detects estimated feature information is the same as that described in the first embodiment.
[0069] In step S304, the loss calculation unit 150 uses the estimated feature information and the correct feature information to calculate the loss L. In this modification, one loss L is calculated for each training data.
[0070] The method by which the loss calculation unit 150 calculates the loss L will be described below. The loss calculation unit 150 calculates the distance D between estimated feature information and supervised feature information corresponding to the estimated feature information, in the same manner as in step S201 of Fig. 7. The loss calculation unit 150 also treats the calculated distance D as the distance d and calculates the loss L in the same manner as in the first embodiment.
[0071] After step S105, the same processing as in FIG. 6 of the first embodiment is performed, and the parameters of the detection model 132 are updated and saved.
[0072] In this modification, the same functions and effects as those of the first embodiment can be obtained.
[0073] Second Embodiment An information processing device 10 according to a second embodiment is the same as the information processing device 10 according to the first embodiment except for the points described below. An information processing method according to the second embodiment is the same as the information processing method according to the first embodiment except for the points described below.
[0074] As explained in the first embodiment, in this embodiment, the correct feature information is the correct feature vector k a The estimated feature information includes an estimated feature vector k p The loss calculation unit 150 according to this embodiment calculates the correct feature vector k a and estimated feature vector k p The distance between the target and the target object is normalized by the size of the target object. normThen, the loss calculation unit 150 calculates the loss L using the distance d.
[0075] FIGS. 11( a ) to 11 ( c ) are diagrams illustrating errors in correct feature information. When a person assigns correct feature information to a target image 90, for example, the person performs an operation such as drawing a predetermined shape superimposed on the target image 90. Correct feature information indicating the drawn shape is then generated. In FIGS. 11( a ) to 11 ( c ), a circle C1 indicating the outer edge of the pupil and a circle C2 indicating the outer edge of the iris are superimposed on the target image 90. In FIG. 11( a ), the circles C1 and C2 closely match the outer edges of the pupil and iris, respectively, and it can be said that there is little error. On the other hand, in FIGS. 11( b ) and 11 ( c ), the circle C1 is offset from the outer edge of the pupil, and the circle C2 is offset from the outer edge of the iris. In the examples of FIGS. 11( b ) and 11 ( c ), the correct feature information contains errors resulting from these offsets. When a person annotates correct feature information, such errors are unavoidable.
[0076] 11(c) is smaller than that in FIG. 11(b). Therefore, even if the amount of deviation is approximately the same in terms of size between FIG. 11(c) and FIG. 11(b), the effect of the error on the size of the detection target is greater in FIG. 11(c) than in FIG. 11(b).
[0077] In contrast, the loss calculation unit 150 according to this embodiment calculates the correct feature vector k a and estimated feature vector k p The distance between the target and the target object is normalized by the size of the target object. norm Then, the loss calculation unit 150 calculates the loss L using the distance d. Therefore, even when learning is performed using a plurality of target images 90 containing detection targets of different sizes, these target images 90 can be treated uniformly.
[0078] The loss calculation unit 150 calculates the normalized distance D norm can be calculated using the following formula (4-1) or formula (4-2), where s is the size of the detection target, and 1≦n<∞. norm = || (kp -k a ) / s || n ... (4-1) D norm = || (k p -k a ) | | n / s...(4-2)
[0079] The loss calculation unit 150 can specify, for example, the size of the detection target using the correct feature information. a Any one of the elements may indicate the size of the detection object, or the loss calculation section 150 may calculate the size of the detection object using one or more elements.
[0080] The size of the detection target may be, for example, the radius, diameter, or area of a circle indicated in the correct feature information, or the major axis, minor axis, or area of an ellipse indicated in the correct feature information. The size of the detection target may be the length of any side of a figure indicated in the correct feature information, the radius, diameter, or area of a circumscribing circle, or the distance between any two points indicated in the correct feature information. When the detection target is an iris, the size of the detection target is, for example, the radius, diameter, or area of a circle indicating the outer edge of the iris.
[0081] 12 is a flowchart illustrating the flow of processing performed by the loss calculation unit 150 according to this embodiment. In this embodiment, similar to the first embodiment, the acquisition unit 110 acquires a plurality of training data sets, each set consisting of a target image 90 and correct feature information, as training data batches. Then, in step S401, the loss calculation unit 150 calculates the normalized distance D norm Furthermore, in step S402, the loss calculation unit 150 calculates a plurality of normalized distances D norm The distance d is calculated as the sum or average of the above. By updating the parameters of the detection model 132 for each training data batch in this way, the accuracy of each parameter update can be improved.
[0082] The loss calculation unit 150 according to this embodiment uses the distance d calculated in this manner to calculate the loss L in the same manner as described in the first embodiment. Steps S203 to S206 in Fig. 12 are the same as steps S203 to S206 described in the first embodiment.
[0083] The update unit 170 according to this embodiment uses the loss L calculated by the loss calculation unit 150 to update the parameters of the detection model 132 in the same manner as described in the first embodiment.
[0084] Next, the operation and effect of this embodiment will be described. In this embodiment, the same operation and effect as in the first embodiment can be obtained. In addition, the loss calculation unit 150 according to this embodiment calculates the correct feature vector k a and estimated feature vector k p The distance between the target and the target object is normalized by the size of the target object. norm Therefore, even when learning is performed using a plurality of target images 90 with different sizes of detection targets, these target images 90 can be treated uniformly.
[0085] (Modification 2) Modification 2 is a modification of the second embodiment. The information processing device 10 according to this modification is the same as the information processing device 10 according to the second embodiment except for the points described below. The information processing method according to this modification is the same as the information processing method according to the second embodiment except for the points described below.
[0086] In the second embodiment, the acquisition unit 110 acquires a plurality of sets of training data, each set being a target image 90 and correct feature information, as training data batches. In the second embodiment, the loss calculation unit 150 calculates the normalized distance D norm Calculate multiple normalized distances D norm The sum or average of these values was calculated as the distance d.
[0087] In this modification, the distance d and the loss L are calculated for each training data. In this case, the parameters of the detection model 132 are updated for each training data.
[0088] The flow of processing executed by the information processing device 10 according to this modification is shown in FIG. 10 , similarly to the first modification. The loss calculation unit 150 according to this modification calculates the normalized distance D between estimated feature information and correct feature information corresponding to the estimated feature information, similarly to step S401 in FIG. norm Furthermore, the loss calculation unit 150 calculates the calculated normalized distance D norm is treated as the distance d as it is, and the loss L is calculated in the same manner as in the second embodiment.
[0089] Then, the same processing as that from step S105 in FIG. 6 described in the first embodiment is performed, and the parameters of the detection model 132 are updated and saved.
[0090] In this modified example, the same functions and effects as those of the second embodiment can be obtained.
[0091] (Third Embodiment) An information processing device 10 according to a third embodiment is the same as the information processing device 10 according to any one of the first embodiment, second embodiment, modified example 1, and modified example 2, except for the points described below. An information processing method according to the third embodiment is the same as the information processing method according to any one of the first embodiment, second embodiment, modified example 1, and modified example 2, except for the points described below.
[0092] 13 is a diagram showing a first example of the relationship between the distance d and the loss L according to this embodiment. When the distance d is less than the reference value ε, the loss calculation unit 150 according to this embodiment sets a predetermined value α as the loss L, and when the distance d is equal to or greater than the reference value ε, calculates the L1 loss as the loss L. That is, the first rule is to set the predetermined value α as the loss L, and the second rule is to set the L1 loss as the loss L. By doing so, even if an error is included in the correct feature information, the influence of the error on learning can be reduced.
[0093] The predetermined value α is determined in advance. There are no particular limitations on the predetermined value α, but it is preferable that the predetermined value α is smaller than any of the losses L obtained by distances d equal to or greater than the reference value ε. Furthermore, it is preferable that the change in loss L relative to the change in distance d is continuous at the reference value ε. In other words, it is preferable that a graph with one axis representing the distance d and the other axis representing the loss L is continuous at the reference value ε.
[0094] 13, the predetermined value α is 0. In this case, a distance d less than the reference value ε is allowed in training the detection model 132.
[0095] Fig. 14 is a diagram showing a second example of the relationship between the distance d and the loss L according to this embodiment. In the example of Fig. 14, the predetermined value α is not zero but a positive value. However, the predetermined value α is not particularly limited and may be a positive value, a negative value, or even zero.
[0096] The L1 loss is the same as the L1 loss described in the first embodiment. The formula used by the loss calculation unit 150 to calculate the L1 loss may be determined in advance. However, when the reference value ε is specified by the information processing device 10, the slope a of the linear function used to calculate the L1 loss may be determined in advance. 1 and intercept b 1 The loss calculation unit 150 may specify at least one of the slope a and the distance d based on the reference value ε. In this case, the loss calculation unit 150 may, for example, determine the slope a so that the change in the loss L with respect to the change in the distance d is continuous at the reference value ε. 1 and intercept b 1 Identify at least one of the following:
[0097] Next, the operation and effect of this embodiment will be described. In this embodiment, the same operation and effect as in the first or second embodiment can be obtained.
[0098] (Fourth embodiment) Except for the points described below, the information processing device 10 according to the fourth embodiment is the same as the information processing device 10 according to any of the first to third embodiments, Modification 1, and Modification 2. Except for the points described below, the information processing method according to the fourth embodiment is the same as the information processing method according to any of the first to third embodiments, Modification 1, and Modification 2.
[0099] 15 is a block diagram illustrating a functional configuration of an information processing device 10 according to a fourth embodiment. The information processing device 10 according to this embodiment further includes a feature extraction unit 140. The feature extraction unit 140 extracts an estimated object feature vector f p The acquisition unit 110 also extracts a correct object feature vector f a Then, when the distance d is less than the reference value ε, the loss calculation unit 150 according to this embodiment further obtains the estimation target feature vector f p and the correct target feature vector f a On the other hand, when the distance d is equal to or greater than the reference value ε, the loss calculation unit 150 according to this embodiment calculates the loss L by using the estimated feature vector k p and the correct feature vector k a The L1 loss based on the above is calculated as the loss L.
[0100] That is, the first rule is to estimate the feature vector f p and the correct target feature vector f a The first rule is to calculate the loss L using the L1 loss, and the second rule is to set the L1 loss as the loss L. By doing so, machine learning of the detection model 132 can be performed so as to improve the accuracy of feature extraction of the detection target.
[0101] 16 is a flowchart illustrating the flow of processing executed by the feature extraction unit 140 and the loss calculation unit 150 when the distance d is less than the reference value ε. The processing executed by the feature extraction unit 140 and the loss calculation unit 150 when the distance d is less than the reference value ε corresponds to the content of the first rule.
[0102] As described in the first embodiment, it is possible to use feature information to identify a target region in the target image 90. In step S501, the feature extraction unit 140 according to this embodiment uses estimated feature information detected by the detection unit 130 to identify a target region in the image.
[0103] For example, if the detection target is an iris, the feature extraction section 140 identifies the iris region in the target image 90 based on the estimation target information.
[0104] Then, in step S502, the feature extraction unit 140 extracts a target region from the target image 90, and in step S503, extracts features of the target region, i.e., features of the detection target. The features of the detection target are extracted as an estimated target feature vector f p The feature extraction unit 140 extracts an estimated target feature vector f from the image of the target region by using an extraction model including a neural network. p The extraction model used here is preferably a trained model. During the training process of the detection model 132, the parameters of the neural network of the extraction model may be fixed.
[0105] For example, if the detection target is an iris, the feature extraction unit 140 obtains an iris image by cutting out an iris region from the target image 90. The feature extraction unit 140 then inputs the iris image to an extraction model and outputs an estimated target feature vector f p get.
[0106] Then, in step S504, the loss calculation unit 150 calculates the estimation target feature vector f p and the correct target feature vector f a and the feature extraction loss L f Then, the loss calculation unit 150 calculates the obtained feature extraction loss L f Let L be the loss.
[0107] Feature extraction loss L f The correct target feature vector f ais included in the training data in a state associated with the target image 90. Alternatively, the training data may include personal identification information associated with the target image 90. In this case, the personal identification information and the correct target feature vector f a and are stored in a storage device accessible by the acquisition unit 110 in a state where they are associated with each other. The acquisition unit 110 then acquires a correct answer target feature vector f a can be obtained by reading it from the storage device.
[0108] Feature extraction loss L f An existing method can be used to calculate the feature extraction loss L f For example, the Softmax function and cross-entropy loss are used to calculate the feature extraction loss L f Other loss calculation methods such as the L2-Softmax function, cosface, or arcface may be used to calculate .
[0109] Feature extraction loss L f is preferably smaller than any loss L calculated by a distance d equal to or greater than the reference value ε. Furthermore, it is preferable that the change in loss L relative to the change in distance d is continuous at the reference value ε. In other words, it is preferable that a graph with one axis representing the distance d and the other axis representing the loss L is continuous at the reference value ε. To satisfy these conditions, the feature extraction loss L f A range may be set.
[0110] As described above, when the distance d is equal to or greater than the reference value ε, the loss calculation unit 150 according to this embodiment calculates the estimated feature vector k p and the correct feature vector k a The L1 loss based on the above is calculated as the loss L. The L1 loss is the same as the L1 loss described in the first embodiment.
[0111] The formula used by the loss calculation unit 150 to calculate the L1 loss may be determined in advance. However, when the reference value ε is specified by the information processing device 10, the slope a of the linear function used to calculate the L1 loss may be determined in advance. 1 and intercept b 1The loss calculation unit 150 may specify at least one of the slope a and the distance d based on the reference value ε. In this case, the loss calculation unit 150 may, for example, determine the slope a so that the change in the loss L with respect to the change in the distance d is continuous at the reference value ε. 1 and intercept b 1 Identify at least one of the following:
[0112] The hardware configuration of the computer that realizes the information processing device 10 according to this embodiment is shown in Fig. 9, for example, similar to the information processing device 10. However, the storage device 1080 of the computer 1000 that realizes the information processing device 10 according to this embodiment further stores a program module that realizes the function of the feature extraction unit 140.
[0113] Next, the operation and effect of this embodiment will be described. In this embodiment, the same operation and effect as at least one of the first to third embodiments can be obtained. In addition, when the distance d is less than the reference value ε, the loss calculation unit 150 according to this embodiment calculates the estimation target feature vector f p and the correct target feature vector f a The loss L is calculated using [mathematical formula - see original document]. Therefore, machine learning of the detection model 132 can be performed so as to improve the accuracy of feature extraction of the detection target.
[0114] Fifth Embodiment Fig. 17 is a flowchart illustrating the flow of processing executed by a loss calculation unit 150 according to a fifth embodiment. The information processing device 10 according to this embodiment is the same as the information processing device 10 according to any of the first to fourth embodiments, Modification 1, and Modification 2, except for the points described below. The information processing method according to this embodiment is the same as the information processing method according to any of the first to fourth embodiments, Modification 1, and Modification 2, except for the points described below.
[0115] The information processing device 10 according to this embodiment includes a feature extraction unit 140, similar to the information processing device 10 according to the fourth embodiment. The acquisition unit 110 also acquires a target feature vector f aIn this embodiment, the loss L calculated by the loss calculation unit 150 is expressed by the following equation (5): Loss L=β×(1−g(d−ε))×E 2 + g(d-ε) × E 1 ...(5)
[0116] In equation (5), β is a predetermined weight, ε is a predetermined reference value ε, and d is the distance d. g(d-ε) is a monotonically increasing function with an output between 0 and 1. E 1 is the correct feature vector k a and estimated feature vector k p It is a value calculated based on the above, for example, L1 loss. 2 is the feature vector f p and the correct target feature vector f a The value is calculated using E 2 is, for example, the feature extraction loss L f This is explained in detail below.
[0117] In this embodiment, when the detection unit 130 detects estimated feature information of the target image 90, the loss calculation unit 150 calculates the distance d in step S601. The distance d and the method by which the loss calculation unit 150 calculates the distance d are the same as the distance d and the method by which the loss calculation unit 150 calculates the distance d described in any one of the first embodiment, the second embodiment, the first modification, and the second modification, respectively.
[0118] Next, the loss calculation unit 150 calculates the L1 loss in step S602. The L1 loss and the method by which the loss calculation unit 150 calculates the L1 loss are the same as those described in the first embodiment and the method by which the loss calculation unit 150 calculates the L1 loss, respectively.
[0119] Next, in step S603, the loss calculation unit 150 calculates the feature extraction loss L f Calculate the feature extraction loss L f The loss calculation unit 150 calculates the feature extraction loss L f The method for calculating the feature extraction loss L f The loss calculation unit 150 calculates the feature extraction loss L fThe loss calculation section 150 may perform step S603 before at least one of S601 and S602.
[0120] Distance d, L1 loss, and feature extraction loss L f Once calculated, in step S604, the loss calculation unit 150 calculates the loss L using the above-mentioned formula (5). Note that the loss calculation unit 150 according to this embodiment does not need to perform a determination using the reference value ε, as in step S203 of FIG.
[0121] 18 is a diagram illustrating an example of the function g(x). In the example of FIG. 18, g(x) is a sigmoid function.
[0122] 19 is a diagram illustrating the function g(d−ε). In equation (5), by using the function g(d−ε), E 1 and E 2 The ratio of the contribution of the function g(d−ε) to the loss L is continuously changed. The so-called temperature parameter indicating the spread of the function g(d−ε) can be determined in advance as a part of the function g(d−ε).
[0123] When d=ε, the function g(d-ε) is 0.5. That is, E 1 and E 2 On the other hand, when d<ε, the function g(d-ε) is smaller than 0.5. That is, E 2 The contribution of E 1 The smaller d is, the greater the contribution of E to the loss L is. 1 The contribution of E becomes smaller. Also, when d>ε, the function g(d-ε) is greater than 0.5. That is, E 1 The contribution of E 2 The larger d is, the larger the contribution of E 1 In this embodiment, the change in the loss L relative to the change in the distance d is continuous at the reference value ε.
[0124] By calculating the loss L using Equation (5), the L1 loss and the feature extraction loss L, which are losses from different perspectives, are obtained. fThe contribution rate of the loss L to the
[0125] The function g(d-ε) may be any monotonically increasing function having an output between 0 and 1, and is not limited to a sigmoid function. 1 is the correct feature vector k a and estimated feature vector k p The loss is not limited to the L1 loss as long as it is calculated based on the above. 2 is the feature vector f p and the correct target feature vector f a There are no particular limitations as long as the loss is calculated using the above.
[0126] After calculating the loss L, the loss calculation unit 150 outputs the calculated loss L to the update unit 170 in step S605.
[0127] Next, the operation and effect of this embodiment will be described. In this embodiment, the same operation and effect as at least one of the first to fourth embodiments can be obtained. In addition, according to this embodiment, the loss L calculated by the detection unit 130 is expressed by the above-mentioned formula (5). Therefore, the L1 loss and the feature extraction loss L, which are losses from different perspectives, can be calculated. f The contribution rate of the loss L to the
[0128] 20 is a block diagram illustrating the functional configuration of an information processing device 10 according to a sixth embodiment. The information processing device 10 according to this embodiment is the same as the information processing device 10 according to any of the first to fifth embodiments, Modification 1, and Modification 2, except for the points described below. The information processing method according to this embodiment is the same as the information processing method according to any of the first to fifth embodiments, Modification 1, and Modification 2, except for the points described below.
[0129] The information processing device 10 according to this embodiment further includes an identification unit 160. The identification unit 160 identifies a reference value ε based on the accuracy required by the authentication unit 20 that performs authentication using the estimated feature information output from the detection model 132.
[0130] Although FIG. 20 shows an example in which the information processing device 10 includes the feature extraction unit 140, the information processing device 10 does not have to include the feature extraction unit 140.
[0131] 21 is a diagram illustrating the authentication unit 20. The authentication unit 20 includes a feature extraction unit 240. The feature extraction unit 240 performs the same processing as the feature extraction unit 140. When the information processing device 10 includes the feature extraction unit 140, as in the information processing device 10 according to the fourth embodiment, the feature extraction unit 240 included in the authentication unit 20 may be the same as the feature extraction unit 140, or may be a feature extraction unit for authentication different from the feature extraction unit 140. The feature extraction unit 240 extracts estimated target feature information indicating the features of the detection target using estimated feature information output from the detection unit 130.
[0132] The authentication unit 20 further includes a matching unit 210. A plurality of pairs of identification information (for example, personal identification information) and target characteristic information are stored in an authentication information storage unit 200 accessible from the matching unit 210. The authentication information storage unit 200 may be included in the authentication unit 20 or may be provided outside the authentication unit 20.
[0133] The matching unit 210 calculates the degree of match between the estimated target feature information extracted by the feature extraction unit 240 and each piece of target feature information stored in the authentication information storage unit 200. Then, the matching unit 210 identifies target feature information whose degree of match satisfies a predetermined condition, and identifies identification information corresponding to the identified target feature information, i.e., identification information included in the same set as the identified target feature information.
[0134] The authentication unit 20 outputs an authentication result. The authentication result output by the authentication unit 20 may be the identification information specified by the matching unit 210, or may be information indicating whether or not there is target feature information whose degree of match satisfies a predetermined condition.
[0135] Here, the detection accuracy of feature information required by the authentication unit 20 is determined in advance. This detection accuracy may be the detection accuracy required by the feature extraction unit 240. Prior to training the detection model 132, the user of the information processing device 10 inputs the detection accuracy required by the authentication unit 20 to the information processing device 10. Then, the specification unit 160 of the information processing device 10 according to this embodiment specifies a reference value ε to be used in training based on the input detection accuracy. Then, the information processing device 10 performs training of the detection model 132 using the specified reference value ε.
[0136] It can be said that the reference value ε indicates the degree to which deviation in the estimated feature information from the correct feature information is tolerated in learning the detection model 132. In other words, the smaller the reference value ε, the smaller the degree to which deviation in the estimated feature information from the correct feature information is tolerated. This also means that the influence of errors in the correct feature information becomes greater. Furthermore, the larger the reference value ε, the greater the degree to which deviation in the estimated feature information from the correct feature information is tolerated. This also means that the influence of errors in the correct feature information becomes smaller. Therefore, it is important to appropriately set the reference value ε according to the intended use of the detection model 132.
[0137] The identification unit 160 can calculate the reference value ε using, for example, a formula that indicates the relationship between the detection accuracy and the reference value ε. The formula that indicates the relationship between the detection accuracy and the reference value ε is determined in advance.
[0138] The mathematical expression showing the relationship between the detection accuracy and the reference value ε can be determined, for example, based on a relationship between the reference value ε and the accuracy of the obtained detection model 132 that has been experimentally investigated in advance.
[0139] The hardware configuration of the computer that realizes the information processing device 10 according to this embodiment is shown in Fig. 9, for example, similar to the information processing device 10. However, a program module that realizes the function of the identification unit 160 is further stored in the storage device 1080 of the computer 1000 that realizes the information processing device 10 according to this embodiment.
[0140] Next, the operation and effect of this embodiment will be described. In this embodiment, the same operation and effect as at least one of the first to fifth embodiments can be obtained. In addition, in the information processing device 10 according to this embodiment, the identification unit 160 identifies the reference value ε based on the accuracy required by the authentication unit 20 that performs authentication using estimated feature information output from the detection model 132. Therefore, a detection model 132 having appropriate detection accuracy depending on the intended use of the detection model 132 can be obtained.
[0141] Seventh Embodiment Fig. 22 is a diagram illustrating an example of the functional configuration of an authentication device 30 according to a seventh embodiment. Fig. 23 is a flowchart illustrating a processing flow performed by the authentication device 30 according to this embodiment. The authentication device 30 according to this embodiment is an authentication device that uses a detection model 132 that has been trained using an information processing device 10 according to any one of the first to sixth embodiments, Modification 1, and Modification 2. Furthermore, an authentication method according to this embodiment is an authentication method that uses a detection model 132 that has been trained using an information processing method according to any one of the first to sixth embodiments, Modification 1, and Modification 2.
[0142] The authentication device 30 according to this embodiment includes an image acquisition unit 310, a detection unit 330, and an authentication unit 350. The authentication method according to this embodiment can be executed by the authentication device 30 according to this embodiment.
[0143] In step S701, the image acquisition unit 310 acquires an image to be used for authentication. The image to be used for authentication is the same as the target image 90 described in the first embodiment. The image acquisition unit 310 may acquire the image to be used for authentication from an imaging device such as a camera, or may read and acquire the image stored in a storage device accessible from the image acquisition unit 310.
[0144] Then, in step S702, the detection unit 330 detects estimated feature information in the image acquired by the image acquisition unit 310. Specifically, the detection unit 330 inputs the image acquired by the image acquisition unit 310 (input image) to the detection model 132. The detection model 132 used here is a trained model by the information processing device 10. The detection unit 330 obtains estimated feature information for the input image as an output of the detection model 132.
[0145] Next, in step S703, the authentication unit 350 performs authentication processing using the estimated feature information and the authentication information stored in the authentication information storage unit 300. Specifically, the authentication unit 350 extracts estimation target feature information indicating the features of the detection target using the estimated feature information. The authentication unit 350 can extract the estimation target feature information using an extraction model such as that described in the fourth embodiment. The extraction model used by the authentication unit 350 may be the extraction model used in training the detection model 132 used by the detection unit 330, or may be a model different from the extraction model used in training the detection model 132 used by the detection unit 330.
[0146] The authentication information storage unit 300, which is accessible from the authentication unit 350, holds a plurality of pairs of identification information (e.g., personal identification information) and target feature information as authentication information. The authentication information storage unit 300 may be included in the authentication device 30 or may be provided external to the authentication device 30. The authentication unit 350 calculates the degree of match between the estimated target feature information extracted by the detection unit 330 and each piece of target feature information held in the authentication information storage unit 300. The authentication unit 350 then identifies target feature information whose degree of match satisfies a predetermined condition, and identifies identification information corresponding to the identified target feature information, i.e., identification information included in the same pair as the identified target feature information.
[0147] The authentication device 30 outputs an authentication result. The authentication result output by the authentication device 30 may be identification information specified by the authentication unit 350, or may be information indicating whether or not there is target feature information whose degree of match satisfies a predetermined condition. The authentication result may be displayed on a display connected to the authentication device 30, or may be output to another device.
[0148] An example of the case where the authentication device 30 performs iris authentication will be described below. The image acquisition unit 310 acquires an image including an eye as an image to be used for authentication. The detection unit 330 then inputs this image into the detection model 132 to obtain estimated feature information for identifying the iris region of the image, as shown in FIG. 4 .
[0149] The authentication unit 350 then extracts an iris image from the image including the eye, using the estimated feature information output from the detection unit 330. The authentication unit 350 then inputs the iris image to an extraction model, and obtains iris information indicating the features of the iris (estimated feature information) as an output of the extraction model.
[0150] A plurality of pairs of personal identification information and iris information are stored in advance in the authentication information storage unit 300. The authentication unit 350 calculates the degree of match between the iris information extracted by the detection unit 330 and each piece of iris information stored in the authentication information storage unit 300. The authentication unit 350 then identifies iris information whose degree of match satisfies a predetermined condition, and identifies personal identification information corresponding to the iris information identified in the authentication information.
[0151] The authentication device 30 may output the identified personal identification information, or may output information indicating whether or not there is authenticated personal identification information.
[0152] It is preferable that the information processing device 10 according to this embodiment includes the identification unit 160. The reference value ε used in the information processing device 10 according to this embodiment is preferably identified based on the detection accuracy required by the authentication unit 350, as described in the sixth embodiment.
[0153] The hardware configuration of a computer that realizes the authentication device 30 according to this embodiment is shown in Fig. 9, for example, similar to that of the information processing device 10. However, a storage device 1080 of a computer 1000 that realizes the authentication device 30 according to this embodiment stores program modules that realize the functional components of the authentication device 30 (image acquisition unit 310, detection unit 330, and authentication unit 350).
[0154] When the authentication device 30 includes the authentication information storage unit 300, the authentication information storage unit 300 is realized by the storage device 1080 of the computer 1000 that realizes the authentication device 30 according to this embodiment.
[0155] Next, the operation and effect of this embodiment will be described. According to the authentication device 30 of this embodiment, authentication is performed using the detection model 132 that has been trained using the information processing device 10 according to any one of the first to sixth embodiments, Modification 1, and Modification 2. Therefore, highly accurate authentication can be performed.
[0156] Although the embodiments and modifications of this disclosure have been described above with reference to the drawings, these are merely examples of this disclosure, and various configurations other than those described above can also be adopted.
[0157] In addition, although the flowcharts used in the above description describe multiple steps (processes) in a sequential order, the execution order of the steps performed in each embodiment and each modified example is not limited to the order described. In each embodiment and each modified example, the order of the steps shown in the drawings can be changed to the extent that the content is not affected. Furthermore, the above-described embodiments and modified examples can be combined to the extent that the content is not contradictory.
[0158] Some or all of the above embodiments can be described as in the following supplementary notes, but are not limited to them. 1-1. An information processing device comprising: an acquisition means for acquiring a target image and ground truth feature information; a detection means for detecting estimated feature information related to a detection target from the target image using a detection model; and a loss calculation means for calculating a loss to update parameters of the detection model so as to generate a detection model that can tolerate deviations in the estimated feature information from the ground truth feature information. 1-2. The information processing device described in 1-1., further comprising an update means for updating parameters of the detection model using the loss. 1-3. The information processing device described in 1-2., wherein the update means updates parameters of the detection model by an error backpropagation method using the loss. 1-4. The information processing device described in 1-1. to 1-3. 1-5. An information processing device according to any one of the above, wherein the correct feature information includes a correct feature vector, and the estimated feature information includes an estimated feature vector, and the loss calculation means calculates a distance d based on a normalized distance obtained by normalizing the distance between the correct feature vector and the estimated feature vector by the size of the detection target, and calculates the loss using the distance d. 1-5. An information processing device according to 1-4., wherein the acquisition means acquires a plurality of pairs of the target image and the correct feature information, and the loss calculation means calculates the normalized distance for each of the plurality of pairs, and the distance d is the sum or average of the plurality of normalized distances. 1-6. An information processing device according to 1-4. or 1-5., wherein the loss calculation means calculates the loss based on a first rule if the distance d is less than a predetermined reference value, and calculates the loss based on a second rule different from the first rule if the distance d is equal to or greater than the reference value. 1-7. An information processing device according to 1-6. wherein the loss calculation means calculates an L2 loss as the loss when the distance d is less than the reference value, and calculates an L1 loss as the loss when the distance d is equal to or greater than the reference value.1-8. The information processing device described in 1-6., wherein the loss calculation means sets a predetermined value as the loss when the distance d is less than the reference value, and calculates an L1 loss as the loss when the distance d is equal to or greater than the reference value. 1-9. The information processing device described in 1-6., further comprising feature extraction means for extracting an estimated target feature vector indicating a feature of the detection target indicated by the estimated feature information, wherein the acquisition means further acquires a supervised target feature vector indicating a feature of the detection target, and the loss calculation means calculates the loss using the estimated target feature vector and the supervised target feature vector when the distance d is less than the reference value, and calculates an L1 loss based on the estimated feature vector and the supervised feature vector as the loss when the distance d is equal to or greater than the reference value. 1-10. 1-4. or 1-5. The information processing device described in the above item further comprises a feature extraction means for extracting an estimated object feature vector indicating the features of the detection object indicated by the estimated feature information, wherein the acquisition means further acquires a ground truth object feature vector indicating the features of the detection object, and the loss is expressed by the following formula: Loss = β × (1 - g (d - ε)) × E. 2 + g(d-ε) × E 1 ... (Equation) β is a predetermined weight, ε is a predetermined reference value, d is the distance d, g(d-ε) is a monotonically increasing function having an output between 0 and 1, and E 1 is the L1 loss calculated based on the correct feature vector and the estimated feature vector, and E 2is a value calculated using the estimated target feature vector and the supervised target feature vector. 1-11. An information processing device according to any one of 1-6. to 1-10., wherein a change in the loss with respect to a change in the distance d is continuous at the reference value. 1-12. An information processing device according to any one of 1-6. to 1-11., further comprising: a specifying means for specifying the reference value based on the accuracy required by an authentication means that performs authentication using the estimated feature information output from the detection model. 2-1. An information processing method in which one or more computers acquire a target image and supervised feature information, detect estimated feature information related to the detection target from the target image using a detection model, and calculate a loss for updating parameters of the detection model to generate a detection model that can tolerate deviations in the estimated feature information from the supervised feature information. 2-2. An information processing method according to 2-1., wherein the one or more computers further update parameters of the detection model using the loss. 2-3. 2-2. 2-4. The information processing method described in any one of 2-1. to 2-3., wherein the correct feature information includes a correct feature vector, and the estimated feature information includes an estimated feature vector, and the one or more computers calculate a distance d based on a normalized distance obtained by normalizing the distance between the correct feature vector and the estimated feature vector by the size of the detection target, and calculate the loss using the distance d. 2-5. The information processing method described in 2-4., wherein the one or more computers obtain a plurality of pairs of the target image and the correct feature information, and calculate the normalized distance for each of the plurality of pairs, and the distance d is the sum or average of the plurality of normalized distances.2-6. The information processing method described in 2-4. or 2-5., wherein the one or more computers calculate the loss based on a first rule when the distance d is less than a predetermined reference value, and calculate the loss based on a second rule different from the first rule when the distance d is equal to or greater than the reference value. 2-7. The information processing method described in 2-6., wherein the one or more computers calculate an L2 loss as the loss when the distance d is less than the reference value, and calculate an L1 loss as the loss when the distance d is equal to or greater than the reference value. 2-8. The information processing method described in 2-6., wherein the one or more computers set a predetermined value as the loss when the distance d is less than the reference value, and calculate an L1 loss as the loss when the distance d is equal to or greater than the reference value. 2-9. 2-6. the one or more computers further extract an estimated target feature vector indicating features of the detection target indicated by the estimated feature information, the acquisition means further acquires a supervised target feature vector indicating features of the detection target, and the one or more computers calculate the loss using the estimated target feature vector and the supervised target feature vector if the distance d is less than the reference value, and calculate an L1 loss based on the estimated feature vector and the supervised feature vector as the loss if the distance d is equal to or greater than the reference value. The information processing method described in 2-10.2-4. or 2-5., the one or more computers further extract an estimated target feature vector indicating features of the detection target indicated by the estimated feature information, and the one or more computers further acquire a supervised target feature vector indicating features of the detection target, and the loss is expressed by the following formula: Loss=β×(1−g(d−ε))×E. 2 + g(d-ε) × E 1 ... (Equation) β is a predetermined weight, ε is a predetermined reference value, d is the distance d, g(d-ε) is a monotonically increasing function having an output between 0 and 1, and E 1is the L1 loss calculated based on the correct feature vector and the estimated feature vector, and E 2is a value calculated using the estimated target feature vector and the supervised target feature vector. 2-11. An information processing method according to any one of 2-6. to 2-10., wherein a change in the loss with respect to a change in the distance d is continuous at the reference value. 2-12. An information processing method according to any one of 2-6. to 2-11., wherein the one or more computers further specify the reference value based on the accuracy required by an authentication means that performs authentication using the estimated feature information output from the detection model. 3-1. A computer-readable recording medium having a program recorded thereon, the program causing a computer to function as: acquisition means that acquires a target image and supervised feature information; detection means that detects estimated feature information related to the detection target from the target image using a detection model; and loss calculation means that calculates a loss for updating parameters of the detection model so as to generate a detection model that can tolerate deviations in the estimated feature information from the supervised feature information. 3-2. 3-1. 3-3. A recording medium described in 3-2., wherein the program further causes a computer to function as an update means for updating parameters of the detection model using the loss. 3-3. A recording medium described in 3-2., wherein the update means updates parameters of the detection model by backpropagation using the loss. 3-4. A recording medium described in any one of 3-1. to 3-3., wherein the correct feature information includes a correct feature vector, and the estimated feature information includes an estimated feature vector, and the loss calculation means calculates a distance d based on a normalized distance obtained by normalizing the distance between the correct feature vector and the estimated feature vector by the size of the detection target, and calculates the loss using the distance d. 3-5. A recording medium described in 3-4., wherein the acquisition means acquires a plurality of pairs of the target image and the correct feature information, and the loss calculation means calculates the normalized distance for each of the plurality of pairs, and the distance d is the sum or average of the plurality of normalized distances.3-6. A recording medium according to 3-4. or 3-5., wherein the loss calculation means calculates the loss based on a first rule when the distance d is less than a predetermined reference value, and calculates the loss based on a second rule different from the first rule when the distance d is equal to or greater than the reference value. 3-7. A recording medium according to 3-6., wherein the loss calculation means calculates an L2 loss as the loss when the distance d is less than the reference value, and calculates an L1 loss as the loss when the distance d is equal to or greater than the reference value. 3-8. A recording medium according to 3-6., wherein the loss calculation means sets a predetermined value as the loss when the distance d is less than the reference value, and calculates an L1 loss as the loss when the distance d is equal to or greater than the reference value. 3-9. A recording medium according to 3-6. 3-10.The recording medium described in 3-4. or 3-5., wherein the program further causes a computer to function as feature extraction means that extracts an estimated target feature vector that indicates a feature of the detection target indicated by the estimated feature information, the acquisition means further acquires a supervised target feature vector that indicates a feature of the detection target, and the loss calculation means, when the distance d is less than the reference value, calculates the loss using the estimated target feature vector and the supervised target feature vector, and when the distance d is equal to or greater than the reference value, calculates an L1 loss based on the estimated feature vector and the supervised feature vector as the loss. 2 + g(d-ε) × E 1 ... (Equation) β is a predetermined weight, ε is a predetermined reference value, d is the distance d, g(d-ε) is a monotonically increasing function having an output between 0 and 1, and E 1is the L1 loss calculated based on the correct feature vector and the estimated feature vector, and E 2is a value calculated using the estimated target feature vector and the supervised target feature vector. 3-11. A recording medium according to any one of 3-6. to 3-10., wherein a change in the loss with respect to a change in the distance d is continuous at the reference value. 3-12. A recording medium according to any one of 3-6. to 3-11., wherein the program further causes a computer to function as specifying means for specifying the reference value based on the accuracy required by authentication means that performs authentication using the estimated feature information output from the detection model. 4-1. A program that causes a computer to function as: acquisition means that acquires a target image and supervised feature information; detection means that detects estimated feature information related to the detection target from the target image using a detection model; and loss calculation means that calculates a loss for updating parameters of the detection model to generate a detection model that can tolerate deviations in the estimated feature information from the supervised feature information. 4-2. A program according to 4-1., wherein the program further causes a computer to function as update means that updates parameters of the detection model using the loss. 4-3. A program in accordance with 4-2, wherein the updating means updates parameters of the detection model by backpropagation using the loss. 4-4. A program in accordance with any one of 4-1 to 4-3, wherein the correct feature information includes a correct feature vector, and the estimated feature information includes an estimated feature vector, and the loss calculation means calculates a distance d based on a normalized distance obtained by normalizing the distance between the correct feature vector and the estimated feature vector by the size of the detection target, and calculates the loss using the distance d. 4-5. A program in accordance with 4-4, wherein the acquisition means acquires a plurality of pairs of the target image and the correct feature information, and the loss calculation means calculates the normalized distance for each of the plurality of pairs, and the distance d is the sum or average of the plurality of normalized distances.4-6. A program in accordance with 4-4. or 4-5., wherein the loss calculation means calculates the loss based on a first rule when the distance d is less than a predetermined reference value, and calculates the loss based on a second rule different from the first rule when the distance d is equal to or greater than the reference value. 4-7. A program in accordance with 4-6., wherein the loss calculation means calculates an L2 loss as the loss when the distance d is less than the reference value, and calculates an L1 loss as the loss when the distance d is equal to or greater than the reference value. 4-8. A program in accordance with 4-6., wherein the loss calculation means sets a predetermined value as the loss when the distance d is less than the reference value, and calculates an L1 loss as the loss when the distance d is equal to or greater than the reference value. 4-9. 4-6. 4-10.A program according to claim 4-10, wherein the program further causes a computer to function as feature extraction means that extracts an estimated target feature vector that indicates a feature of the detection target indicated by the estimated feature information, the acquisition means further acquires a supervised target feature vector that indicates a feature of the detection target, and the loss calculation means, if the distance d is less than the reference value, calculates the loss using the estimated target feature vector and the supervised target feature vector, and if the distance d is equal to or greater than the reference value, calculates an L1 loss based on the estimated feature vector and the supervised feature vector as the loss. 4-10.A program according to claim 4-10, wherein the program further causes a computer to function as feature extraction means that extracts an estimated target feature vector that indicates a feature of the detection target indicated by the estimated feature information, the acquisition means further acquires a supervised target feature vector that indicates a feature of the detection target, and the loss is expressed by the following formula: Loss=β×(1−g(d−ε))×E. 2 + g(d-ε) × E 1 ... (Equation) β is a predetermined weight, ε is a predetermined reference value, d is the distance d, g(d-ε) is a monotonically increasing function having an output between 0 and 1, and E 1is the L1 loss calculated based on the correct feature vector and the estimated feature vector, and E 2 is a value calculated using the estimated target feature vector and the correct target feature vector. 4-11. A program according to any one of 4-6. to 4-10., wherein a change in the loss with respect to a change in the distance d is continuous at the reference value. 4-12. A program according to any one of 4-6. to 4-11., wherein the program causes a computer to further function as specifying means for specifying the reference value based on the accuracy required by authentication means that performs authentication using the estimated feature information output from the detection model.
[0159] 10 Information processing device 20 Authentication unit 30 Authentication device 110 Acquisition unit 130, 330 Detection unit 132 Detection model 140, 240 Feature extraction unit 150 Loss calculation unit 160 Identification unit 170 Update unit 200 Authentication information storage unit 210 Matching unit 300 Authentication information storage unit 310 Image acquisition unit 350 Authentication unit 1000 Computer 1020 Bus 1040 Processor 1060 Memory 1080 Storage device 1100 Input / output interface 1120 Network interface
Claims
1. An acquisition means for acquiring a target image and correct feature information; A detection means for detecting estimated feature information regarding a detection target from the target image using a detection model; A loss calculation means for calculating a loss for updating parameters of the detection model so as to generate a detection model that allows a deviation from the correct feature information in the estimated feature information, and an information processing apparatus comprising the same.
2. In the information processing apparatus according to Claim 1, the correct feature information includes a correct feature vector, the estimated feature information includes an estimated feature vector, the loss calculation means calculates a distance d based on a normalized distance obtained by normalizing the distance between the correct feature vector and the estimated feature vector by the size of the detection target, and calculates the loss using the distance d.
3. In the information processing apparatus according to Claim 2, the acquisition means acquires a plurality of sets of the target image and the correct feature information, the loss calculation means calculates the normalized distance for each of the plurality of sets, and the distance d is the sum or average of the plurality of normalized distances.
4. In the information processing apparatus according to Claim 2 or 3, the loss calculation means calculates the loss based on a first rule when the distance d is less than a predetermined reference value, and calculates the loss based on a second rule different from the first rule when the distance d is greater than or equal to the reference value.
5. In the information processing apparatus according to Claim 4, the loss calculation means calculates the L2 loss as the loss when the distance d is less than the reference value, and calculates the L1 loss as the loss when the distance d is greater than or equal to the reference value.
6. In the information processing apparatus according to Claim 4, the loss calculation means uses a predetermined value as the loss when the distance d is less than the reference value, and calculates the L1 loss as the loss when the distance d is greater than or equal to the reference value.
7. In the information processing apparatus according to Claim 4, further comprising a feature extraction means for extracting an estimated target feature vector indicating a feature of the detection target indicated by the estimated feature information, the acquisition means further acquires a correct target feature vector indicating a feature of the detection target, and the loss calculation means calculates the loss using the estimated target feature vector and the correct target feature vector when the distance d is less than the reference value. When the distance d is greater than or equal to the reference value, calculate the L1 loss based on the estimated feature vector and the correct feature vector as the loss. Information processing apparatus. **Claim 8** In the information processing apparatus according to claim 2 or 3, further comprising feature extraction means for extracting an estimated target feature vector indicating a feature of the detection target indicated by the estimated feature information, the acquisition means further acquires a correct target feature vector indicating a feature of the detection target, the loss is represented by the following (formula), Loss = β × (1 - g(d - ε)) × E 2 + g(d - ε) × E 1 ・・・(Equation) β is a predetermined weight, ε is a predetermined reference value, d is the distance d, g(d - ε) is a monotonically increasing function having an output of 0 or more and 1 or less, E 1 is the L1 loss calculated based on the correct feature vector and the estimated feature vector, E 2 is a value calculated using the estimated target feature vector and the correct target feature vector Information processing apparatus. **Claim 9** One or more computers acquire a target image and correct feature information, detect estimated feature information regarding a detection target from the target image using a detection model, calculate a loss for updating parameters of the detection model so as to generate a detection model that allows a deviation from the correct feature information in the estimated feature information Information processing method. **Claim 10** A computer acquisition means for acquiring a target image and correct feature information, detection means for detecting estimated feature information regarding a detection target from the target image using a detection model, and function as loss calculation means for calculating a loss for updating parameters of the detection model so as to generate a detection model that allows a deviation from the correct feature information in the estimated feature information Program.