Image processing device, learning device, method, and program
Patent Information
- Application Number
- JP2022152804
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-09-25
AI Technical Summary
Existing face recognition systems face challenges in accurately authenticating faces due to differences in modalities of input data, such as lighting conditions, facial expressions, and face orientation, especially when converting 2D to 3D face shape data, which introduces errors and reduces recognition accuracy.
An image processing device that uses a neural network to extract features from input data of different modalities, such as RGB images and depth information, and calculates similarity scores based on these features without intermediate 3D conversion, ensuring accurate face authentication.
The system improves face authentication accuracy by directly comparing features from varying data types, reducing errors associated with 3D conversion and enhancing robustness to lighting and orientation changes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an image processing device, a learning device, a learning method, and a program. [Background technology]
[0002] As a face authentication system, there is an image processing device that compares a face image registered in advance with a face image to be matched, and determines whether or not the face in the registered face image is the same as the face in the face image to be matched. Specifically, the image processing device acquires features from the registered face image and the face image to be matched using the same feature conversion unit, and performs face authentication by comparing the features. However, when the face image to be matched is an RGB image, the image processing device may not be able to perform face authentication because it is affected by the lighting conditions at the time of shooting, differences in facial expressions, and the direction of the face. In response to this, there is a method of performing face authentication using a face image to be matched other than an RGB image. For example, the image processing device can perform matching between the face image to be registered and the face image to be matched with higher accuracy by using three-dimensional face shape data. Here, there is a problem that a dedicated device is required to generate three-dimensional face shape data, which increases the introduction cost. Therefore, Patent Document 1 proposes face authentication when either the face data to be registered or the face data to be matched is not three-dimensional face shape data.
[0003] For example, Patent Document 1 uses a two-dimensional image as face data for registration, and a combination of a black-and-white image and depth information, also called 2.5-dimensional data, as face data for matching. Patent Document 1 then converts the two-dimensional face data and the 2.5-dimensional face data into three-dimensional face shape data, and matches the three-dimensional face shape data with each other, taking into account the difference between images having different modalities. Here, images having different modalities refer to images with different physical quantities and qualities of information that are the basis of the images. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2015-162012 A [Non-patent literature]
[0005] [Non-Patent Document 1] Deng, et. Al., ArcFace: Additive Angular Margin Loss for Deep Face Recognition. In CVPR, 2019 [Non-Patent Document 2] Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clustering. In CVPR, 2015 Summary of the Invention [Problem to be solved by the invention]
[0006] However, in Patent Document 1, a 3D face shape generator is required to perform face recognition using a method of generating 3D face shape data from 2D face data and 2.5D face data. In general, the problem of estimating a three-dimensional shape from 2D face data is an ill-posed problem, so the generated 3D face shape data contains errors. Thus, the technology in Patent Document 1 has a problem that the accuracy of face recognition is low because the accuracy of the 3D face shape data generated from the registration face data and the matching face data, which are modal different from each other, is low.
[0007] Therefore, an object of the present invention is to provide a technique for improving the accuracy of object authentication when authenticating an object using registered data and matching data that have different modalities. [Means for solving the problem]
[0008] In order to achieve the object of the present invention, an image processing device according to an embodiment of the present invention has the following configuration: That is, the image processing device includes an extraction means for extracting a first feature from first data of a first modal including information of a registered first object, and extracting a second feature from second data of a second modal different from the first modal including information of a second object to be matched, and a determination means for determining whether the first object and the second object are identical based on the first feature and the second feature, and the extraction means has been trained so that the first feature and the second feature are similar when the first object and the second object are identical. Effect of the Invention
[0009] According to the present invention, it is possible to provide a technique for improving the accuracy of object authentication when authenticating an object using registered data and matching data that have different modalities. [Brief description of the drawings]
[0010] [Figure 1] FIG. 2 is a diagram showing an example of a hardware configuration of an image processing apparatus. [Diagram 2] FIG. 2 is a block diagram showing an example of a functional configuration of the first embodiment. [Diagram 3] 1A and 1B are schematic diagrams showing a comparison between a conventional matching process and a matching process of the present invention. [Figure 4] 5 is a flowchart showing the procedure of a matching process according to the first embodiment. [Diagram 5] FIG. 4 is a diagram showing a procedure of feature conversion according to the first embodiment. [Figure 6] 5 is a flowchart showing the procedure of a learning process according to the first embodiment. [Figure 7] FIG. 4 is a schematic diagram showing the operation of a learning process according to the first embodiment. [Figure 8] 10 is a flowchart showing the procedure of a learning process according to a second embodiment. [Figure 9] FIG. 13 is a block diagram showing an example of a functional configuration of a third embodiment. [Figure 10]13 is a flowchart showing the procedure of a matching process according to a third embodiment. [Figure 11] 13 is a flowchart showing the procedure of a learning process according to a third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Hereinafter, the embodiments will be described in detail with reference to the attached drawings. Note that the following embodiments do not limit the invention according to the claims. Although the embodiments describe a number of features, not all of these features are essential to the invention, and the features may be combined in any manner. Furthermore, in the attached drawings, the same reference numbers are used for the same or similar configurations, and duplicated descriptions are omitted.
[0012] First Embodiment 1 is a diagram showing an example of a hardware configuration of an image processing device. The image processing device 10 includes a CPU 101, a ROM 102, a RAM 103, a storage unit 104, an input unit 105, a display unit 106, and a communication unit 107. Note that, although face recognition will be described in this embodiment, the image processing device 10 can perform authentication of any object, not limited to face recognition.
[0013] The CPU 101 executes a control program stored in the ROM 102 to generally control the image processing device 10 .
[0014] The ROM 102 is a non-volatile memory and stores various data and programs.
[0015] The RAM 103 temporarily stores various data from each component of the image processing device 10. The RAM 103 also expands programs so that the CPU 101 can execute the programs.
[0016] Parameters for performing feature conversion are stored in the storage unit 104. The storage unit 104 includes, for example, a hard disk drive (HDD), a flash memory, and various optical media.
[0017] The input unit 105 is a device that accepts input from a user, and includes, for example, a keyboard, a touch panel, a dial, etc. The input unit 105 is used in settings for reconstructing information such as the texture, shape, temperature, and movement of a person's face.
[0018] The display unit 106 is a device that displays the reconstruction results of the person's face information, and includes, for example, a liquid crystal display (LCD) and an organic EL display.
[0019] The communication unit 107 enables the image processing device 10 to communicate with an imaging device (not shown), an external device (not shown), and the like.
[0020] FIG. 2 is a block diagram illustrating an example of a functional configuration of the first embodiment.
[0021] The image processing device 10 includes a first input unit 201 , a second input unit 202 , a confirmation unit 203 , a storage unit 204 , a first conversion unit 205 , a second conversion unit 206 , and a matching unit 207 .
[0022] The first input unit 201 accepts a first data set.
[0023] The second input unit 202 accepts a second data set.
[0024] The verification unit 203 verifies the modality of the first data set and the modality of the second data set.
[0025] The storage unit 204 stores parameters for performing feature conversion.
[0026] The first conversion unit 205 converts the first data set into features using conversion parameters read from the storage unit 204 based on the confirmation result of the confirmation unit 203.
[0027] The second conversion unit 206 converts the second data set into features using the conversion parameters read from the storage unit 204 based on the confirmation result of the confirmation unit 203.
[0028] A matching unit 207 performs face authentication by comparing the feature amount extracted from the first data set with the feature amount extracted from the second data set.
[0029] <Input data verification processing phase> FIG. 3 is a schematic diagram showing a comparison between the conventional matching process and the matching process of the first embodiment.
[0030] FIG. 3(A) is a diagram showing a conventional matching process. In FIG. 3(A), a three-dimensional face shape generating unit 303 generates three-dimensional face shape data from first input data 301 including a face to be registered. A three-dimensional face shape generating unit 304 generates three-dimensional face shape data from second input data 302 including a face to be matched. In this case, if the first input data 301 is a two-dimensional face image, the accuracy of the three-dimensional face shape data is low because there is little depth information of the face shape for generating the three-dimensional face shape data. In addition, the resolution of the three-dimensional face shape data is reduced by converting the two-dimensional image into three-dimensional face shape data, and detailed features of the face to be registered are lost. On the other hand, if the second input data 302 is a 2.5-dimensional face image, the accuracy of the three-dimensional face shape data generated from the 2.5-dimensional face image is high. Therefore, when matching 305 (face recognition) each of the three-dimensional face shape data generated by the three-dimensional face shape generating unit 303 and the three-dimensional face shape generating unit 304, the accuracy of face recognition is reduced. In this specification, "matching" and "face recognition" are used as terms having the same meaning.
[0031] Fig. 3(B) is a diagram showing the matching process of the present invention. In Fig. 3(B), the confirmation unit 203 confirms the modality of the first input data 310. The first conversion unit 205 converts the first input data 310 into features using conversion parameters read from the storage unit 204 based on the confirmation result of the confirmation unit 203. The confirmation unit 203 confirms the modality of the second input data 311. The second conversion unit 206 converts the second input data 311 into features using conversion parameters read from the storage unit 204 based on the confirmation result of the confirmation unit 203.
[0032] Here, the conversion parameters are the results of having the neural network learn the features of the input data according to the modality of the input data. Therefore, the image processing device 10 can perform highly accurate face recognition even when images of different modalities are input. Furthermore, the image processing device 10 can extract facial features from face data of various modalities. The facial features include, for example, texture features extracted from an RGB image, shape features such as facial unevenness, and temperature distribution features.
[0033] In the present invention, the neural network is trained so that the similarity between face images containing the same face is increased. Therefore, the matching unit 207 can calculate the similarity between face images based on the inner product and angle between the feature amounts extracted from each face image. Therefore, the present invention does not require any special processing other than the calculation processing of the similarity between face images. In this way, the matching unit 207 can perform face authentication using one type of similarity, independent of the modality of the face image. This advantage is one of the characteristics of the present invention.
[0034] FIG. 4 is a flowchart showing the procedure of the matching process according to the first embodiment.
[0035] When two face images are input, the image processing device 10 determines whether or not a face in one face image is the same as a face in the other face image based on a comparison between a feature amount extracted from one face image and a feature amount extracted from the other face image. Each of the two face images may be any of a two-dimensional RGB image, a normal vector having three-dimensional face shape information, a curvature, a stereo image, and a depth image. For example, the face data for matching may be a combination of a two-dimensional image and a depth image, or a combination of the above multiple images.
[0036] In S401, the first input unit 201 accepts first input data.
[0037] In S402, the confirmation unit 203 confirms the modality of the first input data.
[0038] In S403, the first conversion unit 205 reads from the memory unit 204 conversion parameters for converting the first input data confirmed in S402 into features corresponding to the modality, and sets the read conversion parameters in the DNN described below.
[0039] In S404, the first conversion unit 205 converts the first input data into features. Here, the first conversion unit 205 has, for example, a deep neural network (hereinafter, DNN) called the convolutional neural network of Non-Patent Document 1 and the Transformer network of Non-Patent Document 2. Note that the conversion parameters include parameters such as the number of neuron layers, the number of neurons, and connection weights.
[0040] FIG. 5 is a diagram showing a procedure of feature conversion according to the first embodiment.
[0041] 5(A) shows a flow of converting a two-dimensional image and depth information input to the first input unit 201 into feature quantities. The first conversion unit 205 obtains feature quantities 505 by superimposing a two-dimensional image 501 and depth information 502 in the channel direction (S503 to S504).
[0042] 5B shows a flow of converting the two-dimensional image and depth information input to the first input unit 201 into feature quantities. The first conversion unit 205 converts the two-dimensional image 511 into feature quantities using conversion parameters (S513), and converts the depth information 512 into feature quantities using conversion parameters different from the above conversion parameters (S514). The first conversion unit 205 then superimposes the respective feature quantities converted in S513 and S514 in the channel direction (S516) to obtain feature quantity 516.
[0043] In addition, the DNN that converts a two-dimensional image into a feature and the DNN that converts depth information into a feature may share the previous layer of the DNN and may share only the subsequent layer partially depending on the state of the person's face. As described above, there are multiple methods for converting input data into features, and the method shown in FIG. 5 is an example of a method for converting input data into features. In addition, the second conversion unit 206 can convert the second input data into features using the same method as the first conversion unit 205, so a detailed description will be omitted.
[0044] In S405 to S408, the image processing device 10 performs the same processes as those in S401 to S404 on the second input data, and therefore the description thereof will be omitted. However, the second input unit 202 performs the process in S405, and the second conversion unit 206 performs the processes in S406 to S408.
[0045] As a result, the first input data and the second input data are each converted into a feature. The feature of the first input data is represented by f1, and the feature of the second input data is represented by f2. f1 and f2 are one-dimensional vectors. f1 and f2 are converted into one-dimensional vectors through processing of the fully connected layer of the DNN. In addition, the conversion parameters of the DNN of the first conversion unit 205 and the conversion parameters of the DNN of the second conversion unit 206 do not need to be the same. However, the number of output channels of the neurons in the final layer of the DNN of the first conversion unit 205 and the second conversion unit 206 is the same. As a result, the dimensions of f1 and f2 become the same.
[0046] In S409, the matching unit 207 calculates a similarity score between the feature amount f1 and the feature amount f2 using the following formula 1. Here, an index indicating the similarity between the feature amounts (i.e., the similarity score) is expressed as an angle between the feature amount vectors (see Non-Patent Document 1).
[0047] Similarity score(f1,f2):=cos(θ 12 ) =<f1,f2> ÷(|f1|·|f2|) (Equation 1) Here, θ 12 is the angle between feature vectors f1 and f2.<f1,f2> is the dot product of f1 and f2. |f1| is the length of f1. |f2| is the length of f2.
[0048] If the calculated similarity score is equal to or less than the threshold, the matching unit 207 determines that the face included in the first input data and the face included in the second input data are the same, and ends the matching process. On the other hand, if the calculated similarity score is not equal to or less than the threshold, the matching unit 207 determines that the face included in the first input data and the face included in the second input data are not the same, and ends the matching process.
[0049] <Learning process phase> Fig. 6 is a flowchart showing the procedure of the learning process according to the first embodiment. Fig. 7 is a schematic diagram showing the operation of the learning process according to the first embodiment.
[0050] Here, we will explain the learning of DNN using the representative vector method. The representative vector method is a face recognition learning method that improves the learning efficiency of DNN by setting a feature vector that represents each person (see Non-Patent Document 1).
[0051] In S601 of FIG. 6, the first conversion unit 205 converts the parameters of the DNN and the representative vectors v1 to v n is initialized with a random number. 1 to n are the IDs of all people included in the training image. Each representative vector v is a d-dimensional vector, where d is a predetermined value.
[0052] In S602, the first input unit 201 randomly selects images I1 to I m The first input data set includes a plurality of image data groups. The image data groups include one or more pieces of image data each depicting only one person. Each image data group includes information on the person's ID. Here, the image data refers to image data of various modalities. The image data of this embodiment includes, for example, an RGB image acquired by a digital camera, a black-and-white image captured by an infrared camera at night, and depth information acquired by a TOF sensor simultaneously with the black-and-white image.
[0053] For example, ID#1 in Fig. 7 indicates an image data group composed of only multiple RGB images (e.g., RGB image a, RGB image b). ID#2 indicates an image data group composed of only pairs of black and white images and depth information (e.g., black and white image p and depth information p, black and white image q and depth information q). ID#3 indicates an image data group including pairs of RGB images, black and white images and depth information (e.g., RGB image i, black and white image j and depth information j). In this way, the image data group of each ID may be composed of only images of a single modality, or may be composed of images of multiple modalities.
[0054] In S603, the confirmation unit 203 confirms the modality of the first input data set.
[0055] In S604, the first conversion unit 205 reads out the conversion parameters corresponding to the modality of the first input data confirmed by the confirmation unit 203 from the storage unit 204. As a result, the conversion parameters of the DNN in the first conversion unit 205 are changed.
[0056] In S605, the first conversion unit 205 converts each data I of the first input data set into a DNN to which conversion parameters corresponding to the modality of the first input data set are applied. i (i.e., face image data) as feature f i Here, the feature value f i is a d-dimensional vector.
[0057] In S606, the first conversion unit 205 calculates the similarity of features between each person's face image and representative vector (intra-class similarity) and the similarity of features between each person's face image and other people's representative vectors (inter-class similarity) using the following Equations 2 and 3.
[0058] Intra-class similarity score (f i ) = similarity score(f i ,V y(i) ) (Formula 2) Inter-class similarity score (f i )=Σ j≠y(i) Similarity score (f i ,V j ) (Formula 3) Here, y(i) is the input data I i is the person's ID number.
[0059] The first conversion unit 205 uses the intra-class similarity score, the inter-class similarity score, and Equation 4 to calculate a loss value to be used for training the DNN.
[0060] Loss value = Σ i Inter-class similarity score (f i )-λ × intra-class similarity score (f i ) (Formula 4) Here, λ is a weighting parameter for balancing the learning. Note that the loss value above is an example, and can be calculated by various known methods such as a similarity score with margin and cross entropy.
[0061] In S607 to S608, the first conversion unit 205 updates the conversion parameters so as to reduce the calculated loss value.
[0062] In S607, the first conversion unit 205 updates the value of the representative vector. In S608, the first conversion unit 205 updates the parameters of the DNN. Here, the conversion parameters of the DNN to be updated are only those that correspond to the modality of the first input data set read in S604. In addition, the first conversion unit 205 updates the conversion parameters of the DNN using a general backpropagation method. This allows the representative vector to function better as a value that represents the facial features of each person. The DNN of the first conversion unit learns to make the facial features of the same person closer to each other.
[0063] In S608, the first conversion unit 205 determines whether the learning of the DNN has converged, for example, based on whether the loss value is equal to or less than a threshold. If the first conversion unit 205 determines that the loss value is equal to or less than the threshold (Yes in S608), the process proceeds to S610. If the first conversion unit 205 determines that the loss value is not equal to or less than the threshold (No in S608), the process returns to S602.
[0064] In S610, the storage unit 204 stores the representative vectors V1 to V n The value of is stored.
[0065] In S611, the memory unit 204 stores the parameters of the DNN of the first conversion unit 205.
[0066] The feature space 700 in FIG. 7 shows a schematic result at the time when the learning process of the DNN is completed. The representative vector 701, the representative vector 702, and the representative vector 703 on the feature space 700 are feature vectors that respectively represent the persons of ID#1 to ID#3. The DNN of the first conversion unit 205 learns the representative vector 701 of the person of ID#1 to be located in the vicinity of the feature a and the feature b, and learns the representative vector 702 of the person of ID#2 to be located in the vicinity of the feature p and the feature q. The feature a and the feature b are the features of the person of ID#1, and are respectively illustrated by black circles in FIG. 7. The feature p and the feature q are the features of the person of ID#2, and are respectively illustrated by black circles in FIG. 7. Note that the features of the person of ID#3 are not shown in FIG. 7, but the DNN of the first conversion unit 205 learns the representative vector 703 of the person of ID#3 to be located in the vicinity of the features of the person of ID#3.
[0067] The first conversion unit 205 mainly extracts features of the positions and contours of facial organ points (e.g., eyes, nose, and mouth) of a person from the RGB image. The first conversion unit 205 also extracts features related to the contours of the face, such as the hollows of the eyes and the height of the nose, based on information such as color changes and shading in the image.
[0068] The first conversion unit 205 extracts the positions of organ points and contour features of a person's face from the black-and-white image, similar to the RGB image. Because it is difficult to extract facial color features from a black-and-white image, the first conversion unit 205 uses detailed concave / convex features based on depth information to complement the features extracted from the black-and-white image. This allows the first conversion unit 205 to extract detailed shape information for identifying a person's face.
[0069] In this way, the DNN is trained by changing the conversion parameters of the first conversion unit 205 according to the modality of the input data, while arranging the facial features of various people extracted from the input data in the same feature space 600. This allows the image processing device 10 to perform highly accurate face recognition even when input data of various modalities is received.
[0070] <Derived form of input data modal> The modality of input data accepted by the first input unit 201 and the second input unit 202 includes, but is not limited to, a two-dimensional RGB image and three-dimensional shape information. Here, a specific example of input data (image data) having a modality other than a two-dimensional RGB image and three-dimensional shape information will be described.
[0071] The input data includes a plurality of images captured by switching the imaging settings of the digital camera to various settings. The plurality of images are input to the first input unit 201 or the second input unit 202. The imaging settings are, for example, settings of shutter speed, exposure, aperture, white balance, and ISO sensitivity. The input data may be, for example, a pair of a noisy image captured with an imaging setting of a fast shutter speed and underexposure, and an image captured with an imaging setting of sufficient exposure but blurred. The input data may also be a pair of an image captured with supplemental lighting such as a strobe, and an image captured without supplemental lighting.
[0072] Furthermore, the input data may be an infrared image captured by an infrared camera, an image captured using infrared pattern projection, a spectroscopic image acquired using a spectrometer, or a polarized image captured by a polarized camera. The input data may be a combination of an RGB image and the above images. Each of the images described above may have different resolutions, or may be scaled to have the same resolution.
[0073] The input data may be a plurality of images separated and imaged for each incident angle of the light beam by a microlens array or the like, a camera having a compound lens, or a plurality of images captured by a group of cameras with different optical axes. These parallax images contain distance information, and therefore have features of the concave and convex shape of the face, compared to images without parallax. This allows the first conversion unit 205 and the second conversion unit 206 to extract features of the concave and convex shape of the face from the parallax images with high accuracy.
[0074] <Advantages of the First Embodiment> Even if the modality of the registration face data and the modality of the matching face data are different from each other, the first conversion unit 205 can learn the features extracted from the registration face data and the features extracted from the matching face data in the same feature space. As a result, according to this embodiment, highly accurate face authentication can be performed without intermediately generating a three-dimensional face shape, etc.
[0075] In addition, there are cases where face authentication becomes easier by combining features extracted from a two-dimensional face image with depth features. For example, the operation will be described when input data including a set of a two-dimensional face image with a face facing forward and depth information as face data for registration, and a two-dimensional face image with a face facing sideways as face data for matching is input to the image processing device 10. In this case, the first conversion unit 205 has difficulty extracting so-called facial feature depth, such as nose height, from the two-dimensional face image (face image with a face facing forward). On the other hand, the second conversion unit 206 can extract the facial feature depth of a person's face from the two-dimensional face image (i.e., face image with a face facing sideways).
[0076] However, when a set of a two-dimensional face image and depth information exists as face data for registration, the first conversion unit 205 can extract the feature of the depth of the person's face from the depth information. Therefore, the first conversion unit 205 has both the feature extracted from the two-dimensional face image and the feature extracted from the depth information. As a result, when the matching unit 207 matches the feature extracted by the first conversion unit 205 from the set of the two-dimensional face image and the depth information with the feature extracted by the second conversion unit 206 from the two-dimensional face image, the matching can be performed while taking into consideration the feature including the depth of the face. This further improves the accuracy of face authentication.
[0077] <Second embodiment> In general, due to the ease of obtaining data, most of the data used in face recognition is images, and there is little non-image data. Due to an imbalance between the number of images and the number of non-image data, face recognition accuracy may be reduced when face recognition is performed using non-image data. To prevent this, the DNN is made to perform learning twice. In the first learning, the DNN is made to learn representative vectors using only images. In the second learning, the representative vectors are fixed, and the DNN is made to learn data other than images. The first learning of the DNN is performed by the learning process shown in FIG. 6. The second learning of the DNN is performed by the learning process shown in FIG. 8, which will be described later.
[0078] FIG. 8 is a flowchart showing the procedure of the learning process according to the second embodiment.
[0079] In the first DNN training, a set of only a single image (first input dataset) is used to train the DNN on feature conversion specialized for only a single image. Here, a single image is an RGB image showing the texture of a face input to the DNN, and is not a set of multiple images or an image showing depth information. In the second DNN training, a dataset other than images (second input dataset) is used to train the DNN of the second conversion unit 206 on feature conversion specialized for data other than images.
[0080] The details of the first DNN learning have been described in the first embodiment. However, the first input data set is only a single image, and the conversion parameters of the first conversion unit 205 are only for a single image. Hereinafter, the second DNN learning will be described with reference to FIG. 8.
[0081] In S801, the conversion parameters obtained by duplicating the conversion parameters of the DNN of the first conversion unit 205 are set as the initial values of the conversion parameters of the DNN of the second conversion unit 206.
[0082] The processes of S802 to S808 are the same as the processes of S602 to S609 in Fig. 6. However, as in S607 in Fig. 6, the representative vectors V1 to V nThe update process is not performed in Fig. 8. In other words, the value of the representative vector saved in S610 in Fig. 6 is used in Fig. 8. This allows DN to learn the feature amount extracted from data other than images so as to approach the representative vector learned using a single image.
[0083] In S808, the second conversion unit 206 determines whether the learning has converged, for example, based on whether the loss value is equal to or less than a threshold. If the loss value is equal to or less than the threshold (Yes in S808), the second conversion unit 206 proceeds to S809. On the other hand, if the loss value is not equal to or less than the threshold (No in S808), the second conversion unit 206 returns to S802.
[0084] In S809, the second conversion unit 206 saves the parameters of the DNN and ends the process. Note that the value of the representative vector is used only during learning of the DNN, and the value of the representative vector is not used during face matching.
[0085] Although the above-mentioned method for learning DNN using the representative vector has been described, DNN learning may be performed using a method that does not use the representative vector. For example, in the first learning process, in S606, the second conversion unit 206 calculates only intra-class and inter-class loss values, and generates a feature space without performing the process of S607. Then, in the second learning process, in S602, the second input unit 202 receives a first dataset that is paired with the second dataset. Then, the second conversion unit 206 performs feature conversion on the first dataset in the same manner as the second dataset, calculates losses based on the first feature and the second feature, and adjusts parameters using the backpropagation method.
[0086] <Effects of the second embodiment> In the first DNN training, the first conversion unit 205 can generate a fully trained face feature space and representative vector by training the DNN using a large number of single images. As a result, in the second DNN training, the second conversion unit 206 can train the DNN using a small amount of data other than single images.
[0087] <Third embodiment> When an imbalance in the number of data between images and non-image data occurs, the accuracy of face recognition is significantly reduced when face recognition is performed using a small number of images or non-image data. In order to prevent a reduction in the accuracy of face recognition, the enrollment face data and the matching face data are each set to either an image or non-image data having a sufficient number of data.
[0088] For example, the first input unit 201 accepts only two-dimensional RGB images, and the second input unit 202 accepts only sets of black-and-white images and depth information. In this case, the first conversion unit 205 trains the DNN using only RGB images including faces facing forward. On the other hand, the second conversion unit 206 trains the DNN using black-and-white images and depth information including faces facing in various directions. Then, the first conversion unit 205 calculates the loss between the RGB images and the sets of black-and-white images and depth information for the same ID. Here, the first conversion unit 205 does not calculate the loss value between RGB images and between sets of black-and-white images and depth information.
[0089] <Effects of the third embodiment> When there is sufficient learning data for each of the RGB image received by the first input unit 201 and the set of the black-and-white image and depth information received by the second input unit 202, the first conversion unit 205 and the second conversion unit 206 can cause the DNN to perform learning limited to the combination of the learning data described above. At that time, the RGB image parameters of the DNN of the first input unit 201 are adjusted so that the feature amount of the RGB image can be easily matched with the feature amount of the set of the black-and-white image and depth information. In addition, the black-and-white image and depth information set parameters of the DNN of the second input unit 202 are adjusted so that the feature amount of the set of the black-and-white image and depth information can be easily matched with the feature amount of the RGB image. This allows the image processing device 10 to perform highly accurate face recognition based on a limited combination of images and data other than images.
[0090] <Fourth embodiment> A person other than the person registered in the face recognition system may pass the face recognition by using a non-biometric object such as a photograph of the person. This behavior is called "spoofing" and is likely to occur in face recognition systems that use two-dimensional images to recognize faces. To prevent spoofing, a judgment device that judges spoofing is provided separately from the image processing device that only performs face recognition. Specifically, the judgment device acquires a near-infrared image in addition to the visible light image, extracts depth information from the near-infrared image, and judges whether the person in the image is a biological body or not.
[0091] If the determination device determines that the person in the image is a living body, the image processing device performs face authentication of the person. In this way, since the spoofing determination and face authentication are processed in series, it takes a long time to complete all the processing. In addition, since the determination device determines spoofing using only depth information from the near-infrared image, it cannot deal with spoofing when another person is disguised to impersonate the person.
[0092] Therefore, the image processing device 10 of the present invention performs spoofing judgment simultaneously with face authentication when the matching data holds any one of three-dimensional, temperature, and movement information. Here, the three-dimensional information is, for example, information including a pair of stereo images, depth information, normal vectors, point cloud coordinates, and face shape such as curvature. The temperature information is, for example, the temperature of the face measured when photographed with a thermal camera or the like. The movement information is, for example, a video and optical flow including information on the movement of an object. This allows spoofing judgment and face authentication to be performed simultaneously, thereby shortening the time required for all processing. In addition, the image processing device 10 can further improve the accuracy of spoofing judgment by using information other than depth information. The image processing device 10 can also use texture features such as skin texture and color as features of two-dimensional images for spoofing judgment. Therefore, the image processing device 10 can judge spoofing by another person with high accuracy even if the other person is disguised to impersonate the person.
[0093] FIG. 9 is a block diagram illustrating an example of a functional configuration of the third embodiment.
[0094] The image processing device 10 includes a first input unit 201 , a second input unit 202 , a confirmation unit 203 , a storage unit 204 , a first conversion unit 205 , a second conversion unit 206 , a matching unit 207 , and a determination unit 901 .
[0095] <Input data validity determination phase> FIG. 10 is a flowchart showing the procedure of the matching process according to the third embodiment.
[0096] The processes of S1001 to S1009 are the same as those of S101 to S109 in Fig. 4 of the first embodiment. The type of first data accepted by the first input unit 201 in S1001 is not particularly limited. The second data accepted by the second input unit 202 in S1005 includes at least one of three-dimensional face shape information, temperature information, and movement information.
[0097] In S1110, the determination unit 901 determines the validity of the second data based on whether the likelihood based on the feature amount of the second data and the correct answer data exceeds a threshold. Here, validity is an index indicating whether the second data is valid data or not when the matching unit 207 matches the feature amount of the first data with the feature amount of the second data. In other words, when the likelihood is less than the threshold, the determination unit 901 determines the validity of the second data to be "invalid". On the other hand, when the likelihood exceeds the threshold, the determination unit 901 determines the validity of the second data to be "valid".
[0098] The determination unit 901 includes, for example, a DNN. The DNN includes a fully connected layer and a sigmoid function, and outputs a probability (i.e., likelihood) indicating that the second data is valid and a probability indicating that the second data is invalid. At this time, the determination unit 901 determines that the second data is valid when the probability indicating that the second data is valid is greater than the probability indicating that the second data is invalid. On the other hand, the determination unit 901 determines that the second data is invalid when the probability indicating that the second data is invalid is greater than the probability indicating that the second data is valid. Note that the configuration of the DNN is not limited to the above, and the DNN may include a convolution layer, an activation function other than the above, and a softmax function. The number of layers constituting the DNN may be three or more.
[0099] <Learning process phase> FIG. 11 is a flowchart showing the procedure of the learning process according to the third embodiment.
[0100] The processing of S1101 to S1106 is similar to the processing of S601 to S606 in the first embodiment shown in Fig. 6. Note that the first input data set input in S1102 includes at least one of three-dimensional face shape information, temperature information, and movement information.
[0101] In S1107, the determination unit 901 converts the first feature amount into a probability indicating that the first input data set is valid (valid probability) and a probability indicating that the first input data set is invalid (invalid probability) using the DNN. Here, the determination unit 901 includes, for example, a fully connected layer and a sigmoid function. The valid probability and the invalid probability are each expressed as a real number between 0 and 1, so that the sum of the valid probability and the invalid probability is 1.
[0102] In S1108, the determination unit 901 calculates a loss value based on the valid probability or invalid probability and the correct answer data. Here, the correct answer data is data represented by 0 or 1, and the loss value is calculated using a binary cross-entropy error. Note that the loss value is not limited to this, and may be a normal cross-entropy error.
[0103] In S1109, the determination unit 901 sums up the loss value calculated in S1106 and the loss value calculated in S1108. Here, the sum may be the sum of the respective loss values, or may be the sum of weighted loss values.
[0104] The processes of S1110 and S1111 are similar to the processes of S207 and S208 in FIG. 6 of the first embodiment.
[0105] In S1112, the image processing device 10 updates the parameters of the first conversion unit 205 and the DNN of the determination unit 901 so as to reduce the calculated loss value. The update method is an error backpropagation method that is common in DNNs. As a result, the DNN of the determination unit 901 is improved so that it can determine spoofing.
[0106] In S1113, the image processing device 10 determines whether the learning of the DNN has converged based on, for example, whether the loss value is equal to or less than a threshold. If the loss value is equal to or less than the threshold (Yes in S1113), the image processing device 10 determines that the learning of the DNN has converged, and the process proceeds to S1114. On the other hand, if the loss value is not equal to or less than the threshold (No in S1113), the image processing device 10 determines that the learning of the DNN has not converged, and the process returns to S1102.
[0107] In S1114, the storage unit 204 stores the representative vectors V1 to V n The value of is stored.
[0108] In S1115, the memory unit 204 stores the conversion parameters of the DNN of the first conversion unit 205.
[0109] In S1116, the memory unit 204 stores the conversion parameters of the DNN of the determination unit 901.
[0110] <Effects of the Fourth Embodiment> When the face data for matching has at least one of three-dimensional face shape information, temperature information, and movement information, the image processing device is able to perform face authentication and impersonation determination in parallel, thereby further improving the accuracy of face authentication.
[0111] (Other Examples) The present invention can also be realized by a process in which a program for implementing one or more of the functions of the above-described embodiments is supplied to a system or device via a network or a storage medium, and one or more processors in a computer of the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) that implements one or more of the functions.
[0112] The disclosure of this specification includes the following image processing device, method, and program. (Item 1) an extraction means for extracting a first feature from first data of a first modality including information of a registered first object, and extracting a second feature from second data of a second modality different from the first modality including information of a second object to be matched; a determination means for determining whether or not the first object and the second object are the same object based on the first feature and the second feature, The extraction means has been trained so that the first feature and the second feature are similar when the first object and the second object are the same. 13. An image processing device comprising: (Item 2) a selection unit for selecting parameters of the extraction unit corresponding to each of the first modal and the second modal, the extraction means extracts the first feature from the first data based on a parameter corresponding to the first modal, and extracts the second feature from the second data based on a parameter corresponding to the second modal; the determination means determines that the first object and the second object are the same when a similarity between the first feature and the second feature is equal to or less than a threshold value; 2. The image processing device according to item 1, (Item 3) When the degree of similarity is not equal to or less than the threshold value, the determination means determines that the first object and the second object are not identical. 3. The image processing device according to item 2, (Item 4) a validity determination means for determining whether the second data is valid when the second data has predetermined information; the validity determination means determines that the second data is valid when a likelihood of the second object based on the second features and the ground truth data exceeds a threshold value. 4. The image processing device according to any one of items 1 to 3, (Item 5) The validity determination means determines that the second data is not valid when the likelihood does not exceed a threshold value. 5. The image processing device according to item 4, (Item 6) The predetermined information is at least one of three-dimensional shape information, temperature information, and movement information of the second object. 5. The image processing device according to item 4, (Item 7) The first data and the second data include at least one of a two-dimensional RGB image, an image having three-dimensional shape information, a stereo image, a depth image, and a monochrome image. the first object and the second object are human faces; 7. The image processing device according to any one of items 1 to 6, (Item 8) an extraction means for extracting a third feature from third data of a third modal including information of a third object, and extracting a fourth feature from fourth data of a fourth modal different from the third modal including information of a fourth object; an update means for updating a third parameter corresponding to the third modal and a fourth parameter corresponding to the fourth modal based on the third feature and the fourth feature, when the third object and the fourth object are identical, the updating means updates the third parameter and the fourth parameter so that the third feature and the fourth feature become similar to each other; A learning device characterized by: (Item 9) the updating means updates the third parameter based on an intra-class similarity between the third feature and a third representative vector representing a representative feature of the third object, and an inter-class similarity between the third feature and a fourth representative vector representing a representative feature of the fourth object. 9. The learning device according to item 8, (Item 10) the updating means further updates the third representative vector based on the intra-class similarity and the inter-class similarity. 10. The learning device according to item 9, (Item 11) The extraction means extracts a fifth feature from the fourth data based on the third parameter updated by the update means; and the updating means updates the fourth parameter based on an intra-class similarity between the fifth feature and the fourth representative vector and an inter-class similarity between the fifth feature and the third representative vector. 11. The learning device according to item 10, (Item 12) The number of the fourth data is less than the number of the third data. 12. A learning device according to any one of items 8 to 11, (Item 13) the third data is an RGB image; The fourth data is a black and white image and depth information. the third object and the fourth object are human faces; 13. The learning device according to item 12, (Item 14) an extraction step in which an extraction means of the image processing device extracts a first feature from first data of a first modality including information of a registered first object, and extracts a second feature from second data of a second modality different from the first modality including information of a second object to be matched; a determination step in which a determination means of the image processing device determines whether or not the first object and the second object are the same based on the first feature and the second feature, In the extraction step, when the first object and the second object are identical, the first feature and the second feature have been trained to be similar to each other. A method comprising: (Item 15) A program for causing a computer to execute each step of a method, the method comprising: an extraction step in which an extraction means of the image processing device extracts a first feature from first data of a first modality including information of a registered first object, and extracts a second feature from second data of a second modality different from the first modality including information of a second object to be matched; a determination step in which a determination means of the image processing device determines whether or not the first object and the second object are the same based on the first feature and the second feature, In the extraction step, when the first object and the second object are identical, the first feature and the second feature have been trained to be similar to each other. A program characterized by:
[0113] The invention is not limited to the above-described embodiments, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]
[0114] 10: Image processing device 101:CPU 102:ROM 103:RAM 104: Storage section 105: Input section 106: Display section 107: Communications Department 201: First input section 202: Second input section 203: Verification department 204: Storage section 205: First conversion section 206: Second conversion section 207: Matching section
Claims
1. an extraction means for extracting a first feature from first data of a first modality including information on a registered first object, and extracting a second feature from second data of a second modality different from the first modality including information on a second object to be matched; a determination means for determining whether the first object and the second object are the same based on the first feature and the second feature, the extraction means has been trained so that, when the first object and the second object are the same, the first feature and the second feature are similar to each other.
1. An image processing device comprising:
2. a selection unit for selecting parameters of the extraction unit corresponding to the first modal and the second modal, respectively; the extraction means extracts the first feature from the first data based on a parameter corresponding to the first modal, and extracts the second feature from the second data based on a parameter corresponding to the second modal; the determination means determines that the first object and the second object are the same when the similarity between the first feature and the second feature is equal to or less than a threshold value; 2. The image processing device according to claim 1, wherein:
3. the determination means determines that the first object and the second object are not identical when the similarity is not equal to or less than the threshold value; 3. The image processing device according to claim 2.
4. a validity determination means for determining whether the second data is valid when the second data contains predetermined information; the validity determination means determines that the second data is valid when the likelihood of the second object based on the second features and the ground truth data exceeds a threshold.
2. The image processing device according to claim 1, wherein:
5. the validity determination means determines that the second data is not valid when the likelihood does not exceed a threshold value.
5. The image processing device according to claim 4.
6. the predetermined information is at least one of information on a three-dimensional shape of the second object, temperature information, and movement information; 5. The image processing device according to claim 4.
7. the first data and the second data include at least one of a two-dimensional RGB image, an image having three-dimensional shape information, a stereo image, a depth image, and a monochrome image; the first object and the second object are human faces; 7. The image processing device according to claim 1, wherein the image processing device is a computer.
8. extraction means for extracting a third feature from third data of a third modal including information on a third object, and extracting a fourth feature from fourth data of a fourth modal different from the third modal including information on a fourth object; updating means for updating a third parameter corresponding to the third modal and a fourth parameter corresponding to the fourth modal based on the third feature and the fourth feature, respectively; when the third object and the fourth object are identical, the updating means updates the third parameter and the fourth parameter, respectively, so that the third feature and the fourth feature become similar; A learning device characterized by:
9. the updating means updates the third parameter based on an intra-class similarity between the third feature and a third representative vector representing a representative feature of the third object, and an inter-class similarity between the third feature and a fourth representative vector representing a representative feature of the fourth object.
9. The learning device according to claim 8.
10. the updating means further updates the third representative vector based on the intra-class similarity and the inter-class similarity.
10. The learning device according to claim 9.
11. the extracting means extracts a fifth feature from the fourth data based on the third parameter updated by the updating means; the updating means updates the fourth parameter based on an intra-class similarity between the fifth feature and the fourth representative vector and an inter-class similarity between the fifth feature and the third representative vector. The learning device according to claim 10 .
12. the number of the fourth data is less than the number of the third data; The learning device according to claim 11 .
13. the third data is an RGB image, the fourth data is a black and white image and depth information; the third object and the fourth object are human faces; 13. The learning device according to claim 8, wherein the learning device is a learning device for learning a plurality of learning operations.
14. an extraction step in which an extraction means of the image processing device extracts a first feature from first data of a first modality including information on a registered first object, and extracts a second feature from second data of a second modality different from the first modality including information on a second object to be matched; a determination step in which a determination means of the image processing device determines whether or not the first object and the second object are the same based on the first feature and the second feature, the extraction means has been trained so that, when the first object and the second object are the same, the first feature and the second feature are similar to each other. A method characterized by:
15. A program for causing a computer to execute each step of a method, the method comprising: an extraction step in which an extraction means of the image processing device extracts a first feature from first data of a first modality including information on a registered first object, and extracts a second feature from second data of a second modality different from the first modality including information on a second object to be matched; a determination step in which a determination means of the image processing device determines whether or not the first object and the second object are the same based on the first feature and the second feature, the extraction means has been trained so that, when the first object and the second object are the same, the first feature and the second feature are similar to each other. A program characterized by: