A determination device, a mobile terminal including the determination device, and a program for the determination device.
The determination device uses a single image capture with specific color light and neural networks to quickly and securely identify real faces, addressing the inefficiencies and security issues of existing face recognition technologies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ELEMENTS INC
- Filing Date
- 2024-11-26
- Publication Date
- 2026-04-17
AI Technical Summary
Existing face recognition technologies require multiple imaging cycles, are inconvenient due to method switching, and struggle with preventing unauthorized use.
A determination device that captures a human face with a single image using a light-emitting device emitting light lacking a specific color, analyzes the imaging data to distinguish between real faces and photographs, and employs convolutional neural networks to enhance security and accuracy.
Enables rapid and reliable identification of real faces, preventing misuse by distinguishing between live individuals and photographic images, enhancing security through unpredictable light emissions and advanced neural network analysis.
Smart Images

Figure 0007847635000001 
Figure 0007847635000002 
Figure 0007847635000003
Abstract
Description
Technical Field
[0001] The present invention relates to a determination device for determining a human face, a mobile terminal including the determination device, and a program for the determination device.
Background Art
[0002] Conventionally, research and development have been carried out on a determination device for determining a human face. For example, Patent Document 1 (Japanese Patent Application Laid-Open No. 2007-148968) discloses a face recognition technology that can set security levels of different intensities according to the purpose of use without forcing the user to perform complicated operations.
[0003] The face recognition device described in Patent Document 1 (Japanese Patent Application Laid-Open No. 2007-148968) includes a storage unit that acquires and stores feature amounts from a face image of a registrant, an authentication unit that acquires a face feature amount from an input image and performs authentication to determine whether it matches the registrant stored in the storage unit, a security setting unit that allows the user to set the security strength for each purpose of use, and a control unit that restricts storage in the storage unit when the face image of the registrant does not satisfy predetermined shooting conditions based on the security strength, and changes the content of the authentication process performed by the storage unit and the authentication unit based on the security strength.
[0004] In addition, Patent Document 2 (Japanese Patent Application Laid-Open No. 2004-260240) discloses a mobile phone in which a security mode can be set according to the level of awareness of security for each user.
[0005] The mobile phone described in Patent Document 2 (Japanese Patent Application Laid-Open No. 2004-260240) is a mobile phone having a plurality of security execution units that execute the security function of the mobile phone and a selection unit that selects one security execution unit from the plurality of security execution units.
[0006] Furthermore, Patent Document 3 (Japanese Unexamined Patent Publication No. 2004-80080) discloses a mobile phone with a personal authentication function that allows setting of security levels according to usage patterns.
[0007] The mobile phone with personal authentication function described in Patent Document 3 (Japanese Patent Publication No. 2004-80080) comprises: an authentication method selection storage unit that selects at least one personal authentication method from a plurality of personal authentication methods and stores the selected personal authentication method; an information storage unit that stores input information used for the personal authentication method stored in the authentication method selection storage unit; a feature extraction unit that extracts feature information as authentication information from the input information stored in the information storage unit; a feature recording unit that records the user's feature information in advance; and a feature authentication unit that performs personal authentication of the user by comparing the feature information recorded in advance in the feature recording unit with the authentication information extracted by the feature extraction unit.
[0008] Furthermore, Patent Document 4 (Japanese Patent Publication No. 2002-112340) discloses a mobile device user authentication system and method thereof, which allows the user to freely change the authentication means according to the purpose of use, and therefore makes it possible to set a security level that satisfies the user for each purpose of use.
[0009] The mobile device authentication system described in Patent Document 4 (Japanese Patent Publication No. 2002-112340) is an authentication system that authenticates the user of a mobile device by communicating with a base station, and comprises a plurality of user authentication means for authenticating that the user of the mobile device is the person in question, an authentication setting means for setting the correspondence between the user authentication means and the purpose of use of the mobile device by the person in question, an authentication information input means for receiving authentication information from the mobile device for judgment by the user authentication means, and a user verification means for comparing the authentication information input by this means with pre-stored authentication information to confirm whether or not the person is the person in question, and changes the user authentication means according to the purpose of use of the user based on the correspondence set by the authentication setting means.
[0010] Furthermore, Patent Document 5 (Japanese Patent Publication No. 2000-215308) discloses how to improve convenience by making it possible to adjust the balance between the performance and comfort of personal authentication, as well as its resilience to environmental changes, according to the situation.
[0011] The biometric authentication device described in Patent Document 5 (Japanese Patent Publication No. 2000-215308) is a biometric authentication device comprising: a biometric information receiving means that acquires the biometric information of a person to be authenticated and extracts the characteristics of the person to be authenticated from the input information to form biometric characteristic information; and an identification means that compares the biometric characteristic information with registration information prepared in advance by category to determine the category to which the person to be authenticated belongs. The biometric authentication device is provided with multiple biometric information receiving means and identification means, and when performing biometric authentication using an authentication logic configured by combining them, it has a mode management means that can control the biometric information receiving means and identification means by selecting one from a plurality of authentication logic modes. [Prior art documents] [Patent Documents]
[0012] [Patent Document 1] Japanese Patent Publication No. 2007-148968 [Patent Document 2] Japanese Patent Publication No. 2004-260240 [Patent Document 3] Japanese Patent Publication No. 2004-80080 [Patent Document 4] Japanese Patent Publication No. 2002-112340 [Patent Document 5] Japanese Patent Publication No. 2000-215308 [Overview of the project] [Problems that the invention aims to solve]
[0013] The technology described in Patent Document 1 discloses a technology that allows setting different security levels of intensity depending on the purpose of use, but this presents the problem of requiring multiple imaging cycles. Furthermore, the technologies described in Patent Documents 2, 3, and 4 have the problem of low convenience because they require switching or selecting an authentication method for authentication. Furthermore, the technology described in Patent Document 5 presents the problem of difficulty in preventing unauthorized use.
[0014] The main objective of the present invention is to provide a determination device, a portable terminal including the determination device, and a program for the determination device that can prevent misuse and identify a person with a single image capture. Another object of the present invention is to provide a determination device, a portable terminal including the determination device, and a program for the determination device that can increase processing speed, prevent misuse with a single image, and reliably identify a person. [Means for solving the problem]
[0015] (1) The one-sided determination device is a determination device for determining a human face, and includes an imaging device that takes an image of a human face only once, a light-emitting device that emits light lacking a specific color when the imaging device takes an image, an analysis device that analyzes the imaging data from the imaging device, and a control device that controls the imaging device, the light-emitting device, and the analysis device. The control device decomposes the imaging data into components of each color using the analysis device, and determines whether the imaging data is imaging data of a real person's face based on the difference between the component data of the color emitted by the light-emitting device and the component data of the color that the light-emitting device does not emit.
[0016] In this case, it is possible to reliably identify a person's face. In particular, since it is not necessary to take multiple images, it is possible to identify a large number of people in a short amount of time. Furthermore, the ability to reliably identify a human face here means being able to determine whether it is the face of a real person or a human face captured in a photograph and printed. Furthermore, light lacking a specific color is light lacking the spectrum and frequency of that specific color, and does not necessarily mean white light. For example, light lacking a specific color may be such that among RGB, G and B, only G, or only B is the dominant color with respect to R. Therefore, it includes both a state where no red (R) is included at all and a state where a little red (R) is included.
[0017] (2) The determination device according to the second invention is a determination device according to one aspect, and the control device may determine whether the imaging data is the imaging data of a real person's face at a specific part of the person's face.
[0018] In this case, it is possible to determine a real person's face using only a part of the face, that is, a characteristic part which is a specific part.
[0019] (3) The determination device according to the third invention is the determination device according to the second invention, and the specific part may be at least any one of the eyes, nose, and cheeks.
[0020] In this case, it is possible to determine a person's face not by the whole face but by characteristic parts such as the nose, mouth, eyes, cheeks, and contour.
[0021] (4) The determination device according to the fourth invention is a determination device according to one aspect to the fourth invention, and the control device may instruct the light emitting device to emit light lacking a specific color a plurality of times, or to emit a plurality of lights lacking different colors from each other.
[0022] In this case, since the user cannot recognize which color - lacking light emission the determination is being made in response to by receiving various light emissions from the light emitting device, security can be enhanced and unauthorized use can be prevented. For example, in an RGB color code, by emitting light in the order of #00FF00, #008000, #0000FF, #FFFF00, #00080, #00FFFF, and #800080, and using only the data captured at the moment of #0000FF, the user will not know which emission timing was used for the determination, thus enhancing security. Note that the above RGB color code represents the intensity of light in the order of RGB from left to right, each represented by a two-digit hexadecimal number.
[0023] (5) The determination device according to the fifth invention, in one aspect, the emission of light lacking a specific color may consist of one of the following: light lacking one color from RGB, light lacking one color from the XYZ color system, or light lacking one color from the Munsell color system.
[0024] In this case, the emission of light lacking a specific color consists of one of the following: light lacking one color from RGB, light lacking one color from the XYZ color system, or light lacking one color from the Munsell color system. Therefore, users cannot recognize the details, thus enhancing security and preventing misuse.
[0025] (6) The determination device according to the sixth invention, in one aspect, the determination device may determine that the image data is image data of a real person's face if, from the data decomposed into components of each color, the brightness of the eye area is greater than the brightness of the surrounding area, based on the data obtained by subtracting the component data of colors that the light-emitting device does not emit from the component data of colors that the light-emitting device emits.
[0026] In this case, it is possible to reliably determine that the image data represents the face of a real person.
[0027] (7) The determination device according to the seventh invention, from one perspective, the determination device, in the determination device according to the seventh invention, steps as follows: 1) A step in which, among the data decomposed into components for each color, the component data of colors that the light-emitting device does not emit light from the component data of colors that the light-emitting device emits light. 2) A step to find the peak value of the subtracted data, 3) A step in which the subtracted data is divided into first half data and second half data, using the peak value as the boundary. 4) A step to reverse the order of the second half of the data so that the boundary comes last. 5) A step of calculating the Pearson product-moment correlation coefficient between the first half of the data and the second half of the data with the order reversed. 6) The step of determining whether the image data is image data of a real person's face if Pearson's product-moment correlation coefficient is 0.5 or higher may be used to determine whether the image data is image data of a real person's face.
[0028] In this case, it is possible to easily and quickly determine whether or not the image data is of a real person's face. It is assumed that the first and second halves of the data have the same amount of data. If the peak value is not the central value, the correlation coefficient is calculated from the peak value to a predetermined amount of data.
[0029] (8) The determination device according to the eighth invention is a determination device that follows one aspect, which calculates the difference between the component data of a color emitted by the light-emitting device and the component data of a color not emitted by the light-emitting device for the part of the person's face and the part of the background, and compares the magnitude of the difference for the part of the person's face with the magnitude of the difference for the part of the background to determine whether or not the image data is image data of the face of a real person.
[0030] In the case of image data of a real person's face, the color component emitted by the light-emitting device has high brightness in the face area (close to the light-emitting device) due to reflection from the face, and low brightness in the background area (outside the face). In contrast, the color component not emitted by the light-emitting device shows little difference between the face area and the background area. Therefore, the difference between the color component data emitted by the light-emitting device and the color component data not emitted by the light-emitting device is greater in the face area than in the background area. On the other hand, in the case of image data such as photographs, which do not represent the faces of real people, the light from the light-emitting device is reflected from the background, so the difference between the color components that the device emits and the color components that it does not emit is small between the face and the background. Therefore, the magnitude of the difference between the color component data that the device emits and the color component data that it does not emit is small between the face and the background. Based on the above, by comparing the magnitude of the difference between the component data of the color emitted by the light-emitting device and the component data of the color not emitted by the light-emitting device, specifically the difference in the human face area and the difference in the background area, if the ratio or difference in the magnitude of the difference is greater than or equal to a predetermined value, it can be determined that the image data is image data of a real person's face.
[0031] (9) The determination device of the invention according to other aspects is a determination device for determining a human face, and includes an imaging device that takes an image of a human face, a light-emitting device that emits light of a specific color when the imaging device takes an image, an analysis device equipped with a convolutional neural network that analyzes the image data from the imaging device, and a control device that controls the imaging device, the light-emitting device, and the analysis device, wherein the convolutional neural network is trained to determine whether the input data is image data of a real person's face or not, using image data of a real person's face and image data of a photograph of the same person's face when the light-emitting device is emitting light, and the determination device takes an image of a human face when the light-emitting device is emitting light, inputs the image data extracted from the area around the face into the convolutional neural network, and determines whether the captured human face is the face of a real person or not.
[0032] When the light-emitting device is emitting light, there are differences in brightness between image data of a real person's face of a specific color and image data of a photograph of the same person's face. These differences include differences in the brightness of the eyes, the nose, and the face and background. By comparing these differences with the brightness of image data of a non-emitting color, it is possible to determine whether the captured face is that of a real person. In a convolutional neural network, by training it using image data of real people's faces and image data of the same person's face, it can comprehensively determine whether or not a face is that of a real person, including the differences in the eyes, nose, background, etc. However, to improve the accuracy of the judgment, it is necessary to appropriately select the width (number of neurons), number of layers, and number of samples to be trained on in the convolutional neural network. For example, a width of 64 channels of 3x3 filters, 10 to 50 layers, and 1000 or more samples are desirable. Alternatively, publicly available pre-trained models such as VGG, ResNet, or denseNet may be used through transfer learning to determine whether the captured human face is a real human face or not.
[0033] (10) The determination device according to the tenth invention is a determination device according to the other aspects, comprising a first convolutional neural network trained using imaging data when the light-emitting device is emitting light and a second convolutional neural network trained using imaging data when the light-emitting device is not emitting light. The determination device may input the image data taken when the light-emitting device is emitting light into a first convolutional neural network, and the image data taken when the light-emitting device is not emitting light into a second convolutional neural network, and determine whether the captured human face is the face of a real person or not based on the determination result of the first convolutional neural network and the determination result of the second convolutional neural network.
[0034] In this case, combining the results of the assessment with the imaging data when the light-emitting device is not emitting light has the following two effects, enabling more accurate assessment. Firstly, when imaging an actual human being, the image data from when the person is emitting light includes reflections due to the contours of the eyes and face, whereas the image data from when the person is not emitting light does not include reflections due to the contours of the eyes and face. On the other hand, when a photograph is taken while the light-emitting device is emitting light, the image data taken while the device is emitting light shows that not only the contours of the eyes and face, but the entire photograph reflects the light from the device, whereas the image data taken while the device is not emitting light shows only the reflection from the contours of the eyes and face. From these differences, it is possible to reliably determine whether the face of the person in the photograph is that of a real person, even when the photograph is taken while the light-emitting device is emitting light. Furthermore, when a photograph is taken, in addition to the reflections caused by the contours of the eyes and face mentioned above, there is an unnaturalness that arises from the fact that the photograph was taken. The first convolutional neural network, which was trained using image data taken while the subject was emitting light, naturally recognizes and judges this unnaturalness as a difference. However, by using a second convolutional neural network, which was trained using image data taken while the subject was not emitting light, it is possible to more reliably determine whether the captured human face is the face of a real person or not.
[0035] (11) The determination device according to the 11th invention is a determination device according to the 10th invention, wherein the analysis device further comprises a third convolutional neural network trained using image data of a real person's face and image data of a display of the same person's face when the light-emitting device is emitting light, and a fourth convolutional neural network trained using image data of a real person's face and image data of a display of the same person's face when the light-emitting device is not emitting light. The determination device may input image data when the light-emitting device is emitting light into the first and third convolutional neural networks, input image data when the light-emitting device is not emitting light into the second and fourth convolutional neural networks, and determine whether the captured person's face is the face of a real person based on the determination results of the first to fourth convolutional neural networks.
[0036] Recently, it has become necessary to distinguish between image data of faces displayed on LCD or mobile displays (also known as display attack data) and image data of real people's faces, rather than printed photographs of faces. In this case, training a convolutional neural network using image data of real people's faces and image data of the same person's face displayed on a screen can improve the accuracy of the discrimination. Then, by comprehensively determining whether or not it is the face of a real person based on the judgment results of the first and second convolutional neural networks trained using image data of printed photographs of faces, and the judgment results of the third and fourth convolutional neural networks trained using image data of screen displays, it is possible to make a more accurate judgment even when the input data includes display attack data.
[0037] (12) The determination device according to the 12th invention is, in one aspect, a determination device according to the 11th invention, further including a notification device, wherein the control device may cause the imaging device to capture a person's face before performing the determination, and the notification device may notify the person of their standing position so that the size of the person's face is suitable for determination.
[0038] In this case, the notification device can be used to adjust the position of the person standing. Alternatively, the imaging device may recognize the face and automatically enlarge or reduce the image so that the captured face reaches a predetermined size. Alternatively, the notification device may include a display device that shows the captured image and a frame on the display device, and automatically capture an image when the size and position of the face match the frame.
[0039] (13) The determination device according to the 13th invention is, in one aspect, the determination device according to the 12th invention, wherein the control device may cause the imaging device to capture a person's face before performing the determination, and from the captured imaging data, select the optimal light lacking a specific color and emit it from the light-emitting device.
[0040] In this case, the light emitted from the light-emitting device, which lacks a specific color, can be optimized. For example, if a person is wearing contact lenses of a specific color, the device can emit light in a way that does not produce the color of those contact lenses. As a result, accurate judgment can be made.
[0041] (14) Other aspects of the portable terminal include a determination device according to the 13th invention, starting from one aspect.
[0042] In this case, since the mobile device includes a detection device, human detection can be performed easily and reliably. As a result, the misuse of the mobile device can be prevented.
[0043] (15) Furthermore, a program for a judgment device relating to other aspects is a program for a judgment device that determines a human face, and includes an imaging process that takes an image of a human face once, a light emission process that emits light lacking a specific color when the imaging process takes an image, an analysis process that analyzes the imaging data from the imaging process, and a control process that controls the imaging process, the light emission process and the analysis process, wherein the control process decomposes the imaging data into components of each color by the analysis process and determines whether the imaging data is imaging data of a real person's face according to the difference between the component data of the color emitted by the light emission process and the component data of the color that the light emission process does not emit.
[0044] In this case, it is possible to reliably identify a person's face. In particular, since it is not necessary to take multiple images, it is possible to identify a large number of people in a short amount of time. Furthermore, the ability to reliably identify a human face here means being able to determine whether it is the face of a real person or a human face captured in a photograph and printed. Furthermore, light lacking a specific color is light lacking the spectrum and frequency of that specific color, and does not necessarily mean white light. [Brief explanation of the drawing]
[0045] [Figure 1] This is a schematic diagram showing an example of the configuration of a mobile terminal including a determination device according to the first embodiment. [Figure 2] This is a schematic diagram showing an example of the configuration of a judgment device. [Figure 3] This is a flowchart showing an example of the control flow of the control unit. [Figure 4] This is a flowchart showing an example of the control flow of the control unit. [Figure 5] This is a schematic diagram showing an example of facial data captured by an imaging device and a light-emitting device. [Figure 6] Figure 5 is a schematic diagram showing an example of data analyzed within an analytical instrument. [Figure 7] Figure 5 is a schematic diagram showing an example of data analyzed within an analytical instrument. [Figure 8] Figure 5 is a schematic diagram showing an example of data analyzed within an analytical instrument. [Figure 9] This is a schematic diagram illustrating an example of the difference in data analyzed within the analytical instrument. [Figure 10] This is a schematic diagram illustrating an example of the difference in data analyzed within the analytical instrument. [Figure 11] The data captured in steps S1 and S2 is obtained from a printed photograph of a face, and then processed from steps S3 to S5. [Figure 12] This figure shows an example of extracting facial features in step S11. [Figure 13] This figure shows an example of data representing only the red (R) part, only the green (G) part, and only the blue (B) part of a facial feature when a flash of light with RGB code #00FFFF is shone onto the face of a real person and photographed. [Figure 14] This figure shows an example of the result of the process in step S12. [Figure 15] This figure shows an example of how Pearson's product-moment correlation coefficient can be calculated. [Figure 16]This figure shows an example of data obtained by re-imaging a paper printout of a facial photograph taken using flash light, in the process of steps S1 and S2. [Figure 17] This figure shows an example of data obtained by re-imaging a paper printout of a facial photograph taken using flash light, in the process of steps S1 and S2. [Figure 18] This figure shows an example of data obtained by re-imaging a paper printout of a facial photograph taken using flash light, in the process of steps S1 and S2. [Figure 19] This flowchart shows another example of the control flow for the control unit. [Figure 20] Here is another flowchart showing a control flow for a control unit. [Figure 21] Figure 2 is a schematic diagram showing another example of the determination device. [Figure 22] This flowchart shows an example of indicating the standing position of a person's face when imaging a face with the imaging device 200 using the determination device 100 shown in Figure 21. [Figure 23] This is a schematic diagram showing an example of a notification device display for indicating the facial position. [Figure 24] This flowchart shows an example of a process in which, before the processing of step S1 shown in Figure 3, a real person's face is captured in advance, and a specific color of light lacking is selected to be emitted from the light-emitting device. [Modes for carrying out the invention]
[0046] [Embodiment] Embodiments of the present invention will be described below with reference to the drawings. In the following description, identical parts are denoted by the same reference numerals. Their names and functions are also the same. Therefore, detailed descriptions of them will not be repeated.
[0047] (Overall configuration of the mobile terminal 900 including the judgment device 100) Figure 1 is a schematic structural diagram showing an example of the overall configuration of a mobile terminal 900 including the determination device 100 according to this embodiment, and Figure 2 is a schematic diagram showing an example of the configuration of the determination device 100.
[0048] As shown in Figure 1, the mobile terminal 900 includes a determination device 100, a display screen 920, and an operation unit 930. Furthermore, the determination device 100 in Figure 2 includes an imaging device 200, a light-emitting device 300, an analysis device 400, and a control unit 500.
[0049] Next, the light emitted from the light-emitting device 300 (hereinafter referred to as flash light) has an RGB color code of #00FFFF. In other words, it is a flash of light from which at least the red light (R) of the three primary colors R, G, and B has been removed. The reason for removing the red light in particular will be explained later. In this embodiment, the flash light uses the RGB color code #00FFFF, but it is not limited to this, and the flash light may use one or more of #FFFF00, #0000FF, or #00FF00. Furthermore, light lacking a specific color can be defined as light in which G and B, G only, or B only are dominant over R in the RGB spectrum. Therefore, this includes both states where red (R) is completely absent and states where a small amount of red (R) is present.
[0050] Figures 3 and 4 are flowcharts showing an example of the control flow of the control unit 500. Figure 5 is a schematic diagram showing an example of face data captured by the imaging device 200 and the light-emitting device 300. Next, Figures 6, 7, and 8 are schematic diagrams showing an example of data analyzed within the analysis device 400 using the data from Figure 5, and Figures 9 and 10 are schematic diagrams showing an example of the difference in data analyzed within the analysis device 400.
[0051] First, as shown in Figure 3, the control unit 500 instructs the imaging device 200 and the light-emitting device 300 to take an image (step S1). In step S1 of this embodiment, the control unit 500 is unaware whether the data captured by the imaging device 200 is the face of a real person or the face of a printed photograph of a person. However, for the purposes of this explanation, we will assume that it is the face of a real person. Furthermore, the control unit 500 instructs the light-emitting device 300 to take images with a predetermined flash of light during imaging (step S2). Furthermore, the user interface may include a system that automatically presses the shutter when the face is in a predetermined position. Alternatively, if the face is not in the predetermined position, the system may notify the user to move closer to the imaging device 200. Here, the predetermined flash light in this embodiment is the light with the color code #00FFFF described above. Alternatively, the system may recognize the face of a real person before imaging, and adjust the predetermined flash light according to the color of the face, eye color, etc.
[0052] (First analysis) The control unit 500 passes the data captured in step S2 to the analysis device 400 and instructs it to perform the first analysis (step S3). Here, as shown in Figure 5, based on the instructions of the control unit 500, the data captured by the light-emitting device 300 using the color code #00FFFF of the flash light will be data with almost no red (R) color. However, since ambient light, such as light from fluorescent lamps, may be present, even if there is no red (R) in the flash light, a small amount of red (R) may be included.
[0053] Next, the control unit 500 uses the analysis device 400 to split the face data shown in Figure 5, received from the imaging device 200, into R-only data, G-only data, and B-only data, respectively (step S4). Figure 6 shows an example of data containing only R, Figure 7 shows an example of data containing only G, and Figure 8 shows an example of data containing only B.
[0054] Next, the control unit 500 causes the analyzer 400 to perform a difference calculation (step S5). Here, the analyzer 400 subtracts the data for R only in Figure 6 from the data for B only in Figure 8, as shown in Figure 9, and subtracts the data for R only in Figure 6 from the data for G only in Figure 9, as shown in Figure 10.
[0055] Here, Figure 11 shows the data obtained by processing a printed facial photograph in steps S1 and S2, followed by processing from steps S3 to S5. This data is known as photoattack data.
[0056] Here, the control unit 500 checks the eye's reflexes based on the data from the analyzer 400 and makes a determination (step S6). For example, if we consider the face of a real person, that is, if a person is actually present and the processing in step S2 is performed, the data for G only and B only will have high brightness in the eye area, but the data for R only will not have high brightness in the eye area. Therefore, as shown in Figures 9 and 10, the BR and GR data will have high brightness only in the eye area of the flash light. On the other hand, as shown in Figure 11, when a face photograph is printed out on paper, the bright areas of the eyes disappear.
[0057] In other words, in the case of a real person's face, the reflection from the eyes is particularly large, but in the case of a printed photograph of a face, in step S2, there is no difference in the reflection of light from the eyes and other parts, so the bright parts of the eyes disappear.
[0058] Based on the above, the control unit 500 can easily determine whether the face data captured by the imaging device 200 is of a real person or a printed photograph of a face on paper.
[0059] (Second analysis) Next, I will explain the second determination. In the first determination, we decided to exclude paper with printed facial photographs. However, even among paper with printed facial photographs, the inventor found that if the same flash light was used for the original facial photograph, there was a possibility that the brighter parts of the eyes would remain visible. Therefore, a second assessment will be conducted.
[0060] First, as shown in Figure 4, the control unit 500 instructs the analysis device 400 to perform a second analysis using the data used in step S6 (step S10). The control unit 500 extracts facial features from the data (step S11). For example, in this embodiment, the nose is extracted. In addition to the nose, if the cheeks are used as a feature, since it is unlikely that the flash light will hit both cheeks evenly, the unevenness of the cheeks may also be used as a feature. In the above embodiment, the human face portion is extracted, but the invention is not limited to this. The outline of the face may be extracted, and the area outside of it, the so-called background portion, may be extracted, and the processing in steps S12 to S15 below may be performed.
[0061] Figure 12 shows an example of extracting facial features in step S11, Figure 13 shows an example of R-only data, G-only data, and B-only data of facial features when a real person's face is illuminated with a flash of light with RGB code #00FFFF and photographed, and Figure 14 shows an example of the result of processing in step S12.
[0062] As shown in Figure 12, the control unit 500 extracts the tip portion P, the left end portion L, and the right end portion R of the nose. The inventors discovered that there is a region in the tip portion P of the nose where specular reflection occurs. As shown in Figure 13, in this embodiment, the data for G only and the data for B only have a bell-shaped luminance distribution at the tip of the nose P, the left end L, and the right end R, while the data for R only does not have a bell-shaped luminance distribution because it is not included in the flash light and is generated by the reflection of background light with a wide angle of incidence.
[0063] Next, the control unit 500 instructs the analyzer 400 to subtract the minimum value of each luminance distribution in the R-only data, G-only data, and B-only data, and subtracts the R-only data from the G-only data, or subtracts the R-only data from the B-only data (step S12). Here, the process in step S12 is to adjust the average value of the data to 0, and a similar method may be used to adjust the average value of each data point to 0.
[0064] As shown in Figure 14, when the data of R only is subtracted from the data of G only as a result of the processing in step S12, a peak value appears.
[0065] Next, as shown in Figure 4, the control unit 500 causes the analyzer 400 to detect the peak value and divides it into the first half of the peak value and the second half of the peak value (step S13). Finally, the control unit 500 instructs the analyzer 400 to calculate the Pearson product-moment correlation coefficient for the two data series obtained by folding the first half data and the second half data (step S14).
[0066] In this embodiment, the data was divided into two halves based on the peak value, but the invention is not limited to this, and the first and second halves may be determined at any point based on the amount of data. For example, each half may be divided into two parts, each equaling 50% of the data amount.
[0067] Figure 15 is a graph obtained by calculating the average value of the GR in the graph of Figure 14, subtracting the average value from the data for each GR, and then dividing the graph into two halves, the first half and the second half, with the center point or peak value as the boundary. The second half of the GR graph is the portion to the right of the center point or peak value of the graph in Figure 14, folded in the opposite direction from the center point. In the case of Figure 15, there is a positive correlation between the first half and the second half of the GR, and the Pearson product-moment correlation coefficient is calculated to be 0.82.
[0068] Figures 16, 17, and 18 show examples of data obtained by printing out a facial photograph taken using flash light and then re-imaging that paper using the processes in steps S1 and S2. Figure 16 shows an example of data showing only R, only G, and only B of the facial feature portion of data obtained by re-imaging a printout of a facial photograph taken using flash light. Figure 17 shows an example of the results of processing step S12 of the facial feature portion of data obtained by re-imaging a printout of a facial photograph taken using flash light. Figure 18 is a graph obtained by calculating the average value of GR in the graph of Figure 17, subtracting the average value from each GR data, and then dividing the graph into the first and second halves with the center point as the boundary. The second half of the GR graph is obtained by folding the part to the right of the center of the graph of Figure 17 in the opposite direction from the center point. In the case of Figure 18, the correlation between the first half of GR and the second half of GR is small, and the Pearson product-moment correlation coefficient is calculated to be 0.08.
[0069] As shown in Figure 4, the control unit 500 causes the analysis device 400 to perform the determination (step S15). Here, the control unit 500 determines that the face data belongs to a real person if the Pearson product-moment correlation coefficient is 0.5 or greater. For example, in the case of Figure 15, the Pearson product-moment correlation coefficient is 0.82. On the other hand, if the Pearson product-moment correlation coefficient is less than 0.5, the control unit 500 determines that the data is from a photograph of a face that was captured using flash light and then printed out and photographed again. For example, in the case of Figure 18, the Pearson product-moment correlation coefficient is 0.08. Therefore, the Pearson product-moment correlation coefficient is calculated between the data from the first half and the data obtained by folding the data from the second half. If the correlation coefficient is 0.5 or higher, it is determined that the captured data is the face of a real person; if the correlation coefficient is less than 0.5, it is determined that the data is the image of a printed photograph of a face that has been photographed again.
[0070] [Embodiments of other determination methods] Next, we will describe another embodiment for determining whether or not the image data is of a real person's face. Figure 19 is a schematic flowchart of an embodiment of another determination method.
[0071] First, the control unit 500 instructs the imaging device 200 and the light-emitting device 300 to perform imaging (step S20). Furthermore, the control unit 500 instructs the light-emitting device 300 to take images with a predetermined flash of light during imaging (step S21). The control unit 500 extracts the inner and outer regions of the facial contour from the data (step S22). Next, similar to S12 in Figure 4, the subtraction value between the component data of the color emitted by the light-emitting device 300 and the component data of the color not emitted by the light-emitting device 300, for example, GR or BR, is calculated (step S23). Furthermore, before performing the subtraction, it is desirable to correct the data so that the minimum value of the component data for the color emitted by the light-emitting device 300 matches the minimum value (or the average values of the average values) of the component data for the color that the light-emitting device 300 does not emit. Next, the subtraction values are compared between the area inside the contour, i.e., the face area, and the area outside of it, i.e., the background area (step S24). Note that when comparing the face area and the background area, for example, the average values of the brightness of each area may be compared. Finally, the comparison result between the face and the background, for example, the ratio or the difference between them, is compared with a predetermined threshold. If the ratio or the difference is greater than the predetermined threshold, it is determined that the image data is image data of a real person's face (step S25).
[0072] In the case of image data of a real person's face, the color component emitted by the light-emitting device 300 has high brightness in the face area (close to the light-emitting device 300) due to reflection from the face, and low brightness in the background area (outside the face). In contrast, the color component that the light-emitting device 300 does not emit has little difference between the face area and the background area. Therefore, the difference between the color component data emitted by the light-emitting device 300 and the color component data that the light-emitting device 300 does not emit is greater in the face area than in the background area. On the other hand, in the case of image data such as photographs, which are not of a real person's face, the light from the light-emitting device 300 is reflected from the background, so the difference between the color components that the light-emitting device 300 emits and the color components that the light-emitting device 300 does not emit is small between the face and the background. Therefore, the magnitude of the difference between the color component data that the light-emitting device 300 emits and the color component data that the light-emitting device 300 does not emit is small between the face and the background. As described above, by comparing the size of the human face portion and the size of the background portion of the difference between the component data of the color emitted by the light-emitting device 300 and the component data of the color not emitted by the light-emitting device 300, it is possible to determine whether the image data is image data of a real person's face.
[0073] [Further embodiments of the determination method] Furthermore, in another embodiment of the determination method, machine learning is used to determine whether the image data is of a real person's face. Specifically, first, a convolutional neural network is trained by inputting multiple images of real people's faces and images of the same person's face (also called photoattack data) (hereafter, the trained network is referred to as the photo-learning network). Subsequently, the image data of faces captured by the imaging device 200 is input to the photo learning network, which is then used to determine whether the input data is image data of a real person's face or image data of a photograph of a face.
[0074] Figure 20 is a schematic flowchart of another embodiment of the determination method. First, the camera captures the person's face when the light-emitting device 300 emits cyan light (hereinafter referred to as cyan data) and the person's face when the light-emitting device 300 does not emit any light (hereinafter referred to as black data) (steps S30, S30'). Next, the facial contour is extracted using the facial outline and characteristic features such as the eyes and lips (steps S32, S32'). Furthermore, the cyan and black data around the face are resized to, for example, 224 x 224 pixels. The pixel values of the resized data may also be normalized to a range between 0 and 1 (steps S33, S33'). Note that the same preprocessing, such as extracting the face contour, resizing the data around the face, and normalization, must be applied to the training data as well. The cyan data of faces captured by the imaging device 200 is input to a convolutional neural network trained using cyan data as input data (hereinafter referred to as the cyan photo training network) and used for judgment. The black data of faces captured by the imaging device 200 is input to a convolutional neural network trained using black data (hereinafter referred to as the black photo training network) and used for judgment (steps S34, S34').
[0075] The system determines whether the image data represents a real person based on the results of the cyan data and the black data. If both the cyan data and black data results indicate a real person, the system may determine that the image data represents a real person. Alternatively, if either the cyan data or black data result indicates a real person, the system may determine that the image data represents a real person. Furthermore, if the results of each determination are output as an intermediate value rather than 0 or 1, the system may use the sum of these results to make the determination (step S35).
[0076] In this embodiment, the determination of whether the image data represents a real person is made based on the determination results of the cyan data and the black data. However, the determination of whether the image data represents a real person may also be made based solely on the determination results of the cyan data. Furthermore, the light-emitting color of the light-emitting device 300 may be magenta, yellow, red, blue, or green instead of cyan.
[0077] The purpose of using the black data determination results is as follows: Firstly, when photographing an actual person, cyan data shows reflections due to the contours of the eyes and face, whereas black data does not. On the other hand, when a photograph is taken with the light-emitting device emitting light, the cyan data shows reflections from the entire photograph, not just the contours of the eyes and face, while the black data shows reflections only due to the contours of the eyes and face. From these differences, it is possible to reliably determine whether the face of the person in the photograph is that of a real person, even when the photograph is taken with the light-emitting device emitting light. Furthermore, when a photograph is taken, in addition to the reflections caused by the unevenness of the eyes and face mentioned above, there is an unnaturalness that arises from the fact that the photograph was taken. The first convolutional neural network, which was trained using cyan data, naturally recognizes and judges this unnaturalness as a difference, but by using a second convolutional neural network, which was trained using black data, it is possible to more reliably determine whether the captured human face is the face of a real person or not.
[0078] Recently, it has also become necessary to distinguish between image data of faces displayed on LCD screens (also called display attack data) and image data of real people's faces, rather than printed photographs of faces. In this case as well, first, the third convolutional neural network is trained by inputting multiple cyan data of real people's faces and cyan data of the same person's face displayed on a screen (cyan display learning network). Next, the fourth convolutional neural network is trained by inputting multiple black data of real people's faces and black data of the same person's face displayed on a screen (black display learning network). Then, the cyan data of faces captured by the imaging device 200 is input to the cyan display learning network, and the black data of faces captured by the imaging device 200 is input to the black display learning network, and the network is made to determine whether the input data is image data of a real person's face or not.
[0079] The input to the detection device 100 may include not only image data of real people's faces and image data of printed face photographs, but also image data of faces displayed on an LCD screen or the like. In order to reliably determine that the image data is of real people's faces from among these inputs, the following procedure is followed. First, the cyan data and black data captured by the imaging device 200 are input to the cyan photo learning network and the black photo learning network, respectively. Next, the cyan data and black data captured by the imaging device 200 are input to the cyan display learning network and the black display learning network, respectively. Furthermore, the output of each of the four learning networks is set to 1 if it is determined to be the face of a real person, and to 0 if it is determined not to be the face of a real person. If the sum of the outputs of the four learning networks is greater than or equal to a predetermined value, it is determined to be the face of a real person. In addition, if the output of each determination is an intermediate value rather than 0 or 1, the determination may be made by adding up those determination results.
[0080] (Configuration of a convolutional neural network) Various configurations are possible for the width (number of neurons) and number of layers of a convolutional neural network. For example, if image data is resized to 224x224 and divided into RGB, the number of input signals will be 224x224x3. The output signals consist of two signals: one that is 1 when it is determined to be a real person's face and 0 when it is not, and the opposite signal. The width and number of hidden layers between these can be selected based on the required detection accuracy and the acceptable circuit size. Recently, a new type of convolutional neural network called ResNet (Deep Residual Network) has been made publicly available as a training model. In this embodiment, for example, an 18-layer ResNet (ResNet-18) may be used through transfer learning to determine whether the captured human face is the face of a real person or not. Alternatively, other publicly available pre-trained models such as VGG and denseNet may also be used through transfer learning to determine whether the captured human face is the face of a real person or not.
[0081] In this embodiment, the control unit 500 performs all the decisions, but this is not limited to this, and the analyzer 400 may have a separate control unit, and the analyzer 400 may perform the decisions.
[0082] [Embodiment of notification device 600] Figure 21 is a schematic diagram showing another example of the determination device 100 shown in Figure 2. The determination device 100 shown in Figure 21 is further equipped with a notification device 600 in addition to the determination device 100 shown in Figure 2. Here, the notification device 600 may have a sound generating device, a display device, or both, or may include any other notification device 600.
[0083] Figure 22 is a flowchart showing an example of indicating the standing position of a person's face when the imaging device 200 captures a person's face using the determination device 100 shown in Figure 21.
[0084] As shown in Figure 22, the control unit 500 gives an instruction to the imaging device 200 to take an image (step S41). The imaging device 200 takes an image of a real person's face in advance. In this case, the imaging device 200 may take continuous images of the real person's face, take a video, or take multiple images intermittently.
[0085] Next, as shown in Figure 22, the control unit 500 determines whether the face is within a predetermined frame (step S42). If the face is not within the predetermined frame, the notification device 600 notifies the subject to change their standing position (step S43). Next, as shown in Figure 22, the control unit 500 determines whether the size of the face captured by the imaging device 200 is the optimal value (step S44). In this embodiment, only the size of the face is determined to be the optimal value, and no determination is made as to whether it is a real face or a photoattack. In other words, if the determination result is that the size of the face is smaller than the optimal value (No. 1 in step S44), the control unit 500 will notify the subject to be imaged via the notification device 600 to move closer to the imaging device 200 (step S45). Figure 23 shows an example of the display on the display device of the notification device 600. In Figure 23, 610 is the screen of the display device, 620 is a frame indicating the desired face size, and 630 is an example of notification from the notification device 600.
[0086] In step S45, the notification device 600 may be used to either or both of the following actions be taken: the voice generator of the notification device 600 may say, "Please bring your face closer to the imaging device 200," or the display device of the notification device 600 may say, "Please move closer so that your face fits within the size of the frame displayed on the display unit," so that the captured face becomes the optimal size.
[0087] On the other hand, if the determination result is that the size of the face is larger than the optimal value (No. 2 in step S44), the control unit 500 will notify the subject to be imaged via the notification device 600 to move away from the imaging device 200 (step S46).
[0088] In step S46, the notification device 600 may be used to either or both of the following actions be taken: the voice generator of the notification device 600 may say, "Please move your face away from the imaging device 200," or the display device of the notification device 600 may say, "Please move your face away so that it fits within the size of the circle displayed on the display unit," so that the captured face becomes the optimal size.
[0089] Finally, if the determination result is that the face size is at the optimal value (Yes in step S44), the control unit 500 may proceed to the process of step S1 shown in Figure 3 (step S47). Alternatively, the display device's screen 610 may display the captured image and the frame 620, and the device may be configured to automatically capture an image when the size and position of the face match those of the frame 620.
[0090] [An embodiment for selecting light of a specific color missing from a light-emitting device] Figure 24 is a flowchart illustrating an example in which, before the processing of step S1 shown in Figure 3, a real person's face is captured in advance, and a specific color of light lacking is selected to be emitted from the light-emitting device 300. In other words, Figure 24 shows the process of selecting a specific color of missing light to be emitted from the light-emitting device 300 before a single image is taken by the imaging device 200.
[0091] First, as shown in Figure 24, the control unit 500 instructs the imaging device 200 to image the real object to be photographed (step S51). Following the instructions of the control unit 500, the imaging device 200 images the real face. In this case, the imaging device 200 may continuously capture images of a real person's face, capture images in video, or capture images intermittently multiple times.
[0092] Next, the imaging device 200 sends the imaging data to the control unit 500. The control unit 500 receives the imaging data (step S52). The control unit 500 extracts the pupils of the person from the imaging data (step S53). The control unit 500 selects the color of the pupils of the person from the imaging data (step S54).
[0093] Next, the control unit 500 selects a specific color of missing light to be emitted according to the color of the pupil (step S55). Subsequently, the control unit 500 performs the process shown in step S2 of Figure 3.
[0094] The reason for selecting specific color-deficient lights to emit based on pupil color is to take into account the presence of users wearing colored contact lenses. For example, specifically for users wearing #0000FF colored contact lenses, it becomes easier to determine whether they are real people or not by using #FFFF00, #FF8800, and #88FF00 lights as specific color-deficient lights. Furthermore, since pupil color varies depending on race, gender, and place of origin, not limited to colored contact lenses, it is preferable to perform the treatment shown in Figure 24.
[0095] As described above, by using the portable terminal 900 and the analysis device 400 and control unit 500 included in the determination device 100 according to the present invention, it is possible to reliably determine the faces of real people. In particular, since it is not necessary to perform multiple imaging, it is possible to determine the faces of a large number of people in a short amount of time. Furthermore, since data that is not a real human face can be excluded in the first determination, processing speed can be increased.
[0096] Furthermore, security can be further enhanced by using various types of flash light to irradiate the subject, making it unclear to the subject which data the mobile terminal 900 and the judgment device 100 are using.
[0097] In this embodiment, the determination device 100 corresponds to the "determination device," the imaging device 200 corresponds to the "imaging device," the flash light corresponds to "colorless emission," the light emission device 300 corresponds to the "light emission device," the analysis device 400 corresponds to the "analysis device," the control unit 500 corresponds to the "control device," the notification device 600 corresponds to the "notification device," the BR data and GR data correspond to "data differences," and the mobile terminal 900 corresponds to the "mobile terminal."
[0098] While the above describes a preferred embodiment of the present invention, the invention is not limited thereto. It will be understood that various other embodiments can be made without departing from the spirit and scope of the invention. Furthermore, although the operation and effects of the configuration of the present invention are described in this embodiment, these operation and effects are examples and do not limit the invention. [Explanation of Symbols]
[0099] 100 Judgment device 200 Imaging device 300 Light-emitting devices 400 Analyzer 500 Control Unit 600 notification device 900 mobile devices
Claims
1. A recognition device for determining a person's face, An imaging device for capturing images of the aforementioned person's face, A light-emitting device that emits light of a specific color when the imaging device performs imaging, An analysis device comprising a first convolutional neural network and a second convolutional neural network, which analyzes imaging data from the imaging device, The system includes a control device for controlling the imaging device, the light-emitting device, and the analysis device, The first convolutional neural network is trained to determine whether the input data is image data of a real person's face or not, using image data of a real person's face and image data of a photograph of the same person's face while the light-emitting device is emitting light, and the second convolutional neural network is trained to determine whether the input data is image data of a real person's face or not, using image data of a real person's face and image data of a photograph of the same person's face while the light-emitting device is not emitting light. The determination device captures the person's face while the light-emitting device is emitting light, and inputs the image data extracted from the area around the face into the first convolutional neural network, and also captures the person's face while the light-emitting device is not emitting light, and inputs the image data extracted from the area around the face into the second convolutional neural network. A determination device that determines whether the captured human face is the face of a real person or not, based on the determination result of the first convolutional neural network and the determination result of the second convolutional neural network.
2. The analysis device further comprises a third convolutional neural network trained using image data of a real person's face and image data of the same person's face displayed on a screen while the light-emitting device is emitting light, and a fourth convolutional neural network trained using image data of a real person's face and image data of the same person's face displayed on a screen while the light-emitting device is not emitting light. The determination device inputs the image data when the light-emitting device is emitting light into the first and third convolutional neural networks, and inputs the image data when the light-emitting device is not emitting light into the second and fourth convolutional neural networks. The determination device according to claim 1, which determines whether the captured human face is the face of a real person based on the determination results of the first to fourth convolutional neural networks.
3. Further including a notification device, The determination device according to claim 1 or 2, wherein the control device causes the imaging device to capture a person's face before performing the determination, and the notification device notifies the person of their standing position so that the size of the person's face is such that the determination can be performed.
4. A recognition device for determining a person's face, An imaging device for capturing images of the aforementioned person's face, A light-emitting device that emits light of a specific color when the imaging device performs imaging, An analysis device equipped with a convolutional neural network for analyzing imaging data from the imaging device, The system includes a control device for controlling the imaging device, the light-emitting device, and the analysis device, The aforementioned convolutional neural network has been trained to determine whether the input data is image data of a real person's face or not, using image data of a real person's face and image data of a photograph of the same person's face while the light-emitting device is emitting light. The determination device captures an image of the person's face while the light-emitting device is emitting light, inputs the image data extracted from the area around the face into the convolutional neural network, and determines whether the captured face is the face of a real person or not. The control device is a determination device that, before performing the determination, causes the imaging device to capture a person's face, and from the captured imaging data, causes the light-emitting device to emit light lacking the optimal second specific color as light of the specific color.
5. A portable terminal including the determination device described in claims 1 to 4.
6. A program for a face recognition device that identifies human faces, An imaging process for capturing the face of the person, The aforementioned imaging process includes a light emission process that emits a specific color of light when imaging is performed, An analysis process comprising a first convolutional neural network and a second convolutional neural network, which analyzes imaging data from the imaging process, Includes control processing for controlling the imaging process, the light emission process, and the analysis process, The first convolutional neural network is trained to determine whether the input data is image data of a real person's face or not, using image data of a real person's face while the light-emitting device is emitting light and image data of a photograph of the same person's face. The second convolutional neural network is trained to determine whether the input data is image data of a real person's face or not, using image data of a real person's face while the light-emitting device is not emitting light and image data of a photograph of the same person's face. The determination process involves capturing an image of the person's face while the light-emitting device is emitting light, and inputting the image data extracted from the area around the face into the first convolutional neural network, and capturing an image of the person's face while the light-emitting device is not emitting light, and inputting the image data extracted from the area around the face into the second convolutional neural network. A program for a determination device that determines whether the captured human face is the face of a real person or not, based on the determination result of the first convolutional neural network and the determination result of the second convolutional neural network.
Citation Information
Patent Citations
Device and method for authenticating biological information
JP2000215308A
Personal authentication system for mobile device and its method
JP2002112340A
Portable telephone with personal authenticating function set
JP2004080080A
Mobile phone
JP2004260240A
Face authentication device, and method and program for changing security strength
JP2007148968A