Personal authentication device, personal authentication method, and program

The personal authentication device enhances accuracy by estimating eyeglass information and applying correction coefficients, addressing the issue of ghost light interference in eyeglass-wearing subjects.

JP7766432B2Active Publication Date: 2025-11-10CANON KK
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021148078
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-10
Publication Date
2025-11-10
Estimated Expiration
2041-09-10

AI Technical Summary

Technical Problem

Existing personal authentication methods using eyeball images are affected by ghost light from eyeglasses, leading to reduced accuracy in gaze detection and personal authentication.

Method used

A personal authentication device that acquires an eyeball image, estimates eyeglass information using a ghost image, and performs personal authentication with infrared illumination and correction coefficients for individual differences.

Benefits of technology

The device maintains accuracy in personal authentication even when subjects wear eyeglasses by utilizing ghost images for improved authentication through deep learning and correction coefficients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007766432000003
    Figure 0007766432000003
  • Figure 0007766432000004
    Figure 0007766432000004
  • Figure 0007766432000005
    Figure 0007766432000005
Patent Text Reader

Abstract

To suppress reduction in accuracy of eyeball image-based personal authentication even when a subject is wearing eyeglasses.SOLUTION: A personal authentication device is provided, comprising an acquisition unit for acquiring an eyeball image of a user, an estimation unit configured to estimate information on eyeglasses worn by the user based on a ghost image appearing in the eyeball image, and an authentication unit configured to perform personal authentication of the user based on the eyeball image and the information on eyeglasses.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a personal authentication technique using an eyeball image. [Background technology]

[0002] Conventionally, there are known techniques for detecting the direction of a person's gaze and for performing personal authentication using eyeball images. However, when a subject wears eyeglasses, ghost light from the eyeglasses can affect the accuracy of gaze detection and personal authentication.

[0003] As a solution to this problem, Patent Document 1 discloses a method of determining the refractive index of glasses during calibration of gaze detection and correcting gaze point detection analysis.

[0004] Furthermore, Patent Document 2 discloses a method for detecting the gaze by photographing the eyeball from three directions and using polarization information from the three directions to compensate for areas where the image is not visible due to light reflected from the surface of the glasses using multiple pieces of information. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] International Publication No. 2014 / 046206 [Patent Document 2] International Publication No. 2017 / 014137 Summary of the Invention [Problem to be solved by the invention]

[0006] However, in the method disclosed in Patent Document 1, when ghost light is incident, saturated light is reflected on the pupil, and the eyeball image is missing, resulting in a problem of reduced accuracy in gaze detection.

[0007] Furthermore, in Patent Document 2, the device becomes complicated, expensive, and large, and the range of ghost light is large, so that it is not always possible to eliminate the influence of coast light in images from any direction.

[0008] These problems also occur when the techniques described in Patent Documents 1 and 2 are applied to personal authentication.

[0009] The present invention has been made in consideration of the above-mentioned problems, and its purpose is to suppress a decrease in the accuracy of personal authentication when performing personal authentication using an eyeball image, even when the subject is wearing glasses. [Means for solving the problem]

[0010] The personal authentication device according to the present invention comprises: an acquisition means for acquiring an image of a user's eyeball; an estimation means for estimating information about eyeglasses worn by the user based on a ghost image captured in the image of the eyeball; and an authentication means for performing personal authentication of the user based on the image of the eyeball and the information about the eyeglasses. an illumination means for illuminating the user's eyeball with infrared light from at least two light sources; a gaze detection means for detecting the gaze from the image of the eyeball illuminated by the illumination means; and a storage means for storing a correction coefficient according to individual differences of the user and information about the eyeglasses. Equipped with The gaze detection means detects the gaze using the correction coefficient according to information about the glasses when the same user wears different glasses. It is characterized by: [Effects of the Invention]

[0011] According to the present invention, when performing personal authentication using an eyeball image, it is possible to suppress a decrease in accuracy of personal authentication even when the subject is wearing eyeglasses. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a diagram showing the appearance of a digital still camera that is a first embodiment of a personal authentication device of the present invention. [Figure 2] FIG. 1 is a schematic diagram illustrating the configuration of a digital still camera. [Figure 3] FIG. 1 is a diagram showing the block configuration of a digital still camera. [Figure 4] 1 is a schematic diagram of an eyeball image projected onto an eyeball image sensor. [Figure 5] FIG. 4 is an explanatory diagram of a visual recognition state of a display element by a user. [Figure 6] 4A and 4B are diagrams illustrating the reflected light of illumination light from the surface of glasses. [Figure 7] FIG. 10 is an explanatory diagram showing a state in which a ghost occurs in an eyeball image. [Figure 8] FIG. 1 is an explanatory diagram showing the basic configuration of a CNN that estimates a gaze point position from input image data. [Figure 9] FIG. 10 is an explanatory diagram of feature detection processing at the feature detection cell plane and feature integration processing at the feature integration cell plane. [Figure 10] 4 is a flowchart of personal authentication processing in the first embodiment. [Figure 11] 10 is a flowchart of a personal authentication process according to the second embodiment. [Figure 12] FIG. 11 is an explanatory diagram showing the field of view in a finder according to a third embodiment. [Figure 13] FIG. 10 is a diagram illustrating the principle of a gaze detection method according to a third embodiment. [Figure 14] 11 is a flowchart of a gaze detection process using eyeglasses individual information according to the third embodiment. [Figure 15] FIG. 11 is a diagram showing the relationship between individuals and correction coefficient data in the third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention claimed. Although multiple features are described in the embodiments, not all of these multiple features are necessarily essential to the invention, and multiple features may be combined arbitrarily. Furthermore, in the accompanying drawings, the same reference numerals are used to designate the same or similar components, and redundant explanations will be omitted.

[0014] (First embodiment) 1A and 1B are diagrams showing the appearance of a digital still camera 100, which is a first embodiment of a personal authentication device of the present invention. Fig. 1A is a front perspective view of the digital still camera 100, and Fig. 1B is a rear perspective view of the digital still camera 100.

[0015] 1(a), the digital still camera 100 in this embodiment is configured by detachably attaching a photographing lens 100B to a camera body 100A. A release button 105, which is an operating member that receives image capturing operations from the user, is disposed on the camera body 100A.

[0016] 1(b), an eyepiece window frame 121 and an eyepiece 112 are arranged on the back of the digital still camera 100, allowing the user to look into a display element 110 (described later) contained inside the camera. A plurality of illumination light sources 113a, 113b, 113c, and 113d that illuminate the eyeball are arranged around the eyepiece 112.

[0017] Fig. 2 is a cross-sectional view of the camera housing cut along the YZ plane defined by the Y axis and Z axis shown in Fig. 1(a), and shows a schematic configuration of the digital still camera 100. In Fig. 1 and Fig. 2, corresponding parts are denoted by the same reference numerals.

[0018] 2, a photographing lens 100B in an interchangeable lens camera is attached to a camera body 100A. In this embodiment, for the sake of convenience, the interior of photographing lens 100B is shown as being composed of two lenses 151 and 152, but in reality, as is well known, it is composed of many more lenses.

[0019] Camera body 100A is equipped with an image sensor 102 arranged at the intended imaging plane of photographing lens 100B. Camera body 100A contains a CPU 103 that controls the entire camera, and a memory unit 104 that records images captured by image sensor 102. Also arranged within eyepiece window frame 121 are a display element 110 made up of a liquid crystal or the like for displaying the captured image, a display element drive circuit 111 that drives it, and an eyepiece 112 for observing the subject image displayed on display element 110.

[0020] By looking into the interior of the eyepiece window frame 121, the user can see a virtual image 501 of the display element 110 in a state as shown in Fig. 5. Specifically, the virtual image 501 of the display element 110 is enlarged by the eyepiece 112 and formed at a position approximately 50 cm to 2 m away from the eyepiece 112, with the image being larger than the actual size. The user will view this virtual image 501. Fig. 5 illustrates the image being formed at a position 1 m away.

[0021] Illumination light sources 113a to 113d are light sources for illuminating the photographer's eyeball 114, and are made up of infrared light-emitting diodes and arranged around the eyepiece 112. Illumination light sources 113a to 113d are used to detect the gaze direction from the relationship between the pupil and the reflection of the light source due to corneal reflection.

[0022] The illuminated eyeball image and images resulting from the corneal reflection of illumination light sources 113a to 113d pass through eyepiece 112, are reflected by beam splitter 115, and are then focused by light-receiving lens 116 on eyeball image-capturing element 117, which is a two-dimensional array of photoelectric conversion elements such as CCDs. Light-receiving lens 116 is positioned so that the pupil of photographer's eyeball 114 and eyeball image-capturing element 117 form a conjugate imaging relationship. From the positional relationship between the eyeball image focused on eyeball image-capturing element 117 and the images resulting from the corneal reflection of light sources 113a to 113b, the gaze direction can be detected using a predetermined algorithm, which will be described later. In this embodiment, the photographer (subject) wears an optical element such as eyeglasses 144, for example, and this optical element is positioned between eyeball 114 and eyepiece window frame 121.

[0023] On the other hand, the photographic lens 100B includes an aperture 161, an aperture drive device 162, a lens drive motor 163, a lens drive member 164, and a photocoupler 165. The lens drive member 164 is made up of gears and the like. The photocoupler 165 detects the rotation of a pulse plate 166 that rotates in conjunction with the lens drive member 164, and transmits this to a focus adjustment circuit 168.

[0024] The focus adjustment circuit 168 drives the lens drive motor 163 a predetermined amount based on information about the amount of rotation of the lens drive member 164 and information about the amount of lens drive from the camera side, and adjusts the photographic lens 100B to a focused state. Note that the photographic lens 100B exchanges signals with the camera body 100A via the mount contacts 147 of the camera body 100A.

[0025] 3 is a block diagram showing the electrical configuration of digital still camera 100. A gaze detection circuit 301, a photometry circuit 302, an autofocus detection circuit 303, a signal input circuit 304, a display element drive circuit 111, and an illumination light source drive circuit 305 are connected to CPU 103 built into camera body 100A. CPU 103 is also connected via mount contacts 147 to focus adjustment circuit 168 disposed in photographic lens 100B and aperture control circuit 306 included in aperture drive device 162. Memory unit 104 attached to CPU 103 has a storage area for image signals from image sensor 102 and eye image sensor 117, and a storage area for gaze correction data that corrects for individual differences in gaze, as described below.

[0026] The gaze detection circuit 301 A / D converts the eyeball image signal from the eyeball image sensor 117 and transmits this image information to the CPU 103. The CPU 103 extracts each feature point of the eyeball image required for gaze detection in accordance with a predetermined algorithm described later, and further calculates the photographer's gaze from the position of each feature point.

[0027] The photometry circuit 302 acquires a luminance signal corresponding to the brightness of the field based on a signal obtained from the image sensor 102, which also functions as a photometry sensor, and performs amplification, logarithmic compression, and A / D conversion on the signal before sending it to the CPU 103 as field luminance information.

[0028] The autofocus detection circuit 303 A / D converts signals from multiple pixels used for phase difference detection in the image sensor 102 and sends the converted signals to the CPU 103. The CPU 103 calculates the defocus amount corresponding to each focus detection point from the signals from these multiple pixels. This is a well-known technique known as image plane phase difference AF.

[0029] Fig. 4 is a diagram showing the output of the eyeball image pickup element 117. Fig. 4(a) shows an example of a reflected image obtained from the eyeball image pickup element 117, and Fig. 4(b) shows the signal output intensity of the eyeball image pickup element 117 in an area α of this example image. In Fig. 4, the horizontal direction is the X-axis, and the vertical direction is the Y-axis.

[0030] In this case, the coordinates in the X-axis direction (horizontal direction) of images Pd' and Pe' formed by the corneal reflection images of illumination light sources 113a and 113b are defined as Xd and Xe, respectively. Also, the coordinates in the X-axis direction of images a' and b' formed by the light beams from edges 401a and 401b of pupil 401 are defined as Xa and Xb.

[0031] In the example of luminance information (example of signal intensity) in FIG. 4(b), extremely high levels of luminance are obtained at positions Xd and Xe corresponding to images Pd' and Pe' formed by corneal reflection light from illumination light sources 113a and 113b. In the region between coordinates Xa and Xb corresponding to the region of pupil 401, excluding the positions Xd and Xe, extremely low levels of luminance are obtained. In contrast, in regions corresponding to the region of iris 403 outside pupil 401, with X coordinate values ​​smaller than Xa and X coordinate values ​​larger than Xb, intermediate values ​​between these two types of luminance levels are obtained. From such information about fluctuations in luminance level with respect to X coordinate position, the X coordinates Xd and Xe of images Pd' and Pe' formed by corneal reflection light from illumination light sources 113a and 113b and the X coordinates Xa and Xb of images a' and b' at the pupil edge can be obtained.

[0032] Furthermore, when the rotation angle θx of the optical axis of the eyeball 114 relative to the optical axis of the light receiving lens 116 is small, the coordinate Xc of the point (let's call it c') corresponding to the pupil center c imaged on the eyeball imaging element 117 can be expressed as Xc ≒ (Xa + Xb) / 2.

[0033] From these, it is possible to estimate the X coordinate of the point c' corresponding to the pupil center imaged on the eyeball image sensor 117 and the coordinates of the corneal reflection images Pd' and Pe' of the illumination light sources 113a and 113b.

[0034] FIG. 4(a) shows an example of calculating the rotation angle θx when the user's eyeball rotates in a plane perpendicular to the Y axis, but the method for calculating the rotation angle θy when the user's eyeball rotates in a plane perpendicular to the X axis is similar.

[0035] The state of looking into the eyepiece window frame 121 will be described with reference to Fig. 5. Fig. 5 is a schematic top view seen from the positive direction of the Y axis, showing a state in which a user views a virtual image 501 of the display element 110 through the eyepiece window frame 121 and the eyepiece 112. Note that, for ease of understanding, the eyepiece 112 shown in Fig. 1 is omitted from Fig. 5.

[0036] In Fig. 5, the user views a virtual image 501 that is enlarged from the actual size of the display element 110 by the eyepiece 112 (not shown). Typically, the optical system of a viewfinder is adjusted so that the virtual image is formed at a distance of several tens of centimeters to 2 meters from the eyepiece 112. In this embodiment, the virtual image is illustrated as being formed at a distance of 1 meter. Fig. 5 also shows a situation in which the user is wearing eyeglasses 144, and the eyeball 114 and the eyepiece window frame 121 are positioned apart, with the eyeglasses 144 sandwiched between them.

[0037] 5(a) shows a state in which a user is gazing at approximately the center of the screen, with the center of their eyeballs positioned at approximately the same position as the optical axis centers of the eyepiece window frame 121 and the display element 110. In this state, the user can see an area indicated by field of view range β1. In this state, if the user wants to see a range γ1 at the edge of the screen that is currently not visible, the user tends to translate their eyeballs and head together by a large amount in the positive direction of the X axis (downward on the page), as shown in FIG. 5(b).

[0038] This translational movement shifts the center position of the eyeball 114 in a direction perpendicular to the optical axis, and a new field of view range β2 is formed by the lines OC and OC', which connect the pupil 401 of the eyeball after the shift and the edge of the eyepiece window frame 121. In field of view range β2, the visible range shifts toward the negative direction of the X axis (upward on the page), and range γ1, which was not visible before the translational movement, is included in the field of view. Therefore, the user has successfully seen range γ1 through the translational movement. However, conversely, the range outside the field of view in the downward direction on the page becomes range γ2', which is wider than range γ2 in Figure 5(a).

[0039] Similarly, if the user wants to see the opposite screen edge area γ2, which is not visible in the state shown in Figure 5(a), the user translates their head in the negative direction of the X axis (upward on the paper), in the opposite direction to Figure 5(b), resulting in the state shown in Figure 5(c). By shifting their head in the opposite direction to Figure 5(b), a new field of view area β3 is formed in Figure 5(c), sandwiched between lines OD and OD'. The visible area of ​​field of view area β3 shifts toward the positive direction of the X axis (downward on the paper), and area γ2, which was outside the field of view in Figure 5(a), is included in the field of view. Therefore, by translating their head, the user has successfully seen area γ2. However, the area outside the field of view in the upward direction on the paper becomes area γ1', which is wider than area γ1 in Figure 5(a).

[0040] Furthermore, in the viewing state where the head is translated as described above and the eyepiece window frame 121 is peered at from an oblique direction, as mentioned above, not only do the eyeballs 114 rotate, but the head itself often tilts as well. This will be explained below.

[0041] In FIGS. 5(a) to 5(c), the user is viewing a virtual image 501 of the display element 110 through the eyepiece window frame 121 with the eyeball (right eye) facing upward on the paper.

[0042] As described above, when the user enters the peering viewing state, the state changes from, for example, FIG. 5(a) to FIG. 5(b). At this time, not only does the eyeball (right eye) facing upward in the page, which is looking into the eyepiece window frame 121, rotate, but the user often also tilts their head as shown by the symbol θh to peer into the image. If the user is wearing optical components such as eyeglasses at this time, the tilt of the eyeglasses 144 also changes in the same direction as the tilt of the head θh, as shown in FIG. 5(b). As a result, the tilt of the eyeglasses 144 with respect to the optical axis changes, and the position of the ghost image generated when the light from the illumination light sources 113a to 113d is reflected on the surface of the eyeglasses 144 also changes accordingly.

[0043] 6, the above-described ghosts are caused when light emitted from illumination light sources 113a to 113d is reflected by eyeglasses 144 and the reflected light is incident on eye image pickup element 117 as shown by the arrows. Although Fig. 6 shows the path of light from illumination light source 113a or 113b as an example, light from illumination light sources 113c and 113d can also be incident on eye image pickup element 117 in the same way.

[0044] The above ghosts appear as, for example, Ga, Gb, Gc, and Gd in the eyeball image shown in Fig. 7(a). These ghosts are generated when light emitted from illumination light sources 113a, 113b, 113c, and 113d is reflected on the surface of eyeglasses 144, and appear separately from the Purkinje images generated when light emitted from each illumination light source is reflected on the corneal surface of eyeball 114. These ghosts exist approximately symmetrically when the optical components, such as eyeglasses 144, worn by the user are facing forward.

[0045] In contrast, when the head tilts θh as described above, the ghosts reflected on the eyes move toward the right on the page, as shown in Figure 7(b). At this time, a ghost Ga' caused by, for example, illumination light source 113a may overlap with the image of pupil 401 located in the area near the center of the eyeball image, as shown in the figure, thereby obscuring part of the pupil image. When part of the pupil image is obscured, the eyeball image is missing, which causes a problem of reduced accuracy in personal authentication.

[0046] In this embodiment, ghosts, which have been a factor in reducing accuracy in conventional personal authentication, are utilized in personal authentication to improve authentication accuracy. In this embodiment, estimation of the type of eyeglasses using ghosts is performed using a trained model obtained through deep learning. Note that in this embodiment, an example of estimating information on the type of eyeglasses will be described, but other information about eyeglasses, such as the reflectance, refractive index, and color of the eyeglasses, may also be estimated from the ghost.

[0047] The following describes a method for analyzing and estimating the type of eyeglasses by inputting the image obtained from the eyeball image sensor 117 as input information to a convolutional neural network (hereinafter referred to as CNN). By using CNN, it becomes possible to perform detection with higher accuracy.

[0048] The basic configuration of CNN will be explained with reference to FIGS. 8 and 9.

[0049] Fig. 8(a) is a diagram showing the basic configuration of a CNN that estimates the type of eyeglasses from input 2D image data. The input 2D image data may be an eyeball image obtained from the eyeball image sensor 117, but it is better to use an image in which ghosts are highlighted by methods such as reducing the overall brightness of the eyeball image or applying a filter. This is because, as mentioned above, ghosts are generated by reflection from the illumination light source and therefore have high brightness, and because the CNN in Fig. 8(a) aims to determine the type of eyeglasses, the accuracy of the determination can be improved by reducing signals other than ghosts as much as possible.

[0050] The processing flow of CNN is that the input is on the left and processing proceeds to the right. CNN is composed of a set of two layers called the feature detection layer (S layer) and the feature integration layer (C layer), which are arranged hierarchically.

[0051] In CNN, the next feature is detected in the S layer based on the features detected in the previous layer. The features detected in the S layer are then integrated in the C layer and sent to the next layer as the detection result for that layer.

[0052] The S layer consists of feature detection cell planes, each of which detects a different feature. The C layer consists of feature integration cell planes, which pool the detection results from the previous feature detection cell planes. Hereinafter, unless there is a need to distinguish between them, the feature detection cell planes and feature integration cell planes will be collectively referred to as feature planes. In this embodiment, the output layer, which is the final layer, does not use the C layer and is composed only of the S layer.

[0053] The feature detection process at the feature detection cell plane and the feature integration process at the feature integration cell plane will be described in detail with reference to FIG.

[0054] The feature detection cell surface is composed of multiple feature detection neurons, which are connected to the C layer in the previous layer in a specific structure. The feature integration cell surface is composed of multiple feature integration neurons, which are connected to the S layer in the same layer in a specific structure.

[0055] In the Mth cell plane of the Sth layer of the Lth hierarchy shown in Figure 9, the output value of the feature detection neuron at position (ξ,ζ) is expressed as y LS M (ξ,ζ), in the Mth cell plane of the C layer of the Lth layer, the output value of the feature integration neuron at position (ξ,ζ) is y LC M (ξ,ζ). In this case, the coupling coefficient of each neuron is expressed as w LS M (n,u,v), w LC M Assuming (u,v), each output value can be expressed as follows:

[0056]

number

[0057] In (Equation 1), f is an activation function, which can be any sigmoid function such as a logistic function or a hyperbolic tangent function, and can be realized by, for example, a tanh function. LSM (ξ,ζ) is the internal state of the feature detection neuron at position (ξ,ζ) on the Mth cell plane of the Sth layer of the Lth hierarchy. (Equation 2) is a simple linear sum without using an activation function. When no activation function is used as in (Equation 2), the internal state of the neuron u LS M (ξ,ζ) and output value y LC M (ξ,ζ) are equal. Also, y in (Equation 1) L-1C n (ξ+u,ζ+v), y in (Equation 2) LS M (ξ+u,ζ+v) are called the output values ​​of the feature detection neuron and the feature integration neuron, respectively.

[0058] The following describes ξ, ζ, u, v, and n in (Equation 1) and (Equation 2).

[0059] The position (ξ,ζ) corresponds to the position coordinate in the input image, e.g., y LS M If (ξ,ζ) has a high output value, it means that there is a high possibility that the feature to be detected on the Mth cell plane of the Sth layer of the Lth hierarchical level exists at pixel position (ξ,ζ) of the input image. Furthermore, n in (Equation 2) refers to the nth cell plane of the Cth layer of the L-1th hierarchical level, and is called the target feature number. Basically, a product-sum operation is performed on all cell planes that exist on the Cth layer of the L-1th hierarchical level. (u,v) are the relative position coordinates of the coupling coefficient, and the product-sum operation is performed within a finite range (u,v) depending on the size of the feature to be detected. This finite range of (u,v) is called the receptive field. The size of the receptive field is hereinafter called the receptive field size, and is expressed as the number of horizontal pixels x the number of vertical pixels in the coupled range.

[0060] Also, in (Equation 1), L=1, that is, the first S layer, y L-1C n (ξ+u,ζ+v) is the input image y in_image (ξ+u,ζ+v) or input position map y in_posi_map(ξ+u,ζ+v). Incidentally, the distribution of neurons and pixels is discrete, and the connection feature numbers are also discrete, so ξ,ζ,u,v,n are not continuous variables but take discrete values. Here, ξ and ζ are non-negative integers, n is a natural number, and u and v are integers, all of which have a finite range.

[0061] (Equation 1) LS M (n,u,v) is the distribution of coupling coefficients for detecting a specific feature, and by adjusting it to an appropriate value, it becomes possible to detect the specific feature. This adjustment of the coupling coefficient distribution is deep learning, and in the construction of CNN, various test patterns are presented and y LS M The coupling coefficients are adjusted by repeatedly correcting them gradually so that (ξ,ζ) becomes an appropriate output value.

[0062] Next, w in (Equation 2) LC M (u, v) uses a two-dimensional Gaussian function and can be expressed as shown in the following (Equation 3).

[0063]

number

[0064] Here again, (u,v) is assumed to have a finite range, so just as in the explanation of feature detection neurons, this finite range is called the receptive field, and the size of the range is called the receptive field size. Here, this receptive field size can be set to an appropriate value depending on the size of the Mth feature in the Sth layer of the Lth hierarchy. In (Equation 3), σ is the feature size factor, and can be set to an appropriate constant depending on the receptive field size. Specifically, it is best to set it so that the outermost values ​​of the receptive field can be considered to be almost 0.

[0065] The CNN of this embodiment shown in FIG. 8(a) performs the above-described calculations at each layer, and determines the type of glasses at the final layer, S layer.

[0066] By performing the above-mentioned determination, it becomes possible to determine the type of glasses from the ghost. Note that in this embodiment, the learning model configured by CNN that estimates the type of glasses described above is a trained model that has previously completed deep learning using two-dimensional image data as input and the type of glasses as training data.

[0067] Fig. 8(b) is a diagram showing the configuration of a CNN that performs personal authentication judgment. The CNN in Fig. 8(b) receives as input an eyeball image obtained from the eyeball image sensor 117 and the judgment result of the eyeglasses type in Fig. 8(a) as eyeglasses individual information. The configuration and processing flow of the CNN are the same as those in Fig. 8(a) although there are differences in input, and therefore a description thereof will be omitted.

[0068] In personal authentication using only an eyeball image, the accuracy of identification decreases if the eyeball image is obscured by ghosts. However, by using both information on the type of eyeglasses and the eyeball image as described above, the accuracy of personal authentication can be improved. In this embodiment, the learning model configured with CNN that performs the above-described personal authentication determination is a trained model that has previously completed deep learning using an eyeball image and type of eyeglasses as input and information for identifying individuals as training data. Since the output of the trained model outputs information that identifies individuals identical to the training data, personal authentication is achieved by comparing the data with pre-registered information for identifying individuals and confirming a match. However, the method of personal authentication is not limited to this procedure. For example, a comparison of the output of a later trained model with pre-registered information for identifying individuals may be incorporated into the trained model. In this case, the learning process may involve, for example, using an eyeball image and type of eyeglasses as input and learning information indicating whether or not personal authentication was successful as training data.

[0069] FIG. 10 is a flowchart showing the control flow for performing personal authentication in camera body 100A.

[0070] In step S1001, CPU 103 turns on illumination light sources 113a and 113b to emit infrared light toward user's eyeball 114. Light reflected from user's eyeball 114 illuminated by this infrared light is formed into an image on eyeball image sensor 117 through light receiving lens 116. Then, eyeball image sensor 117 performs photoelectric conversion, making it possible to process the eyeball image as an electrical signal.

[0071] In step S1002, CPU 103 determines whether or not a ghost exists from the obtained eyeball image signal. The presence or absence of a ghost can be determined by checking the presence or absence of high-intensity pixels other than the Purkinje image in the eyeball image signal. More specifically, according to the principle of ghost occurrence described above, ghosts differ in size and shape from Purkinje images, so the presence or absence of a ghost can be determined by checking the presence or absence of high-intensity pixels different from the Purkinje image.

[0072] In step S1003, if CPU 103 determines that a ghost exists, it proceeds to step S1004, and if it determines that a ghost does not exist, it proceeds to step S1006.

[0073] In step S1004, the CPU 103 inputs the eyeball image signal to the CNN shown in FIG. 8(a) and estimates the type of eyeglasses.

[0074] In step S1005, the CPU 103 inputs information about the type of eyeglasses as eyeglasses individual information as a first input and an eyeball image as a second input to the CNN shown in FIG. 8(b), and performs personal authentication.

[0075] On the other hand, since it is determined in step S1006 that there is no ghost, CPU 103 inputs the eyeball image signal to the CNN shown in FIG. 8(b) to perform personal authentication. Note that the CNN in FIG. 8(b) requires input of eyeglasses-specific information, but in step S1006, it is sufficient to input that there is no eyeglasses-specific information. As an alternative method, a CNN for no-glasses use may be provided separately from the CNN in FIG. 8(b) to perform personal authentication. The CNN for no-glasses use may be configured to identify individuals using only the input image, without using eyeglasses-specific information as input.

[0076] As described above, even when a ghost image caused by glasses occurs, it is possible to prevent a decrease in the accuracy of personal authentication by performing personal authentication using information on the type of glasses based on the characteristics of the ghost image in addition to the eyeball image.

[0077] (Second embodiment) In this embodiment, the result of personal authentication performed from an eyeball image and the result of personal authentication performed from information on ghost characteristics are used to perform final personal authentication, thereby improving authentication accuracy when a ghost occurs.

[0078] In this embodiment, the type of eyeglasses is estimated using a CNN similar to that shown in FIG. 8(a). Furthermore, a CNN similar to that shown in FIG. 8(b) is used to estimate an individual using only an eyeball image, without using individual eyeglasses information as input, as in step S1006 of FIG. 10. Furthermore, a CNN shown in FIG. 8(c) is used to estimate an individual using the eyeglasses type information obtained by the CNN in FIG. 8(a). Finally, individual authentication is performed using the individual estimation results obtained by the CNN in FIG. 8(b) and the individual estimation results obtained by the CNN in FIG. 8(c). Note that the configuration and processing flow of the CNN in FIG. 8(c) are the same as those in FIG. 8(a) despite differences in input, and therefore will not be described here.

[0079] FIG. 11 is a flowchart showing the control flow for performing personal authentication in camera body 100A in the second embodiment.

[0080] Steps S1001 to S1004 and step S1006, which are given the same step numbers as in FIG. 10, are the same as those in FIG. 10 showing the first embodiment.

[0081] In step S1101, the CPU 103 inputs the type of eyeglasses acquired by the CNN in FIG. 8(a) in step S1004 to the CNN in FIG. 8(c) to estimate the individual.

[0082] In step S1102, a CNN similar to that in FIG. 8(b) is used to estimate an individual using only the eyeball image, in the same manner as in step S1006 in FIG. 10, without using individual eyeglass information as input.

[0083] In step S1103, CPU 103 performs personal authentication using the individual estimation result in step S1101 and the individual estimation result in step S1102. A specific example is a method of performing personal authentication by weighting the two estimation results, multiplying each estimation result by a coefficient, and adding the weighted results. However, the method of determining personal authentication using the two estimation results is not limited to this and may be changed as appropriate.

[0084] On the other hand, if there is no ghost in step S1003, the eyeball image is not missing, so in step S1006, CPU 103 performs individual estimation using the CNN in FIG. 8(b) with only the eyeball image as input, similar to step S1006 in FIG. 10.

[0085] As described above, when a ghost occurs due to glasses, the accuracy of authentication can be improved by separately estimating the individual using an eyeball image and estimating the individual using information on the characteristics of the ghost, and then performing individual authentication using the results of each estimation.

[0086] (Third embodiment) In this embodiment, a method for solving the problem regarding the line of sight correction coefficient in line of sight detection when wearing glasses, using the personal authentication method described in the first embodiment, will be described.

[0087] 4 of the first embodiment, it has been explained that the rotation angles θx and θy of the optical axis of the eyeball 114 can be obtained. Here, a method for calculating the gaze point coordinates and the line of sight correction coefficient will be further explained.

[0088] The rotation angles θx and θy of the eyeball optical axis are used to determine the position of the user's line of sight (position of the point of gaze, hereinafter referred to as the gaze point) on the display element 110. When the gaze point position is expressed by coordinates (Hx, Hy) corresponding to the center c of the pupil 401 on the display element 110, it is expressed as follows: Hx=m×( Ax×θx + Bx ) Hy=m×( Ay×θy + By ) Here, coefficient m is a constant determined by the configuration of the viewfinder optical system of the camera, and is a conversion coefficient that converts rotation angles θx and θy into position coordinates corresponding to the center c of the pupil 401 on the display element 110. Coefficient m is assumed to be determined in advance and stored in memory unit 104. Also, Ax, Bx, Ay, and By are gaze correction coefficients that correct for individual differences in the user's gaze, and are assumed to be acquired by performing a calibration operation, which will be described later, and stored in memory unit 104 before the gaze detection routine is started.

[0089] For users who do not wear glasses, the gaze correction coefficient can be calculated as described above. However, users who wear glasses may use multiple pairs of glasses. Since the lens shape differs for each pair of glasses, the appropriate gaze correction coefficient naturally differs for each pair of glasses. Therefore, for users who wear glasses, the appropriate coefficient must be selected from the multiple gaze correction coefficients for each pair of glasses to perform gaze detection.

[0090] The method for achieving this is described below.

[0091] First, a basic method for detecting the line of sight and the calibration required for calculating the line of sight correction coefficient will be described.

[0092] FIG. 12 is a diagram showing the field of view in the finder, showing the state in which the display element 110 is in operation.

[0093] 12, reference numeral 1210 denotes a field of view mask, and reference numeral 1220 denotes a focus detection area. Reference numerals 1230-1 to 1230-180 denote 180 ranging point targets, which are positions corresponding to multiple focus detection points on the imaging surface and are superimposed on the through image displayed on the display element 110. Of these targets, the target frame corresponding to the current estimated gaze point position is displayed as estimated gaze point A.

[0094] FIG. 13 is a diagram illustrating the principle of the line of sight detection method, and shows a schematic configuration of an optical system for performing the line of sight detection shown in FIG.

[0095] 13, illumination light sources 113a and 113b, such as light-emitting diodes, emit infrared light that is insensitive to the user. Illumination light sources 113a and 113b are arranged approximately symmetrically with respect to the optical axis of light-receiving lens 116, and illuminate user's eyeball 114. A portion of the illumination light reflected by eyeball 114 is collected by light-receiving lens 116 onto eyeball image sensor 117.

[0096] As explained in FIG. 4 of the first embodiment, in gaze detection, the rotation angles θx and θy of the eyeball optical axis are obtained from the eyeball image, and the position of the pupil center is subjected to coordinate conversion to a corresponding position on the display element 110 to estimate the position of the gaze point.

[0097] However, due to factors such as individual differences in the shape of human eyeballs and differences in the shape of eyeglasses, unless the values ​​of the gaze correction coefficients Ax, Ay, Bx, and By are adjusted to appropriate values, a discrepancy will occur between the actual gaze position B and the calculated estimated gaze point C, as shown in Fig. 12(b). In the example of Fig. 12(b), even though the person is gazing at the person at position B, the camera erroneously estimates that the person is gazing at the background, making it impossible to perform appropriate focus detection and adjustment.

[0098] Therefore, before capturing an image with the camera, it is necessary to perform a calibration operation to obtain a line-of-sight correction coefficient value appropriate for the user and the glasses being used, and store the value in the camera.

[0099] Conventionally, calibration work is performed by highlighting multiple indices at different positions in the viewfinder field of view before capturing an image, as shown in Figure 12(c), and having the user look at the indices. A known technique is to execute a gaze point detection flow when gazing at each indices, and then calculate an appropriate gaze correction coefficient value from the calculated multiple estimated gaze point coordinates and each indices coordinate.

[0100] By performing the above calibration for each combination of a user and the eyeglasses worn by that user, it becomes possible to maintain an appropriate gaze correction coefficient during use. In other words, it is desirable for a user who wears multiple eyeglasses to perform calibration for each pair of eyeglasses.

[0101] Next, FIG. 14 is a flowchart showing a method for selecting a line-of-sight correction coefficient using the result of personal authentication and an operation for line-of-sight detection.

[0102] Steps S1001 to S1006, which are given the same step numbers as in FIG. 10, are operations for performing personal authentication in the first embodiment, and are the same operations as in FIG.

[0103] In step S1401, the CPU 103 reads out the associated gaze correction coefficient from the memory unit 104 based on the personal authentication result of step S1005 or S1006. The relationship between the personal authentication result and the gaze correction coefficient will be described with reference to FIG. 15. The problem with the third embodiment is that when a user who wears glasses uses multiple pairs of glasses, the appropriate gaze correction coefficient will differ. Therefore, as shown in FIG. 15, in addition to the personal authentication result, the type of glasses is associated with the gaze correction coefficient. This makes it possible to manage correction coefficients that can be used even when multiple pairs of glasses are used, and to read out the gaze correction coefficient.

[0104] In step S1402, the CPU 103 uses the line-of-sight correction coefficient obtained in step S1401 to calculate the gaze point coordinates (Hx, Hy) of the center c of the pupil 401 on the display element 110 as described above.

[0105] In step S1403, the CPU 103 stores the calculated gaze point coordinates in the memory unit 104, and ends the gaze detection routine.

[0106] As described above, according to this embodiment, by using an appropriate line-of-sight correction coefficient associated with each individual and the type of eyeglasses, it is possible to obtain appropriate gaze point coordinates.

[0107] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention.

[0108] For example, in the above embodiment, CNN is used for personal authentication, but the present invention is not limited to this. For example, information on iris characteristics, such as an iris code, may be extracted from an eyeball image, and this information may be used in conjunction with information on the type of eyeglasses to identify the individual through pattern matching. The type of eyeglasses may also be estimated using pattern matching.

[0109] (Other embodiments) The present invention can also be realized by providing a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having the computer of the system or device read and execute the program. The computer has one or more processors or circuits, and may include multiple separate computers or a network of multiple separate processors or circuits to read and execute computer-executable instructions.

[0110] The processor or circuitry may include a central processing unit (CPU), a microprocessing unit (MPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gateway (FPGA), a digital signal processor (DSP), a data flow processor (DFP), or a neural processing unit (NPU).

[0111] The invention is not limited to the above-described embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention. Accordingly, the following claims are appended to apprise the public of the scope of the invention. [Explanation of symbols]

[0112] 100: digital still camera, 100A: camera body, 100B: photographing lens, 102: image sensor, 103: CPU, 104: memory section, 110: display element, 113a to 113d: illumination light source, 114: eyeball, 115: light splitter, 116: light receiving lens, 117: eyeball image sensor, 401: pupil, 403: iris, 144: glasses

Claims

1. an acquisition means for acquiring an image of the user's eyeball; an estimation means for estimating information about the eyeglasses worn by the user based on a ghost image captured in the eyeball image; an authentication means for performing personal authentication of the user based on the eyeball image and information about the eyeglasses; an illumination means for illuminating the user's eye with infrared light from at least two light sources; a gaze detection means for detecting a gaze from the eyeball image illuminated by the illumination means; a storage means for storing correction coefficients according to individual differences of users and information about their eyeglasses; Equipped with The personal authentication device is characterized in that the gaze detection means detects the gaze using the correction coefficient according to information about the glasses when the same user wears different glasses.

2. 2. The personal authentication device according to claim 1, wherein the information about the glasses is information about the type of glasses.

3. 3. The personal authentication device according to claim 1, further comprising a determination means for determining whether a ghost image is present in the eyeball image, and when the determination means determines that no ghost image is present, the authentication means performs personal authentication of the user based only on the eyeball image.

4. 4. The personal authentication device according to claim 1, wherein the authentication means authenticates the individual using a deep learning neural network.

5. The personal authentication device according to claim 1 , wherein the estimation unit estimates information about the eyeglasses using a deep learning neural network.

6. 4. The personal authentication device according to claim 1, wherein the authentication means extracts features from the eyeball image and authenticates the individual by using pattern matching.

7. 4. The personal authentication device according to claim 1, wherein the estimation means estimates information about the eyeglasses based on the eyeball image by using pattern matching.

8. 8. The personal authentication device according to claim 1, wherein the eyeball image is an image that includes at least an iris.

9. The personal authentication device according to any one of claims 1 to 8, characterized in that the authentication means performs a first identification to identify an individual from information about the glasses and a second identification to identify an individual from the eyeball image, and authenticates the individual using the results of the first identification and the results of the second identification.

10. 10. The personal authentication device according to claim 9, wherein the authentication means authenticates the individual by weighting and adding the result of the first identification and the result of the second identification.

11. The personal authentication device according to any one of claims 1 to 10, wherein information about the glasses is used as input information for deep learning for personal authentication.

12. an acquisition step of acquiring an image of the user's eyeball; an estimation step of estimating information about the eyeglasses worn by the user based on the ghosts captured in the eyeball image; an authentication step of performing personal authentication of the user based on the eyeball image and information about the eyeglasses; illuminating the user's eye with infrared light from at least two light sources; a gaze detection step of detecting a gaze from the eyeball image illuminated in the illumination step; a storage step of storing correction coefficients according to individual differences of the user and information about the glasses; and A personal authentication method characterized in that, in the gaze detection step, when the same user wears different glasses, the gaze is detected using the correction coefficient according to information about the glasses.

13. A program for causing a computer to function as each of the means of the personal authentication device according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Gaze determination using glare as input

    JP2021195124A

  • Identity authentication using lens features

    US20200233202A1

  • Gaze determination using glare as input

    US20210181837A1

  • Line of sight detection device, display method, line of sight detection device calibration method, spectacle lens design method, spectacle lens selection method, spectacle lens manufacturing method, printed matter, spectacle lens sales method, optical device, line of sight information detection method, optical instrument design method, optical instrument, optical instrument selection method, and optical instrument production method

    WO2014046206A1

  • Eyeball observation device, eyewear terminal, gaze detection method, and program

    WO2017014137A1