Imaging device, control method thereof, and program
The imaging device improves gaze position detection accuracy by using eye image data, reliability assessment, and neural networks to infer gaze positions, addressing discrepancies in conventional systems and enhancing focus control.
Patent Information
- Application Number
- JP2021070512
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-04-19
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-04-19
AI Technical Summary
Conventional imaging devices face challenges in accurately detecting the user's intended gaze position due to discrepancies caused by variations in pupil diameter and distance during calibration and shooting, leading to a heavy burden on users and potential misalignment in focus control.
The imaging device includes a gaze detection system that uses eye image data to estimate the gaze position, incorporates a reliability determination mechanism to assess accuracy, and employs a neural network (CNN) to infer the gaze position when reliability is low, ensuring accurate focus adjustment through user interaction and learning.
This approach enhances the accuracy of gaze position detection, reducing user burden and improving focus control by adaptively adjusting to varying conditions, thereby enhancing imaging device performance.
Smart Images

Figure 0007739026000004 
Figure 0007739026000005 
Figure 0007739026000006
Abstract
Description
[Technical Field]
[0001] The present invention relates to an imaging device, a control method thereof, and a program, and more particularly to an imaging device, a control method thereof, and a program that perform focus control based on information on a detected line-of-sight position. [Background technology]
[0002] In recent years, imaging devices have become more automated and intelligent, and imaging devices have been proposed that can recognize the subject intended by the user based on information about the gaze position of the user looking through the viewfinder and perform focus control without the need to manually input the subject position.In this case, when the imaging device detects the user's gaze position, a discrepancy occurs between the user's intended gaze position and the user's gaze position recognized by the imaging device, and it may not be possible to focus on the subject intended by the user.
[0003] In response to this, a technique is known in which an index is displayed in the viewfinder before shooting, the user is instructed to gaze at the index, the user's gaze position is detected while the user is gazing at the index, and calibration is performed to detect the amount of deviation from the index position.Then, during shooting, the user's gaze position recognized by the imaging device is corrected by the detected amount of deviation, so that the corrected gaze position is closer to the gaze position intended by the user (see, for example, Patent Document 1).
[0004] Also, a technique is known in which the detection accuracy of the gaze position is determined, and display objects are sparsely displayed in areas where the determined detection accuracy is low, and display objects are densely displayed in areas where the gaze detection accuracy is high, thereby preventing the selection of a gaze position unintended by the user (see, for example, Patent Document 2). [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-008323 [Patent Document 2] Japanese Patent Application Laid-Open No. 2015-152938 Summary of the Invention [Problem to be solved by the invention]
[0006] However, with the conventional technology disclosed in Patent Document 1, if conditions such as the pupil diameter and distance of the user looking through the viewfinder differ between when shooting and when calibrating, the accuracy of detecting the gaze position drops. Performing calibration again every time the accuracy of detecting the gaze position drops in this way places a heavy burden on the user.
[0007] Furthermore, as described in Patent Document 2, it is possible to prevent erroneous detection of gaze position by enlarging the focus frame in areas where detection accuracy is poor, but if the focus frame becomes too large, there is a risk that the focus control desired by the user will not be performed.
[0008] SUMMARY OF THE INVENTION It is therefore an object of the present invention to provide an imaging device, a control method thereof, and a program that can improve the accuracy of detecting the position of a user's gaze. [Means for solving the problem]
[0009] The imaging device according to claim 1 of the present invention is an imaging device that displays a through image in an internal finder, and includes a generating unit that captures an image of a user's eyeball looking through the finder and generates eye image data, and a generating unit that acquires the eye image data and generates the through image of the finder based on the acquired eye image data. View a gaze detection means for detecting a gaze position of a user; a display control means for displaying the detected gaze position on the finder so that the gaze position can be moved to another position by a first user operation; and a collection means for collecting the other position as a correct position when a second user operation is performed to determine the other position as a focus position. a reliability determination means for determining the reliability of the detected gaze position; a first focus means for focusing the imaging device using the detected gaze position as the focus position when the reliability is high; and a second focus means for focusing the imaging device using the gaze position estimated by an inferential device for estimating the gaze position as the focus position when the reliability is low. The correct position is determined by using the eye image data acquired by the gaze detection means as input data, Memorandum It is characterized by being used for learning to create logic tools. [Effects of the Invention]
[0010] According to the present invention, it is possible to improve the accuracy of detecting the gaze position of a user. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a diagram illustrating an outline of the internal configuration of an imaging device according to a first embodiment. [Figure 2] FIG. 1 is a diagram illustrating the appearance of an imaging device. [Figure 3] FIG. 2 is a block diagram showing an electrical configuration built into the imaging device. [Figure 4] 4 is a diagram showing the field of view in the finder in FIG. 3 when the finder is in operation. FIG. [Figure 5] FIG. 1 is a diagram for explaining the principle of a gaze detection method. [Figure 6] 10A and 10B are diagrams for explaining a method for detecting coordinates corresponding to a corneal reflection image and a pupil center from eye image data. [Figure 7] 10 is a flowchart of a gaze detection process. [Figure 8] 4A and 4B are diagrams showing examples of eye image data at the time of calibration and at the time of photography, which are determined to have low reliability by the gaze detection reliability determination circuit in FIG. 3. [Figure 9] 10 is a flowchart of a process for collecting correct answer data used for learning when creating an inference device. [Figure 10] 10 is a flowchart of a focus process during shooting. [Figure 11] 10 is a flowchart of a process for determining a factor that reduces the reliability of a first estimated point of gaze position according to a second embodiment. [Figure 12] 12 is a diagram for explaining a method of generating differential eye image data in step S1104 of FIG. 11. FIG. [Figure 13] 1 is a schematic diagram illustrating an example of the overall configuration of a CNN according to a first embodiment. FIG. [Figure 14] FIG. 14 is a schematic diagram illustrating an example of a partial configuration of the CNN in FIG. 13. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the present invention, and not all of the combinations of features described in the embodiments are necessarily essential to the solution of the present invention.
[0013] Example 1 Hereinafter, with reference to FIGS. 1 to 10, 13 and 14, a learning method and an inference method executed to improve the detection accuracy of the gaze position in the imaging device 1 according to the first embodiment of the present invention will be described.
[0014] The configuration of the imaging device 1 will be described with reference to FIGS.
[0015] 2A and 2B are diagrams showing the appearance of the imaging device 1, where FIG. 2A is a front perspective view, FIG. 2B is a rear perspective view, and FIG. 2C is a diagram for explaining the operating member 42 of FIG. 2B.
[0016] In this embodiment, the image pickup device 1 is made up of a camera housing 1B and a photographing lens 1A that is detachably attached to the camera housing 1B.
[0017] As shown in FIG. 2(a), a release button 5 is provided on the front surface of the camera housing 1B.
[0018] The release button 5 is an operating member that receives an image capturing operation from the user.
[0019] As shown in FIG. 2(b), an eyepiece window 6 and operation members 41 to 43 are provided on the rear surface of the camera housing 1B.
[0020] The eyepiece window 6 is a window through which the user can view an image for visual confirmation displayed on a viewfinder 10, which will be described later with reference to FIG. 1 and is included inside the camera housing 1B.
[0021] The operation member 41 is a touch panel compatible liquid crystal display, the operation member 42 is a lever type operation member, and the operation member 43 is a button type cross key. In this embodiment, the operation members 41 to 43 used for camera operations such as manual movement control of an estimated gaze point position, which will be described later, are provided on the camera housing 1B, but are not limited to this. For example, other operation members such as electronic dials may be provided on the camera housing 1B in addition to or instead of the operation members 41 to 43.
[0022] Fig. 1 is a cross-sectional view of the camera housing B cut along the YZ plane formed by the Y axis and Z axis shown in Fig. 2(a), and shows an outline of the internal configuration of the imaging device 1. In Fig. 1, the same components as those in Fig. 2 are assigned the same reference numerals.
[0023] 1, photographic lens 1A is a photographic lens that is detachably attached to camera body 1B. For convenience, in this embodiment, only two lenses 101 and 102 are shown as lenses inside photographic lens 1A, but it is well known that in reality, photographic lens 1A is made up of many more lenses.
[0024] Camera housing 1B includes therein imaging element 2, CPU 3, memory section 4, viewfinder 10, viewfinder drive circuit 11, eyepiece 12, light sources 13a to 13b, beam splitter 15, light receiving lens 16, and eye imaging element 17.
[0025] The image sensor 2 is disposed on the intended image plane of the photographic lens 1A and captures an image. The image sensor 2 also functions as a photometric sensor.
[0026] The CPU 3 is a central processing unit of a microcomputer that controls the entire imaging device 1 .
[0027] The memory unit 4 records images captured by the imaging element 2. The memory unit 4 also has a function of storing imaging signals from the imaging element 2 and the eye imaging element 17, and stores line-of-sight correction data that corrects for individual differences in line of sight, which will be described later.
[0028] The finder 10 is configured with a liquid crystal display or the like for displaying an image (through image) captured by the image sensor 2.
[0029] The viewfinder drive circuit 11 is a circuit that drives the viewfinder 10.
[0030] The eyepiece 12 is a lens through which the user peers into the eyepiece window 6 (FIG. 2) to observe the visual image displayed on the viewfinder 10.
[0031] Light sources 13a-13b are infrared light-emitting diodes arranged around the eyepiece window 6 (FIG. 2) to illuminate the user's eyeball 14 and detect the user's line of sight. When light sources 13a-13b are turned on, corneal reflection images (Purkinje images) Pd, Pe (FIG. 5) of light sources 13a-13b are formed on the eyeball 14. In this state, light from the eyeball 14 passes through the eyepiece 12 and is reflected by the beam splitter 15. An eye image including an eyeball image is formed on an ocular imaging element 17 (generation means) consisting of a two-dimensional array of photoelectric elements such as a CMOS by the light-receiving lens 16, thereby generating eye image data. The light-receiving lens 16 positions the pupil of the user's eyeball 14 and the ocular imaging element 17 in a conjugate imaging relationship. Using a predetermined algorithm described later, the gaze detection circuit 201 (gaze detection means: Figure 3) detects the gaze direction (the user's viewpoint fixed on the viewing image, hereinafter referred to as the first estimated gaze point position) from the position of the corneal reflection image in the eyeball image formed on the ocular imaging element 17.
[0032] The light splitter 15 reflects the light that has passed through the eyepiece 12 and forms an image on the ocular imaging element 17 via the light receiving lens 16, and also transmits the light from the viewfinder 10 so that the user can see the visual image displayed on the viewfinder 10.
[0033] The photographic lens 1A includes an aperture 111, an aperture drive device 112, a lens drive motor 113, a lens drive member 114 including a drive gear and the like, a photocoupler 115, a pulse plate 116, a mount contact 117, and a focus adjustment circuit 118.
[0034] A photocoupler 115 detects the rotation of a pulse plate 116 that is linked to the lens driving member 114 and transmits the detection result to a focus adjustment circuit 118 .
[0035] The focus adjustment circuit 118 drives the lens drive motor 113 by a predetermined amount based on information from the photocoupler 115 and information on the lens drive amount from the camera body 1B, and moves the photographic lens 1A to the in-focus position.
[0036] The mount contact 117 is an interface between the camera housing 1B and the photographic lens 1A and has a known configuration. Signals are transmitted between the camera housing 1B and the photographic lens 1A via the mount contact 117. The CPU 3 of the camera housing 1B acquires type information and optical information of the photographic lens 1A to determine the focusable range of the photographic lens 1A attached to the camera housing 1B.
[0037] Fig. 3 is a block diagram showing the electrical configuration built into the imaging device 1. In Fig. 3, the same components as those in Figs. 1 and 2 are assigned the same numbers.
[0038] The camera housing 1B includes a line-of-sight detection circuit 201, a photometry circuit 202, an autofocus detection circuit 203, a signal input circuit 204, a viewfinder drive circuit 11, a light source drive circuit 205, a line-of-sight detection reliability determination circuit 31, and a communication circuit 32, each of which is connected to the CPU 3. The photographic lens 1A also includes a focus adjustment circuit 118 and an aperture control circuit 206 included in the aperture drive device 112 (FIG. 1), each of which transmits signals to the CPU 3 of the camera housing 1B via mount contacts 117.
[0039] The gaze detection circuit 201 A / D converts the eye image data formed and output on the eye imaging element 17 and transmits this eye image data to the CPU 3. The CPU 3 extracts each feature point of the eye image required for gaze detection from the eye image data in accordance with a predetermined algorithm described later, and further calculates the gaze position of the user estimated from the position of each extracted feature point (first estimated gaze point position).
[0040] Based on the signal obtained from the image sensor 2, which also functions as a photometric sensor, the photometry circuit 202 amplifies the luminance signal output corresponding to the brightness of the field, then performs logarithmic compression and A / D conversion, and sends it to the CPU 3 as field luminance information.
[0041] The autofocus detection circuit 203 A / D converts signal voltages from multiple pixels included in the image sensor 2 that are used for phase difference detection, and sends the converted signal to the CPU 3. The CPU 3 calculates the distance to the subject corresponding to each focus detection point from the signal voltages from the multiple pixels. This is a well-known technique known as image plane phase difference AF. In this embodiment, there are 180 focus detection points on the image plane of the viewfinder 10, as shown in the viewfinder field of view image (visual image) in Figure 4.
[0042] The signal input circuit 204 is connected to switches SW1 and SW2 (not shown). Switch SW1 is turned on by the first stroke of the release button 5 (FIG. 2(a)) to start photometry, distance measurement, line of sight detection, and other operations of the imaging device 1. Switch SW2 is turned on by the second stroke of the release button 5 to start the release operation. Signals from switches SW1 and SW2 are input to the signal input circuit 204 and sent to the CPU 3.
[0043] The gaze detection reliability determination circuit 31 (reliability determination means) determines the reliability of the first estimated gaze point position calculated by the CPU 3. This determination is performed based on the differences between two sets of eye image data: eye image data acquired during calibration (described later) and eye image data acquired during photography. Specifically, the differences here are differences in pupil diameter, the number of corneal reflections, and the amount of external light detected from each of the two sets of eye image data. More specifically, the gaze detection method described later in FIGS. 5 to 7 calculates the pupil edge. For example, if the number of extracted pupil edges is equal to or greater than a threshold, the reliability is determined to be high; otherwise, the reliability is determined to be low. This is because the pupil 141 (FIG. 5) of the user's eyeball 14 is estimated by connecting the pupil edges, and the more pupil edges that can be extracted, the higher the estimation accuracy. Alternatively, the reliability may be determined based on the degree to which the pupil 141 calculated by connecting the pupil edges is distorted relative to a circle. As another method, the reliability may be determined to be high near the index that the user gazes at during calibration, which will be described later, and low the reliability as the user gazes away from the index. When the gaze position information of the user calculated by the gaze detection circuit 201 is transmitted to the CPU 3, the gaze detection reliability determination circuit 31 transmits the reliability of the gaze position information to the CPU 3.
[0044] Under the control of the CPU 3, the communication circuit 32 communicates with a PC (not shown) on a server via a network (not shown) such as a LAN or the Internet.
[0045] The above-mentioned operation members 41 to 43 are configured to transmit their operation signals to the CPU 3, and in response to these signals, movement control by manual operation of the first estimated point of gaze position, which will be described later, is performed.
[0046] FIG. 4 is a diagram showing the field of view within the finder, showing the state in which the finder 10 is in operation (the state in which the visual image is displayed).
[0047] As shown in FIG. 4, the field of view within the finder includes a field mask 300, a focus detection area 400, 180 distance measurement point indices 4001 to 4180, and the like.
[0048] Each of the ranging point indices 4001 to 4180 is superimposed on the through image (live view image) displayed on the finder 10 so as to be displayed at a position corresponding to one of the multiple focus detection points on the imaging surface of the finder 10. Furthermore, of the ranging point indices 4001 to 4180, the indice that coincides with position A, which is the current first estimated point of interest position, is highlighted on the finder 10.
[0049] Next, a method for detecting a line of sight using the imaging device 1 will be described with reference to FIGS.
[0050] FIG. 5 is a diagram for explaining the principle of the line-of-sight detection method, and is a schematic diagram of an optical system for performing line-of-sight detection.
[0051] 5, light sources 13a and 13b are light sources such as light-emitting diodes that emit infrared light that is insensitive to the user, and each light source is arranged approximately symmetrically with respect to the optical axis of light-receiving lens 16 to illuminate user's eyeball 14. A portion of the illumination light emitted from light sources 13a and 13b and reflected by eyeball 14 is collected by light-receiving lens 16 onto ocular imaging element 17.
[0052] Figure 6(a) is a schematic diagram of an eye image captured by the eye imaging element 17 (eye image projected onto the eye imaging element 17), and Figure 6(b) is a diagram showing the output intensity of the photoelectric element array in the eye imaging element 17.
[0053] 7 is a flowchart of the gaze detection process, which is executed by the CPU 3 reading out a program stored in a ROM (not shown in FIG. 3).
[0054] 7, when the gaze detection process starts, in step S701, CPU 3 causes light sources 13a and 13b to emit infrared light toward user's eyeball 14. An image of the user's eye illuminated by the infrared light is formed on eye image sensor 17 through light receiving lens 16 and is photoelectrically converted by eye image sensor 17. As a result, a processable electrical signal of the eye image (eye image data) is obtained.
[0055] In step S702, the CPU 3 acquires from the eye imaging element 17 the eye image data obtained from the eye imaging element 17 as described above.
[0056] In step S703, the CPU 3 detects the coordinates corresponding to the corneal reflection images Pd and Pe of the light sources 13a and 13b and the pupil center c from the eye image data obtained in step S702.
[0057] Infrared light emitted from light sources 13a and 13b illuminates cornea 142 of user's eyeball 14. At this time, corneal reflection images Pd and Pe formed by part of the infrared light reflected from the surface of cornea 142 are condensed by light receiving lens 16 and formed on ocular imaging element 17 as corneal reflection images Pd' and Pe'. Similarly, light beams from edges a and b of pupil 141 are also formed on ocular imaging element 17 as pupil edge images a' and b'.
[0058] FIG. 6(b) shows luminance information (luminance distribution) of region α in the eye image of FIG. 6(a). In FIG. 6(b), the horizontal direction of the eye image is the X axis and the vertical direction is the Y axis, and the luminance distribution in the X axis direction is shown. In this embodiment, the X axis (horizontal direction) coordinates of the corneal reflection images Pd' and Pe' are designated Xd and Xe, and the X axis coordinates of the pupil edge images a' and b' are designated Xa and Xb. As shown in FIG. 6(b), an extremely high level of luminance is obtained at the coordinates Xd and Xe of the corneal reflection images Pd' and Pe'. In the range greater than the coordinate Xa and smaller than the coordinate Xb, which corresponds to the region of the pupil 141 (the region of the pupil image 141' obtained by focusing the light beam from the pupil 141 on the ocular imaging element 17), an extremely low level of luminance is obtained except for the coordinates Xd and Xe. In contrast, a luminance intermediate between the two types of luminance described above is obtained in the region of iris 143 outside pupil 141 (the region of iris image 143' outside pupil image 141' obtained by focusing the light beam from iris 143). Specifically, a luminance intermediate between the two types of luminance described above is obtained in a region where the X coordinate (coordinate in the X-axis direction) is smaller than coordinate Xa and a region where the X coordinate is larger than coordinate Xb.
[0059] From the luminance distribution shown in FIG. 6(b), the X-coordinates Xd and Xe of the corneal reflection images Pd' and Pe' and the X-coordinates Xa and Xb of the pupil edge images a' and b' can be obtained. Specifically, the coordinates of the corneal reflection images Pd' and Pe' can be obtained as the coordinates of extremely high luminance, and the coordinates of the pupil edge images a' and b' can be obtained as the coordinates of extremely low luminance. Furthermore, when the rotation angle θx of the optical axis of the eyeball 14 relative to the optical axis of the light receiving lens 16 is small, the coordinate Xc of the pupil center image c' (center of the pupil image 141') obtained when the light beam from the pupil center c is focused on the ocular imaging element 17 can be expressed as Xc ≒ (Xa + Xb) / 2. In other words, the X-coordinate Xc of the pupil center image c' can be calculated from the X-coordinates Xa and Xb of the pupil edge images a' and b'. In this way, the X coordinates of the corneal reflection images Pd' and Pe' and the X coordinate of the pupil center image c' can be estimated.
[0060] 7, in step S704, CPU 3 calculates the imaging magnification β of the eyeball image. The imaging magnification β is determined by the position of eyeball 14 relative to light receiving lens 16, and can be calculated as a function of the distance (Xd-Xe) between corneal reflection images Pd' and Pe'.
[0061] In step S705, the CPU 3 calculates the distance between the optical axis of the eyeball 14 and the optical axis of the light receiving lens 16. The rotation angle is calculated. The X coordinate of the midpoint between the corneal reflection image Pd and the corneal reflection image Pe and the X coordinate of the center of curvature O of the cornea 142 are almost the same. Therefore, if the standard distance from the center of curvature O of the cornea 142 to the center c of the pupil 141 is Oc, the rotation angle θ of the eyeball 14 in the ZX plane (plane perpendicular to the Y axis) is X can be calculated using the following formula 1. The rotation angle θy of the eyeball 14 in the ZY plane (plane perpendicular to the X axis) can also be calculated in the same way as the rotation angle θx.
[0062] β×Oc×SINθ X ≒{(Xd + Xe) / 2}-Xc (Equation 1) In step S706, the CPU 3 acquires correction coefficients (coefficient m and line-of-sight correction coefficients Ax, Bx, Ay, By) from the memory unit 4. The coefficient m is a constant determined by the configuration of the finder optical system (light receiving lens 16, etc.) of the imaging device 1, and is a conversion coefficient that converts the rotation angles θx, θy into coordinates corresponding to the pupil center c in the visual image, and is determined in advance and stored in the memory unit 4. The line-of-sight correction coefficients Ax, Bx, Ay, By are parameters that correct individual differences in the eyeball, and are acquired by performing a calibration operation (to be described later), and are stored in the memory unit 4 before this process starts.
[0063] In step S707, the CPU 3 instructs the line-of-sight detection circuit 201 to calculate the position of the user's gaze fixed on the visual image displayed on the finder 10 (first estimated gaze position). Specifically, the line-of-sight detection circuit 201 calculates the first estimated gaze position using the rotation angles θx, θy of the eyeball 14 calculated in step S705 and the correction coefficient data acquired in step S706. If the coordinates (Hx, Hy) of the first estimated gaze position are coordinates corresponding to the pupil center c, the coordinates (Hx, Hy) of the first estimated gaze position can be calculated by the following equations 2 and 3.
[0064] Hx=m×(Ax×θx+Bx) (Formula 2) Hy=m×(Ay×θy+By) (Formula 3) In step S708, the CPU 3 stores the coordinates (Hx, Hy) of the first estimated gaze point position calculated in step S706 in the memory unit 4, and then ends this process.
[0065] As described above, in the gaze detection process of this embodiment, the first estimated gaze point position was calculated using the rotation angles θx, θy of the eyeball 14 and the correction coefficients (coefficient m and gaze correction coefficients Ax, Bx, Ay, By) previously obtained by the calibration work described below.
[0066] However, due to factors such as individual differences in the shape of human eyeballs, it may not be possible to estimate the first estimated gaze point position with high accuracy. Specifically, unless the values of the gaze correction coefficients Ax, Ay, Bx, and By are adjusted to values suitable for the user, a discrepancy will occur between position B where the user actually gazes and position C, which is the first estimated gaze point position calculated in step S707, as shown in FIG. 4(b). In FIG. 4(b), the user is gazing at a person at position B, but the imaging device 1 erroneously estimates that the user is gazing at the background at position C, which is the first estimated gaze point position, resulting in a state in which appropriate focus detection and adjustment cannot be performed.
[0067] Therefore, the CPU 3 (calibration means) performs a calibration operation before the imaging device 1 performs imaging (focus detection), obtains line-of-sight correction coefficients Ax, Ay, Bx, and By suitable for the user, and stores them in the memory unit 4.
[0068] Conventionally, calibration work has been performed by highlighting multiple indices D1 to D5 at different positions as shown in Figure 4(c) in a visual confirmation image before imaging and having the user look at the indices. A publicly known technique involves performing a gaze detection process when the user gazes at each target, and calculating gaze correction coefficients Ax, Ay, Bx, and By suitable for the user from the calculated coordinates of multiple first estimated gaze point positions and the coordinates of each indices. Note that as long as the position where the user should look is suggested, it is not necessary to display the indices; the position may be highlighted by changing the brightness or color.
[0069] However, as mentioned above, the accuracy of gaze detection can decrease depending on the differences in conditions between when shooting and when calibration is performed, such as when external light enters the camera or when the distance between the user's eyes looking through the viewfinder 10 differs between when shooting and when calibration is performed.
[0070] FIG. 8 shows examples of eye image data during calibration and during photography, which are determined to have low reliability by the gaze detection reliability determination circuit 31. FIG. 8(a) is an example of eye image data acquired from the gaze detection circuit 201 during calibration, and FIG. 8(b) is an example of eye image data acquired from the gaze detection circuit 201 during photography. Here, FIG. 8(b) shows a case where the position of the user's eyes is farther from the eyepiece window 6 during photography than during calibration, resulting in a smaller eye size detected from the eye image data. For example, eye image data such as that shown in FIG. 8 may be obtained when calibration is performed with the imaging device 1 held horizontally in the optical axis direction, and then an image of a flower blooming on the ground is captured with the imaging device 1 pointed downward in the optical axis direction.
[0071] In such a case, the reliability of the first estimated gaze point position output from the gaze detection reliability determination circuit 31 becomes low, so in this embodiment, the CPU 3 estimates the second estimated gaze point position using an inference device that uses a neural network, more specifically, a CNN. Here, CNN is an abbreviation for Convolutional Neural Network, which is often used particularly for image recognition.
[0072] In this embodiment, the CPU 3 performs calculations in an inference unit using CNN. The basic configuration of CNN will be described with reference to FIGS.
[0073] FIG. 13 shows the basic configuration of a CNN that estimates the second estimated gaze point position from the eye image data output from the gaze detection circuit 201 to the CPU 3.
[0074] The processing flow is such that the left end is the input and processing proceeds to the right. CNN is structured hierarchically, with a set of two layers called the feature detection layer (S layer) and the feature integration layer (C layer).
[0075] In CNN, the next feature is first detected in layer S based on the features detected in the previous layer. The features detected in layer S are then integrated in layer C and sent to the next layer as the detection result for that layer.
[0076] The S layer consists of feature detection cell planes, each of which detects a different feature. The C layer consists of feature integration cell planes, which pool the detection results from the previous feature detection cell planes. Hereinafter, unless there is a need to distinguish between them, the feature detection cell planes and feature integration cell planes will be collectively referred to as feature planes. In this embodiment, the output layer, which is the final layer, does not use the C layer and is composed only of the S layer.
[0077] The details of the feature detection process on the feature detection cell plane and the feature integration process on the feature integration cell plane will be explained using Figure 14. The feature detection cell plane is composed of multiple feature detection neurons, which are connected to the C layer of the previous layer in a predetermined structure. The feature integration cell plane is composed of multiple feature integration neurons, which are connected to the S layer of the same layer in a predetermined structure. In the Mth cell plane of the S layer of the Lth layer shown in Figure 14, the output value of the feature detection neuron at position (ξ,ζ) is expressed as y M LS (ξ,ζ), in the Mth cell plane of the C layer of the Lth layer, the output value of the feature integration neuron at position (ξ,ζ) is y M LC (ξ,ζ). At this time, the coupling coefficient of each neuron is expressed as w M LS (n,u,v), w M LC Assuming that the output values are (u, v), each output value can be expressed as in the following equations 4 and 5.
[0078]
number
[0079]
number
[0080] In Equation 4, f is an activation function, and can be any sigmoid function such as a logistic function or a hyperbolic tangent function. M LS (ξ,ζ) is the internal state of the feature detection neuron at position (ξ,ζ) on the Mth cell surface of the S layer of the Lth hierarchy. On the other hand, since Equation 5 uses a simple linear sum without using an activation function, the internal state of the feature integration neuron at position (ξ,ζ) on the Mth cell surface of the C layer of the Lth hierarchy, u M LC (ξ,ζ) is the output value y calculated by Equation 5 M LC (ξ,ζ) are equal. Also, y in Eq. n L-1C(ξ+u,ζ+v), y in Eq. M LS (ξ+u, ζ+v) are called the output values of the feature detection neuron and the feature integration neuron, respectively.
[0081] The following explains ξ, ζ, u, v, and n in equations 4 and 5.
[0082] The position (ξ,ζ) corresponds to the position coordinates in the input image. For example, the output value y calculated by Equation 4 M LS If (ξ, ζ) has a high output value, it means that there is a high possibility that the feature to be detected in the Mth cell plane of the Sth layer of the Lth hierarchy exists at the pixel position (ξ, ζ) of the input image.
[0083] In addition, n in Equation 4 means the n-th cell surface in the C layer of the L-1th layer, and is called the integration target feature number. Basically, a product-sum operation is performed on all cell surfaces that exist in the C layer of the L-1th layer.
[0084] (u,v) are the relative position coordinates of the coupling coefficient, and the product-sum operation is performed within a finite range (u,v) depending on the size of the feature to be detected. This finite range (u,v) is called the receptive field. The size of the receptive field is hereinafter referred to as the receptive field size, and is expressed as the number of horizontal pixels x the number of vertical pixels in the coupled range.
[0085] In addition, in Equation 4, L=1, that is, the first S layer, y n L-1C (ξ+u,ζ+v) is the input image y in_image (ξ+u,ζ+v). Incidentally, the distribution of neurons and pixels is discrete, and the connection feature numbers are also discrete, so ξ,ζ,u,v,n are not continuous variables but take discrete values. Here, ξ and ζ are non-negative integers, n is a natural number, and u and v are integers, all of which have a finite range.
[0086] w in Equation 4 M LS(n,u,v) is the distribution of coupling coefficients for detecting a specific feature, and by adjusting it to an appropriate value, it becomes possible to detect the specific feature. Adjusting this distribution of coupling coefficients is called learning. In building a CNN, various test patterns are presented to learn the output value y calculated by Equation 4. M LS The coupling coefficients are adjusted by repeatedly correcting them gradually so that (ξ,ζ) becomes an appropriate output value.
[0087] Next, w in Equation 5 M LC (u,v) uses a two-dimensional Gaussian function and can be expressed as the following equation 6.
[0088]
number
[0089] Here, (u, v) is also a finite range, so similar to the explanation of feature detection neurons, this finite range is called the receptive field, and the size of the range is called the receptive field size. Here, this receptive field size can be set to an appropriate value depending on the size of the Mth feature in the S layer of the Lth layer. In Equation 6, σ is a feature size factor, which can be set to an appropriate constant depending on the receptive field size. Specifically, it is best to set it so that the outermost value of the receptive field can be considered to be approximately 0. The CNN of this embodiment is configured to estimate the second estimated gaze point position in the S layer of the final layer by performing the above-mentioned calculations at each layer.
[0090] Here, when creating an inference device that uses image data of the user's eye looking through the viewfinder 10 as input data and outputs the second estimated gaze point position as an inference result, how to define correct data (correct position) becomes important. If the reliability of the first estimated gaze point position calculated by the gaze detection reliability determination circuit 31 is high, the correct position may be used as the first estimated gaze point position. However, if this is not the case and the first estimated gaze point position is used as the correct position, the accuracy rate through learning will not increase. Therefore, in this embodiment, if the reliability of the first estimated gaze point position is low, other information obtained from the imaging device 1 is collected as correct data.
[0091] 9 is a flowchart of the process of collecting correct answer data used for learning when creating an inference device. This process is executed by the CPU 3 reading out a program recorded in a ROM (not shown in FIG. 3).
[0092] In step S901, CPU 3 monitors whether the user is looking through viewfinder 10. This can be determined, for example, by checking whether the image data output from gaze detection circuit 201 is eye image data. Note that any method can be used to monitor whether the user is looking through viewfinder 10; for example, an optical sensor (not shown) provided around eyepiece 12 may be used to detect whether the eye is placed in eyepiece 12. If it is determined that the user is looking through viewfinder 10, the process proceeds to step S902.
[0093] In step S902, the CPU 3 calculates a first estimated gaze point position from the eye image data output from the gaze detection circuit 201, and acquires the reliability output from the gaze detection reliability determination circuit 31. Then, the process proceeds to step S903.
[0094] In step S903, if the reliability output from the gaze detection reliability determination circuit 31 is high, the CPU 3 proceeds to step S904, whereas if the reliability is low, the CPU 3 proceeds to step S905.
[0095] In step S904, the CPU 3 collects the first estimated gaze point position as the correct position, and then ends this process.
[0096] In step S905, CPU 3 determines whether a subject is present near the first estimated point of gaze position. In this determination, a person or an eye may be detected as the subject. If the result of this determination is that a subject is detected near the first estimated point of gaze position, the process proceeds to step S906; otherwise, the process proceeds to step S907.
[0097] In step S906, the CPU 3 (collection means) collects the coordinates of the subject near the first estimated gaze point position as the correct position, and then ends this process.
[0098] In step S907, the CPU 3 (display control means) highlights the first estimated point of gaze position (hereinafter, in this process, position C in FIG. 4B) in the viewfinder 10 so that it can be moved to another position (hereinafter, in this process, position B in FIG. 4B) by a first user operation. Here, the first user operation refers to a manual operation by the user using any of the operation members 41 to 43. After that, the CPU 3 determines whether or not an image has been captured with the imaging device 1 using position B as the focus position (whether the release button 5 has been pressed (second user operation)) after position C, which was highlighted in the viewfinder 10, has been moved to position B by the first user operation. Only if such an image has been captured, proceed to step S908. Note that the second user operation may be any user operation other than pressing the release button 5, as long as it determines another position on the viewfinder 10 selected by the user as the focus position.
[0099] In step S908, the CPU 3 (collection means) collects the coordinates of the focus position (other position) of the photographed image as the correct position, and then ends this process.
[0100] 9 (input data) and the correct positions (correct data) at that time, using the communication circuit 32 to a PC on a server (not shown) via a network such as a LAN or the Internet. The PC on the server performs CNN machine learning using this data, and transmits an "inferer" generated as a result of the learning to the imaging device 1. Note that the imaging device 1 may have a high-performance GPU, and the GPU (or cooperation with the CPU 3) may perform the CNN machine learning.
[0101] Next, we will explain how to use the generated inference machine generated by CNN machine learning performed on a PC on the server.
[0102] 10 is a flowchart of the focus process during photography. This process is executed by the CPU 3 reading out a program recorded in a ROM (not shown in FIG. 3).
[0103] 10, first, the CPU 3 performs the processes of steps S901 to S903. These processes have been described above in the description of FIG. 9, so a duplicated description will be omitted.
[0104] In step S903, if the reliability of the first estimated gaze point position is high, the CPU 3 proceeds to step S1004, and if the reliability is low, the CPU 3 proceeds to step S1005.
[0105] In step S1004, the CPU 3 (first focus means) determines that the first estimated point of interest position is the focus position desired by the user, and performs focusing based on the first estimated point of interest position. This process is realized by operating the autofocus detection circuit 203 and focus adjustment circuit 118 in response to instructions from the CPU 3. Specifically, the autofocus detection circuit 203 first calculates the distance to the subject corresponding to the focus detection point that coincides with the first estimated point of interest position. Then, the focus adjustment circuit 118 drives the lens drive motor 113 a predetermined amount based on this information, moving the photographic lens 1A to the in-focus position. Then, this process ends.
[0106] In step S1005, the CPU 3 inputs the eye image data output from the gaze detection circuit 201 in step S902 into the inference device transmitted from the PC on the server, and estimates the second estimated gaze point position. In this embodiment, the CPU 3 estimates the second estimated gaze point position, but this is not limiting. For example, the PC on the server may estimate the second estimated gaze point position. In this case, the CPU 3 transmits the eye image data output from the gaze detection circuit 201 in step S902 to the PC on the server, and the PC on the server uses the inference device to estimate the second estimated gaze point position and output the inference result to the CPU 3 of the imaging device 1. Then, the process proceeds to step S1006.
[0107] In step S1006, if the CPU 3 determines that the reliability of the second estimated point of gaze position estimated in step S1005 is high, the process proceeds to step S1007. If the CPU 3 determines that the reliability is low, the process proceeds to step S1008. In the inference device of this embodiment, the likelihood is calculated for each of the 180 focus detection points. Therefore, the focus detection point for which the highest likelihood is calculated is determined to be the second estimated point of gaze position. Furthermore, if the value of the highest likelihood is equal to or greater than a threshold, the reliability of the second estimated point of gaze position is determined to be high.
[0108] In step S1007, the CPU 3 (second focus means) determines that the second estimated point of interest position is the focus position desired by the user, and performs focusing based on the second estimated point of interest position. This process is realized by operating the auto focus detection circuit 203 and the focus adjustment circuit 118 in response to an instruction from the CPU 3. Then, this process ends.
[0109] In step S1008, CPU 3 determines which of the first and second estimated gaze point positions has a higher reliability. If it determines that the reliability of the first estimated gaze point position is higher, the process proceeds to step S1009, whereas if it determines that the reliability of the second estimated gaze point position is higher, the process proceeds to step S1010.
[0110] In step S1009, CPU 3 determines that the focus position desired by the user is near the first estimated point of gaze position, detects an object near the first estimated point of gaze position, and focuses on the detected object as the focus point, and then ends this process.
[0111] In step S1010, CPU 3 determines that the focus position desired by the user is near the second estimated point of gaze position, detects an object near the second estimated point of gaze position, and focuses on the detected object as the focus point, and then ends this process.
[0112] Furthermore, even if the reliability of the first estimated point of gaze position is high as a result of progress in the machine learning of the CNN on the PC on the server, if the reliability of the second estimated point of gaze position estimated by the inference device becomes equal to or higher than that of the first estimated point of gaze position, focus processing during shooting may always be performed using the inference device. Specifically, steps S902, S903, S1004, S1008, and S1009 in FIG. 10 are unnecessary, and if YES in step S901, the process proceeds directly to step S1005, and if NO in step S1006, the process proceeds directly to step S1010. As a result, the focus position is determined based only on the second estimated point of gaze position.
[0113] In this embodiment, the position of the user's eyes looking through the viewfinder 10 during photography is farther from the eyepiece window 6 than during calibration, but the present invention is not limited to this. In other words, this embodiment can be applied when focus processing is performed under various conditions that reduce the reliability output by the gaze detection reliability determination circuit 31.
[0114] As described above, in this embodiment, meaningful learning can be performed even when the accuracy of gaze detection is reduced by defining the correct position during learning using information from the image capture device 1. Then, by switching between the gaze detection circuit 201 and the inference unit depending on the reliability of the first and second estimated gaze point positions, the accuracy of gaze detection can be improved even in shooting conditions that differ from those at the time of calibration.
[0115] Example 2 Hereinafter, a method for collecting input data during optimal learning when the reliability of gaze detection is low under various conditions according to a second embodiment of the present invention will be described with reference to Figures 11 and 12. In this embodiment, the same components as those in the first embodiment are assigned the same numbers, and duplicated descriptions will be omitted.
[0116] As explained in the first embodiment, the decrease in reliability of the first estimated gaze point position becomes more pronounced when the eye information differs between calibration and shooting. However, this can be caused by multiple factors, such as differences in the distance between the eye of the user looking through the viewfinder 10 and the eyepiece 12, or differences in the amount of light entering the eye. Therefore, in this embodiment, the CPU 3 (means for determining reliability decrease) determines these multiple factors that cause a decrease in reliability and collects input data for neural network training that is optimal for each factor.
[0117] 11 is a flowchart of a process for determining a factor that reduces the reliability of the first estimated point of gaze position according to this embodiment. This process is executed by the CPU 3 reading out a program recorded in a ROM (not shown in FIG. 3).
[0118] 11, first, the CPU 3 performs the processes of steps S901 to S903 when an image is captured by the imaging device 1. These processes have been described above in the description of FIG. 9, so a duplicated description will be omitted.
[0119] In step S903, if the reliability of the first estimated gaze point position is low, the CPU 3 proceeds to step S1101, and if the reliability is high, the CPU 3 ends this process.
[0120] In step S1101, CPU 3 determines whether the eye size has changed since calibration, i.e., whether a change in the distance between the user's eye looking through viewfinder 10 and eyepiece 12 has occurred, causing a decrease in reliability. The distance between eyepiece 12 and the user's eye can be calculated based on the current eye size relative to the eye size during calibration. For example, when detecting the pupil using the gaze detection method shown in FIGS. 5 to 7, the calculated pupil diameter can be calculated as the eye size, and the distance between eyepiece 12 and the user's eye from calibration can be estimated. Note that any method can be used as long as it can determine whether the eye size has changed since calibration. For example, an optical sensor provided around eyepiece 12 can be used to calculate the distance between eyepiece 12 and the user's eye during calibration and when capturing an image. If it is determined that the eye size is the same as during calibration, the process proceeds to step S1103. If it is determined that the eye size is different from that during calibration, the process proceeds to step S1102.
[0121] In step S1102, the CPU 3 stores information indicating that the distance between the viewfinder 10 and the photographer (user's eyes) is different as a factor in reducing reliability in the memory unit 4. Thereafter, the process proceeds to step S1103.
[0122] In step S1103, CPU 3 determines whether the brightness is equal to or greater than a predetermined value. This can be determined by checking the brightness of the eye image data of the user looking through viewfinder 10, acquired by ocular imaging element 17. If the brightness is equal to or greater than the predetermined value, it is determined that external light has entered, and the process proceeds to step S1104. If the brightness is less than the predetermined value, it is determined that external light has not entered, and the process ends.
[0123] In step S1104, the CPU 3 stores information indicating that external light is entering as a factor in reducing reliability in the memory unit 4. The CPU 3 also generates differential eye image data shown in Fig. 12(c) (to be described later) as input data. After that, the process ends.
[0124] Generally, improving the accuracy of inference by an inference device requires training using a large amount of input data. Therefore, when the input data is small, a technique called padding is often used in CNNs, in which the amount of data is increased by applying transformations to the original input data. While various padding techniques exist, such as increasing noise, scaling the image, masking parts, and inverting the image, some padding techniques may reduce the accuracy of inference by the inference device. This is because the padded data used during training may be of poor quality, such as data that is not actually possible or data that the inference device cannot infer. In this embodiment, by storing the factors that caused the reliability to decrease according to the process of Figure 11, it is possible to perform appropriate padding according to the original input data obtained.
[0125] The CPU 3 transmits the eye image data and the correct positions (correct data) at that time collected in the process of Fig. 9, as well as the reliability reduction factors acquired in the process of Fig. 11, to a PC on a server via a network such as a LAN or the Internet using the communication circuit 32. The PC on the server performs CNN machine learning using these data, and transmits an "inferer" generated as a learning result to the imaging device 1.
[0126] When the server PC receives information indicating that the distance between the viewfinder 10 and the photographer is different as a factor in reducing reliability, it enlarges or reduces the eye image data (original input data) at the time of image capture when it is transmitted from the imaging device 1, and uses the enlarged or reduced data as padded data during learning. More specifically, the server PC also acquires eye image data at the time of calibration from the imaging device 1 and detects the pupil diameter from each of the eye image data at the time of calibration and image capture. The server PC creates padded data by enlarging or reducing the original input data so that the pupil diameter detected from the padded data falls within the range of these detected pupil diameters. Creating padded data in this manner enables optimal padding while suppressing the generation of poor-quality padded data. Furthermore, it is possible to prevent the generation of infeasible data, such as masking a portion of the user's pupil image, as padded data.
[0127] On the other hand, when reliability is reduced due to the intrusion of external light, the upper or lower part of the eye image captured by the ocular imaging device 17 may be blown out, depending on the external light conditions. The former occurs when sunlight enters during daytime photography, for example, while the latter occurs when sunlight is reflected off snow at a ski resort. In addition, sunlight may enter from the side of the eyepiece 12 depending on the shooting position of the imaging device 1. In other words, the parts of the eye image where blown out white occur vary widely depending on the external light conditions.
[0128] Even if eye image data with such whiteout defects is used as input data for learning, the accuracy of inference by the inference device will not improve easily if the external light conditions are different. Therefore, in this embodiment, in such cases, differential eye image data generated by the method shown in Figure 12 is collected as input data.
[0129] 12(a) is eye image data output from the eye imaging element 17, and external light 1200 enters above it. During this imaging, the light sources 13a and 13b are turned on, and corneal reflection images 1201a-c are formed on the user's cornea (eyeball).
[0130] The eye image data in Fig. 12(b) is the eye image data output by the eye imaging element 17, and like Fig. 12(a), external light 1200 enters the upper part of the image. However, at the time of capturing this image, the light sources 13a and 13b are turned off, and the corneal reflection images 1201a-c are not formed on the user's cornea (eyeball). The eye image shown in Fig. 12(b) is captured slightly darker overall than the eye image shown in Fig. 12(a) because the light sources 13a and 13b are turned off.
[0131] The differential eye image data of FIG. 12(c) is differential data obtained by subtracting FIG. 12(b) from FIG. 12(a), and external light 1200 that has entered the upper part of the eye image has been removed.
[0132] Thus, when the reliability reduction factor is information indicating the presence of external light, the server PC receives differential eye image data such as that shown in FIG. 12(c) as input data, rather than eye image data in which a corneal reflex image is formed as shown in FIG. 12(a). This allows learning to separate external light conditions to a certain extent. However, because the differential eye image data in FIG. 12(c) is an image obtained by subtracting the strong external light, the edges around the pupil are blurred, resulting in poor gaze detection accuracy. The server PC performs CNN machine learning using the differential eye image data and padded data created based on it, thereby creating an inference machine that is effective against edge blur around the pupil, enabling inference even when external light conditions change. In this case, data in which a portion of the pupil boundary of the differential eye image data transmitted from the imaging device 1 is masked or data in which noise around the pupil is increased is created as padded data for learning.
[0133] As described above, in this embodiment, when the reliability of the first estimated gaze point position is reduced during input data collection during learning, the CPU 3 transmits a flag indicating the cause of the reduced reliability together with the input data to the PC on the server. This allows the PC on the server to create appropriate padded data based on the input data, thereby improving the accuracy of the inference device with fewer learning iterations.
[0134] Although the preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention.
[0135] (Other embodiments) The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more functions. [Explanation of symbols]
[0136] 1. Imaging device 3 CPU 10 Finder 13a~13b Light source 17 Eye imaging device 31 Eye gaze detection reliability determination circuit 41-43 Operating member 118 Focus adjustment circuit 201 Gaze detection circuit 203 Auto focus detection circuit
Claims
1. An imaging device that displays a through image in an internal finder, a generating means for capturing an image of the eyeball of a user looking through the viewfinder and generating eye image data; a gaze detection means for acquiring the eye image data and detecting a gaze position of a user viewing the through image of the finder based on the acquired eye image data; a display control means for displaying the detected line of sight position on the viewfinder in a manner that allows the line of sight position to be moved to another position by a first user operation; a collection means for collecting the other position as a correct position when a second user operation is performed to determine the other position as a focus position; a reliability determination means for determining the reliability of the detected gaze position; a first focusing unit that focuses the imaging device using the detected gaze position as the focus position when the reliability is high; a second focusing means for focusing the imaging device using the gaze position estimated by the inferring device that estimates the gaze position as the focus position when the reliability is low, The imaging device is characterized in that the correct position is used for learning to create the inference device, using the eye image data acquired by the gaze detection means as input data.
2. 2. The imaging device according to claim 1, wherein, when a subject is present near the detected gaze position, the collecting means collects the position of the subject as the correct position.
3. a calibration unit that acquires the eye image data before detecting the gaze position by the gaze detection unit and corrects individual differences in the eyeballs based on the acquired eye image data; The imaging device according to claim 1, characterized in that the reliability determination means determines the reliability based on the difference between two sets of eye image data: the eye image data acquired by the calibration means and the eye image data acquired by the gaze detection means.
4. 4. The imaging device according to claim 3, wherein the difference is a difference in pupil diameter detected from each of the two sets of eye image data.
5. 5. The imaging device according to claim 4, wherein the difference is a difference in external light detected from each of the two sets of eye image data.
6. a light source for illuminating the user's eyeball; 6. The imaging device according to claim 5, wherein the difference is a difference in the number of corneal reflection images formed on the user's eyeball by turning on the light source, detected from each of the two eye image data.
7. The imaging device of claim 6, characterized in that when the reliability discrimination means determines that the reliability has decreased due to the difference in external light, the learning is performed using differential data between the eye image data of the user's eyeball in which the corneal reflection image is formed by turning on the light source and the eye image data of the user's eyeball in which the corneal reflection image is not formed by turning off the light source as the input data instead of the eye image data acquired by the gaze detection means.
8. The method further includes a reliability reduction factor determining means for determining a factor causing a reduction in reliability when the reliability determining means determines that the reliability is low, 8. The imaging device according to claim 7, wherein padded data corresponding to the reliability reduction factors is created during the learning process.
9. An imaging device as described in claim 8, characterized in that if the factor of the decrease in reliability is a difference in distance between the viewfinder and the user, the size of the pupil diameter is detected from each of the two eye image data, and the input data is enlarged or reduced so that the size of the pupil diameter detected from the padded data is within the range of the detected pupil diameter size, thereby creating the padded data.
10. The imaging device described in claim 8, characterized in that when the factor that reduces reliability is the intrusion of external light, the padding is performed by creating data in which a portion of the pupil boundary of the difference data as the input data is masked or by creating data in which noise around the pupil is increased.
11. A control method for an imaging device that displays a through image in an internal viewfinder, comprising: a generating step of capturing an image of the user's eyeball looking through the viewfinder to generate eye image data; a gaze detection step of acquiring the eye image data and detecting a gaze position of a user viewing the through image of the finder based on the acquired eye image data; a display control step of displaying the detected gaze position on the viewfinder in a manner that allows the gaze position to be moved to another position by a first user operation; a collection step of collecting the other position as a correct position when a second user operation to determine the other position as a focus position is performed; a reliability determination step of determining the reliability of the detected gaze position; a first focusing step of focusing the imaging device using the detected gaze position as the focus position when the reliability is high; a second focusing step of focusing the imaging device using the gaze position estimated by the inferrer that estimates the gaze position as the focus position when the reliability is low, A control method characterized in that the correct position is used for learning to create the inference device, using the eye image data acquired in the gaze detection step as input data.
12. A computer-executable program that causes a computer to function as each of the means of the imaging device according to any one of claims 1 to 10.
Citation Information
Patent Citations
Optical device and camera
JP2002287011A
Optical device with visual axis function
JP2004008323A
Glance detector
JP2004129927A
Apparatus, method and program for measuring visual axis
JP2007136000A
Information processing apparatus, information processing method, and program
JP2015152938A